docs(jarvis): create iteration stories 1 (ready) + 2/3 (refine)

- jarvis 1 (chat loop) brainstormed to ready with forks AUTO-APPROVED for
  autonomous execution and flagged in `review_pending` frontmatter for the
  developer's second review: Anthropic Messages API backend, env-var API key,
  actor-per-conversation SSE relay, durable @table history, session-gated routes
- jarvis 2 (tool use) + 3 (retrieval/RAG) created at refine with forks named
- 00-story iterations table linked to the new files
- blocked until rv2 9 TLS reaches phase F; pure .wo on porch 2/3/6/7 + the seam

(cherry picked from commit 8e160c3fcf9c50a05af073c036c693b192c9e595)
This commit is contained in:
shoney.arickathil 2026-09-08 17:36:38 +02:00
parent ca572086ee
commit 9d1689be55
4 changed files with 200 additions and 3 deletions

View file

@ -60,9 +60,9 @@ past it is worth building until that seam is proven.
| # | Iteration | Delivers | Needs |
| --- | --- | --- | --- |
| 1 | the chat loop | a prompt sent to one LLM, tokens streamed back to the browser, the conversation persisted durably | the outbound seam (language 38 + TLS); porch 2/3/6/7; wo-html |
| 2 | tool use / the agent loop | function-calling and multi-step orchestration through actors — where "assistant" becomes "agent" | 1 |
| 3 | retrieval (RAG) | embeddings + vector search over a document set; carries its own sub-gap — an embeddings call over the same outbound path, plus a vector store (pure-`.wo` or a new primitive, decided in that story) | 1, and the embeddings/vector decision |
| 1 | [the chat loop](01-chat-loop.md) — ✅ `ready` (forks auto-approved, `review_pending`) | a prompt sent to one LLM, tokens streamed back to the browser, the conversation persisted durably | the outbound seam (TLS phase F); porch 2/3/6/7; wo-html |
| 2 | [tool use / the agent loop](02-tool-use.md) — `refine` | function-calling and multi-step orchestration through actors — where "assistant" becomes "agent" | 1 |
| 3 | [retrieval (RAG)](03-retrieval.md) — `refine` | embeddings + vector search over a document set; carries its own sub-gap — an embeddings call over the same outbound path, plus a vector store (pure-`.wo` or a new primitive, decided in that story) | 1, and the embeddings/vector decision |
Sketched, not committed — named so the shape is visible, not to schedule them:
**model routing / multi-model** (choose a backend per request) and an **MCP

View file

@ -0,0 +1,96 @@
---
track: jarvis
iteration: "1"
status: pending
readiness: ready
review_pending: "forks auto-approved 2026-09-08 for autonomous execution — developer second review before code lands"
---
# jarvis 1 — the chat loop: a prompt in, streamed tokens out, durable history
> Part of [Story — jarvis, the writeonce AI assistant](00-story.md).
> The whole end-to-end seam: a browser sends a prompt, jarvis relays it to an
> LLM over HTTPS, streams the reply back token by token, and persists the
> conversation durably. Nothing past this rung is worth building until it runs.
>
> **Forks auto-approved (2026-09-08) for autonomous execution and flagged in
> `review_pending` for the developer's second review** — the defaults below are
> reasonable but were not individually confirmed.
## Blocked until the outbound seam lands
This iteration cannot run until `net.connect` (✅ landed) and runtime-v2
[9](../runtime-v2/09-in-process-tls.md) TLS reach **phase F** (`net.connect_tls`
+ the handshake). Crypto A–D are landed; E (X.509) and F (handshake) remain.
Everything below is buildable `.wo` on top of that seam plus porch 2/3/6/7.
## Decisions locked (auto-approved, review pending)
1. **Backend: the Anthropic Messages API over HTTPS**, streaming (SSE). One
backend for v1; a thin adapter boundary so an OpenAI-compatible backend can
slot later without touching the chat loop. Endpoint, model id and version
header come from config.
2. **Auth: an API key from the environment/config**, sent as the `x-api-key`
header — never in a URL, never logged. Missing key is a startup refusal, not
a runtime surprise.
3. **The relay is an actor per conversation.** A conversation actor owns the
upstream `net.connect_tls` connection, parses the LLM's SSE
(`content_block_delta` text deltas), and forwards each delta to the browser
through porch [7](../porch/07-sse-and-compression.md)'s SSE. Backpressure and
disconnect cleanup follow porch 7's contract (a write to a gone client frees
the fiber, the upstream fd and the actor).
4. **Durable history in two `@table` classes**: `Conversation { id @unique,
principal, created_at }` and `Message { conv_id (indexed), seq, role,
content, created_at }`. Keyed to the session principal (porch
[3](../porch/03-sessions.md)); history replays after a restart with no
external store — the writeonce differentiator.
5. **Route surface, session-gated**: `GET /` (chat UI via wo-html/writeonce-view),
`POST /message` (accept a prompt, append it, start the stream), `GET /stream`
(SSE of the reply). The POST is CSRF-protected once porch
[4](../porch/04-csrf.md) is built; until then it is bearer/session-gated and
the gap is stated.
6. **One turn at a time, no tools.** No function-calling, no retrieval — those
are iterations 2 and 3. The reply is a single streamed assistant message.
## Phases
- **A — the backend client.** Over `net.connect_tls`, issue the Messages
request and parse the streamed SSE deltas into text. Adapter boundary isolated
so the wire format is one file.
- **B — the conversation store.** The two `@table` classes, append-message,
load-history, list-conversations; durable-replay proven.
- **C — the relay + web surface.** The conversation actor, the porch routes, the
wo-html chat page, browser SSE of the reply; session-gated.
- **D — the gate and the ledger.** `just` a jarvis sample end to end against a
**local stub** LLM server (no network in the gate): prompt → streamed reply →
durable history → restart-replay. Both the happy path and a mid-stream
disconnect. Record what landed.
## Acceptance Criteria
- **Given** a signed-in user and a prompt, **when** it is sent, **then** the
reply streams back token by token and both prompt and reply are persisted.
- **Given** a restart, **when** the user returns, **then** their conversation
history replays from the `@table`, no external store.
- **Given** the client disconnecting mid-stream, **when** the next delta
arrives, **then** the upstream fd, the fiber and the actor are freed.
- **Given** a missing API key, **when** the app starts, **then** it refuses
loudly rather than failing at first request.
- **Given** the gate, **when** it runs, **then** it exercises the whole path
against a local stub with no live network call.
## Out Of Scope
- **Tool use / function calling** — iteration [2](00-story.md).
- **Retrieval / embeddings** — iteration [3](00-story.md).
- **Multiple backends at once / model routing** — sketched in the overview,
later.
- **The live network in the gate** — the gate uses a local stub server; a real
API smoke test is a manual, keyed, out-of-gate step.
## Info
Pure `.wo` on porch 2/3/6/7 + the TLS seam; no new runtime work of its own. The
one hard dependency is runtime-v2 9 reaching phase F. The auto-approved forks
(backend choice, auth transport, store schema, route surface) are the developer
second-review items flagged in `review_pending`.

View file

@ -0,0 +1,50 @@
---
track: jarvis
iteration: "2"
status: pending
readiness: refine
---
# jarvis 2 — tool use / the agent loop
> Part of [Story — jarvis, the writeonce AI assistant](00-story.md).
> Where the assistant becomes an agent: the model may call tools, and jarvis
> runs them and feeds results back until the model answers. Needs jarvis
> [1](01-chat-loop.md).
## Why this exists
A chat loop answers from the model alone. An agent can act — read a file, query
a `@table`, call a `.wo` function — by the model emitting a tool call, jarvis
executing it, and returning the result for another turn. The actor model fits
the multi-step orchestration naturally.
## What it should deliver (to be refined)
- A **tool registry**: named tools, each a `.wo` class with a typed input and a
handler, satisfying a `Tool` interface structurally (the porch middleware
precedent).
- The **agent loop**: relay the tool-use request, dispatch to the tool actor,
append the tool result, continue the turn — bounded to a max step count.
- **Tool-call framing** in the backend adapter (the Messages API `tool_use` /
`tool_result` blocks), isolated in the same adapter file jarvis 1 established.
## Forks the brainstorm must settle
1. **Tool declaration shape** — a `.wo` class + interface (no reflection, per
principle 13) versus a declarative schema; and how the input JSON maps to a
typed value without `@derive` (language 29).
2. **Sandboxing / permissions** — which tools are allowed, and whether a tool
that touches the filesystem or spawns a process (runtime-v2) needs an
explicit grant.
3. **Step bound + loop safety** — the max tool-call depth and how a runaway loop
is stopped.
## Out of scope
- Retrieval (iteration [3](00-story.md)); multi-model routing (later).
## Info
Pure `.wo` on iteration 1. No new runtime primitive expected unless a tool needs
one (each such tool names its own dependency).

View file

@ -0,0 +1,51 @@
---
track: jarvis
iteration: "3"
status: pending
readiness: refine
---
# jarvis 3 — retrieval (RAG): grounded answers over a document set
> Part of [Story — jarvis, the writeonce AI assistant](00-story.md).
> Answers grounded in a corpus: embed documents, embed the query, retrieve the
> nearest chunks, and put them in the prompt. Needs jarvis [1](01-chat-loop.md)
> and the embeddings/vector decision below.
## Why this exists
A chat/agent loop knows only the model's training and the tools it can call.
Retrieval lets jarvis answer over the developer's own documents — the on-brand
"ask my codebase / my notes" use case — without fine-tuning.
## What it should deliver (to be refined)
- **Embeddings** for documents and queries — an embeddings API call over the
same outbound TLS path jarvis 1 uses.
- **A vector store**: chunk text, store `{chunk, vector}` durably, and a
similarity search (cosine / dot-product) over it.
- **Retrieval into the prompt**: top-k chunks prepended as context in the chat
loop.
## Forks the brainstorm must settle
1. **The vector store: pure `.wo` or a new runtime primitive.** A brute-force
cosine scan over a `@table` of `Bytes` vectors is pure `.wo` and fine for a
modest corpus; an approximate-nearest-neighbour index or a SIMD dot-product
builtin is a runtime iteration if scale demands it. Decide on a measured
need, not up front.
2. **Chunking strategy** — fixed-size vs semantic; overlap; where metadata
(source, offset) lives.
3. **Embedding storage** — vectors as `Bytes` (packed floats) in a `@table`, and
whether Float arrays need a better carrier than iteration 19's `Bytes`.
## Out of scope
- Re-ranking models, hybrid keyword+vector search, and multi-corpus tenancy —
each its own later slice.
## Info
Depends on jarvis 1's outbound path (for embeddings) and the vector-store fork.
The only possible new runtime work is fork 1's ANN/SIMD option, deferred until a
corpus size measures the need.