From 9d1689be5584fd03f7cc82f9414a98a39f5e203b Mon Sep 17 00:00:00 2001 From: "shoney.arickathil" Date: Tue, 8 Sep 2026 17:36:38 +0200 Subject: [PATCH] docs(jarvis): create iteration stories 1 (ready) + 2/3 (refine) - jarvis 1 (chat loop) brainstormed to ready with forks AUTO-APPROVED for autonomous execution and flagged in `review_pending` frontmatter for the developer's second review: Anthropic Messages API backend, env-var API key, actor-per-conversation SSE relay, durable @table history, session-gated routes - jarvis 2 (tool use) + 3 (retrieval/RAG) created at refine with forks named - 00-story iterations table linked to the new files - blocked until rv2 9 TLS reaches phase F; pure .wo on porch 2/3/6/7 + the seam (cherry picked from commit 8e160c3fcf9c50a05af073c036c693b192c9e595) --- docs/stories/jarvis/00-story.md | 6 +- docs/stories/jarvis/01-chat-loop.md | 96 +++++++++++++++++++++++++++++ docs/stories/jarvis/02-tool-use.md | 50 +++++++++++++++ docs/stories/jarvis/03-retrieval.md | 51 +++++++++++++++ 4 files changed, 200 insertions(+), 3 deletions(-) create mode 100644 docs/stories/jarvis/01-chat-loop.md create mode 100644 docs/stories/jarvis/02-tool-use.md create mode 100644 docs/stories/jarvis/03-retrieval.md diff --git a/docs/stories/jarvis/00-story.md b/docs/stories/jarvis/00-story.md index c0b8225..e6f7ec9 100644 --- a/docs/stories/jarvis/00-story.md +++ b/docs/stories/jarvis/00-story.md @@ -60,9 +60,9 @@ past it is worth building until that seam is proven. | # | Iteration | Delivers | Needs | | --- | --- | --- | --- | -| 1 | the chat loop | a prompt sent to one LLM, tokens streamed back to the browser, the conversation persisted durably | the outbound seam (language 38 + TLS); porch 2/3/6/7; wo-html | -| 2 | tool use / the agent loop | function-calling and multi-step orchestration through actors — where "assistant" becomes "agent" | 1 | -| 3 | retrieval (RAG) | embeddings + vector search over a document set; carries its own sub-gap — an embeddings call over the same outbound path, plus a vector store (pure-`.wo` or a new primitive, decided in that story) | 1, and the embeddings/vector decision | +| 1 | [the chat loop](01-chat-loop.md) — ✅ `ready` (forks auto-approved, `review_pending`) | a prompt sent to one LLM, tokens streamed back to the browser, the conversation persisted durably | the outbound seam (TLS phase F); porch 2/3/6/7; wo-html | +| 2 | [tool use / the agent loop](02-tool-use.md) — `refine` | function-calling and multi-step orchestration through actors — where "assistant" becomes "agent" | 1 | +| 3 | [retrieval (RAG)](03-retrieval.md) — `refine` | embeddings + vector search over a document set; carries its own sub-gap — an embeddings call over the same outbound path, plus a vector store (pure-`.wo` or a new primitive, decided in that story) | 1, and the embeddings/vector decision | Sketched, not committed — named so the shape is visible, not to schedule them: **model routing / multi-model** (choose a backend per request) and an **MCP diff --git a/docs/stories/jarvis/01-chat-loop.md b/docs/stories/jarvis/01-chat-loop.md new file mode 100644 index 0000000..04d6936 --- /dev/null +++ b/docs/stories/jarvis/01-chat-loop.md @@ -0,0 +1,96 @@ +--- +track: jarvis +iteration: "1" +status: pending +readiness: ready +review_pending: "forks auto-approved 2026-09-08 for autonomous execution — developer second review before code lands" +--- + +# jarvis 1 — the chat loop: a prompt in, streamed tokens out, durable history + +> Part of [Story — jarvis, the writeonce AI assistant](00-story.md). +> The whole end-to-end seam: a browser sends a prompt, jarvis relays it to an +> LLM over HTTPS, streams the reply back token by token, and persists the +> conversation durably. Nothing past this rung is worth building until it runs. +> +> **Forks auto-approved (2026-09-08) for autonomous execution and flagged in +> `review_pending` for the developer's second review** — the defaults below are +> reasonable but were not individually confirmed. + +## Blocked until the outbound seam lands + +This iteration cannot run until `net.connect` (✅ landed) and runtime-v2 +[9](../runtime-v2/09-in-process-tls.md) TLS reach **phase F** (`net.connect_tls` ++ the handshake). Crypto A–D are landed; E (X.509) and F (handshake) remain. +Everything below is buildable `.wo` on top of that seam plus porch 2/3/6/7. + +## Decisions locked (auto-approved, review pending) + +1. **Backend: the Anthropic Messages API over HTTPS**, streaming (SSE). One + backend for v1; a thin adapter boundary so an OpenAI-compatible backend can + slot later without touching the chat loop. Endpoint, model id and version + header come from config. +2. **Auth: an API key from the environment/config**, sent as the `x-api-key` + header — never in a URL, never logged. Missing key is a startup refusal, not + a runtime surprise. +3. **The relay is an actor per conversation.** A conversation actor owns the + upstream `net.connect_tls` connection, parses the LLM's SSE + (`content_block_delta` text deltas), and forwards each delta to the browser + through porch [7](../porch/07-sse-and-compression.md)'s SSE. Backpressure and + disconnect cleanup follow porch 7's contract (a write to a gone client frees + the fiber, the upstream fd and the actor). +4. **Durable history in two `@table` classes**: `Conversation { id @unique, + principal, created_at }` and `Message { conv_id (indexed), seq, role, + content, created_at }`. Keyed to the session principal (porch + [3](../porch/03-sessions.md)); history replays after a restart with no + external store — the writeonce differentiator. +5. **Route surface, session-gated**: `GET /` (chat UI via wo-html/writeonce-view), + `POST /message` (accept a prompt, append it, start the stream), `GET /stream` + (SSE of the reply). The POST is CSRF-protected once porch + [4](../porch/04-csrf.md) is built; until then it is bearer/session-gated and + the gap is stated. +6. **One turn at a time, no tools.** No function-calling, no retrieval — those + are iterations 2 and 3. The reply is a single streamed assistant message. + +## Phases + +- **A — the backend client.** Over `net.connect_tls`, issue the Messages + request and parse the streamed SSE deltas into text. Adapter boundary isolated + so the wire format is one file. +- **B — the conversation store.** The two `@table` classes, append-message, + load-history, list-conversations; durable-replay proven. +- **C — the relay + web surface.** The conversation actor, the porch routes, the + wo-html chat page, browser SSE of the reply; session-gated. +- **D — the gate and the ledger.** `just` a jarvis sample end to end against a + **local stub** LLM server (no network in the gate): prompt → streamed reply → + durable history → restart-replay. Both the happy path and a mid-stream + disconnect. Record what landed. + +## Acceptance Criteria + +- **Given** a signed-in user and a prompt, **when** it is sent, **then** the + reply streams back token by token and both prompt and reply are persisted. +- **Given** a restart, **when** the user returns, **then** their conversation + history replays from the `@table`, no external store. +- **Given** the client disconnecting mid-stream, **when** the next delta + arrives, **then** the upstream fd, the fiber and the actor are freed. +- **Given** a missing API key, **when** the app starts, **then** it refuses + loudly rather than failing at first request. +- **Given** the gate, **when** it runs, **then** it exercises the whole path + against a local stub with no live network call. + +## Out Of Scope + +- **Tool use / function calling** — iteration [2](00-story.md). +- **Retrieval / embeddings** — iteration [3](00-story.md). +- **Multiple backends at once / model routing** — sketched in the overview, + later. +- **The live network in the gate** — the gate uses a local stub server; a real + API smoke test is a manual, keyed, out-of-gate step. + +## Info + +Pure `.wo` on porch 2/3/6/7 + the TLS seam; no new runtime work of its own. The +one hard dependency is runtime-v2 9 reaching phase F. The auto-approved forks +(backend choice, auth transport, store schema, route surface) are the developer +second-review items flagged in `review_pending`. diff --git a/docs/stories/jarvis/02-tool-use.md b/docs/stories/jarvis/02-tool-use.md new file mode 100644 index 0000000..1fd4b06 --- /dev/null +++ b/docs/stories/jarvis/02-tool-use.md @@ -0,0 +1,50 @@ +--- +track: jarvis +iteration: "2" +status: pending +readiness: refine +--- + +# jarvis 2 — tool use / the agent loop + +> Part of [Story — jarvis, the writeonce AI assistant](00-story.md). +> Where the assistant becomes an agent: the model may call tools, and jarvis +> runs them and feeds results back until the model answers. Needs jarvis +> [1](01-chat-loop.md). + +## Why this exists + +A chat loop answers from the model alone. An agent can act — read a file, query +a `@table`, call a `.wo` function — by the model emitting a tool call, jarvis +executing it, and returning the result for another turn. The actor model fits +the multi-step orchestration naturally. + +## What it should deliver (to be refined) + +- A **tool registry**: named tools, each a `.wo` class with a typed input and a + handler, satisfying a `Tool` interface structurally (the porch middleware + precedent). +- The **agent loop**: relay the tool-use request, dispatch to the tool actor, + append the tool result, continue the turn — bounded to a max step count. +- **Tool-call framing** in the backend adapter (the Messages API `tool_use` / + `tool_result` blocks), isolated in the same adapter file jarvis 1 established. + +## Forks the brainstorm must settle + +1. **Tool declaration shape** — a `.wo` class + interface (no reflection, per + principle 13) versus a declarative schema; and how the input JSON maps to a + typed value without `@derive` (language 29). +2. **Sandboxing / permissions** — which tools are allowed, and whether a tool + that touches the filesystem or spawns a process (runtime-v2) needs an + explicit grant. +3. **Step bound + loop safety** — the max tool-call depth and how a runaway loop + is stopped. + +## Out of scope + +- Retrieval (iteration [3](00-story.md)); multi-model routing (later). + +## Info + +Pure `.wo` on iteration 1. No new runtime primitive expected unless a tool needs +one (each such tool names its own dependency). diff --git a/docs/stories/jarvis/03-retrieval.md b/docs/stories/jarvis/03-retrieval.md new file mode 100644 index 0000000..c686850 --- /dev/null +++ b/docs/stories/jarvis/03-retrieval.md @@ -0,0 +1,51 @@ +--- +track: jarvis +iteration: "3" +status: pending +readiness: refine +--- + +# jarvis 3 — retrieval (RAG): grounded answers over a document set + +> Part of [Story — jarvis, the writeonce AI assistant](00-story.md). +> Answers grounded in a corpus: embed documents, embed the query, retrieve the +> nearest chunks, and put them in the prompt. Needs jarvis [1](01-chat-loop.md) +> and the embeddings/vector decision below. + +## Why this exists + +A chat/agent loop knows only the model's training and the tools it can call. +Retrieval lets jarvis answer over the developer's own documents — the on-brand +"ask my codebase / my notes" use case — without fine-tuning. + +## What it should deliver (to be refined) + +- **Embeddings** for documents and queries — an embeddings API call over the + same outbound TLS path jarvis 1 uses. +- **A vector store**: chunk text, store `{chunk, vector}` durably, and a + similarity search (cosine / dot-product) over it. +- **Retrieval into the prompt**: top-k chunks prepended as context in the chat + loop. + +## Forks the brainstorm must settle + +1. **The vector store: pure `.wo` or a new runtime primitive.** A brute-force + cosine scan over a `@table` of `Bytes` vectors is pure `.wo` and fine for a + modest corpus; an approximate-nearest-neighbour index or a SIMD dot-product + builtin is a runtime iteration if scale demands it. Decide on a measured + need, not up front. +2. **Chunking strategy** — fixed-size vs semantic; overlap; where metadata + (source, offset) lives. +3. **Embedding storage** — vectors as `Bytes` (packed floats) in a `@table`, and + whether Float arrays need a better carrier than iteration 19's `Bytes`. + +## Out of scope + +- Re-ranking models, hybrid keyword+vector search, and multi-corpus tenancy — + each its own later slice. + +## Info + +Depends on jarvis 1's outbound path (for embeddings) and the vector-store fork. +The only possible new runtime work is fork 1's ANN/SIMD option, deferred until a +corpus size measures the need.