writeonce/docs/stories/jarvis/03-retrieval.md
shoney.arickathil 9d1689be55 docs(jarvis): create iteration stories 1 (ready) + 2/3 (refine)
- jarvis 1 (chat loop) brainstormed to ready with forks AUTO-APPROVED for
  autonomous execution and flagged in `review_pending` frontmatter for the
  developer's second review: Anthropic Messages API backend, env-var API key,
  actor-per-conversation SSE relay, durable @table history, session-gated routes
- jarvis 2 (tool use) + 3 (retrieval/RAG) created at refine with forks named
- 00-story iterations table linked to the new files
- blocked until rv2 9 TLS reaches phase F; pure .wo on porch 2/3/6/7 + the seam

(cherry picked from commit 8e160c3fcf9c50a05af073c036c693b192c9e595)
2026-09-15 01:15:31 +02:00

51 lines
1.9 KiB
Markdown

---
track: jarvis
iteration: "3"
status: pending
readiness: refine
---
# jarvis 3 — retrieval (RAG): grounded answers over a document set
> Part of [Story — jarvis, the writeonce AI assistant](00-story.md).
> Answers grounded in a corpus: embed documents, embed the query, retrieve the
> nearest chunks, and put them in the prompt. Needs jarvis [1](01-chat-loop.md)
> and the embeddings/vector decision below.
## Why this exists
A chat/agent loop knows only the model's training and the tools it can call.
Retrieval lets jarvis answer over the developer's own documents — the on-brand
"ask my codebase / my notes" use case — without fine-tuning.
## What it should deliver (to be refined)
- **Embeddings** for documents and queries — an embeddings API call over the
same outbound TLS path jarvis 1 uses.
- **A vector store**: chunk text, store `{chunk, vector}` durably, and a
similarity search (cosine / dot-product) over it.
- **Retrieval into the prompt**: top-k chunks prepended as context in the chat
loop.
## Forks the brainstorm must settle
1. **The vector store: pure `.wo` or a new runtime primitive.** A brute-force
cosine scan over a `@table` of `Bytes` vectors is pure `.wo` and fine for a
modest corpus; an approximate-nearest-neighbour index or a SIMD dot-product
builtin is a runtime iteration if scale demands it. Decide on a measured
need, not up front.
2. **Chunking strategy** — fixed-size vs semantic; overlap; where metadata
(source, offset) lives.
3. **Embedding storage** — vectors as `Bytes` (packed floats) in a `@table`, and
whether Float arrays need a better carrier than iteration 19's `Bytes`.
## Out of scope
- Re-ranking models, hybrid keyword+vector search, and multi-corpus tenancy —
each its own later slice.
## Info
Depends on jarvis 1's outbound path (for embeddings) and the vector-store fork.
The only possible new runtime work is fork 1's ANN/SIMD option, deferred until a
corpus size measures the need.