docs: six-bucket status board; recover lost vision doc; stop ignoring docs/

- 00-kanban.md rebuilt: ▶ NEXT PLAN pointer (iteration 4 — emitter, corpus,
  `woc build`) then six buckets — stories, in progress, done, pending,
  discarded, learnings. It tracked only the Rust runtime before, so the whole
  OOP track (wovm shipped, woc Tasks 1-8 shipped) was invisible.
- Board now records iteration 3's known gaps instead of silently owing them:
  `?T` plumbed but unenforced; E205/E201/E203/E204 dead, so structural
  interface satisfaction is unchecked.
- New discarded.md — settled rejections with reasons so they are not
  re-proposed: inheritance, `abstract` newtypes, Money/SKU/Float, Dynamic/cast/
  macro/extern, AOT-to-C, Menhir, shared engine state, external deployer.
- RECOVERED docs/plan/exploration/blue-green-vm/00-vision.md — gone from disk,
  never committed (gitignored path), cited by five docs incl. principle 12.
- Root cause was broader: all seven forward-roadmap plans in
  docs/superpowers/plans/ were untracked and ignored, on one disk only. Rules
  were half-fiction — 33 of 34 exploration files were already tracked, so they
  swallowed only *new* files.
- Dropped the docs ignore rules (exploration, oop-vm, superpowers/plans,
  examples/agent-loop, examples/mcp-think) with a do-not-re-add comment; added
  __pycache__/*.pyc. `tests/` stays ignored but warns that the next plan lands
  the corpus there.
This commit is contained in:
shoney.arickathil 2026-08-10 21:52:59 +02:00
parent 3b4372b913
commit d5c76d1dcf
21 changed files with 2246 additions and 149 deletions

35
.gitignore vendored
View file

@ -64,14 +64,29 @@
/runtime/bench/goref/goref
prototypes/llama-moe-stream/
prototypes/wo-db/
# NOTE (2026-08-10): `tests/` stays ignored, but story iteration 4 (plan 3)
# lands the conformance corpus under `tests/corpus/` — un-ignore it in that
# change, or the corpus will be invisible to git exactly as the docs below
# were. See docs/plan/learnings.md, "check that a new document is actually
# tracked".
tests/
docs/examples/agent-loop/
docs/examples/mcp-think/
docs/plan/exploration/
docs/plan/oop-vm/
# docs/plan/oop-vm/ carries the compiler<->VM normative contracts
# (.wob format, error catalog) both plan tracks cite -- carve it back
# out of the blanket docs/plan/ ignore above; the other rules here are
# untouched.
!docs/plan/oop-vm/
docs/superpowers/plans/
# Documentation is version-controlled — repo doctrine puts docs under `docs/`,
# and ignoring them there defeats the point. These directories were ignored
# until 2026-08-10, which silently cost the blue-green vision doc (recovered
# from a session transcript) and left all seven forward-roadmap plan docs in
# `docs/superpowers/plans/` existing only on one developer's disk. The rules
# were also half-fiction: 33 of the 34 files under `docs/plan/exploration/`
# were already tracked, so the rule only swallowed *new* files — the worst
# possible failure mode. Do not re-add them.
# docs/examples/agent-loop/
# docs/examples/mcp-think/
# docs/plan/exploration/
# docs/plan/oop-vm/ (carries the .wob format + error-catalog contracts)
# docs/superpowers/plans/
# Python bytecode — the agent-loop / mcp-think samples are Python, and
# un-ignoring their directories above exposed these.
__pycache__/
*.pyc

View file

@ -0,0 +1,148 @@
# agent-loop — build the loop that turns a model into an agent
A guided, working implementation of an **agent loop** — the mechanism at the
core of Claude Code, opencode, Cursor, and every other "agentic" tool — in one
Python file, standard library only, running entirely on the **local model**
served by [`prototypes/llama-moe-stream/start-local-agents.sh`](../../../prototypes/llama-moe-stream/start-local-agents.sh)
(or any OpenAI-compatible endpoint, e.g. Ollama).
You already run opencode against the local Qwen3-Coder server. This example is
what opencode *is*, with the product stripped away: read
[`agent.py`](agent.py) top to bottom and you know how every coding agent works.
## 1. The idea
A language model only ever produces text. What makes it an *agent* is a loop
around it that (a) tells it what tools exist, (b) executes the tool calls it
emits, and (c) feeds the results back — until it stops asking:
```
messages = [system, user question]
loop:
reply = POST /v1/chat/completions (messages + TOOL SCHEMAS)
append reply to messages
if reply has no tool_calls: # the model is done
print reply.content; stop
for each tool_call in reply:
result = run it locally # YOUR code — the model never executes anything
append {role: "tool", content: result} to messages
```
Two properties fall out of this shape, and they are the whole mental model:
- **The transcript is the only state.** `messages` grows append-only; the
model re-reads the entire history every round. There is no other memory —
which is why the assistant's own tool-call message must be appended too, not
just the results (the model has to see *what it asked for* next round).
- **The model proposes, your process disposes.** Tool calls are requests in
JSON. The executor decides what actually happens — it is the security
boundary, so the agent is exactly as dangerous as its tools, never more.
## 2. The five pieces (each maps to a section of `agent.py`)
| # | Piece | In `agent.py` |
| --- | --- | --- |
| 1 | **A tool-calling endpoint** — `llama-server --jinja` applies Qwen's chat template so tool schemas go in and structured `tool_calls` come out | `chat()` |
| 2 | **Tool schemas** — the JSON contract shown to the model; descriptions are prompts, write them like documentation | `TOOLS` |
| 3 | **The executor** — dispatch, argument parsing, sandboxing; every failure returned as words, never raised | `execute()` |
| 4 | **The transcript** — one append-only `messages` list | `run_turn()` |
| 5 | **Stop conditions** — natural (no `tool_calls`) and budgeted (`MAX_TURNS`) | `run_turn()` |
The tools here are deliberately read-only (`list_dir`, `read_file`, `search`)
and confined to `AGENT_ROOT` — enough to make a useful repo-Q&A agent with
zero risk while you study the loop.
## 3. Run it
```bash
# 1. Start the local model (from prototypes/llama-moe-stream)
prototypes/llama-moe-stream/start-local-agents.sh # Qwen3-Coder on :8080
# 2. One-shot, over this repo
cd /path/to/writeonce-all
python3 docs/examples/agent-loop/agent.py "which file implements the shard bus, and how do cross-shard reads work?"
# 3. Interactive — the conversation persists across questions
python3 docs/examples/agent-loop/agent.py
```
Tool calls print to stderr as they happen (`⚙ search({"pattern": ...})`), so
you watch the loop investigate before it answers.
Against Ollama instead: `LLM_URL=http://localhost:11434/v1 LLM_MODEL=qwen3:4b
python3 agent.py ...` (Ollama serves the same OpenAI-compatible `/v1`; small
models tool-call noticeably worse than Qwen3-Coder-30B — that difference is
itself instructive).
## 4. The guards that keep the loop alive
The naive loop works until the model misbehaves — and it will. The one rule:
**never raise at the model; return the failure as the tool result.** A
tool-calling model reads the error and corrects itself next round; an
exception just kills the conversation.
| Failure | Guard in `agent.py` |
| --- | --- |
| Arguments aren't valid JSON | error string back: "resend the call with corrected JSON" |
| Tool name doesn't exist | error string back, listing the real tools |
| Wrong/missing/extra parameters | `TypeError` caught → described back |
| Tool output too big for the context | truncated at 8k chars with an instruction to narrow |
| Path outside the project | `_resolve()` refuses (symlinks resolved first) |
| Model never stops calling tools | `MAX_TURNS` budget ends the turn with a readable note |
## 5. Smoke-test without a model
```bash
python3 docs/examples/agent-loop/test_loop.py
```
Stands up a canned `/chat/completions` on a local port and scripts three
rounds — a real `read_file`, a tool call with deliberately broken JSON
arguments, then a final answer — asserting that results are fed back, errors
self-correct, and the loop terminates. This is the loop's mechanics verified
in milliseconds, no GGUF required.
## Configuration
| Variable | Default | Meaning |
| --- | --- | --- |
| `LLM_URL` | `http://127.0.0.1:8080/v1` | OpenAI-compatible base URL |
| `LLM_MODEL` | `qwen3-coder` | model name (`--alias` of the llama-server) |
| `LLM_TIMEOUT` | `300` | per-request timeout, seconds |
| `AGENT_ROOT` | current directory | sandbox root for all three tools |
| `MAX_TURNS` | `20` | tool rounds per user message |
## How this relates to writeonce
This is the third corner of the local-agent triangle in this repo:
- [`mcp-think`](../mcp-think/) is the **tool side** — a local model offered
*as tools to* an agent (Claude).
- [`prototypes/llama-moe-stream`](../../../prototypes/llama-moe-stream/) is
the **model side** — serving the MoE this loop drives.
- **This example is the agent side** — the harness itself.
The natural next step joins them to [plan 15](../../plan/15-mcp-streamable-http.md):
once the writeonce runtime speaks MCP over Streamable HTTP, the hardcoded
`TOOLS` list here gets replaced by an **MCP client** — `tools/list` supplies
the schemas, `tools/call` becomes the executor (the JSON-RPC lifecycle is
already demonstrated in [`mcp-think/test_client.py`](../mcp-think/test_client.py)).
Then this same ~40-line loop lets a fully local model operate a `.wo`
application: `article_create`, `set_price`, live resources — one catalog, one
engine, one port, no cloud.
## Ideas to build next
Each is a small, self-contained extension of the loop — and each corresponds
to a feature you use daily in Claude Code/opencode:
- **MCP tool source** — fetch schemas from an MCP server at startup instead of
hardcoding `TOOLS`; route `execute()` through `tools/call`. (= MCP support)
- **A `write_file` tool behind a y/n prompt** — the executor asks *you* before
acting. (= permission modes)
- **A `spawn_agent` tool** that runs a fresh `run_turn()` with its own
transcript and returns only the final answer. (= sub-agents)
- **Transcript compaction** — when `messages` outgrows the context window,
summarize the older rounds into one message. (= auto-compact)
- **Parallel tool execution** — a reply may carry several `tool_calls`; run
them concurrently, append results in order. (= parallel tool use)

View file

@ -0,0 +1,285 @@
#!/usr/bin/env python3
"""agent-loop — a complete agent in one file, on a local model.
The whole trick behind Claude Code, opencode, Cursor and every other
"agentic" tool is one loop:
send the transcript + tool schemas to the model
while the model answers with tool calls:
run the tools, append the results to the transcript
send again
print the final text
Everything else those tools add (permissions, context management,
sub-agents) is elaboration on that loop. This file IS the loop, small
enough to read in one sitting: an OpenAI-compatible chat endpoint
(llama-server from prototypes/llama-moe-stream, or Ollama) + three
read-only repo tools + the guards that keep the loop alive when the
model misbehaves.
Run (one-shot): python3 agent.py "what does crates/rt/src/shard.rs do?"
Run (interactive): python3 agent.py
Configuration (environment):
LLM_URL OpenAI-compatible base URL (default http://127.0.0.1:8080/v1)
LLM_MODEL model name / alias (default qwen3-coder)
LLM_TIMEOUT per-request timeout in seconds (default 300)
AGENT_ROOT directory the tools may touch (default: current directory)
MAX_TURNS tool rounds per user message (default 20)
Dependencies: Python standard library only.
"""
import json
import os
import re
import sys
import urllib.error
import urllib.request
LLM_URL = os.environ.get("LLM_URL", "http://127.0.0.1:8080/v1").rstrip("/")
LLM_MODEL = os.environ.get("LLM_MODEL", "qwen3-coder")
LLM_TIMEOUT = int(os.environ.get("LLM_TIMEOUT", "300"))
ROOT = os.path.realpath(os.environ.get("AGENT_ROOT", os.getcwd()))
MAX_TURNS = int(os.environ.get("MAX_TURNS", "20"))
MAX_RESULT = 8_000 # chars of tool output fed back per call
SKIP_DIRS = {".git", "target", "node_modules", "build", "__pycache__", ".cache"}
# ---------------------------------------------------------------- the tools
# Schemas are the contract shown to the model; implementations are the
# security boundary. Read-only on purpose — an agent is exactly as dangerous
# as its tools, never more.
TOOLS = [
{"type": "function", "function": {
"name": "list_dir",
"description": "List one directory: entries with a trailing / for "
"subdirectories and a byte size for files.",
"parameters": {"type": "object", "properties": {
"path": {"type": "string",
"description": "directory, relative to the project root"},
}, "required": ["path"]}}},
{"type": "function", "function": {
"name": "read_file",
"description": "Read a text file with line numbers. Large files are "
"windowed — pass offset (1-based first line) and limit "
"(max lines) to page through.",
"parameters": {"type": "object", "properties": {
"path": {"type": "string", "description": "file, relative to the project root"},
"offset": {"type": "integer", "description": "first line to show, 1-based (default 1)"},
"limit": {"type": "integer", "description": "max lines to show (default 200)"},
}, "required": ["path"]}}},
{"type": "function", "function": {
"name": "search",
"description": "Search file contents under a directory with a Python "
"regular expression. Returns path:line: text matches. "
"Use a specific pattern — results cap at 100 matches.",
"parameters": {"type": "object", "properties": {
"pattern": {"type": "string", "description": "Python regex to find"},
"path": {"type": "string", "description": "directory to search, relative to the project root (default: whole root)"},
}, "required": ["pattern"]}}},
]
class ToolError(Exception):
"""A tool refusing to do something — reported to the model, never fatal."""
def _resolve(path: str) -> str:
"""Confine every path the model asks for to ROOT (symlinks resolved)."""
full = os.path.realpath(os.path.join(ROOT, path))
if full != ROOT and not full.startswith(ROOT + os.sep):
raise ToolError(f"path escapes the project root: {path}")
return full
def list_dir(path: str = ".") -> str:
full = _resolve(path)
if not os.path.isdir(full):
raise ToolError(f"not a directory: {path}")
rows = []
for name in sorted(os.listdir(full)):
p = os.path.join(full, name)
rows.append(f"{name}/" if os.path.isdir(p)
else f"{name} ({os.path.getsize(p)} bytes)")
return "\n".join(rows) or "(empty directory)"
def read_file(path: str, offset: int = 1, limit: int = 200) -> str:
full = _resolve(path)
if os.path.isdir(full):
raise ToolError(f"{path} is a directory — use list_dir")
try:
with open(full, errors="replace") as f:
lines = f.readlines()
except FileNotFoundError:
raise ToolError(f"no such file: {path}")
if not lines:
return "(empty file)"
offset = max(1, int(offset))
limit = max(1, min(int(limit), 1000))
window = lines[offset - 1: offset - 1 + limit]
if not window:
raise ToolError(f"{path} has only {len(lines)} lines; offset {offset} is past the end")
out = "".join(f"{i}\t{line}" for i, line in enumerate(window, offset))
last = offset + len(window) - 1
if last < len(lines):
out += (f"\n[agent] showing lines {offset}-{last} of {len(lines)} — "
f"call again with offset={last + 1} for more")
return out
def search(pattern: str, path: str = ".") -> str:
try:
rx = re.compile(pattern)
except re.error as e:
raise ToolError(f"bad regex {pattern!r}: {e}")
start = _resolve(path)
hits: list[str] = []
for dirpath, dirnames, filenames in os.walk(start): # never follows symlinks
dirnames[:] = sorted(d for d in dirnames if d not in SKIP_DIRS)
for fname in sorted(filenames):
full = os.path.join(dirpath, fname)
if os.path.islink(full) or os.path.getsize(full) > 2_000_000:
continue
try:
with open(full, errors="replace") as f:
head = f.read(1024)
if "\0" in head: # binary — skip
continue
f.seek(0)
for i, line in enumerate(f, 1):
if rx.search(line):
rel = os.path.relpath(full, ROOT)
hits.append(f"{rel}:{i}: {line.rstrip()[:200]}")
if len(hits) >= 100:
hits.append("[agent] 100-match cap hit — tighten the pattern or narrow the path")
return "\n".join(hits)
except OSError:
continue
return "\n".join(hits) or f"no matches for {pattern!r} under {path}"
TOOL_IMPLS = {"list_dir": list_dir, "read_file": read_file, "search": search}
# ---------------------------------------------------------- executing calls
# The one rule that keeps the loop alive: NEVER raise at the model. Whatever
# goes wrong — unknown tool, broken JSON, missing file — comes back as the
# tool result, in words. A tool-calling model reads the error and corrects
# itself on the next round; an exception would just kill the conversation.
def execute(call: dict) -> str:
fn_block = call.get("function") or {}
name = fn_block.get("name", "")
impl = TOOL_IMPLS.get(name)
if impl is None:
return f"[agent] unknown tool {name!r} — available: {', '.join(TOOL_IMPLS)}"
try:
args = json.loads(fn_block.get("arguments") or "{}")
except json.JSONDecodeError as e:
return f"[agent] arguments were not valid JSON ({e}) — resend the call with corrected JSON"
if not isinstance(args, dict):
return "[agent] arguments must be a JSON object"
try:
out = impl(**args)
except ToolError as e:
return f"[agent] {e}"
except TypeError as e:
return f"[agent] bad arguments for {name}: {e}"
except OSError as e:
return f"[agent] {name} failed: {e}"
if len(out) > MAX_RESULT:
out = (out[:MAX_RESULT] + f"\n[agent] truncated — {len(out)} chars total; "
"narrow the request (offset/limit, tighter pattern)")
return out
# ------------------------------------------------------------------ the loop
def chat(messages: list[dict]) -> dict:
"""One request to the model. Returns the assistant message verbatim."""
body = json.dumps({
"model": LLM_MODEL,
"messages": messages,
"tools": TOOLS,
"tool_choice": "auto",
}).encode()
req = urllib.request.Request(
f"{LLM_URL}/chat/completions",
data=body,
headers={"Content-Type": "application/json"},
)
try:
with urllib.request.urlopen(req, timeout=LLM_TIMEOUT) as resp:
data = json.load(resp)
except urllib.error.HTTPError as e:
detail = e.read().decode(errors="replace")[:400]
raise SystemExit(f"the model endpoint rejected the request (HTTP {e.code}): {detail}")
except (urllib.error.URLError, TimeoutError, OSError) as e:
raise SystemExit(
f"cannot reach the model at {LLM_URL} ({e}).\n"
"Start one first:\n"
" prototypes/llama-moe-stream/start-local-agents.sh # llama-server on :8080\n"
" LLM_URL=http://localhost:11434/v1 LLM_MODEL=qwen3:4b … # or a pulled Ollama model")
return data["choices"][0]["message"]
def run_turn(messages: list[dict]) -> str:
"""THE AGENT LOOP. Everything above exists to serve these few lines."""
for _ in range(MAX_TURNS):
msg = chat(messages)
messages.append(msg) # the transcript is the only state
calls = msg.get("tool_calls") or []
if not calls: # no tool call = the model is done
return _strip_think(msg.get("content") or "")
for call in calls:
fn = call.get("function") or {}
print(f" ⚙ {fn.get('name', '?')}({(fn.get('arguments') or '')[:120]})",
file=sys.stderr)
messages.append({
"role": "tool",
"tool_call_id": call.get("id", ""),
"content": execute(call),
})
return (f"[agent] stopped after {MAX_TURNS} tool rounds without a final answer — "
"ask a narrower question or raise MAX_TURNS")
def _strip_think(text: str) -> str:
# Reasoning models may inline chain-of-thought as <think>…</think>;
# the user wants the conclusion, not the scratchpad.
return re.sub(r"<think>.*?</think>", "", text, flags=re.DOTALL).strip()
# ----------------------------------------------------------------- the shell
SYSTEM = (f"You are a code assistant working inside the project rooted at {ROOT}. "
"Use the tools to look at real files before answering; never invent "
"file contents or paths. When you have enough evidence, answer "
"concisely and cite locations as path:line.")
def main() -> None:
messages: list[dict] = [{"role": "system", "content": SYSTEM}]
if len(sys.argv) > 1: # one-shot
messages.append({"role": "user", "content": " ".join(sys.argv[1:])})
print(run_turn(messages))
return
print(f"agent-loop — model {LLM_MODEL} at {LLM_URL}\n"
f"root {ROOT} (Ctrl-D to exit; the conversation persists across questions)")
while True: # interactive
try:
line = input("\nyou> ").strip()
except EOFError:
print()
return
if not line:
continue
messages.append({"role": "user", "content": line})
print(run_turn(messages))
if __name__ == "__main__":
main()

View file

@ -0,0 +1,105 @@
#!/usr/bin/env python3
"""Smoke-test the agent loop without any model.
Stands up a canned OpenAI-compatible /chat/completions endpoint on a local
port and drives agent.run_turn() through a scripted conversation:
round 1: the "model" calls read_file on notes.txt -> executed for real
round 2: it sends a tool call with broken JSON args -> fed back as an error, not a crash
round 3: it returns a final answer
This verifies the three load-bearing behaviours of the loop: tool execution
with result feedback, errors-as-results self-correction, and termination on
a plain (tool-call-free) answer.
Usage: python3 test_loop.py # standard library only, no model needed
"""
import json
import os
import sys
import tempfile
import threading
from http.server import BaseHTTPRequestHandler, ThreadingHTTPServer
HERE = os.path.dirname(os.path.abspath(__file__))
MARKER = "the WAL fsyncs before acknowledging the commit"
# What the fake model answers, round by round.
SCRIPT = [
{"role": "assistant", "content": None, "tool_calls": [
{"id": "call_1", "type": "function", "function": {
"name": "read_file",
"arguments": json.dumps({"path": "notes.txt"})}}]},
{"role": "assistant", "content": None, "tool_calls": [
{"id": "call_2", "type": "function", "function": {
"name": "search",
"arguments": '{"pattern": '}}]}, # deliberately broken JSON
{"role": "assistant",
"content": f"FINAL: per notes.txt, {MARKER}."},
]
REQUESTS: list[dict] = []
class MockModel(BaseHTTPRequestHandler):
def do_POST(self):
REQUESTS.append(json.loads(self.rfile.read(int(self.headers["Content-Length"]))))
body = json.dumps({"choices": [{"message": SCRIPT[len(REQUESTS) - 1]}]}).encode()
self.send_response(200)
self.send_header("Content-Type", "application/json")
self.send_header("Content-Length", str(len(body)))
self.end_headers()
self.wfile.write(body)
def log_message(self, format, *args):
pass
def main() -> int:
server = ThreadingHTTPServer(("127.0.0.1", 0), MockModel)
threading.Thread(target=server.serve_forever, daemon=True).start()
workdir = tempfile.mkdtemp(prefix="agent-loop-test-")
with open(os.path.join(workdir, "notes.txt"), "w") as f:
f.write(f"Durability rule: {MARKER}.\n")
# agent.py reads its configuration at import time — set env first.
os.environ["LLM_URL"] = f"http://127.0.0.1:{server.server_address[1]}/v1"
os.environ["LLM_MODEL"] = "mock"
os.environ["AGENT_ROOT"] = workdir
sys.path.insert(0, HERE)
import agent
answer = agent.run_turn([
{"role": "system", "content": "test"},
{"role": "user", "content": "what is the durability rule?"},
])
server.shutdown()
def tool_results(request: dict) -> list[str]:
return [m["content"] for m in request["messages"] if m.get("role") == "tool"]
assert len(REQUESTS) == 3, f"expected 3 model rounds, got {len(REQUESTS)}"
# Round 2's request must carry the real file content back as a tool result.
round2 = tool_results(REQUESTS[1])
assert any(MARKER in r for r in round2), f"file content not fed back: {round2}"
print("ok — tool call executed, result appended to the transcript")
# Round 3's request must carry the JSON error as a result, not a crash.
round3 = tool_results(REQUESTS[2])
assert any("[agent]" in r and "JSON" in r for r in round3), \
f"broken arguments not reported back: {round3}"
print("ok — malformed tool arguments came back as an error result")
assert answer.startswith("FINAL:") and MARKER in answer, f"unexpected answer: {answer}"
print("ok — loop terminated on the tool-call-free answer")
print(f"\nOK — the agent loop round-trips tools, self-corrects, and stops.\n"
f"Now run it against a real model: python3 {os.path.join(HERE, 'agent.py')}")
return 0
if __name__ == "__main__":
sys.exit(main())

View file

@ -0,0 +1,92 @@
# mcp-think — a local-model MCP tool for Claude
A minimal, working **MCP server** that Claude (Claude Code or Claude Desktop) can call as a tool — and whose work runs entirely on **your machine**, on a local model served by [Ollama](https://ollama.com). One Python file, stdio transport, standard library only (plus the official `mcp` package).
Why offload from Claude to a local model:
| Reason | Example |
| --- | --- |
| **Privacy** | `summarize` a confidential document — the text never leaves the box |
| **Cost** | `brainstorm` 20 ideas or condense a 50k-line log at zero token cost |
| **Diversity** | `critique` from a different model family — an independent second opinion |
| **Bandwidth** | Claude stays on the main task while the local model grinds a side job |
## Tools
| Tool | What it does |
| --- | --- |
| `think(task, context?)` | Reason through a problem; conclusion first, reasoning after |
| `critique(work, focus?)` | Adversarial review — strongest objections, ranked |
| `brainstorm(topic, n?)` | `n` genuinely distinct ideas |
| `summarize(text, max_words?)` | Faithful local summary of large/sensitive material |
## Prerequisites
```bash
# 1. Ollama, serving a model
ollama serve # if not already running as a service
ollama pull qwen3:4b # the default model (or any other — see Configuration)
# 2. uv (provisions the `mcp` package on the fly) — or `pip install mcp`
```
For a quick low-disk trial, `ollama pull qwen3:0.6b` (~500 MB) works — set `THINK_MODEL=qwen3:0.6b`.
## Add to Claude Code
```bash
claude mcp add think -- uv run --with mcp python3 \
/abs/path/to/docs/examples/mcp-think/server.py
```
(replace with your absolute path; add `-e THINK_MODEL=...` before `--` to override the model). Or declare it in a project's `.mcp.json`:
```json
{
"mcpServers": {
"think": {
"command": "uv",
"args": ["run", "--with", "mcp", "python3",
"/abs/path/to/docs/examples/mcp-think/server.py"],
"env": { "THINK_MODEL": "qwen3:4b" }
}
}
}
```
Claude Desktop: the same `command`/`args`/`env` block goes under `mcpServers` in `claude_desktop_config.json`.
Then just ask: *"use the think tool to weigh epoll vs io_uring for the WAL path"*, *"brainstorm 10 names for this feature"*, *"summarize this log with the local model"* — or let Claude reach for the tools on its own (the tool descriptions say when each applies).
## Configuration
| Variable | Default | Meaning |
| --- | --- | --- |
| `OLLAMA_URL` | `http://localhost:11434` | Ollama daemon base URL |
| `THINK_MODEL` | `qwen3:4b` | any pulled Ollama model |
| `THINK_TIMEOUT` | `300` | per-call timeout, seconds |
Errors (daemon down, model not pulled) come back as readable tool results, so Claude can tell you exactly what to run.
## Smoke test without Claude
```bash
python3 docs/examples/mcp-think/test_client.py # uses uv + THINK_MODEL if set
```
Drives the server over stdio exactly like Claude does — `initialize` → `tools/list` → one `think` call — and prints the local model's answer.
## How this relates to writeonce
This example is the **consumer side** of MCP: Claude ← stdio → this server → local model. The **server side** is [plan 15](../../plan/15-mcp-streamable-http.md): the writeonce runtime itself becomes an MCP server over Streamable HTTP, generating tools/resources from `.wo` declarations — after 15d, a class method like [`Product.set_price`](../pricing/types/product.wo) is callable as an MCP tool with no Python in between. The two meet naturally: a writeonce app exposes its data as MCP tools, and a local model (through a server like this one) works over it.
## Ideas to build next
Same pattern (one file, stdio, local model), different capability:
- **mcp-redact** — strip PII/secrets from text locally *before* it is sent to any cloud model.
- **mcp-embed** — local embeddings (`ollama embed`) + a small vector index over a repo or notes; gives Claude semantic search without a cloud vector DB.
- **mcp-vision** — describe screenshots/diagrams via a local vision model (`llava`, `qwen2.5-vl`) for machines where images must stay local.
- **mcp-translate** — private translation of documents.
- **mcp-testgen** — bulk-generate test fixtures/edge cases on the local model, free of token cost.
- **mcp-judge** — a local LLM-as-judge for scoring outputs in eval loops, so the judge is independent of the model being judged.

View file

@ -0,0 +1,147 @@
#!/usr/bin/env python3
"""mcp-think — offload thinking to a LOCAL model, exposed to Claude over MCP.
A minimal MCP server (stdio transport) whose tools run entirely on your
machine: each tool sends a prompt to a local model served by Ollama and
returns the answer. Nothing in the tool inputs ever leaves the box.
Why offload to a local model from Claude?
* privacy — summarize/critique content that must not leave the machine
* cost — bulk work (summaries, drafts, idea generation) at zero token cost
* diversity — a second opinion from a different model family
* bandwidth — Claude stays on the main task while the local model grinds
Run (Claude Code):
claude mcp add think -- uv run --with mcp python3 /abs/path/to/server.py
Configuration (environment):
OLLAMA_URL base URL of the Ollama daemon (default http://localhost:11434)
THINK_MODEL model to use (default qwen3:4b)
THINK_TIMEOUT per-call timeout in seconds (default 300)
Dependencies: the `mcp` package only (`uv run --with mcp` provisions it).
The Ollama call uses the Python standard library — no SDK, no httpx.
"""
import json
import os
import re
import urllib.error
import urllib.request
# The SDK is renaming FastMCP → MCPServer; support both so the example works
# with the released `mcp` package today and with the renamed one later.
try:
from mcp.server.mcpserver import MCPServer as Server # unreleased SDK head
except ImportError:
from mcp.server.fastmcp import FastMCP as Server # mcp <= 1.x (PyPI)
OLLAMA_URL = os.environ.get("OLLAMA_URL", "http://localhost:11434").rstrip("/")
THINK_MODEL = os.environ.get("THINK_MODEL", "qwen3:4b")
THINK_TIMEOUT = int(os.environ.get("THINK_TIMEOUT", "300"))
mcp = Server("think")
def ask_local_model(system: str, prompt: str) -> str:
"""One chat round against the local model. Errors come back as text so
the calling model can read them and tell the user what to fix."""
body = json.dumps({
"model": THINK_MODEL,
"messages": [
{"role": "system", "content": system},
{"role": "user", "content": prompt},
],
"stream": False,
}).encode()
req = urllib.request.Request(
f"{OLLAMA_URL}/api/chat",
data=body,
headers={"Content-Type": "application/json"},
)
try:
with urllib.request.urlopen(req, timeout=THINK_TIMEOUT) as resp:
data = json.load(resp)
except urllib.error.HTTPError as e:
detail = e.read().decode(errors="replace")[:300]
return (f"[mcp-think] the local model rejected the request "
f"(HTTP {e.code}): {detail}\n"
f"Is the model pulled? Try: ollama pull {THINK_MODEL}")
except (urllib.error.URLError, TimeoutError, OSError) as e:
return (f"[mcp-think] cannot reach the local model at {OLLAMA_URL} ({e}).\n"
f"Start it with `ollama serve`, pull the model with "
f"`ollama pull {THINK_MODEL}`, then retry.")
text = (data.get("message") or {}).get("content", "")
# Reasoning models may inline their chain of thought as <think>…</think>;
# strip it — the caller wants the conclusion, not the scratchpad.
text = re.sub(r"<think>.*?</think>", "", text, flags=re.DOTALL).strip()
return text or "[mcp-think] the local model returned an empty response"
@mcp.tool()
def think(task: str, context: str = "") -> str:
"""Reason through a problem on the local model and return its conclusion.
Call this to get an independent, private, zero-cost analysis of a
question — weighing a design trade-off, sanity-checking a plan, working
through logic — especially when a second opinion from a different model
family is valuable. `context` carries any background the task needs
(code, notes, constraints); it never leaves this machine.
"""
system = ("You are a careful reasoning assistant. Think the problem "
"through step by step, then answer with your conclusion first, "
"followed by the key reasoning in a few short paragraphs.")
prompt = f"{task}\n\n--- context ---\n{context}" if context else task
return ask_local_model(system, prompt)
@mcp.tool()
def critique(work: str, focus: str = "") -> str:
"""Adversarially review a piece of work (code, prose, a plan, a design)
on the local model and return the strongest objections.
Call this before committing to a decision or shipping a draft, when an
independent devil's advocate is useful, or when the material is private
and must be reviewed without leaving the machine. `focus` optionally
narrows the review (e.g. "error handling", "argument structure").
"""
system = ("You are a rigorous, adversarial reviewer. Find the strongest "
"objections: errors, gaps, risks, unstated assumptions. Rank "
"them most-severe first. Be specific — point at the exact "
"part you object to and say why. Do not pad with praise.")
prompt = f"Review the following{f', focusing on {focus}' if focus else ''}:\n\n{work}"
return ask_local_model(system, prompt)
@mcp.tool()
def brainstorm(topic: str, n: int = 5) -> str:
"""Generate `n` distinct ideas on a topic using the local model.
Call this to widen the option space cheaply before converging — naming,
approaches, test cases, failure modes, feature ideas. The local model's
different training makes its ideas usefully different from yours.
"""
system = ("You are a prolific idea generator. Produce genuinely distinct "
"ideas — different mechanisms, not rephrasings. One line of "
"pitch plus one line of how it would work, per idea.")
return ask_local_model(system, f"Generate {n} distinct ideas for: {topic}")
@mcp.tool()
def summarize(text: str, max_words: int = 200) -> str:
"""Summarize text on the local model — the content never leaves this
machine and costs no API tokens.
Call this to condense large material (logs, documents, transcripts,
diffs) before reasoning about it, or when the content is sensitive and
must stay local. The summary comes back; the original stays here.
"""
system = (f"Summarize faithfully in at most {max_words} words. Keep "
"concrete facts, numbers, names, and conclusions; drop filler. "
"Note anything surprising or anomalous explicitly.")
return ask_local_model(system, text)
if __name__ == "__main__":
mcp.run() # stdio — Claude launches this as a subprocess

View file

@ -0,0 +1,85 @@
#!/usr/bin/env python3
"""Smoke-test mcp-think over stdio, exactly the way Claude drives it.
Launches server.py as a subprocess (via `uv run --with mcp`), performs the
MCP lifecycle (initialize → initialized), lists the tools, and calls
`think` once. Requires Ollama running with the configured model pulled.
Usage: python3 test_client.py # standard library only
THINK_MODEL=qwen3:0.6b python3 test_client.py
"""
import json
import os
import subprocess
import sys
HERE = os.path.dirname(os.path.abspath(__file__))
def main() -> int:
proc = subprocess.Popen(
["uv", "run", "--with", "mcp", "python3", os.path.join(HERE, "server.py")],
stdin=subprocess.PIPE, stdout=subprocess.PIPE, stderr=subprocess.DEVNULL,
env=os.environ.copy(), text=True,
)
stdin, stdout = proc.stdin, proc.stdout
assert stdin is not None and stdout is not None
msg_id = 0
def send(method: str, params: dict | None = None, notify: bool = False) -> dict:
nonlocal msg_id
msg: dict = {"jsonrpc": "2.0", "method": method}
if params is not None:
msg["params"] = params
if not notify:
msg_id += 1
msg["id"] = msg_id
stdin.write(json.dumps(msg) + "\n")
stdin.flush()
if notify:
return {}
# stdio transport: one JSON-RPC message per line
line = stdout.readline()
if not line:
raise RuntimeError("server closed the pipe")
return json.loads(line)
try:
r = send("initialize", {
"protocolVersion": "2025-06-18",
"capabilities": {},
"clientInfo": {"name": "mcp-think-smoke", "version": "0"},
})
server = r["result"]["serverInfo"]
print(f"initialize ok — server: {server['name']}")
send("notifications/initialized", {}, notify=True)
r = send("tools/list", {})
tools = [t["name"] for t in r["result"]["tools"]]
print(f"tools/list ok — {tools}")
assert {"think", "critique", "brainstorm", "summarize"} <= set(tools)
print("tools/call think(...) — waiting on the local model…")
r = send("tools/call", {
"name": "think",
"arguments": {"task": "In one sentence: why do write-ahead logs "
"fsync before acknowledging a commit?"},
})
content = r["result"]["content"][0]["text"]
print(f"\n--- local model says ---\n{content}\n")
if content.startswith("[mcp-think]"):
print("NOT OK — the server answered, but the local model is unavailable.")
return 1
print("OK — MCP lifecycle, tool discovery, and local-model call all work.")
return 0
finally:
stdin.close()
proc.terminate()
proc.wait(timeout=5)
if __name__ == "__main__":
sys.exit(main())

View file

@ -1,94 +1,189 @@
# Plan kanban — backend: runtime, database, API
# Status board — what is done, what is next
Status board for every phase doc under `docs/plan/`. **Focus: the backend** — the runtime, the database engine, and the REST/`.wo`-language API. Frontend phases are parked, not deleted. Each phase file carries a matching status banner; this board is the index.
The single place to learn where this project stands. Organised in six buckets:
**stories** (the narrative arc), **in progress**, **done**, **pending**,
**discarded**, **learnings**. Every phase doc carries a matching status banner;
this board is the index.
Related planning surfaces this board indexes across: the mission story + review-in-order iterations at [`docs/stories/language-runtime-database/`](../stories/language-runtime-database/00-story.md); approved designs at [`docs/superpowers/specs/`](../superpowers/specs/); their implementation plans at [`docs/superpowers/plans/`](../superpowers/plans/) and [`docs/plan/compiler/`](compiler/architecture.md). Entry chain: [`docs/00-principles.md`](../00-principles.md) → [`docs/08-project-structure.md`](../08-project-structure.md) → this board.
Update this board in the same change that finishes work — move the item to done
with *what actually landed*, set the next in-progress item, and record any
rejection in [`discarded.md`](discarded.md) with its reason.
Statuses: ✅ **done** · 🔄 **in progress** · ⬜ **not started** · ⏸ **parked (out of backend focus)**
Statuses: ✅ **done** · 🔄 **in progress** · ⬜ **pending** · ⏸ **parked**
## Board
---
### Track 1 — Runtime foundations (dependency removal, kernel primitives)
## ▶ NEXT PLAN
**Story iteration 4 — single binary end-to-end.**
Plan: [`compiler/plan/2026-08-01-wob-emit-e2e-single-binary.md`](compiler/2026-08-01-wob-emit-e2e-single-binary.md) ·
Story slice: [`docs/stories/language-runtime-database/04-single-binary-e2e.md`](../stories/language-runtime-database/04-single-binary-e2e.md)
The bytecode emitter, the three-kind conformance corpus, and `woc build`. This
is the milestone where `.wo` source becomes a running self-contained binary —
compiler front (iteration 3) and VM core (iteration 2) both shipped, so it is
unblocked. Its inputs are the four ownership tables `owner.ml` now produces;
`dump.ml`'s format-contract comments are normative for it, **including the
requirement to coalesce borrow guards per operand**.
Two tracks run in this repo. The critical path is the **language track**:
iterations 3 → 4 → 5 → 6 → 7, ending at *compile and run log-watcher*. The
Rust-runtime track is shipped-and-maintained, not advancing.
---
## Stories
[`docs/stories/language-runtime-database/`](../stories/language-runtime-database/00-story.md)
— one language, one runtime, one database, one binary. Twelve iterations, each
an unsplittable slice with Given/When/Then acceptance and a pointer to the plan
that sequences its tasks. Read one, approve, then the next starts.
| # | Iteration | State |
| --- | --- | --- |
| 1 | [Principles doc](../stories/language-runtime-database/01-principles-doc.md) | ✅ |
| 2 | [VM core (`wovm`)](../stories/language-runtime-database/02-vm-core.md) | ✅ |
| 3 | [Compiler front (`woc`)](../stories/language-runtime-database/03-compiler-front.md) | ✅ (known gaps below) |
| 4 | [Single binary end-to-end](../stories/language-runtime-database/04-single-binary-e2e.md) | 🔄 **next** |
| 5 | [Language surface](../stories/language-runtime-database/05-language-surface.md) | ⬜ |
| 6 | [Program mode + stdlib](../stories/language-runtime-database/06-program-mode-stdlib.md) | ⬜ |
| 7 | [log-watcher proof](../stories/language-runtime-database/07-logwatcher-proof.md) | ⬜ acceptance |
| 8 | [Shard-actor runtime](../stories/language-runtime-database/08-shard-actor-runtime.md) | ⬜ |
| 9 | [Database engine](../stories/language-runtime-database/09-database-engine.md) | ⬜ |
| 10 | [HTTP service layer](../stories/language-runtime-database/10-http-service.md) | ⬜ |
| 11 | [Fibers](../stories/language-runtime-database/11-fibers.md) | ⬜ |
| 12 | [Blue-green deploy](../stories/language-runtime-database/12-blue-green-deploy.md) | ⬜ |
---
## In progress
| Track | Item | Where |
| --- | --- | --- |
| Language | Iteration 4 — emitter, conformance corpus, `woc build` | [plan 3](compiler/2026-08-01-wob-emit-e2e-single-binary.md) |
Nothing else should be started until iteration 4 lands. Off-critical-path work
is parked by explicit scope directive (2026-08-08).
---
## Done
### Language track — compiler + VM (OOP track)
| Status | Item | Doc | What actually landed |
| --- | --- | --- | --- |
| ✅ | Principles | [`../00-principles.md`](../00-principles.md) | 13 principles, each with a why and a link to the doc that enforces it |
| ✅ | `wovm` VM core | [plan 1](../superpowers/plans/2026-08-01-wob-format-and-vm-core.md) | `.wob` v1 loader with full static validation, register interpreter (computed-goto + ISO-C fallback), arena with size-class free lists, borrow word, RC + budgeted Bacon–Rajan cycle collector, drop-map trap unwinding, containers, builtins, ICALL, CLI. 13 suites × 2 dispatch flavors + CLI smoke, ASan/UBSan clean |
| ✅ | `.wob` format contract | [`oop-vm/00-wob-format.md`](oop-vm/00-wob-format.md) | Normative; twinned with `runtime/src/wob.h` |
| ✅ | `woc` compiler front | [plan 2](compiler/2026-08-01-woc-compiler-front.md) | Tasks 1–8: dune scaffold, `diag` (WO-E codes, two-site related errors, ordered dedup), newline-significant lexer at rt parity, declaration + statement/expression parser with skip-on-block and multi-error recovery, typechecker (field kinds, `?T` plumbing, W201, E225, E214), MVS ownership pass with the four emitter tables, driver with directory discovery + cross-file programs. 14 + 264 checks |
| ✅ | Error catalog | [`oop-vm/01-error-catalog.md`](oop-vm/01-error-catalog.md) | 14 emitted codes + 10 reserved, each with the reason it is not yet emitted |
| ✅ | log-watcher `.wo` sample | [`../examples/log-watcher/`](../examples/log-watcher/README.md) | Eight-file port authored docs-first with its `.hx` mapping table; compiles for real in iteration 7 |
| ✅ | Scalar cleanup | [`discarded.md`](discarded.md) | `Money`/`SKU`/`Float` and the abstract allowlist removed; `abstract` flipped adopt → reject |
**Known gaps carried out of iteration 3** — recorded, not silently owed:
- **`?T` is plumbed but unenforced.** Lexer/token/AST/parser/dump all handle
`?T`; the semantics do not exist (`WO-E211`/`E212`/`E213` declared, never
emitted — a probe returning `?Int` as `Int` exits 0). Owned by iteration 5,
plan 8 Task 6, which is that iteration's first task because it blocks the
log-watcher port. See [`compiler/nullable-types-implementation.md`](compiler/nullable-types-implementation.md).
- **Structural interface satisfaction is not checked** (`WO-E205` dead), along
with type mismatch, bad arity, and unknown-fn (`E201`/`E203`/`E204`) — all
named in plan 2 Task 6's own must-fail list. Gaps in shipped work, catalogued
as reserved.
- Six further narrowings (W201 heuristic, E225 reach, dead code after `return`,
unresolved-callee drops, RC table ordering, residual b-side role) are listed
in the plan-2 SDD ledger and in the affected files' own comments.
### Rust runtime track — Stage 2 shipped, maintained
| Status | Phase | Doc | Notes |
| --- | --- | --- | --- |
| ✅ | 01 crate scaffolding | [done/01](done/01-scafolding-crates.md) | 15 crates |
| ✅ | 02 epoll event loop | [done/02](done/02-event-loop-epoll.md) | `runtime/netpoll_epoll.rs` |
| ✅ | 03 hand-rolled HTTP | [done/03](done/03-hand-rolled-http.md) | + keep-alive & pipelining (follow-up under plan 09) |
| ✅ | 03 hand-rolled HTTP | [done/03](done/03-hand-rolled-http.md) | + keep-alive & pipelining |
| ✅ | 04 tokio/axum cutover | [done/04](done/04-cutover-remove-tokio-axum.md) | deps now: anyhow, serde, serde_json, libc |
| ⬜ | 05 hand-rolled JSON | [05](05-hand-rolled-json.md) | removes serde/serde_json — **next in this track** |
| ✅ | 09a thread-per-core | [09](09-concurrency-scaleout.md) | `scheduler.rs`, `SO_REUSEPORT`, pinned `wo-shard-<t>` workers |
| ✅ | 09b sharded engine | [09](09-concurrency-scaleout.md) | `shard.rs` bus; `Arc<Mutex<Engine>>` deleted; interleaved ids |
| ✅ | 09c per-shard WAL | [09](09-concurrency-scaleout.md) | ack-after-fsync; boot replay; `meta` shard guard |
| ✅ | — keep-alive follow-up | [09](09-concurrency-scaleout.md) | reads ×3.4 → 770k/s |
| ✅ | — io_uring group commit | [09](09-concurrency-scaleout.md) | raw ring; 4.7× durable writes on real disk |
| ✅ | 16a PG wire client | [16](16-postgres-mirror.md) | hand-rolled protocol v3, zero crates |
| ✅ | 16b PG backup mirror | [16](16-postgres-mirror.md) | async JSONB upserts behind the WAL ack; RAM authoritative |
| ✅ | 13a class surface | [13](13-class-model-live-pricing.md) | `class` parses, CRUD serves |
| ✅ | 13b method execution | [13](13-class-model-live-pricing.md) | row-scoped txn per call; abort → 409 rollback |
| ✅ | — `@table` + indexed DML | [13](13-class-model-live-pricing.md) | secondary indexes, `find_by`, `select Type{…}`, REST filters |
| ✅ | C proving ground A–F | [exploration/c-runtime/00-plan.md](exploration/c-runtime/00-plan.md) | 859k reads/s, 618k durable commits/s; found the ack-ordering + fd-ABA bugs the Rust port avoided |
Ecommerce sample (verified 2026-06-13): `api.rest` 17/17 expected statuses pass.
---
## Pending
### Language track — sequenced, on the critical path
| # | Item | Plan |
| --- | --- | --- |
| 5 | Haxe-parity language surface — **`?T` forced handling first**, then switch expressions, records, enum payloads, try/catch, statics, `using`, modules, `is`, `pub(read)`, `#if` | [plan 8](compiler/2026-08-01-haxe-parity-language.md) |
| 6 | Program mode + systems stdlib — `fn main`, exit codes, `fs`/`proc`/`net`/`time`/`json` | [plan 9](../superpowers/plans/2026-08-01-program-mode-stdlib.md) |
| 7 | log-watcher proof — the sample compiles and detects a silent death live | [plan 10](../superpowers/plans/2026-08-01-log-watcher-sample.md) |
| 8 | Shard-actor runtime | [plan 4](../superpowers/plans/2026-08-01-shard-actor-vm-runtime.md) |
| 9 | Database engine binding | [plan 5](../superpowers/plans/2026-08-01-db-engine-binding.md) |
| 10 | HTTP service layer | [plan 6](../superpowers/plans/2026-08-01-http-service-layer.md) |
| 11 | Fibers | vision §3, [blue-green exploration](exploration/blue-green-vm/00-vision.md) |
| 12 | Blue-green deploy | [spec](../superpowers/specs/2026-08-03-blue-green-vm-design.md) — plan authored after iterations 9–10 |
### Language track — parked until after iteration 12
Recorded 2026-08-08 by scope directive; nothing here lands before the
log-watcher proof.
- `WO-W201` `@gc`-suggestion refinement beyond the self-reference heuristic
- `WO-E225` broadened to `ref`/`multi`/`map` element types and fn signatures
- ADT container roster adoption (Stack, Queue, Set, Tree, Graph, …) — see the
roster in [`compiler/nullable-types-implementation.md`](compiler/nullable-types-implementation.md)
- Web framework as a `.wo` library; UI (`##ui` SSR + live patches);
script-based destructive migrations; MCP/agent wrapper over the management plane
### Rust runtime track — not advancing while the language track runs
| Status | Phase | Doc | Notes |
| --- | --- | --- | --- |
| ⬜ | 05 hand-rolled JSON | [05](05-hand-rolled-json.md) | removes serde/serde_json |
| ⬜ | 06 bespoke error type | [06](06-bespoke-error.md) | removes anyhow |
| ⬜ | 07 inotify content watcher | [07](07-inotify-content-watcher.md) | `wo dev` hot reload |
| ⬜ | 08 sendfile static assets | [08](08-sendfile-static-assets.md) | needed by parked UI track too |
| ⬜ | 08 sendfile static assets | [08](08-sendfile-static-assets.md) | needed by the parked UI track |
| ⬜ | 09d cross-shard subscriptions | [09](09-concurrency-scaleout.md) | LIVE fan-out; pairs with Stage 3 |
| ⬜ | 09e cross-shard transactions (2PC) | [09](09-concurrency-scaleout.md) | needed by `fn checkout` spanning shards |
| ⬜ | 09f observability & reshard | [09](09-concurrency-scaleout.md) | per-shard metrics, `WO_RESHARD` |
| ⬜ | 10–12 storage completion | [10](10-storage-foundations.md), [11](11-wal-and-recovery.md), [12](12-engine-disk-cutover.md) | snapshots, compaction, WAL rotation, mmap arena engine |
| ⬜ | 13c LIVE pricing push · Stage 3 wire layer · 13e at scale | [13](13-class-model-live-pricing.md) | replaces the 501 stub |
| ⬜ | 15a–15e MCP over streamable HTTP | [15](15-mcp-streamable-http.md) | 15e needs 13c + 09d |
| ⬜ | 16c–16f typed columns, lossless resync, restore, SCRAM | [16](16-postgres-mirror.md) | |
### Track 2 — Concurrency & scale-out (plan 09) — 🔄 in progress
| Status | Sub-phase | Notes |
| --- | --- | --- |
| ✅ | 09a thread-per-core | `runtime/scheduler.rs`, `SO_REUSEPORT`, pinned `wo-shard-<t>` workers |
| ✅ | 09b sharded engine | `shard.rs` bus; `Arc<Mutex<Engine>>` deleted; interleaved ids; fan-out lists |
| ✅ | 09c per-shard WAL | `wal.rs`; ack-after-fsync; boot replay; `meta` shard guard |
| ✅ | — keep-alive (follow-up) | reads ×3.4 → 770k/s; C phase-C sequence |
| ✅ | — io_uring group commit (follow-up) | `netpoll_io_uring.rs` raw ring; 4.7× durable writes on real disk |
| ⬜ | 09d cross-shard subscriptions | LIVE fan-out, one message per shard — pairs with Stage 3 |
| ⬜ | 09e cross-shard transactions (2PC) | needed by `fn checkout` spanning shards |
| ⬜ | 09f observability & reshard | per-shard metrics, `WO_RESHARD` |
All numbers + find-and-fix stories: [09-concurrency-scaleout.md](09-concurrency-scaleout.md) shipped notes and the [benchmark table](../../runtime/README.md).
### Track 3 — Storage & durability (plans 10–12, 16) — 🔄 in progress
| Status | Phase | Doc | Notes |
| --- | --- | --- | --- |
| ⬜ | 10 storage foundations | [10](10-storage-foundations.md) | scope reduced: WAL framing/fallocate landed via 09c; `@table(name:, index:)` surface + RAM secondary indexes landed via 13 follow-up |
| ⬜ | 11 WAL & recovery | [11](11-wal-and-recovery.md) | remaining: snapshots (`.data`), compaction, WAL rotation — replay core shipped in 09c |
| ⬜ | 12 engine disk cutover | [12](12-engine-disk-cutover.md) | mmap arena engine (C phase B is the proving ground) |
| ✅ | 16a PG wire client | [16](16-postgres-mirror.md) | hand-rolled protocol v3 (`pg.rs`): trust/password/md5 auth, simple query — zero crates |
| ✅ | 16b PG backup mirror | [16](16-postgres-mirror.md) | `WO_PG=…`: async JSONB upserts behind the WAL ack; boot = full resync; RAM authoritative, reads never touch PG |
| ⬜ | 16c typed columns | [16](16-postgres-mirror.md) | catalog → typed columns; `@table` indexes → `CREATE INDEX` |
| ⬜ | 16d lossless resync | [16](16-postgres-mirror.md) | dirty-flag repair without restart; lag metrics on `/` |
| ⬜ | 16e restore from PG | [16](16-postgres-mirror.md) | `WO_PG_RESTORE=1` boot when the WAL is gone |
| ⬜ | 16f SCRAM auth | [16](16-postgres-mirror.md) | hand-rolled SHA-256/HMAC/PBKDF2 |
### Track 4 — Language & API (`.wo` on the wire) — 🔄 in progress
| Status | Phase | Doc | Notes |
| --- | --- | --- | --- |
| ✅ | 13a class surface | [13](13-class-model-live-pricing.md) | `class` parses, CRUD serves, spec amended |
| ✅ | 13b method execution | [13](13-class-model-live-pricing.md) | `POST /api/<t>/:id/<method>`; row-scoped txn; one `WalRec::Txn` frame per call; abort → 409 rollback |
| ✅ | — `@table` + indexed DML (follow-up) | [13](13-class-model-live-pricing.md) | `@table(name:, index:)`; engine secondary indexes + `find_by`; `select Type{…}` in methods; `?field=value` REST filters |
| ⬜ | 13c LIVE pricing push | [13](13-class-model-live-pricing.md) | Stage 3 scoped: subscription registry, WS at `/api/<t>/live`, replaces the 501 stub |
| ⬜ | Stage 3 wire layer | [../runtime/database/04-client-api.md](../runtime/database/04-client-api.md) | full subscription engine + wire protocol; 13c is its beachhead |
| ⬜ | 13e pricing at scale | [13](13-class-model-live-pricing.md) | wires demo to 09; hot-row reads |
| ⬜ | 15a MCP core + tools | [15](15-mcp-streamable-http.md) | JSON-RPC 2.0 on `POST /mcp`; catalog-generated tools; durable-ack gated; stateless |
| ⬜ | 15b MCP resources | [15](15-mcp-streamable-http.md) | `wo://` URIs, templates, cursors |
| ⬜ | 15c MCP SSE + sessions | [15](15-mcp-streamable-http.md) | first streaming response in `rt`; `Mcp-Session-Id`; GET stream |
| ⬜ | 15d `service mcp` surface | [15](15-mcp-streamable-http.md) | `ServiceKind::Mcp`; 13b methods become tools |
| ⬜ | 15e MCP live subscriptions | [15](15-mcp-streamable-http.md) | `resources/subscribe` + `Last-Event-ID` replay — needs 13c + 09d |
Ecommerce sample status (verified 2026-06-13, `api.rest` **17/17 expected statuses pass**): route wiring, empty lists, 405 for un-exposed ops, 404 for unregistered routes, 501 Stage-3 stubs — all exactly as documented. `fn checkout`, `on startup` seeding, and status-lifecycle triggers await 13b-style execution + 09e.
### Track 5 — C proving ground — ✅ done (A–F)
### Frontend — parked
| Status | Phase | Doc |
| --- | --- | --- |
| ✅ | A threads · B arena · C io_uring · D WAL · E recovery · F bench+ACID | [exploration/c-runtime/00-plan.md](exploration/c-runtime/00-plan.md) |
| ⏸ | 13d pricing UI | [13](13-class-model-live-pricing.md) |
| ⏸ | 14 MVC UI implementation (14a–f) | [14](14-mvc-ui-implementation.md) |
| ⏸ | UI exploration track | [exploration/ui/00-overview.md](exploration/ui/00-overview.md) |
859k reads/s, 618k durable commits/s; found the ack-ordering + fd-ABA bugs the Rust port then avoided. Optional phase G (splice `wo-db` C++ engine) remains an idea.
---
### Track 6 — Frontend — ⏸ parked (out of backend focus)
## Discarded
| Status | Phase | Doc | Notes |
| --- | --- | --- | --- |
| ⏸ | 13d pricing UI | [13](13-class-model-live-pricing.md) | MVC triplet exists as design artifact |
| ⏸ | 14 MVC UI implementation (14a–f) | [14](14-mvc-ui-implementation.md) | htmlx engine, SCSS subset, controllers, SSR, actions, live patching |
| ⏸ | UI exploration track | [exploration/ui/00-overview.md](exploration/ui/00-overview.md) | 00–08 design docs stay current |
Settled rejections with their reasons live in [`discarded.md`](discarded.md) —
inheritance, `abstract` newtypes, `Money`/`SKU`/`Float`, `Dynamic`/`cast`/
`macro`/`extern`, AOT-to-C, Menhir, shared mutable engine state, external
deployer daemon, destructive migrations in v1, and more. Argue against the
recorded reason rather than re-opening an entry as new.
## Suggested order of play (backend)
## Learnings
1. ~~**13b — method execution**~~ ✅ shipped — methods run as row-scoped transactions over RPC
2. **13c / Stage 3 LIVE** with **09d** fan-out (turns every 501 stub real; the ecommerce order-ops board's backend)
3. **09e — 2PC** (cross-shard `fn checkout` — the canonical ACID demo end-to-end; 13b's single-shard txn is its building block)
4. **15a/15b — MCP core, tools, resources** (needs nothing unshipped; makes every app agent-callable — 13b methods become MCP tools in 15d; 15e waits on 13c + 09d)
5. **05/06** dependency removal (mechanical, any time)
6. **10–12** storage completion (snapshots/compaction; arena engine)
What attempts taught, shipped or not, in [`learnings.md`](learnings.md) —
plumbed-is-not-enforced, vacuously-passing goldens, exit-0-with-wrong-output,
the malloc-path ASan trick, deferred checks that never reach the runtime,
validate-once-at-the-boundary, and reference-implement-in-C-first.

View file

@ -77,7 +77,7 @@ docs/plan/oop-vm/00-wob-format.md grows with the three VM changes
**Concept & reason:** the null-safety story. `?T` admits nil; `T` never does — the diagnostic-enforced boundary. Representation: heap kinds use the zero word (the VM's existing null checks already trap on it — optionals make those unreachable by typing); scalar optionals box into a one-field cell (the VM piece — obj.c gains the box; format doc notes the convention). Narrowing: comparing against `null` narrows in the branch (`if x != null` makes `x` a `T` inside — Haxe's exact idiom); using a `?T` un-narrowed where `T` is required diagnoses. Stdlib returns (plan 9) and record `?fields` (Task 4) type as `?T` from here on.
**Next work item:** this task is next up — `?T` is plumbed (lexer/token/AST/parser/dump) but unenforced (E211/E212/E213 dead, probe exits 0 with zero diagnostics per `docs/plan/compiler/nullable-types-implementation.md`), and it blocks the log-watcher port (story iterations 5–6), which uses optionals throughout in place of the Haxe original's sentinel values.
**First task of this plan:** iteration 4 (plan 3 — emitter, corpus, `woc build`) precedes plan 8; within plan 8 this task goes first — `?T` is plumbed (lexer/token/AST/parser/dump) but unenforced (E211/E212/E213 dead, probe exits 0 with zero diagnostics per `docs/plan/compiler/nullable-types-implementation.md`), and it blocks the log-watcher port (story iterations 5–6), which uses optionals throughout in place of the Haxe original's sentinel values.
- [ ] Failing fixtures: narrowing goldens; un-narrowed-use must-fail; nil propagation through record optional fields; boxed scalar optional round-trip; assignment of null to plain `T` must-fail.
- [ ] Implement; green.

54
docs/plan/discarded.md Normal file
View file

@ -0,0 +1,54 @@
# Discarded — settled rejections
Decisions that were considered and **rejected**, with the reason. This file
exists so a settled question is not re-proposed. If you want to revisit an
entry, argue against the reason recorded here — do not re-open it as if it were
new.
Status board: [`00-kanban.md`](00-kanban.md) · Doctrine: [`../00-principles.md`](../00-principles.md)
## Language surface
| Rejected | Date | Reason |
| --- | --- | --- |
| **Inheritance** — `extends`, `super`, `override`, `implements` | plan 13 / OOP spec | No hierarchies, ever. Is-a is a tagged union, has-a is composition, polymorphism is structural interfaces. Hierarchies fossilize early guesses and make dispatch, ownership, and diagnostics all harder. Principle 4. |
| **`abstract` newtypes** (`abstract Money = Int`) | 2026-08-10 | Flipped **adopt → reject** in the systems-track verdict table. A distinct scalar type adds a conversion surface without buying safety this language needs; domain scalars stay plain `Int`/`Text`. The `Money`/`SKU` stopgap allowlist that stood in for the unbuilt feature proved the cost was real. Haxe-parity Task 7 keeps only `is`. |
| **`Money`, `SKU` as language types** | 2026-08-10 | Removed from `builtin_scalars` and from the abstract allowlist. `Money` → `Int` (minor units), `SKU` → `Text` everywhere. |
| **`Float` as a builtin scalar** | 2026-08-10 | Phantom: it was in `builtin_scalars` and documented as f64, but `token.ml` has no float literal and `wob.h` has no float kind — `ratio: Float` typechecked while no `Float` value could ever be written or represented. Re-add only when literals **and** a wob float kind land together. |
| **`Dynamic` / `untyped`** | systems-track spec | Static typing is the untagged-register VM's foundation. Typed `json.decode … as T -> ?T` covers the real use. Principle 13. |
| **`cast`** | systems-track spec | No unsafe casts. Conversions are typed; `as` exists only in the decode-target position. |
| **`macro`** | systems-track spec | Kills the fast-compile promise. Codegen belongs to `wo gen` tooling. |
| **`extern` / FFI** | systems-track spec | One FFI hole voids the whole memory-safety story. Capabilities are audited typed builtins. Principle 10. |
| **`operator` overloading, `overload`** | systems-track spec | One name, one signature. Keeps dispatch and diagnostics simple. |
| **`inline` functions** | systems-track spec | Optimization is the compiler's job. `const` compile-time values are adopted; inline *functions* are not. |
| **User-definable generic containers** | OOP spec §3 | `multi` and `map` are runtime-provided native classes implemented in C. The VM picks the backing data structure; the language exposes only the ADT's operations. |
## Compiler / VM architecture
| Rejected | Reason |
| --- | --- |
| **AOT compilation to C** | Kills hot reload and makes builds slow. A register bytecode interpreter with computed-goto dispatch ships first; a JIT stays possible later. |
| **Approach B — "Lua-shaped minimal" VM** | Defers the project's core risk (the hybrid borrow VM) and adds a Menhir dependency. |
| **Approach C — "Rust-lite static regions"** | Research-grade complexity; recreates exactly the Rust ergonomics pain that mutable value semantics exists to avoid. |
| **Menhir, ppx, any opam package** | OCaml stdlib only; handwritten lexer and recursive-descent parser. dune is a build runner, nothing more. |
| **First-class borrows** (storable, returnable) | Second-class borrows (Hylo/Val mutable value semantics) eliminate full lifetime inference — the compiler needs only per-function flow analysis. Long-lived links go through `ref T` ids or `@gc`. |
| **Empty-list stub for `is_abstract_type`** | 2026-08-10: a function that can never return true is dead code. Deleted outright instead. |
## Runtime / deployment
| Rejected | Reason |
| --- | --- |
| **`Arc<Mutex<Engine>>` / any shared mutable engine state** | Thread-per-core shards each own their engine, heap, and event loop; cross-shard work is an ownership-moving message send. Principle 5. |
| **External deployer daemon** | Violates two-VMs-in-one-runtime and splits the data plane. The management plane is a runtime module with a thin CLI. |
| **Self-hosted `.wo` deploy logic** | Bootstrap problem — the deploy path cannot be written in the language whose deployment it implements. Recorded as a much-later possibility. |
| **Script-based / destructive schema migrations (v1)** | Additive-only auto-diff is what makes rollback unconditional: the previous version ignores fields and classes it never knew. Destructive changes reject at propose time. |
| **Portability abstractions over Linux** | Targeting one kernel lets the runtime use its sharpest primitives directly instead of the lowest common denominator. Principle 9. |
| **Mirrors on the commit path** | The Postgres mirror is a reconstructible backup. RAM is authoritative; reads and acks never depend on it. Principle 7. |
## Process / docs
| Rejected | Reason |
| --- | --- |
| **Per-example `principle.md` files** | 2026-08-08: one canonical repo-level [`docs/00-principles.md`](../00-principles.md) instead; examples link to it. |
| **Minimal 3-file log-watcher sample** | Breaks the file-for-file `.hx` → `.wo` mapping and leaves the "could not express" column unproven — which is the sample's entire acceptance criterion. |
| **Raw code in plan documents** | Plans carry concept, reason, and required behavior in words; the executor writes the code. |

View file

@ -0,0 +1,180 @@
# Blue/Green VMs — a self-hosting, agent-managed runtime
> **Partially superseded (2026-08-03):** the deployment subsystem (§5–§6 here)
> is now specified in
> [`docs/superpowers/specs/2026-08-03-blue-green-vm-design.md`](../../../superpowers/specs/2026-08-03-blue-green-vm-design.md)
> — developer + `wo` CLI as the management client (agent/MCP becomes a later
> wrapper), schema migration folded into the approval step (additive-only
> auto-diff in v1), fixed slots with alternating activity, HTTP+JSON+SSE.
> §1–§4 (transports, recipe box, fibers, source-in-binary) remain current
> thinking feeding plans 3/4/6.
> Thought-process capture (2026-08-02). Not a phase plan yet — the vision that
> shapes how the wovm runtime grows past milestone 1, recorded before the
> details harden. Related: the OOP spec
> ([`../../../superpowers/specs/2026-08-01-oop-compiler-vm-design.md`](../../../superpowers/specs/2026-08-01-oop-compiler-vm-design.md)),
> plan 4 (shard-actor runtime), plan 15 (MCP streamable HTTP), and the
> single-binary trailer already shipped by `woc build`.
## The idea, in five sentences
The writeonce executable is a **systemd service that never stops**. It embeds
its own **source code**, not just its bytecode. An external **Claude agent**
reads and edits that source through a managed channel; an approved change is
compiled **inside the runtime** and loaded into the idle VM slot. The runtime
holds **exactly two VMs — Blue (active) and Green (previous version)** — and
deployment is an atomic switch between them. Rollback is the same switch in
reverse, because the previous version never left memory.
## 1. A runtime is not a port
The runtime core is the VM pair + engine + scheduler — it must run with zero
listeners. Ports are **transports**, attached at boot like modules: an HTTP
listener, a unix socket, an MCP endpoint, stdio. Consequences:
- The same binary serves as web app, CLI batch runner, or agent-managed
service depending on which transports the deployment attaches — one of the
recipes a custom web framework builds from (§2).
- **systemd socket activation** fits exactly: the unit owns the socket
(`LISTEN_FDS`), the runtime accepts on whatever fds it inherits. The
"always running" property (§6) and the "no port of its own" property come
from the same mechanism.
## 2. The runtime is a recipe box for web frameworks
Everything a custom web framework needs in later phases must exist as a
separable runtime capability, not a monolith: transports (§1), fibers (§3),
routing surface (plan 6), the subscription registry (plan 7), the DB engine
(plan 5), and the deploy/rollback machinery (§5). A "framework" in a later
phase is a `.wo` library that composes these recipes — the runtime itself
stays framework-agnostic.
## 3. Fibers (green threads)
Concurrency inside a shard is **cooperative fibers scheduled by the VM**, not
OS threads — the Erlang shape on the wovm substrate:
- A fiber is exactly the execution state `wo_vm` already isolates: a register
window stack + frame stack + a current pc. Making that state per-fiber
instead of per-VM turns the interpreter into a fiber scheduler almost for
free.
- **Preemption by reduction budget**: the dispatch loop decrements a counter
per instruction (or per call/back-edge); at zero, the fiber parks and the
scheduler picks the next runnable one. No signals, no stack switching
tricks, deterministic and debuggable.
- Fibers **park on I/O**: a blocked read hands the fd to the shard's event
loop (`wo-rt.c`'s epoll/io_uring machinery) and the fiber resumes when the
completion arrives. One OS thread per core (plan 4's shard), thousands of
fibers per shard.
- Fits the ownership model: a fiber is an actor mailbox owner; cross-fiber
sends follow the same ownership-move rule as cross-shard sends.
## 4. The binary contains its source
`woc build` already appends the `.wob` image to a copy of `wovm` with an
offset trailer. The trailer grows one more section: **the `.wo` source tree**
(paths + contents, compressed). Why:
- The deployed artifact is self-describing — no "which commit is prod
running?" class of question. `wovm --dump-source` can always reproduce
exactly what is executing.
- The agent workflow (§5) needs a source of truth that travels with the
binary, not a checkout that can drift from it.
- After a deployment, the runtime rewrites its own source section (write to
temp, fsync, rename) so the artifact on disk always matches the Blue VM.
## 5. Agent-managed source — how Claude fits
The runtime exposes a **management transport** (MCP over streamable HTTP —
plan 15's machinery, localhost + bearer token, the log-watcher posture).
Claude Code connects as an MCP client. Tools the runtime serves:
| Tool | What it does |
| --- | --- |
| `source_list` / `source_read` | browse the embedded source tree of the running (Blue) version |
| `source_propose` | submit a changed file set as a **proposal** — staged, never applied |
| `proposal_diff` | render the pending proposal against Blue's source |
| `proposal_check` | run `woc check` on the proposal inside the runtime — diagnostics come back to the agent |
| `proposal_approve` | **human-only gate** (separate credential or out-of-band confirmation) — approval triggers compile + green-slot load |
| `deploy_switch` | atomic Blue↔Green switch after health checks |
| `deploy_rollback` | the same switch back — Green still holds the previous version |
| `deploy_status` | which version is Blue, which is Green, in-flight drain state |
Properties worth pinning now:
- **The agent proposes; a human approves.** `proposal_approve` is not
reachable with the agent's token. Approval is the compile trigger, not the
edit.
- **Every step is WAL-logged** — proposals, diagnostics, approvals, switches,
rollbacks form an audit trail that survives crashes like any other commit.
- **The compiler lives with the runtime** for this loop to work: either
`woc` embedded in the binary (adds OCaml runtime weight) or shipped beside
it in the service directory (lighter; the systemd unit owns both files).
Open question in §8 — start with "beside it".
## 6. Blue/Green VM lifecycle
Exactly **two VM slots** per runtime, never more:
- **Blue** — the active VM: all new requests/fibers dispatch into it.
- **Green** — the previous version, loaded and warm: the instant-rollback
target. After a successful deploy the roles swap; the old Blue becomes the
new Green.
The critical separation: **VMs own code, the engine owns data.** Tables,
WAL, subscriptions, and the arena slabs live in the engine layer beneath both
VMs; a switch swaps which bytecode handles requests, never the data. That is
what makes the switch cheap and rollback safe — no state migration on the
happy path (and schema changes are exactly the hard part, §8).
Deploy sequence:
1. Approved proposal compiles (`woc emit`) — failure ends the deploy,
Blue untouched.
2. New image loads + validates into the idle slot (loader is the same
validation battery as always — a bad image cannot boot).
3. Health gate: entry smoke / conformance subset runs against the idle VM.
4. **Switch at the dispatch boundary**: new work enters the new Blue;
in-flight fibers on the old VM drain to completion (bounded timeout).
5. Old Blue becomes Green (rollback target); the binary's source section is
rewritten to match (§4).
6. `deploy_rollback` at any later point is step 4 in reverse — no compile,
no load, the code is already resident.
## 7. Always running
The executable maps to a **systemd service**: `Restart=always`, socket
activation for the transports (§1), the hardening posture proven in the
log-watcher units (unprivileged user, read-only system, `StateDirectory`
for WAL/data). Deployment never restarts the unit — that is the whole point
of the VM pair. The unit restarting (crash, host reboot) boots Blue from the
binary's current source/bytecode section and reloads Green only when the
next deploy happens.
## 8. Open questions (deliberately unresolved here)
1. **Schema migrations.** Code switches atomically; data does not. A
proposal that changes a class's fields needs a migration story between
Green-shaped and Blue-shaped rows — the wo-seg migration doc's
dual-write thinking applies inside one process. Hardest problem in this
vision; needs its own exploration.
2. **Live subscriptions across a switch.** Do WebSocket subscribers survive
a deploy (registry lives in the engine layer → yes, by design), and what
do they see mid-drain?
3. **`woc` placement** — beside the binary vs embedded (§5).
4. **Fiber preemption granularity** — per-instruction counter vs
call/back-edge only (cheaper, coarser).
5. **Does Green count against the heap budget** (two arenas resident) or
does Green hibernate (bytecode resident, heap lazily rebuilt on
rollback)?
## 9. Where this lands in the plan sequence
- Fibers (§3): extends **plan 4** (shard-actor runtime) — same scheduler
work, one more scheduling unit.
- Transports-not-ports (§1): shapes **plan 6** (HTTP/service layer) — the
listener becomes one attachable transport among several.
- Management MCP (§5): builds on **plan 15**'s streamable-HTTP machinery.
- Source-in-binary (§4): extends plan 3's `woc build` trailer.
- Blue/Green switch (§6) + agent loop (§5): a new phase after those land —
needs spec + plan of its own once this vision stabilizes.

109
docs/plan/learnings.md Normal file
View file

@ -0,0 +1,109 @@
# Learnings — what attempts taught
What the work actually taught, independent of whether it shipped. Recorded so
the same wall is not hit twice. Newest first within each section.
Status board: [`00-kanban.md`](00-kanban.md) · Rejections: [`discarded.md`](discarded.md)
## Testing and verification
**Plumbed is not enforced — and a status table will happily claim otherwise.**
`?T` passed every lexer, parser, and dump test while its semantics did not
exist: `WO-E211`/`E212`/`E213` were declared and never emitted, so
`fn take_it(b: Box) -> Int { return b.v; }` with `v: ?Int` exited **0**. The
plan doc had marked the typechecker ✅. Assert on *behavior* — did the
diagnostic fire, what was the exit code — never on the presence of plumbing.
(2026-08-10, nullable-types audit.)
**A declared error-code constant is not a feature.** Ten `WO-E2xx` constants
existed in `types.ml` with no emission site anywhere, including
`unsatisfied_interface` — so structural interface satisfaction was unenforced
while the module's own doc comment claimed the satisfaction set was produced.
The error catalog now lists emitted and reserved codes separately.
**A golden test can pass vacuously.** The first diagnostic-ordering fixture put
both errors in the same pipeline stage, so the collector's insertion order
already equalled the required output order — the fixture would have passed with
the sort deleted. A fixture must **fail** when its mechanism is removed; prove
that by removing it once. The replacement splits the errors across stages so
insertion order is the reverse of output order.
**Exit-0-with-wrong-output is the worst failure mode, and only absence-testing
catches it.** Two instances in one plan: a class field named `on`, `service`, or
`policy` was silently absorbed by the skip-on-block dispatch (field gone, exit
0, empty stderr), and a dangling backslash at EOF inside a string was swallowed
with no diagnostic. Both were found by review, not by the suite, because no test
asserted that something *should* appear.
**Make memory correctness machine-checked.** Test classes given ~130 fields
exceed the arena's 1024-byte size-class ceiling and take the malloc path, so any
missed free becomes a hard ASan report instead of an invisible slop. Used
throughout the VM's drop, RC, cycle-collector, and trap-unwinding tests.
**A blind `bless` absorbs regressions.** `WOC_BLESS=1` rewrites every
`.expected` in every stage, not just the fixture you were thinking about. Prefer
hand-editing a predictable golden — then a passing test *proves* the prediction
— and if you do bless, diff the changed-file list against what you intended.
## Analysis and design
**"The runtime check will catch it" is false if nothing reaches the runtime.**
A double-`mut` reached through let-bound aliases (`let r = bag.items[i]; let s =
bag.items[k]; swap(r, s)`) escaped the static check *and* produced no residual
site — so the VM's borrow word, which only guards sites the table names, was
never engaged. Fixed by canonicalizing places through borrow bindings before
asking the aliasing question. Lesson: when a hybrid design defers a check to
runtime, verify the deferral actually lands in the table that drives it.
**Removals that look load-bearing may not be.** `owner.ml`'s
`is_abstract_type` call sat in a branch returning `Copy` — and its `else` branch
already returned `Copy` for unknown names, so deleting it was provably
behavior-neutral. Read the fallthrough before assuming a call site matters.
**Conservative joins leak.** Marking a conditionally-moved value `Moved` at an
`if`-join is safe against double-free but drops it from every later drop set, so
the not-moved path leaks — which an ASan gate would have caught only in plan 3.
Normalizing instead (drop at the non-moving branch's end) keeps the table shape
and costs only an earlier death on that path.
**Contract notes must live where the consumer will read them.** Two obligations
for the bytecode emitter — coalesce borrow guards per operand, and the
conditional-move drop rule — were first disclosed only in agent report files
that plan 3 will never open. They now live in `dump.ml`'s format-contract
comments beside the tables they constrain.
## Process and tooling
**Check that a new document is actually tracked.** `docs/plan/oop-vm/` was
swallowed by a blanket `docs/plan/` ignore rule, so the error catalog — a
normative compiler↔VM contract both plan tracks cite — existed only on local
disk and appeared in no diff. `.gitignore` now carves that directory back out.
**A status board that covers one track hides the other.** The kanban tracked
only the four Rust-runtime tracks while the entire OOP track (VM core shipped,
compiler front shipped) was invisible on it — so "what is next?" required
reading code and ledgers. Hence the six-bucket format and the next-plan pointer
at the top of the board.
**Prose reports are not a handoff.** Agent-written reports and SDD ledgers hold
the reasoning, but only files a developer opens by habit — the board, the plan,
the format contract — actually transfer it.
## Runtime, from the C proving ground
**Reference-implement first in C, then port.** The C proving ground (phases A–F)
hit 859k reads/s and 618k durable commits/s, and found the ack-ordering and
fd-ABA bugs the Rust port then avoided entirely. Building the risky thing twice,
cheaply first, was faster than building it once carefully.
**Budget the collector, don't stop the world.** Per-shard heaps plus per-shard
cycle-candidate buffers mean no global pause can even be expressed — the
worst case is a bounded slice of one shard's tick. Deferring all frees until
after the trial-deletion phases removed every dangling-candidate hazard that
incremental freeing had introduced.
**Validate once at the trust boundary, then trust it.** The `.wob` loader
bounds-checks every index and copies everything out into aligned structures, so
the interpreter's hot loop carries no static checks at all. Only the checks that
*cannot* be static — the borrow word, runtime-indexed bounds, map keys — remain,
and those always trap rather than corrupt.

View file

@ -50,7 +50,8 @@ iterations); no commits by agents — drafts go to `.dev/commit.md`.
| 8 | [Shard-actor runtime](08-shard-actor-runtime.md) | thread-per-core shards, per-shard heaps, ownership-move messaging |
| 9 | [Database engine](09-database-engine.md) | class-shaped tables, typed WAL + recovery, `insert`/`select` execute |
| 10 | [HTTP service layer](10-http-service.md) | `service` blocks route to VM methods; REST parity with Stage 2 |
| 11 | [Blue-green deploy](11-blue-green-deploy.md) | two VM slots, in-runtime compile, atomic switch, resident rollback |
| 11 | [Fibers](11-fibers.md) | green threads on the shard scheduler: reduction-budget preemption, park on I/O |
| 12 | [Blue-green deploy](12-blue-green-deploy.md) | two VM slots, in-runtime compile, atomic switch, resident rollback |
Review protocol: the developer reads one iteration, approves or amends;
the next starts only after approval. Each iteration is an unsplittable

View file

@ -1,65 +0,0 @@
# Iteration 11 — blue-green in-runtime deployment
> Format: fiberloom `product/story-iteration-template`. Part of
> [Story — one language, one runtime, one database, one binary](00-story.md).
## Goals
- The story's closing promise: the running binary updates its own code.
Two fixed VM slots (activity alternating); a proposal pipeline —
propose → approve → in-runtime compile → additive schema migration →
load → health → atomic switch — with the previous version staying
resident as the instant rollback target.
- The developer drives it remotely: `wo remote pull / propose / diff /
approve / rollback / status` over a loopback management transport,
every stage streaming live and WAL-audited.
## Acceptance Criteria
- What to achieve?
- **Given** a running fixture app and an additive code+schema change,
- **when** the developer proposes and approves it,
- **then** the deploy completes with zero dropped requests (in-flight
work drains on the old slot), and the binary's embedded source
trailer matches the new active version afterward.
- What to achieve?
- **Given** a failure at any pipeline stage (compile diagnostic,
destructive-change rejection, load failure, health failure, drain
timeout),
- **when** it occurs,
- **then** the active slot keeps serving untouched, the failure is
visible in the SSE stream and the WAL trail, and a destructive
change was rejected at propose time naming the offending
declaration.
- What to achieve?
- **Given** a completed deploy,
- **when** `wo remote rollback` runs,
- **then** the previous version serves again in under one second with
no compile and no data change — additive-only migration guarantees
old code runs correctly against the migrated schema.
- What to achieve?
- **Given** kill -9 during COMPILING / MIGRATING / SWITCHING /
trailer-rewrite,
- **when** the unit restarts,
- **then** it serves one consistent version and the WAL shows whole
migrations only.
## Out Of Scope
- Script-based/destructive migrations, in-runtime editing workspace,
MCP/agent wrapper over the management plane — all recorded follow-ups.
- Fibers (the vision's §3; a scheduler concern, not a deploy concern).
## Info
- Approved spec: `docs/superpowers/specs/2026-08-03-blue-green-vm-design.md`;
its implementation plan is deliberately authored only after iterations
9–10 ship (prerequisites: a catalog to diff, HTTP machinery to build on).
- VMs own code; the engine owns data — the separation that makes the
switch cheap and rollback unconditional.
## Proposed Solution
- Author the implementation plan from the approved spec once iterations
9–10 land, then execute it (slots, deploy state machine, additive
differ, management surface + SSE, `wo remote` verbs, crash battery).

View file

@ -0,0 +1,92 @@
# DB Engine Binding Implementation Plan
> **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (`- [ ]`) syntax for tracking.
>
> **Style rule (user convention):** concept, reason, and required behavior in words only; the executor writes the code.
**Goal:** Sub-project 3 of the OOP spec — objects become rows: every class is a table, `insert`/`select` execute against a per-shard RAM engine with WAL durability and boot replay, and the `DB_STUB` trap disappears from compiled programs.
**Architecture:** Plan 5 of 7. Depends on plans 1–4 (VM, compiler, emit/corpus, shard runtime). The storage substrate is again already proven in C: `runtime/wo-rt.c` phases B/D/E shipped the mmap arena with per-shard slices, framed CRC WAL records with group-commit fsync, torn-tail detection, and parallel boot replay (`docs/plan/exploration/c-runtime/00-plan.md`, exits measured). This plan generalizes those patterns from the demo's fixed row shape to class-shaped rows driven by the `.wob` class table, and wires the language's DB statements through them. The C++ prototype `prototypes/wo-db/` is the semantic reference for the query subset. The Rust engine's index doctrine carries over verbatim: secondary indexes are maintained ONLY through the engine's row-insert/row-remove path — nothing touches table storage directly.
**Tech Stack:** C11 + libc (pwrite/fdatasync or io_uring per the shipped phase-D pattern, fallocate, mmap). Hand-rolled CRC32 (already exists in wo-rt.c to port).
## Global Constraints
- All plan-1/plan-4 constraints carry over (libc only, no commits — drafts to `.dev/commit.md`, ASan/TSan gates, docs under `docs/`).
- **Every class IS a table** — `@table(...)` configures storage (name, indexes), never toggles table-ness (CLAUDE.md doctrine).
- **Per-shard ACID at the c-runtime plan's definition:** atomicity via framed replay-whole-or-not-at-all WAL records; isolation by the one-thread-one-stream execution model; durability = ack only after the covering fsync; group commit per tick.
- **Id discipline:** ids interleave per shard (`t+1, t+1+N, …`); a row's owner shard is `(id-1) % N`; point ops on a foreign row hop once via the plan-4 mailbox — creates are always local.
- **Index doctrine:** secondary indexes update only inside the engine's insert/remove path; direct storage mutation is a defect by definition.
- **RAM is authoritative:** reads never touch a file descriptor (phase-B doctrine); disk exists for durability and boot.
- **Format changes go through the format doc:** DB operations extend the builtin table (ids appended to `docs/plan/oop-vm/00-wob-format.md`); no opcode-space or version change.
---
## File Structure
```
runtime/src/
table.c table.h class-shaped row storage: slabs, slots, id alloc, indexes (Tasks 1, 4)
wal.c wal.h typed-row WAL records, group commit, replay (Task 2)
db.c db.h statement executors: insert/select/update-point (Tasks 3, 5)
compiler/src/ DbStub nodes become typed DB AST + lowering (Task 3)
tests/corpus/db/ DB fixtures incl. crash/replay (Task 6)
docs/plan/oop-vm/04-db-binding.md row format, WAL record layout, query subset (Task 1)
```
---
### Task 1: Class-shaped row storage
**Concept & reason:** generalize phase B. Per shard, per class: a slab of fixed-size row slots sized from the class's field count (16-byte row header — id, class, flags — plus the same 8-byte slots the VM object layout uses, so a row and an object share their field encoding; text and container fields store engine-owned copies, not VM pointers). An allocation bitmap per slab; slab growth by arena extension. Id allocation interleaved per shard for coordination-free global uniqueness (shipped phase-A behavior). Row create/read/remove go through one API that Task 4's indexes hook — the doctrine choke point. The binding doc pins the row format, the field-encoding rules (what happens to each of the six kinds when a value crosses from VM heap to row storage — scalars copy, texts copy, owned objects flatten by value, `@gc` references are a compile error in stored fields already, `ref` is an id, containers copy element-wise), and the query subset promised by Task 5.
- [ ] Failing tests: create/read/remove round-trips across kinds; id interleave across shards; slab growth; removal reuses slots.
- [ ] Implement; ASan green. Write the binding doc.
- [ ] Record commit draft: `feat(runtime): class-shaped row storage — per-shard per-class slabs from .wob class table, VM-compatible field encoding, interleaved id allocation, single choke-point row API; docs/plan/oop-vm/04-db-binding.md.`
### Task 2: Typed WAL + boot replay
**Concept & reason:** port phase D/E to typed rows. Records frame `length | crc | payload | commit-mark` (replay-whole-or-not-at-all); payload = record kind (insert/remove/update), class id, row id, encoded fields. Per-shard WAL files, fallocate-preallocated, appended on commit AFTER the RAM apply, one fdatasync covering the tick's batch, ack after the sync completes — the shipped commit order, verbatim. Boot: per-shard parallel replay before listeners open; torn tails detected and dropped whole (phase E behavior). The offline `wal-check` verification mode ports too — it is the crash test's oracle.
- [ ] Failing tests: commit-then-kill crash battery (the phase-D test shape: concurrent writes, SIGKILL mid-stream, offline verification proves every acked write present and CRC-valid, zero acked-but-missing); torn-tail drop; parallel replay rebuilds identical RAM state (deep-compare against pre-crash snapshot dump).
- [ ] Implement; green.
- [ ] Record commit draft: `feat(runtime): typed-row WAL + replay — framed CRC records over class rows, group-commit ack-after-fsync, parallel boot replay with torn-tail drop, wal-check oracle; crash battery green.`
### Task 3: `insert` executes
**Concept & reason:** the compiler's DbStub node for `insert` becomes a typed AST: target class, field initializer list (defaults applied for omitted fields with defaults — the explicit now() form computes at execution), returning the new id. Typechecking validates fields against the class exactly like constructor literals. Lowering emits DB builtins (ids appended to the format doc's builtin table): the executor allocates the id, encodes fields from registers, applies to RAM through the Task-1 API, stages the WAL record; the VM sees the id as the result. Inserts targeting the local shard complete inline; there is no remote insert — creates are always local by the id discipline. The pricing corpus's `set_price` fixture flips from expecting the DB trap to expecting success — the milestone's most satisfying diff.
- [ ] Failing tests: compiler goldens (typed insert AST, emitted builtins); runtime fixtures (insert then read back through select-by-id once Task 5 lands — interim: through a test hook on the row API); default-value application; the flipped pricing fixture.
- [ ] Implement both halves; green.
- [ ] Record commit draft: `feat: insert executes — DbStub becomes typed insert AST with constructor-grade field checking, DB builtins apply RAM-then-WAL through the row API; pricing set_price fixture flips from trap to green.`
### Task 4: Secondary indexes
**Concept & reason:** `@table(index: [a, b])` becomes real. Per-shard hash indexes (open addressing; text keys by content) from indexed field value to row id, maintained exclusively inside the row API's insert/remove — the doctrine choke point built in Task 1 pays off here. Unique constraints (`@unique` fields) enforce at insert with a constraint trap (new trap code, format doc updated). Index rebuild on boot replay happens through the same path for free.
- [ ] Failing tests: indexed lookup hits; unique violation traps; replay rebuilds indexes (crash battery re-run asserting post-replay index lookups); index maintenance survives remove-then-reinsert.
- [ ] Implement; green.
- [ ] Record commit draft: `feat(runtime): secondary indexes — per-shard hash indexes maintained only inside the row API, @unique constraint trap, replay rebuild; doctrine enforced by construction.`
### Task 5: `select` subset + cross-shard point reads
**Concept & reason:** the milestone query subset, semantics per the C++ `wo-db` reference where they overlap: select-by-id; select with a WHERE conjunction over indexed fields (index-backed) or a full shard scan (explicitly allowed, explicitly slower); dotted-path field access in the projection; results materialize as VM objects (rows decode back through the field-encoding rules — the Task-1 doc's table read in reverse). Local rows resolve inline; a by-id read of a foreign row hops once via the plan-4 mailbox (point ops hop once; the requesting job pumps its inbox while waiting — never blocks). List queries stay shard-local in this plan; scatter-gather fan-out is future work, stated in the doc. `update` limited to point-by-id field sets (the method-transaction pattern the pricing demo uses); no joins, no aggregations beyond the existing builtins, no RETURNING chaining — all named as out-of-scope in the binding doc.
- [ ] Failing tests: by-id local and cross-shard (deterministic two-shard fixture); WHERE over an index vs scan parity (same results both paths); projection decoding across kinds; point update round-trip with WAL coverage.
- [ ] Implement; green.
- [ ] Record commit draft: `feat: select subset — by-id (cross-shard hop-once), indexed/scan WHERE conjunctions, projection decode to VM objects, point update; scatter-gather and joins explicitly deferred.`
### Task 6: DB corpus + acceptance
**Concept & reason:** the corpus grows a `db/` kind wired into the standard runner: insert/select round-trip programs with exact stdout; the crash/replay battery as a scripted scenario (run, kill, reboot, assert identical query results); a unique-violation trap fixture; the C++ `wo-db` smoke overlap — where `prototypes/wo-db`'s `.wo` smoke files exercise semantics this subset implements, run both and compare (manifest-scoped like the plan-3 parity harness). `just oop-accept` gains the DB corpus and the crash scenario. The kanban and CLAUDE.md sync: Stage-3-adjacent language ("insert/select execute in the C runtime") replaces the DB_STUB story.
- [ ] Add fixtures + scenario + manifest; all green under ASan; docs synced.
- [ ] Record commit draft: `test(corpus): db suite — round-trips, crash/replay scenario, unique-violation trap, wo-db overlap manifest; oop-accept gains the DB gate; docs sync.`
---
## Plan self-review notes
- **Spec coverage (sub-project 3):** objects↔tables, DB_STUB retired, per-shard ACID with the shipped WAL discipline, id/owner discipline, index doctrine — all tasked. Explicitly deferred and documented: scatter-gather lists, joins/aggregations, RETURNING chaining, cross-shard transactions, `LIVE` (plan 7), the Postgres mirror (a Rust-runtime feature; revisit after parity).
- **Order rationale:** storage before WAL before statements (each layer is the next one's substrate); indexes after insert exists but before select needs them; corpus last.
- **Consistency check:** field-encoding rules defined once (Task 1 doc) and cited by insert (encode) and select (decode); the row API choke point defined in Task 1 is the only mutation path Tasks 3–5 use.

View file

@ -0,0 +1,92 @@
# HTTP Service Layer Implementation Plan
> **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (`- [ ]`) syntax for tracking.
>
> **Style rule (user convention):** concept, reason, and required behavior in words only; the executor writes the code.
**Goal:** Sub-project 4 of the OOP spec — `service rest` blocks become live routes on the C runtime: generated CRUD over the DB binding, method-RPC endpoints, one trap-to-HTTP mapping, hand-rolled JSON — the blog sample's REST surface served by the new stack.
**Architecture:** Plan 6 of 7. Depends on plans 1–5. The HTTP substrate is shipped C: `runtime/wo-rt.c` phase C's io_uring loop with keep-alive and a per-connection write queue (strace-verified: zero epoll/recv/send syscalls on the request path). This plan ports that pattern into `runtime/src/` and mounts a router on it, per shard (routers are shard-local state, mirroring the Rust runtime's thread-local routers — `HandlerFn` not shared). The compiler stops skipping `service` blocks and compiles them into a route section the runtime loads. JSON is hand-rolled per the existing dependency-removal doc (`docs/plan/05-hand-rolled-json.md`) — the C sibling of what the Rust track already specified.
**Tech Stack:** C11 + libc (io_uring per the shipped phase-C pattern, SO_REUSEPORT per phase A). OCaml (service-block parsing + route emission). No TLS in this plan — stated non-scope, terminate elsewhere.
## Global Constraints
- All plan-1/4/5 constraints carry over (libc only, no commits — drafts to `.dev/commit.md`, ASan/TSan gates, docs under `docs/`).
- **Routers are shard-local** — built per worker at boot from the module's route section; no shared routing state.
- **Reads and acks never leave the shard's serial stream** — handlers run as jobs on the owning shard; cross-shard point ops use the plan-4/5 hop-once machinery.
- **The `.wob` format change (route section) bumps the version to 2** — one coordinated change to the format doc, loader, builder, and emitter, in one task, with the loader still rejecting v1-invalid images identically ("format changes go through the format doc" rule honored by doing it once, visibly).
- **Stage-3 endpoints stay honest:** `/api/<type>/live` returns 501 exactly as `.dev/reference/rest/*.rest` documents — LIVE lands in plan 7; don't "fix" the 501s.
- **JSON is hand-rolled** — no library, bounded depth, UTF-8 safe, byte-length (not char-length) content framing.
---
## File Structure
```
runtime/src/
http.c http.h request parse, keep-alive state machine, response write (Task 1)
json.c json.h hand-rolled encoder/decoder (Task 2)
router.c router.h per-shard route table, path/method match (Task 3)
handlers.c handlers.h CRUD + method-RPC handlers over db.h (Tasks 4–5)
compiler/src/ service-block parsing + route-section emission (Task 3)
tests/corpus/http/ request/response fixtures driven over real sockets (Task 6)
docs/plan/oop-vm/05-http-service.md route section format, trap→HTTP table, JSON subset (Task 1)
```
---
### Task 1: HTTP module port
**Concept & reason:** lift phase C's proven connection machine out of the single file into a module the scheduler mounts per shard: multishot accept, keep-alive request parsing (method, path, headers subset, content-length bodies), per-connection write queue, pipelined-tail carry-over — behaviors phase C already measured; the port's test evidence must match (single connection serving many requests, syscall profile clean). The doc opens with the layer's contract: how handlers receive a parsed request and return status/headers/body, and the trap→HTTP table Task 5 implements.
- [ ] Failing tests: socket-level harness — keep-alive sequence on one connection, bad-request handling (malformed start line → 400 and close), body framing by content-length, oversized request → 413 and close.
- [ ] Implement the port; green; syscall profile spot-checked against the phase-C evidence.
- [ ] Record commit draft: `feat(runtime): http module — phase-C io_uring keep-alive machine as a per-shard module (multishot accept, write queues, pipelined tails); socket-harness tests; docs/plan/oop-vm/05-http-service.md contract.`
### Task 2: Hand-rolled JSON
**Concept & reason:** the codec both directions of every endpoint use, built to the existing plan-05 doc's discipline: encoder streams rows/objects using the six field kinds (scalars, texts with escaping, containers as arrays/objects, ids as numbers); decoder parses request bodies into field initializers with bounded nesting depth, exact UTF-8 validation, and duplicate-key rejection; numbers are i64-safe (no doubles in milestone types). Errors carry positions for 400 responses that name the byte offset. Content-Length is computed from encoded byte length — the log-watcher-documented non-ASCII truncation bug class, prevented by rule here.
- [ ] Failing tests: round-trip across all kinds; escaping/UTF-8 edges (embedded quotes, multibyte, invalid sequences rejected); depth bomb rejected; duplicate keys rejected; byte-length framing with multibyte content.
- [ ] Implement; green.
- [ ] Record commit draft: `feat(runtime): hand-rolled JSON — kind-driven encode, strict decode (UTF-8, depth, dup keys) with byte-offset errors, byte-length framing; edge-case battery.`
### Task 3: Service blocks compile to a route section
**Concept & reason:** the compiler's skip-on-block for `service rest "..." expose list, get, create, update, delete` ends: the parser builds a service AST (path prefix, exposed verbs, owning class), the typechecker validates verbs against the class (subscribe stays legal to declare — it routes to the 501 stub), and the emitter writes a route section into the `.wob` — the coordinated **format v2** change: header gains the section, builder and loader gain support with full validation (paths well-formed, class ids in range, verbs known), version bumps, format doc updated, and every plan-1 loader test re-run to prove v1-shaped rejects still reject. The runtime's router loads the section per shard into a match table (exact-prefix plus `:id` segment).
- [ ] Failing tests: compiler goldens (service AST, route section in the disassembler); loader validation of malformed route sections; router unit tests (match/miss/verb table incl. 405-shaped responses per the `.rest` conventions).
- [ ] Implement across compiler, builder, loader, router; all prior gates re-run green.
- [ ] Record commit draft: `feat: service rest compiles — service AST + .wob v2 route section (coordinated builder/loader/emitter bump, format doc updated), per-shard router with :id matching and 405 semantics.`
### Task 4: Generated CRUD over the DB binding
**Concept & reason:** the six-endpoint contract the README promises, on the new stack: list (shard-local per the plan-5 scope, documented), get by id (hop-once cross-shard), create (JSON body → insert path with defaults, 201 with the row), update (PATCH partial-set → point update), delete (row remove through the choke-point API), each encoding responses through Task 2 and running as a job on the owning shard. Behavior parity target is the Rust runtime's Stage-2 semantics as documented by `.dev/reference/rest/blog.rest` — including auto-id, default seeding, and partial-update PATCH.
- [ ] Failing tests: socket-level CRUD round-trips against a compiled blog-shaped fixture; PATCH partial semantics; 404 on missing ids; create-on-foreign-shard impossible by construction (create is local — test proves ids from the accepting shard).
- [ ] Implement; green.
- [ ] Record commit draft: `feat(runtime): generated CRUD — list/get/create/update/delete over the row API with hop-once foreign reads, Stage-2 parity semantics (auto-id, defaults, PATCH).`
### Task 5: Method RPC + trap→HTTP mapping
**Concept & reason:** the 13b story on the new stack: exposed methods get POST routes (path per the Rust runtime's method-RPC convention), the handler decodes arguments, runs the method on the row's owning shard as a row-scoped job, and encodes the return value. The trap table becomes the error contract, implemented once in the handler layer and documented in the Task-1 doc: BOUNDS/KEY on missing rows or fields → 404/400 as appropriate; BORROW (residual aliasing) → 409; DIV0 and EXPLICIT → 500 with the structured `{code, method, line, message}` body the spec promised in section 6; DB → 501 (anything still unbound); STACK/OOM → 500 with no body detail. The `subscribe` verb's `/live` route returns the honest 501.
- [ ] Failing tests: method round-trip on the pricing fixture (set_price over HTTP mutates, current_price reads back); each trap class mapped (fixtures rig each trap) with the structured error body asserted; /live 501.
- [ ] Implement; green.
- [ ] Record commit draft: `feat(runtime): method RPC + trap contract — exposed methods as shard-local POST jobs, one trap→HTTP table (409 borrow, 404/400 bounds, structured 500 bodies), honest /live 501.`
### Task 6: Blog-sample smoke + acceptance
**Concept & reason:** the end-to-end proof the stack means something: the blog sample's milestone-compatible subset (types + service blocks; policies/triggers still parse-and-discard) compiles with `woc`, serves with the sharded runtime, and a scripted run of the `.dev/reference/rest/blog.rest` request sequence (curl-driven, per that directory's README) passes — expected statuses including the documented 501s and policy-shaped 405/404s where applicable. A light bench recipe (reusing the phase-C bench harness shape) records requests/sec for the record, not as a gate. `just oop-accept` gains the HTTP corpus and the blog smoke; CLAUDE.md/kanban sync.
- [ ] Wire smoke + bench + gate; green; docs synced.
- [ ] Record commit draft: `test: blog-sample smoke on the C stack — scripted blog.rest sequence green (incl. honest 501/405 semantics), bench recipe for the record; oop-accept gains the HTTP gate; docs sync.`
---
## Plan self-review notes
- **Spec coverage (sub-project 4):** service blocks routed, trap surface → HTTP exactly as spec section 6 promised ("one trap surface forever"), CRUD parity with Stage 2, method RPC, shard-local routers — all tasked. Non-scope, stated: TLS, LIVE/WebSocket (plan 7), policies/triggers (still parse-and-discard), scatter-gather list.
- **Order rationale:** transport before codec before routing before handlers; the format-v2 bump isolated in one task with full regression re-run; end-to-end smoke last.
- **Consistency check:** the trap→HTTP table lives in one doc section and one handler-layer implementation; CRUD and RPC both route through it. Route section validated with the same loader rigor as every other section.

View file

@ -0,0 +1,87 @@
# log-watcher .wo Sample Implementation Plan
> **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (`- [ ]`) syntax for tracking.
>
> **Style rule (user convention):** concept, reason, and required behavior in words only; the executor writes the code.
**Goal:** Plan 10 — re-express `~/projects/log-watcher` in `.wo` at `docs/examples/log-watcher/`, file-for-file, ported test fixtures passing — the systems track's acceptance workload (spec criteria 4–5).
**Architecture:** Plan 10 of the roadmap, the proof plan. Depends on plans 8 (language) and 9 (stdlib) complete. No new language or runtime features may land here — a task that cannot express its file has found a defect in plans 8/9 and stops (that feedback loop is this plan's purpose; the repo pattern "samples force the grammar" runs in verification direction now). The Haxe original at `~/projects/log-watcher/src/` is the behavioral reference; its test suite's cases (`test/TestMain.hx`) are the fixture source. The port is behavior-faithful, not line-faithful — `.wo` idioms (optionals over sentinel values, records, switch expressions, RAII handles) where they read better, with the README mapping table recording every deliberate divergence.
**Tech Stack:** `.wo` only, plus corpus fixtures. The MCP scope is the original's hand-rolled subset: stateless Streamable-HTTP, tools-only, no SSE — matching `docs/plan/15-mcp-streamable-http.md`'s neighborhood but implemented in-sample over `net`.
## Global Constraints
- All track constraints carry over (no commits — drafts to `.dev/commit.md`; docs under `docs/`; sample keeps only an orientation README beside code).
- **No new features in this plan** — expressiveness gaps stop the task and report against plans 8/9.
- **Behavior parity is fixture-defined:** every ported fixture states which Haxe test case it mirrors; divergences (improvements included) are README-tabled, never silent.
- **The pure-core discipline is preserved:** tail state machine, cron math, and MCP `handle` stay socket-free and clock-injected, exactly like the original — testability was its best design decision.
- **Acceptance = spec criteria 4 and 5:** empty "could not express" column; live silent-death detection on a real tempfile.
---
## File Structure
```
docs/examples/log-watcher/
README.md mapping table: .wo file ↔ .hx sibling ↔ divergences ↔ could-not-express
main.wo subcommand dispatch, config decode (Task 5)
logtail.wo TailState record + bounded tail poll (Task 1)
watcher.wo quiet-period alert state machine (Task 1)
cron.wo cron.d parse + next-fire (Task 2)
probes.wo flock/pgrep probes as static fns (Task 3)
supervisor.wo tick loop, scheduled/active watches, detections sink (Task 3)
mcp.wo typed records, pure handle(), serve loop (Task 4)
tools.wo the MCP tool implementations over fs (Task 4)
tests/corpus/sample-logwatcher/ ported fixtures per task
```
---
### Task 1: `logtail.wo` + `watcher.wo` — the tail state machine
**Concept & reason:** the heart of the original: `TailState` (offset, inode, last level, last-newline clock) as a typedef record; poll semantics ported exactly — first-sight starts a bounded chunk before EOF, inode change or shrink = rotation restart, burst jumps to tail, torn final line held back, complete lines classified by level prefix (the relaxed timestamp-aware rule the original converged on). `watcher.wo` layers the alert rule: last entry error + quiet period elapsed → alert transition. Both take injected clocks (`now` parameters) — the original's testability discipline. Divergence expected and tabled: `?TailState` and `?stat` optionals replace the `-1`-inode and exists-flag sentinels.
- [ ] Port fixtures from the Haxe suite's tail/watcher groups: first-sight window, rotation by rename, truncate restart, torn-line holdback, level classification incl. timestamped lines, quiet-period alert timing.
- [ ] Write the two files; fixtures green.
- [ ] Record commit draft: `docs(examples): log-watcher port — logtail/watcher (records, bounded read_at polls, rotation by inode, torn-line holdback, injected clocks); tail fixture group green.`
### Task 2: `cron.wo` — cron.d parsing + next-fire
**Concept & reason:** the original's `Cron.hx`: parse `/etc/cron.d`-format entries (five-field schedules, user column, command with `>> logfile` redirection extraction — the zero-config trick that derives what to watch), collapse same-log entries, compute next-fire from a schedule list. Unreadable directory reports as a skipped entry, never a throw (the production fix the original carries — preserved via `?` returns). Switch expressions over field patterns replace the original's if-chains where clearer (tabled divergence).
- [ ] Port fixtures: schedule parsing edges (steps, ranges, lists, weekday names), redirection extraction, same-log collapse, next-fire across day/week boundaries, unreadable-dir skip.
- [ ] Write the file; green.
- [ ] Record commit draft: `docs(examples): log-watcher port — cron.d parse (redirection-derived watch list, unreadable-dir skip as data), next-fire math; cron fixture group green.`
### Task 3: `probes.wo` + `supervisor.wo` — the daemon loop
**Concept & reason:** probes as `static fn`s over `proc.run` — flock's exit-1-means-held with exists-guard, pgrep's exit-0-means-alive, every unknown code falling in the safe direction (the original's comment-documented contract, now in the README table). The supervisor: the single-threaded tick loop verbatim — rescan on interval, pre-fire lock probes with the PROBE_LEAD constant, watch activation/completion, service watches with alert-transition detections appended as JSONL through `fs.append`, all clock-injected. The daemon `run()` wraps tick in the `while !env.stopping() { tick; time.sleep }` idiom — the original's `while(true)` improved by the shutdown flag (tabled).
- [ ] Port fixtures: probe exit-code table; supervisor tick scenarios (activation at fire, skip-locked window, completion pruning, rescan on dir change, detection line shape).
- [ ] Write both files; green.
- [ ] Record commit draft: `docs(examples): log-watcher port — flock/pgrep safe-direction probes, supervisor tick loop (lock-lead probes, JSONL detections, stopping-flag daemon idiom); supervisor fixture group green.`
### Task 4: `mcp.wo` + `tools.wo` — the MCP server
**Concept & reason:** the crown piece: hand-rolled MCP-over-HTTP in `.wo`. Typed request/response records; `handle(req) -> resp` stays a pure function — auth-first Bearer check, method/path/size gates, JSON-RPC envelope (initialize/ping/tools-list/tools-call, notifications answered 202), tool dispatch returning isError results for model-recoverable failures — all decoded/encoded through typed `json` records (the `Dynamic`-free rewrite is the port's most instructive diff). The serve loop: `net.listen` on 127.0.0.1, one request per connection, read with the body cap, write with byte-length framing (the original's UTF-8 lesson holds by construction — lengths are byte lengths in the stdlib). `tools.wo` implements the tool subset that needs only shipped capability: list_logs, tail_log, search_log (bounded windows over `fs.read_at`); the sqlite-backed minilog tools are OUT — tabled as "expressible when the DB track's in-RAM SQL lands", not a could-not-express row (the spec scoped embedded SQL out).
- [ ] Port fixtures from the MCP test group: envelope cases (auth 401, wrong method 405, oversized 413, parse error -32700, unknown method -32601, notification 202), tool-call round-trips, socket-level smoke (scripted client, one connection).
- [ ] Write both files; green.
- [ ] Record commit draft: `docs(examples): log-watcher port — MCP subset in .wo (pure handle() over typed json records, Bearer auth, tools list/tail/search over fs), net serve loop; envelope + socket fixtures green.`
### Task 5: `main.wo` + README + acceptance
**Concept & reason:** close the loop. `main.wo`: the three subcommands — `watch` (single-watcher poll loop), `run` (supervisor + optional config), `mcp` (config + env-fallback API key, required-field errors exit 1 with usage) — config decoded via `json.decode as` into the config record, usage text on anything else. The README mapping table: every `.wo` file, its `.hx` sibling, tabled divergences, and the **could-not-express column — acceptance demands it empty** (criterion 4). The live test (criterion 5): a scripted scenario starts the built sample in watch mode against a tempfile, feeds timestamped lines ending in an error, waits past the quiet period, asserts exactly one detection — the original's measured behavior, reproduced. `just oop-accept` gains the sample build + fixture groups + the live scenario; kanban and the systems spec get their shipped-status notes.
- [ ] Write main.wo + README table; port config-loading fixtures (defaults, partial config, missing-key mcp errors).
- [ ] Wire the live scenario + gate; run acceptance: criteria 4 and 5 checked against the spec.
- [ ] Record commit draft: `docs(examples): log-watcher port complete — main.wo subcommands + typed config, README mapping table (could-not-express: empty), live silent-death scenario in oop-accept; systems-track criteria 4-5 checked.`
---
## Plan self-review notes
- **Spec coverage (Part 4, criteria 4–5):** all five `.hx→.wo` mappings from the spec's table have tasks; the pure-core discipline, the README table, and both acceptance criteria are explicit task outputs. The minilog/sqlite tools exclusion matches the spec's out-of-scope list and is recorded as scoped-out, not inexpressible.
- **Feedback-loop honesty:** the no-new-features constraint plus stop-on-gap rule makes this plan the verification instrument for plans 8/9 — its failure mode is a defect report, not a workaround.
- **Order rationale:** pure cores first (tail, cron) — testable without any daemon; probes/supervisor next (compose them); MCP after json/net are proven by earlier tasks; main last, wiring everything.

View file

@ -0,0 +1,100 @@
# Program Mode + Systems Stdlib Implementation Plan
> **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (`- [ ]`) syntax for tracking.
>
> **Style rule (user convention):** concept, reason, and required behavior in words only; the executor writes the code.
**Goal:** Plan 9 — `.wo` becomes a systems language: `fn main` programs with exit codes, and the five capability modules (`env`, `fs`, `proc`, `net`, `time`, plus typed `json`) as safe wovm builtins with RAII handles.
**Architecture:** Plan 9 of the roadmap. Depends on plans 1–3 (toolchain) and plan 8 Task 1 (module resolver knows the stdlib namespaces) and Task 6 (`?T` — most stdlib returns are optional-typed). Runtime work is C builtin families in new `runtime/src/` modules; compiler work is thin (main detection, namespace binding). The spec's Part 2/3 tables are normative. Blocking discipline: program mode runs one shard where blocking builtins are legal; the same surface loop-integrates on server shards later (the shard-actor and HTTP plans own that side — this plan implements program mode only and keeps the builtin layer's seam clean for the other discipline).
**Tech Stack:** C11 + libc (openat/fstat/pread, fork/execvp/waitpid + pipes, socket/bind/listen/accept, clock_gettime/nanosleep, sigprocmask+signalfd), OCaml (driver + typing).
## Global Constraints
- All OOP-track constraints carry over (libc only, no commits — drafts to `.dev/commit.md`, ASan gate, docs under `docs/`).
- **Handles are owned objects; drop closes** — file/socket/process resources ride the MVS destructor path; an fd leak is a test failure, not a review comment.
- **Expected absence is `?T`, faults are traps** (spec error-handling rule): missing file → nil, permission denied on an existing path → trap; the split is documented per function in the module doc.
- **`proc` takes args arrays only** — no shell-string form exists; injection unrepresentable.
- **`fs` has no write/truncate/delete in v1** — `append` is the only mutation (read-only posture default).
- **Builtin ids extend the format doc's table** in each module's task; no opcode changes.
- **Program mode = one shard, blocking legal; server shards unchanged** — nothing in this plan touches the server path.
---
## File Structure
```
runtime/src/
sys_env.c sys_fs.c sys_proc.c sys_net.c sys_time.c one C module per namespace (Tasks 2-6)
builtin.c dispatch table grows per module
compiler/src/ main detection, stdlib namespace typing (Task 1)
compiler/bin/main.ml program-mode driver behavior (Task 1)
tests/corpus/sys/ stdlib fixtures against real tempdirs/processes/sockets
docs/plan/oop-vm/07-systems-stdlib.md per-function contracts: types, nil-vs-trap, bounds (Task 1)
```
---
### Task 1: Program mode — `fn main`, exit codes, `env` module
**Concept & reason:** the mode switch. A project with a free `fn main(args: multi Text) -> Int` compiles as a program: `wo run` (the C stack's runner) executes main on one shard and exits with its return; `woc build` packages it single-binary (plan-3 trailer, unchanged). Service blocks without main = server (existing); both = main runs and decides (spec's log-watcher shape). The `env` module lands here because main is unusable without it: `env.args() -> multi Text`, `env.get(name) -> ?Text`, `env.exit(code)` (typed as never-returning), `env.stopping() -> Bool` — the runtime blocks SIGTERM/SIGINT, owns a signalfd, and flips the flag; programs poll it (no callbacks, spec rule). `print_err` joins the print family; stdout/stderr flush on newline. The module doc opens with the nil-vs-trap contract table every later task extends.
- [ ] Failing fixtures: main runs with args and its return becomes the exit code; env.get present/absent; a stopping-flag fixture (send SIGTERM to the child, assert clean flagged exit); print_err lands on stderr.
- [ ] Implement driver + sys_env + builtins; green.
- [ ] Record commit draft: `feat: program mode — fn main entry with exit codes on the C stack, env module (args/get/exit/stopping via signalfd flag), print_err, flush-on-newline; docs/plan/oop-vm/07-systems-stdlib.md contract table.`
### Task 2: `time` module
**Concept & reason:** smallest module, unblocks every poll-loop fixture after it. `time.now()` exists (wall ms); add `time.mono() -> Int` (monotonic ms, CLOCK_MONOTONIC — interval math must not jump with wall-clock changes) and `time.sleep(ms)` (nanosleep; in program mode it blocks the shard, which is the point; EINTR from the shutdown signal returns early — the daemon loop's exit path).
- [ ] Failing fixtures: mono monotonicity across a sleep; sleep duration lower-bound; sleep cut short by SIGTERM with stopping() true after.
- [ ] Implement; green.
- [ ] Record commit draft: `feat(runtime): time module — mono (CLOCK_MONOTONIC ms), sleep (EINTR-aware, shutdown cuts it short); poll-loop idiom complete.`
### Task 3: `fs` module
**Concept & reason:** the log-watcher workhorse. Per the spec table: `exists`, `stat -> ?{size, inode, mtime}` (a record from plan 8 Task 4 — inode present because rotation detection depends on it), `read_at(path, offset, max) -> Text` (open/pread/close inside the builtin — bounded tail-chunk reads; no persistent file handles in v1 since every LogTail poll is open-read-close by design), `read_all(path, cap)`, `append(path, text)` (O_APPEND open-write-close — the JSONL sink), `list(dir) -> multi Text`. Nil-vs-trap per the contract: missing path → nil/false variants; EACCES on an existing path, or reads past cap → traps. All texts are bytes-faithful (logs contain ANSI junk; the language's Text carries it — sanitizing is the program's job, as log-watcher itself learned in production).
- [ ] Failing fixtures (real tempdir): stat fields incl. inode change across rename-rotation; read_at windows (offset, max, EOF clamp); append accumulates JSONL lines; list ordering defined (sorted — determinism for tests); missing-path nils; permission trap (chmod 000 fixture, skipped when running as root).
- [ ] Implement; green under ASan.
- [ ] Record commit draft: `feat(runtime): fs module — stat with inode/mtime, bounded read_at/read_all, O_APPEND append, sorted list; nil-vs-trap per contract; rotation fixture via rename.`
### Task 4: `proc` module
**Concept & reason:** the probe capability. `proc.run(cmd, args: multi Text) -> {code: Int, out: Text, err: Text}`: fork/execvp with an args array (PATH search yes, shell never), both pipes captured with a per-stream byte cap (truncation flagged in the record — a fourth field, `truncated: Bool`, small spec addition recorded in the module doc), waitpid for the exit code; signal-death reported as conventional 128+signal. Exec failure (missing binary) is expected absence territory: code 127 in the record, not a trap — matching the Haxe original's "unknown → safe direction" probes. Blocking by definition; program mode only (a server-shard call diagnoses at compile time until the loop-integrated discipline lands in its own track).
- [ ] Failing fixtures: true/false exit codes; output capture both streams; cap truncation flag; missing binary → 127; args with spaces pass verbatim (no shell proof); zombie-free after many runs (fixture loops 100 spawns, asserts no defunct children via /proc scan).
- [ ] Implement; green.
- [ ] Record commit draft: `feat(runtime): proc module — fork/execvp args-only run with capped dual capture + truncated flag, waitpid codes (128+sig, 127 missing), zombie-free; server-shard use diagnosed.`
### Task 5: `net` module
**Concept & reason:** the serve capability, and the RAII showcase. `net.listen(addr, port) -> Listener` (SO_REUSEADDR, backlog sane), `net.accept(listener) -> ?Conn` (blocking; nil when interrupted by shutdown — the accept-loop exit path), `net.read(conn, max) -> ?Text` (nil on peer close), `net.write(conn, text)`. Listener and Conn are owned handle objects — class-table natives whose drop closes the fd (the plan-1 native-sentinel pattern; format doc grows the two handle classes). One request per connection is the supported v1 shape (the MCP pattern); keep-alive service belongs to the server stack, not here.
- [ ] Failing fixtures (loopback): listen/accept/read/write echo round-trip driven by a scripted client; peer-close nil; shutdown interrupts accept with nil + stopping(); **fd RAII battery** — accept and drop many conns in a loop, assert stable fd count via /proc/self/fd (the spec's leak test, ASan-adjacent but fd-specific).
- [ ] Implement; green.
- [ ] Record commit draft: `feat(runtime): net module — TCP listen/accept/read/write with owned Listener/Conn handles (drop closes), shutdown-aware accept, peer-close nil; fd-count RAII battery.`
### Task 6: `json` module
**Concept & reason:** typed decode, the `Dynamic` replacement. `json.decode(text) as RecordType -> ?RecordType`: the emitter passes the target class/record id alongside the builtin call; the C side walks the plan-6 codec's parse events against the class table's field kinds — matching fields fill, `?fields` absent stay nil, unknown JSON keys skip, any shape mismatch (wrong type, missing required field) yields nil overall (never a trap — config errors are expected absence). Nested records, `multi` of scalars/records, and `map<Text, scalar>` decode; anything else in the target type diagnoses at compile time as undecodable. `json.encode(value) -> Text` walks kinds in reverse (the plan-6 encoder generalized). If plan 6 has not landed when this executes, the codec lands here and plan 6 consumes it — the doc notes the either-order seam.
- [ ] Failing fixtures: config-shaped decode (the SupConfig pattern: numbers, optional string, array of strings); missing-required nil; unknown-keys-skip golden; nested + multi + map decode; undecodable-target must-fail; encode round-trip.
- [ ] Implement; green.
- [ ] Record commit draft: `feat: typed json — decode-as against class-table kinds (?fields nil, unknown keys skip, mismatch = nil never trap), encode by kinds; undecodable targets diagnosed at compile time.`
### Task 7: Corpus battery + acceptance
**Concept & reason:** criteria 2 and 3 of the spec, gated. The `sys/` corpus runs everything above end to end plus the cross-module fixtures that mimic log-watcher's composites: a tail-poll fixture (write to a tempfile between polls, assert offset math via read_at), a probe fixture (flock-style: proc.run against a held lock file — using the real flock binary when present, skipped cleanly otherwise), a mini serve-loop fixture (accept one request, respond, exit on stopping). `just oop-accept` gains the sys corpus; the module doc's contract table gets a shipped-status column; CLAUDE.md commands note program mode.
- [ ] Wire fixtures + gate; green; docs synced.
- [ ] Record commit draft: `test(corpus): sys battery — tail-poll offset math, real-flock probe (skip-aware), mini serve loop with shutdown; oop-accept gains the sys gate; docs synced.`
---
## Plan self-review notes
- **Spec coverage (Parts 2–3, criteria 2–3):** program mode T1, all five modules T2–T6 matching the spec tables exactly (one addition: `proc` result's `truncated` flag, recorded in the module doc), RAII proofs in T5 (fd battery) + ASan everywhere, corpus gate T7.
- **Dependency honesty:** needs plan 8's modules (`use`), records, and `?T`; json (T6) shares the plan-6 codec with an either-order seam noted.
- **Order rationale:** env/main first (nothing testable without an entry point), time second (fixtures need sleep/mono), fs/proc/net by increasing machinery, json last (needs records + codec), battery at the end.

View file

@ -0,0 +1,100 @@
# Shard-Actor VM Runtime Implementation Plan
> **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (`- [ ]`) syntax for tracking.
>
> **Style rule (user convention):** concept, reason, and required behavior in words only; the executor writes the code.
**Goal:** Sub-project 2 of the OOP spec — multithread the VM: one pinned worker per core, one VM heap per shard, `spawn`/`send` in the language with ownership-transfer semantics, and GC pacing on the event-loop tick — no locks on the data path, no stop-the-world, ever.
**Architecture:** Plan 4 of 7. Depends on plans 1–3 (wovm, woc, emit/corpus). The concurrency substrate already exists and is *shipped*: `runtime/wo-rt.c` phases A–F (thread-per-core pthreads pinned via sched_setaffinity, SO_REUSEPORT accept spreading, per-thread io_uring loops, one mmap arena sharded by address, per-shard WAL + boot replay — see `docs/plan/exploration/c-runtime/00-plan.md`, all exit criteria met). This plan does NOT rewrite that; it ports the patterns into `runtime/src/` modules and mounts one `wo_vm` per shard on top. `wo-rt.c` stays the untouched reference.
**Tech Stack:** C11 + libc, raw syscalls (pthreads, sched_setaffinity, eventfd, epoll now / io_uring when the loop module ports phase C). C11 atomics allowed ONLY in the mailbox ring — the data path stays lock-free and atomic-free per doctrine.
## Global Constraints
- All plan-1 constraints carry over (libc only, no commits — drafts to `.dev/commit.md`, ASan gate, docs under `docs/`).
- **Thread-per-core, shared-nothing** (plan 09 doctrine, c-runtime decision 1): no work stealing, no connection migration, no locks/atomics on the data path.
- **One heap per shard:** every `wo_vm` owns its arena; an object's `shard_id` header field (reserved since plan 1) names its owner. Cross-shard access is always a message, never a pointer dereference.
- **Send = ownership move, not copy** (spec concurrency decision): transferred objects hand over the pointer; memory returns to its origin shard's arena for freeing.
- **Jobs never block; waiters pump their own inbox** (deadlock-freedom doctrine from `crates/rt` shard bus).
- **GC is per-shard and budgeted per tick** — replacing plan 3's post-exit pump; no cross-shard tracing exists by construction.
- **ThreadSanitizer joins the gate:** the mailbox ring and shutdown path must be TSan-clean in addition to ASan.
---
## File Structure
```
runtime/src/
sched.c sched.h worker spawn/pin/join, signalfd shutdown broadcast (Task 1)
mailbox.c mailbox.h MPSC ring + mail eventfd per shard (Task 3)
shard.c shard.h per-shard state: vm, loop, mailbox, gc pacing (Tasks 2, 5)
compiler/src/ grammar/typing/owner additions for spawn/send (Task 6)
tests/corpus/actor/ concurrency fixtures (Task 7)
docs/plan/oop-vm/03-shard-actor.md the runtime contract this plan pins (Task 1)
```
---
### Task 1: Scheduler skeleton + contract doc
**Concept & reason:** port phase A's shipped pattern into modules: `WO_THREADS` workers (default = online cores), each pinned, each owning an event loop and — new — one `wo_vm` instance with its own heap. Signals blocked before spawn; worker 0 owns the signalfd and broadcasts shutdown via per-worker eventfds; clean join. The contract doc records what every later task obeys: shard ownership rules, mailbox semantics, send-as-move, free-routing, gc pacing — the C sibling of plan 09's doctrine section.
- [ ] Failing test: a harness boots N workers, proves pinning (`/proc/self/task/*/stat` cpu affinity), proves per-worker VM isolation (each runs a trivial method concurrently, results independent), proves clean SIGTERM join.
- [ ] Implement; TSan + ASan clean. Write the contract doc.
- [ ] Record commit draft: `feat(runtime): shard scheduler — pinned workers each owning a wo_vm, signalfd shutdown broadcast (phase-A port into src/); docs/plan/oop-vm/03-shard-actor.md contract.`
### Task 2: Per-shard heaps + ownership stamping
**Concept & reason:** every allocation stamps the allocating shard's id into the header (plan 1 reserved the field, always 0 until now). Debug builds assert on every object access that the toucher is the owner shard — turning silent cross-shard races into loud failures during development; release builds compile the assert out. The arena stays per-VM as plan 1 built it; what this task adds is identity and enforcement.
- [ ] Failing test: allocation on shard *t* stamps *t*; a rigged cross-shard touch aborts under the debug assert; release build tolerates the same touch (documented as UB caught only in debug).
- [ ] Implement; green.
- [ ] Record commit draft: `feat(runtime): shard ownership — allocations stamp shard_id, debug-build access asserts turn cross-shard races into aborts.`
### Task 3: Mailboxes
**Concept & reason:** the only cross-thread structure in the system, so it carries the whole correctness burden. Per-shard MPSC ring (fixed capacity, C11 atomics for the producer side, single consumer = the owner shard) plus a mail eventfd the owner's event loop watches — the same wake pattern the Rust shard bus and phase A's shutdown already use. Messages are small fixed structs: kind, sender shard, payload words. Full ring applies backpressure by failing the send — the sender re-queues or traps — never by blocking (doctrine). Two message kinds exist from day one: USER (a sent value) and FREE (route a transferred object's memory home, Task 4).
- [ ] Failing tests: single-producer and multi-producer ordering; full-ring send fails without blocking; eventfd wake fires exactly when an empty ring becomes non-empty; TSan-clean under a hammering stress test.
- [ ] Implement; green under ASan + TSan.
- [ ] Record commit draft: `feat(runtime): per-shard MPSC mailbox ring + mail eventfd — non-blocking sends with backpressure, USER/FREE message kinds; TSan-hammered.`
### Task 4: Send as ownership move + free routing
**Concept & reason:** the spec's "no copy, no locks" promise made mechanical. Sending an owned object: the sender's register is dead after the send (compiler enforces — Task 6), the pointer travels in the message, the receiver restamps `shard_id` to itself. The memory, though, still belongs to the origin shard's arena — freeing it on the wrong thread would corrupt the free lists. So `wo_drop_any` grows a routing check: an object whose *allocation home* (recorded beside the arena, not the mutable `shard_id`) differs from the current shard is not freed in place — a FREE message carries it home, and the origin shard frees it on its next tick. `@gc` objects do not transfer in milestone one: sending one is a compile error (Task 6) — cross-shard reference counting is a distributed-GC problem the spec's model deliberately avoids.
- [ ] Failing tests: transfer round-trip (shard A builds an owned graph, sends to B, B reads/mutates/drops it, memory returns to A's arena — ASan-proven across a full drain-and-join); FREE routing under load; deep graphs transfer whole (children move with the root).
- [ ] Implement; green.
- [ ] Record commit draft: `feat(runtime): ownership-transfer send — pointer handoff with shard restamp, allocation-home free routing via FREE messages, @gc transfer excluded; ASan-proven round-trips.`
### Task 5: GC pacing on the loop tick
**Concept & reason:** the spec's no-stop-the-world claim becomes an event-loop property. Each shard's loop calls the budgeted collector step (plan 1's `wo_gc_step` semantics) once per tick when the candidate buffer is non-empty, budget from configuration, between draining I/O and draining the mailbox — so collection interleaves with work and its cost per tick is bounded by construction. Plan 3's post-exit pump in the CLI stays for single-shot `wovm file.wob` runs; the scheduler path supersedes it. Stats surface per shard: candidates, freed, steps, max step time.
- [ ] Failing test: a shard under continuous method traffic that churns `@gc` cycles keeps collecting (buffer drains over ticks) while request latencies stay bounded — the test asserts the buffer does not grow monotonically and no single tick exceeds the budgeted slice by more than the documented component overshoot.
- [ ] Implement; green.
- [ ] Record commit draft: `feat(runtime): gc pacing — budgeted collector step per shard loop tick between io and mailbox drains, per-shard stats; latency-bounded under cycle churn.`
### Task 6: Language surface — `spawn` and `send`
**Concept & reason:** the smallest surface that exposes the model, mirroring the spec's deferral note. `spawn free_fn(args)` schedules a free fn onto a chosen-by-runtime shard (round-robin), arguments move; `send(handle, value)` moves a value to the shard owning `handle`; `receive` binds delivered values in a free fn designated as an actor body. Grammar keeps the repo's keyword discipline (these parse positionally where possible). The typechecker forbids sending `@gc` values and borrow-mode parameters across; the owner pass treats send/spawn argument positions exactly like `take` (move sites, two-site errors on reuse). The emitter lowers to SEND/SPAWN builtins (builtin table extension — the format doc gains the ids; no opcode change, no version bump).
- [ ] Failing tests: golden parses; must-fail fixtures (use-after-send, sending @gc, sending a borrow); emit goldens for the new builtins; format doc updated in the same change.
- [ ] Implement compiler + wovm builtin halves; green.
- [ ] Record commit draft: `feat: spawn/send surface — parser/typing/owner treat send args as moves (@gc and borrows rejected), SEND/SPAWN builtins in wovm + format doc; ownership corpus grows the send suite.`
### Task 7: Concurrency corpus + acceptance
**Concept & reason:** determinism first: fixtures pin `WO_THREADS`, and assertions observe *joined* results (a spawner collects replies and prints a sorted summary), never interleaving-dependent output. Suite: ping-pong transfer between two shards; fan-out/fan-in over all shards; transfer-heavy stress under ASan and TSan; a mailbox-backpressure fixture proving non-blocking behavior surfaces as a trap, not a hang. Acceptance extends `just oop-accept` with the actor corpus and the TSan gate; the contract doc gets the measured evidence appended, in the style the c-runtime plan's phases record their exits.
- [ ] Add fixtures + gate wiring; all green; evidence recorded.
- [ ] Record commit draft: `test(corpus): actor suite — deterministic ping-pong/fan-in/stress under ASan+TSan, backpressure-as-trap fixture; oop-accept gains the concurrency gate.`
---
## Plan self-review notes
- **Spec coverage (sub-project 2):** shard-actor with per-shard heaps, ownership-transfer sends, no-GC-pause-by-construction, header shard id activated, deadlock-freedom doctrine — all tasked. Deliberately excluded, stated where: cross-shard `@gc` (rejected at compile time), cross-shard transactions (Rust plan 09e territory), io_uring port of the VM loop (the phase-C pattern exists; this plan runs on the epoll loop and the port is a later mechanical swap).
- **Order rationale:** scheduler before heaps before mailboxes before send (each is the next one's substrate); language surface only after the runtime halves exist; corpus last.
- **Risk called out:** free-routing (Task 4) is the subtle piece — its tests are the ones ASan must own end to end.

View file

@ -0,0 +1,92 @@
# UI — .htmlx SSR + LIVE Subscriptions Implementation Plan
> **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (`- [ ]`) syntax for tracking.
>
> **Style rule (user convention):** concept, reason, and required behavior in words only; the executor writes the code.
**Goal:** Sub-project 5 of the OOP spec — the single binary serves a user interface: `##ui` screens compile to `.htmlx` templates rendered server-side, LIVE subscriptions push delta frames over WebSocket on commit, and the pricing demo's driving workload runs end to end — a price update on the server patches subscribed browsers' cells in place.
**Architecture:** Plan 7 of 7. Depends on plans 1–6. Direction is already locked by the UI track (`docs/plan/exploration/ui/00-overview.md`): **`.htmlx` is the template format; `##ui` is the DSL that emits it** — the v1 engine at `.dev/reference/crates/wo-htmlx/` (bindings, each-blocks, partials, data-bind attributes) is the semantic reference the C renderer ports. The subscription machinery is the Stage-3 story (`docs/runtime/database/04-client-api.md`, plan 13c): a per-shard registry hooked into the commit path emits delta frames to WebSocket subscribers — replacing the honest 501 that plan 6 preserved. Static assets (the client runtime JS) serve per the sendfile doctrine (`docs/plan/08-sendfile-static-assets.md`). The UI track's larger workspace story (apps/, per-app binaries, shared DB daemon — ui docs 04–06) is explicitly OUT of this plan: one binary, its own screens, first.
**Tech Stack:** C11 + libc (WebSocket framing hand-rolled; SHA-1 + base64 for the upgrade handshake hand-rolled — the only crypto in the runtime, ~150 lines, documented). OCaml (##ui parsing, .htmlx emission). Vanilla JS client runtime (~20 KB target per the UI track), no build tooling.
## Global Constraints
- All prior plans' constraints carry over (libc only, no commits — drafts to `.dev/commit.md`, gates, docs under `docs/`).
- **`.htmlx` semantics follow the v1 engine** where features overlap (`{{path}}` bindings, `{{#each}}`, `{{> partial}}`, `data-bind`); the live-subtree extension (`<wo:live source="...">`) is this plan's addition, specified in the format doc it writes.
- **Deltas ride commits:** subscription taps sit AFTER the WAL accepts, never in the read or ack path — a slow subscriber can never delay a commit (the Postgres-mirror tap discipline, applied to WebSockets: overflow drops the subscriber loudly, never blocks the shard).
- **Subscriptions are shard-local:** a socket subscribes on the shard that owns its connection; queries against rows on that shard push directly. Cross-shard live queries are out of scope, stated in the doc.
- **No JS frameworks, no bundlers** — the client runtime is one hand-written file served as a static asset.
---
## File Structure
```
runtime/src/
ws.c ws.h WebSocket upgrade (SHA-1/base64), frame codec, ping/pong (Task 1)
sub.c sub.h per-shard subscription registry + commit tap + delta encode (Task 2)
htmlx.c htmlx.h template parse + SSR render (Task 4)
assets.c assets.h static asset serving, sendfile path (Task 5)
compiler/src/ ##ui parsing + .htmlx emission (Task 3)
client/wo-live.js the DOM-patching client runtime (Task 5)
tests/corpus/ui/ render goldens + live end-to-end scripts (Task 6)
docs/plan/oop-vm/06-ui-live.md .htmlx subset, wo:live semantics, delta frame format (Task 1)
```
---
### Task 1: WebSocket transport
**Concept & reason:** the push channel, hand-rolled to the zero-dep bar: HTTP upgrade handshake (the fixed-GUID SHA-1/base64 accept key — the two primitives implemented locally with test vectors from their RFCs), frame codec (text frames, masking rules, fragmentation tolerated on receive, close handshake, ping/pong), mounted on plan 6's connection machine so a socket upgrades in place and joins the shard's loop. The doc this task writes pins everything downstream: the delta frame JSON shape (kind: insert/update/delete, class, id, changed fields), the subscribe message a client sends, and the `wo:live` template semantics Task 3–5 implement.
- [ ] Failing tests: RFC test vectors for the accept key; frame codec round-trips incl. masked payloads and close; a socket-level upgrade-then-echo harness on the real loop.
- [ ] Implement; green.
- [ ] Record commit draft: `feat(runtime): hand-rolled WebSocket — upgrade handshake (local SHA-1/base64 with RFC vectors), frame codec, ping/pong/close, mounted on the shard loop; docs/plan/oop-vm/06-ui-live.md pins delta/subscribe/wo:live formats.`
### Task 2: Subscription registry + commit taps
**Concept & reason:** Stage 3's engine, per shard. A registry maps subscription keys (class + optional indexed-field filter, the plan-5 WHERE subset) to subscriber lists (socket + subscription id). The commit path gains a tap after the WAL accept: each committed mutation consults the registry, encodes one delta frame, and enqueues it on matching subscribers' write queues with try-semantics — a full queue drops that subscriber with a loud close frame and a log line (the mirror discipline: RAM-side progress never waits on a consumer). The `/api/<type>/live` route flips from 501 to the upgrade + subscribe flow; unsubscribe and disconnect clean the registry.
- [ ] Failing tests: registry match/miss across filters; commit-to-frame flow on a rigged shard (insert/update/delete each produce the right frame); overflow drops the subscriber and only the subscriber; disconnect cleanup; the 501 fixture from plan 6 flips to expecting an upgrade.
- [ ] Implement; green under ASan/TSan (sockets and registry are shard-local — the tests prove no cross-thread traffic exists).
- [ ] Record commit draft: `feat(runtime): LIVE subscriptions — per-shard registry with filter matching, post-WAL commit taps with try-enqueue drop-loudly discipline, /live flips from 501 to upgrade+subscribe; delta frames per the pinned format.`
### Task 3: `##ui` compiles to `.htmlx`
**Concept & reason:** the compiler's side of the locked UI decision. The parser accepts the `##ui` block form the samples use (screen name, source class, projection/columns, filters) and the typechecker validates it against the class (fields exist, filter fields indexed where the live path requires it). Emission writes an `.htmlx` file per screen into the build output beside the `.wob`: bindings for projected fields, an each-block over the source rows, and the screen's live subtree wrapped in `<wo:live source="...">` carrying the subscription key the runtime will register. Hand-written `.htmlx` files in the project pass through untouched (first-class authoring alternative, per the locked decision) — the compiler only validates their `wo:live` sources against the schema.
- [ ] Failing tests: golden `.htmlx` output for a pricing screen fixture; validation diagnostics (unknown field, unfilterable live source); pass-through of a hand-written template with source validation.
- [ ] Implement; green.
- [ ] Record commit draft: `feat(compiler): ##ui emits .htmlx — screen DSL parse/typecheck, binding+each+wo:live template generation with subscription keys, hand-written .htmlx pass-through with source validation.`
### Task 4: `.htmlx` SSR renderer
**Concept & reason:** the C port of the v1 engine's render semantics, scoped to the compiled subset: parse the template once at boot into a node tree (static chunks, bindings, each-blocks, partials, live-subtree markers); render a screen by walking the tree against query results from the plan-5 select path, HTML-escaping bound values, expanding each-blocks per row, and stamping each live subtree with the ids the client runtime needs to patch later (stable per-row element ids derived from class + row id — the contract the delta patcher relies on). Rendered pages route like any handler; the screen's route comes from the service/UI declarations.
- [ ] Failing tests: render goldens (template + fixture rows → exact HTML) covering escaping, each over rows, partials, and live-subtree id stamping; a malformed-template diagnostic at boot, not at request time.
- [ ] Implement; green.
- [ ] Record commit draft: `feat(runtime): .htmlx SSR renderer — boot-time parse to node tree, escaped binding render over select results, stable per-row ids in live subtrees; render goldens.`
### Task 5: Client runtime + static assets
**Concept & reason:** the last mile. `wo-live.js` (hand-written, one file, ~20 KB budget): on load, find `wo:live` subtrees, open the WebSocket, send subscribe messages from the stamped keys, and patch on frames — update replaces bound cell contents by stable id, insert appends a row rendered from a client-side row template the SSR emitted, delete removes the row's element; reconnect with backoff; a visible stale indicator when the socket is down (honesty over silence). Static serving: the asset module serves the JS (and any project assets) with correct content types and the sendfile-doctrine path for regular files.
- [ ] Failing tests: asset serving (content type, byte-exact body, 404 miss); client runtime exercised by the Task-6 end-to-end (no separate JS test harness — stated tradeoff: the e2e is the test).
- [ ] Implement; green.
- [ ] Record commit draft: `feat: wo-live.js client runtime (subscribe from stamped keys, patch update/insert/delete by stable ids, reconnect+stale indicator) + static asset serving on the sendfile path.`
### Task 6: Pricing live demo end to end + acceptance
**Concept & reason:** the spec's driving workload, closed: the pricing project compiles to one binary; a scripted browser-less client (a test WebSocket client speaking the pinned protocol) loads the SSR page, subscribes, then a method RPC (`set_price` over plan 6's route) commits an insert — the script asserts the delta frame arrives with the new amount, and that a second subscriber sees it too (fan-out). A manual demo recipe (`just pricing-live-demo`) serves it for human eyes. Render goldens + the scripted live scenario + the asset tests join `just oop-accept`. Docs closeout: kanban (13c-equivalent milestone on the C stack), CLAUDE.md stage table amendment, the UI track doc gains a status note that the format/live layer shipped and the workspace/per-app-binary layers (ui docs 04–06) remain open.
- [ ] Wire the scenario + recipe + gate; green; docs synced.
- [ ] Record commit draft: `test: pricing live demo e2e — SSR load, subscribe, set_price RPC commit, delta fan-out asserted by scripted WS clients; just pricing-live-demo; oop-accept gains the UI gate; UI-track status synced.`
---
## Plan self-review notes
- **Spec coverage (sub-project 5 + Stage 3):** `##ui` SSR, live subscriptions with delta frames on commit, the pricing live workload, single binary serving UI + API + DB — the spec's "single binary that is the database, the web API, and the UI" sentence is fully mechanized after this plan. Non-scope, stated: cross-shard live queries, the workspace/per-app-binaries/shared-daemon layers (ui docs 04–06), auth/`me` sessions, TLS.
- **Order rationale:** transport before registry before templates before renderer before client — each is the next one's substrate; the e2e needs all five.
- **Consistency check:** delta frame shape, subscribe message, stable-id contract, and `wo:live` semantics are pinned once (Task 1 doc) and consumed by Tasks 2–5; subscription keys originate in the compiler (Task 3) and terminate in the registry (Task 2) — same key format, one doc section.

View file

@ -0,0 +1,283 @@
# .wob Format + wovm VM Core Implementation Plan
> **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (`- [ ]`) syntax for tracking.
>
> **Style rule (user convention):** this plan states concept, reason, and required behavior in words. The executor writes the actual code at implementation time; nothing here is copy-paste source.
**Goal:** Build the C register VM (`wovm`) and pin the `.wob` bytecode format so it runs hand-assembled bytecode with the full milestone-1 memory model: owned objects with a runtime borrow word, `@gc` reference counting with budgeted cycle collection, and trap unwinding that never leaks.
**Architecture:** Plan 1 of 3 for the approved spec `docs/superpowers/specs/2026-08-01-oop-compiler-vm-design.md`. The existing `prototypes/wo-rt-c` moves to root-level `runtime/`; VM modules land in `runtime/src/`, C unit tests in `runtime/test/`, driven by an in-memory `.wob` assembler helper so the VM is fully testable before any compiler exists. Plan 2 (OCaml `woc` compiler front) and plan 3 (bytecode emit + end-to-end conformance corpus + single binary) follow and consume the format pinned here. Plan 2 needs OCaml + dune installed (not present on the dev box today).
**Tech Stack:** C11, libc only. gcc 13 with AddressSanitizer + UBSan for the test suite (valgrind is not installed — ASan is the primary checker). `make` inside `runtime/`, `just` recipes at the repo root.
## Global Constraints
- **libc only** in `runtime/` — no external C libraries (spec dependency doctrine).
- **C11**, `-Wall -Wextra -Werror` in test builds.
- **Little-endian on-disk format**; x86-64/ARM64 Linux only in milestone 1.
- **Limits, loader-enforced:** 64 registers per method, 4096 value-stack slots, 256 frames.
- **No undefined behavior on any input:** the loader validates every static index once; residual dynamic checks (borrow word, field bounds, map keys) always trap, never corrupt (spec section 6).
- **Traps never leak:** unwinding applies drop maps; the whole suite must be ASan-clean on success and trap paths alike.
- **Dispatch:** computed goto under GNU C, `switch` fallback under a `WO_ISO_C` define (spec section 5); both flavors stay under test.
- **Git: the executing agent NEVER runs `git commit`** (user convention). "Record commit draft" means: append the given message to `.dev/commit.md` at the repo root and move on; the user commits at their own checkpoints. `git add`/`git mv`/`git status` are fine.
- **Docs rule:** documentation lives under `docs/`; `runtime/README.md` stays an orientation README.
---
## File Structure
```
runtime/ (moved from prototypes/wo-rt-c — Task 1)
Makefile existing + wovm/test targets (grows per task)
README.md orientation note updated
wo-rt.c UNTOUCHED — event-loop reference for sub-project 2
bench/ moved as-is
src/
wob.h format constants, limits, opcodes, shared structs (Task 2)
obj.h obj.c arena allocator, object headers, strings (Tasks 3–4)
borrow.h borrow.c borrow-word acquire/release (Task 5)
cont.h cont.c native containers: multi, map (Task 6)
gc.h gc.c RC, drop plans, cycle collector (Tasks 7–8)
loader.h loader.c .wob loader + full validation (Task 10)
vm.h vm.c register interpreter, frames, traps, unwinding (Tasks 11–13)
builtin.h builtin.c builtin table (Task 14)
main.c wovm CLI (Task 16)
test/
t.h tiny assert harness (Task 2)
wob_build.h wob_build.c in-memory .wob assembler for tests (Task 9)
test_*.c one binary per module/topic (Tasks 3–15)
cli_smoke.sh end-to-end CLI check (Task 16)
docs/plan/oop-vm/00-wob-format.md normative format reference (Task 2)
justfile paths fixed (Task 1), wovm recipes (Task 16)
```
Every `test_*.c` compiles to its own binary and exits nonzero on failure; `make -C runtime test` builds and runs all of them as an ASan build.
---
## The `.wob` format v1 (normative — Task 2 copies this section into `docs/plan/oop-vm/00-wob-format.md`)
All integers little-endian; offsets are absolute file offsets.
**Header (44 bytes):** magic `"WOB1"`, version 1, then offset/count u32 pairs for the constant pool, class table, interface section, and method table, then a u32 entry-method index (all-ones = none).
**Constant pool** — sequential entries: one tag byte; tag 0 = i64 follows; tag 1 = text (u32 length + bytes, no NUL).
**Class table** — per class: name constant index, flags u32 (bit0 = instances are `@gc`), field count, then one kind byte per field padded to a 4-byte boundary. Field kinds: 0 SCALAR, 1 OWNED, 2 GCREF, 3 TEXT, 4 MULTI, 5 MAP. Runtime object layout: 16-byte header then one 8-byte slot per field, in declaration order.
**Interface section** — per interface: name constant index, method count. Global *slot ids* are assigned sequentially across interfaces in declaration order. Then a vtable row count and rows: class id, interface id, one method index per interface method.
**Method table** — per method: name constant index, class id (all-ones = free fn), arg count u8, register count u8, reserved u16, code length in bytes (multiple of 4), the u32 instructions, a line table (count + ascending pc→line pairs), and a drop table (count + ascending entries of pc, owned-register bitmask u64, gc-register bitmask u64). Drop-table lookup = last entry with pc ≤ current pc; no entry means nothing live.
**Instructions** — fixed 32-bit, Lua-style fields: opcode byte, A byte, then either B and C bytes or a 16-bit Bx (signed jumps encode as Bx − 32768).
| op | name | semantics (in words) |
| --- | --- | --- |
| 0 | NOP | nothing |
| 1 | LOADK A Bx | register A = constant Bx (int inline; text = pointer to interned const string) |
| 2 | MOVE A B | copy register; for owned values this IS the move — compiler guarantees the source is dead |
| 3–7 | ADD/SUB/MUL/DIV/NEG | i64 arithmetic, two's-complement wrapping (no signed-overflow UB); DIV traps on zero divisor and on INT64_MIN ÷ −1 |
| 8 | CONCAT A B C | new owned text from two texts |
| 9–12 | EQ/LT/LE/EQS | i64 compares and text-content equality, result 0/1 |
| 13–14 | JMP / JZ | relative jump (JZ when register A is zero) |
| 15 | CALL A Bx | call method Bx; callee's register window starts at caller base + A (register-window overlap, Lua-style); args sit at A, A+1, …; return value lands back in slot A |
| 16 | ICALL A Bx | interface call by global slot id Bx; receiver in A; vtable lookup by the receiver's class |
| 17–18 | RET A / RET0 | return value from register A (or zero), pop frame |
| 19 | NEW A Bx | new zeroed instance of class Bx |
| 20–21 | GETF / SETF | field read/write with runtime null/native/bounds checks (trap T_BOUNDS); overwriting a non-scalar field does NOT auto-drop the old value — the compiler emits the drop |
| 22 | DROP A | recursively drop the owned value in A per its class drop plan, null the register |
| 23–26 | BORROW_S/BORROW_X/RELEASE_S/RELEASE_X | borrow-word ops on the object in A; violation traps T_BORROW |
| 27–28 | RC_INC / RC_DEC | refcount ops on the `@gc` object in A |
| 29 | BUILTIN A B C | register A = builtin C applied to args starting at register B (fixed arity per builtin; `multi_new`/`map_new` carry kind immediates in B instead) |
| 30 | DB_STUB | trap T_DB "engine not linked" (spec: SQL-layer statements in milestone 1) |
| 31 | TRAP Bx | explicit trap with code Bx |
**Builtins:** now (ms), print (text), print_int, words (whitespace token count), multi_new/multi_push/multi_get/count/latest, map_new/map_set/map_get/map_has.
**Trap codes:** DIV0, BORROW, STACK, OOM, DB, BOUNDS, KEY, EXPLICIT.
---
### Task 1: Move `prototypes/wo-rt-c` → `runtime/`
**Files:** git-mv the directory; modify `justfile` (rt-c-demo, rt-c-bench paths), `runtime/README.md`, `CLAUDE.md` path references.
**Concept & reason:** the spec's monorepo decision — root-level `runtime/` is the real C runtime now, seeded by the wo-rt-c event-loop reference. Move first so every later task lands in the final location. Recipe names stay (muscle memory); only paths change. `wo-rt.c` itself is not touched.
- [x] Move with `git mv`, delete any stale built binary.
- [x] Update every `prototypes/wo-rt-c` reference in `justfile`, `CLAUDE.md`, and the README's self-description; add a one-line note in `runtime/README.md` explaining the move and the coming `src/` VM core.
- [x] Verify: `make -C runtime` still builds `wo-rt`; `just rt-c-demo` round-trips.
- [x] Record commit draft: `refactor: move prototypes/wo-rt-c to runtime/ (monorepo root dir per OOP spec); justfile/CLAUDE.md/README paths updated, wo-rt.c untouched.`
### Task 2: `wob.h` format header + test harness + format doc
**Files:** create `runtime/src/wob.h`, `runtime/test/t.h`, `docs/plan/oop-vm/00-wob-format.md`; modify `runtime/Makefile`.
**Concept & reason:** one header is the single source of truth for the compiler↔VM contract: magic/version, header field offsets, field kinds, object-header struct (16 bytes, static-asserted), flags (GC, cycle-buffer, const, two color bits for the collector), native class-id sentinels (string/multi/map/free-fn), borrow-word sentinels, trap codes, the opcode enum, instruction encode/decode helpers, builtin ids, limits, and a class-descriptor struct (name, flags, field count, kind array) shared by loader and runtime. `t.h` is a ~20-line assert harness (counters + report macro) — no external test framework, per doctrine. The Makefile gains a generic rule: each `test/test_*.c` builds into its own ASan binary linked against all `src/*.c` except `main.c`, and a `test` target runs them all.
- [x] Write a syntax-only compile check for the not-yet-existing header; see it fail.
- [x] Write `wob.h` with all constants above and a static assert that the object header is exactly 16 bytes; write `t.h`; extend the Makefile.
- [x] Verify the header compiles clean under `-Wall -Wextra -Werror`.
- [x] Create the format doc by copying this plan's normative format section verbatim; link it from the spec's plan references.
- [x] Record commit draft: `feat(runtime): wob.h .wob v1 contract (opcodes, kinds, traps, builtins, 16-byte header), t.h harness, Makefile test scaffold; docs/plan/oop-vm/00-wob-format.md normative reference.`
### Task 3: Arena allocator
**Files:** create `runtime/src/obj.h`, `runtime/src/obj.c`; test `runtime/test/test_arena.c`.
**Concept & reason:** spec section 4 — per-shard arena with size-class free lists. One malloc'd region; allocations round to 16 bytes; sizes up to 1024 use per-size free lists (freed blocks chain through their own first word); larger sizes fall through to plain malloc/free (caller always passes the size back on free, so no size headers are needed). Region exhaustion returns null — the VM maps that to the OOM trap, never aborts.
**Interface (in words):** init with a byte capacity, destroy, alloc(size)→pointer-or-null, free(pointer, size).
- [x] Write failing tests: same-size-class reuse returns the freed block; exceeding the region returns null; >1024 sizes succeed regardless of region cap.
- [x] See them fail (missing header), implement, see them pass under ASan.
- [x] Record commit draft: `feat(runtime): arena allocator — 16-byte size-class free lists to 1024B, malloc fallback above, NULL on region OOM; test_arena.`
### Task 4: Object model, strings, runtime context
**Files:** extend `obj.h`/`obj.c`; test `runtime/test/test_obj.c`.
**Concept & reason:** everything on the heap carries the 16-byte header — class objects, strings, containers — so drop/GC/borrow logic can dispatch on any value uniformly. A `wo_rt` context struct bundles what every module needs: the arena, the class-descriptor table, the cycle-candidate buffer (filled by Task 8), and an output stream pointer (so tests can capture builtin `print`). Object creation zeroes all field slots and, for `@gc` classes, sets the GC flag and refcount 1. Strings are header + length + inline bytes; concat allocates a new owned string; freeing a const-flagged string is a no-op (const strings are interned by the loader and outlive everything).
**Interface (in words):** rt init/destroy; object-size-of-class helper; object-new by class id (null = OOM); string new/concat/content-equality/free; a fields-accessor giving the slot array behind a header.
- [x] Failing tests: owned object comes back zeroed with free borrow word; `@gc` object starts at rc 1 with the GC flag; string round-trip + concat + equality; const-flagged string survives a free call.
- [x] Implement; suite green under ASan.
- [x] Record commit draft: `feat(runtime): object model — wo_rt context, wo_obj_new (zeroed, @gc rc=1), strings with const interning; test_obj.`
### Task 5: Borrow-word runtime
**Files:** create `runtime/src/borrow.h`, `borrow.c`; test `runtime/test/test_borrow.c`.
**Concept & reason:** spec section 4's residual-check primitive. The header's borrow word counts shared readers; an all-ones sentinel means exclusively borrowed. Acquire-shared fails only against exclusive; acquire-exclusive fails unless completely free; releases are unconditional (the compiler emits them balanced). Failures return an error code — the VM (Task 13) turns that into the borrow trap. ~20 lines of C; kept its own module because it is the semantic heart of the hybrid model and the compiler's emit rules will cite it.
- [x] Failing tests: the full state machine — stacked shared readers block exclusive; exclusive blocks both; releases restore.
- [x] Implement; green.
- [x] Record commit draft: `feat(runtime): borrow-word runtime — shared counter / exclusive sentinel, -1 on violation (VM maps to T_BORROW); test_borrow.`
### Task 6: Native containers — multi and map (data ops)
**Files:** create `runtime/src/cont.h`, `cont.c`; test `runtime/test/test_cont.c`.
**Concept & reason:** spec section 3 — `multi T` and `map<K,V>` are runtime-provided native classes, not user generics. Multi: header + length/capacity + element-kind tag + growable u64 array (malloc/realloc-backed). Map: header + parallel key/value arrays + kind tags for both, **linear scan lookup** — deliberate milestone-1 KISS, documented; keys compare by content when the key kind is text, by bits otherwise. Map insert-or-replace hands the replaced old value back to the caller instead of dropping it, because dropping needs Task 7's machinery — the builtin layer (Task 14) does the drop. Container freeing (element drops) also lands in Task 7 for the same reason; this task's tests free backing arrays by hand.
- [x] Failing tests: multi push/get across growth + bounds miss; map insert/replace/get/has with int keys; map text keys hit by content through a different pointer.
- [x] Implement; green.
- [x] Record commit draft: `feat(runtime): native containers — multi (growable, elem-kind tag), map (linear scan, content keys, replace returns old value); test_cont.`
### Task 7: RC + drop plans (deterministic destruction)
**Files:** create `runtime/src/gc.h`, `gc.c`; test `runtime/test/test_rc.c`.
**Concept & reason:** spec section 4 — owned objects die deterministically; `@gc` objects die at refcount zero. One kind-directed dispatcher ("drop this 8-byte value known to be of field kind K") is the workhorse: scalars ignored, owned values drop recursively, gc refs decrement, texts free, containers free element-wise then their backing. A class object's drop walks its kind array over its field slots, then frees itself; native headers (string/multi/map sentinels) route to their own frees. Refcount decrement to zero releases the same way (with a guard for objects sitting in the cycle buffer — the collector owns their death, Task 8).
**Testing trick (used by all memory tests from here on):** classes given ~130 fields exceed the 1024-byte size-class ceiling, so instances take the malloc path — any missed free is a hard ASan leak report, making destructor correctness machine-checked instead of eyeballed.
- [x] Failing tests: owned-tree recursive drop; gc-ref field decrement then final decrement frees; text and multi-of-text fields freed with the holder.
- [x] Implement; whole suite ASan-clean.
- [x] Record commit draft: `feat(runtime): RC + drop plans — kind-directed drop dispatcher, recursive class drops, container element drops, rc inc/dec with zombie guard; test_rc (malloc-path classes make ASan prove every free).`
### Task 8: Budgeted cycle collector
**Files:** extend `gc.c`; test `runtime/test/test_cycle.c`.
**Concept & reason:** spec section 4 — Bacon–Rajan trial deletion, per-shard, budgeted, never global. Decrements that leave a possibly-cyclic object alive (class has any gcref/multi/map field) buffer it as a candidate, deduplicated by a header flag. A collection step takes up to *budget* buffered roots and runs the classic three phases — mark-gray (trial-delete internal edges by decrementing through them), scan (restore subgraphs that still have external count), collect-white (gather the dead) — using the two header color bits. Two deliberate simplifications, both documented in code: (1) whole strongly-connected components process atomically, so budget bounds *roots started*, with bounded overshoot; (2) all gathered whites are freed together *after* the phases (contents first — skipping already-trial-deleted gcref edges — then the objects), which removes every dangling-candidate/zombie hazard that plagues incremental freeing. Child traversal descends through container elements whose kind is gcref. The step returns how many objects it freed, which is what tests observe.
- [x] Failing tests: a two-object cycle collects fully; two independent cycles under budget 2 collect one per step; an externally-held cycle survives with counts restored and flags cleared, then collects once truly dead; a cycle routed through a multi's elements collects.
- [x] Implement; ASan-clean (cycle objects use the malloc-path trick).
- [x] Record commit draft: `feat(runtime): budgeted cycle collector — Bacon-Rajan trial deletion in epochs, candidate buffer with flag dedup, deferred freeing kills zombie hazards, container edges traversed; wo_gc_step(budget) returns freed count; test_cycle.`
### Task 9: `.wob` assembler test helper
**Files:** create `runtime/test/wob_build.h`, `wob_build.c`; test `runtime/test/test_wobbuild.c`; Makefile links the helper into every test binary.
**Concept & reason:** the VM must be testable without a compiler (plan 2 is months of OCaml away). A small builder assembles valid images in memory: append constants (int/text), classes (flags + kind bytes + padding), interfaces (slot ids accumulate in declaration order), vtable rows, and methods (code array + optional line pairs + optional drop entries of pc/owned-mask/gc-mask), then finish by computing the section offsets into a header and concatenating — returning one malloc'd image. Every later VM test and the Task-16 smoke generator is written against this helper, which also makes it the second, independent encoding of the format — disagreements between builder and loader surface as test failures, effectively cross-checking the spec.
- [x] Failing test: build a minimal one-method image and assert raw header bytes — magic, version, section counts, entry index, first constant's tag/length/content at its stated offset.
- [x] Implement (dynamic byte buffers per section; sizes computed at finish); green.
- [x] Record commit draft: `test(runtime): wob_build in-memory .wob assembler (all sections, line+drop tables, offset computation); test_wobbuild checks raw header bytes.`
### Task 10: Loader with full validation
**Files:** create `runtime/src/loader.h`, `loader.c`; test `runtime/test/test_loader.c`.
**Concept & reason:** the "no UB on any input" constraint lives or dies here. The loader parses a byte buffer through a bounds-checked cursor (every read validates remaining length), **copies everything out** into aligned, malloc'd structures — so misaligned files and lifetime coupling to the input buffer are non-issues — and interns text constants as const-flagged strings. File loading is mmap → parse → munmap. Output module: constant array, class-descriptor table (kind bytes pooled), interface slot count, vtable rows expanded to flat (class, slot, method) triples sorted for binary search, and per-method records (code, line table, drop table).
**Validation contract (what the interpreter is allowed to assume forever):** magic/version match; register counts in 1..64 with args ≤ registers; every opcode known; every static register operand within the method's register count; constant/class/callee/slot indexes in range; CALL argument windows fit the caller's frame; jump targets inside the code; the **last instruction is a terminator** (return/trap/db-stub/jump — nothing falls off the end); builtin ids in range with per-builtin arity fitting the frame and kind-immediates in range; line/drop tables strictly ascending and in range; vtable ids in range; entry (if present) is a zero-arg free fn. Field indexes on GETF/SETF are deliberately NOT static-checkable (registers are untyped) — they stay runtime checks (Task 12), matching the spec's residual-check doctrine. Every rejection fills a human-readable error naming method and pc; every failure path frees everything parsed so far.
- [x] Failing tests: a valid builder image loads with correct counts, interned const string, line entries; then a rejection battery — corrupted magic, unknown opcode, constant index out of range, non-terminator tail, register out of range, truncated buffer — each must fail with a nonempty error and no leak.
- [x] Implement; green, including ASan on the failure paths.
- [x] Record commit draft: `feat(runtime): .wob loader — bounds-checked parse, aligned copies, const-string interning, sorted vtable expansion, full static validation (opcode/reg/index/jump/terminator/arity), mmap file path; test_loader happy + 6 rejects, leak-free failures.`
### Task 11: Interpreter core — arithmetic, control flow, calls, traps
**Files:** create `runtime/src/vm.h`, `vm.c`; test `runtime/test/test_vm.c`; Makefile gains a `test-iso` target.
**Concept & reason:** the register machine per spec section 5. A VM instance owns the module pointer, a runtime context, the 4096-slot value stack, and a 256-deep frame stack (method, pc, base). Calls use Lua-style window overlap: the callee's register 0 is the caller's slot A, so argument passing and value return are the same copy — returning writes the result into the call slot and pops. Non-argument callee registers are zeroed on entry (drop masks must never see stale bits). Dispatch is the dual-flavor macro pattern: computed goto under GNU C, plain switch under the ISO define, one shared case-body text — and the ISO flavor gets its own Makefile target run in every gate so the fallback can never rot. Arithmetic is two's-complement wrapping via unsigned math (no signed-overflow UB); division traps on zero and on the INT64_MIN ÷ −1 corner. The public entry ("call method with these argument words") validates method index and arity, seeds frame zero, runs to completion, and on any trap fills the structured error — code, method name (from the name constant), source line (from the line table, looked up at the trapping pc), message — then unwinds. In this task unwinding just pops frames; drop maps arrive in Task 13. Ops owned by Tasks 12–15 exist as cases that trap "op not wired yet" — loader-legal, semantically explicit, replaced task by task.
- [x] Failing tests: two-arg add returns through the window convention; recursive fib(10) = 55 (exercises call/return, jumps, compares); division by zero traps with the right code AND the right line; self-call-forever traps stack overflow at the frame cap; depth is zero after every trap.
- [x] Implement; both dispatch flavors green (`test` + `test-iso`).
- [x] Record commit draft: `feat(runtime): interpreter core — dual dispatch (computed goto / WO_ISO_C switch), register-window calls, wrapping i64 arithmetic, structured trap errors with line lookup; unwired ops trap explicitly; test_vm (add, fib, div0+line, stack cap) + test-iso gate.`
### Task 12: Object opcodes — NEW / GETF / SETF / DROP
**Files:** modify `runtime/src/vm.c`; test `runtime/test/test_objops.c`.
**Concept & reason:** wire the object model into the loop. One shared receiver-check helper enforces the residual runtime checks the loader cannot do statically: non-null, an actual class object (not a native sentinel), field index inside the class — any failure traps BOUNDS with a message naming the reason. NEW allocates by class id (OOM traps, never aborts). SETF stores the raw word — overwriting a non-scalar field does not auto-drop the old value; that is the compiler's obligation (already documented in the format section). DROP delegates to the Task-7 dispatcher and nulls the register so unwinding cannot double-free.
- [x] Failing tests: new → set field → get field → drop → return round-trip; out-of-range field index traps BOUNDS; null receiver traps BOUNDS.
- [x] Implement; suite + ISO flavor green.
- [x] Record commit draft: `feat(runtime): object opcodes — NEW (T_OOM), GETF/SETF null/native/bounds residual checks (T_BOUNDS), DROP with register nulling; test_objops.`
### Task 13: Borrow/RC opcodes + drop-map trap unwinding
**Files:** modify `runtime/src/vm.c`; test `runtime/test/test_unwind.c`.
**Concept & reason:** the spec's core promise — "traps never leak" — becomes real. The four borrow ops and two rc ops call the Task-5/Task-7 primitives after the same receiver checks; acquire failures trap BORROW. The Task-11 unwinding stub is replaced: walk frames innermost to outermost; in each, look up the drop-table entry for that frame's current instruction (the trap pc for the innermost frame; the instruction before the saved resume pc — i.e. the CALL — for outer frames); apply the owned mask by recursive drop and the gc mask by decrement, nulling registers as they go. A frame with no matching entry drops nothing. A borrow held by a dying register does not block its drop — the borrower IS the dying frame (noted in a comment).
- [x] Failing tests: double-exclusive borrow traps BORROW with correct line, and the owned object named in the drop mask is freed (malloc-path class, ASan-proven); a two-frame program where the child traps mid-body frees owned objects in BOTH frames; rc inc/dec through opcodes frees at zero.
- [x] Implement; suite + ISO green, ASan silent on all trap paths.
- [x] Record commit draft: `feat(runtime): borrow/rc opcodes + drop-map unwinding — T_BORROW on violation, frame walk applies owned/gc masks so traps never leak (spec section 6); test_unwind, ASan-proven.`
### Task 14: Builtins + DB_STUB + TRAP
**Files:** create `runtime/src/builtin.h`, `builtin.c`; modify `vm.c` (three cases); test `runtime/test/test_builtin.c`.
**Concept & reason:** the runtime services bytecode can't express. One dispatcher takes the VM, the current frame's registers, and the instruction; returns success or a trap code plus message. Behaviors: `now` = wall-clock milliseconds; `print` writes a text plus newline to the runtime context's output stream (tests point that stream at a temp file and diff it — the same observation seam plan 3's conformance corpus will use); `print_int` likewise; `words` counts whitespace-separated tokens. Container builtins bridge to Task 6 with type checks on the receiver header (wrong native class traps BOUNDS): multi new/push/get/count/latest (get/latest bounds-trap; push OOM-traps), map new/set/get/has (get of a missing key traps KEY; set drops a replaced old value via the Task-7 dispatcher — closing the loose end Task 6 left open; new decodes its two kind nibbles from the immediate). DB_STUB traps DB with "engine not linked" — the spec's parse-but-trap story for SQL statements. TRAP raises its immediate as the trap code.
- [x] Failing tests: print/print_int output captured and diffed; multi push/count/latest and map set/get/has driven from bytecode; missing map key traps KEY; empty-multi latest traps BOUNDS; words counts correctly; DB_STUB traps DB.
- [x] Implement; green. (CONCAT/EQS text ops wired here too — same test file covers them.)
- [x] Record commit draft: `feat(runtime): builtins (now/print/print_int/words/multi_*/map_* with replaced-value drop) + DB_STUB (T_DB engine-not-linked) + TRAP; test_builtin with captured-stdout diffs.`
### Task 15: ICALL — interface dispatch
**Files:** modify `runtime/src/vm.c`; test `runtime/test/test_icall.c`.
**Concept & reason:** spec section 2's structural interfaces reach the VM as data: the loader already produces sorted (class, slot, method) triples. ICALL checks the receiver (null/native → BOUNDS), binary-searches the triple table by the receiver's class and the instruction's slot id, and on a hit performs exactly the CALL sequence at the same window base (receiver already sits in slot A = callee's `self`). A miss — class doesn't implement the interface — traps BOUNDS with a "no vtable entry" message; the compiler's type checker makes that unreachable in compiled code, the VM keeps it as defense (spec section 6).
- [x] Failing tests: two classes implementing one interface through different methods return different results via the same ICALL site (real dynamic dispatch, chosen by receiver class); a class without a vtable entry traps with the no-vtable message.
- [x] Implement; green both flavors.
- [x] Record commit draft: `feat(runtime): ICALL — binary search over sorted (class,slot,method) vtable triples, same window convention as CALL, defensive no-vtable trap; test_icall.`
### Task 16: `wovm` CLI, smoke test, just recipes, final gate
**Files:** create `runtime/src/main.c`, `runtime/test/mkwob.c`, `runtime/test/cli_smoke.sh`; modify `runtime/Makefile` (wovm + mkwob targets), `justfile`.
**Concept & reason:** make the VM a real program with the exit-code contract plans 2–3 script against: usage or load failure → exit 2 with the loader's message; a trap → exit 1 with one stderr line in the fixed shape "trap CODE in METHOD at line N: MESSAGE"; success → exit 0. The heap cap defaults to 64 MiB, overridable via a `WO_HEAP_MB` environment variable (the spec's configurable-arena note). A tiny generator tool built from the test-side assembler writes two fixture files: a hello program (print text, print int) and a trapping program — the smoke script builds everything, runs both, diffs stdout for the first, and asserts exit code and stderr shape for the second. Root justfile gains `wovm-build` and `wovm-test` (the full make gate: unit suite + ISO flavor + CLI smoke). The final step runs that whole gate as this plan's acceptance run.
- [x] Failing smoke: script exists and fails (no binary yet).
- [x] Implement CLI + generator + recipes; smoke green.
- [x] Run the full gate: `just wovm-test` — every unit test, the ISO dispatch flavor, and the CLI smoke, all ASan-clean. (13 suites × 2 flavors, 327 checks each, + smoke: all green.)
- [x] Record commit draft: `feat(runtime): wovm CLI (exit 0/1/2 contract, trap line format, WO_HEAP_MB) + mkwob fixture generator + cli_smoke.sh; justfile wovm-build/wovm-test full gate.`
---
## Plan self-review notes
- **Spec coverage (plan-1 slice):** header layout, borrow word, hybrid residual checks, `@gc` RC + budgeted no-pause cycle collection, deterministic owned drops, drop-map unwinding with structured errors, containers, structural-interface dispatch, DB_STUB story, computed-goto + ISO dispatch, arena OOM-as-trap — all mapped to tasks above. Deliberately out of scope here (plan 2/3): everything OCaml, ownership inference, `.wo` parsing, conformance corpus, single-binary append trick, parity harness.
- **Order rationale:** memory model before interpreter (Tasks 3–8 testable without bytecode), assembler before loader (Task 10's tests need images), loader before interpreter (validation contract is what keeps the hot loop check-free).
- **Known accepted simplifications, all documented in their tasks:** map linear scan; cycle-collector component-atomic budget; SETF no auto-drop; big-endian unsupported.
## Execution note
Executor needs only gcc/make/just (present on the dev box). Valgrind optional; ASan is the gate. OCaml is NOT needed for this plan.