Pivot to the .wo language runtime: design docs, phase plans

This commit is contained in:
shoney.arickathil 2026-04-21 02:48:45 +02:00
parent 43b9bc689a
commit c93915a3cd
40 changed files with 5731 additions and 63 deletions

26
.gitignore vendored
View file

@ -1,3 +1,29 @@
# Cargo build artifacts
/target
/reference/crates/target
# C++ prototype build output
/prototypes/*/build
# Runtime data directories for the sample projects.
# `wo.toml` points at `./data` which holds the per-project engine state.
/data
/docs/examples/*/data
# Legacy blog content and data (v1 writeonce storage)
/content
# Symlink to the Linux kernel source tree for research — user-specific
# absolute path; each contributor sets their own via
# ln -s <path-to-linux-src> reference/linux
/reference/linux
# Editor / OS noise — left broad on purpose so a contributor doesn't
# accidentally commit their IDE scratch or macOS metadata.
.DS_Store
*.swp
*.swo
/.idea/
/.vscode/*
!/.vscode/settings.json.example
!/.vscode/extensions.json

138
CLAUDE.md Normal file
View file

@ -0,0 +1,138 @@
# CLAUDE.md
This file provides guidance to Claude Code (claude.ai/code) when working with code in this repository.
## What this repo is
writeonce is a **declarative full-stack programming language**. You write `.wo` files; the runtime compiles them into a single binary that owns the database, serves REST, and (Stage 3+) pushes live subscriptions. Think: Go + Postgres + `net/http` + Phoenix LiveView folded into one language.
## The end goal — zero external deps, kernel primitives only
The runtime's north star — documented in [`docs/01-problem.md`](docs/01-problem.md), [`docs/02-recovery.md`](docs/02-recovery.md), and [`docs/plan/linux/00-linux.md`](docs/plan/linux/00-linux.md) — is **one binary, no external Rust crates, all I/O driven directly by Linux kernel primitives**. `epoll` (or `io_uring`), `inotify`, `eventfd`, `timerfd`, `signalfd`, `sendfile`, `mmap` — the kernel IS the subscription engine, the async runtime, and the file watcher.
Target end state of `crates/rt/Cargo.toml`:
```toml
[dependencies]
libc = "0.2" # the unavoidable FFI bridge to syscalls
```
Stage 2 today carries six transitional deps (`anyhow`, `serde`, `serde_json`, `tokio`, `axum`, `tower`). [`docs/plan/02`](docs/plan/02-event-loop-epoll.md) through [`docs/plan/08`](docs/plan/08-sendfile-static-assets.md) sequence the removal of each one, replaced by hand-rolled modules ported from the v1 crates that already did exactly this (`reference/crates/wo-event`, `wo-http`, `wo-route`, `wo-serve`, `wo-watch`). **When working on the runtime, default to direct-syscall solutions over reaching for new crates** — the `docs/plan/` docs name the port source for every module.
## Layout
The repo holds three cuts of the same project plus one research reference:
1. **`crates/rt/`** — the **active Rust runtime** (Stage 2 shipped). Monolithic on purpose for now: lexer, parser, AST, in-memory engine, axum REST server all in one crate. The `wo` binary lives at `crates/rt/src/bin/wo.rs`.
2. **`crates/{ql,value,engine,txn,db,wal,sub,http,gen,policy,logic,service,ui,app}/`** — 14 **empty sibling crates** scaffolded to match the 7-phase design. Each has a `Cargo.toml` + `src/lib.rs` with just a doc comment pointing at its phase spec. Real code moves in from `rt` as each phase activates; do NOT refactor `rt` to use these today — it would break Stage 2.
3. **`reference/crates/`** — the **v1 writeonce blog** (13 crates: `wo-seg`, `wo-index`, `wo-store`, `wo-htmlx`, etc.). This is a **nested Cargo workspace**, deliberately excluded from the root workspace. The v1 crates keep their `wo-` prefix; the new runtime crates dropped theirs. `cd reference/crates && cargo build` builds v1 standalone. See `docs/runtime/database/07-wo-seg-migration.md` for the plan replacing v1 with the new runtime.
4. **`reference/linux/`** — **symlink to the Linux kernel source tree** (`/home/shoney/projects/linux`). Not committed (see `.gitignore`). Research resource for the kernel-primitives work: read `io_uring/`, `fs/notify/inotify/`, `kernel/eventfd.c`, `include/uapi/linux/*.h` when designing the runtime's kernel-facing modules. Each contributor sets their own target via `ln -s <path-to-linux-src> reference/linux`.
There is also **`prototypes/wo-db/`** — a ~2k-line **C++ prototype** of the query-layer engine (SQL + Cypher + document paths, `RETURNING` aliases, `LIVE` stub). It keeps its `wo-db` directory name (C++ project, separate from the Rust crate `db`). It's the reference implementation the Rust port follows; `make test` still passes.
## Commands
```bash
# Build + test the runtime
cargo build # compiles all 15 crates
cargo test --lib # 14 unit tests (all in rt today)
cargo test --lib parses_inline_struct -- --nocapture # single named test
# Run the runtime against a sample project
cargo run --bin wo -- run docs/examples/blog # :8080 — blog sample
cargo run --bin wo -- run docs/examples/ecommerce # :8080 — ecommerce sample
WO_LISTEN=127.0.0.1:9000 cargo run --bin wo -- run docs/examples/blog # override port
# v1 blog codebase (nested workspace — must cd first)
cd reference/crates && cargo build && cargo test
# C++ prototype of the query-layer engine
cd prototypes/wo-db && make test # smoke.wo + checkout.wo
cd prototypes/wo-db && make run # interactive REPL
# Manual HTTP smoke against a running `wo run ...`
# Open reference/rest/blog.rest or ecommerce.rest in VS Code (with REST Client)
# or JetBrains (built-in HTTP client). Or run curl per reference/rest/README.md.
```
## Architecture — what requires reading multiple files to understand
### Naming convention
- **New runtime crates are unprefixed.** `ql`, `value`, `engine`, `txn`, `db`, `wal`, `sub`, `http`, `gen`, `policy`, `logic`, `service`, `ui`, `app`, `rt`. Internal imports read cleanly: `use ql::Parser`, `use db::Tx`, `use http::router`.
- **V1 crates keep the `wo-` prefix.** `wo-seg`, `wo-index`, `wo-store`, `wo-htmlx`, `wo-md`, and v1's own `wo-rt`/`wo-http`/`wo-sub`. These live in `reference/crates/`.
- **The C++ prototype directory is `prototypes/wo-db/`** — unchanged, not a Rust crate.
- **The binary is `wo`** — defined in `crates/rt/Cargo.toml` `[[bin]]`. Independent of the crate name.
### The two-layer `.wo` language
Covered in `docs/runtime/database/02-wo-language.md`:
- **Schema layer** — unified `type Name { ... }` DSL (fields, embedded structs, `ref`, `multi @edge`, `multi via`, `backlink`, tagged unions, `policy`, `on <event>`, `service`). This is what developers write day-to-day.
- **Query layer** — hybrid SQL + Cypher with five "fixed-glue" rules that make the three grammars share semantics: `$name` parameters everywhere, cross-paradigm `RETURNING col AS alias` visible to later statements in the same `BEGIN … COMMIT`, dotted paths identical in SQL/doc/Cypher, one transaction block syntax, one `LIVE` prefix on subscriptions.
Both layers are `.wo` files. The schema layer compiles down to query-layer operations — but only when Phase 5 codegen and Phase 6 full-stack blocks need a single authoritative input. Stage 2 ships with the schema layer only.
### Single-threaded event loop
Covered in `docs/runtime/database/02-wo-language.md § Concurrency Model` and `03-inmemory-engine.md`. The runtime is Redis/TigerBeetle-style: **one userland thread owns everything** — connection accept, parser, engine, subscription registry. The only non-userland thread is the kernel-owned io_uring SQPOLL helper. This is pinned architecturally — group commit still applies (loop drains many commits into one fsync SQE per tick), and scaling past one core is done by **sharding** independent engine processes, not by adding worker threads. Keep this in mind before proposing `Arc<Mutex<...>>` anything beyond what's already there.
### What's in `rt` today vs. what the empty crates promise
`rt`'s modules deliberately mirror the future crate names so the extraction is mechanical when each phase activates:
| `rt` module | Will move to | Phase |
| --- | --- | --- |
| `token.rs` + `lexer.rs` + `ast.rs` + `parser.rs` | `ql` | 2 |
| `engine.rs` (Value + Row helpers) | `value` | 2 |
| `engine.rs` (Engine + Catalog) | `engine` | 2 |
| `compile.rs` | `engine` | 2 |
| `server.rs` | `http` + `service` | 4 / 6 |
| `bin/wo.rs` | stays in `rt` (the binary) | — |
The `sub`, `wal`, `txn`, `policy`, `logic`, `ui`, `app`, `gen` crates have no `rt` counterpart yet — they land when their phase activates.
### Sample projects drive the grammar
`docs/examples/blog/` and `docs/examples/ecommerce/` are **both docs artifacts and the de facto integration tests**. The parser survives these because specific features in them (nested `{...}` object literals inside trigger actions, `count(...)` / `words(...)` computed defaults, unions like `Pending | Paid | Shipped`) forced real fixes. When changing the parser, run the full end-to-end against both samples, not just `cargo test`.
The ecommerce sample in particular uses features that are deliberately **parse-and-discard** in Stage 2: `fn checkout(...) in txn snapshot`, type-attached `on update` triggers with multi-line `do` actions, `policy read for role ...`. These are part of the `.wo` language but Stage 2 does not execute them.
### The migration story
`docs/runtime/database/07-wo-seg-migration.md` specifies **phased coexistence**: abstract the v1 article store behind an `ArticleStore` trait, stand up the `.wo` engine as a second implementation, dual-write, cut over, decommission v1. Phase A (trait abstraction) hasn't started — the plan is on paper, the v1 code is still monolithic in `reference/crates/wo-store/`. Do not remove anything from `reference/crates/` without checking the migration doc.
### Stage progress
| Stage | Status |
| --- | --- |
| 1 — `wo run <dir>` discovers `.wo` files | ✅ shipped |
| 2 — parser + engine + REST CRUD | ✅ shipped (`cargo run -- run docs/examples/blog`) |
| 3 — LIVE subscriptions over WebSocket | pending — `/api/<type>/live` returns 501 as a stub |
| 4+ — transactional `fn`, policies, triggers, `##ui`, WAL, codegen | design-only (see `docs/runtime/database.md`) |
Stage-3 stubs (501) and policy-shaped 405/404 responses are **intentional and documented** in `reference/rest/*.rest`. Don't "fix" them without checking those files first.
## Non-obvious gotchas
- **`rt` is monolithic on purpose.** Splitting it into the 14 sibling crates is Phase-by-Phase work, not a Stage-2 refactor.
- **`reference/crates/` is its own workspace.** Running `cargo build` at the root does not build v1. Running it in `reference/crates/` does.
- **Parser identifiers vs. keywords.** `subscribe`, `receive`, `expect_abort`, `me` are NOT keywords in the lexer — they stay as plain idents so they can appear as operation names in `expose` lists. Adding them to the keyword map breaks `service rest "..." expose subscribe`.
- **Parser skip-on-block.** Unknown triggers (`on update do ...`) are parsed-and-discarded by brace-depth-aware skipping. Object literals like `{ article_id: self.id }` inside trigger actions contain `}` that must not be mistaken for the type's outer close brace — the depth counter exists specifically because of this.
- **Newline significance.** The lexer emits `Kind::Newline` tokens and the parser uses them to end policy/trigger lines. Do not filter newlines globally.
- **Default-value parsing.** `= now()` is recognised explicitly as `DefaultExpr::Now`; anything else falls into an opaque-expression path that `engine::eval_default` then **omits from created rows** (computed fields display as empty, not as debug-printed tokens).
- **Binary variable shadowing.** `crates/rt/src/bin/wo.rs` has `let rt = ...` (a tokio runtime handle) inside `run()` that shadows the crate named `rt`. Inside `run()` the variable wins; inside `serve()` (a different function) `rt::` refers to the crate. Don't rename the variable without also auditing the crate-path references.
## Where to read next
- `docs/runtime/wo-language.md` — user-facing language overview
- `docs/runtime/database.md` — 7-phase engineering series index
- `docs/plan/linux/00-linux.md` — catalogue of kernel primitives the runtime leans on
- `docs/plan/02-event-loop-epoll.md` through `08-sendfile-static-assets.md` — the dependency-removal phase sequence
- `docs/plan/done/01-scafolding-crates.md` — the completed crate-scaffolding phase
- `docs/examples/blog/README.md` — the canonical worked example
- `prototypes/wo-db/README.md` — the C++ prototype that shows the query layer
- `reference/rest/README.md` — how to exercise the running prototype
- `reference/README.md` — what's in the v1 archive and why it's preserved
- `reference/linux/` (symlink) — the Linux kernel source tree itself; grep `io_uring/`, `fs/notify/inotify/`, `include/uapi/linux/*.h` when designing kernel-facing modules
- `crates/README.md` — inventory of all 15 crates with phase assignments

110
README.md
View file

@ -1,64 +1,72 @@
# writeonce
A single self-contained binary that serves a content platform — no external database, no cloud pipeline, no JavaScript framework. Built in Rust on raw Linux kernel primitives.
A declarative full-stack programming language. You write `.wo` files; the runtime compiles them into a binary that owns the database, serves REST, and pushes live subscriptions — no external database, no external web server, no frontend framework.
## Why
Think **Go + Postgres + `net/http` + Phoenix LiveView, folded into one language and one binary.**
The original writeonce system spread across five repositories, four languages, AWS infrastructure (S3, Lambda, API Gateway), PostgreSQL, and an Angular frontend. All of that to serve articles from local files. This project collapses everything into one process that owns storage, serves content, and pushes real-time updates.
## Quickstart
## Architecture
- **Single process** — one binary replaces S3 + Lambda + Rust API + PostgreSQL + Angular
- **Embedded storage** — custom `.seg` segment files with positional indexing, no external database
- **Real-time subscriptions** — route-based SSE streams push content diffs to connected clients
- **Server-rendered HTML** — `.htmlx` templates with data bindings, minimal client-side JS
- **Markdown-first content** — `.md` files are the source of truth, JSON holds only metadata
- **Linux kernel I/O** — `epoll`, `inotify`, `eventfd`, `timerfd`, `sendfile` — no tokio, no async runtime
## Workspace Crates
| Crate | Purpose |
|-------|---------|
| `wo-model` | Article and metadata types |
| `wo-seg` | Segment file reader/writer (.seg format) |
| `wo-index` | Title hash map, date sorted array, tags inverted index |
| `wo-store` | Query API over segments and indexes |
| `wo-watch` | `inotify`-based content directory watcher |
| `wo-event` | `epoll` event loop, `eventfd`, `timerfd`, `signalfd` |
| `wo-sub` | Subscription manager and diff delivery |
| `wo-rt` | Single-threaded runtime tying I/O sources together |
| `wo-http` | HTTP request parsing and response writing |
| `wo-route` | URL routing and handler dispatch |
| `wo-htmlx` | Template engine for `.htmlx` files |
| `wo-md` | Markdown to HTML rendering |
| `wo-serve` | Binary entry point — wires everything together |
## Build
```sh
cargo build --release
```bash
git clone https://github.com/shoneyJ/writeonce
cd writeonce
cargo run --bin wo -- run docs/examples/blog # serve the sample blog on :8080
curl http://127.0.0.1:8080/api/articles # it's a real REST API now
```
## Deploy
See [`reference/rest/blog.rest`](reference/rest/blog.rest) for a preconfigured HTTP-request file that drives the whole sample — open it in VS Code (with the REST Client extension) or JetBrains and click "Send Request" on each block.
The binary runs behind nginx with Let's Encrypt SSL. See `infra/setup.sh` for first-time server setup and `docs/07-ssl.md` for the full deployment walkthrough.
## What this repository contains
```sh
# Build, copy binary, sync content, restart service
./infra/deploy.sh
| Path | What it is |
| --- | --- |
| [`crates/rt/`](crates/rt/) | The new `.wo` language runtime — lexer, type-DSL parser, in-memory engine, axum REST server. Produces the `wo` binary. |
| [`crates/{ql,value,engine,txn,db,wal,sub,http,gen,policy,logic,service,ui,app}/`](crates/) | 14 empty placeholder crates scaffolded for Phases 2–6. Real code extracts from `rt/` as each phase activates. |
| [`docs/runtime/wo-language.md`](docs/runtime/wo-language.md) | **Start here.** The language overview: toolchain, hello-world, stdlib, client model. |
| [`docs/runtime/database.md`](docs/runtime/database.md) | The 7-phase engineering series that drives the runtime's design. |
| [`docs/examples/blog/`](docs/examples/blog/) | Sample `.wo` project: blog with articles, authors, tags, comments. ~200 lines. |
| [`docs/examples/ecommerce/`](docs/examples/ecommerce/) | Sample `.wo` project: storefront + live order-ops table + cross-paradigm checkout. ~300 lines. |
| [`prototypes/wo-db/`](prototypes/wo-db/) | C++ prototype of the query-layer engine (SQL + Cypher + document paths, `RETURNING` aliases, `LIVE` stub). ~2k lines, smoke tests pass. Reference implementation the Rust port follows. |
| [`reference/rest/`](reference/rest/) | `.rest` files (VS Code REST Client / JetBrains HTTP format) for manually testing the running prototype. |
| [`reference/crates/`](reference/crates/) | The v1 writeonce blog — 13 Rust crates implementing the original `.seg` + sidecar-index storage engine and `.htmlx` templating. Preserved as a nested workspace; see [`reference/README.md`](reference/README.md). |
## Current stage
The runtime is under active development. Each stage lands as an independently shippable cut:
| Stage | What works | Status |
| --- | --- | --- |
| **1** | `wo run <dir>` discovers every `.wo` file under a directory | ✅ shipped |
| **2** | Type-DSL parser, in-memory engine, REST CRUD (`list` / `get` / `create` / `update` / `delete`) generated from `service rest` blocks, JSON bodies with auto-id, default-value seeding, partial-update PATCH | ✅ shipped — `cargo run -- run docs/examples/blog` |
| **3** | LIVE subscriptions over WebSocket, delta frames on commit, `me` / session layer | pending |
| **4+** | Transactional fns (`fn checkout in txn snapshot`), row-level policies, type-attached triggers, `##ui` SSR, WAL durability, codegen | see [docs/runtime/database.md](docs/runtime/database.md) |
`cargo test --lib` at the root runs 14 unit tests covering the lexer, parser, compiler, and engine. Stage-3 endpoints respond `501 Not Implemented` until they land.
## Build & test
```bash
cargo build # builds all 15 crates (only `rt` has real code)
cargo test --lib # 14 unit tests
cargo run --bin wo -- run docs/examples/blog # serve the blog sample
cargo run --bin wo -- run docs/examples/ecommerce # serve the ecommerce sample
# Override the listen address
WO_LISTEN=127.0.0.1:9000 cargo run --bin wo -- run docs/examples/blog
```
## Documentation
## The v1 codebase (reference)
Design documents live in `docs/`:
The original writeonce blog engine — 13 crates, flat-file `.seg` storage, sidecar indexes, `.htmlx` templates, hand-rolled `epoll` event loop — moved to [`reference/crates/`](reference/crates/) when the new runtime was scaffolded. It's a nested Cargo workspace:
- `00-linux.md` — Linux kernel primitives used
- `01-problem.md` — Problem statement and motivation
- `02-recovery.md` — Target architecture
- `03-data.md` — Embedded storage and subscription model
- `04-ui.md` — Server-rendered HTMLX templates
- `05-datalayer.md` — Data layer implementation status
- `06-markdown-render.md` — Markdown-first content model
- `07-ssl.md` — SSL, nginx, and deployment
- `runtime/` — Deep dives on async runtimes, fibers, and Rust's ownership model
- `future-scope/` — Planned features including AI agent content management
```bash
cd reference/crates
cargo build # all 13 v1 crates still compile
cargo test # 12 unit tests, 1 ignored integration test
```
V1 crates keep the `wo-` prefix (`wo-seg`, `wo-store`, …). The new runtime crates dropped it (`ql`, `value`, `engine`, …). [`docs/runtime/database/07-wo-seg-migration.md`](docs/runtime/database/07-wo-seg-migration.md) is the phased coexistence plan for replacing v1 with the new runtime — abstract behind a trait, dual-write, cut over, decommission.
## License & status
Work in progress. Nothing here is stable. Read the language overview in [`docs/runtime/wo-language.md`](docs/runtime/wo-language.md) if you want to know the shape; read the phase docs if you want to see the engineering plan; look in [`docs/examples/`](docs/examples/) if you want to see what the end product feels like.

54
crates/README.md Normal file
View file

@ -0,0 +1,54 @@
# `crates/` — the `.wo` runtime
Fifteen crates make up the new runtime. Only `rt/` carries real code today (Stage 2); the other fourteen are **empty placeholders** scaffolded to match the 7-phase design so each phase's extraction work becomes a mechanical code move into an existing home.
> The crate-name prefix `wo-` was dropped when the active project namespaced itself under `wo` (the binary, the file extension, the language). Internal imports read cleanly: `use ql::Parser`, `use db::Tx`, `use http::router`. The v1 codebase keeps its `wo-*` prefix in [`reference/crates/`](../reference/crates/) to distinguish the generations.
## Map
| Phase | Crate | Purpose | Status |
| --- | --- | --- | --- |
| 2 | [`ql`](./ql/) | `.wo` grammar — lexer, parser, AST | placeholder |
| 2 | [`value`](./value/) | tagged `Value` + dotted-path helpers | placeholder |
| 2 | [`engine`](./engine/) | in-memory executor (rel / doc / graph) + schema catalog | placeholder |
| 2 | [`txn`](./txn/) | transaction coordinator — MVCC, `RETURNING` alias table | placeholder |
| 2 | [`db`](./db/) | top-level facade — `open()`, `Tx`, `Query`, `Subscribe` | placeholder |
| 3 | [`wal`](./wal/) | write-ahead log — io_uring + fsync + recovery | placeholder |
| 4 | [`sub`](./sub/) | live subscriptions — delta frames on commit | placeholder |
| 4 | [`http`](./http/) | wire protocol — REST / GraphQL-over-WS / native codec | placeholder |
| 5 | [`gen`](./gen/) | codegen — `.wo type` → Go / TS / Rust / Python clients | placeholder |
| 6 | [`policy`](./policy/) | RBAC + row-level rules compiled into planner rewrites | placeholder |
| 6 | [`logic`](./logic/) | `on <event>` triggers + `fn ... in txn` interpreter | placeholder |
| 6 | [`service`](./service/) | `service rest/graphql/native` endpoint dispatch | placeholder |
| 6 | [`ui`](./ui/) | `##ui` screens → SSR HTML + client runtime | placeholder |
| 6 | [`app`](./app/) | `##app` route manifest + startup hooks | placeholder |
| — | [`rt`](./rt/) | **active** — Stage-2 monolith + the `wo` binary | **shipped** |
## Why `rt/` is monolithic right now
`rt/` currently holds every module the runtime needs — lexer, parser, AST, in-memory engine, axum REST server — because **shipping working Stage 2 was more important than hitting the final crate layout on day one**. Each module inside `rt/src/` is written with a target home in mind:
| `rt` module | Moves to | Phase |
| --- | --- | --- |
| `token.rs` + `lexer.rs` + `ast.rs` + `parser.rs` | `ql/` | 2 |
| `engine.rs` (Value + Row helpers) | `value/` | 2 |
| `engine.rs` (Engine + Catalog) | `engine/` | 2 |
| `compile.rs` | `engine/` | 2 |
| `server.rs` | `http/` + `service/` | 4 / 6 |
| `bin/wo.rs` | stays in `rt/` (the binary) | — |
Extractions happen phase-by-phase — first one lands when a second caller appears (likely when Stage 3 needs the parser for raw-`.wo` HTTP requests).
## Build & test
```bash
cargo build # compiles all 15 crates
cargo test --lib # 14 unit tests (all in rt today)
cargo run --bin wo -- run docs/examples/blog # serve the blog sample
```
## What's outside this directory
- [`../reference/crates/`](../reference/crates/) — the v1 writeonce blog (13 crates, nested workspace). Preserved for reference per [docs/runtime/database/07-wo-seg-migration.md](../docs/runtime/database/07-wo-seg-migration.md). Keeps its `wo-*` prefix.
- [`../prototypes/wo-db/`](../prototypes/wo-db/) — C++ prototype of the query-layer engine (~2k lines). The reference implementation this Rust port follows at the language level.
- [`../docs/plan/`](../docs/plan/) — planning documents for in-flight work (the `.md` files directly under `plan/` are upcoming phases; `plan/done/` holds completed ones). [`plan/done/01-scafolding-crates.md`](../docs/plan/done/01-scafolding-crates.md) is the authoritative scope doc for the 14 new placeholders.

View file

@ -0,0 +1,175 @@
# `blog` — a sample writeonce app
A complete blogging website in **~200 lines of `.wo`** that creates a database, exposes REST + live-subscription endpoints, renders HTML pages, enforces row-level policies, and emits typed client SDKs.
> This project is a **docs artifact** — it illustrates the shape of a real `wo init`'d project. The `wo` toolchain referenced here is the one specified in [`../../runtime/wo-language.md`](../../runtime/wo-language.md); the engine is at prototype stage in [`../../../prototypes/wo-db/`](../../../prototypes/wo-db/).
## What it does
| Thing | How |
| --- | --- |
| Persists articles, authors, tags, comments | `type` declarations compiled to relational rows + embedded documents + graph edges |
| Serves 24 REST endpoints (CRUD + subscribe × 4 types) | `service rest` blocks on each type |
| Serves 4 web pages (list, detail, tag, admin) | `##ui` screens + route table in `app.wo` |
| Pushes live updates on every commit | `live: true` on screens + `LIVE` queries under the hood |
| Enforces "drafts hidden from anonymous readers" | `policy read anyone when published == true` |
| Bumps `published_at` automatically | `on update` trigger inside the transaction |
| Generates a typed Go client | `wo gen sdk --lang go` |
## Project layout
```
blog/
├── wo.toml # project manifest (like go.mod)
├── app.wo # routes, theme, startup hooks
├── types/
│ ├── author.wo # Author type + per-type service/policy
│ ├── article.wo # Article — all three paradigms in one type
│ ├── tag.wo # Tag taxonomy
│ └── comment.wo # Reader comments
├── ui/
│ ├── article_list.wo # home page list view (live)
│ └── article_detail.wo # per-article page with comments + related
└── tests/
└── article_test.wo # `wo test` picks this up
```
No `main.wo` is needed — a pure type+service app auto-generates its entry point. Add `main.wo` if you need CLI args, background workers, or custom startup logic beyond the `on startup` hook in `app.wo`.
## Run it
```bash
$ cd docs/examples/blog
$ wo run
[wo] parsing: 7 files, 4 types, 2 ui screens
[wo] compiling schema: 4 sql tables, 1 doc collection, 3 graph edge types
[wo] starting runtime (engine: in-memory, data_dir: ./data)
[wo] on startup: seed_admin() — inserted admin@example.com
[wo] HTTP listening on :8080
GET /api/articles list
GET /api/articles/:id get
POST /api/articles create
PATCH /api/articles/:id update
DELETE /api/articles/:id delete
WS /api/articles/live subscribe
GET /api/authors list
GET /api/authors/me me
WS /api/authors/live subscribe
GET /api/tags list
GET /api/comments list
POST /api/comments create
WS /api/comments/live subscribe
... (and the rest)
GET / ui.article-list
GET /article/:slug ui.article-detail
GET /tag/:slug ui.article-list (filtered)
GET /admin ui.article-list (role: Admin)
```
## Exercise the REST API
```bash
# Create an author (requires admin session — see auth docs; stub'd here for brevity)
$ curl -X POST localhost:8080/api/authors \
-H "Content-Type: application/json" \
-H "Authorization: Bearer $ADMIN_TOKEN" \
-d '{"email":"alice@example.com","handle":"alice","display":"Alice","role":"Author"}'
{"id":2,"email":"alice@example.com","handle":"alice",...}
# Create an article as that author
$ curl -X POST localhost:8080/api/articles \
-H "Authorization: Bearer $ALICE_TOKEN" \
-d '{
"slug": "hello",
"title": "Hello, writeonce",
"author": 2,
"meta": {"excerpt":"First post","body_md":"# Hi\n\nHello."},
"published": true
}'
{"id":1,"slug":"hello","title":"Hello, writeonce","published_at":"2026-04-17T12:00:00Z",...}
# List published articles (public — no token)
$ curl localhost:8080/api/articles
[{"id":1,"slug":"hello","title":"Hello, writeonce",...}]
# Filter by tag (via the query layer)
$ curl 'localhost:8080/api/articles?tags.slug=rust'
[...]
```
## Subscribe to live updates
```bash
$ websocat ws://localhost:8080/api/articles/live?published=eq.true
{"kind":"snapshot","rows":[{"id":1,"slug":"hello",...}]}
# Now in another terminal, update article 1. The open socket receives:
{"kind":"update","id":1,"old":{"title":"Hello, writeonce"},"new":{"title":"Hello!"}}
```
No polling. The subscription predicate was registered at connect time; the engine's commit path emits the delta directly.
## Generate a Go client
```bash
$ wo gen sdk --lang go --out ./client
[wo] reading types from ./types/
[wo] writing ./client/sdk.go (4 types, 16 endpoints, 4 subscriptions)
```
Use it:
```go
import "github.com/you/blog/client"
c, _ := client.Connect(ctx, "wo://localhost:8080", client.WithToken(token))
// Typed query
articles, _ := c.Articles.List(ctx, client.Where{Published: ptr(true)})
// Typed subscription — deltas arrive on a channel
sub, _ := c.Articles.Subscribe(ctx, client.Where{Published: ptr(true)})
for d := range sub.C {
switch d.Kind {
case client.Insert:
fmt.Printf("new article: %s\n", d.Row.Title)
case client.Update:
fmt.Printf("updated: %s\n", d.Row.Slug)
}
}
```
## Run the tests
```bash
$ wo test
=== tests/article_test.wo ===
create and fetch by slug OK (3ms)
policy blocks public read of unpublished drafts OK (4ms)
graph traversal: related articles OK (7ms)
live subscription receives delta on commit OK (12ms)
PASS 4/4 tests, 0 failures (26ms)
```
Each `test` block runs against an isolated engine snapshot that's rolled back at the end — no setup/teardown code needed.
## Build a production binary
```bash
$ wo build --target linux-amd64 --out bin/blog
[wo] static binary: bin/blog (14 MB, database + HTTP + subscription engine embedded)
$ ./bin/blog
[wo] HTTP listening on :8080
```
One binary, no dependencies. Copy it to a server, run it, done. The database file lives in `./data/` relative to the binary; the WAL ensures crash safety ([Phase 3](../../runtime/database/03-inmemory-engine.md)).
## What to read next
- [`../../runtime/wo-language.md`](../../runtime/wo-language.md) — the user-facing language overview this project builds on
- [`../../runtime/database/02-wo-language.md`](../../runtime/database/02-wo-language.md) — the two-layer language spec (schema + query layers)
- [`../../runtime/database/06-lowcode-fullstack.md`](../../runtime/database/06-lowcode-fullstack.md) — the `##ui`/`##policy`/`##service`/`##app` block spec
- [`../../../prototypes/wo-db/`](../../../prototypes/wo-db/) — the C++ prototype that runs the query-layer subset today

View file

@ -0,0 +1,251 @@
# `ecommerce` — a sample writeonce e-commerce app
A storefront + checkout + live ops dashboard in **~300 lines of `.wo`**. Exercises the features that make `.wo` distinct from a plain REST app: **cross-paradigm ACID transactions**, **type-attached lifecycle triggers**, **link types with properties**, and a **live-updating operations table**.
> Like the [blog sample](../blog/), this is a **docs artifact** — illustrative `.wo` source showing the shape of a real `wo init`'d project. Toolchain specified in [`../../runtime/wo-language.md`](../../runtime/wo-language.md).
## What's here
| File | What it shows |
| --- | --- |
| [`types/product.wo`](./types/product.wo) | Relational scalars + embedded doc (`meta`, `inventory`) + computed field (`available`) + graph edge (`similar_to`) + inventory-low trigger |
| [`types/order.wo`](./types/order.wo) | Tagged union status, array-of-struct `line_items`, computed `total`, four lifecycle triggers setting timestamp columns atomically |
| [`types/customer.wo`](./types/customer.wo) | Role union + `multi Product via Purchase` (link with properties) + `backlink Order.customer` |
| [`types/purchase.wo`](./types/purchase.wo) | `link Customer -> Product` — a graph edge **type** with its own columns (`order`, `qty`, `unit_price`) |
| [`logic/checkout.wo`](./logic/checkout.wo) | The canonical cross-paradigm transaction: reserve inventory + insert order + create graph edge, atomic across all three engines |
| [`ui/admin_orders.wo`](./ui/admin_orders.wo) | **The live order-ops table** — role-gated, auto-subscribes, delta-in-place updates |
| [`ui/storefront.wo`](./ui/storefront.wo) | Customer-facing product list with live inventory |
| [`ui/order_tracker.wo`](./ui/order_tracker.wo) | Customer-facing order history, same live engine, policy-filtered source |
| [`app.wo`](./app.wo) | Route table, Admin/Ops bypass policy, idempotent `seed()` |
| [`tests/checkout_test.wo`](./tests/checkout_test.wo) | Three tests covering the atomic checkout, the abort-without-partial-state guarantee, and the live-subscription delta stream |
## Project layout
```
ecommerce/
├── wo.toml
├── app.wo
├── types/
│ ├── customer.wo
│ ├── product.wo
│ ├── order.wo
│ └── purchase.wo # link type — graph edge with properties
├── logic/
│ └── checkout.wo # transactional functions (fn … in txn snapshot)
├── ui/
│ ├── storefront.wo
│ ├── order_tracker.wo
│ └── admin_orders.wo # the live ops table
└── tests/
└── checkout_test.wo
```
## Run it
```bash
$ cd docs/examples/ecommerce
$ wo run
[wo] parsing: 10 files, 4 types + 1 link type, 3 ui screens, 4 fns
[wo] compiling schema: 3 sql tables, 2 doc collections, 2 graph edge types
[wo] starting runtime (engine: in-memory, data_dir: ./data, isolation: snapshot)
[wo] on startup: seed() — 1 customer, 2 products
[wo] HTTP listening on :8080
GET /api/products list
GET /api/products/:id get
WS /api/products/live subscribe
GET /api/orders list
GET /api/orders/:id get
WS /api/orders/live subscribe
GET /api/customers/:id get
GET /api/customers/me me
PATCH /api/customers/:id update
WS /api/customers/live subscribe
POST /api/fn/checkout fn checkout(customer, product, qty) -> Order
POST /api/fn/mark_paid fn mark_paid(order)
POST /api/fn/mark_shipped fn mark_shipped(order)
GET / ui.storefront
GET /product/:sku ui.product-detail
GET /orders ui.order-tracker
GET /admin/orders ui.admin-orders (Admin | Ops)
```
> **Runtime model.** The engine is a single-threaded event loop today ([Phase 2 concurrency](../../runtime/database/02-wo-language.md#concurrency-model)). Snapshot isolation is trivially correct because there are no concurrent writers — the checkout, mark_paid, and mark_shipped fns run sequentially even when fired in quick succession. The throughput ceiling is ~one core (plenty for the sample); sharding across independent engine processes is the horizontal-scale path.
## Exercise the cross-paradigm checkout
The `fn checkout(...)` in [`logic/checkout.wo`](./logic/checkout.wo) is the canonical Phase 2 test case: one transaction that mutates relational, document, and graph state atomically.
```bash
# Place an order — one HTTP call runs the whole BEGIN ... COMMIT block
$ curl -X POST localhost:8080/api/fn/checkout \
-H "Authorization: Bearer $CUSTOMER_TOKEN" \
-d '{"customer":1, "product":2, "qty":3}'
{"id":1, "status":"Pending", "total":5997, "line_items":[{...}], "placed_at":"..."}
# Verify inventory was reserved (not yet decremented)
$ curl localhost:8080/api/products/2
{"sku":"SKU-GIZMO", "inventory":{"on_hand":12, "reserved":3, "reorder_at":3}, "available":9, ...}
# Verify the graph edge was created in the same transaction
$ curl localhost:8080/api/customers/1/purchased
[{"target":{"sku":"SKU-GIZMO"}, "order":1, "qty":3, "unit_price":1999, "at":"..."}]
```
If the inventory check failed inside `checkout`, **none** of the above writes happen — the order isn't created, the reservation isn't made, and the graph edge doesn't exist. That atomicity is the whole point of building your own engine instead of stitching Postgres + Neo4j.
## Watch the live admin ops table
Open the admin orders UI in a browser:
```bash
$ open http://localhost:8080/admin/orders # authenticated as Admin or Ops
```
The page renders a table with the columns declared in [`ui/admin_orders.wo`](./ui/admin_orders.wo). Behind the scenes, the client runtime has opened one WebSocket to the engine's subscription endpoint:
```
WS /api/orders/live ? status!=Cancelled
```
Now, from another terminal, fire a sequence of state changes:
```bash
# 1. New customer places an order — admin table gains a row, highlighted for 2s
$ curl -X POST localhost:8080/api/fn/checkout -d '{"customer":2,"product":1,"qty":1}'
# 2. Payment webhook flips status Pending → Paid — row updates in place, paid_at fills in
$ curl -X POST localhost:8080/api/fn/mark_paid -d '{"order":2}'
# 3. Ops ships the order — status → Shipped, shipped_at fills in
$ curl -X POST localhost:8080/api/fn/mark_shipped -d '{"order":2}'
```
The browser table re-renders each row delta as it arrives, without a full list refetch. `status` cell swaps its pill colour; timestamp cells populate. No polling anywhere in the path — the deltas are emitted by the transaction coordinator on commit, routed through the subscription registry, and pushed down the socket ([Phase 4](../../runtime/database/04-client-api.md)).
A filtered subscription — the admin clicking the **"Ready to ship"** quick-filter — doesn't rebuild state client-side. It sends the new predicate to the server, which replies with a `SNAPSHOT` frame of just the matching rows, then streams deltas that match the new predicate. Also zero-polling.
## Generate a typed Go client
```bash
$ wo gen sdk --lang go --out ./client
[wo] reading types from ./types/ and fns from ./logic/
[wo] writing ./client/sdk.go (4 types, 13 endpoints, 4 fns, 3 subscriptions)
```
The generated client speaks the native wire protocol:
```go
import "myshop/client"
c, _ := client.Connect(ctx, "wo://localhost:8080", client.WithToken(token))
// Typed transactional function call
order, err := c.Checkout(ctx, client.CheckoutArgs{
Customer: 1,
Product: 2,
Qty: 3,
})
// Typed live subscription — same wire as the admin UI uses
sub, _ := c.Orders.Subscribe(ctx, client.Where{Status: client.Ne(client.Cancelled)})
for d := range sub.C {
switch d.Kind {
case client.Insert:
fmt.Printf("new order #%d from %s — $%.2f\n", d.Row.ID, d.Row.Customer.Name, float64(d.Row.Total)/100)
case client.Update:
fmt.Printf("order #%d → %s\n", d.Row.ID, d.Row.Status)
}
}
```
## Checkout from Go without codegen
For ad-hoc scripts, admin tools, or client paths not on the app's hot loop, send raw `.wo` DML with `client.Wo(...)`. The server parses the block exactly like `wo run` would — same parser, same transaction coordinator, same `RETURNING` alias table — so the **cross-paradigm checkout runs in one round trip**:
```go
import "go.writeonce.dev/wo"
c, _ := wo.Connect(ctx, "wo://localhost:8080", wo.WithToken(token))
// Same logic as fn checkout(), but authored at the Go call site.
// BEGIN SNAPSHOT ... COMMIT runs server-side; RETURNING aliases ($pid, $oid)
// thread from the SQL UPDATE/INSERT into the Cypher CREATE within the txn.
result, err := c.Wo(ctx, `
BEGIN SNAPSHOT;
UPDATE products
SET inventory.reserved = inventory.reserved + $qty
WHERE id = $pid AND available >= $qty
RETURNING id AS pid;
INSERT INTO orders (customer, status, line_items)
VALUES ($uid, 'Pending', [{product: $pid, qty: $qty, unit_price: $unit}])
RETURNING id AS oid;
MATCH (u:Customer {id: $uid}), (p:Product {id: $pid})
CREATE (u)-[:PURCHASED {order: $oid, qty: $qty, unit_price: $unit}]->(p);
COMMIT;
`, wo.Params{"uid": 1, "pid": 2, "qty": 3, "unit": 4999})
if err != nil { log.Fatal(err) }
orderID := result.Aliases["oid"].(int64)
fmt.Printf("created order #%d\n", orderID)
```
**When to reach for this form** — see the [raw-vs-typed guidance in the Go SDK doc](../../runtime/database/05-go-sdk.md#when-to-use-raw-wo-vs-typed-codegen). Rule of thumb: typed `c.Checkout(...)` for the app's storefront; raw `c.Wo(...)` for an ops console that runs a custom report, or when you want to paste a block from [`logic/checkout.wo`](./logic/checkout.wo) straight into Go.
## Run the tests
```bash
$ wo test
=== tests/checkout_test.wo ===
checkout atomically reserves inventory, creates order, and creates graph edge OK (8ms)
checkout aborts without partial state when inventory is insufficient OK (4ms)
admin live-orders subscription receives deltas across the order lifecycle OK (14ms)
PASS 3/3 tests, 0 failures (26ms)
```
The third test is the important one for the docs: it proves that the same engine that serves `/admin/orders` in the browser delivers deltas in commit order through a programmatic `subscribe live` handle. One engine, one delta stream, two consumers (the browser and the test).
## Build a production binary
```bash
$ wo build --target linux-amd64 --out bin/shop
[wo] static binary: bin/shop (15 MB — database + HTTP + subscription engine embedded)
$ ./bin/shop
[wo] HTTP listening on :8080
```
Drop the binary on a server, give it a writable directory for `./data/` (WAL + engine state), run it behind nginx or let it terminate TLS itself. The admin ops table works on the first page-load — no Redis, no Kafka, no separate DB process, no ORM-and-migration dance.
## Compare to the blog sample
Both projects use the same language and runtime. They showcase different slices:
| Feature | [blog](../blog/) | ecommerce (this project) |
| --- | --- | --- |
| Embedded document | `article.meta` | `product.meta`, `product.inventory` |
| Graph edges (zero-prop) | tags, related | similar_to |
| Graph edges **with** properties | — | `type Purchase link Customer -> Product` |
| Tagged union | — | `Pending \| Paid \| Shipped \| ...` |
| Computed field | `word_count` | `available`, `total` (sum over line_items) |
| Array-of-struct column | — | `line_items: [{product, qty, unit_price}]` |
| Stored procedure (`fn ... in txn`) | seed only | full checkout + fulfillment |
| Cross-paradigm transaction | — | `checkout` (relational + doc + graph atomic) |
| Live subscription | list view | live **ops** table with in-place delta updates |
| Row-level policy | draft hiding | customer sees own orders; ops/admin sees all |
If the blog shows **what a CRUD app looks like in `.wo`**, the ecommerce sample shows **what a transactional business app looks like in `.wo`** — and why building the engine as part of the language is the differentiator.
## What to read next
- [`../../runtime/wo-language.md`](../../runtime/wo-language.md) — language overview and toolchain
- [`../../runtime/database/02-wo-language.md`](../../runtime/database/02-wo-language.md) — the schema/query two-layer spec
- [`../../runtime/database/04-client-api.md`](../../runtime/database/04-client-api.md) — the wire protocol and subscription engine behind the live ops table
- [`../../runtime/database/06-lowcode-fullstack.md`](../../runtime/database/06-lowcode-fullstack.md) — `##ui` / `##app` block spec
- [`../../../prototypes/wo-db/tests/checkout.wo`](../../../prototypes/wo-db/tests/checkout.wo) — the C++ prototype's smoke test that exercises the same cross-paradigm transaction at the query layer

View file

@ -47,7 +47,7 @@ Per [06-markdown-render.md](./06-markdown-render.md), the JSON metadata is minim
### Mapping Types
| Type | Meaning | Agent Use |
|------|---------|-----------|
| -------------- | ----------------------------------------------- | --------------------------------------------------------------------------------------- |
| `related` | Topically related articles | Agent loads these for cross-reference when editing |
| `prerequisite` | Articles the reader should read first | Agent ensures no concept duplication, references prerequisites instead of re-explaining |
| `series` | Articles that form an ordered sequence | Agent maintains narrative continuity across the series |
@ -76,6 +76,7 @@ content/
The author asks an agent: "Write an article about deploying GitLab Runner on ECS."
The agent:
1. Scans the content directory for existing articles with tags `gitlab`, `ci-cd`, `aws`
2. Finds `gitlab-runner-with-kubernetes-executor` and `auto-scale-gitlab-runner-using-aws-spot-instance`
3. Reads their `.md` files to understand what's already covered
@ -87,6 +88,7 @@ The agent:
The author asks: "Update the Kubernetes executor article with the new runner token format."
The agent:
1. Reads the article's JSON metadata and `.md` content
2. Reads the `mappings.related` articles to check for consistency
3. Makes the update in the `.md` file
@ -98,6 +100,7 @@ The agent:
The author asks: "Which articles reference outdated AWS configurations?"
The agent:
1. Loads all article metadata (the Store already indexes everything)
2. Follows `mappings` to build a dependency graph
3. Reads the `.md` files of articles tagged with `aws`
@ -109,6 +112,7 @@ The agent:
The author asks: "Add a new part to the gitlab-runner series."
The agent:
1. Finds all articles with `mappings.series.name == "gitlab-runner"`
2. Reads them in order to understand the narrative arc
3. Writes the new article continuing from where the series left off
@ -140,9 +144,7 @@ The `article.htmlx` template can render related articles:
```html
<article>
<h1>{{article.title}}</h1>
{{article.content_html}}
{{#each article.related}}
{{article.content_html}} {{#each article.related}}
<aside class="related">
<h3>Related</h3>
<ul>
@ -173,3 +175,8 @@ For agents to use the mappings effectively, the project can include an agent ins
```
This turns the content directory into an agent-navigable knowledge graph where the metadata provides the edges and the markdown files provide the nodes.
## queryable graph database
- traversable knowledge graphs available on RAM.
- which linux kernels, develop in C ++. User wants to learn it.

View file

@ -0,0 +1,95 @@
# 02 — Event Loop on `epoll`
**Context sources:** [`../01-problem.md`](../01-problem.md), [`../02-recovery.md`](../02-recovery.md), [`./linux/00-linux.md`](./linux/00-linux.md), [`./done/01-scafolding-crates.md`](./done/01-scafolding-crates.md).
## Goal
Land a hand-rolled single-threaded event loop inside `crates/rt/` that wraps Linux's file-descriptor primitives directly — **without** touching `tokio`, `axum`, or any other async runtime crate. This is the foundation every later phase builds on: phase 03 puts an HTTP server on top of it, phase 04 retires tokio + axum, phase 07 registers `inotify` watches on it, phase 08 drives `sendfile` through it.
Nothing is removed in this phase. The module sits alongside the tokio-backed axum server, unused by the `wo` binary until phase 04 flips the switch.
## Design decisions (locked)
1. **`epoll`, not `io_uring`, on day one.** `epoll` is ubiquitous (Linux 2.6+), well-understood, and every primitive we need (eventfd, timerfd, signalfd, inotify, accepted sockets) already integrates with it via `epoll_ctl`. `io_uring` is a natural follow-on phase once the event-loop abstraction exists — [`00-linux.md`](./linux/00-linux.md) calls it out for that role.
2. **Single-threaded, edge-triggered.** Matches [02-wo-language.md § Concurrency Model](../runtime/database/02-wo-language.md#concurrency-model). Every fd registered with `EPOLLET`; the loop reads until `EAGAIN`. No worker pool, no cross-thread state.
3. **`libc` is the only new dependency.** `libc = "0.2"` added to `crates/rt/Cargo.toml`. No `nix`, no `mio`. Direct `unsafe extern "C"` calls against the kernel surface.
4. **Module, not crate (yet).** Lives at `crates/rt/src/event/` so phase 03 can call into it cheaply. Extraction to the empty `crates/event/` sibling is deferred until a second caller appears outside `rt` — likely when [`sub`](../../crates/sub/) starts consuming the loop for subscription delivery.
## Scope
### New files inside `crates/rt/src/event/`
| File | Responsibility | Port source |
| --- | --- | --- |
| `mod.rs` | Re-exports `EventLoop`, `Event`, `Interest`, `Token`, `EventFd`, `TimerFd`, `SignalFd` | [`reference/crates/wo-event/src/lib.rs`](../../reference/crates/wo-event/src/lib.rs) (9 LOC) |
| `epoll.rs` | `EventLoop { fd, events }` — `new()`, `register(raw_fd, interest, token)`, `wait_once(timeout) -> &[Event]`, `deregister(raw_fd)` | [`reference/crates/wo-event/src/epoll.rs`](../../reference/crates/wo-event/src/epoll.rs) (183 LOC) |
| `eventfd.rs` | `EventFd { fd }` — counter semaphore for cross-fd wake-up (subscription dispatch, shutdown signal) | [`reference/crates/wo-event/src/eventfd.rs`](../../reference/crates/wo-event/src/eventfd.rs) (66 LOC) |
| `timerfd.rs` | `TimerFd { fd }` — oneshot + periodic timers as fds for the loop | [`reference/crates/wo-event/src/timerfd.rs`](../../reference/crates/wo-event/src/timerfd.rs) (91 LOC) |
| `signalfd.rs` | `SignalFd { fd }` — SIGINT / SIGTERM / SIGHUP delivered as fd reads for graceful shutdown without a tokio signal handler | [`reference/crates/wo-event/src/signalfd.rs`](../../reference/crates/wo-event/src/signalfd.rs) (62 LOC) |
Total: ~410 LOC lifted and adapted. The v1 code already compiles standalone in `reference/crates/wo-event/` and has unit tests; the port is near-verbatim plus namespace cleanups.
### `Cargo.toml` change
```toml
[dependencies]
anyhow = "1"
serde = { version = "1", features = ["derive"] }
serde_json = "1"
tokio = { version = "1", features = ["rt", "macros", "net", "signal", "sync", "time"] }
axum = "0.7"
tower = "0.4"
libc = "0.2" # NEW — see docs/plan/02-event-loop-epoll.md
```
## API shape (target — validate against v1 when porting)
```rust
use rt::event::{EventLoop, EventFd, Interest, Token};
let mut loop_ = EventLoop::new()?;
let ev = EventFd::new()?;
loop_.register(ev.as_raw_fd(), Interest::READABLE, Token(0))?;
ev.write(1)?; // wake the loop from another flow
for event in loop_.wait_once(Some(Duration::from_millis(100)))? {
match event.token() {
Token(0) => { let n = ev.read()?; /* ... */ }
_ => unreachable!(),
}
}
```
## Exit criteria
1. `cargo build` at root compiles cleanly.
2. A new unit test in `crates/rt/src/event/epoll.rs`:
- create an `EventLoop`,
- register an `EventFd`,
- `write(1)` to the eventfd from the same thread,
- `wait_once(timeout)` returns an `Event` for the correct token,
- `read()` on the eventfd returns `1`.
3. A second unit test validates `TimerFd::oneshot(100ms)` fires within a `wait_once(500ms)` window.
4. All 14 existing `rt` tests still pass. `cargo run --bin wo -- run docs/examples/blog` still serves (tokio path unchanged).
5. `cd reference/crates && cargo build && cargo test` still green (nothing touched).
## Non-scope
- **No cutover.** The `wo` binary keeps calling `tokio::runtime::Builder::new_current_thread()`. That happens in phase 04.
- **No HTTP.** Accepting connections is phase 03's problem. This phase is pure kernel-primitive plumbing.
- **No subscription dispatch.** The `sub` crate doesn't exist yet as real code; phase 07 (inotify) is the first real loop consumer after phase 03.
- **No crate extraction.** Stays at `crates/rt/src/event/`. Pulling to `crates/event/` waits for a second consumer.
- **No `io_uring`.** Separate follow-on once the abstraction solidifies.
## Verification
```bash
cargo build
cargo test --lib event # new tests in crates/rt/src/event/
cargo test --lib # all 14 existing + new epoll/eventfd/timerfd tests green
cargo run --bin wo -- run docs/examples/blog # axum path unchanged, still serves
cd reference/crates && cargo build && cargo test # v1 untouched
```
## After this phase
Phase 03 puts a non-blocking HTTP/1.1 listener on top of the `EventLoop` and proves end-to-end I/O without tokio. The two phases together give phase 04 everything it needs to delete the tokio + axum dependencies.

View file

@ -0,0 +1,100 @@
# 03 — Hand-Rolled HTTP/1.1
**Context sources:** [`./02-event-loop-epoll.md`](./02-event-loop-epoll.md), [`./linux/00-linux.md`](./linux/00-linux.md), [`../02-recovery.md`](../02-recovery.md).
## Goal
A non-blocking HTTP/1.1 server module that accepts connections, parses requests, and writes responses, **driven by the phase-02 `EventLoop`** — no axum, no hyper, no tokio. Still additive: the existing axum router keeps serving `wo run` until phase 04 cuts over.
## Design decisions (locked)
1. **HTTP/1.1 only, keep-alive supported.** HTTP/2 and HTTP/3 are not on the roadmap for Stage 2 — they need ALPN / TLS support we don't have a plan for yet. HTTP/1.1 covers every endpoint the blog + ecommerce samples exercise.
2. **Per-connection state machine.** Each accepted socket fd is registered on the event loop with its own `Connection { state: Reading | Writing | Idle, parser, pending_response }`. Edge-triggered `EPOLLIN`/`EPOLLOUT` drive state transitions. Matches [v1 wo-http](../../reference/crates/wo-http/src/connection.rs)'s model verbatim.
3. **Router is pattern-matched at registration.** `Router::new().route("/api/articles/:id", Method::GET, handler)` resolves to a trie at boot. Per-request dispatch is a single trie walk — no axum-style type-erased layers.
4. **Handlers are `fn(&Request, &Engine) -> Response`.** Synchronous. The single-threaded event loop means a handler blocking is a bug; each handler must be a pure transformation over engine state.
5. **Module, not crate (yet).** Lives at `crates/rt/src/http/` with the same "extract when a second consumer shows up" rule as phase 02. The eventual home is the empty [`crates/http/`](../../crates/http/) sibling — but not in this phase.
## Scope
### New files inside `crates/rt/src/http/`
| File | Responsibility | Port source |
| --- | --- | --- |
| `mod.rs` | Re-exports `Listener`, `Connection`, `Request`, `Response`, `Router`, `Method`, `Status` | [`reference/crates/wo-http/src/lib.rs`](../../reference/crates/wo-http/src/lib.rs) (4 LOC) |
| `listener.rs` | `Listener { fd }` wrapping `socket + bind + listen + accept4(SOCK_NONBLOCK \| SOCK_CLOEXEC)`; integrates with `EventLoop` | [`reference/crates/wo-http/src/listener.rs`](../../reference/crates/wo-http/src/listener.rs) (202 LOC) |
| `connection.rs` | Per-fd state machine: drain request bytes, parse, dispatch, drain response bytes, keep-alive or close | [`reference/crates/wo-http/src/connection.rs`](../../reference/crates/wo-http/src/connection.rs) (202 LOC) |
| `request.rs` | Incremental HTTP/1.1 request parser: request line, headers, optional body. `Content-Length` only (no chunked request bodies in Stage 2 — they don't appear in the samples) | [`reference/crates/wo-http/src/request.rs`](../../reference/crates/wo-http/src/request.rs) (158 LOC) |
| `response.rs` | Response builder + writer: status line, headers, body (fixed or chunked) | [`reference/crates/wo-http/src/response.rs`](../../reference/crates/wo-http/src/response.rs) (110 LOC) |
| `route.rs` | Trie-based router: static paths + `:param` segments. `Router::route(method, path, handler) -> Router` | [`reference/crates/wo-route/src/router.rs`](../../reference/crates/wo-route/src/router.rs) (127 LOC) + [`pattern.rs`](../../reference/crates/wo-route/src/pattern.rs) (146 LOC) |
Total: ~949 LOC ported. Most of it is mechanical adaptation from v1; the namespace + the `Interest` enum change from phase 02 are the only non-trivial edits.
### `Cargo.toml` change
None. `libc` already in from phase 02 covers the raw syscalls.
## API shape (target)
```rust
use rt::event::EventLoop;
use rt::http::{Listener, Router, Method, Status, Response};
let mut loop_ = EventLoop::new()?;
let router = Router::new()
.route(Method::GET, "/healthz", |_req, _eng| Response::ok().body("ok"))
.route(Method::GET, "/api/articles", list_articles)
.route(Method::GET, "/api/articles/:id", get_article)
.route(Method::POST,"/api/articles", create_article);
let listener = Listener::bind("127.0.0.1:8080")?;
loop_.register(listener.as_raw_fd(), Interest::READABLE, Token::LISTENER)?;
let mut conns: HashMap<RawFd, Connection> = HashMap::new();
loop {
for ev in loop_.wait_once(None)? {
match ev.token() {
Token::LISTENER => {
while let Some(stream) = listener.accept_nonblocking()? {
let fd = stream.as_raw_fd();
loop_.register(fd, Interest::READABLE, Token::CONN(fd))?;
conns.insert(fd, Connection::new(stream));
}
}
Token::CONN(fd) => {
conns.get_mut(&fd).unwrap().drive(ev, &router, &engine)?;
if conns[&fd].is_closed() { conns.remove(&fd); }
}
_ => {}
}
}
}
```
## Exit criteria
1. A test binary `crates/rt/src/bin/http-smoke.rs` (`[[bin]] name = "http-smoke"` in `rt/Cargo.toml`) binds on `127.0.0.1:0` (auto-assigned port), registers routes for `/healthz`, `/echo/:name`, `/counter`, and services them via the phase-02 event loop.
2. An integration test (also in `crates/rt/tests/http_smoke.rs` or similar) spawns the binary, sends three `curl` equivalents using `std::net::TcpStream`, validates status codes and bodies.
3. All 14 existing `rt` tests still pass.
4. `wo run docs/examples/blog` unchanged — axum path still drives the real CLI.
5. `cargo build` at root; `cd reference/crates && cargo build` still green.
## Non-scope
- **No TLS.** Deferred. When it lands, it's a wrapper around `Connection` that `read`/`write`s through `rustls` or (ideally) kTLS. Not this phase.
- **No HTTP/2.** See design decision 1.
- **No chunked request bodies.** Every sample's `POST /api/X` uses `Content-Length`. If a future sample needs chunked, it's a small extension to `request.rs`.
- **No middleware.** axum's `tower::Layer` idiom has no direct analog. Cross-cutting concerns (logging, auth) live in the handler or in a wrapper fn — phase 04 re-integrates with the existing axum state handling when the cutover happens.
- **Not wired into the `wo` binary yet.** That's phase 04.
## Verification
```bash
cargo build # root workspace compiles
cargo test --bin http-smoke # the bundled test binary
cargo test --lib # 14 existing rt tests still green
cargo run --bin wo -- run docs/examples/blog # axum path unchanged
```
## After this phase
Phase 04 takes the same in-memory `Engine` that the axum router serves and points the phase-03 router at it instead. Removing `tokio`, `axum`, `tower` is a consequence; the behaviour visible to `reference/rest/blog.rest` does not change.

View file

@ -0,0 +1,119 @@
# 04 — Cutover: Remove tokio, axum, tower
**Context sources:** [`./02-event-loop-epoll.md`](./02-event-loop-epoll.md), [`./03-hand-rolled-http.md`](./03-hand-rolled-http.md), [`../01-problem.md`](../01-problem.md).
## Goal
Flip the `wo` binary off the tokio + axum stack and onto the phase-02 event loop + phase-03 HTTP server. Delete three dependencies from `crates/rt/Cargo.toml`. REST behaviour visible to [`reference/rest/blog.rest`](../../reference/rest/blog.rest) does not change — same status codes, same response bodies, same endpoint paths.
This is the first phase where the dependency count goes *down*. Phases 02 and 03 were additive; this one is the switch.
## Design decisions (locked)
1. **Atomic swap, single commit.** Don't run tokio and the new loop in parallel in production. Flip the binary's `main()` in one change. Phase 03 already gave us confidence the new stack works end-to-end via `http-smoke`.
2. **Preserve the `Engine` trait surface.** `Arc<Mutex<Engine>>` stays exactly as `crates/rt/src/engine.rs` has it today. The routing layer in `crates/rt/src/server.rs` — the function that maps `service rest` blocks to axum `MethodRouter` — gets rewritten to emit phase-03 `Router::route(...)` calls instead. Same data flow, different transport.
3. **No tokio — no async.** Handlers become synchronous `fn(&Request, &Engine) -> Response`. The single-threaded event loop [already assumes this](../runtime/database/02-wo-language.md#concurrency-model); removing `async fn` plumbing simplifies the code. `tokio::sync::Mutex` becomes `std::sync::Mutex` (fine in a single-threaded loop since lock contention is impossible).
4. **`signalfd` replaces `tokio::signal::ctrl_c()`.** Registered as another fd on the loop; reading a SIGINT cleanly exits the loop and closes outstanding connections.
5. **`WO_LISTEN` env var semantics unchanged.** The `127.0.0.1:8080` default + the `WO_LISTEN=...` override stays exactly as today. Operators don't notice the change.
## Scope
### Files rewritten inside `crates/rt/`
| File | Change | Notes |
| --- | --- | --- |
| `src/bin/wo.rs` | Replace `tokio::runtime::Builder::new_current_thread` + `axum::serve` with `EventLoop` + `Listener` + `Router` wiring | The `run()` / `serve()` fns fuse into a single synchronous `run()` that drives the loop |
| `src/server.rs` | Replace `axum::Router` construction + axum handlers (`async fn list_h(State(st): State<TypeState>) -> impl IntoResponse`) with phase-03 `Router::route(...)` + sync handlers | Handler bodies are otherwise untouched: `engine.lock().list(&ty).map(Json)` logic flows through |
| `src/engine.rs` | `tokio::sync::Mutex` → `std::sync::Mutex`; `.lock().await` → `.lock().unwrap()` | Only the wrapper changes; row logic intact |
### `Cargo.toml` delta
```diff
[dependencies]
anyhow = "1"
serde = { version = "1", features = ["derive"] }
serde_json = "1"
-tokio = { version = "1", features = ["rt", "macros", "net", "signal", "sync", "time"] }
-axum = "0.7"
-tower = "0.4"
libc = "0.2"
```
Three deps gone. Four remaining: `anyhow`, `serde`, `serde_json`, `libc`.
### Files deleted
- None. The phase-02 `event/` module and phase-03 `http/` module stay in place and now become the primary code path.
## Handler signature change
**Before (axum + tokio):**
```rust
async fn list_h(State(st): State<TypeState>) -> impl IntoResponse {
let eng = st.engine.lock().await;
match eng.list(&st.ty) {
Ok(rows) => (StatusCode::OK, Json(json!(rows))).into_response(),
Err(e) => (StatusCode::INTERNAL_SERVER_ERROR, e.to_string()).into_response(),
}
}
```
**After (sync, event loop):**
```rust
fn list_h(_req: &Request, st: &TypeState) -> Response {
let eng = st.engine.lock().unwrap();
match eng.list(&st.ty) {
Ok(rows) => Response::ok().json(&serde_json::json!(rows)),
Err(e) => Response::status(Status::INTERNAL_SERVER_ERROR).body(e.to_string()),
}
}
```
Twelve handlers total — one pair per `{list, get, create, update, delete}` × four types. Mechanical rewrite.
## Exit criteria
1. **`cargo build`** at root — compiles with four deps (not seven).
2. **`cargo test --lib`** — all 14 existing `rt` unit tests still pass. A new test in `src/server.rs` exercises the router build from a compiled catalog (no HTTP, just static registration).
3. **End-to-end REST smoke — the 20-assertion battery from [`reference/rest/blog.rest`](../../reference/rest/blog.rest)** must pass byte-identical to Stage 2 today. Script:
```bash
WO_LISTEN=127.0.0.1:8765 cargo run --bin wo -- run docs/examples/blog &
# ... curl each block, check expected status
```
4. **Graceful shutdown.** SIGINT on the process exits cleanly (no panic, no orphan fds). Validate with `strace -f -e signalfd4,close` on shutdown.
5. **`cd reference/crates && cargo build && cargo test`** still green.
6. **Dep audit.** `cargo tree -p rt --depth 1` shows `libc` as the only non-transitive external dep beyond `anyhow`, `serde`, `serde_json`.
## Non-scope
- **No JSON replacement.** `serde` + `serde_json` are still imported and used. Phase 05 removes them.
- **No inotify / sendfile.** Stage 3 capabilities. Phases 07 and 08.
- **No `io_uring`.** The `EventLoop` keeps using `epoll` here; swapping is a later phase.
- **No crate extraction.** `event/` and `http/` stay inside `crates/rt/src/`. The empty `crates/event/` and `crates/http/` sibling crates wait for second consumers.
## Risk
The `.rest` files are the safety net — 20 assertions that every Stage 2 endpoint returns the expected status. If one breaks after the cutover, the fix is almost always in the handler rewrite (sync semantics + the new `Response::json(...)` helper). No transport-layer regression should survive phase 03's `http-smoke` passing.
## Verification
```bash
cargo build # 4 deps, no tokio/axum/tower
cargo test --lib # 14 + any new server.rs tests green
# end-to-end
WO_LISTEN=127.0.0.1:8765 cargo run --bin wo -- run docs/examples/blog &
PID=$!
sleep 2
# every block in reference/rest/blog.rest, via curl, checking %{http_code}
# (copy-paste the 20-assertion script from the Stage 2 turn that verified blog.rest)
kill $PID
cd reference/crates && cargo build && cargo test # v1 untouched
```
## After this phase
Phase 05 removes `serde` + `serde_json` by writing a minimal JSON parser + emitter against the new `http::Response::json()` surface. At the end of phase 06, `crates/rt/Cargo.toml` is down to `libc` alone — the stated end goal.

View file

@ -0,0 +1,114 @@
# 05 — Hand-Rolled JSON
**Context sources:** [`./04-cutover-remove-tokio-axum.md`](./04-cutover-remove-tokio-axum.md), [`../../prototypes/wo-db/src/value.hpp`](../../prototypes/wo-db/src/value.hpp).
## Goal
Replace `serde_json::Value` / `serde_json::Map` with a hand-rolled `Value` type covering exactly the shapes the runtime reads and writes: request bodies, response bodies, and the `Engine`'s in-memory `Row`. Remove `serde` + `serde_json` from `crates/rt/Cargo.toml`. After this phase the dep list is `anyhow` + `libc`.
## Design decisions (locked)
1. **Minimal surface.** The runtime's actual JSON needs are small:
- Parse request body bytes → `Value::Object` (single top-level object on every sample endpoint).
- Emit `Value::Object` / `Value::Array` → bytes for the response.
- Pretty-printing is **not** required. Operators reach for `| python3 -m json.tool` if they want it.
2. **No `#[derive(Serialize/Deserialize)]`.** `Value` is the union type; every `Row`, `Product`, `Article` is already a `Value::Object` at the boundary. The only thing that "serializes" is `Value`. The ecosystem of derive-based types doesn't exist in `rt` today — `engine::Row` is `HashMap<String, Value>` via `serde_json` today, becomes `HashMap<String, Value>` via the new module tomorrow.
3. **RFC 8259 compliant, but strict.** No unquoted keys, no trailing commas, no comments. Standard JSON. The runtime isn't serving JSON5.
4. **Parser is recursive descent, zero-copy where possible.** String values borrow from the input buffer unless they contain escapes; objects own their keys. Preserves the "no heavy abstraction" pattern of phase 02 / 03.
5. **Module at `crates/rt/src/json/`.** Same "extract when a second consumer appears" rule. Eventual home is the empty [`crates/value/`](../../crates/value/) sibling — not this phase.
## Scope
### New files inside `crates/rt/src/json/`
| File | Responsibility | Approx LOC |
| --- | --- | --- |
| `mod.rs` | Re-exports `Value`, `Object`, `Array`, `parse`, `emit` | ~10 |
| `value.rs` | `pub enum Value { Null, Bool(bool), Int(i64), Float(f64), Str(String), Array(Vec<Value>), Object(BTreeMap<String, Value>) }` + `impl Value` helpers (`as_str`, `as_i64`, `get`, indexing) | ~200 |
| `parse.rs` | `parse(&[u8]) -> Result<Value, ParseError>` — recursive descent: `parse_value` → `parse_object` / `parse_array` / `parse_string` / `parse_number` / `parse_keyword`. Single-pass, no backtracking. | ~300 |
| `emit.rs` | `emit(value: &Value, buf: &mut Vec<u8>)` — iterative-ish writer, escapes strings per RFC 8259 §7 | ~150 |
Total: ~660 LOC. No v1 precedent — no sample parser to port. Reference the target shape against [`prototypes/wo-db/src/value.hpp`](../../prototypes/wo-db/src/value.hpp) for the `Value` variants (same six kinds as the C++ prototype, minus `Float` which that prototype folds into `Int` but we need for HTTP request bodies like `{"qty": 2.5}`).
### `Cargo.toml` delta
```diff
[dependencies]
anyhow = "1"
-serde = { version = "1", features = ["derive"] }
-serde_json = "1"
libc = "0.2"
```
### Consumers to update
Search: `rg 'serde_json|serde::' crates/rt/src | wc -l` — expected ~20 call sites. Each is a mechanical swap:
| Current | After |
| --- | --- |
| `serde_json::json!({"key": value})` | `json::Value::Object(…)` or a small `json!` macro we ship |
| `serde_json::Value` | `json::Value` |
| `serde_json::Map<String, Value>` | `BTreeMap<String, json::Value>` (the runtime already uses BTreeMap for stable order) |
| `serde_json::from_slice::<Value>(&bytes)?` | `json::parse(&bytes)?` |
| `Json(json!(rows)).into_response()` | `Response::ok().json_body(&rows)` (new helper on phase-03 `Response`) |
| `#[derive(Serialize, Deserialize)]` on any `rt` struct | deleted — no consumer after this phase |
The biggest consumer is `crates/rt/src/engine.rs` — `Row` is a `serde_json::Map<String, Value>` today. It becomes `BTreeMap<String, json::Value>`. The `eval_default()` function's `json!(n)` / `json!(b)` calls become `Value::Int(n)` / `Value::Bool(b)`. Minor, all local.
## A compact `json!` macro (for ergonomics)
Without `serde_json::json!`, the most-used construction pattern (`json!({"runtime": "wo", "stage": 2})`) gets verbose. Ship a minimal macro:
```rust
#[macro_export]
macro_rules! json {
(null) => ($crate::json::Value::Null);
(true) => ($crate::json::Value::Bool(true));
(false) => ($crate::json::Value::Bool(false));
([$($e:tt),* $(,)?]) => (
$crate::json::Value::Array(vec![$($crate::json!($e)),*])
);
({$($k:tt : $v:tt),* $(,)?}) => ({
let mut m = std::collections::BTreeMap::new();
$( m.insert(stringify!($k).trim_matches('"').to_string(), $crate::json!($v)); )*
$crate::json::Value::Object(m)
});
($e:expr) => ($crate::json::Value::from($e));
}
```
Covers 95% of current `serde_json::json!(...)` uses in the codebase. For the other 5%, build `Value` by hand.
## Exit criteria
1. **`cargo build`** — compiles with three deps (`anyhow`, `libc` + `Cargo.toml` itself).
2. **New tests in `crates/rt/src/json/`:**
- `parse_object_simple` — `{"a":1,"b":"x"}` round-trips.
- `parse_nested_and_array` — `{"xs":[1,2,3],"meta":{"k":"v"}}` round-trips.
- `parse_escapes` — `"\\n\\t\\\"\\u0041"` → `"\n\t\"A"`.
- `emit_stable_key_order` — emitting a `BTreeMap`-backed object produces keys in sorted order (matters for `.rest` expected-body stability).
- `parse_errors` — unterminated string, trailing comma, missing comma, unclosed object all return `ParseError` with line/col.
3. **All 14 existing `rt` tests pass** after the swap (the `engine::Engine` and `server::*` tests most affected).
4. **`reference/rest/blog.rest`** — 20 assertions all return the same HTTP status AND the same response body shape (may differ in key ordering if `BTreeMap` ordering differs from `serde_json`'s insertion order — document the shift).
5. **Dep audit.** `cargo tree -p rt --depth 1` shows zero `serde*` lines.
## Non-scope
- **No streaming parse.** Request bodies are small (< 1 MB on every sample endpoint). A buffered full-body parse is fine.
- **No `serde_json`-compat feature flag.** Clean break; this is the only consumer that matters, and we control it.
- **No JSON Pointer, no JSON Schema, no JSON Patch.** If needed later, layer on top.
- **No crate extraction to `crates/value/`.** Same rule as phase 02/03: wait for a second consumer.
## Verification
```bash
cargo build # three deps
cargo test --lib json # new parser/emitter tests
cargo test --lib # 14 existing tests still green
# full .rest smoke — same script as phase 04 exit criterion 3
cd reference/crates && cargo build && cargo test
```
## After this phase
Two deps left: `anyhow` and `libc`. Phase 06 removes `anyhow`. After that, `libc` is the only external crate — the stated end goal.

View file

@ -0,0 +1,120 @@
# 06 — Bespoke Error Type
**Context sources:** [`./05-hand-rolled-json.md`](./05-hand-rolled-json.md), [`../01-problem.md`](../01-problem.md).
## Goal
Replace `anyhow::Error` / `anyhow::Result` with a single `Error` enum rooted in `crates/rt/src/error.rs`. Remove `anyhow` from `crates/rt/Cargo.toml`. At the end of this phase, `[dependencies]` contains only `libc` — the **stated end goal** of this plan sequence.
## Design decisions (locked)
1. **One enum per crate.** `rt::Error` covers everything the runtime produces — parse errors, compile errors, engine errors, I/O errors, lex errors, HTTP framing errors, JSON parse errors (phase 05), signal-handler failures. The v1 crates use `std::io::Result` throughout, which is fine for I/O-heavy code but loses context for parse and compile failures. We split the difference with a tagged variant enum.
2. **`From` impls for standard errors.** `io::Error`, `ParseIntError`, `FromUtf8Error`, `Utf8Error` — automatically convert via `?`. Everything else wraps explicitly through constructors like `Error::parse(line, col, msg)`.
3. **`Display` composes a line + context.** No chained backtrace. Error messages stay compact: `parse error at line 42: expected ':', got '}'`. This is what every caller prints today after `anyhow::Error` formats.
4. **`Result<T>` alias.** Shorthand: `pub type Result<T> = core::result::Result<T, Error>`. Replaces `anyhow::Result<T>` every existing call site uses.
5. **No macros.** `anyhow::anyhow!("...")` becomes `Error::msg("...")`. `anyhow::bail!(...)` becomes `return Err(Error::msg(...))`. `anyhow::Context::context(err, "...")` becomes `err.with_context(|| "...")` via a tiny inherent method.
## Scope
### New file
| File | Responsibility | Approx LOC |
| --- | --- | --- |
| `crates/rt/src/error.rs` | `pub enum Error`, `pub type Result`, `From` impls, `Display`, `fn msg`, `fn parse`, `fn with_context` | ~120 |
### Enum shape (target)
```rust
#[derive(Debug)]
pub enum Error {
Io(io::Error),
Parse { line: u32, col: u32, message: String },
Lex { line: u32, col: u32, message: String },
Compile { message: String },
Engine { message: String },
Http { message: String },
Json { message: String },
NotFound { ty: String, id: i64 },
Config { message: String },
Msg (String), // anyhow::anyhow!-style catch-all
}
impl Error {
pub fn msg(s: impl Into<String>) -> Self { Error::Msg(s.into()) }
pub fn parse(line: u32, col: u32, m: impl Into<String>) -> Self
{ Error::Parse { line, col, message: m.into() } }
// ... one constructor per non-Io variant
pub fn with_context<F>(self, ctx: F) -> Error
where F: FnOnce() -> String
{
Error::Msg(format!("{}: {}", ctx(), self))
}
}
impl From<io::Error> for Error { ... }
impl From<ParseIntError> for Error { ... }
impl From<Utf8Error> for Error { ... }
impl From<FromUtf8Error> for Error { ... }
impl core::fmt::Display for Error { ... }
impl core::error::Error for Error {}
pub type Result<T> = core::result::Result<T, Error>;
```
### Call-site mechanical sweep
Find: `rg 'anyhow::|anyhow!|bail!|\.context\(' crates/rt/src | wc -l` — expected ~40 sites.
| Current | After |
| --- | --- |
| `use anyhow::{Context, Result};` | `use crate::error::{Error, Result};` (from `rt`) |
| `fn foo() -> anyhow::Result<T>` | `fn foo() -> Result<T>` |
| `anyhow::anyhow!("no such table: {}", name)` | `Error::msg(format!("no such table: {}", name))` |
| `anyhow::bail!("...")` | `return Err(Error::msg("..."))` |
| `result.context("while doing X")?` | `result.map_err(\|e\| e.with_context(\|\| "while doing X".into()))?` |
| `Err(anyhow::anyhow!("parse error at line {line}: ..."))` | `Err(Error::parse(line, col, "..."))` |
`crates/rt/src/parser.rs` and `crates/rt/src/compile.rs` are the biggest consumers. Most edits are verbatim substitutions.
### `Cargo.toml` delta
```diff
[dependencies]
-anyhow = "1"
libc = "0.2"
```
**After this phase:** `[dependencies]` has one line. The stated end goal.
## Exit criteria
1. **`cargo build`** — compiles with exactly one external dep.
2. **`cargo test --lib`** — all 14 existing `rt` tests pass. A new test in `error.rs` exercises `From<io::Error>`, `with_context`, and `Display` formatting.
3. **No `anyhow::` references anywhere in the repo.** `rg 'anyhow' crates/ docs/` returns zero hits (docs updated by this phase too).
4. **`reference/rest/blog.rest`** — 20 assertions still pass. Error paths (404, 400) still produce the same response body format (plain-text error message from the handler's `.to_string()`).
5. **`cd reference/crates && cargo build && cargo test`** unchanged. V1 doesn't use `anyhow` — nothing to touch there.
6. **`cargo tree -p rt --depth 1`** lists only `libc` as an external dep (plus transitive ones brought in by libc itself, all of which are kernel-facing).
## Non-scope
- **No custom `#[derive(Error)]` macro.** `thiserror` would be nicer but it's a dep. Hand-writing the enum + impls is ~120 lines, done once, maintained rarely.
- **No source-chain traversal.** `Error` stores messages, not source errors (except `Io` which wraps). If we need chained context later, extend the enum then — don't over-engineer now.
- **No `backtrace` crate.** If a panic-style backtrace is ever needed, `RUST_BACKTRACE=1` on `panic!()` gives it. Errors don't carry them.
## Verification
```bash
cargo build # one external dep
cargo test --lib # 14 + error.rs test green
rg 'anyhow' crates/ docs/ # zero hits
# full .rest smoke — same script as phase 04
cd reference/crates && cargo build && cargo test # v1 untouched
cat crates/rt/Cargo.toml | grep -A 20 '\[dependencies\]' # libc is the only line
```
## After this phase
`crates/rt/Cargo.toml` is at its minimum. The runtime drives every I/O operation through direct kernel primitives: `socket`, `bind`, `listen`, `accept4`, `epoll_wait`, `read`, `write`, `signalfd4`, `close`. Nothing between the code and the kernel except `libc`.
Phases 07 and 08 extend the kernel-primitive surface without adding deps — they're pure feature work on top of the foundation this sequence laid.

View file

@ -0,0 +1,124 @@
# 07 — `inotify` Content Watcher
**Context sources:** [`./02-event-loop-epoll.md`](./02-event-loop-epoll.md), [`./linux/00-linux.md`](./linux/00-linux.md) § File Watching, [`../02-recovery.md`](../02-recovery.md) § No AWS Infrastructure.
## Goal
First Stage-3 capability. When a `.wo` source file under the active project directory changes, `inotify` fires on the phase-02 event loop and the runtime hot-reloads the affected schema — parser re-run, catalog refreshed, live routes updated in-place. Maps to [`00-linux.md`](./linux/00-linux.md)'s "watch the content directory for file creates, modifications, and deletes. Triggers re-indexing and subscriber notification when articles change. **Replaces the S3 + Lambda event pipeline entirely.**"
Also the first real second consumer of the phase-02 `EventLoop` beyond the HTTP listener — validates the abstraction under cross-feature load.
## Design decisions (locked)
1. **`inotify_init1(IN_CLOEXEC | IN_NONBLOCK)` + `inotify_add_watch`.** The fd is registered on the phase-02 event loop alongside the HTTP listener. No polling. No cross-platform fallback (`kqueue` on macOS, `ReadDirectoryChangesW` on Windows) — writeonce targets Linux only.
2. **Per-directory watches, not per-file.** `types/`, `ui/`, `logic/`, `tests/`, and `app.wo`'s parent get a watch each; individual files are resolved from the event's `wd` + `name` fields. Prevents fd exhaustion on large projects (the default `fs.inotify.max_user_watches` is 8192 on most distros, but we'd rather spend watches carefully).
3. **Debounce at 150 ms.** Editors issue multiple events per save (create tempfile → write → rename → delete old). A `TimerFd::oneshot(150ms)` per-watch absorbs the burst; only the final "settled" state triggers a recompile.
4. **Full recompile, not incremental.** A file change invalidates the full schema catalog — re-run `rt::discover()` → `rt::parser::parse()` → `rt::compile::Catalog::from_schemas()`. The sample projects are small (8 files for the blog); a full parse is < 50 ms. Incremental type-graph invalidation is a phase 09+ optimization.
5. **Atomic catalog swap.** The `Engine`'s catalog is behind an `Arc<ArcSwap<Catalog>>` or equivalent — a new catalog replaces the old under a single pointer write, and in-flight HTTP handlers finish with the old one while new ones see the new. Under single-threaded execution this is essentially free; under sharding it becomes the per-shard atomic.
6. **Module at `crates/rt/src/watch/`.** Same extraction-deferred rule as earlier modules.
## Scope
### New files inside `crates/rt/src/watch/`
| File | Responsibility | Port source |
| --- | --- | --- |
| `mod.rs` | Re-exports `Watcher`, `WatchEvent` | — |
| `inotify.rs` | Raw wrappers: `init()`, `add_watch(path, mask)`, `read_events() -> Vec<RawEvent>`. Registers on the `EventLoop`. | [`reference/crates/wo-watch/src/lib.rs`](../../reference/crates/wo-watch/src/lib.rs) (280 LOC) — v1 already does exactly this |
| `recursive.rs` | Walks the project root, calls `add_watch` for every directory matching `types/\|ui/\|logic/\|tests/` or containing `*.wo` | ~80 new LOC |
| `debounce.rs` | Coalesces bursts per-watch-descriptor, fires a `TimerFd` for the 150 ms settle window | ~100 new LOC |
| `reload.rs` | On debounced fire: re-discover, re-parse, re-compile, `ArcSwap::store(new_catalog)` | ~80 new LOC |
Total: ~540 LOC (280 ported + ~260 new).
### `Cargo.toml` change
None. `libc` already covers `inotify_init1` / `inotify_add_watch` / `inotify_rm_watch`.
### Routing change in `crates/rt/src/server.rs`
The router needs to re-resolve the catalog on each request rather than close over a snapshot at boot:
```rust
// before
let router = Router::new().route("/api/articles", list_h_bound_to_catalog_snapshot);
// after
let shared = Arc::new(ArcSwap::from_pointee(catalog));
let router = Router::new().route("/api/articles", move |req, st| {
let cat = shared.load();
list_h(req, &cat, st)
});
```
One-time rewrite of the 12 handlers (4 types × 3 ops). Mechanical.
## API shape (target)
```rust
use rt::event::EventLoop;
use rt::watch::Watcher;
let mut loop_ = EventLoop::new()?;
let mut watcher = Watcher::recursive(Path::new("docs/examples/blog"), Duration::from_millis(150))?;
watcher.register(&mut loop_)?;
for event in loop_.wait_once(None)? {
if event.token() == watcher.token() {
for change in watcher.drain() {
eprintln!("[wo] content change: {} ({})", change.path.display(), change.kind);
// reload pipeline fires here
}
}
}
```
## Exit criteria
1. **`cargo build`** green. No new deps.
2. **Unit test:** create a temp dir, write `a.wo`, spin up a `Watcher` on a loop in a test thread, modify `a.wo`, assert the debounced `WatchEvent::Modified(path)` arrives within 250 ms.
3. **End-to-end manual:**
```bash
cargo run --bin wo -- run docs/examples/blog &
# observe: `curl :8080/api/articles` returns [...]
# edit docs/examples/blog/types/article.wo — add a `nickname: Text?` field
# wait 200 ms
# observe: `curl :8080/api/articles` response shape reflects new field (no restart)
```
4. **`[wo]` log lines** match the spec in [00-linux.md](./linux/00-linux.md) — one line per debounced change, showing the relative path and event kind.
5. **All 14 `rt` unit tests still pass.** `reference/rest/blog.rest` 20-assertion battery still green.
6. **No fd leak** — `ls -la /proc/$PID/fd` before and after ten consecutive edits shows the same count.
## Non-scope
- **No cross-platform fallback.** `kqueue` and `ReadDirectoryChangesW` are not on the roadmap. Linux only.
- **No `fanotify`.** [00-linux.md](./linux/00-linux.md) lists it as "useful if watching needs to span mount points" — writeonce projects live in one directory tree; `inotify` is enough.
- **No incremental reparse.** Full recompile per settled change. If a real project hits the full-recompile wall, phase 09+ can add a dependency-graph-aware rebuilder.
- **No subscription push.** Phase 07 only detects and reloads. Notifying connected clients (the `register! { #{blog-title} => notify(fd) }` model in [00-linux.md](./linux/00-linux.md)) is phase 09 once `sub` activates.
## Verification
```bash
cargo build
cargo test --lib watch
cargo test --lib # 14 existing + watcher tests green
# manual hot-reload check
cargo run --bin wo -- run docs/examples/blog &
PID=$!
sleep 2
curl -s http://127.0.0.1:8080/api/articles
echo ' policy read anyone' >> docs/examples/blog/types/article.wo
sleep 0.5
curl -s http://127.0.0.1:8080/api/articles # server did not restart; catalog refreshed
kill $PID
git checkout docs/examples/blog/types/article.wo # undo the edit
cd reference/crates && cargo build && cargo test # v1 untouched
```
## After this phase
The runtime now does what `docs/02-recovery.md` originally promised: the binary watches its own content directory with `inotify` and re-indexes on change. The S3 + Lambda + sync-trigger pipeline is fully replaced by one fd on one event loop in one process.
Phase 08 adds the other half of the v1 kernel-primitive story — `sendfile` for zero-copy static serving.

View file

@ -0,0 +1,108 @@
# 08 — `sendfile` Zero-Copy Static Serving
**Context sources:** [`./03-hand-rolled-http.md`](./03-hand-rolled-http.md), [`./linux/00-linux.md`](./linux/00-linux.md) § Efficient File Serving, [`../02-recovery.md`](../02-recovery.md).
## Goal
Serve static file bytes — eventually `##ui`-emitted HTML + CSS + JS bundle, today anything put under a project's `static/` directory — via `sendfile(sock_fd, file_fd, NULL, count)`. Zero userspace copy on the payload path: the kernel moves bytes from the page cache directly to the socket's send buffer. Completes the v1 kernel-primitive port started in phase 02.
## Design decisions (locked)
1. **`sendfile(2)` only.** Not `splice`, not `vmsplice`. `sendfile` handles "file fd → socket fd" exactly, which is 100% of the use case here. `splice`-through-pipe is ~30% more code for cases we don't have (non-regular-file sources).
2. **`GET /static/...` is the only route mounted.** Hard-coded for Stage 3. When `##ui` arrives in phase 6+ it'll emit bundles into this path; when typed SDK codegen arrives in phase 5 the generated JS client goes here too.
3. **No path traversal.** Canonicalise the requested path; reject anything that escapes the configured static root. Standard directory-traversal defence — `..` segments already stripped by the HTTP request parser from phase 03, but the static resolver double-checks with a realpath comparison.
4. **`open + fstat + sendfile` chain.** No `mmap`. `mmap` wins for repeated reads of the same file (where the page cache warming pays off), but `sendfile` is strictly faster for one-shot delivery since the kernel manages the page cache itself. [00-linux.md](./linux/00-linux.md) lists both; the runtime's static-asset pattern is one-shot, pick `sendfile`.
5. **`EAGAIN` backoff through the event loop.** If `sendfile` returns partial bytes (send buffer full), re-arm `EPOLLOUT` for the socket and resume when the kernel signals writable. Matches v1 wo-serve's flow.
6. **MIME by extension table.** Compact `match` on `.html`/`.css`/`.js`/`.json`/`.svg`/`.png`/`.jpg`/`.woff2`/`.wasm` covers every asset the SSR layer will emit. Unknown extensions default to `application/octet-stream`.
## Scope
### New files inside `crates/rt/src/static_files/`
| File | Responsibility | Port source |
| --- | --- | --- |
| `mod.rs` | Re-exports `StaticHandler`, `resolve` | — |
| `sendfile.rs` | Raw `sendfile(2)` wrapper + non-blocking `send_all` that co-operates with `EPOLLOUT` | [`reference/crates/wo-serve/src/sendfile.rs`](../../reference/crates/wo-serve/src/sendfile.rs) (109 LOC) |
| `resolve.rs` | Path canonicalisation + traversal defence + file existence check | [`reference/crates/wo-serve/src/resolve.rs`](../../reference/crates/wo-serve/src/resolve.rs) (80 LOC) |
| `mime.rs` | Extension → `Content-Type` table | [`reference/crates/wo-serve/src/mime.rs`](../../reference/crates/wo-serve/src/mime.rs) (44 LOC) |
| `handler.rs` | `StaticHandler` — integrates the three with phase-03's `Response` builder; returns 404 / 403 / 200 as appropriate | ~120 new LOC |
Total: ~350 LOC (233 ported + ~120 new).
### `Cargo.toml` change
None.
### Router change in `crates/rt/src/server.rs`
One new route per project:
```rust
let static_root = project_dir.join("static");
let handler = StaticHandler::new(static_root);
router.route(Method::GET, "/static/*path", move |req, _| handler.serve(req));
```
`/static/*path` is a new wildcard pattern in the phase-03 router — add it to `route.rs` if not already supported.
## API shape (target)
```rust
use rt::static_files::StaticHandler;
let handler = StaticHandler::new("/app/static");
handler.serve(&request)?; // returns a Response that streams via sendfile()
```
The `Response` returned by `handler.serve()` owns the open `File` fd. The phase-03 connection writer notices it's a `sendfile`-backed response and uses the non-blocking `send_all` path instead of `write`.
## Exit criteria
1. **`cargo build`** green. No new deps.
2. **Unit test:** place a 10 MB file under a temp static root, `StaticHandler::serve` on a mock request, assert the `Response` reports 200 / correct `Content-Length` / correct `Content-Type`. A separate integration test validates the actual `sendfile` path using an accepted socket.
3. **`strace` validation:**
```bash
cargo run --bin wo -- run docs/examples/blog &
PID=$!
# ship a 10 MB file into docs/examples/blog/static/big.bin
strace -p $PID -f -e sendfile,read,write 2>&1 | tee /tmp/strace.log &
curl -s -o /dev/null http://127.0.0.1:8080/static/big.bin
# assert /tmp/strace.log shows sendfile(...) calls and zero read/write
# of the file content
```
4. **Path traversal attempts fail closed.** `curl :8080/static/../Cargo.toml` returns 403. `curl :8080/static/nonexistent.png` returns 404.
5. **`EAGAIN` handling.** A test that rate-limits the socket sendbuf to force a partial write exercises the `EPOLLOUT` re-arm path; the full payload still arrives.
6. **All 14 `rt` tests** + phase-02/03/04/05/06/07 additions pass. `reference/rest/blog.rest` 20 assertions still green (no regressions on the JSON endpoints).
## Non-scope
- **No `sendfile64`.** Modern glibc aliases `sendfile` to `sendfile64` transparently; explicit 64-bit selection isn't needed.
- **No range requests.** `GET /static/big.bin` with `Range: bytes=...` returns 200 + full payload in Stage 3; proper range-request handling is a follow-on. Nothing in the blog or ecommerce samples uses ranges.
- **No in-memory cache.** The kernel page cache is the only cache. Re-opening the file on every request is cheap; if future profiling says otherwise, add an LRU fd cache — but not pre-emptively.
- **No TLS.** `sendfile` over TLS requires kTLS (`setsockopt(TCP_ULP, "tls")` + kernel 4.13+ and the right cipher suites). Worth doing when TLS lands as its own phase; out of scope here.
- **No compression.** The HTTP response writer in phase 03 doesn't gzip; `sendfile` can't gzip on the fly either. Pre-compress (`.br` / `.gz` sibling files) is a future phase — for now the table lists `.br`/`.gz` extensions with the correct `Content-Encoding` but the caller has to produce the pre-compressed file itself.
## Verification
```bash
cargo build
cargo test --lib static_files
cargo test --lib # all existing tests green
# strace-backed zero-copy proof
cargo run --bin wo -- run docs/examples/blog &
# ... (full script from exit criterion 3)
# .rest smoke unchanged
# full 20-assertion battery against reference/rest/blog.rest
cd reference/crates && cargo build && cargo test # v1 untouched
```
## After this phase
The runtime covers every kernel primitive listed in [`00-linux.md`](./linux/00-linux.md) except `io_uring`, `mmap`, `fallocate`, and `memfd_create` — which all belong to the storage engine (phase 3 of the database series), not the runtime per se.
Next natural phase: **`09-native-subscriptions.md`** — the `register! { #{blog-title} => notify(fd) }` model from [00-linux.md](./linux/00-linux.md). Takes the `inotify` watcher from phase 07 and wires it into a subscription table that dispatches delta writes directly to subscriber sockets over the phase-03 HTTP connection. That replaces the Stage-3 `501` stub the `/api/<type>/live` endpoint currently returns.
After that phase, `crates/rt/` is feature-complete for Stages 1–3 of the runtime, with exactly one external dependency.

View file

@ -0,0 +1,68 @@
# 01 — Scaffolding the Crate Tree
**Context sources:** [`../../CLAUDE.md`](../../CLAUDE.md), [`../../README.md`](../../README.md), [`../../crates/README.md`](../../crates/README.md), [`../runtime/database.md`](../runtime/database.md), [`../runtime/database/07-wo-seg-migration.md`](../runtime/database/07-wo-seg-migration.md).
## Goal
Lay out the full `crates/` directory tree that the 7-phase `.wo` runtime design implies, **without moving any active code**. Every target crate the design docs name gets an empty-but-compilable home so that:
- every doc reference to `ql`, `wal`, `policy`, etc. resolves to a real directory
- each phase's extraction work (move X out of `rt` into its target crate) is a mechanical copy into an existing skeleton rather than a net-new crate creation
- IDE workspace views, dependency graphs, and `cargo doc` show the project's intended shape on day one
## Design decisions (locked)
1. **Scope: all 7 phases.** 14 new empty library crates covering Phases 2 → 6 land in one pass. `rt` (already shipping Stage 2) is the 15th.
2. **`rt` stays monolithic.** Today's Stage 2 code — lexer / parser / AST / compile / engine / server — stays inside `crates/rt/` and continues to satisfy the 14 existing unit tests. Code migrates into the new crates as each phase activates, not in this pass.
3. **Contents: `Cargo.toml` + `src/lib.rs` doc-comment only.** Each `lib.rs` is one module-level `//!` doc block pointing at the phase doc, naming the responsibilities, and flagging which modules in `rt` migrate here later. No placeholder types, no stub traits.
4. **No `wo-` prefix.** New runtime crates are `ql`, `value`, `engine`, etc. — not `wo-ql`, `wo-value`. The prefix is redundant inside the project's own `wo` namespace and noisy in imports (`use ql::Parser` beats `use wo_ql::Parser`). The v1 crates in `reference/crates/` keep their `wo-` prefix — the distinct prefix makes the v1/v2 split visible at a glance.
5. **Workspace membership: root `Cargo.toml` lists every new crate as a member.** `reference/crates` stays `exclude`-d (nested workspace, separate v1 code).
Rationale and alternatives considered: see [`../../CLAUDE.md`](../../CLAUDE.md) "What's in `rt` today vs. what the empty crates promise" and the recorded `AskUserQuestion` answers that preceded this plan.
## Crate map
All names are stable — documented in [`../runtime/database/07-wo-seg-migration.md`](../runtime/database/07-wo-seg-migration.md) (Phase 2–5) and derived from [`../runtime/database/06-lowcode-fullstack.md`](../runtime/database/06-lowcode-fullstack.md) component tables (Phase 6).
| Phase | Crate | One-line purpose |
| --- | --- | --- |
| 2 | `ql` | `.wo` grammar — lexer, parser, AST |
| 2 | `value` | tagged `Value` + dotted-path helpers |
| 2 | `engine` | in-memory executor (rel / doc / graph) + schema catalog |
| 2 | `txn` | transaction coordinator — MVCC, `RETURNING` alias table |
| 2 | `db` | top-level facade — `open()`, `Tx`, `Query`, `Subscribe` |
| 3 | `wal` | write-ahead log — io_uring + fsync + recovery |
| 4 | `sub` | live subscriptions — delta frames on commit |
| 4 | `http` | wire protocol — REST / GraphQL-over-WS / native codec |
| 5 | `gen` | codegen — `.wo type` → Go / TS / Rust / Python clients |
| 6 | `policy` | RBAC + row-level rules compiled into planner rewrites |
| 6 | `logic` | `on <event>` triggers + `fn ... in txn` interpreter |
| 6 | `service` | `service rest/graphql/native` endpoint dispatch |
| 6 | `ui` | `##ui` screens → SSR HTML + client runtime |
| 6 | `app` | `##app` route manifest + startup hooks |
| — | `rt` | **existing** — Stage-2 monolith + the `wo` binary |
## Status
✅ **Done.** All 15 crates exist, the root workspace members list includes them, `cargo build` and `cargo test --lib` both pass, and `cargo run --bin wo -- run docs/examples/blog` still serves the blog sample (Stage 2 behaviour unchanged).
| Artifact | Status |
| --- | --- |
| 14 new crate skeletons (`Cargo.toml` + `src/lib.rs`) | ✅ |
| Root `Cargo.toml` lists all 15 crates as members | ✅ |
| `crates/README.md` with the phase-mapped inventory | ✅ |
| `cargo build` at root (compiles 15 crates) | ✅ |
| `cargo test --lib` at root (14 existing `rt` tests) | ✅ |
| `cd reference/crates && cargo build && cargo test` (v1 still green) | ✅ |
| `cargo run --bin wo -- run docs/examples/blog` (Stage 2 still serves) | ✅ |
## Non-scope
- **No code extraction.** Moving `lexer.rs` / `parser.rs` / `engine.rs` out of `rt` into `ql` / `engine` is explicitly deferred. That happens incrementally as each phase activates.
- **No `gen` binary target.** `gen` ships as a library-only crate in this pass. The `[[bin]]` lands when Phase 5 starts.
- **No test scaffolding.** The empty crates don't get unit-test stubs. When a crate gets real code, it gets real tests.
- **No cross-crate `pub use` reexports from `db`.** The facade crate documents its future surface in its doc comment but doesn't import anything yet.
## After this plan lands
Next planning documents in this directory should describe the first real extraction — likely `02-extract-ql.md` when Stage 3 begins and the subscription engine needs the parser from a second call site. Until then, the 14 placeholders sit unmodified alongside `rt`.

View file

@ -1,6 +1,26 @@
## Linux Kernel Features
Kernel primitives that the writeonce binary can leverage, mapped to the architectural needs identified in [01-problem.md](./01-problem.md) and [02-recovery.md](./02-recovery.md).
Kernel primitives that the writeonce binary can leverage, mapped to the architectural needs identified in [docs/01-problem.md](../../01-problem.md) and [docs/02-recovery.md](../../02-recovery.md).
### Per-primitive reference cards
Each primitive has its own numbered file with the kernel source path (into [`reference/linux/`](../../../reference/linux/)), Rust FFI signature via `libc`, a minimal direct-syscall example, and the v1 port source. Use these when implementing the phase docs under [`docs/plan/`](../).
| # | Primitive | Used by |
| --- | --- | --- |
| [01](./01-epoll.md) | `epoll` — event-driven I/O multiplexing | every runtime phase |
| [02](./02-eventfd.md) | `eventfd` — counter as fd, cross-flow wake | phase 02, subscription wakeup |
| [03](./03-timerfd.md) | `timerfd` — timers as fds | phase 02, phase 07 debounce |
| [04](./04-signalfd.md) | `signalfd` — signals as fds, graceful shutdown | phase 04 |
| [05](./05-inotify.md) | `inotify` — filesystem events as fds | phase 07, future register! subscription |
| [06](./06-sendfile.md) | `sendfile` — zero-copy file → socket | phase 08 |
| [07](./07-io_uring.md) | `io_uring` — async I/O ring buffers | phase 3 (WAL fsync), future HTTP |
| [08](./08-mmap.md) | `mmap` + `madvise` — memory-mapped files, page-cache hints | phase 3 (storage engine) |
| [09](./09-fallocate.md) | `fallocate` + `pread` + `pwritev2` — positional I/O & pre-allocation | phase 3 (WAL + SSTables) |
| [10](./10-pidfd.md) | `pidfd` — process as fd, race-free supervision | future supervisor |
| [11](./11-memfd_create.md) | `memfd_create` — anonymous shared memory | phase 3 (index build) |
The list below is the original overview kept for context and for a handful of adjacent primitives (`fanotify`, `splice`/`tee`) that don't yet have their own reference card.
### File Watching — Content Directory

View file

@ -0,0 +1,75 @@
# 01 — `epoll`
Event-driven I/O multiplexing. One `epoll_fd` watches many fds for readiness; `epoll_wait` blocks the loop until at least one fires. Foundation of the whole runtime — every other primitive on this list is an fd that lands on an `epoll` loop.
## Kernel source
| Path | What |
| --- | --- |
| [`reference/linux/fs/eventpoll.c`](../../../reference/linux/fs/eventpoll.c) | All three syscalls (`epoll_create1`, `epoll_ctl`, `epoll_wait`) live here. Grep for `SYSCALL_DEFINE`. |
| [`reference/linux/include/uapi/linux/eventpoll.h`](../../../reference/linux/include/uapi/linux/eventpoll.h) | `struct epoll_event`, `EPOLL_*` flags, the userspace-facing ABI. |
## Man pages
`man 7 epoll` (overview), `man 2 epoll_create1`, `man 2 epoll_ctl`, `man 2 epoll_wait`.
## Rust FFI via `libc`
```rust
use libc::{epoll_create1, epoll_ctl, epoll_wait, epoll_event, EFD_CLOEXEC};
use libc::{EPOLL_CLOEXEC, EPOLL_CTL_ADD, EPOLL_CTL_MOD, EPOLL_CTL_DEL};
use libc::{EPOLLIN, EPOLLOUT, EPOLLET, EPOLLRDHUP, EPOLLHUP, EPOLLERR};
extern "C" {
// all three already wrapped in libc — no SYS_* workaround needed
}
```
## Direct-syscall example
```rust
unsafe {
let epfd = libc::epoll_create1(libc::EPOLL_CLOEXEC);
if epfd < 0 { return Err(io::Error::last_os_error()); }
let mut ev = libc::epoll_event {
events: (libc::EPOLLIN | libc::EPOLLET) as u32,
u64: token_value, // your own dispatch key
};
if libc::epoll_ctl(epfd, libc::EPOLL_CTL_ADD, watched_fd, &mut ev) < 0 {
return Err(io::Error::last_os_error());
}
let mut events: [libc::epoll_event; 64] = std::mem::zeroed();
let n = libc::epoll_wait(epfd, events.as_mut_ptr(), 64, timeout_ms);
for ev in &events[..n as usize] {
let token = ev.u64;
// dispatch based on token
}
}
```
## Key flags
| Flag | Meaning |
| --- | --- |
| `EPOLL_CLOEXEC` | Close the `epoll_fd` on `exec` — always set it. |
| `EPOLLIN` / `EPOLLOUT` | Readable / writable readiness. |
| `EPOLLET` | **Edge-triggered.** Reads must drain until `EAGAIN`; no re-arm needed. Required for the single-threaded loop's correctness. |
| `EPOLLRDHUP` | Peer half-closed — distinguishes "client closed" from transient. |
| `EPOLLHUP` / `EPOLLERR` | Always implicitly set; you don't ask for them but you must handle. |
## Gotchas
- **Edge-triggered means drain.** A handler that reads once and returns leaks readiness; the next `epoll_wait` won't re-fire until more bytes arrive. Always `read()` in a loop until `EAGAIN`.
- **`timeout = -1` blocks forever**; `0` polls; positive milliseconds cap the wait.
- **`epoll_wait` can return fewer events than you asked.** Normal. Don't assume full batches.
- **Modifying a watched fd's interest requires `EPOLL_CTL_MOD`**, not delete-then-add — the atomic update avoids races with pending events.
## Used by
Every runtime phase that touches I/O: [`02-event-loop-epoll.md`](../02-event-loop-epoll.md), [`03-hand-rolled-http.md`](../03-hand-rolled-http.md), [`07-inotify-content-watcher.md`](../07-inotify-content-watcher.md), [`08-sendfile-static-assets.md`](../08-sendfile-static-assets.md).
## v1 port source
[`reference/crates/wo-event/src/epoll.rs`](../../../reference/crates/wo-event/src/epoll.rs) (183 LOC) — already wraps all three syscalls with a safe `EventLoop { register, deregister, wait_once }` facade.

View file

@ -0,0 +1,66 @@
# 02 — `eventfd`
Counter as a file descriptor. `write(fd, &n, 8)` adds `n` to the counter; `read(fd, &buf, 8)` drains it to zero (or decrements by one in semaphore mode). Paired with `epoll`, it's the cheapest way to wake the event loop from another flow — cross-thread signalling, scheduled work, shutdown requests.
## Kernel source
| Path | What |
| --- | --- |
| [`reference/linux/fs/eventfd.c`](../../../reference/linux/fs/eventfd.c) | `SYSCALL_DEFINE2(eventfd, ...)`, `struct eventfd_ctx`, read/write handlers. |
| [`reference/linux/include/uapi/linux/eventfd.h`](../../../reference/linux/include/uapi/linux/eventfd.h) | `EFD_*` flags. |
## Man pages
`man 2 eventfd`.
## Rust FFI via `libc`
```rust
use libc::{eventfd, EFD_CLOEXEC, EFD_NONBLOCK, EFD_SEMAPHORE};
// read / write go through plain libc::read / libc::write
extern "C" {
// eventfd is already wrapped in libc
}
```
## Direct-syscall example
```rust
unsafe {
let fd = libc::eventfd(0, libc::EFD_CLOEXEC | libc::EFD_NONBLOCK);
if fd < 0 { return Err(io::Error::last_os_error()); }
// wake the loop from anywhere
let one: u64 = 1;
libc::write(fd, &one as *const _ as *const _, 8);
// inside the loop, on EPOLLIN:
let mut buf: u64 = 0;
let n = libc::read(fd, &mut buf as *mut _ as *mut _, 8);
// buf now holds the accumulated count (or 1 if EFD_SEMAPHORE)
}
```
## Key flags
| Flag | Meaning |
| --- | --- |
| `EFD_CLOEXEC` | Close on exec. Always set. |
| `EFD_NONBLOCK` | `read` returns `EAGAIN` when counter is zero instead of blocking. Required when the fd is on an `epoll` loop. |
| `EFD_SEMAPHORE` | Each `read` decrements by one (otherwise reads drain the whole counter). Useful as a bounded work queue. |
## Gotchas
- **Always 8-byte `read` / `write`.** Short reads/writes return `EINVAL` — the counter is `u64`, full word or nothing.
- **Writing `u64::MAX`** returns `EINVAL`; the counter can't hold more than `u64::MAX - 1`.
- **Multiple writers are OK**; the kernel serialises. But reads race — use `EFD_SEMAPHORE` if you want one consumer per write.
- **Not async-signal-safe.** Don't `write(fd, ...)` from a signal handler; use `signalfd` instead (see [04-signalfd.md](./04-signalfd.md)).
## Used by
[`02-event-loop-epoll.md`](../02-event-loop-epoll.md) — wake the loop for shutdown or internal work. Future `sub` crate ([`09-native-subscriptions`], not yet planned) uses it to signal that a subscriber queue has drained.
## v1 port source
[`reference/crates/wo-event/src/eventfd.rs`](../../../reference/crates/wo-event/src/eventfd.rs) (66 LOC) — `EventFd { new, write, read, as_raw_fd }`.

View file

@ -0,0 +1,77 @@
# 03 — `timerfd`
Timers as file descriptors. Set an expiry with `timerfd_settime`, `read` the fd to retrieve the number of expirations, `epoll` notifies the loop when the timer fires. Drives everything in the runtime that needs a deadline without a separate timer thread: keepalives, debounce windows, checkpoint intervals.
## Kernel source
| Path | What |
| --- | --- |
| [`reference/linux/fs/timerfd.c`](../../../reference/linux/fs/timerfd.c) | All three syscalls (`timerfd_create`, `timerfd_settime`, `timerfd_gettime`). |
| [`reference/linux/include/uapi/linux/timerfd.h`](../../../reference/linux/include/uapi/linux/timerfd.h) | `TFD_*` flags. |
## Man pages
`man 2 timerfd_create`, `man 2 timerfd_settime`, `man 2 timerfd_gettime`.
## Rust FFI via `libc`
```rust
use libc::{timerfd_create, timerfd_settime, timerfd_gettime};
use libc::{itimerspec, timespec};
use libc::{TFD_CLOEXEC, TFD_NONBLOCK, TFD_TIMER_ABSTIME};
use libc::{CLOCK_MONOTONIC, CLOCK_REALTIME};
```
## Direct-syscall example
```rust
unsafe {
let fd = libc::timerfd_create(
libc::CLOCK_MONOTONIC,
libc::TFD_CLOEXEC | libc::TFD_NONBLOCK,
);
if fd < 0 { return Err(io::Error::last_os_error()); }
// one-shot: fire in 150ms, no interval
let spec = libc::itimerspec {
it_value: libc::timespec { tv_sec: 0, tv_nsec: 150_000_000 },
it_interval: libc::timespec { tv_sec: 0, tv_nsec: 0 },
};
// periodic: replace it_interval with the period, e.g. { tv_sec: 1, tv_nsec: 0 }
if libc::timerfd_settime(fd, 0, &spec, std::ptr::null_mut()) < 0 {
return Err(io::Error::last_os_error());
}
// register fd on the epoll loop; on EPOLLIN:
let mut expirations: u64 = 0;
libc::read(fd, &mut expirations as *mut _ as *mut _, 8);
// expirations > 0 means the timer fired. Usually 1 for a oneshot,
// could be >1 for a periodic timer whose consumer missed ticks.
}
```
## Key flags
| Flag | Meaning |
| --- | --- |
| `CLOCK_MONOTONIC` | Steady clock; immune to wall-clock jumps. **Default choice for runtime timers.** |
| `CLOCK_REALTIME` | Wall clock; jumps when NTP corrects or user sets time. Avoid unless you need calendar semantics. |
| `TFD_CLOEXEC` | Close on exec. Always set. |
| `TFD_NONBLOCK` | `read` doesn't block when expirations == 0. Required when on `epoll`. |
| `TFD_TIMER_ABSTIME` | Interpret `it_value` as absolute (not relative). Passed to `timerfd_settime`'s `flags`, not at create. |
## Gotchas
- **Reads are always 8 bytes.** The value is an expiration count.
- **Disarming is `timerfd_settime(fd, 0, &{0,0,0,0}, NULL)`** — a zero spec cancels pending fires.
- **Re-arming a periodic timer** replaces the spec atomically; no "pause then resume" surface.
- **Resolution is nanosecond-granular but scheduling can slip.** Don't use for sub-millisecond accuracy — use a busy loop or hardware timer for that.
## Used by
[`02-event-loop-epoll.md`](../02-event-loop-epoll.md) — housekeeping timers. [`07-inotify-content-watcher.md`](../07-inotify-content-watcher.md) — 150 ms debounce window after an inotify burst. Future subscription phase — keepalive pings to long-lived connections.
## v1 port source
[`reference/crates/wo-event/src/timerfd.rs`](../../../reference/crates/wo-event/src/timerfd.rs) (91 LOC) — `TimerFd { oneshot(dur), periodic(dur), disarm, read_expirations }`.

View file

@ -0,0 +1,77 @@
# 04 — `signalfd`
Unix signals as file descriptors. `signalfd(fd, mask)` installs a mask on the process and returns an fd that becomes readable when any masked signal arrives. Replaces signal handlers entirely — no `sigaction`, no re-entrancy minefield, no interrupted syscalls. The whole `wo` binary's shutdown path is one `signalfd` on the event loop.
## Kernel source
| Path | What |
| --- | --- |
| [`reference/linux/fs/signalfd.c`](../../../reference/linux/fs/signalfd.c) | `SYSCALL_DEFINE4(signalfd4, ...)` + `signalfd_dequeue`. |
| [`reference/linux/include/uapi/linux/signalfd.h`](../../../reference/linux/include/uapi/linux/signalfd.h) | `struct signalfd_siginfo`, `SFD_*` flags. |
| [`reference/linux/kernel/signal.c`](../../../reference/linux/kernel/signal.c) | Background: `sigprocmask`, pending-signal dequeue. |
## Man pages
`man 2 signalfd`, `man 7 signal`.
## Rust FFI via `libc`
```rust
use libc::{signalfd, signalfd_siginfo, sigset_t};
use libc::{sigemptyset, sigaddset, sigprocmask};
use libc::{SFD_CLOEXEC, SFD_NONBLOCK, SIG_BLOCK, SIG_UNBLOCK};
use libc::{SIGINT, SIGTERM, SIGHUP, SIGQUIT};
```
## Direct-syscall example
```rust
unsafe {
// 1. Block the signals in the thread so the kernel delivers them via the fd instead
let mut mask: sigset_t = std::mem::zeroed();
libc::sigemptyset(&mut mask);
libc::sigaddset(&mut mask, libc::SIGINT);
libc::sigaddset(&mut mask, libc::SIGTERM);
if libc::sigprocmask(libc::SIG_BLOCK, &mask, std::ptr::null_mut()) < 0 {
return Err(io::Error::last_os_error());
}
// 2. Create the fd
let fd = libc::signalfd(-1, &mask, libc::SFD_CLOEXEC | libc::SFD_NONBLOCK);
if fd < 0 { return Err(io::Error::last_os_error()); }
// 3. Register fd on epoll. On EPOLLIN:
let mut info: signalfd_siginfo = std::mem::zeroed();
let n = libc::read(fd, &mut info as *mut _ as *mut _, std::mem::size_of::<signalfd_siginfo>());
if n as usize == std::mem::size_of::<signalfd_siginfo>() {
match info.ssi_signo as i32 {
libc::SIGINT | libc::SIGTERM => begin_graceful_shutdown(),
_ => {}
}
}
}
```
## Key flags
| Flag | Meaning |
| --- | --- |
| `SFD_CLOEXEC` | Close on exec. Always set. |
| `SFD_NONBLOCK` | Non-blocking reads. Required when on `epoll`. |
| `SIG_BLOCK` | Passed to `sigprocmask` to add the set to the current mask. The corresponding `SIG_UNBLOCK` / `SIG_SETMASK` are available if you need them. |
## Gotchas
- **Blocking the signal with `sigprocmask` is mandatory.** Otherwise the default handler runs and kills the process before the fd ever becomes readable. Block once at boot, never unblock.
- **Per-thread mask.** `sigprocmask` is per-thread. In a single-threaded runtime that's fine; if you ever spawn threads, use `pthread_sigmask` on each so the main loop gets the signals.
- **`signalfd_siginfo` is big** (128 bytes). Read into an aligned buffer; partial reads return `EINVAL`.
- **Doesn't catch `SIGKILL` or `SIGSTOP`.** Nothing does. Those bypass everything.
- **Child exit notifications (`SIGCHLD`)** work via signalfd but `pidfd` (see [10-pidfd.md](./10-pidfd.md)) is usually the better fit for clean child supervision.
## Used by
[`04-cutover-remove-tokio-axum.md`](../04-cutover-remove-tokio-axum.md) — replaces `tokio::signal::ctrl_c()` for graceful shutdown. Every subsequent phase inherits this pattern.
## v1 port source
[`reference/crates/wo-event/src/signalfd.rs`](../../../reference/crates/wo-event/src/signalfd.rs) (62 LOC) — `SignalFd::new(&[SIGINT, SIGTERM]) -> SignalFd` with a safe `read_signo()` helper.

View file

@ -0,0 +1,88 @@
# 05 — `inotify`
Filesystem event notifications as a file descriptor. `inotify_add_watch(dir, mask)` installs a watch; reading the fd returns variable-length `inotify_event` records each time a matching file creates, modifies, or disappears. Replaces the S3 + Lambda pipeline from the v1 architecture — one fd on the event loop IS the content-sync engine.
## Kernel source
| Path | What |
| --- | --- |
| [`reference/linux/fs/notify/inotify/inotify_user.c`](../../../reference/linux/fs/notify/inotify/inotify_user.c) | `SYSCALL_DEFINE1(inotify_init1, ...)`, `SYSCALL_DEFINE3(inotify_add_watch, ...)`, `SYSCALL_DEFINE2(inotify_rm_watch, ...)`. |
| [`reference/linux/fs/notify/inotify/inotify_fsnotify.c`](../../../reference/linux/fs/notify/inotify/inotify_fsnotify.c) | The fsnotify backend that feeds events into the fd. |
| [`reference/linux/include/uapi/linux/inotify.h`](../../../reference/linux/include/uapi/linux/inotify.h) | `struct inotify_event`, `IN_*` masks. |
## Man pages
`man 7 inotify` (overview + event semantics), `man 2 inotify_init1`, `man 2 inotify_add_watch`, `man 2 inotify_rm_watch`.
## Rust FFI via `libc`
```rust
use libc::{inotify_init1, inotify_add_watch, inotify_rm_watch, inotify_event};
use libc::{IN_CLOEXEC, IN_NONBLOCK};
use libc::{IN_MODIFY, IN_CREATE, IN_DELETE, IN_CLOSE_WRITE};
use libc::{IN_MOVED_FROM, IN_MOVED_TO, IN_ISDIR, IN_Q_OVERFLOW};
```
## Direct-syscall example
```rust
unsafe {
let fd = libc::inotify_init1(libc::IN_CLOEXEC | libc::IN_NONBLOCK);
if fd < 0 { return Err(io::Error::last_os_error()); }
let dir = std::ffi::CString::new("docs/examples/blog/types").unwrap();
let wd = libc::inotify_add_watch(
fd,
dir.as_ptr(),
(libc::IN_MODIFY | libc::IN_CLOSE_WRITE
| libc::IN_CREATE | libc::IN_DELETE
| libc::IN_MOVED_FROM | libc::IN_MOVED_TO) as u32,
);
if wd < 0 { return Err(io::Error::last_os_error()); }
// register fd on epoll. On EPOLLIN:
let mut buf = [0u8; 4096];
let n = libc::read(fd, buf.as_mut_ptr() as *mut _, buf.len());
// parse the buffer: a packed sequence of inotify_event records,
// each followed by a variable-length name field (event.len bytes)
let mut offset = 0usize;
while offset < n as usize {
let ev = &*(buf.as_ptr().add(offset) as *const inotify_event);
let name_len = ev.len as usize;
let name_start = offset + std::mem::size_of::<inotify_event>();
let name = std::str::from_utf8(&buf[name_start..name_start + name_len])
.unwrap().trim_end_matches('\0');
// dispatch based on ev.mask + ev.wd → directory → full path
offset = name_start + name_len;
}
}
```
## Key flags
| Flag | Meaning |
| --- | --- |
| `IN_CLOEXEC` / `IN_NONBLOCK` | Close on exec, non-blocking reads. Always set. |
| `IN_MODIFY` | File content written. Fires per-`write(2)` — noisy; prefer `IN_CLOSE_WRITE`. |
| `IN_CLOSE_WRITE` | File opened for writing was closed. **Usual choice** — one event per editor save. |
| `IN_CREATE` / `IN_DELETE` | Entry created / deleted inside a watched directory. |
| `IN_MOVED_FROM` / `IN_MOVED_TO` | The two halves of a rename. Paired by `cookie`. Editors often write tmp → rename → delete; you get both halves. |
| `IN_ISDIR` | Set on the event when the target is a directory. |
| `IN_Q_OVERFLOW` | Kernel event queue overflowed; `wd = -1`, rescan from scratch. **Must handle.** |
## Gotchas
- **Per-directory watches, not per-file.** Watching individual files wastes descriptors and misses `IN_CREATE`/`IN_DELETE` for new entries. Watch the directory; filter by event `name` in userspace.
- **Recursive watching is manual.** Walk the tree at init and add a watch per directory. React to `IN_CREATE | IN_ISDIR` by adding a watch for the new subdirectory — and to `IN_MOVED_TO | IN_ISDIR` too.
- **`fs.inotify.max_user_watches`** defaults to 8192 on most distros. Recursive watches over a big node_modules or target dir exhaust it fast. Filter aggressively before adding.
- **Editors burst events.** Tmp-file + rename + delete is 3–4 events per logical save. Debounce 100–200 ms with [`timerfd`](./03-timerfd.md).
- **Reading less than a full event is an `EINVAL`.** Use a buffer ≥ `sizeof(inotify_event) + NAME_MAX + 1` (≈ 4 KiB is a safe size).
- **`wd` is stable per-watch but reused after `rm_watch`.** Keep a `wd → path` map; remove from it on `IN_IGNORED`.
## Used by
[`07-inotify-content-watcher.md`](../07-inotify-content-watcher.md) — the Stage-3 hot-reload feature. Future `sub` crate — the register-macro subscription model in [`00-linux.md § Database Subscription`](./00-linux.md#database-subscription).
## v1 port source
[`reference/crates/wo-watch/src/lib.rs`](../../../reference/crates/wo-watch/src/lib.rs) (280 LOC) — already does recursive watch setup, event parsing, and path resolution via a `wd → PathBuf` map.

View file

@ -0,0 +1,78 @@
# 06 — `sendfile`
Zero-copy transfer from a file fd to a socket fd. The kernel splices pages directly from the page cache into the socket's send buffer — userspace never touches the bytes. One syscall, one copy (DMA → NIC), no userspace buffer.
## Kernel source
| Path | What |
| --- | --- |
| [`reference/linux/fs/read_write.c`](../../../reference/linux/fs/read_write.c) | `SYSCALL_DEFINE4(sendfile, ...)` and `SYSCALL_DEFINE4(sendfile64, ...)`. Modern glibc aliases the first to the second; the syscalls are distinguished by the offset type. |
| [`reference/linux/fs/splice.c`](../../../reference/linux/fs/splice.c) | Internally `sendfile` delegates to `splice_direct_to_actor`. Related — see [07-splice.md](./07-splice.md) if you ever need the more general fd-to-fd pipe path. |
## Man pages
`man 2 sendfile`.
## Rust FFI via `libc`
```rust
use libc::{sendfile, off_t};
// sendfile64 is the same syscall on 64-bit Linux; libc::sendfile already uses it
```
## Direct-syscall example
```rust
unsafe {
// out_fd must be a socket; in_fd must be a regular file opened O_RDONLY.
// Pass NULL for the offset pointer to advance the file's own file offset;
// pass a &mut offset to keep the file position untouched and step through explicitly.
let mut sent: isize = 0;
let mut remaining = file_len;
let mut offset: off_t = 0;
while remaining > 0 {
let n = libc::sendfile(sock_fd, file_fd, &mut offset as *mut _, remaining as usize);
if n < 0 {
let e = io::Error::last_os_error();
match e.raw_os_error() {
Some(libc::EAGAIN) | Some(libc::EWOULDBLOCK) => {
// re-arm EPOLLOUT on sock_fd and yield back to the loop;
// resume this call when the loop wakes us
return Err(e);
}
_ => return Err(e),
}
}
sent += n;
remaining -= n as u64;
if n == 0 { break; } // peer closed
}
}
```
## Key behaviour
| Detail | Notes |
| --- | --- |
| Input fd | Must support `mmap`-like access — regular files, shared memory, some block devices. **Not** sockets, pipes, or character devices. |
| Output fd | Must be a socket (kernel 2.6.33+ lifted the restriction to anything, but practically: sockets). |
| Max per-call | Kernel caps at ~2 GB per syscall regardless of what you request. Loop for bigger files. |
| `offset` pointer | If NULL, updates the input fd's internal offset (like `read` does). If non-NULL, updates only the pointed-to variable. **Always use non-NULL** when sharing the fd across concurrent readers. |
| Return | Bytes sent (possibly less than requested — partial send; re-arm `EPOLLOUT` and resume). |
## Gotchas
- **`EAGAIN` is the common partial-write case.** The socket's send buffer filled; register `EPOLLOUT`, drain the CQ when the loop fires, keep calling `sendfile` with the updated `offset`.
- **TLS and `sendfile` don't mix** (without kTLS). Encryption requires a userspace copy. If / when TLS lands, use kTLS (`setsockopt(TCP_ULP, "tls")`, kernel 4.13+, AES-GCM only in practice).
- **Chunked transfer encoding and `sendfile` also don't mix** — the chunk framing has to go around the payload. Send headers with `write`, then `sendfile` the body, then write the trailing zero-length chunk.
- **Compression can't happen on the wire** — the kernel doesn't gzip. Pre-compressed sibling files (`.gz`, `.br`) + `Content-Encoding` header is the idiom.
- **`fstat` before `sendfile`** to get the file length for `Content-Length` — spares the client from guessing when the body ends.
## Used by
[`08-sendfile-static-assets.md`](../08-sendfile-static-assets.md) — the `GET /static/...` handler. Future `##ui` SSR output bundles go through the same path.
## v1 port source
[`reference/crates/wo-serve/src/sendfile.rs`](../../../reference/crates/wo-serve/src/sendfile.rs) (109 LOC) — `send_file(sock, path) -> Result` wrapping the loop + `EAGAIN` handling.

View file

@ -0,0 +1,97 @@
# 07 — `io_uring`
Ring-buffer based async I/O (Linux 5.1+, mature 5.11+). Two lock-free SPSC rings shared between userspace and kernel: submissions (SQEs) go in one, completions (CQEs) come out of the other. Batched, zero-syscall submission (with SQPOLL), zero-copy where the underlying op allows. Successor to `epoll` + `libaio` for the storage engine's WAL fsync path and — eventually — the HTTP server's accept/recv/send path.
**Not on the runtime's critical path in phases 02–08.** Phase 02 uses `epoll`. `io_uring` comes in during [Phase 3 — In-Memory Engine](../runtime/database/03-inmemory-engine.md) for the WAL's group-commit fsync loop. This card is the reference for that phase.
## Kernel source
| Path | What |
| --- | --- |
| [`reference/linux/io_uring/`](../../../reference/linux/io_uring/) | Whole subsystem. Start with `io_uring.c` (ring setup + submission/completion) and `fs.c` (fsync op). |
| [`reference/linux/io_uring/io_uring.c`](../../../reference/linux/io_uring/io_uring.c) | `SYSCALL_DEFINE2(io_uring_setup, ...)`, `SYSCALL_DEFINE6(io_uring_enter, ...)`, `SYSCALL_DEFINE4(io_uring_register, ...)`. |
| [`reference/linux/include/uapi/linux/io_uring.h`](../../../reference/linux/include/uapi/linux/io_uring.h) | `struct io_uring_sqe`, `io_uring_cqe`, `io_uring_params`, every `IORING_*` flag. |
## Man pages
`man 7 io_uring` (overview + entire ring model), `man 2 io_uring_setup`, `man 2 io_uring_enter`, `man 2 io_uring_register`. Also the `liburing` manual — Axboe's C library — useful for the per-op surface even if we don't link it.
## Rust FFI via `libc`
`libc` currently exposes the **constants** and the raw **syscall numbers** (`SYS_io_uring_setup`, `SYS_io_uring_enter`, `SYS_io_uring_register`), not wrapper functions. Invoke via `libc::syscall`:
```rust
use libc::{syscall, SYS_io_uring_setup, SYS_io_uring_enter, SYS_io_uring_register};
use libc::{mmap, munmap, MAP_SHARED, MAP_POPULATE, PROT_READ, PROT_WRITE};
// struct layouts from include/uapi/linux/io_uring.h — must mirror exactly
```
## Direct-syscall example (minimum viable ring)
```rust
// 1. setup — size is the number of SQEs; kernel rounds to power of 2
let mut params: io_uring_params = std::mem::zeroed();
// params.flags |= IORING_SETUP_SQPOLL; // kernel polls SQ — zero-syscall submit
let ring_fd = libc::syscall(SYS_io_uring_setup, 256u32, &mut params as *mut _) as i32;
// 2. mmap the three regions the kernel allocated
let sq_ring = libc::mmap(
std::ptr::null_mut(),
params.sq_off.array as usize + params.sq_entries as usize * 4,
PROT_READ | PROT_WRITE,
MAP_SHARED | MAP_POPULATE,
ring_fd,
IORING_OFF_SQ_RING,
);
let cq_ring = libc::mmap(..., IORING_OFF_CQ_RING);
let sqes = libc::mmap(..., IORING_OFF_SQES);
// 3. submit an fsync — fill an SQE and bump the SQ tail
let idx = *sq_tail & ring_mask;
let sqe = &mut *(sqes as *mut io_uring_sqe).add(idx as usize);
sqe.opcode = IORING_OP_FSYNC as u8;
sqe.fd = wal_fd;
sqe.user_data = commit_lsn; // your correlation key
*sq_tail = sq_tail.wrapping_add(1);
// 4. enter — tell the kernel to process N SQEs, optionally wait for completions
libc::syscall(SYS_io_uring_enter, ring_fd, 1u32, 1u32, IORING_ENTER_GETEVENTS, 0, 0);
// 5. reap a CQE
let idx = *cq_head & ring_mask;
let cqe = &*(cq_ring.add(params.cq_off.cqes as usize) as *const io_uring_cqe).add(idx as usize);
let lsn = cqe.user_data;
let err = cqe.res; // < 0 is -errno
*cq_head = cq_head.wrapping_add(1);
```
Full working code is ~200 LOC including error handling — see `liburing` source for the canonical shape.
## Key flags + ops
| | |
| --- | --- |
| `IORING_SETUP_SQPOLL` | Kernel thread polls the SQ — userspace writes SQEs with no syscall. One pinned kernel thread per ring. Needs `CAP_SYS_NICE` before 5.11. |
| `IORING_SETUP_IOPOLL` | Busy-poll completions from the NVMe device (no interrupts). Lower latency, higher CPU. Requires `O_DIRECT`. |
| `IORING_SETUP_SINGLE_ISSUER` | Optimisation when only one thread submits (Linux 6.0+). **Always set** in the single-threaded runtime. |
| `IORING_REGISTER_FILES` | Pre-register a set of fds with the ring — skips per-op fd-table lookup. Use it for the WAL fd. |
| `IORING_REGISTER_BUFFERS` | Pre-register userspace pages — skips per-op page pinning. Use for the WAL ring buffer. |
| `IOSQE_IO_LINK` | Chain SQEs — the second doesn't start until the first completes. Essential for WAL: `WRITE` linked to `FSYNC`. |
| `IORING_OP_WRITE`, `IORING_OP_FSYNC`, `IORING_OP_READ`, `IORING_OP_ACCEPT`, `IORING_OP_SEND`, `IORING_OP_RECV` | The ops that replace the phase-02 `epoll` + `read`/`write` dance. |
## Gotchas
- **Ring memory layout is ABI.** The kernel writes via the mmap'd regions; the `params.sq_off.*` / `cq_off.*` fields tell you the exact byte offsets. Hard-coding offsets breaks across kernel versions.
- **`user_data` is the correlation key.** The kernel echoes it back on the CQE untouched. Use it to thread whatever identifier you need (LSN, request id, subscriber id).
- **No ordering between unlinked SQEs.** Independent writes can complete in any order. Use `IOSQE_IO_LINK` for ordering (write-then-fsync) or per-fd serialization (one fd at a time).
- **CQE `res` is `-errno` on failure**, not `-1` + `errno`. Sign-extend it as `i32`, negate for the error code.
- **Always check `sq_ring_mask` from `params.sq_off.ring_mask`** before indexing. Never assume size 256.
- **`io_uring` has had CVE fights.** Some hosting providers and container runtimes disable it (`io_uring_disabled=2`). Detect at runtime and fall back to `epoll` — the phase-02 event loop stays useful forever as a compatibility path.
## Used by
Phase 3 of the database series — see [`docs/runtime/database/03-inmemory-engine.md`](../runtime/database/03-inmemory-engine.md). Specifically the WAL fsync path: link `WRITE` → `FSYNC` SQEs, submit many per tick, reap completions to ack committed transactions. Also the natural upgrade target for the HTTP server once Phase 4 adds the native wire protocol.
## v1 port source
**None.** The v1 crates predate `io_uring` and use `epoll` + blocking `fsync` on a WAL-writer thread. This crate will be new code in `crates/wal/` when Phase 3 activates.

100
docs/plan/linux/08-mmap.md Normal file
View file

@ -0,0 +1,100 @@
# 08 — `mmap` + `madvise`
Memory-map a file (or anonymous region) into the process's address space. The kernel manages the page cache; your code sees a `&[u8]` slice. `madvise` hints the kernel about access patterns so it can pre-fetch sequentially, evict aggressively after scans, or map huge pages.
Central to Phase 3's storage engine: segment files are `mmap`ed read-only for O(1)/O(log n) indexed lookups without copying bytes into heap memory.
## Kernel source
| Path | What |
| --- | --- |
| [`reference/linux/mm/mmap.c`](../../../reference/linux/mm/mmap.c) | VMA creation, `SYSCALL_DEFINE6(mmap, ...)`, `SYSCALL_DEFINE2(munmap, ...)`. |
| [`reference/linux/mm/madvise.c`](../../../reference/linux/mm/madvise.c) | `SYSCALL_DEFINE3(madvise, ...)` + every `MADV_*` handler. |
| [`reference/linux/include/uapi/linux/mman.h`](../../../reference/linux/include/uapi/linux/mman.h) | `MAP_*` flags, huge-page sizing macros. |
| POSIX `<sys/mman.h>` | The other half of the constants (`PROT_*`, `MADV_*`). Usually folded into `linux/mman.h` by libc. |
## Man pages
`man 2 mmap`, `man 2 madvise`, `man 2 munmap`, `man 2 msync`, `man 2 mprotect`.
## Rust FFI via `libc`
```rust
use libc::{mmap, munmap, madvise, msync, mprotect};
use libc::{PROT_READ, PROT_WRITE, PROT_NONE, PROT_EXEC};
use libc::{MAP_SHARED, MAP_PRIVATE, MAP_ANONYMOUS, MAP_FIXED};
use libc::{MAP_POPULATE, MAP_HUGETLB, MAP_HUGE_2MB, MAP_HUGE_1GB};
use libc::{MADV_SEQUENTIAL, MADV_RANDOM, MADV_WILLNEED, MADV_DONTNEED};
use libc::{MADV_HUGEPAGE, MADV_NOHUGEPAGE, MS_SYNC, MS_ASYNC};
```
## Direct-syscall example
```rust
unsafe {
// 1. Map a segment file read-only. Use MAP_POPULATE to pre-fault all pages
// so lookups don't hit a minor page fault mid-request.
let fd = libc::open(path.as_ptr(), libc::O_RDONLY);
let len = libc::lseek(fd, 0, libc::SEEK_END) as usize;
let ptr = libc::mmap(
std::ptr::null_mut(),
len,
libc::PROT_READ,
libc::MAP_SHARED | libc::MAP_POPULATE,
fd,
0,
);
if ptr == libc::MAP_FAILED {
return Err(io::Error::last_os_error());
}
// 2. Hint access pattern — sequential scan for a full compaction pass,
// random for indexed lookups. MADV_DONTNEED after a scan releases page cache pressure.
libc::madvise(ptr, len, libc::MADV_RANDOM);
// 3. Use it as a byte slice
let slice: &[u8] = std::slice::from_raw_parts(ptr as *const u8, len);
let record = &slice[offset..offset + record_len];
// 4. Clean up
libc::munmap(ptr, len);
libc::close(fd);
}
```
## Key flags
| Flag | Meaning |
| --- | --- |
| `PROT_READ` / `PROT_WRITE` | Obvious. Combine as needed. `PROT_NONE` makes a guard page. |
| `MAP_SHARED` | Writes go back to the file. Required for write-through semantics (WAL staging into a `mmap`ed region). |
| `MAP_PRIVATE` | Copy-on-write. Writes never hit the file. Use for read-only segments where you want CoW safety. |
| `MAP_POPULATE` | Pre-fault the whole mapping at `mmap` time. Trades boot latency for zero-fault request path. **Use it for hot segments.** |
| `MAP_HUGETLB` / `MAP_HUGE_2MB` | Back with huge pages. 512× fewer TLB entries for a 64 GB arena. Requires `vm.nr_hugepages` configured. |
| `MAP_ANONYMOUS` | Not file-backed — just zero-initialised pages. Used for arenas the engine allocates internally. |
| `MAP_FIXED` | Place at the exact requested address. Dangerous — will silently overwrite existing mappings. Only when you know what you're doing (e.g. placing guard pages). |
| `madvise` | Meaning |
| --- | --- |
| `MADV_SEQUENTIAL` | "I'll read sequentially." Kernel prefetches ahead, drops pages behind. Full scans, compaction. |
| `MADV_RANDOM` | "Lookups will be random." Kernel disables read-ahead. Index lookups. |
| `MADV_WILLNEED` | "Bring these pages in now." Async prefetch for an upcoming working set. |
| `MADV_DONTNEED` | "I'm done; drop these pages." Frees page-cache slots immediately — good after a scan to avoid polluting the cache. |
| `MADV_HUGEPAGE` | Opt this range into Transparent Huge Pages. |
## Gotchas
- **Shared writable mappings and `fsync`.** Writes to a `MAP_SHARED` region are *not* durable until you `msync(MS_SYNC)` or `fsync` the underlying fd. For write paths that need durability, prefer explicit `pwrite` — don't rely on `msync` for the hot path.
- **`SIGBUS` on truncated files.** If the file shrinks beneath your mapping, accesses past the new end raise `SIGBUS`. Arrange for sealed files (`memfd_create(MFD_ALLOW_SEALING)` + `F_SEAL_SHRINK`) or just don't truncate.
- **Page faults block the single thread.** In a single-threaded runtime, a minor fault during a request freezes the whole loop. `MAP_POPULATE` at boot sidesteps this for hot data. Use `mlockall(MCL_CURRENT \| MCL_FUTURE)` if faults must never happen — but that requires `CAP_IPC_LOCK` or `RLIMIT_MEMLOCK` headroom.
- **`madvise` hints are advice, not commands.** The kernel may ignore them under memory pressure. Don't rely on them for correctness; only for perf.
- **Huge pages need config.** `vm.nr_hugepages` has to have enough entries for your arenas. Startup-time check, not request-time.
## Used by
Phase 3 of the database series — see [`docs/runtime/database/03-inmemory-engine.md`](../runtime/database/03-inmemory-engine.md) § Linux Tuning Checklist. The relational B+ tree, the LSM memtables' on-disk segments, and the document store's arenas all live behind `mmap`.
## v1 port source
**None directly** — v1's wo-seg reads with `pread`, not `mmap`. This is new code territory for the new storage engine.

View file

@ -0,0 +1,96 @@
# 09 — `fallocate` + positional I/O (`pread`, `pwritev2`)
`fallocate` pre-allocates disk space for a file without writing any bytes — lets the filesystem commit to a contiguous extent, so later writes don't fragment and can't fail mid-operation due to disk pressure. `pread` / `pwritev2` read and write at an explicit offset without touching the file's cursor — letting many concurrent readers share one fd safely.
Together they form the backbone of the storage engine's on-disk layout: segment files are pre-allocated to their target size at creation, then written into via `pwritev2`; readers hit them via `pread` or `mmap` (see [08-mmap.md](./08-mmap.md)).
## Kernel source
| Path | What |
| --- | --- |
| [`reference/linux/fs/open.c`](../../../reference/linux/fs/open.c) | `SYSCALL_DEFINE4(fallocate, ...)`. The syscall delegates to `file->f_op->fallocate` — per-filesystem. |
| [`reference/linux/fs/read_write.c`](../../../reference/linux/fs/read_write.c) | `SYSCALL_DEFINE4(pread64, ...)`, `SYSCALL_DEFINE4(pwrite64, ...)`, `SYSCALL_DEFINE6(pwritev2, ...)`. |
| [`reference/linux/include/uapi/linux/falloc.h`](../../../reference/linux/include/uapi/linux/falloc.h) | `FALLOC_FL_*` flags. |
## Man pages
`man 2 fallocate`, `man 2 pread`, `man 2 pwrite`, `man 2 pwritev2`.
## Rust FFI via `libc`
```rust
use libc::{fallocate, pread, pread64, pwrite, pwrite64, pwritev2, iovec, off_t};
use libc::{FALLOC_FL_KEEP_SIZE, FALLOC_FL_PUNCH_HOLE, FALLOC_FL_ZERO_RANGE,
FALLOC_FL_COLLAPSE_RANGE, FALLOC_FL_INSERT_RANGE};
// pwritev2 has its own flags:
use libc::{RWF_SYNC, RWF_DSYNC, RWF_HIPRI, RWF_NOWAIT, RWF_APPEND};
```
## Direct-syscall example
```rust
unsafe {
let fd = libc::open(path.as_ptr(), libc::O_RDWR | libc::O_CREAT, 0o644);
// 1. Pre-allocate 64 MB so writes can't fail with ENOSPC later.
// Omit FALLOC_FL_KEEP_SIZE to make the size reflect the allocation
// (common for WAL rings); include it to reserve space without growing
// the file's apparent size (common for LSM SSTables before finalization).
if libc::fallocate(fd, 0, 0, 64 * 1024 * 1024) < 0 {
return Err(io::Error::last_os_error());
}
// 2. Positional write from a scattered set of buffers — no shared cursor,
// no extra copy to concat. pwritev2 also accepts per-call flags like
// RWF_SYNC for integrity barriers on a specific write.
let iovs = [
iovec { iov_base: header.as_ptr() as *mut _, iov_len: header.len() },
iovec { iov_base: body.as_ptr() as *mut _, iov_len: body.len() },
];
let offset: off_t = 4096;
let written = libc::pwritev2(fd, iovs.as_ptr(), iovs.len() as i32, offset, libc::RWF_DSYNC);
// 3. Concurrent readers hit the same fd with pread — no locking needed,
// no interference with the writer's implicit cursor (there isn't one).
let mut buf = vec![0u8; 8192];
let n = libc::pread(fd, buf.as_mut_ptr() as *mut _, buf.len(), record_offset as off_t);
}
```
## Key flags
### `fallocate` modes (first arg after fd)
| Flag (bitwise-OR into `mode`) | Meaning |
| --- | --- |
| `0` (default) | Allocate and extend the file if offset+len > size. WAL growth. |
| `FALLOC_FL_KEEP_SIZE` | Allocate without changing the reported file size. SSTables-in-progress. |
| `FALLOC_FL_PUNCH_HOLE` (+ `KEEP_SIZE`) | Release blocks in a range. Sparse-file compaction. |
| `FALLOC_FL_ZERO_RANGE` | Zero a byte range efficiently (filesystem marks it unwritten). Faster than `pwrite(zeros)` for segment reset. |
| `FALLOC_FL_COLLAPSE_RANGE` / `FALLOC_FL_INSERT_RANGE` | Move extents — remove or create holes without re-writing. Log compaction. Requires filesystem support (ext4 / xfs). |
### `pwritev2` flags (6th arg)
| Flag | Meaning |
| --- | --- |
| `RWF_SYNC` | `O_SYNC` semantics for this call only — data + metadata barrier. |
| `RWF_DSYNC` | `O_DSYNC` semantics — data barrier, metadata not guaranteed. **WAL commits.** |
| `RWF_HIPRI` | Best-effort high priority; polls for completion on NVMe. Pair with `IOPOLL` rings. |
| `RWF_NOWAIT` | Return `EAGAIN` rather than blocking if the kernel would sleep. Useful for async paths. |
| `RWF_APPEND` | Equivalent to `O_APPEND` for this call, even if the fd wasn't opened with it. |
## Gotchas
- **`fallocate` is per-filesystem.** ext4 and xfs support every flag above; tmpfs supports `0` but not `PUNCH_HOLE`; network filesystems may silently no-op. Check `statfs(2) / f_type` at startup if cross-fs portability matters — or just require ext4/xfs.
- **`pread`/`pwrite` don't update the fd's file offset.** Great for concurrent readers. If you have code that alternates seek+read, don't mix it with `pread`-based readers on the same fd — it'll work but the mental model gets confusing.
- **`pwritev2` is Linux 4.6+.** Older kernels need `pwritev` + `fdatasync`. Not a concern for the runtime's target kernel (5.1+ for `io_uring` anyway).
- **`RWF_DSYNC` ≠ fsync.** It's a per-call data barrier. If you've opened with `O_DIRECT`, pages bypass the cache and the barrier is cheap. Otherwise still cheaper than a full `fsync` because only this call's metadata barrier is enforced.
- **Filesystem ENOSPC is silent in `fallocate` on some fs**. It can return 0 then fail at first write. Test against your target filesystem; don't assume the guarantee.
## Used by
Phase 3 of the database series — WAL pre-allocation, SSTable extent reservation, segment punching for compaction. Also [`07-io_uring.md`](./07-io_uring.md) pairs beautifully with positional I/O: `IORING_OP_WRITE` / `IORING_OP_READ` take an offset, so they're `pwrite`/`pread` under the hood.
## v1 port source
**None.** V1's wo-seg writes sequentially with `write` + `sync_all`; no pre-allocation. New territory for the v2 storage engine.

View file

@ -0,0 +1,93 @@
# 10 — `pidfd`
Process identifier as a file descriptor. `pidfd_open(pid)` returns an fd that becomes readable when the process exits — reaping is a `read`, not a `waitpid` race. Lets the supervisor track child processes on the same `epoll` loop that drives everything else, with no PID-reuse bugs (an fd can't be recycled to point at a different process).
Not on the runtime's critical path today; useful when the runtime grows a supervisor (spawning workers, running `wo build` subprocesses, managing a sharded engine's child processes). Worth knowing the shape now so Phase 9+ doesn't reinvent `waitpid`.
## Kernel source
| Path | What |
| --- | --- |
| [`reference/linux/kernel/pid.c`](../../../reference/linux/kernel/pid.c) | `SYSCALL_DEFINE2(pidfd_open, ...)`, `pidfd_create`, `pidfd_pid`. |
| [`reference/linux/kernel/signal.c`](../../../reference/linux/kernel/signal.c) | `SYSCALL_DEFINE4(pidfd_send_signal, ...)`. |
| [`reference/linux/kernel/fork.c`](../../../reference/linux/kernel/fork.c) | `clone3` — the only way to get a pidfd atomically with spawn. |
| [`reference/linux/include/uapi/linux/pidfd.h`](../../../reference/linux/include/uapi/linux/pidfd.h) | `PIDFD_*` flags. |
## Man pages
`man 2 pidfd_open`, `man 2 pidfd_send_signal`, `man 2 pidfd_getfd`, `man 2 clone3`.
## Rust FFI via `libc`
`libc` doesn't have direct wrappers for every pidfd syscall. Use `libc::syscall` with the numeric id:
```rust
use libc::{syscall, SYS_pidfd_open, SYS_pidfd_send_signal, SYS_pidfd_getfd};
use libc::{SYS_clone3, clone_args}; // clone3 also goes via syscall(SYS_clone3, ...)
use libc::{PIDFD_NONBLOCK, PIDFD_THREAD}; // Linux 5.10+
```
## Direct-syscall example
```rust
unsafe {
// Open a pidfd for an already-running child (race-prone: the child could
// have exited and the PID been reused before this call — fine for the
// self-pid, risky for arbitrary children)
let pidfd = libc::syscall(libc::SYS_pidfd_open, child_pid, 0);
if pidfd < 0 { return Err(io::Error::last_os_error()); }
// Register on epoll. EPOLLIN fires exactly once, when the process exits.
let mut ev = libc::epoll_event {
events: libc::EPOLLIN as u32,
u64: child_pid as u64, // your correlation key
};
libc::epoll_ctl(epfd, libc::EPOLL_CTL_ADD, pidfd as i32, &mut ev);
// On the readiness event, waitid collects the exit status
let mut info: libc::siginfo_t = std::mem::zeroed();
libc::waitid(libc::P_PIDFD, pidfd as u32, &mut info, libc::WEXITED);
// Send a signal via the fd — no PID race
libc::syscall(libc::SYS_pidfd_send_signal, pidfd, libc::SIGTERM, std::ptr::null::<libc::siginfo_t>(), 0);
libc::close(pidfd as i32);
}
```
For race-free child spawn, use `clone3(CLONE_PIDFD)`:
```rust
let mut pidfd: i32 = -1;
let args = libc::clone_args {
flags: libc::CLONE_PIDFD as u64,
pidfd: &mut pidfd as *mut _ as u64,
// ... stack, tls, etc.
..std::mem::zeroed()
};
let child = libc::syscall(libc::SYS_clone3, &args, std::mem::size_of::<libc::clone_args>());
```
## Key flags
| Flag | Meaning |
| --- | --- |
| `PIDFD_NONBLOCK` | Non-blocking reads; combine with epoll. Linux 5.10+. |
| `PIDFD_THREAD` | Open a pidfd for a TID, not just a PID. Rarely needed. |
| `CLONE_PIDFD` | Passed to `clone3` — kernel stores the new pidfd at `args.pidfd`. The atomic way to get a pidfd without a race window. |
## Gotchas
- **Kernel version matters.** `pidfd_open` is 5.3+; `PIDFD_NONBLOCK` is 5.10+; `pidfd_getfd` (steal an fd from another process) is 5.6+. Check your target range.
- **PID reuse race on manual `pidfd_open`.** If the child exited and something else was spawned with the same PID between `fork` and `pidfd_open`, you hold a pidfd for the wrong process. `clone3(CLONE_PIDFD)` eliminates the window; for arbitrary external processes, `pidfd_open` is best-effort.
- **`epoll` fires once per exit.** The pidfd stays readable forever after, which is sometimes useful (always-ready means always-wake-me), sometimes annoying (you must `EPOLL_CTL_DEL` or the loop spins).
- **`pidfd_send_signal` refuses to signal the init process** (`pid == 1`). Not a concern unless running as PID 1 in a container.
- **Permissions.** You can only pidfd-open a child of yours, or a process in the same session, or with `CAP_KILL`.
## Used by
Not used in phases 02–08. Future supervisor work: the `wo build` subprocess, a hypothetical `wo dev` hot-reload supervisor, or a sharded engine's child monitoring — all would register child pidfds on the existing event loop.
## v1 port source
**None.** V1 doesn't spawn processes.

View file

@ -0,0 +1,102 @@
# 11 — `memfd_create`
Anonymous memory-backed file descriptor. `memfd_create(name, flags)` returns an fd that points at a region of RAM (tmpfs-like) with no filesystem path. Size it with `ftruncate`, fill with writes or `mmap`, pass the fd across processes for zero-copy IPC, or use it as a scratchpad that vanishes on close. With `MFD_ALLOW_SEALING`, you can freeze the mapping so consumers can safely `mmap` it without worrying about shrinks or truncations.
Useful for the storage engine's transient work: building an index in memory before atomically renaming into place, staging a large response body that needs to be `sendfile`'d, or sharing a read-only snapshot with a child process.
## Kernel source
| Path | What |
| --- | --- |
| [`reference/linux/mm/memfd.c`](../../../reference/linux/mm/memfd.c) | `SYSCALL_DEFINE2(memfd_create, ...)` + seal ops. |
| [`reference/linux/include/uapi/linux/memfd.h`](../../../reference/linux/include/uapi/linux/memfd.h) | `MFD_*` flags. |
| [`reference/linux/include/uapi/linux/fcntl.h`](../../../reference/linux/include/uapi/linux/fcntl.h) | `F_ADD_SEALS`, `F_GET_SEALS`, `F_SEAL_*` constants. Sealing is a `fcntl(F_ADD_SEALS, ...)` operation on the memfd. |
## Man pages
`man 2 memfd_create`, `man 2 fcntl` (for the seal operations).
## Rust FFI via `libc`
```rust
use libc::{memfd_create, ftruncate, mmap, munmap};
use libc::{MFD_CLOEXEC, MFD_ALLOW_SEALING, MFD_HUGETLB, MFD_NOEXEC_SEAL};
use libc::{F_ADD_SEALS, F_GET_SEALS};
use libc::{F_SEAL_SEAL, F_SEAL_SHRINK, F_SEAL_GROW, F_SEAL_WRITE, F_SEAL_FUTURE_WRITE};
```
## Direct-syscall example
```rust
unsafe {
// 1. Create the anonymous fd. The name is for /proc/self/fd listings; not a path.
let name = std::ffi::CString::new("wo-index-build").unwrap();
let fd = libc::memfd_create(name.as_ptr(), libc::MFD_CLOEXEC | libc::MFD_ALLOW_SEALING);
if fd < 0 { return Err(io::Error::last_os_error()); }
// 2. Size it, then write into it (or mmap and write directly).
libc::ftruncate(fd, 4 * 1024 * 1024); // 4 MiB
let ptr = libc::mmap(
std::ptr::null_mut(),
4 * 1024 * 1024,
libc::PROT_READ | libc::PROT_WRITE,
libc::MAP_SHARED,
fd,
0,
);
// ... build an index into the mapping ...
// 3. Seal it so consumers can mmap RO without races.
// SHRINK prevents ftruncate-down; GROW prevents ftruncate-up;
// WRITE prevents further writes; SEAL prevents more seals from being added.
libc::fcntl(fd, libc::F_ADD_SEALS,
libc::F_SEAL_SHRINK | libc::F_SEAL_GROW
| libc::F_SEAL_WRITE | libc::F_SEAL_SEAL);
// 4. Pass fd to consumers via SCM_RIGHTS or clone3(CLONE_FILES).
// They mmap it RO and treat the contents as immutable.
// 5. Close the final fd: when the last consumer closes theirs, the kernel
// frees the memory. No filesystem cleanup.
libc::munmap(ptr, 4 * 1024 * 1024);
libc::close(fd);
}
```
## Key flags
### `memfd_create` flags
| Flag | Meaning |
| --- | --- |
| `MFD_CLOEXEC` | Close on exec. Always set. |
| `MFD_ALLOW_SEALING` | Permit later `F_ADD_SEALS` calls on this fd. Required if consumers `mmap` read-only and you want to promise immutability. |
| `MFD_HUGETLB` | Back with huge pages. Pair with `MFD_HUGE_2MB` or `MFD_HUGE_1GB`. Good for large index arenas. |
| `MFD_NOEXEC_SEAL` | Linux 6.3+: apply `F_SEAL_EXEC` automatically. Prevents the memfd from ever being mmap'd executable — a mild defence against exploitation if you ever accept untrusted data. |
### Seals (`fcntl(F_ADD_SEALS, ...)`)
| Seal | Meaning |
| --- | --- |
| `F_SEAL_SEAL` | No more seals can be added. Always the last one you apply. |
| `F_SEAL_SHRINK` | File size cannot decrease. Required before safe `mmap` by other processes. |
| `F_SEAL_GROW` | File size cannot increase. |
| `F_SEAL_WRITE` | No further writes permitted. Turns the memfd into a read-only shared region. |
| `F_SEAL_FUTURE_WRITE` | Linux 5.1+: prevents future writes, but keeps existing writable mappings functional. Softer than `F_SEAL_WRITE`. |
| `F_SEAL_EXEC` | Linux 6.3+: prevents mapping with `PROT_EXEC`. Security hardening. |
## Gotchas
- **Not on disk — the memory counts against `RLIMIT_MEMLOCK` / cgroup memory.** A 64 GB memfd is a 64 GB RAM commitment. Plan capacity.
- **Sealing is one-way.** Once `F_SEAL_WRITE` is on, the fd is read-only forever. The usual pattern: build → seal → share.
- **Existing writable mappings survive `F_SEAL_WRITE`.** The seal blocks *new* `mmap(PROT_WRITE)` and `write()`. Existing `MAP_SHARED` mappings still let you write. Use `F_SEAL_FUTURE_WRITE` if you want the existing writers to keep working while preventing new ones.
- **Sharing between processes.** The standard mechanisms: `SCM_RIGHTS` over a Unix socket, or `clone3(CLONE_FILES)` to inherit the fd table. Both preserve fd identity — the receiver sees the same memfd.
- **Not for durable data.** The memfd dies with its last fd. If you need the contents persisted, write them to a real file before closing.
## Used by
Phase 3 of the database series — index-build-then-swap (mentioned in [03-inmemory-engine.md § Recovery](../runtime/database/03-inmemory-engine.md#recovery) as "build a .seg index in memory before atomically swapping it to disk"). Also any future IPC story with worker processes (Phase 6 full-stack with multiple render workers, say).
## v1 port source
**None.** V1 doesn't use anonymous memory — all index work hits the filesystem directly.

65
docs/runtime/database.md Normal file
View file

@ -0,0 +1,65 @@
# Document & Graph Database — Design Series
A seven-phase design series that starts with "should writeonce use a document or graph database?" and arrives at a full-stack declarative application platform — then plans the migration from writeonce's current flat-file store to that platform.
> **Start here if you're new:** [wo-language.md](./wo-language.md) — the user-facing overview of what writeonce *is* (a programming language with DB + HTTP in its runtime, Go-style toolchain). This series is the engineering plan that gets you there.
Each phase is self-contained and shippable on its own. Every phase after Phase 1 builds on the previous ones.
## Phases
| Phase | Document | Summary |
| --- | --- | --- |
| **1** | [Database Evaluation](./database/01-evaluation.md) | Evaluate CouchDB, Postgres, Neo4j, petgraph against writeonce's needs. Decision: no external DB; add petgraph-backed `mappings.idx`. |
| **2** | [The `.wo` Language & ACID Engine](./database/02-wo-language.md) | Design a two-layer `.wo` language (unified `type` schema layer + hybrid SQL/Cypher query layer with fixed glue) for an e-commerce platform with ACID transactions across relational, document, and graph storage. |
| **3** | [In-Memory Engine](./database/03-inmemory-engine.md) | RAM-primary, SSD-durable storage engine using `io_uring`, `mlockall`, group commit. 64 GB Linux machine. |
| **4** | [Client API: Wire Protocol & Subscriptions](./database/04-client-api.md) | Native binary protocol + GraphQL over WebSocket. Subscription engine inside the transaction coordinator — no polling anywhere. |
| **5** | [Go Client SDK](./database/05-go-sdk.md) | Typed Go client with subscription-first design. `gen` codegen from `.wo` schema. Subscribe to a live query in 5 lines. |
| **6** | [Low-Code Full-Stack](./database/06-lowcode-fullstack.md) | Expand `.wo` into an application language (like SAP CDS): `##ui`, `##logic`, `##policy`, `##service` blocks compiled into a single binary. |
| **7** | [Replacing `wo-seg`](./database/07-wo-seg-migration.md) | Phased coexistence plan: abstract the article store behind a trait, stand up the `.wo` engine as a second impl, dual-run, cut over, decommission `wo-seg`. |
## Build Order
```
Phase 1: Evaluation ← writeonce today (blog, flat files)
│
▼
Phase 2: .wo Language ← query language + ACID engine design
│
▼
Phase 3: In-Memory Engine ← RAM-primary storage, io_uring, WAL
│
▼
Phase 4: Client API ← wire protocol, subscriptions, GraphQL
│
▼
Phase 5: Go SDK ← typed client, codegen, subscription channels
│
▼
Phase 6: Low-Code Full-Stack ← .wo as application DSL, UI generation, CLI
│
▼
Phase 7: Replace wo-seg ← trait abstraction, dual-run, cutover, decommission
```
## Cumulative Scope
| After Phase | What exists | Rough cumulative effort |
| --- | --- | --- |
| 1 | petgraph `mappings.idx` in writeonce | ~200 lines |
| 2 | `.wo` parser + planner + ACID engine prototype | 5–10 months |
| 3 | In-memory engine with io_uring durability | 8–16 months |
| 4 | Wire protocol + subscription engine | 14–28 months |
| 5 | Go SDK with typed subscriptions | 16–32 months |
| 6 | Full-stack low-code platform | 28–56 months |
| 7 | `wo-seg` replaced by the `.wo` engine in writeonce | +2–4 months on top of Phase 2 arrival |
## Related Documents
- [wo-language.md](./wo-language.md) — writeonce as a programming language: toolchain, hello-world, stdlib, client model
- [surreal-case-study.md](./surreal-case-study.md) — SurrealDB runtime analysis; why writeonce does not use a multi-model DB for live queries
- [async.md](./async.md) — custom async runtime using Linux kernel primitives
- [05-datalayer.md](../05-datalayer.md) — current `.seg` + `.idx` implementation (8 crates, 44 tests)
- [03-data.md](../03-data.md) — data layer design with subscription model
- [06-markdown-render.md](../06-markdown-render.md) — markdown-first content model
- [ai-agents-content-management.md](../future-scope/ai-agents-content-management.md) — the `mappings` feature that started this series

View file

@ -0,0 +1,272 @@
# Phase 1 — Database Evaluation
> Why build a custom database instead of adopting CouchDB, Postgres, Neo4j, or an in-memory graph library?
**Next**: [Phase 2 — The `.wo` Language & ACID Engine](./02-wo-language.md) | **Index**: [database.md](../database.md)
---
Reference repositories:
- [github.com/apache/couchdb](https://github.com/apache/couchdb) — document DB with Mango query language
- [github.com/postgres/postgres](https://github.com/postgres/postgres) — relational DB with JSONB
- [github.com/neo4j/neo4j](https://github.com/neo4j/neo4j) — native graph DB, Cypher query language
- [github.com/petgraph/petgraph](https://github.com/petgraph/petgraph) — Rust in-memory graph library (NetworkX analogue)
## The Question
writeonce articles have two shapes at once: they are **documents** (per-article JSON metadata + markdown body, per [06-markdown-render.md](../../06-markdown-render.md)) and they form a **graph** (the `mappings` field — `related`, `prerequisite`, `series`, `supersedes`, `references` — per [ai-agents-content-management.md](../../future-scope/ai-agents-content-management.md)).
Should writeonce adopt an off-the-shelf document or graph database to back these two shapes, or keep the flat-file `.seg` + `.idx` storage already implemented in [05-datalayer.md](../../05-datalayer.md)?
**Short answer: no external DB.** The dataset is small (hundreds of articles, not millions of rows), single-writer (author commits), and read-heavy. A full rebuild on change is cheap. The `mappings` graph fits entirely in RAM. External databases would add a process, a protocol, a driver, and a failure mode — none of which writeonce needs.
This doc walks through each option, what it buys, and why writeonce does not adopt it.
## The Data Shape
Every article is a pair of files in `content/{sys_title}/`:
```json
{
"sys_title": "linux-misc",
"title": "Linux Miscellaneous",
"published": true,
"author": "Shoney Arickathil",
"tags": ["Linux"],
"published_on": 1740950884,
"mappings": {
"related": ["auto-scale-gitlab-runner-using-aws-spot-instance"],
"prerequisite": ["linux-misc"],
"series": { "name": "gitlab-runner", "order": 2 }
}
}
```
Plus `{sys_title}.md` with the full article body.
The access patterns, per [05-datalayer.md](../../05-datalayer.md):
| Pattern | Frequency | Current Implementation |
| --- | --- | --- |
| `get_by_title(sys_title)` | Every `/blog/:sys_title` hit | `title.idx` hash — O(1) |
| `list_published(skip, limit)` | Homepage, pagination | `date.idx` sorted array — O(log n) |
| `list_by_tag(tag)` | Tag pages | `tags.idx` inverted index — O(1) + scan |
| `list_by_date_range` | Archive views | `date.idx` binary search |
| Graph traversal (`mappings`) | Agent workflows, "related" widget | **Not yet implemented** |
The last row is the gap. Everything else already works on `.seg` + `.idx`.
## Option 1: CouchDB + Mango
CouchDB is an HTTP-native document store. Each article JSON would be a document; Mango queries are JSON selectors expressive enough for most writeonce reads:
```json
{ "selector": { "published": true, "tags": { "$in": ["rust"] } }, "sort": [{ "published_on": "desc" }] }
```
What it buys:
- Built-in revisions (MVCC) — each edit gets a `_rev`
- Multi-master replication, which matters if content is authored from multiple machines
- A _changes feed, conceptually similar to writeonce's subscription model
What it costs:
- Separate Erlang process, HTTP driver, JSON over the wire on every read
- No native graph traversal — `mappings` would require client-side joins or view functions
- Duplicates what inotify + `.seg` already do (change feed, indexing)
### Comparison
| Aspect | CouchDB | writeonce |
| --- | --- | --- |
| Transport | HTTP per query | In-process function call |
| Storage | B-tree per database, append-only | `.seg` append-only, tombstoned records |
| Change feed | `_changes` HTTP long-poll | inotify → epoll → fd write |
| Index | View functions (JavaScript map/reduce), Mango indexes | `title.idx`, `date.idx`, `tags.idx` rebuilt on change |
| Revisions | Every write creates `_rev` | Git already does this for `content/` |
CouchDB's replication and revision model are attractive, but git already covers revision history for `content/`, and the dataset is too small to justify a separate daemon.
## Option 2: PostgreSQL + JSONB
Postgres with a `jsonb` column gives you SQL over the article metadata, GIN indexes for tag queries, and a real query planner. Schema would be roughly:
```sql
CREATE TABLE articles (
sys_title TEXT PRIMARY KEY,
metadata JSONB NOT NULL,
body_md TEXT NOT NULL,
updated TIMESTAMPTZ DEFAULT now()
);
CREATE INDEX ON articles USING GIN ((metadata->'tags'));
CREATE INDEX ON articles (((metadata->>'published_on')::bigint));
```
What it buys:
- Mature: WAL, point-in-time recovery, replication, tooling
- `jsonb_path_query` and `@>` containment make most writeonce queries one-liners
- Recursive CTEs (`WITH RECURSIVE`) can traverse `mappings` — adequate for shallow graphs
What it costs:
- A Postgres process, a driver (tokio-postgres or raw libpq), connection pooling — writeonce's [05-datalayer.md](../../05-datalayer.md) explicitly removed all of this
- JSONB query planning is excellent but still pays per-query cost that an in-process hash does not
- Recursive CTEs on deep mapping chains are slower than a RAM graph walk
### Comparison
| Aspect | Postgres JSONB | writeonce |
| --- | --- | --- |
| Process model | External daemon, TCP or unix socket | Single binary, single process |
| Query language | SQL + JSONB operators | Rust method calls on `Store` |
| Graph traversal | `WITH RECURSIVE` CTE | (Future) in-memory adjacency list |
| Backup | `pg_dump`, WAL archive | `content/` directory + git |
| Failure modes | Connection loss, pool exhaustion, vacuum stalls | File not found |
Postgres is the default reflex for "I have structured data." But writeonce's structured data is ~500 rows that change when the author saves a file. The mismatch is two orders of magnitude on the dataset size and one process on the deployment surface.
## Option 3: Neo4j — Native Graph DB
Neo4j models articles as nodes and `mappings` as typed edges. Cypher makes the traversals natural:
```cypher
MATCH (a:Article {sys_title: 'linux-misc'})-[:PREREQUISITE*1..3]->(p:Article)
RETURN p.sys_title
```
What it buys:
- Native graph storage — constant-time edge traversal regardless of dataset size
- Cypher is the right query language for the `mappings` problem
- Useful when the graph is the primary shape
What it costs:
- JVM process, Bolt protocol, driver — heaviest option on this list
- Document storage is secondary (properties on nodes), so article body and metadata are awkwardly split
- Dataset is tiny — the graph fits in a few KB of RAM; Neo4j's disk-backed adjacency is overkill
### Comparison
| Aspect | Neo4j | writeonce |
| --- | --- | --- |
| Storage | Native adjacency on disk | (Future) `HashMap<sys_title, Vec<Edge>>` in memory |
| Traversal | Cypher over Bolt | Rust iteration over adjacency |
| Transactions | ACID with MVCC | Full rebuild on change |
| Deployment | JVM daemon + driver | Statically linked binary |
| Fit for writeonce dataset | Overprovisioned by 3+ orders of magnitude | Right-sized |
## Option 4: In-Memory Graph — NetworkX / petgraph
NetworkX (Python) and petgraph (Rust) are _libraries_, not databases. You load the graph into process memory and traverse it directly. This is the model gestured at in [ai-agents-content-management.md line 181](../../future-scope/ai-agents-content-management.md) — "traversable knowledge graphs available on RAM."
For writeonce, petgraph is the right shape:
```rust
use petgraph::graph::DiGraph;
let mut g: DiGraph<String, MappingKind> = DiGraph::new();
// nodes: one per sys_title
// edges: one per mapping (related, prerequisite, series, ...)
```
What it buys:
- Zero extra processes — compiles into the `wo-store` crate
- Constant-time neighbor lookup, standard BFS / Dijkstra / SCC algorithms included
- Rebuilt cheaply on any `content/` change by walking the `.seg` and following `mappings`
What it costs:
- No persistence layer — but the graph is derived from JSON, same as `title.idx`, so it rebuilds on cold start for free
- No query language — but the traversals agents need (`all prerequisites of X`, `next article in series Y`) are short Rust functions
### Comparison
| Aspect | petgraph (in-memory) | Neo4j |
| --- | --- | --- |
| Location | Same process as `wo-store` | External JVM |
| Build time | Single pass over `.seg` | Bulk import via CSV / Cypher |
| Cost per traversal | Pointer chase in RAM | Network round trip + disk I/O |
| Query language | Rust | Cypher |
| Scales to | Millions of nodes in RAM (plenty of headroom) | Billions on disk |
For writeonce's hundreds of articles, petgraph is the correct answer. It fits the existing architecture — derived from `content/`, rebuildable, no external process — the same shape as `title.idx` already has.
## Decision Matrix
| Option | Dataset Fit | Process Count | Graph Support | Matches writeonce Philosophy |
| --- | --- | --- | --- | --- |
| CouchDB Mango | Overprovisioned | +1 (Erlang) | Manual joins | No — external daemon |
| Postgres JSONB | Overprovisioned | +1 (Postgres) | Recursive CTE | No — external daemon |
| Neo4j | Overprovisioned | +1 (JVM) | Native, excellent | No — external daemon |
| petgraph in-memory | Right-sized | 0 | Native, in-process | **Yes** |
| Current `.seg` + `.idx` | Right-sized | 0 | None yet | Already here |
## Proposed Addition: `mappings.idx` Backed by petgraph
Per [05-datalayer.md](../../05-datalayer.md), indexes live alongside `.seg`. Add a fourth index:
```
data/
articles.seg
index/
title.idx # existing — O(1) sys_title lookup
date.idx # existing — sorted by published_on
tags.idx # existing — inverted index by tag
mappings.idx # NEW — serialized petgraph adjacency
```
New `wo-graph` crate (or extension to `wo-index`):
```rust
pub struct MappingGraph {
graph: DiGraph<SysTitle, MappingKind>,
by_title: HashMap<String, NodeIndex>,
}
impl MappingGraph {
pub fn neighbors(&self, sys_title: &str, kind: MappingKind) -> Vec<&str> { ... }
pub fn prerequisites_transitive(&self, sys_title: &str) -> Vec<&str> { ... }
pub fn series(&self, name: &str) -> Vec<&str> { ... } // ordered by .order
}
```
Rebuild on every `Store::rebuild()`. Drop and recompute on any `ContentChange` — the graph is small enough that incremental updates are not worth the bug surface.
## Key Takeaways
1. **A document+graph database is the right model — but not the right dependency.** writeonce's articles really are documents with graph edges. That does not mean importing CouchDB, Postgres, or Neo4j. It means writing ~200 lines that give you the document and graph operations the product actually uses.
2. **Dataset size dictates architecture.** SurrealDB, Postgres, and Neo4j all assume millions of rows and concurrent writers. writeonce has hundreds of articles and one author. The gap is where external databases become overhead, not infrastructure.
3. **Git already solves the problems CouchDB's revisions solve.** Content lives in `content/` under version control. `_rev`, replication, and change history are the author's git history.
4. **Query languages are a cost, not a benefit, at this scale.** Mango selectors, JSONB operators, and Cypher exist because production queries are written by humans against large evolving datasets. writeonce's queries are fixed (`get_by_title`, `list_by_tag`, `list_published`) and written once in Rust.
5. **The graph belongs in RAM.** Article `mappings` form a small, mostly static DAG. petgraph holds it with zero protocol overhead and rebuilds from `content/` on cold start — same lifecycle as `title.idx`.
## Reference
If the graph side of this ever grows beyond what petgraph comfortably handles, the reference points are:
```bash
git submodule add https://github.com/neo4j/neo4j.git references/neo4j
git submodule add https://github.com/apache/couchdb.git references/couchdb
git submodule add https://github.com/petgraph/petgraph.git references/petgraph
```
Key files to study:
- `petgraph/src/graph_impl/` — adjacency list implementation, the smallest viable graph backend
- `couchdb/src/mango/` — Mango query compilation, if a selector-style query API ever becomes useful
- `neo4j/community/cypher/` — Cypher planner, for how a real graph query language is structured
See also:
- [surreal-case-study.md](../surreal-case-study.md) — why writeonce does not use a multi-model DB for live queries
- [05-datalayer.md](../../05-datalayer.md) — current `.seg` + `.idx` implementation
- [ai-agents-content-management.md](../../future-scope/ai-agents-content-management.md) — the `mappings` feature this index supports

View file

@ -0,0 +1,435 @@
# Phase 2 — The `.wo` Language & ACID Engine
> A two-layer multi-paradigm language for an e-commerce platform with ACID transactions across relational, document, and graph storage.
**Previous**: [Phase 1 — Database Evaluation](./01-evaluation.md) | **Next**: [Phase 3 — In-Memory Engine](./03-inmemory-engine.md) | **Index**: [database.md](../database.md)
---
## Creating Your Own Runtime Language for Database
- language parser
- file saved as `database.wo`
Early sketch of the schema — three paradigms in one file:
```wo
##sql
#users
id bigint
name varchar
#article
title varchar
meta article-meta
##doc
#article-meta
##graph
#context-map
##sql-article
```
queries (relational)
```wo
SELECT name FROM users
SELECT meta.sys_title FROM article
```
queries (graph)
```wo
MATCH (a:article), (b:article)
WHERE a.meta.sys_title = "Alice" AND b.meta.sys_title = "Bob"
CREATE (a)-[:DEFINES]->(b)
```
That three-paradigm sketch is **the execution substrate** — the thing the engine actually parses and runs. But it is not the right *authoring* surface for a full-stack platform. The language is designed as two layers, described next.
## The Two-Layer Design
```
┌────────────────────────────────────────────────────────┐
│ Schema Layer — unified type DSL (SAP-CDS-style) │ ← source of truth
│ type User { ... } / type Purchase link … / policy … │ (Phases 5 + 6)
└──────────────────────┬─────────────────────────────────┘
│ compiled to ↓
┌──────────────────────▼─────────────────────────────────┐
│ Query Layer — hybrid SQL + Cypher with fixed glue │ ← execution substrate
│ SELECT / UPDATE / MATCH / CREATE / BEGIN … COMMIT │ (Phase 2 prototype)
└────────────────────────────────────────────────────────┘
```
Both layers are `.wo` files — same extension, same tooling, same parser front-end. They differ in role:
- The **schema layer** names the data model once. One `type` declaration per entity covers what the three paradigm blocks cover today (relational columns, embedded documents, graph edges) plus constraints, computed fields, policies, and triggers. It is the source of truth for codegen ([Phase 5](./05-go-sdk.md)) and for the full-stack blocks ([Phase 6](./06-lowcode-fullstack.md)).
- The **query layer** is the operational surface. SQL and Cypher stay as-is — they are universally legible, every backend developer already reads them — but five things are tightened so the three grammars share semantics (parameters, `RETURNING`, dotted paths, transactions, `LIVE`).
The two layers ship on different timelines. The query layer is Phase 2 (already prototyped at [`prototypes/wo-db/`](../../../prototypes/wo-db/)). The schema layer enters when Phase 5 codegen needs a single authoritative input.
### Why Not One Layer?
Two alternatives were considered and rejected:
- **Unified query language only** (EdgeQL-style path algebra replacing SQL and Cypher). Loses the Phase 2 adoption property — SQL+Cypher are universally legible; a novel path language is not. Reinvents 6+ years of EdgeDB planner work.
- **Three paradigm blocks only** (the original sketch). Works for Phase 2 but hits a wall at Phase 5: the SDK codegen has to invent a schema-above-schema layer anyway, because "a relational row with an embedded doc column plus an inverse graph edge" is one Go struct, not three. Also hits a wall at Phase 6: `##ui` / `##policy` / `##logic` want to attach to *entities*, not to tables-vs-collections-vs-edges.
The two-layer split captures the Phase 2 adoption win and the Phase 5/6 coherence win without committing to a single novel query grammar.
## Schema Layer — Unified Type DSL
One `type` construct describes an entity. The compiler chooses physical storage (relational page, document LSM, graph node) from the field declarations.
```wo
type User {
id: Id
email: Email @unique
meta: { name: Text, avatar: Url? } -- inline document
friends: multi User @edge(:FOLLOWS) -- anonymous edge (tag only)
purchased: multi Product via Purchase -- link type with props
orders: backlink Order.user -- inverse scalar ref
}
type Purchase link User -> Product { -- edge with properties
order: ref Order
qty: Int @check(> 0)
at: Timestamp = now()
}
type Order {
id: Id
user: ref User
status: Pending | Paid | Shipped | Refunded -- tagged union
line_items: [{ product: ref Product, qty: Int, unit: Money }]
total: Money = sum(line_items.*.qty * line_items.*.unit) -- computed
}
type Product {
id: Id
sku: SKU @unique
price: Money
meta: { title: Text, description: Markdown, images: [Url], reviews: [Review] }
inventory: { on_hand: Int @check(>= 0), reserved: Int = 0, reorder_at: Int }
}
type Review {
user: ref User
stars: Int @check(between 1 and 5)
body: Markdown
written_at: Timestamp = now()
}
```
**Type-system primitives:**
| Primitive | Example | Compiles to |
| --- | --- | --- |
| Scalar | `Int Float Text Bool Timestamp Id Url Email Markdown Money SKU Slug` | relational column |
| Optional | `Url?` | nullable column |
| Array | `[Url]` | relational array or doc array |
| Struct | `{ k: V, ... }` | embedded document column (doc engine) |
| Scalar ref | `ref User` | foreign-key column |
| Zero-prop edge | `multi User @edge(:FOLLOWS)` | graph edge, no properties |
| Link-with-props | `multi Product via Purchase` | graph edge + linked row (`type Purchase link ...`) |
| Inverse | `backlink Order.user` | computed inverse of `Order.user: ref User` |
| Tagged union | `Pending \| Paid \| Shipped` | enum column |
| Computed | `total: Money = sum(...)` | view or materialized view |
**Annotations:** `@unique @check(...) @default(...) @index @search @immutable`.
**Full-stack blocks attach to types** — no separate `##ui`/`##logic`/`##policy`/`##service` paradigm markers. See [Phase 6](./06-lowcode-fullstack.md) for details; shape:
```wo
type Article {
slug: Slug @unique
title: Text
body: Markdown
author: ref User
policy read anyone
policy write when author == $session.user
on update when old.status != "published" and new.status == "published"
do emit "article.published"(self)
service rest "/articles" expose list, get, subscribe
}
```
## Query Layer — Hybrid SQL + Cypher, Fixed Glue
Keep the syntax developers already know. Fix five things so the three grammars share semantics:
1. **One parameter rule.** `$name` everywhere — SQL, Cypher, document path expressions. Typed at prepare time from the surrounding function signature or session context.
2. **Cross-paradigm `RETURNING`.** `INSERT … RETURNING id AS oid` binds the alias into subsequent statements in the same `BEGIN … COMMIT`. Replaces `LAST_INSERT_ID()`.
3. **One path rule.** `a.b.c[i].d` reads the same inside SQL expressions, document `UPDATE … SET`, and Cypher projections (`RETURN u.meta.name`).
4. **One transaction block.** `BEGIN [SNAPSHOT|SERIALIZABLE] … [SAVEPOINT name; …] … COMMIT|ROLLBACK [TO name]`. No dialect split for stored procedures.
5. **One subscription prefix.** `LIVE <select|match>` returns a subscription handle. Same semantics on both sides.
Every e-commerce query mixes at least two paradigms:
```wo
-- Relational + document: "products under $50 with >4 stars"
SELECT id, meta.title, price_cents
FROM products
WHERE price_cents < 5000
AND AVG(meta.reviews[].stars) > 4.0;
-- Graph + relational: "top sellers among products similar to what I bought"
MATCH (me:user {id: $uid})-[:PURCHASED]->(p:product)-[:SIMILAR_TO]->(rec:product)
WHERE rec.inventory.on_hand > 0
ORDER BY rec.meta.reviews.count DESC
LIMIT 20;
-- All three: atomic checkout, with RETURNING replacing LAST_INSERT_ID()
BEGIN SNAPSHOT
UPDATE products
SET inventory.on_hand = inventory.on_hand - $qty,
inventory.reserved = inventory.reserved + $qty
WHERE id = $pid AND inventory.on_hand >= $qty
RETURNING id AS pid;
INSERT INTO orders (user_id, total_cents, status, line_items)
VALUES ($uid, $total, 'pending',
[{product_id: $pid, qty: $qty, unit_cents: $unit}])
RETURNING id AS oid;
MATCH (u:user {id: $uid}), (p:product {id: $pid})
CREATE (u)-[:PURCHASED {order_id: $oid, qty: $qty, at: now()}]->(p);
COMMIT;
-- Subscription: same predicate language, LIVE prefix
LIVE SELECT id, status, total_cents FROM orders WHERE user_id = $uid;
```
The checkout query is the whole argument. Those three statements **must** commit together or not at all. `RETURNING id AS oid` threads the inserted order's id into the Cypher `CREATE` without inventing a special function call. No off-the-shelf system executes that atomically across SQL + JSONB + a graph store without stitching multiple transaction managers together (or adopting SurrealDB, which is the existence proof that this can be built).
## Option C — Full `.wo` Language for an E-commerce Platform with ACID
The [evaluation phase](./01-evaluation.md) argued against external databases for a small, single-writer content project. An **e-commerce platform inverts every one of those assumptions**:
| writeonce (blog) | E-commerce platform |
| --- | --- |
| Hundreds of articles | Millions of products, orders, sessions |
| Single writer (author) | Thousands of concurrent writers (customers, fulfillment, admin) |
| Read-heavy | Write-heavy on the hot paths (cart, checkout, inventory) |
| Full rebuild on change is free | Full rebuild is impossible — mutations must commit in milliseconds |
| No transactions needed | ACID is the product |
| Data is one shape (article) | Data is genuinely three shapes: relational (orders, inventory), document (product descriptions, reviews), graph (recommendations, categories, affiliations) |
The three-paradigm case that was marginal for a blog becomes **the correct design** for e-commerce. And ACID is not a nice-to-have — an order that decrements inventory but loses the payment record is a lawsuit.
This is Option C: a full `.wo` language backed by a real storage engine, a real transaction manager, and a real planner. It is a database, and writeonce/the parent project becomes the vehicle for building it.
### ACID — What Each Letter Requires
**Atomicity.** The checkout above touches three storage areas. The engine needs a single transaction coordinator that owns writes to all three. On abort, every partial write rolls back. Implementation: one **write-ahead log (WAL)** records intents across all paradigms; commit flips a single on-disk marker; crash recovery replays or discards based on the marker.
**Consistency.** Domain invariants that span paradigms must hold:
- `inventory.on_hand >= 0` (document field, relational-style constraint — expressed as `@check(>= 0)` in the schema layer)
- Every `PURCHASED` edge must reference an existing `orders.id` (graph-to-relational FK — enforced because `Purchase.order: ref Order` in the schema layer)
- `orders.total_cents == sum(line_items[].qty * line_items[].unit_cents)` (denormalization check — expressed as a computed field)
The schema layer captures these as declarations; the compiler emits the planner-level checks. The query layer also supports explicit `CONSTRAINT` statements for invariants that don't fit the type system.
**Isolation.** Thousands of concurrent carts racing for the last unit of inventory. Two realistic models:
| Model | How it works | Trade-off |
| --- | --- | --- |
| **MVCC** (Postgres, SurrealDB, CockroachDB) | Each transaction sees a snapshot; conflicts detected at commit | Readers never block writers; abort rate rises under contention |
| **2PL with row/edge locks** (MySQL InnoDB) | Locks acquired on read/write, released at commit | Lower abort rate; deadlocks must be detected |
MVCC is the modern default and what you'd target. Minimum isolation level for e-commerce: **Snapshot Isolation** (Postgres `REPEATABLE READ`). Anything weaker (`READ COMMITTED`) permits write skew — two customers each passing the "inventory >= 1" check and both succeeding on the last unit.
**Durability.** On `COMMIT`, the WAL record must be `fsync`'d before the client gets acknowledgment. Lose fsync and you lose paid orders on power failure. The WAL is the primary storage commitment; data files are derived and can be rebuilt by replaying the log from the last checkpoint.
### Storage Engine
`.seg` + `.idx` rebuilt-on-change does not survive contact with e-commerce. The engine needs genuine mutable on-disk structures:
| Component | Responsibility | Reference |
| --- | --- | --- |
| **WAL** | Ordered, fsynced log of every committed mutation | Postgres `pg_wal/`, RocksDB `*.log` |
| **Relational pages** | Fixed-size pages (4–16 KB) with slot-directory row layout, B+ tree indexes | Postgres heap + btree, SQLite |
| **Document store** | LSM tree (SSTables + memtable + compaction) for append-friendly JSON blobs | RocksDB, SurrealKV |
| **Graph store** | Adjacency list on disk — doubly-linked edge records per node for O(1) traversal | Neo4j's native store |
| **Buffer pool** | Shared page cache with LRU/CLOCK eviction, dirty page tracking | Postgres `shared_buffers` |
| **Checkpointer** | Periodically flushes dirty pages, truncates WAL | Every major DB has one |
| **Vacuum / compaction** | Reclaim space from MVCC dead tuples / LSM tombstones | Postgres autovacuum, RocksDB compaction |
Three storage backends, one WAL, one transaction coordinator. That is the core of the project. The [In-Memory Engine](./03-inmemory-engine.md) phase details the RAM-primary variant of this.
### Cross-Paradigm Transaction Coordinator
The novel piece — nobody ships this exactly the way `.wo` would need it:
```
┌──────────────────────────────────────┐
│ .wo Query Planner │
└──────────────────────────────────────┘
│
▼
┌──────────────────────────────────────┐
│ Transaction Coordinator (MVCC) │
│ - txn_id allocation │
│ - snapshot timestamp │
│ - commit ordering │
│ - WAL append + fsync │
│ - RETURNING alias table per txn │
└──────────────────────────────────────┘
│ │ │
▼ ▼ ▼
┌────────┐ ┌──────────┐ ┌──────────┐
│ Rel │ │ Doc │ │ Graph │
│ Engine │ │ Engine │ │ Engine │
│ (B+) │ │ (LSM) │ │ (adj) │
└────────┘ └──────────┘ └──────────┘
```
Every engine exposes the same transaction hooks: `begin(snapshot_ts)`, `stage(mutation)`, `prepare()`, `commit(wal_lsn)`, `abort()`. The coordinator drives a **two-phase commit internally** (not distributed 2PC — it's one process, so prepare+commit is cheap and deterministic).
Snapshot read across paradigms: each engine stores per-record `(created_txn_id, deleted_txn_id)` visibility info. A read at snapshot timestamp `T` sees only records visible at `T` — same rule in all three engines.
`RETURNING` aliases live in the transaction's scoped name table; each subsequent statement inside the same `BEGIN … COMMIT` resolves `$oid` etc. against it. This is how SQL results thread into Cypher `CREATE` atomically, without round-tripping to the client.
### Concurrency Model
**One process, one thread, one event loop.** Redis-style. The entire engine — connection accept, parser, planner, executor, buffer pool, subscription registry — runs on a single userland thread; the only non-userland thread is the kernel-owned io_uring SQPOLL helper, which is invisible to engine code.
| Subsystem | Where it runs |
| --- | --- |
| Connection accept | Event loop — non-blocking `accept()` via io_uring |
| Query execution | Event loop — parse, plan, execute inline |
| WAL fsync | Event loop submits SQEs to io_uring; kernel-owned SQPOLL thread drains; loop parks on CQE |
| Subscription dispatch | Event loop — predicates matched on commit, deltas pushed to per-subscription ring buffers |
| Compaction / checkpoint / vacuum | Event loop — scheduled as low-priority tasks between client work |
**Why single-threaded.** Three precedents:
- **Redis** ran single-threaded for its first decade (and still runs command execution single-threaded; 6.0 added threaded I/O only). It hits hundreds of thousands of ops/second on one core.
- **TigerBeetle** is single-threaded by design because determinism beats concurrency for financial workloads.
- **Node.js** is the proof that the event-loop model scales for I/O-bound work at web-scale.
Single-threaded execution removes entire failure modes: no lock ordering, no cross-thread memory ordering hazards, no MVCC visibility logic for reads racing with writes, no torn-page atomics. The transaction coordinator trivially serializes commits because there is only one of them at a time. Snapshot isolation reduces to a logical versioning scheme for live-query delta computation — not a multi-threaded correctness mechanism.
**Group commit still applies.** The loop drains many pending commits into one fsync SQE per tick — same amortization as Postgres's group-commit path, without a dedicated WAL-writer thread. Expected throughput on a modern NVMe server: ~200–500k simple-transaction commits per second per core. More than enough for every writeonce-sized workload.
**Scaling past one core.** The path is **sharding**: partition types across independent single-threaded engine processes (Redis Cluster is the reference). Each shard owns a disjoint set of types; cross-shard transactions use two-phase commit between shards. Revisit only when a production workload actually saturates the core — the multi-threaded single-engine alternative is years of work for the wrong kind of gain.
### Query Language Scope (`.wo` Full Spec)
The minimum grammar for an e-commerce workload is split by layer:
**Schema layer (enters at Phase 5):**
- `type Name { ... }` declarations with scalar fields, embedded structs, `ref`, `multi @edge`, `multi … via LinkType`, `backlink`
- `type Name link A -> B { ... }` for edge-with-properties
- Tagged unions `A | B | C`
- Annotations `@unique @check @default @index @search @immutable`
- Computed fields `total: Money = sum(…)`
- Per-type `policy` / `on <event>` / `service` blocks (full semantics in [Phase 6](./06-lowcode-fullstack.md))
**Query layer (Phase 2 prototype):**
- **DDL (interim)**: the existing `##sql / ##doc / ##graph` block syntax, as the compiler target for the schema layer and the prototype's direct authoring surface
- **DML relational**: `INSERT`, `UPDATE`, `DELETE`, `UPSERT` — all supporting `RETURNING col AS alias`
- **DML document**: path updates (`SET meta.reviews[3].stars = 5`), array operations (append, remove, splice)
- **DML graph**: `CREATE`, `MERGE`, `DELETE` on nodes and edges; variable-length path (`*1..5`)
- **Queries**: `SELECT` with joins, `MATCH` with traversal, subqueries, aggregations (`COUNT`, `SUM`, `AVG`)
- **Transactions**: `BEGIN [SNAPSHOT|SERIALIZABLE]` / `COMMIT` / `ROLLBACK`, `SAVEPOINT name` / `ROLLBACK TO name`
- **Parameters**: `$name` everywhere, typed at prepare
- **Cross-paradigm expressions**: dotted paths descend relational → document (`order.line_items[0].qty`); graph bindings resolve to relational rows (`(u:user {id: $uid})`); `RETURNING` aliases resolve across SQL → Cypher boundary
- **Subscriptions / live queries**: `LIVE SELECT` / `LIVE MATCH` — SurrealDB-style; predicate known at registration time for O(1) lookup matching on commit
- **Prepared statements + parameter binding**: `$name` with typed signatures; required for SQL injection resistance
- **Role-based auth + row-level policies**: schema-layer `policy` declarations compile to planner rewrite rules AND'd into every query at registration time
### Components to Build
```
.wo engine
├── parser Pratt or LALR — handles both layers and all 3 paradigms
├── analyzer name resolution, type checking, cross-paradigm dispatch
├── schema compiler type-DSL → physical schema (Phase 5)
├── planner cost-based: reorder joins, push predicates, choose index,
│ merge policy predicates at registration time
├── executor vectorized over relational, iterator over graph
├── txn manager MVCC, snapshot isolation, group commit, RETURNING aliases
├── wal ordered log, fsync, checkpoint, recovery
├── rel engine B+ tree heap, btree indexes, vacuum
├── doc engine LSM with bloom filters, compaction
├── graph engine native adjacency, property store
├── buffer pool shared page cache, dirty tracking
├── catalog schema metadata (from type DSL), evolvable at runtime
├── wire protocol client connections (pick: Postgres wire, or custom)
├── auth + rbac users, roles, row-level policies from type declarations
├── live queries incremental view maintenance for subscriptions
├── backup / replication physical log shipping; logical streaming for read replicas
└── observability query stats, slow log, lock waits, WAL lag
```
Implementation language is genuinely open. Reasonable picks and what each implies:
| Language | Why | Precedent |
| --- | --- | --- |
| **Rust** | Memory safety without GC, good async story, zero-cost abstractions, fits writeonce's existing stack | SurrealDB, TiKV, Materialize, sled |
| **C++** | Lowest overhead, decades of mature DB internals literature | Postgres (C), MySQL, RocksDB, DuckDB, ClickHouse |
| **Zig** | C-like control with safer semantics, compile-time metaprogramming useful for query planner | TigerBeetle |
| **Go** | Fastest to productive, excellent concurrency primitives, some GC cost on hot paths | CockroachDB, InfluxDB, Dgraph |
| **OCaml / Haskell** | Query planner is a compiler; ML-family languages are excellent at compilers | Irmin, some research DBs |
For e-commerce ACID specifically, **Rust or C++** — GC pauses during a checkout fsync batch are the kind of latency spike that loses money. Go works and has the fastest developer velocity, but CockroachDB has spent years tuning around GC; budget for that.
### Reference Implementations to Study
- **SAP CDS** — <https://cap.cloud.sap/docs/cds/>. The canonical declaration-first application language; direct inspiration for the schema layer.
- **EdgeDB / EdgeQL** — path-based query language over a typed schema; studies the corner cases of computed fields, link properties, and tagged unions. Worth reading their planner learnings before re-implementing.
- **Postgres** — the canonical ACID RDBMS. Read `src/backend/access/transam/` for WAL, `src/backend/storage/buffer/` for the buffer pool, `src/backend/storage/lmgr/` for locking. Decades of battle-tested code.
- **SurrealDB** — closest living example of the query layer. Multi-model, multi-paradigm query language, Rust, embeddable. Already referenced in [surreal-case-study.md](../surreal-case-study.md).
- **SQLite** — smallest complete ACID database in the open-source world. `src/btree.c`, `src/pager.c`, `src/wal.c` are worth reading front-to-back.
- **CockroachDB** — distributed SQL with serializable isolation. Go, but the transaction protocol (Parallel Commits) is well-documented.
- **TigerBeetle** — financial-grade ACID, deterministic, written in Zig. Essay on why they rewrote: <https://tigerbeetle.com/blog/>.
- **DuckDB** — analytics-focused but single-file embeddable C++ engine, excellent reference for a modern vectorized executor.
- **Neo4j** — for the graph-storage side. Native adjacency, transaction log, lock manager.
- **Papers**:
- *Architecture of a Database System* (Hellerstein, Stonebraker, Hamilton, 2007) — the canonical survey
- *The Log-Structured Merge-Tree* (O'Neil 1996) — for the LSM document backend
- *Serializable Snapshot Isolation in PostgreSQL* (Ports & Grittner 2012) — making SI safe
- *A Critique of ANSI SQL Isolation Levels* (Berenson et al. 1995) — so you pick the right default
### Realistic Scope
This is a multi-person-year project. Concrete gates:
| Milestone | What's usable | Rough effort |
| --- | --- | --- |
| Parser + analyzer + in-memory executor | Single-user prototype, no durability (**shipped** at `prototypes/wo-db/`) | 2–4 months |
| Fixed-glue query layer (`$name`, `RETURNING`, `BEGIN/SAVEPOINT/COMMIT`, `LIVE` keyword) | Multi-statement cross-paradigm transactions threadable | +1–2 months |
| WAL + crash recovery + single-table B+ tree | ACID on relational only, one writer | +3–6 months |
| MVCC + concurrent transactions | Multi-writer relational | +3–6 months |
| Document engine (LSM) | Relational + document, transactional | +4–8 months |
| Graph engine + cross-paradigm txns | All three paradigms ACID | +6–12 months |
| Schema-layer compiler (type DSL → physical schema) | Single source of truth for codegen and full-stack blocks | +2–4 months |
| Wire protocol + RBAC + observability | Deployable to production | +3–6 months |
| Replication + backup | Survive a node loss | +6–12 months |
Two to four years for a small team to reach something a real e-commerce business would trust with payment data. Adopting Postgres (with JSONB for the document side and the Apache AGE extension or a separate graph store for the graph side) gets you there in a week.
### Honest Decision Framing
Build `.wo` for an e-commerce platform **only if** at least one of the following is true:
1. The cross-paradigm query atomicity is a business-critical feature the founders want to sell ("our DB does what Postgres + Neo4j glued together cannot"). This is the SurrealDB and EdgeDB thesis.
2. Building the database **is** the product — the e-commerce platform is the test harness, not the goal.
3. You have a team comfortable with the papers listed above and the patience to ship a toy for 18 months before it's useful.
Otherwise, the pragmatic stack for a multi-paradigm ACID e-commerce platform is:
- **Postgres** for relational + document (JSONB) + row-level security. Handles 99% of the workload. Apache AGE extension adds Cypher-compatible graph queries in the same transaction.
- **Redis** for cart/session/rate-limit (expiring, non-durable).
- **Search** (OpenSearch/Meilisearch/Typesense) for product search — specialized workload.
- **Event log** (Kafka/Redpanda) for order events, downstream analytics, fulfillment.
That stack is boring and it works. `.wo` as described is interesting and would take years. Pick based on whether the goal is to ship e-commerce or to ship a database.

View file

@ -0,0 +1,205 @@
# Phase 3 — In-Memory Engine
> RAM-primary, SSD-durable storage using io_uring, designed for OLTP e-commerce workloads on a 64 GB Linux machine.
**Previous**: [Phase 2 — The `.wo` Language & ACID Engine](./02-wo-language.md) | **Next**: [Phase 4 — Client API](./04-client-api.md) | **Index**: [database.md](../database.md)
---
Seed constraints:
- 64 GB RAM — the entire live dataset fits in memory; no page eviction on the hot path
- Dual write — every mutation goes to an in-RAM structure **and** to an on-SSD durable log simultaneously
- Linux-only — Linux kernel APIs are fair game, no portability obligation
- `io_uring` for asynchronous read/write to SSD
This is a **RAM-primary, SSD-durable** engine — the modern OLTP architecture used by TigerBeetle, ScyllaDB (via Seastar), VoltDB/H-Store, SAP HANA, and Redis-with-AOF. Reads never touch disk. Writes touch RAM immediately and SSD asynchronously, with fsync gating commit acknowledgment.
For the e-commerce workload in [Phase 2](./02-wo-language.md), this is the right physical design: checkout latency is dominated by the durability path, not lookup; and a 64 GB live set comfortably holds millions of products + recent orders + active sessions + the entire recommendation graph.
## Memory Layout
All three paradigm engines live in one address space:
```
┌──────────────────────────────────────────────────────────────┐
│ Process address space │
├──────────────────────────────────────────────────────────────┤
│ Relational heap B+ tree pages (row-store) ~20 GB │
│ Document store LSM memtable + sorted runs ~15 GB │
│ Graph store Node + edge arenas, adjacency ~10 GB │
│ Index shards Hash, sorted, inverted ~8 GB │
│ Buffer for WAL Ring buffer staged for SSD ~2 GB │
│ MVCC version chains Per-record visibility history ~5 GB │
│ Connection / query Per-session scratch ~2 GB │
│ Headroom / OS Free for kernel, page tables ~2 GB │
└──────────────────────────────────────────────────────────────┘
```
Key moves:
- **`mlockall(MCL_CURRENT | MCL_FUTURE)`** — pin all pages, guarantee no swap-out. A single swap-in during checkout is a latency catastrophe.
- **`MAP_HUGETLB` / Transparent Huge Pages** — 2 MB pages reduce TLB pressure on hot indexes. For 64 GB of data, 4 KB pages mean 16 M TLB entries; 2 MB pages mean 32 K.
- **`/proc/sys/vm/swappiness = 0`** — belt and braces with `mlockall`.
- **NUMA awareness** — optional on multi-socket hosts. Since the engine is single-threaded ([Phase 2 concurrency model](./02-wo-language.md#concurrency-model)), pin the one event-loop thread and bind the arena to that socket (`numactl --membind=0 --cpunodebind=0`). Cross-socket memory access is 2–3× slower; a single-socket deployment sidesteps it entirely.
- **Slab / arena allocators** — avoid `malloc` in the hot path. Pre-size arenas per paradigm at startup.
## Dual-Write Durability Path
A write is committed only when its WAL record is on the SSD with `fsync` confirmed. The in-memory structure is updated first (fast), then the log write is awaited (slow enough to matter):
```
Transaction COMMIT
│
▼
┌──────────────────────┐
│ 1. Stage mutations │ apply to RAM structures under MVCC
│ to in-memory │ (version chains; readers unaffected)
│ engines │
└──────────────────────┘
│
▼
┌──────────────────────┐
│ 2. Serialize WAL │ append-only ring buffer in RAM
│ record │ header + paradigm deltas + LSN
└──────────────────────┘
│
▼
┌──────────────────────┐
│ 3. io_uring submit │ IORING_OP_WRITE with O_DIRECT
│ WAL write to SSD │ batched with concurrent commits
└──────────────────────┘
│
▼
┌──────────────────────┐
│ 4. io_uring submit │ IORING_OP_FSYNC
│ fsync (barrier) │ linked SQE after the write
└──────────────────────┘
│
▼
┌──────────────────────┐
│ 5. CQE received │ commit marker flipped
│ → ack client │ MVCC snapshot published
└──────────────────────┘
```
Steps 1–2 are synchronous; 3–5 are asynchronous. The event loop submits the SQEs and moves on to the next client; it reaps the CQE on a later tick — so the loop can have thousands of commits in flight without parking on any fsync syscall.
**Group commit**: the loop drains the commit queue into one fsync SQE per tick. If 500 transactions all committed within a 100 μs window, one fsync durable-s the batch. Amortizes SSD latency (~50–100 μs on NVMe) across the batch — throughput approaches `batch_size / fsync_latency`, which on a good NVMe is 500K+ commits/sec.
**What "dual write" means here**: it is *not* a two-database write where both must succeed independently. It is one logical commit that updates RAM (the query surface) and appends to the SSD WAL (the recovery record). On crash, RAM is gone; recovery replays the WAL to rebuild RAM state. The SSD is the source of truth for *durability*; RAM is the source of truth for *reads*.
## io_uring Mechanics
`io_uring` (Linux 5.1+, mature by 5.11) is the replacement for `epoll` + `libaio` for storage I/O. Two lock-free ring buffers shared between user-space and kernel:
| Ring | Direction | Contents |
| --- | --- | --- |
| **SQ** (Submission Queue) | Userland → Kernel | SQEs: `IORING_OP_WRITE`, `IORING_OP_FSYNC`, `IORING_OP_READ`, etc. |
| **CQ** (Completion Queue) | Kernel → Userland | CQEs: result code + user_data pointer back to the request |
Configuration knobs that matter for a database:
- **`IORING_SETUP_SQPOLL`** — a kernel thread polls the SQ. Userland writes SQEs without any syscall. Read/write submission becomes a memory write + memory barrier. Cost: one pinned kernel thread per ring.
- **`IORING_SETUP_IOPOLL`** — busy-poll for completions on the device instead of interrupt-driven. Lower latency on NVMe, higher CPU. Requires `O_DIRECT`.
- **`IORING_REGISTER_BUFFERS`** — pre-register WAL ring-buffer pages with the kernel. Skips per-I/O page pinning.
- **`IORING_REGISTER_FILES`** — pre-register the WAL fd. Skips fd table lookups per I/O.
- **Linked SQEs (`IOSQE_IO_LINK`)** — enforce ordering: write-then-fsync, or WAL-then-commit-marker. Kernel guarantees link order without userland waiting on the intermediate CQE.
- **`O_DIRECT`** on the WAL file — bypass the kernel page cache. The database manages its own buffering; double-caching wastes the 64 GB.
Per-commit path with full optimization: no syscalls at all for submission (SQPOLL), one memory read for completion (IOPOLL), zero page-pinning cost (registered buffers), zero fd-table lookup (registered files). The commit loop is effectively as fast as the NVMe firmware allows.
## Recovery
RAM is volatile; on restart the engine is empty. Recovery rebuilds it:
1. **Open WAL**. Scan forward from the last checkpoint LSN.
2. **Replay committed records.** Apply each to the in-memory engines in LSN order. Skip incomplete transactions (no commit marker).
3. **Load checkpoint snapshot** (optional but standard). Periodically, the engine dumps a consistent snapshot of the RAM state to SSD. On recovery, load snapshot → replay WAL from snapshot LSN forward. Avoids replaying hours of log.
4. **Rebuild indexes.** Indexes are derived from heap data — rebuilt during replay or lazily on first access.
5. **Open for traffic.**
Recovery target: 60 GB of data + a few million WAL records = seconds to a minute on NVMe, not hours. A good checkpointer runs in the background every 5–15 minutes; recovery only replays the delta since the last checkpoint.
## Concurrency in RAM
No page eviction, no buffer pool locks — and because the engine is [single-threaded](./02-wo-language.md#concurrency-model), no cross-thread races either. The concurrency story collapses to "there is no concurrency within the engine; there is a queue of clients being served sequentially by one loop". Every data structure is owned by that one loop:
| Structure | Primitive |
| --- | --- |
| B+ tree (relational) | Plain owned tree; no latches, no optimistic locks |
| LSM memtable (document) | Plain skiplist; sealed memtables still immutable for background compaction SQEs |
| Graph adjacency | Plain hashmap per node label |
| MVCC version chain | Plain singly-linked version list; no CAS |
| WAL ring buffer | Single-producer, single-consumer ring |
| Txn coordinator | Plain `u64` counter — incremented without atomics |
**Readers still see snapshots.** MVCC remains useful but its purpose changes: instead of "readers don't block writers on another thread", it's "a live-query subscriber reading in the same tick sees the pre-commit view; the post-commit delta arrives on the next tick". That semantic is cheap to implement when there is only one mutator.
**The model is sequential.** Clients are served round-robin by the loop; nothing races because nothing runs concurrently inside the engine. When one core isn't enough, [shard](./02-wo-language.md#concurrency-model) rather than bolting multi-threading onto this design.
## Capacity Planning
64 GB is a budget, not a guarantee. Three failure modes to design around:
1. **Working set exceeds RAM.** Solution path: add a tier (warm SSD-backed region for cold rows), or shard across nodes. Neither is in the Phase 2 scope — flag when live data approaches 50 GB.
2. **MVCC version chains bloat.** Long-running transactions hold old versions alive. Solution: aggressive vacuum, transaction timeouts, snapshot horizon tracking. Postgres hits this same wall.
3. **Sudden write bursts flood the WAL.** Solution: admission control — if SSD write queue depth exceeds a threshold, slow down `COMMIT` acknowledgment. Better than OOM'ing the WAL buffer.
## Comparison With Alternatives
| Aspect | In-memory + WAL (this design) | Disk-primary (Postgres) | Pure in-memory (Redis w/o AOF) |
| --- | --- | --- | --- |
| Read latency | ~100 ns (RAM) | ~10 μs (buffer cache hit) to ms (miss) | ~100 ns (RAM) |
| Write latency | ~50–100 μs (fsync) | ~50–100 μs (fsync) | ~100 ns (none) |
| Durability | Full — WAL fsync before ack | Full — WAL fsync before ack | Window of loss (AOF every-sec) |
| Dataset size | Bounded by RAM | Bounded by disk | Bounded by RAM |
| Restart time | Seconds to minutes (WAL replay) | Seconds | Immediate (empty) or minutes (AOF) |
| Ideal workload | OLTP with small-to-medium dataset | General-purpose, large datasets | Cache, session, ephemeral |
This design keeps the durability of Postgres and the read speed of Redis.
## Linux Tuning Checklist
Before production benchmarking:
- `echo 0 > /proc/sys/vm/swappiness`
- `echo never > /sys/kernel/mm/transparent_hugepage/enabled` (databases typically prefer explicit hugepages over THP's defragmentation stalls)
- `vm.nr_hugepages = <enough for the arenas>`
- `ulimit -l unlimited` (for `mlockall`)
- `blk-mq` scheduler: `none` or `mq-deadline` on NVMe (not `cfq`/`bfq`)
- `IORING_SETUP_SINGLE_ISSUER` — always, since the engine is single-threaded (Linux 6.0+)
- NUMA: `numactl --membind=0 --cpunodebind=0` to pin the loop + its arena to one socket. Multi-socket deployments should shard across sockets rather than sharing one engine
- Disable CPU frequency scaling (`cpupower frequency-set -g performance`) — saves microseconds that add up across group-commit batches
- Disable Meltdown/Spectre mitigations only if you control the hardware and understand the trade-off — they cost 10–30% on syscall-heavy paths, but io_uring with SQPOLL largely sidesteps them anyway
## Reference Implementations
- **TigerBeetle** — Zig, in-memory, io_uring end-to-end, deterministic, designed for financial OLTP. The closest living example of this exact architecture. <https://github.com/tigerbeetle/tigerbeetle>
- **ScyllaDB / Seastar** — C++, io_uring (and SPDK), shared-nothing per core, NUMA-aware. Seastar is the framework underneath. <https://github.com/scylladb/seastar>
- **VoltDB (H-Store)** — Java, in-memory OLTP, command-logging for durability. The academic ancestor of this design pattern.
- **Redis (`appendonly yes` + `appendfsync always`)** — simpler but exact same shape: RAM-primary, log-durable.
- **LMDB** — memory-mapped B+ tree; reads are literal pointer chases into mmap'd pages. Not WAL-based but worth studying for RAM-resident read paths.
- **SingleStore (formerly MemSQL)** — commercial in-memory row-store with columnar on-disk secondary. Hybrid of this design and disk-primary.
- **Readings**:
- *The End of an Architectural Era* (Stonebraker et al., 2007) — the H-Store paper that argued disk-primary databases were legacy for OLTP.
- *Efficient Lock-Free Durable Sets* (Zuriel et al.) — for lock-free structures that persist.
- *io_uring by Example* (Jens Axboe) and the `liburing` documentation — the authoritative guide.
## Where This Fits in Phase 2
This replaces the storage-engine block in the [Phase 2](./02-wo-language.md) component list. Specifically:
| Phase 2 component | Becomes (with in-memory design) |
| --- | --- |
| Relational pages + buffer pool | RAM-resident B+ tree / Masstree, no page eviction |
| Document engine (LSM on disk) | LSM memtable in RAM; sealed memtables spilled to SSD only for checkpointing |
| Graph engine (disk adjacency) | RAM adjacency arena; checkpointed, not paged |
| WAL | `io_uring` + `O_DIRECT` append-only log on NVMe |
| Checkpointer | Periodic snapshot of RAM arenas to SSD for fast recovery |
| Buffer pool | **Removed** — all data is in RAM |
| Vacuum | Still needed, but for MVCC chain pruning, not for reclaiming disk pages |
The cross-paradigm transaction coordinator sketched in Phase 2 stays the same — it just drives in-memory engines instead of disk-paged ones, and the WAL append it depends on is the one `io_uring` path.
Net effect: **shorter read paths, identical durability story, same ACID guarantees**, at the cost of a hard dataset ceiling set by RAM.

View file

@ -0,0 +1,279 @@
# Phase 4 — Client API: Wire Protocol and Subscriptions
> How remote clients connect, query, and subscribe to live changes — no polling anywhere in the chain.
**Previous**: [Phase 3 — In-Memory Engine](./03-inmemory-engine.md) | **Next**: [Phase 5 — Go Client SDK](./05-go-sdk.md) | **Index**: [database.md](../database.md)
---
Once the engine is complete it is a **traditional database server**: remote clients connect over the network, issue queries, receive results, and (critically for e-commerce UX) **subscribe to changes without polling**.
The client API has three problems to solve:
1. **Wire protocol** — how bytes move between client and server.
2. **Query surface** — what queries look like from the client's perspective (raw `.wo`, SQL, GraphQL, REST).
3. **Subscriptions** — how the server pushes change notifications to clients when matching data mutates, with no client-side polling.
## The Polling Problem This Must Avoid
Every naive realtime system reaches for polling first. For an e-commerce platform it is disqualifying:
| Polling | Subscriptions |
| --- | --- |
| Client asks "anything new?" every N ms | Server tells client "here's what changed" when it changes |
| Wasted RTTs when nothing changed | Zero traffic when nothing changes |
| Stale data up to N ms old | Sub-ms latency after commit |
| `O(clients × poll_rate)` server load | `O(mutations × matched_subscribers)` — scales with real change, not client count |
| Inventory display lies for up to N ms | Inventory display reflects the commit |
| "Order shipped" email triggered by cron | Fired by a committed status-change |
Every subscription-based design in this section follows the same rule already set by writeonce in [05-datalayer.md](../../05-datalayer.md) and [03-data.md](../../03-data.md): **the client registers a query once, the server pushes deltas on commit, the client never asks again.**
## Protocol Layer — Pick One or Both
Two protocol tiers make sense: a **native binary protocol** for app servers and ORMs that want every microsecond, and a **GraphQL-over-WebSocket layer** for browsers, mobile apps, and third parties. They are not alternatives — they share the same planner and subscription registry underneath.
| Option | Best for | Trade-off |
| --- | --- | --- |
| **Custom binary over TCP** | App servers, in-house clients, highest throughput | Need to ship client libs in every language |
| **Postgres wire protocol** (libpq) | Reuse the Postgres client ecosystem (psql, pgx, node-postgres, JDBC) | Locked into Postgres's shape — no native graph/live-query verbs |
| **gRPC (HTTP/2)** | Cross-language, well-tooled, server-streaming RPC covers subscriptions | Protobuf schema overhead; HTTP/2 stack cost |
| **GraphQL over HTTP + WebSocket** | Web/mobile clients, schema-aware tooling, built-in `subscription` operation | Parser/resolver overhead, N+1 risks |
| **REST + SSE** | Simplest to integrate (curl, browser `fetch`) | Verb-per-endpoint sprawl, SSE is unidirectional |
**Recommended combination:**
- **Native binary protocol** for first-party app servers (cart service, checkout, fulfillment).
- **GraphQL over WebSocket** for everything else (web, mobile, partner APIs).
Both terminate at the same **session layer** inside the server, which delegates to the `.wo` planner.
## Native Binary Protocol — Shape
A minimal framing that's compatible with io_uring on both ends:
```
┌────────┬────────┬──────────┬─────────────────────────────┐
│ opcode │ req_id │ len │ payload │
│ u8 │ u64 │ u32 │ bincode / msgpack │
└────────┴────────┴──────────┴─────────────────────────────┘
```
Opcode set:
| Opcode | Direction | Purpose |
| --- | --- | --- |
| `HELLO` | C → S | Protocol version + auth credentials |
| `WELCOME` | S → C | Session id + server capabilities |
| `PREPARE` | C → S | Compile a `.wo` query, cache plan on server |
| `EXECUTE` | C → S | Run prepared plan with bound parameters |
| `RESULT` | S → C | Full result set for one query |
| `BEGIN` / `COMMIT` / `ROLLBACK` | C → S | Explicit transaction control |
| `SUBSCRIBE` | C → S | Register a live query, get a subscription id |
| `UNSUBSCRIBE` | C → S | Cancel a subscription |
| `DELTA` | S → C | Pushed change matching a subscription |
| `COMPLETE` | S → C | Subscription terminated server-side (schema change, etc.) |
| `ERROR` | S → C | Typed error with query context |
| `PING` / `PONG` | bidirectional | Dead connection detection (no polling for data — just keepalive) |
Multiplexed: many in-flight `req_id`s per connection, responses interleaved. Matches io_uring's async nature naturally — a connection never blocks on a slow query.
## GraphQL — Schema, Queries, Mutations, Subscriptions
GraphQL has the three verbs e-commerce actually uses:
| GraphQL operation | `.wo` mapping |
| --- | --- |
| `query` | `SELECT` / `MATCH` over the engine, single response |
| `mutation` | `INSERT` / `UPDATE` / `DELETE` / `CREATE` inside an implicit transaction |
| `subscription` | `LIVE SELECT` / `LIVE MATCH` — server pushes on match |
**Schema generation.** The `.wo` DDL is the source of truth; the GraphQL SDL is generated from it:
```
##sql #products (id, sku, price_cents, meta, inventory)
##doc #product-meta (title, description, reviews, ...)
##graph (user)-[:PURCHASED]->(product)
│
▼ generator
│
type Product {
id: ID!
sku: String!
priceCents: Int!
meta: ProductMeta!
inventory: InventoryLevel!
similarTo(limit: Int = 10): [Product!]! # graph traversal
purchasedBy: [User!]! # graph traversal
}
type Subscription {
productUpdated(id: ID!): Product!
inventoryChanged(sku: String!): InventoryLevel!
orderStatus(orderId: ID!): Order!
}
```
**Subscription example (e-commerce checkout feedback loop):**
```graphql
subscription CartInventory($skus: [String!]!) {
inventoryChanged(sku_in: $skus) {
sku
onHand
reserved
}
}
```
A web client opens this WebSocket subscription when the cart renders. The server only pushes when a committed transaction changes any of those SKUs' inventory — the cart's "2 left!" badge is always live, no polling.
**Transport: `graphql-ws` protocol over WebSocket.** Standard, well-tooled (Apollo, urql, Relay, Hasura all speak it). Falls back to HTTP POST for plain queries and mutations.
## Subscription Engine — How Push Actually Works
This is the mechanism that makes polling unnecessary. It lives inside the transaction coordinator from [Phase 2](./02-wo-language.md):
```
┌─────────────────────────────────────────────────────┐
│ Transaction Coordinator (MVCC) │
│ │
│ on COMMIT(txn): │
│ delta = collect_changes(txn) │
│ matched = subscription_registry.match(delta) │
│ for (sub, rows) in matched: │
│ sub.writer.push(DELTA { sub.id, rows }) │
└─────────────────────────────────────────────────────┘
│ │
▼ ▼
┌─────────────────────┐ ┌──────────────────────────┐
│ Subscription │ │ Session Writer (per conn)│
│ Registry │ │ - native: io_uring send │
│ │ │ - graphql: ws frame │
│ predicate → [subs] │ │ - grpc: server stream │
└─────────────────────┘ └──────────────────────────┘
```
**Matching strategies**, in order of cost:
| Subscription shape | Matching cost | Example |
| --- | --- | --- |
| Keyed (primary key) | O(1) hash lookup on commit | `productUpdated(id: 42)` |
| Tag / secondary index | O(1) index lookup + scan of matched rows | `orderStatusByUser(userId: 17)` |
| Range | O(log n) index range + filter | `ordersPlaced(between: [start, end])` |
| Graph traversal | O(edges visited) — bound by depth/limit | `recommendationsFor(userId: 17)` |
| Arbitrary predicate | O(subs) — evaluate each against the delta | `LIVE SELECT ... WHERE complex` |
The engine indexes subscriptions by their shape so the common cases (keyed, tag-based) don't pay the arbitrary-predicate price. This is **incremental view maintenance** — the same idea that SurrealDB live queries, Materialize, Hasura, and Feldera all implement at different levels of generality.
## Connection I/O — io_uring All the Way
The same `io_uring` that drives the WAL (per [Phase 3](./03-inmemory-engine.md)) also drives client sockets. One scheduler, not a mix of epoll for networking and io_uring for storage:
| Operation | io_uring opcode |
| --- | --- |
| Accept new client | `IORING_OP_ACCEPT` |
| Read request frame | `IORING_OP_RECV` (with registered buffers) |
| Write result / delta | `IORING_OP_SEND` (with `IOSQE_IO_LINK` to chain writes) |
| TLS handshake | Userland ring integrated with `IORING_OP_RECV`/`SEND` (e.g., rustls or BoringSSL in non-blocking mode) |
| Close | `IORING_OP_CLOSE` |
| Keepalive | `IORING_OP_TIMEOUT` per connection |
A subscription push is one SQE: `SEND(client_fd, delta_frame)`. Thousands of in-flight pushes across thousands of subscribers is just thousands of SQEs — the kernel batches the actual NIC writes. No thread-per-connection, no blocking send.
## Session State
Each connected client has server-side state:
| State | Lifetime | Notes |
| --- | --- | --- |
| Identity / principal | Session | JWT or mTLS validated at `HELLO` |
| Current transaction | One txn at a time per session | Auto-rollback on disconnect |
| Prepared statements | Session | Plan cached, re-parameterized per `EXECUTE` |
| Active subscriptions | Session | All torn down on disconnect (free registry slots, stop pushing) |
| Role / RBAC context | Session | Feeds row-level policies into the planner |
| Back-pressure credits | Per-subscription | Client advertises how many outstanding `DELTA` frames it can buffer |
On disconnect (TCP close, keepalive failure, `EPOLLHUP`-equivalent from io_uring completion): all sessions state is freed, all subscriptions unregistered. Same philosophy as `wo-sub`'s `EPOLLHUP` → automatic `unsubscribe(fd)` from [05-datalayer.md](../../05-datalayer.md), scaled up to a real server.
## Back-Pressure
A slow client cannot be allowed to stall commits. The push path must never block on a socket write:
1. Each subscription has a **bounded outbound queue** (say, 1024 deltas).
2. Writer thread drains the queue via `io_uring_send`.
3. On queue overflow, the engine has three policies:
- **Drop + resync**: mark the subscription as "behind", push a single `RESYNC` marker, client re-requests current state.
- **Coalesce**: fold consecutive deltas for the same key into one (last-writer-wins).
- **Disconnect**: close the connection; clients with a stale subscription reconnect.
4. The coordinator never waits on a subscription — it hands the delta to the writer and moves on.
This is the same trade-off Kafka makes with consumer lag: fast producers, independent consumers, bounded buffer, spillover policy.
## Authentication and Authorization
Covered briefly in [Phase 2](./02-wo-language.md); the wire protocol is where it bites:
- **Transport**: TLS mandatory for any non-loopback connection. Offload to `rustls` / `boringssl` userland; io_uring handles only the underlying sockets.
- **Authentication** at `HELLO`: JWT (stateless), API key (server-validated), or mTLS (cert-based).
- **Authorization**: RBAC + row-level policies evaluated inside the planner. A subscription's registered predicate is **intersected with the user's access policy at registration time** — if the policy says user 17 only sees their own orders, the subscription's effective predicate becomes `(original) AND user_id = 17`. Enforced once, not per push.
- **Rate limiting**: per-session token bucket enforced before any query work. Cheap to implement in the io_uring accept/recv path.
## Comparison: This Design vs. Existing Products
| Aspect | This design | Postgres + Hasura | Supabase Realtime | SurrealDB | Firebase |
| --- | --- | --- | --- | --- | --- |
| Transport | Custom binary + GraphQL/WS | SQL wire + GraphQL/WS | Postgres WAL → WS | HTTP + WS | Custom WS |
| Subscriptions | Native, planner-integrated | Live queries via polling Postgres | Logical replication fan-out | Native live queries | Native |
| Storage coupling | In-process | External Postgres | External Postgres | In-process | Proprietary |
| Cross-paradigm | Yes (`.wo`: rel + doc + graph) | Partial (JSONB, no graph) | Relational only | Yes (rel + doc + graph) | Doc only |
| Polling internally? | No | **Yes** (Hasura polls Postgres) | No (uses WAL) | No | No |
| io_uring throughout | Yes | No | No | Partial | No |
Hasura is the instructive one — it gives clients push subscriptions, but internally it polls Postgres because Postgres has no commit-time subscription hook. Building the subscription engine *inside* the database (as this design does) is what eliminates polling end-to-end.
## Reference Implementations
- **SurrealDB** — the tightest match: custom engine, WebSocket transport, native `LIVE SELECT`. <https://github.com/surrealdb/surrealdb>. Also in [surreal-case-study.md](../surreal-case-study.md).
- **Hasura GraphQL Engine** — production-quality GraphQL over Postgres with subscriptions. Read their `graphql-engine/server/src-lib/Hasura/GraphQL/Transport/` for subscription multiplexing. <https://github.com/hasura/graphql-engine>
- **Supabase Realtime** — Phoenix/Elixir server that tails Postgres logical replication and fans out over WebSocket. Cleanest demo of "subscriptions as a layer over an existing DB." <https://github.com/supabase/realtime>
- **PostgREST** — auto-generated REST from Postgres schema. Simpler than GraphQL, same spirit. <https://github.com/PostgREST/postgrest>
- **EdgeDB** — custom binary protocol, custom query language (EdgeQL), compiles to Postgres underneath. Good reference for protocol framing. <https://github.com/edgedb/edgedb>
- **Materialize** — incremental view maintenance as a product; every query is implicitly a subscription. <https://github.com/MaterializeInc/materialize>
- **Phoenix Channels** (Elixir) — mature pub/sub-over-WebSocket with presence, back-pressure, and reconnection baked in. Worth reading even if the server is Rust/C++.
- **graphql-ws** protocol — <https://github.com/enisdenjo/graphql-ws>. The WebSocket sub-protocol every modern GraphQL client speaks.
- **Apollo Router** — GraphQL gateway with subscription multiplexing, federation. <https://github.com/apollographql/router>
## Scope Addition to Phase 2
The client API is a sizable addition to the [Phase 2](./02-wo-language.md) component list:
| Component | New work |
| --- | --- |
| Native wire codec | Binary framing, opcode dispatch, session lifecycle |
| Postgres-wire compatibility (optional) | libpq protocol v3 parser — reuse clients |
| GraphQL layer | SDL generation from `.wo`, resolver dispatch, `graphql-ws` subscriptions |
| REST/SSE gateway (optional) | Thin translation to native protocol |
| Subscription registry | Indexed by subscription shape; matched on commit |
| Push writer pool | io_uring-backed, per-connection outbound queues, back-pressure policy |
| TLS / auth | rustls or boringssl, JWT/mTLS at connection open |
| Connection manager | Accept, keepalive, graceful shutdown, fd limits |
| Observability | Per-session stats, slow query log, subscription lag, push-queue depth |
Rough incremental effort on top of Phase 2: **6–12 months** for a production-quality client layer with both native and GraphQL protocols, assuming the engine underneath is working.
## Why This Matters for E-commerce
Every hot user-facing screen is a subscription in disguise:
| Screen | Subscription |
| --- | --- |
| Product page | `productUpdated(id)` — price/stock changes reflect instantly |
| Cart | `inventoryChanged(sku_in: cartSkus)` — "out of stock!" appears the moment it's true |
| Order status | `orderStatus(orderId)` — pending → paid → shipped, no refresh |
| Admin dashboard | `LIVE SELECT COUNT(*) FROM orders WHERE placed_at > NOW() - 1h` |
| Recommendations sidebar | `LIVE MATCH (me)-[:VIEWED]->-[:SIMILAR_TO]->(p)` |
| Seller notifications | `LIVE MATCH (order)-[:CONTAINS]->(p) WHERE p.seller_id = $me` |
Each of these is `O(1)` server work per commit — the matching subscription is indexed by the thing that changed. Without subscriptions, every one of those screens would be a polling loop hammering the database. With subscriptions, server load scales with **actual state change**, not with client count × poll rate.
That is the whole argument for building the subscription engine into the database rather than bolting a message bus onto the side: **the engine already knows when something committed. Publishing the delta is a function call, not another system.**

View file

@ -0,0 +1,371 @@
# Phase 5 — Go Client SDK
> A typed Go client with subscription-first design — subscribe to a live query in 5 lines, deltas arrive on a channel.
**Previous**: [Phase 4 — Client API](./04-client-api.md) | **Next**: [Phase 6 — Low-Code Full-Stack](./06-lowcode-fullstack.md) | **Index**: [database.md](../database.md)
---
Concrete scenario: a developer writes a Go backend that serves a web frontend, and uses the `.wo` database as the store. They must be able to **subscribe to a query in ~5 lines of idiomatic Go** and have deltas arrive on a channel.
Everything else the SDK does — connect, query, mutate, transact — is table stakes covered by every existing Go DB driver. Subscriptions are what this SDK has to get right.
## Target API Surface
The engine speaks `.wo` on the wire. Every SDK method — typed or untyped — is a thin wrapper over one primitive: **send a `.wo` source string with `$name` parameters, get back a uniform `Result`**.
```go
import "go.writeonce.dev/wo"
// 1. Connect
client, err := wo.Connect(ctx, "wo://db.example.com:5555",
wo.WithAPIKey(os.Getenv("WO_KEY")),
wo.WithTLS(tlsConfig),
)
defer client.Close()
// 2. Wo — the primitive: run ANY .wo source (one statement or a whole
// BEGIN...COMMIT block mixing SQL, Cypher, and document updates). Params
// use the $name rule from the .wo language spec.
result, err := client.Wo(ctx, `
UPDATE products
SET inventory.on_hand -= $qty
WHERE id = $pid AND inventory.on_hand >= $qty
RETURNING id AS pid;
INSERT INTO orders (user_id, total_cents, status)
VALUES ($uid, $total, 'pending')
RETURNING id AS oid;
MATCH (u:user {id: $uid}), (p:product {id: $pid})
CREATE (u)-[:PURCHASED {order_id: $oid, qty: $qty, at: now()}]->(p);
`, wo.Params{
"uid": uid, "pid": 42, "qty": 2, "total": 9800,
})
// Result layout:
// result.Rows — rows from trailing SELECT/MATCH/RETURNING statements
// result.Aliases — the RETURNING alias table: {"pid": 42, "oid": 17}
// result.Affected — rows touched by INSERT/UPDATE/DELETE, per-statement
// 3. Query — sugar over Wo that scans trailing rows into a typed destination.
var products []Product
err = client.Query(ctx,
"SELECT id, sku, price_cents, meta FROM products WHERE price_cents < $max",
wo.Params{"max": 5000},
).Scan(&products)
// 4. Exec — sugar for writes that don't return rows.
_, err = client.Exec(ctx,
"UPDATE products SET price_cents = $new WHERE id = $id",
wo.Params{"new": 4900, "id": 42},
)
// 5. Tx — wraps Wo/Query/Exec in a server-side BEGIN...COMMIT. The callback's
// return value decides commit vs rollback. Useful when the program needs
// to branch between statements on intermediate results.
err = client.Tx(ctx, func(tx *wo.Tx) error {
r, err := tx.Wo(ctx,
`UPDATE products SET inventory.on_hand -= $qty
WHERE id = $pid AND inventory.on_hand >= $qty
RETURNING id AS pid;`,
wo.Params{"pid": 42, "qty": 2})
if err != nil { return err }
if r.Affected[0] == 0 { return wo.ErrInsufficientInventory }
_, err = tx.Wo(ctx, `
INSERT INTO orders (user_id, status) VALUES ($uid, 'pending') RETURNING id AS oid;
MATCH (u:user {id: $uid}), (p:product {id: $pid})
CREATE (u)-[:PURCHASED {order_id: $oid, qty: $qty}]->(p);
`, wo.Params{"uid": uid, "pid": 42, "qty": 2})
return err
})
```
`Wo` is the primitive; `Query`, `Exec`, and `Subscribe` are typed sugar. If the engine accepts the `.wo` source on disk, `client.Wo` accepts the same string over the wire.
### When to use raw `.wo` vs typed codegen
Both styles coexist in the same program; they share the connection pool.
| Use raw `.wo` (`client.Wo`) when | Use typed codegen (`client.Orders.Create`, etc.) when |
| --- | --- |
| Ad-hoc queries, admin tools, one-off scripts | The app's hot path — compile-time schema checking + IDE autocomplete pay for themselves |
| Cross-cutting queries that join multiple generated types | Per-type CRUD + subscriptions |
| Multi-statement transactions threading `RETURNING` aliases | Single-statement operations |
| DB repair, data migration, ad-hoc analytics | Anything the codegen already covers |
| You want to paste a block from a `.wo` source file straight into Go | You want refactor-safe struct field access |
The canonical rule: **write typed code first, drop to raw `.wo` when the type system gets in the way**. They interleave freely — a typed `client.Orders.Subscribe(...)` can run next to a raw `client.Wo(...)` admin query in the same handler.
## Subscribe — The Primary Use Case
Idiomatic Go for a stream of values is a channel read inside a `for` loop, cancelled by `context.Context`. That is the exact shape a `.wo` subscription should take:
```go
sub, err := client.Subscribe(ctx,
"LIVE SELECT sku, inventory.on_hand FROM products WHERE sku IN $skus",
wo.Params{"skus": cartSkus},
)
if err != nil { return err }
defer sub.Close()
for delta := range sub.Deltas() {
switch d := delta.(type) {
case wo.Insert:
log.Printf("new row: %+v", d.Row)
case wo.Update:
log.Printf("sku=%s on_hand=%d -> %d", d.Key, d.Old["on_hand"], d.New["on_hand"])
case wo.Delete:
log.Printf("removed: %s", d.Key)
case wo.Resync:
// server dropped our queue — refetch and resume
currentState = refetch()
}
}
// loop exits when:
// - ctx cancelled (client shutdown)
// - sub.Close() called (defer)
// - server sent COMPLETE (schema change, permission revoked)
// sub.Err() returns the reason
if err := sub.Err(); err != nil { log.Fatal(err) }
```
**Contract:**
- `sub.Deltas()` returns `<-chan wo.Delta` — standard read-only channel. The SDK closes it when the subscription ends.
- `ctx` cancellation immediately stops deliveries and closes the channel. No leaked goroutines.
- Ordering: deltas arrive in commit order. A `DELTA` on the wire always reflects a committed transaction.
- Back-pressure: the channel has a bounded buffer (default 1024). If it fills, the SDK's policy kicks in (see below).
## Typed SDK via `.wo` Schema Codegen
The `.wo` **schema layer** ([Phase 2](./02-wo-language.md)) declares types. `wo-gen` reads the type DSL — not the underlying `##sql`/`##doc`/`##graph` blocks — as its input; that way one Go struct corresponds to one entity, with embedded documents and graph-edge projections folded in naturally.
```wo
type Product {
id: Id
sku: SKU @unique
price: Money
meta: { title: Text, description: Markdown, images: [Url], reviews: [Review] }
inventory: { on_hand: Int @check(>= 0), reserved: Int = 0, reorder_at: Int }
purchased_by: multi User via Purchase -- inverse graph link
}
```
```bash
wo-gen --schema ./schema.wo --out ./internal/wodb
```
Produces:
```go
package wodb
// from `type Product` — embedded structs compile from the inline `{...}` fields
type Product struct {
ID int64 `wo:"id"`
SKU string `wo:"sku"`
Price Money `wo:"price"`
Meta ProductMeta `wo:"meta"`
Inventory InventoryLvl `wo:"inventory"`
PurchasedBy []PurchaseEdge `wo:"purchased_by"` // link-with-props → edge struct
}
// embedded document inside Product.Meta
type ProductMeta struct {
Title string `wo:"title"`
Description string `wo:"description"`
Images []string `wo:"images"`
Attributes map[string]string `wo:"attributes"`
Reviews []Review `wo:"reviews"`
}
// graph link carrying properties — target + edge props in one struct
type PurchaseEdge struct {
Target User `wo:"target"`
Order int64 `wo:"order"`
Qty int `wo:"qty"`
At time.Time `wo:"at"`
}
// registered live queries become typed helpers
func InventoryChanged(ctx context.Context, c *wo.Client, skus []string) (*wo.TypedSubscription[InventoryLvl], error)
```
One type declaration → one Go struct. Zero-property graph edges (`multi User @edge(:FOLLOWS)`) generate `Friends []User`; link-with-properties types generate `[]EdgeStruct`; computed fields become read-only struct fields populated by the planner.
**Transition path.** While the Phase 2 prototype is still authored directly in `##sql/##doc/##graph` blocks (the query layer), `wo-gen` accepts either — a file of type declarations, or the raw paradigm blocks — and emits the same Go output. The type DSL becomes mandatory only once Phase 6 full-stack blocks (which attach to types) start shipping.
Typed subscription loop loses all `interface{}` ceremony:
```go
sub, err := wodb.InventoryChanged(ctx, client, cartSkus)
if err != nil { return err }
defer sub.Close()
for d := range sub.C {
switch d.Kind {
case wo.DeltaUpdate:
log.Printf("sku=%s now %d in stock", d.Key, d.New.OnHand)
}
}
```
Generics (Go 1.18+) make `TypedSubscription[T]` a single parameterized type — no per-query generated struct. Only the `T` struct itself is generated.
## Connection Lifecycle
```go
type Client struct {
// opaque; holds a connection pool, codec, session registry
}
func Connect(ctx context.Context, dsn string, opts ...Option) (*Client, error)
func (c *Client) Close() error
func (c *Client) Ping(ctx context.Context) error
```
Inside the SDK, one TCP connection per client is fine for native protocol (multiplexed), but a small pool (2–4) helps when one connection's receive goroutine is saturated decoding a large result set. Connection state:
- **Connecting** → `HELLO` sent, waiting for `WELCOME`
- **Ready** → normal operation
- **Reconnecting** → transient network error; automatic exponential backoff; all subscriptions queued for re-registration
- **Closed** → terminal
**Reconnection semantics for subscriptions** (the subtle part): on reconnect, the SDK re-sends every active `SUBSCRIBE` frame. The server replies with a `RESYNC` marker and the current matching state. The app's subscription channel emits a single `wo.Resync{}` value so the consumer knows to rebuild local state. No delta is silently lost, no delta is silently duplicated.
## Options
Fluent options, not a bloated config struct:
```go
wo.WithAPIKey(key string)
wo.WithJWT(token string)
wo.WithMTLS(cert tls.Certificate)
wo.WithTLS(cfg *tls.Config)
wo.WithPoolSize(n int) // default 2
wo.WithSubscriptionBuffer(n int) // default 1024
wo.WithOverflowPolicy(wo.DropAndResync | wo.Coalesce | wo.Disconnect)
wo.WithLogger(l *slog.Logger)
wo.WithRetry(wo.RetryPolicy{...})
wo.WithProtocol(wo.ProtocolNative | wo.ProtocolGraphQL) // native default
```
## Transactions
`client.Tx` maps to [Phase 2's](./02-wo-language.md) `BEGIN ... COMMIT`. The callback's return value decides commit vs rollback:
- `return nil` → `COMMIT` sent, error only if server rejects commit
- `return err` → `ROLLBACK` sent, original `err` surfaced to caller
- `panic` → `ROLLBACK` sent, panic re-raised
- `ctx` cancel → `ROLLBACK` sent, `ctx.Err()` returned
Nested `tx.Wo`/`tx.Query`/`tx.Exec` route to the same server-side transaction — no connection hopping. The SDK enforces this by pinning the transaction to one connection for its lifetime. `RETURNING` aliases bound by one statement in the txn are visible to every later statement in the same txn through the server-side alias table (see [Phase 2 — Transaction Coordinator](./02-wo-language.md#cross-paradigm-transaction-coordinator)), so a Go `tx.Wo` call can leave `$oid` set and the next `tx.Wo` call can use it.
## Back-Pressure Handling in the SDK
The server's back-pressure policy from [Phase 4](./04-client-api.md) is mirrored client-side:
| Client situation | SDK behavior |
| --- | --- |
| Consumer reading channel fast enough | Normal delivery |
| Channel buffer full (1024 unread deltas) | Per `WithOverflowPolicy`: drop buffered + emit `wo.Resync`, or coalesce same-key updates, or close the subscription with `ErrOverflow` |
| Network stalled | `ctx.Deadline` + keepalive `PING` every 10s; stall > 30s → disconnect and reconnect |
| Server closed subscription | Channel closed, `sub.Err()` returns reason (schema change, permission revoked, engine shutdown) |
**Never block the receive goroutine on a full channel.** The SDK drains the socket no matter what; overflow policy decides what to do with the deltas it can't deliver.
## Full Cart-Inventory Example
End-to-end: Go HTTP handler that renders a cart page and keeps its inventory line live via Server-Sent Events to the browser. The SDK drives the upstream subscription to `.wo`:
```go
func (h *Handler) cartInventoryStream(w http.ResponseWriter, r *http.Request) {
skus := parseSkus(r.URL.Query().Get("skus"))
w.Header().Set("Content-Type", "text/event-stream")
w.Header().Set("Cache-Control", "no-cache")
flusher := w.(http.Flusher)
sub, err := wodb.InventoryChanged(r.Context(), h.db, skus)
if err != nil { http.Error(w, err.Error(), 500); return }
defer sub.Close()
for d := range sub.C {
payload, _ := json.Marshal(d)
fmt.Fprintf(w, "event: inventory\ndata: %s\n\n", payload)
flusher.Flush()
}
}
```
The browser connects once with `new EventSource('/cart/inventory?skus=...')`. The Go handler holds one subscription to `.wo`. When inventory commits in the database, the delta flows: engine → subscription registry → Go SDK channel → SSE stream → DOM update. Zero polling anywhere in the chain.
## Go-Specific Design Details
| Go idiom | Application |
| --- | --- |
| `context.Context` threading | Every method takes `ctx` as first arg; cancellation propagates to the wire |
| `io.Closer` | `Client`, `Tx`, `Subscription` all implement `Close() error` |
| Small interfaces | `type Runner interface { Wo(ctx, src, params) (wo.Result, error) }` — `*Client` and `*Tx` both satisfy it, so helper functions compose cleanly |
| `database/sql`-style `Scan` | `client.Query(...).Scan(&dest)` accepts struct, slice of struct, or primitives |
| `sql.Null*` analogues | `wo.NullString`, `wo.NullInt64`, `wo.NullDoc` for optional doc columns |
| Struct tags | `wo:"column_name"` + JSON-style for nested doc fields (`wo:"meta.title"`) |
| `errors.Is` / `errors.As` | `errors.Is(err, wo.ErrConflict)`, `wo.AsError(err, &woErr)` |
| No goroutine leaks | Every background goroutine tied to ctx or a sync.WaitGroup closed in `Client.Close()` |
| Testing via interfaces | `wo.DB` interface; provide `wotest.NewMock()` for unit tests; real embedded engine for integration |
## Comparison With Existing Go DB SDKs
| SDK | Query style | Subscriptions | Transactions | Typed results |
| --- | --- | --- | --- | --- |
| `database/sql` + `pq` | SQL strings | No | Yes | Manual `Scan` |
| `pgx` | SQL strings | `LISTEN/NOTIFY` only (no row-level) | Yes | Manual or `pgxscan` |
| `sqlc` | Generated Go funcs from `.sql` | No | Yes | Generated structs |
| `ent` | ORM | No | Yes | Generated |
| `go-redis` | Commands | Pub/sub + keyspace notifications (no query) | Multi/Exec | Manual |
| `surrealdb/surrealdb.go` | Raw queries | `Live()` returning channel | Yes | Manual |
| `gqlgen` / `machinebox/graphql` | GraphQL docs | WebSocket subscriptions | N/A | Generated |
| **`sa`** (this design) | `.wo` queries | **Native `LIVE` → typed channel** | Yes | Codegen from schema |
The reference points are `sqlc` (for the codegen pipeline) and `surrealdb-go` (for the subscription channel API). Combining their best ideas and tightening the subscription contract is what this SDK is.
## Reference Implementations To Steal From
- **surrealdb/surrealdb.go** — `Live()` returns a channel; closest API precedent. <https://github.com/surrealdb/surrealdb.go>
- **jackc/pgx** — reference quality for a Go database driver. Connection pool, copy protocol, prepared statements all done right. <https://github.com/jackc/pgx>
- **sqlc-dev/sqlc** — codegen from SQL to typed Go. The model for `.wo` → Go. <https://github.com/sqlc-dev/sqlc>
- **Khan/genqlient** — generated typed GraphQL client. Ergonomic precedent for typed query helpers. <https://github.com/Khan/genqlient>
- **nats-io/nats.go** — subscription-first API, back-pressure handled well. `sub.NextMsg(ctx)` and channel-based `ChanSubscribe` both supported. <https://github.com/nats-io/nats.go>
- **hasura/go-graphql-client** — GraphQL subscriptions over WebSocket in Go. <https://github.com/hasura/go-graphql-client>
## SDK Delivery
| Artifact | Purpose |
| --- | --- |
| `go.writeonce.dev/wo` | Runtime package: client, query, subscribe |
| `go.writeonce.dev/wo/wotest` | Mock client + in-memory engine for unit tests |
| `wo-gen` binary | Reads `schema.wo`, emits typed Go code |
| Go module example repo | Cart + inventory demo wired end-to-end |
| Generated docs | `go doc` + hosted examples |
Publishing strategy: semantic versioning, `v0.x` while the wire protocol is unstable, `v1.0` only after the protocol is frozen.
## Why The SDK Matters As Much As The Engine
A database with a beautiful engine and a painful client is a database no one uses. The e-commerce Go backends this targets are built under deadline — if `Subscribe` is not as easy as opening a channel, developers will reach for polling (`time.Tick` + `SELECT`) and defeat the whole architecture.
The success metric is blunt: **a developer who has never seen `.wo` before should have a working subscription to a live query inside 15 minutes**, counting install, schema codegen, and the first delta landing on their channel. If the SDK is any harder than that, the rest of this doc is academic.
## Future SDKs
Same shape, other languages:
- **TypeScript / browser** — fetch + WebSocket for GraphQL subscriptions; types via codegen from `.wo`. Highest priority after Go for a web-first product.
- **Rust** — direct native protocol, `tokio`-friendly, `impl Stream<Item = Delta>` for subscriptions.
- **Python** — async/await, `async for delta in sub` idiom.
- **Java / Kotlin** — Flow (Kotlin) or Reactive Streams (Java) for subscriptions.
Each follows the same rule: subscribe-to-query must be the shortest, most obvious thing in the API.

View file

@ -0,0 +1,389 @@
# Phase 6 — Low-Code Full-Stack: `.wo` as an Application Language
> Expand `.wo` from a query language into a declarative application DSL — schema, services, UI, business logic, and authorization in one language, compiled into a single binary.
**Previous**: [Phase 5 — Go Client SDK](./05-go-sdk.md) | **Index**: [database.md](../database.md)
---
Up to this point `.wo` is a **query language**. The next move is to expand it into an **application language** — a declarative, low-code/no-code DSL in the shape of **SAP Core Data Services (CDS)**: one language, one file extension, one compilation pipeline that produces database schema, service endpoints, UI screens, and business logic from the same source tree.
The reference precedent is SAP CDS, where a small amount of `.cds` code declares:
- Entities (tables), types, associations (relationships)
- Services that project entities to OData/REST
- UI annotations (`@UI.LineItem`, `@UI.Facet`) that drive SAP Fiori rendering
- Actions, functions, and authorization rules
From those declarations, SAP generates a full running application — data model, REST API, CRUD UI, authorization layer — with the developer writing almost no imperative code. `.wo` aims at the same target, for the same reason: **most enterprise and e-commerce applications are 90% CRUD on structured data with live views; declaring what you want and letting the runtime generate the rest is faster than hand-writing it**.
## Two Authoring Styles — Type-Attached vs Standalone
Behavior declarations fall into two groups:
| Group | Blocks | Authoring style |
| --- | --- | --- |
| **Behavior ON an entity** | `policy`, `on <event>`, `service` | **Type-attached** — declared inside the `type` block they govern. One name-resolution rule, zero cross-file coupling for single-entity behavior. |
| **Cross-entity composition** | `##ui`, `##app`, `##logic` that spans entities | **Standalone** — a screen composes multiple entities via `source:`; an app manifest names routes; a workflow that touches both orders and inventory needs its own block. |
Type-attached is the default — and the preferred authoring style because the schema layer ([Phase 2](./02-wo-language.md)) already names entities, and `policy read when author == $session.user` is most legible next to the `author: ref User` declaration it references.
Both forms compile to the same runtime model. A type-attached `policy` block is de-sugared into the same planner rewrite rule as a standalone `##policy` block — splitting is an authoring convenience.
## Project Layout
Convention over configuration. A `.wo` project is a tree of `.wo` files, each in a role-specific directory. The compiler discovers files by path.
```
myapp/
├── app/
│ ├── database/ # schema — `type` declarations (with inline
│ │ ├── article.wo # policy, on, service blocks per type)
│ │ ├── user.wo
│ │ └── order.wo
│ ├── ui/ # standalone screens — ##ui
│ │ ├── list.wo
│ │ ├── detail.wo
│ │ ├── form.wo
│ │ └── dashboard.wo
│ ├── logic/ # cross-entity workflows — ##logic
│ │ └── order-workflow.wo
│ ├── auth/ # cross-entity / session-level policies —
│ │ └── policies.wo # ##policy for things that don't fit on a type
│ ├── api/ # service bundles that expose many types —
│ │ └── services.wo # ##service for multi-entity APIs
│ └── app.wo # root: name, routes, theme, i18n
├── migrations/ # generated, versioned schema migrations
├── static/ # hand-written assets (images, custom CSS)
└── wo.toml # project metadata
```
Most single-entity behavior lives next to the `type` in `database/`; `logic/` and `auth/` and `api/` are for the cross-entity cases. Every file contributes to a **single compiled model**. Splitting is for humans; the runtime sees one graph of declarations.
## File Type Examples
**`app/database/article.wo`** — one `type` declaration covers relational fields, embedded document, graph edges, **and** the per-entity policy/triggers/service:
```wo
type Article {
id: Id
sys_title: Slug @unique
title: Text
published: Bool = false
author: ref User -- foreign key
meta: { -- embedded document
tags: [Text]
excerpt: Text
body_md: Markdown
reviews: [{ user: ref User, stars: Int, body: Markdown, at: Timestamp }]
}
created_at: Timestamp = now()
published_at: Timestamp?
related: multi Article @edge(:RELATED_TO)
prerequisites: multi Article @edge(:PREREQUISITE)
-- type-attached policy — replaces a separate ##policy block for
-- the single-entity case
policy read when published == true
policy read for role editor
policy read for role owner when author == $session.user
policy write for role editor
policy write for role owner when author == $session.user
policy delete for role admin
-- type-attached trigger — fires inside the transaction on commit
on update when old.published == false and new.published == true
do set self.published_at = now()
do emit "article.published"(self)
do enqueue "send-subscriber-emails" with { article_id: self.id }
-- type-attached service — exposes CRUD + subscribe on a REST path
service rest "/api/articles" expose list, get, create, update, delete, subscribe
}
type User { ... } -- authors; AUTHORED is derivable from Article.author via backlink
```
The compiler emits the underlying `##sql` relational row, `##doc` embedded structure, and `##graph` edges from the single `type` declaration. `multi Article @edge(:RELATED_TO)` declares a zero-property graph edge whose direction and label match the original graph sketch; `ref User` emits a foreign-key column in the relational store. Inverses (`User.articles: backlink Article.author`) generate an `AUTHORED` edge — or a plain inverse column, at the planner's discretion.
**`app/ui/list.wo`** — a live list view; renders to HTML, wires subscriptions automatically:
```wo
##ui
#article-list
title: "Articles"
source: article
live: true -- auto-subscribes via LIVE query
filter:
published = true
columns:
- sys_title label: "Slug"
- title label: "Title" searchable
- meta.tags label: "Tags" renderer: tag-chips
- created_at label: "Created" renderer: relative-date
- author.name label: "Author" join: author_id -> user
sort:
default: created_at desc
actions:
row-click: /article/:sys_title
create: /article/new role: editor
row-edit: /article/edit/:id role: editor | owner
row-delete: delete role: editor confirm: true
pagination: 20
```
**`app/ui/detail.wo`** — a detail view composed of nested renderers, including a graph traversal:
```wo
##ui
#article-detail
title: $article.title
source: article
key: sys_title
live: true
sections:
- header:
fields: [title, author.name, created_at]
- body:
renderer: markdown
source: meta.body_md
- related:
title: "Related Articles"
renderer: list
source: Article{ sys_title == $key }.related -- schema-layer path
columns: [title, meta.excerpt]
live: true
```
Screens stay standalone because they compose data from multiple types. The `source:` expression is a schema-layer path (preferred) or a raw `MATCH`/`SELECT` from the query layer — both resolve to the same planner input.
**`app/logic/order-workflow.wo`** — `##logic` is reserved for **cross-entity** triggers that don't belong on a single type. The on-article-published trigger lives on `type Article` (shown above); the on-order-placed trigger touches orders **and** every product in the line items, so it stays standalone:
```wo
##logic
#on-order-placed
when: insert(Order)
do:
- validate: self.total == sum(self.line_items.*.qty * self.line_items.*.unit)
- for-each item in self.line_items:
- update: Product{ id == item.product.id }
set inventory.on_hand -= item.qty
assert inventory.on_hand >= 0
```
**`app/auth/policies.wo`** — reserved for **session-level** or **cross-entity** rules that don't fit on a single type. Most RBAC lives type-attached (see the `policy read ...` block on `type Article` above). A standalone `##policy` is useful for things like "admins bypass all row filters":
```wo
##policy
#admin-bypass
applies_to: Article, Order, User
when: role == admin
effect: skip-row-filters
```
**`app/api/services.wo`** — reserved for **multi-entity** API bundles. Single-entity services live type-attached (see `service rest "/api/articles"` on `type Article` above). A bundle endpoint that exposes a curated subset or a custom aggregation goes here:
```wo
##service
#storefront
path: /api/storefront
protocols: [rest, graphql]
expose:
- Product as products operations: [list, get, subscribe]
- Article as articles operations: [list, get]
- categories: Product{ featured == true }.category -- custom path
```
**`app/app.wo`** — root manifest:
```wo
##app
name: "writeonce"
version: 1
theme: "light"
i18n: [en, de]
routes:
/ -> ui.article-list { filter: { published: true } }
/article/:slug -> ui.article-detail { key: $slug }
/admin/articles -> ui.article-list { role: editor }
```
## Compilation Pipeline
The `.wo` compiler loads every `.wo` file in the tree and emits a single runtime bundle:
```
app/**/*.wo
│
▼
┌─────────────┐
│ Parser │ one grammar, all block types
└─────────────┘
│
▼
┌─────────────┐
│ Analyzer │ name resolution across files, type check, policy check
└─────────────┘
│
▼
┌─────────────────────────────────────────┐
│ Unified Application Model (AST) │
└─────────────────────────────────────────┘
│ │ │ │ │
▼ ▼ ▼ ▼ ▼
┌────────┐ ┌──────────┐ ┌─────────┐ ┌────────┐ ┌──────────┐
│ Schema │ │ Services │ │ UI │ │ Logic │ │ Policies │
│ (DDL) │ │(endpoints)│ │(widgets)│ │(hooks) │ │ (rbac) │
└────────┘ └──────────┘ └─────────┘ └────────┘ └──────────┘
│ │ │ │ │
▼ ▼ ▼ ▼ ▼
Migrations HTTP/GraphQL HTML / JSON Commit- Planner
applied to / native manifest time filters
engine dispatch (SSR or triggers merged into
client) every query
```
Each leaf maps to a runtime component from the earlier phases:
- **Schema → engine**: the `.wo` DDL goes to the in-memory engine from [Phase 3](./03-inmemory-engine.md). Migrations rebuild the schema; data is preserved where possible.
- **Services → endpoints**: HTTP/GraphQL/native dispatch via the wire-protocol layer from [Phase 4](./04-client-api.md).
- **UI → widgets**: a new component — UI declarations compile to a render tree. Default renderer is server-rendered HTML (SSR) with a thin client runtime for subscription wiring. `live: true` on any UI node issues a `LIVE SELECT`/`LIVE MATCH` through the subscription engine and the client runtime swaps DOM fragments on each delta.
- **Logic → hooks**: triggers compile to server-side procedures invoked by the transaction coordinator on matching commits. Same transaction as the mutation — ACID across the hook's writes.
- **Policies → planner**: predicates intersected with every query/subscription at registration time (already covered in [Phase 4](./04-client-api.md)).
## How Subscriptions Wire Themselves
The key low-code payoff: a UI developer never writes subscription code. They write `live: true` on a list or detail, and:
1. The UI compiler inspects the view's `source` (table, document query, or graph `MATCH`).
2. It generates a `LIVE` query that returns exactly the fields the UI displays.
3. It emits a subscription handle in the rendered page.
4. The client runtime opens a WebSocket, registers the subscription, and binds incoming deltas to DOM fragments by key.
5. When a row changes in the engine, the delta flows: engine → subscription registry → client runtime → DOM patch.
Zero hand-written subscription code. Zero polling. Adding a new live column to a list is one line in a `.wo` file.
## Generated Application Stack
For the writeonce schema above, `sa build` produces:
| Layer | Generated From | Output |
| --- | --- | --- |
| Database schema | `app/database/*.wo` | In-memory engine arenas + migrations |
| REST / GraphQL / native endpoints | `app/api/*.wo` + `app/database/*.wo` | HTTP handlers, OpenAPI spec, GraphQL SDL |
| SSR HTML | `app/ui/*.wo` + routes in `app.wo` | Per-route renderers compiled into the server binary |
| Client runtime | `app/ui/*.wo` | Small JS bundle: subscription client + DOM patcher + form binding |
| Admin UI | All of the above | Auto-generated CRUD screens for every `##sql`/`##doc` entity (override any with a `##ui` block) |
| Typed SDKs | `app/database/*.wo` | Go/TypeScript/Rust clients per [Phase 5](./05-go-sdk.md) |
| Migrations | Schema diff vs. current database | Versioned forward/backward migrations in `migrations/` |
| Observability | Everything | Structured logs, query metrics, subscription lag dashboards |
Equivalent hand-written stack: schema (SQL), ORM models, REST controllers, GraphQL schema + resolvers, HTML templates, client JS, subscription plumbing, migrations, admin CRUD, SDKs. Likely **10,000–50,000 lines** for a small e-commerce site. `.wo` target: **~500 lines** across the `app/` tree.
## Developer Experience — The CLI
```bash
sa init myapp # scaffold with sensible defaults
cd myapp
sa dev # live-reload server — edit .wo, see changes instantly
sa build --target linux-amd64 # single static binary with everything inside
sa migrate --plan # preview schema migrations
sa migrate --apply # apply migrations
sa gen sdk --lang go --out ./sdk # emit typed client
sa deploy # upload to a running runtime node
```
`sa dev` is the make-or-break command. It must:
- Detect `.wo` changes via inotify (same mechanism as `wo-watch`)
- Recompile incrementally (~50 ms for a single-file change)
- Hot-swap UI renderers without losing client state
- Run schema migrations in a sandbox, surface conflicts before applying
- Keep open subscriptions alive across reloads (re-register on connect)
## Comparison With Other Declarative Full-Stack Systems
| System | Schema | UI | Logic | Live queries | Single binary |
| --- | --- | --- | --- | --- | --- |
| **SAP CDS** | `.cds` entities | `@UI` annotations → Fiori | Actions, functions | No (request-response) | No (Java/Node runtime) |
| **Hasura** | Reads from Postgres | No (external UI) | Actions, event triggers | Yes (polling-based internally) | No |
| **Supabase** | Postgres schema | Auto-admin UI only | Edge functions, triggers | Yes (logical replication) | No |
| **Retool / AppSmith / Budibase** | External DB | Visual drag-and-drop | JS snippets | Partial | No |
| **Wasp** (`wasp-lang.org`) | `.wasp` + Prisma | React components | JS functions | No | No (Node + React) |
| **RedwoodJS** | `.sdl` + Prisma | React | JS | No | No |
| **Django + Admin** | Python models | Auto-admin only | Python views | No | No |
| **Phoenix LiveView** | Ecto schemas | HEEx templates | Elixir functions | **Yes** (native) | No (BEAM runtime) |
| **Anvil** (`anvil.works`) | Proprietary | Python drag-and-drop | Python | No | No |
| **`.wo`** (this design) | `type` DSL (unified) over `##sql`+`##doc`+`##graph` substrate | `##ui` declarations + type-attached | type-attached `on <event>` + `##logic` | **Yes** (native) | **Yes** |
The differentiators: **cross-paradigm schema** (no competitor unifies SQL + document + graph in one DDL), **engine-native subscriptions** (most bolt on a replication layer or poll), and **single static binary** as the deployment unit (no separate database process, no separate UI server, no separate message bus).
Closest philosophical precedents:
- **SAP CDS** for the declaration-first application language — the explicit inspiration.
- **Phoenix LiveView** for the subscription-wired UI model.
- **Wasp** for the `.wasp` → full stack compilation pipeline.
- **Django Admin** for the "generate CRUD from the model" reflex.
## Scope Addition
This is a compiler + a renderer + a UI toolkit on top of Phases 2–4. Rough new components:
| Component | Work |
| --- | --- |
| `type` DSL parser + schema compiler | Entity declarations → `##sql`/`##doc`/`##graph` physical schema; type-attached `policy`/`on`/`service` → same runtime components as standalone blocks |
| `##ui` grammar + analyzer | UI widget tree, field bindings, renderer dispatch |
| Trigger compiler | Type-attached `on <event>` + standalone `##logic`; both run inside txn coordinator |
| Policy compiler + planner integration | Type-attached `policy` + standalone `##policy` both intersected with every query at registration (some of this exists in Phase 2 already) |
| Service compiler + dispatch table | Type-attached `service` + standalone `##service`; endpoint registration at startup |
| UI render tree → SSR HTML | Template engine, layout system, component library (table, form, chart, etc.) |
| Client runtime (~50 KB JS) | Subscription client, DOM patcher, form binding, validation |
| Auto-admin UI | Generic CRUD screens per entity, override-able with `##ui` |
| Migration engine | Schema diffing, forward/backward migrations, online reshape for the in-memory engine |
| CLI (`sa` binary) | `init`, `dev`, `build`, `migrate`, `gen`, `deploy` |
| Dev-mode live reload | inotify + incremental compiler + client hot-swap |
| Hosted runtime | Optional — for `sa deploy` to work without self-hosting |
Rough effort on top of Phases 2–4: **12–24 months** with a small team, most of it in the UI compiler and client runtime — that's where the complexity lives, not in the language spec.
## Honest Framing
This section is the endgame, not the next step. The sensible build order:
1. Ship the engine ([Phase 2](./02-wo-language.md), in-memory + io_uring durability via [Phase 3](./03-inmemory-engine.md)) — query layer (`##sql`/`##doc`/`##graph`) only.
2. Ship the wire protocol + Go SDK ([Phase 4](./04-client-api.md) + [Phase 5](./05-go-sdk.md)).
3. Ship the **schema-layer `type` DSL** that compiles to the three paradigm blocks. From this point forward, authoring happens against types; the paradigm blocks become an artifact the compiler emits.
4. Add type-attached `service` (and standalone `##service` for bundles) — declarative endpoints.
5. Add type-attached `policy` (and standalone `##policy` for cross-entity rules) — declarative authorization.
6. Add type-attached `on <event>` triggers (and standalone `##logic` for cross-entity workflows) — declarative triggers.
7. Add `##ui` — declarative rendering. **This is where `.wo` becomes a low-code platform.**
8. Add the CLI, live-reload, admin UI.
Each step is shippable on its own. Every step after (3) converts imperative code developers are writing by hand into declarative code they write once. The value compounds: by step (7), a small e-commerce app is a ~500-line `.wo` tree instead of a ~50 KLOC TypeScript/Go/SQL repository.
The risk is the same as every low-code platform: the 20% of use cases outside the declarative model have to have an escape hatch. `.wo` reserves one: any `##ui` node can point to a custom server-rendered template, and any `##logic` block can call out to a host-language plugin (Go/Rust/Wasm). Without that escape hatch, low-code becomes no-code in the pejorative sense — you can build 80% of the app and the rest is impossible.
## Reference Implementations Worth Studying
- **SAP CDS** — <https://cap.cloud.sap/docs/cds/>. Read the CDS Language Reference cover to cover before designing `##ui`. Thirty years of ERP app-generation is condensed into that spec.
- **Wasp** — <https://github.com/wasp-lang/wasp>. Open-source `.wasp` → React + Node + Prisma compiler. Closest living sibling to `.wo`.
- **Phoenix LiveView** — <https://github.com/phoenixframework/phoenix_live_view>. The rendering-subscription loop done right in Elixir.
- **HTMX + Hyperscript** — <https://htmx.org>. Tiny client runtime that consumes server-rendered HTML fragments on events. Good model for `.wo`'s client bundle.
- **Retool / Budibase / Appsmith (open source)** — <https://github.com/Budibase/budibase>, <https://github.com/appsmithorg/appsmith>. Visual low-code; useful to see which UI primitives users actually ask for.
- **SurrealDB `define` syntax** — SurrealQL includes declarative `DEFINE TABLE`, `DEFINE FIELD`, `DEFINE EVENT` that are partway toward an application DSL. Worth studying for how much declaration fits inside a query language.
- **Django admin source** — `django/contrib/admin/`. The canonical "CRUD from models" implementation; read how it introspects schema to generate list/detail/edit views.
- **PocketBase** — <https://github.com/pocketbase/pocketbase>. Single-binary Go app with SQLite, auto-admin UI, realtime subscriptions. Proves the single-binary low-code model is buildable. `.wo` is what PocketBase would look like if its data model were multi-paradigm and its query engine were custom.
## Why This Belongs in This Series
The question that opened [Phase 1](./01-evaluation.md) — "should writeonce use a document or graph database?" — has now inverted. The answer drove past "no, use flat files" through "build your own query language" and "build your own ACID engine" and "build your own wire protocol" to arrive here: **a full-stack declarative application platform where the database, the subscriptions, the UI, and the business logic are one artifact compiled from one language**.
That is the actual ambition. Every earlier phase is a subcomponent of this one. Decide honestly whether the project is a blog engine that needed a graph index, or an application platform that happens to start as a blog engine. The answer determines which phases are scope and which are cautionary.

View file

@ -0,0 +1,209 @@
# Phase 7 — Replacing `wo-seg` with the writeonce Database
> A phased coexistence plan: abstract the article store behind a trait, stand up the `.wo` engine as a second implementation, dual-run, cut over, decommission.
**Previous**: [Phase 6 — Low-Code Full-Stack](./06-lowcode-fullstack.md) | **Index**: [database.md](../database.md)
---
## Context
Today's writeonce runtime stores articles in a hand-rolled append-only file format:
- **`crates/wo-seg`** (~475 LOC) — `.seg` binary file: magic + header + `[u32 length][u8 flags][payload]` records serialized with `bincode`. `SegWriter::append()` returns a byte offset usable as an index pointer. Tombstoning flips a flag byte. No transactions, no MVCC, no concurrent writers.
- **`crates/wo-index`** — sidecar `title.idx`, `date.idx`, `tags.idx` files built from the `.seg`. Queries hit the index to resolve to a byte offset, then the `.seg` to load the record.
- **`crates/wo-store`** — composes the two above, owns cold-start (rebuild from `content/`), and exposes the query API the rest of the system calls: `get_by_title`, `list_published`, `list_by_tag`, `list_by_date_range`, `count_published`, `ingest_article`, `article_version`.
This is the Phase 1 answer ("no external DB — add a petgraph-backed `mappings.idx`") made flesh. It is correct for a single-writer blog and wrong for everything Phases 2–6 want to deliver: no ACID across multiple shapes, no live subscriptions, no cross-paradigm queries, no declarative schema, no codegen.
The six-phase `.wo` design is the replacement. This doc plans the migration — how to swap wo-seg for the `.wo` engine **without halting writeonce** while the engine is built over multiple quarters.
## Intended Outcome
- `crates/wo-seg` is deleted.
- `crates/wo-store` either (a) becomes a thin facade over the `.wo` engine or (b) disappears, with callers depending directly on the engine's Rust SDK.
- Writeonce's serving path is unchanged from the user's perspective throughout the migration.
- The `.wo` engine reaches production-ready status incrementally; each milestone is independently shippable.
## Strategy — Phased Coexistence
Do **not** big-bang. The seg-based store works; replacing it takes many months. Instead:
1. **Abstract** the existing store behind a Rust trait — one weekend of mechanical refactor, zero behavior change.
2. **Build** the `.wo` engine crates next to seg, not in its place. Port the C++ prototype (`prototypes/wo-db/`) to Rust so the engine lives in the same Cargo workspace as the blog.
3. **Dual-run** — writes go to both backends, reads to seg. Compare results in CI and on production data. This surfaces engine bugs without user impact.
4. **Cut over** reads once the engine passes dual-run. Writes still hit seg as a cold standby.
5. **Decommission** seg when enough time has passed without rollback and a restore-from-seg fallback is no longer load-bearing.
```
wo-seg + wo-store ─────────► trait-abstracted ─────► dual-write, read seg ─────► read wo-db, write both ─────► wo-db only, delete wo-seg
(today) (phase A) (phase B) (phase C) (phase D)
```
Each transition is reversible — flip one feature flag or swap one trait object back.
## Proposed Crate Layout
Port the C++ prototype (`prototypes/wo-db/src/*`) to Rust, split along the natural seams. New-runtime crates are unprefixed; v1 crates keep `wo-` in `reference/crates/`.
| Crate | Purpose | Prototype source | Phase |
| --- | --- | --- | --- |
| `ql` | `.wo` grammar: lexer, parser, AST | `src/lexer.*`, `src/parser.*`, `src/ast.hpp` | 2 |
| `value` | tagged `Value` + path utilities | `src/value.hpp`, path helpers in `src/storage.cpp` | 2 |
| `engine` | in-memory executor (sql / doc / graph), schema catalog | `src/storage.*`, `src/executor.*` | 2 |
| `txn` | MVCC, snapshot isolation, `RETURNING` alias table | new (Phase 2 milestone 3) | 2 |
| `wal` | write-ahead log + fsync + crash recovery | new ([Phase 3](./03-inmemory-engine.md)) | 3 |
| `sub` | live subscriptions — delta frames on commit | new ([Phase 4](./04-client-api.md)) | 4 |
| `http` | wire protocol — REST / GraphQL-over-WS / native codec | new ([Phase 4](./04-client-api.md)) | 4 |
| `db` | top-level facade: `open()`, `Tx`, `Query`, `Subscribe` — the Rust SDK | integrates the above | 2–4 |
| `gen` | codegen: `.wo type` → Rust structs, Go structs, TypeScript | `sa-gen`/`wo-gen` in [Phase 5](./05-go-sdk.md) | 5 |
All 15 crates (these 14 plus the existing `rt` binary crate) now exist as empty skeletons in `crates/`. See [`crates/README.md`](../../../crates/README.md) and [`docs/plan/done/01-scafolding-crates.md`](../../plan/done/01-scafolding-crates.md) for the scaffolding plan that landed them.
Today's `reference/crates/wo-seg` and `reference/crates/wo-index` remain in the v1 nested workspace for the entire migration window. They disappear only at the end of Phase D.
`reference/crates/wo-store` evolves but survives — it becomes the writeonce-specific glue layer (trait, article domain model, content-directory cold-start) whose backend is swappable.
## Phase A — Abstract the Article Store
**Goal**: every caller depends on a trait, not on `wo_store::Store` directly. Zero behavior change.
**Work**:
- Define `trait ArticleStore` in `wo-store/src/lib.rs` with the existing public API:
```rust
pub trait ArticleStore: Send + Sync {
fn get_by_title(&self, sys_title: &str) -> io::Result<Option<Article>>;
fn list_published(&self, skip: usize, limit: usize) -> io::Result<Vec<Article>>;
fn list_by_tag(&self, tag: &str) -> io::Result<Vec<Article>>;
fn list_by_date_range(&self, start: i64, end: i64) -> io::Result<Vec<Article>>;
fn count_published(&self) -> io::Result<usize>;
fn ingest_article(&mut self, json_path: &Path) -> io::Result<String>;
fn article_version(&self, sys_title: &str) -> Option<u64>;
fn content_dir(&self) -> &Path;
}
```
- Rename the existing `Store` struct to `SegStore` and implement `ArticleStore` for it. Re-export `SegStore as Store` for one release to avoid churn at call sites.
- Change `wo-route`, `wo-serve`, `wo-sub`, `wo-htmlx` to take `&dyn ArticleStore` (or generic `<S: ArticleStore>`). The trait import stays in `wo-store`; concrete impls move to sibling crates.
- Add a tiny `wo-store::open(content_dir, data_dir) -> Arc<dyn ArticleStore>` factory that picks the backend based on a config env var (`WO_STORE_BACKEND=seg|db|dual`).
**Exit criteria**: `cargo test` passes; `wo serve` boots unchanged; git log shows one PR.
## Phase B — Stand Up `wo-db` in Rust
**Goal**: a Rust `wo-db` crate that speaks the full `.wo` grammar from the prototype, stored in memory, with an `ArticleStore` impl mapping writeonce's `Article` onto the relational paradigm.
**Work**:
- Port `prototypes/wo-db/` (C++) to Rust crates per the layout table above. The `wo` namespace becomes the `wo_*` crate family; the test suites (`tests/smoke.wo`, `tests/checkout.wo`) run as Rust integration tests.
- Define a `.wo` schema for the writeonce domain (in a new file, `crates/wo-store/schema.wo`):
```wo
type Article {
id: Id
sys_title: Slug @unique
title: Text
published: Bool = false
published_at: Timestamp?
author: Text
tags: [Text]
meta: { excerpt: Text, body_md: Markdown }
}
```
- Add `DbStore` — a second `ArticleStore` impl that translates calls into `.wo` queries:
- `get_by_title(t)` → `SELECT * FROM Article WHERE sys_title = $t` (one row)
- `list_published(skip, limit)` → `SELECT * FROM Article WHERE published = true ORDER BY published_at DESC LIMIT $limit OFFSET $skip`
- `list_by_tag(t)` → `SELECT * FROM Article WHERE $t IN tags`
- `list_by_date_range(a, b)` → `SELECT * FROM Article WHERE published_at BETWEEN $a AND $b`
- `count_published` → `SELECT COUNT(*) FROM Article WHERE published = true`
- `ingest_article(path)` → load JSON → `INSERT INTO Article (…)`
- Cold-start path: when the data dir is empty, `DbStore::open` loads all `content/*.json` the same way `SegStore::open` does today and inserts into the engine.
- Gate behind `#[cfg(feature = "db-backend")]` so seg-only builds keep working until Phase C.
**Exit criteria**: `DbStore` passes the same unit tests as `SegStore` (rename `Store` → `ArticleStore` in test assertions). Memory footprint and per-query latency measured against seg; both within an order of magnitude.
## Phase C — Dual-Write, Read Seg
**Goal**: every mutation hits both backends; reads stay on seg; a differ flags mismatches.
**Work**:
- Add `DualStore` — a third `ArticleStore` impl that forwards writes to both `SegStore` and `DbStore` and returns `SegStore` results for reads.
- Add a background task (`wo-store::differ`) that on every ingest runs every query method against both backends and compares results. Mismatches → structured log entry (`store_mismatch` event) + a Prometheus counter.
- Set `WO_STORE_BACKEND=dual` on staging for two weeks, then on prod behind a rollout flag.
**Exit criteria**: zero `store_mismatch` events for 14 consecutive days on production traffic.
## Phase D — Cut Over Reads, Keep Seg as Fallback
**Goal**: reads served from `DbStore`; seg still receives writes and is kept queryable as a cold standby.
**Work**:
- Invert `DualStore`: writes to both, reads from `DbStore`.
- Add an admin command `wo db verify --against seg` that re-runs the differ on demand (for post-incident checks).
- After a stable month, remove `DualStore` entirely. `WO_STORE_BACKEND=db` becomes the only supported value.
**Exit criteria**: one month with no read-path regressions; no active rollback capability needed for routine ops.
## Phase E — Decommission `wo-seg`
**Goal**: delete `crates/wo-seg`, shrink `crates/wo-store` to the trait + content-directory cold-start.
**Work**:
- Delete `crates/wo-seg`. Remove `wo-seg` from `Cargo.toml` workspace members and from `wo-store/Cargo.toml` deps.
- Delete `SegStore` from `wo-store`. The trait `ArticleStore` and `DbStore` remain.
- Delete the `.seg` file from production data directories (via a migration: verify `DbStore` has every record, then `rm`).
- Delete `crates/wo-index` **if and only if** `DbStore` has replaced its indexes with the engine's internal ones. If the LSM/graph indexes inside `wo-db` cover the three sidecar indexes (title, date, tags) — expected — then wo-index goes too. If any index is still load-bearing outside the engine, keep it.
**Exit criteria**: CI is green with the deletions; production runs a release cycle without rollback; `rg "wo-seg\|wo_seg"` returns zero hits.
## Integration Touchpoints
These crates reference the store today and will need light updates for Phase A (trait swap):
| Crate | Current coupling | Change |
| --- | --- | --- |
| `wo-store` | owns `Store`, depends on `wo-seg` + `wo-index` | gains trait + factory + dual-write impl (A–C); shrinks to facade in E |
| `wo-route` | likely consumes `&Store` | accept `&dyn ArticleStore` |
| `wo-serve` | HTTP handlers read the store | accept `Arc<dyn ArticleStore>` |
| `wo-sub` | subscription layer | later — see below |
| `wo-htmlx` | may read article state during render | accept trait or projection |
| `wo-watch` | inotify-driven ingest | unchanged; still calls `ingest_article` |
| `wo-rt` | runtime glue | pass the trait object through |
`wo-sub` is a special case. Today it likely polls or reacts to `article_version` monotonic counters. When Phase 4 activates `LIVE` queries inside `wo-db`, `wo-sub` should stop doing its own diffing and become a pass-through for engine-emitted deltas. That transition happens in Phase C/D, not Phase A — it's not required for the trait refactor.
## Risks
1. **Cold-start cost.** `SegStore` builds its indexes in one pass over `.seg`. `DbStore` has to parse JSON from `content/` the same way but also commit through the engine's WAL. If this is slow, add a `wo db import --from-seg <path>` shortcut that bulk-loads from an existing `.seg` without going through the ingest path.
2. **Memory footprint.** Today's seg-based path `mmap`s the file; the `.wo` engine is RAM-primary. For a blog with hundreds of articles, immaterial; for a larger dataset, Phase 3's SSD-backed variant is what's needed.
3. **Article → `.wo` type drift.** `wo_model::Article` is the canonical domain type today. The `.wo` schema mirrors it, but if the two diverge (a new field is added to `Article` but not to the schema), queries silently drop that field. Mitigation: `wo-gen` should include a `--verify wo_model::Article` mode in Phase 5 that fails CI on drift.
4. **Dual-write contention.** If ingest becomes the bottleneck during Phase C, time-box dual-write: drop it after 14 clean days rather than running it indefinitely.
5. **Feature flag sprawl.** `WO_STORE_BACKEND` should be the only config knob. Resist per-method flags.
## What's Out of Scope for This Doc
- The engine internals themselves — those live in Phases 2–4.
- The Phase 5 SDK (`wo-gen`, typed Go client) — wo-store callers are Rust, and Rust codegen is part of `wo-gen` but not a blocker.
- `##ui` / `##policy` / `##logic` / `##service` — those are Phase 6 and assume the engine is already running.
- Any graph-first features (mappings, `RELATED_TO` traversal) — they become trivially available once `DbStore` is live, but don't need to gate the seg → db cutover.
## Verification
Each phase has its own exit criteria above. End-to-end verification for the whole migration:
1. **Parity** — After Phase B: a shadow script replays one week of production ingest through `DbStore` in a sandbox; every query from the shadow matches seg.
2. **Latency** — After Phase D: p50/p95/p99 of `get_by_title`, `list_published`, `list_by_tag` are at or below the seg baseline. Measured by the existing request-timing middleware, not synthetic benchmarks.
3. **Crash safety** — After Phase 3 WAL ships: `kill -9` during write, reopen, confirm the committed state matches and uncommitted writes are gone. Automated test.
4. **Decommission audit** — After Phase E: `rg 'wo-seg|wo_seg|\.seg\b' crates/` returns zero; `cargo deny check` passes; production restart ingests from `content/` with no `.seg` file present.
## Related Documents
- [02-wo-language.md](./02-wo-language.md) — the two-layer `.wo` language the engine speaks
- [03-inmemory-engine.md](./03-inmemory-engine.md) — the storage engine behind `wo-db`
- [04-client-api.md](./04-client-api.md) — wire protocol and `LIVE` subscriptions
- [05-go-sdk.md](./05-go-sdk.md) — the Go SDK built from `.wo` types via `wo-gen`
- [01-evaluation.md](./01-evaluation.md) — why writeonce built `wo-seg` in the first place, and why that choice still looks right for the blog even as the platform grows past it
- [../05-datalayer.md](../05-datalayer.md) — current `.seg` + `.idx` implementation details
- `prototypes/wo-db/` — the C++ prototype of the `.wo` engine, the reference implementation the Rust port follows

View file

@ -251,3 +251,454 @@ Concurrency: event loop = fibers = async/await >> threads
```
The right choice depends on the workload. For writeonce — an event loop. For a database with millions of queries in flight — fibers or async. For CPU-bound parallel work — kernel threads.
What are fibers? #
Fibers are a lightweight thread of execution similar to OS threads. However, unlike OS threads, they’re cooperatively scheduled as opposed to preemptively scheduled. What this means in plain English is that fibers yield themselves to allow another fiber to run. You may have used something similar to this in your programming language of choice where it’s typically called a coroutine, there’s no real distinction between coroutines and fibers other than that coroutines are usually a language-level construct, while fibers tend to be a systems-level concept.
Other names for fibers you may have heard before include:
green threads
user-space threads
coroutines
tasklets
microthreads
There are very few and minor differences between fibers and the above list. For the purposes of this document, we should consider them equivalent as the distinctions don’t quite matter.
Scheduling #
At any given moment the OS is running multiple processes all with their own OS threads. All of those OS threads need to be making forward progress. There’s two classes of thought when it comes to how you solve this problem.
Cooperative scheduling
Preemptive scheduling
It’s important to note that while you may observe that all processes and OS threads are running in parallel, scheduling is really providing the illusion of that. Not all threads are running in parallel, the scheduler is just switching between them quickly enough that it appears everything is running in parallel. That is they’re concurrent. Threads start, run, and complete in an interleaved fashion.
It is possible for multiple OS threads to be running in parallel with symmetric multiprocessing (SMP) where they’re mapped to multiple hardware threads, but only as many hardware threads as the CPU physically has.
Premptive scheduling #
Most people familiar with threads know that you don’t have to yield to other threads to allow them to run. This is because most operating systems (OS) schedule threads preemptively.
The points at which the OS may decide to preempt a thread include:
IO
sleeps
waits (seen in locking primitives)
interrupts (hardware events mostly)
The first three in particular are often expressed by an application as a system call. These system calls cause the CPU to cease executing the current code and execute the OS’s code registered for that system call. This allows the OS to service the request then resume execution of your application’s calling thread, or another thread entierly.
This is possible because the OS will decide at one of the points listed above to save all the relevant state of that thread then resume some other thread, the idea being that when this thread can run again, the OS can reinstate that thread and continue executing it like nothing ever happened. These transition points where the OS switches a thread are called context switches.
There’s a cost associated with this context switching and all modern operating systems have made great deals of effort to reduce this cost as much as possible. Unfortunately, that overhead begins to show itself when you have a lot of threads. In addition, recent cache side channel attacks like: Spectre, Meltdown, Spoiler, Foreshadow, and Microarchitectural Data Sampling on modern processors has led to a series of both user-space and kernel-space mitigation strategies, some of which increased the overhead of context switches significantly.
You can read more about context switching overhead in this paper.
Cooperative scheduling #
This idea of fibers yielding to each other is what is known as cooperative scheduling. Fibers effectively move the idea of context switching from kernel-space to user-space and then make those switches a fundamental part of computation, that is, they’re a deliberate and explicitly done thing, by the fibers themselves. The benefit of this is that a lot of the previously mentioned overhead can be entierly eliminated while still permitting an excess count of threads of execution, just in the form of these fibers now.
The problem with multi-threading #
There’s many problems related to multi-threading, most obviously that it’s difficult to get right. Most proponents of fibers make false claims about how this problem goes away when you use fibers because you don’t have parallel threads of execution. Instead, you have these cooperatively scheduled fibers which yield to each other. This means it’s not possible to race data, dead lock, live lock, etc. While this statement is true when you look at fibers as a N:1 proposition, the story is entierly different when you introduce M:N.
N:1 and what it means #
Most documentation, libraries, and tutorials on fibers are almost exclusively based around using a single thread given to you by the OS, then sharing it among multiple fibers that cooperatively yield and run all your asynchronous code. This is called N:1 (“N to one”). N fibers to 1 thread, and it’s the most prevalent form of fibers. This is how Lua coroutines work, how Javascript’s and Python’s async/await work, and it’s not what you’re interested in doing if you actually want to take advantage of hardware threads. What you’re interested in is M:N, (“M to N”) M fibers to N threads.
M:N and what it means #
The idea behind M:N is to take the model given to us by N:1 and map it to multiple actual OS threads. Just like we’re familiar to the concept of thread pools where we execute tasks, here we have a pool of threads where we execute fibers and those fibers get to yield more of themselves on that thread.
I should stress that M:N fibers have all the usual problems of multi-threading. You still need to syncronize access to resources shared between multiple fibers because there’s still multiple threads.
The problem with thread pools #
A lot of you may be wondering how this is different from traditional task based parallelism choices seen in many game engines and applications. The model where you have a fixed-size pool of threads you queue tasks on to be executed at some point in the future by one of those threads.
Locality of reference #
The first problem is locality of reference. The data-oriented / cache-aware programmers reading this will have to mind my overloading of that phrase because what I’m really talking about is the resources that a job needs access to are usually local. The job isn’t going to be executed immediately, but rather when the thread pool has a chance to. Any resource that job needs access to, needs to be available for the job at some point in the future. This means local values need their lifetime’s extended for an undefined amount of time.
There’s many ways to solve this lifetime problem, except they all have overhead. Consider this trivial example where I want to asynchronously upload a local file to a webserver.
## Async Runtime for the `.wo` Database
The writeonce event loop (`wo-event` / `wo-rt`, described in [async.md](./async.md)) is a **single-threaded epoll loop with callbacks**. It works because the content workload is microseconds per request and single-writer. The `.wo` database (described in the [database series](./database/02-wo-language.md)) inverts every one of those assumptions:
| writeonce (blog) | `.wo` database |
| --- | --- |
| Single-threaded — one handler at a time | Thousands of queries in flight concurrently |
| epoll — kernel wakes you on fd readiness | io_uring — userland submits I/O, kernel completes asynchronously |
| Callbacks — run to completion, return to loop | Tasks — query execution spans multiple I/O waits (WAL write, network send) |
| No blocking — handlers are microsecond-fast | Graph traversals and join plans can be milliseconds of CPU |
| Single writer — no contention | MVCC — many readers and writers touching shared data structures |
| Read-heavy | Write-heavy on hot paths (checkout, inventory) |
A new runtime is needed. This section designs it, building on the fiber/async/event-loop theory above.
### What the Runtime Must Schedule
Every subsystem in the `.wo` engine produces work with a different shape:
| Subsystem | Work shape | Blocking? | Parallelizable? |
| --- | --- | --- | --- |
| **Client accept** | Wait for incoming TCP connection | I/O-bound | No — one listener fd |
| **Request decode** | Parse binary frame or GraphQL | CPU, microseconds | Yes — per connection |
| **Query planning** | Compile `.wo` to execution plan | CPU, microseconds | Yes — per query |
| **Query execution** | Traverse B+ tree / LSM / graph | CPU, microseconds to milliseconds | Yes — per query |
| **WAL append** | Serialize + io_uring write + fsync | I/O-bound (NVMe) | Batched — group commit |
| **MVCC bookkeeping** | Version chain prepend, snapshot management | Atomic CAS, nanoseconds | Yes — per record |
| **Subscription matching** | Evaluate deltas against registered predicates | CPU, microseconds | Yes — per subscription |
| **Client push (DELTA)** | io_uring send to client socket | I/O-bound | Yes — per connection |
| **Checkpoint** | Snapshot RAM arenas to SSD | I/O-bound, background | Single background task |
| **Vacuum** | Prune old MVCC versions | CPU, background | Parallelizable by range |
The runtime’s job is to keep all of these progressing concurrently — mixing CPU-bound query execution with I/O-bound disk and network operations — without blocking any subsystem on another.
### Four Candidate Architectures
#### 1. Thread-Per-Core / Shared-Nothing (Seastar / ScyllaDB)
Each CPU core runs its own independent event loop. No shared memory between cores. Communication is message-passing over lock-free queues.
```
Core 0 Core 1 Core 2 Core 3
┌──────────┐ ┌──────────┐ ┌──────────┐ ┌──────────┐
│ io_uring │ │ io_uring │ │ io_uring │ │ io_uring │
│ ring │ │ ring │ │ ring │ │ ring │
│ │ │ │ │ │ │ │
│ local │ ←msg→ │ local │ ←msg→ │ local │ ←msg→ │ local │
│ shard of │ │ shard of │ │ shard of │ │ shard of │
│ data │ │ data │ │ data │ │ data │
└──────────┘ └──────────┘ └──────────┘ └──────────┘
```
**How it works**: data is hash-partitioned across cores. A query for `product.id = 42` routes to the core that owns shard `hash(42) % num_cores`. That core executes the query entirely locally — no locks, no contention.
**Pros**: zero contention on the hot path; each core runs flat-out; NUMA-friendly; no lock overhead.
**Cons**: cross-shard queries (joins, graph traversals spanning partitions) require inter-core messaging; range scans hit all shards; transaction coordination across shards is complex; programming model is unfamiliar.
**Used by**: ScyllaDB/Seastar (C++), Redpanda (C++), Glommio (Rust).
#### 2. Work-Stealing Async Runtime (tokio / Rust `async/await`)
M:N scheduling of `async` tasks across a thread pool. Tasks (futures) yield at `.await` points; the scheduler steals work from busy threads.
```
┌─────────────────────────────────────────┐
│ tokio runtime (M:N) │
│ ┌────────┐ ┌────────┐ ┌────────┐ │
│ │ worker │ │ worker │ │ worker │ │
│ │ thread │ │ thread │ │ thread │ │
│ │ │ │ │ │ │ │
│ │ task │ │ task │ │ task │ │
│ │ task │ │ task │←steal──┐ │ │
│ │ task │ │ │ │ task │
│ └────────┘ └────────┘ └────────┘ │
│ shared task queues │
└─────────────────────────────────────────┘
```
**How it works**: each query becomes an `async fn` that `.await`s I/O (network reads, WAL writes). The executor multiplexes thousands of tasks across a small thread pool. Shared state is accessed through `Arc<RwLock<T>>` or lock-free structures.
**Pros**: familiar Rust async model; excellent library ecosystem; handles mixed I/O and CPU work; tokio has 7+ years of production hardening.
**Cons**: `async/await` infects the entire codebase (everything must be async); shared state needs locks or lock-free structures; work-stealing adds scheduling overhead; harder to reason about NUMA locality.
**Used by**: SurrealDB (Rust + tokio), TiKV (Rust + tokio), Materialize (Rust + tokio).
#### 3. M:N Fiber Runtime (Go / Erlang)
User-space fibers with their own stacks, scheduled cooperatively across OS threads. Each query is a fiber; yield points are implicit at I/O boundaries.
```
┌────────────────────────────────────────┐
│ Go-style runtime │
│ │
│ goroutine goroutine goroutine │
│ goroutine goroutine goroutine │
│ goroutine goroutine goroutine │
│ ↓ ↓ ↓ │
│ ┌────────┐ ┌────────┐ ┌────────┐ │
│ │ OS thr │ │ OS thr │ │ OS thr │ │
│ └────────┘ └────────┘ └────────┘ │
│ work-stealing scheduler │
└────────────────────────────────────────┘
```
**How it works**: each query spawns a fiber. The fiber runs synchronously from its own perspective — blocking calls (network read, disk write) are transparently converted to yields by the runtime. The scheduler multiplexes fibers across OS threads.
**Pros**: synchronous programming model (no `async`/`.await` annotations); lightweight (2–8 KB per fiber vs. the state-machine size of a Rust future); Go and Erlang prove this works at massive scale; preemption possible (Go preempts goroutines at function calls since Go 1.14).
**Cons**: stack allocation per fiber (minor); GC pressure if using Go (major for database latency); less ecosystem in Rust (no production-quality M:N fiber runtime); context switch is ~10–100 ns vs. ~1 ns for a Rust future poll.
**Used by**: CockroachDB (Go), Dgraph (Go), Erlang/OTP databases, may_minihttp (Rust).
#### 4. Deterministic io_uring Loop + Thread Pool (TigerBeetle)
A single main thread drives an io_uring instance for all I/O. CPU-bound work (query execution) is dispatched to a thread pool. Results return to the main loop via the io_uring completion queue.
```
┌────────────────────────────────────────────┐
│ Main thread (deterministic) │
│ │
│ io_uring ring: │
│ - ACCEPT (new connections) │
│ - RECV (query frames) │
│ - WRITE (WAL append) │
│ - FSYNC (WAL durable) │
│ - SEND (results / deltas) │
│ │
│ On CQE(RECV, query_bytes): │
│ dispatch to thread pool │
│ │
│ On CQE(thread_pool_result): │
│ submit SEND(client_fd, result_bytes) │
└────────────────────────────────────────────┘
│ ▲
▼ │
┌────────────────────┐ ┌─────────┐
│ Worker threads │ │ eventfd │
│ (query exec) │──→│ signal │
│ │ │ back │
└────────────────────┘ └─────────┘
```
**How it works**: all I/O flows through one io_uring instance on the main thread. When a query arrives (CQE for RECV), the main thread dispatches it to a worker. The worker executes the query (CPU-bound, touching in-memory data structures), then signals completion back to the main loop via eventfd. The main loop submits the SEND SQE to return the result.
**Pros**: fully deterministic — the main loop processes events in a fixed order, which enables replay-based testing and debugging; io_uring handles both disk and network in one scheduler; minimal coordination (workers are fire-and-forget); easiest to reason about.
**Cons**: single main thread is a potential bottleneck for very high connection rates; worker pool dispatch adds one context switch per query; all state mutation in the main loop limits write throughput to one core.
**Used by**: TigerBeetle (Zig), LMAX Disruptor (Java, similar philosophy).
### Decision: Which Model for `.wo`
> **Canonical model: single-threaded event loop.** The rest of this section catalogues a hybrid multi-threaded architecture that was the original proposal; it now lives here as the **scale-out reference** for when one core isn't enough. The Phase 2 spec ([02-wo-language.md § Concurrency Model](./database/02-wo-language.md#concurrency-model)) pins the initial runtime to a single userland thread driving io_uring directly (option 4 below, deterministic io_uring loop), without a worker pool, lock-free shared state, or a dedicated WAL-writer thread. Group commit still happens — the loop drains pending commits into one fsync SQE per tick. Scale past one core by **sharding** independent engine processes, not by reintroducing the multi-threaded hybrid here.
The multi-threaded architecture below remains useful as (a) a reference for the sharded model's internal coordination when cross-shard 2PC is added, and (b) the fallback if a workload profile ever justifies abandoning the single-threaded invariant. Keep reading if you want the full trade-off map; skip to the next section if you only care about what ships.
The architecture maps to the database subsystems (multi-threaded variant — not the current target):
| Subsystem | Best fit | Why |
| --- | --- | --- |
| **I/O scheduling** (accept, recv, send, WAL, fsync) | io_uring on a dedicated I/O thread | One submission queue; batched syscalls; handles disk + network |
| **Query execution** (plan, traverse, filter) | Worker thread pool | CPU-bound; parallelizable per query; no I/O in the hot path (data is in RAM) |
| **WAL writer** | Single dedicated thread | Ordered writes; group commit batching; sequential fsync |
| **Subscription matching** | Worker thread pool (same as query execution) | CPU-bound predicate evaluation; parallelizable per commit |
| **Client push** | io_uring SEND from I/O thread | I/O-bound; batched with other sends |
| **Checkpoint / vacuum** | Background threads | Long-running, low-priority, can yield to production traffic |
| **Transaction coordination** | Lock-free shared state (AtomicU64 for LSN, CAS for version chains) | Must be accessible from any worker |
This multi-threaded variant is a **hybrid** — closest to **option 4 (TigerBeetle) extended with a worker pool and lock-free shared state**. The actual Phase 2 design stops at plain option 4 (main-thread-only) — no worker pool, no cross-thread lock-free structures:
```
┌───────────────────────────────────────────────────────────┐
│ .wo Runtime │
│ │
│ ┌───────────────────────────────────────────────────┐ │
│ │ I/O Thread │ │
│ │ io_uring ring: │ │
│ │ ACCEPT → new session │ │
│ │ RECV → dispatch query to worker pool │ │
│ │ SEND → results / subscription deltas │ │
│ │ WRITE → WAL (via WAL writer thread) │ │
│ │ FSYNC → WAL barrier │ │
│ │ TIMEOUT → keepalive / session expiry │ │
│ └───────────────────────────────────────────────────┘ │
│ │ dispatch ▲ results │
│ ▼ │ │
│ ┌───────────────────────────────────────────────────┐ │
│ │ Worker Pool (N = num_cores - 2) │ │
│ │ │ │
│ │ worker 0: parse → plan → execute → match subs │ │
│ │ worker 1: parse → plan → execute → match subs │ │
│ │ worker 2: parse → plan → execute → match subs │ │
│ │ ... │ │
│ │ │ │
│ │ Shared (lock-free): │ │
│ │ - B+ tree (optimistic lock coupling) │ │
│ │ - LSM memtable (crossbeam-skiplist) │ │
│ │ - Graph adjacency (dashmap) │ │
│ │ - MVCC version chains (atomic CAS) │ │
│ │ - Subscription registry (sharded RwLock) │ │
│ └───────────────────────────────────────────────────┘ │
│ │ WAL records ▲ fsync ack │
│ ▼ │ │
│ ┌───────────────────────────────────────────────────┐ │
│ │ WAL Writer Thread │ │
│ │ │ │
│ │ collect WAL records from worker batch queue │ │
│ │ serialize → io_uring WRITE + FSYNC (linked SQE) │ │
│ │ on CQE: signal workers that commit is durable │ │
│ │ │ │
│ │ Group commit: batch N commits into one fsync │ │
│ └───────────────────────────────────────────────────┘ │
│ │
│ ┌────────────────────┐ ┌────────────────────┐ │
│ │ Checkpoint thread │ │ Vacuum thread │ │
│ │ (background) │ │ (background) │ │
│ └────────────────────┘ └────────────────────┘ │
└───────────────────────────────────────────────────────────┘
```
### Thread Roles — Fixed, Not Dynamic
Each thread has a single role for its lifetime. No work-stealing across roles. This gives predictability and avoids cache-pollution:
| Thread | Count | Role | I/O model |
| --- | --- | --- | --- |
| **I/O thread** | 1 | Accept, recv, send via io_uring; dispatch queries; push subscription deltas | io_uring SQ/CQ poll |
| **Worker threads** | `num_cores - 3` | Parse, plan, execute queries; match subscriptions; stage MVCC mutations | Pure CPU; no I/O; no syscalls |
| **WAL writer** | 1 | Collect committed WAL records; batch; io_uring write + fsync; signal durability | Dedicated io_uring ring for WAL fd |
| **Checkpoint** | 1 | Periodic arena snapshot to SSD | io_uring or plain `pwrite` |
| **Vacuum** | 1 | Prune MVCC version chains; reclaim memtable space | CPU-bound scan |
Total: `num_cores` threads, one per core, pinned via `pthread_setaffinity_np`. No oversubscription, no context switches between roles.
### Query Lifecycle Through the Runtime
A checkout query flows through the runtime like this:
```
1. Client sends EXECUTE frame over TCP
→ io_uring CQE(RECV) on I/O thread
2. I/O thread deserializes frame, identifies session + prepared plan
→ pushes (session, plan, params) onto worker dispatch queue
3. Worker thread picks up the query
→ Acquires MVCC snapshot (AtomicU64 read — nanoseconds)
→ Executes plan against in-memory structures:
UPDATE products SET inventory.on_hand = ... WHERE id = 42
INSERT INTO orders ...
MATCH ... CREATE ...
→ Stages mutations in a local write-set (no global visibility yet)
→ Evaluates constraints
→ Pushes WAL record to WAL writer’s batch queue
4. WAL writer collects this record + records from other workers
→ Serializes batch
→ Submits io_uring WRITE + linked FSYNC to NVMe
→ On CQE(FSYNC): marks all records in the batch as durable
→ Signals each worker’s commit-complete channel
5. Worker receives durability signal
→ Publishes mutations to global MVCC (atomic pointer swaps)
→ Evaluates subscription registry against the delta
→ Pushes DELTA frames to the I/O thread’s send queue
6. I/O thread submits io_uring SEND for each DELTA + the RESULT frame
→ Client receives commit acknowledgment
→ Subscribers receive live inventory update
Total wall time: ~100–200 μs (dominated by step 4: NVMe fsync)
```
### Why Not Pure Async (tokio)?
Tokio is production-proven and SurrealDB uses it. But for this database, the hybrid is better:
| Concern | tokio | Hybrid (this design) |
| --- | --- | --- |
| **io_uring integration** | `tokio-uring` exists but is experimental; tokio’s core is epoll-based | io_uring is the primary I/O model; no epoll fallback |
| **Determinism** | Work-stealing introduces non-deterministic scheduling | Fixed thread roles; deterministic dispatch; replay-testable |
| **GC / allocator pressure** | Futures allocate on the heap; many small allocations per query | Workers use arena allocators; pre-allocated per-query scratch space |
| **Cache locality** | Tasks migrate between cores via work-stealing | Threads are pinned; data stays cache-hot per core |
| **Debugging** | Async stack traces are notoriously hard to read | Each thread has a clear role; stack traces are synchronous |
| **Dependency weight** | tokio + tower + hyper + … | libc + io_uring syscalls |
The trade-off: less library reuse, more manual plumbing. For a database where every microsecond on the commit path matters, that trade-off is correct.
### Why Not Pure Fibers (Go)?
Go’s goroutine scheduler is excellent and CockroachDB proves databases can be built on it. But:
| Concern | Go goroutines | Hybrid (this design) |
| --- | --- | --- |
| **GC pauses** | Stop-the-world pauses during checkout fsync batches lose money | No GC — Rust or C++ with arena allocators |
| **io_uring** | Go has no native io_uring support; falls back to epoll + thread pool for disk I/O | io_uring is first-class |
| **Memory control** | Cannot pin arenas, control huge pages, or use `mlockall` idiomatically | Full control via `libc` bindings |
| **Lock-free structures** | Possible but `sync/atomic` is more limited than Rust’s `crossbeam` | `crossbeam`, `dashmap`, `arc-swap` — mature ecosystem |
If the database were written in Go, goroutines would be the right model. Since the [Phase 2 language analysis](./database/02-wo-language.md) argues for Rust or C++, fibers are not the natural fit.
### Synchronization Between Threads
Workers touch shared data structures. The synchronization budget is strict — any contention on the commit path adds latency to every checkout:
| Shared structure | Accessed by | Primitive | Contention |
| --- | --- | --- | --- |
| **LSN counter** | All workers + WAL writer | `AtomicU64::fetch_add` | One atomic increment per commit — cheapest possible |
| **MVCC version chains** | All workers (read + write) | Atomic CAS to prepend | Per-record; independent records don’t contend |
| **B+ tree internal nodes** | All workers (read) | Optimistic lock coupling (version + retry) | Read-dominant; writes hold latches for microseconds |
| **Worker dispatch queue** | I/O thread (push) + workers (pop) | Lock-free MPSC queue (`crossbeam-channel`) | One producer, N consumers — no contention |
| **WAL batch queue** | Workers (push) + WAL writer (drain) | Lock-free MPMC queue | Drained in bulk every fsync batch (~100 μs) |
| **Subscription registry** | Workers (match) + I/O thread (register/unregister) | Sharded `RwLock` | Readers (match on commit) never block each other |
| **Send queue** | Workers (push DELTA) + I/O thread (drain to io_uring) | Lock-free MPSC per connection | One consumer per connection fd |
No mutex on the commit path. The only serialization point is the LSN counter — a single `fetch_add`.
### Handling Compute-Heavy Queries
Graph traversals and analytical queries can consume milliseconds of CPU. In a fiber model, a long-running fiber starves others (cooperative scheduling). In this hybrid:
- Workers are pre-emptible at the OS level (kernel threads, not fibers).
- Each worker runs one query at a time to completion. If a query takes 5 ms, that worker is busy for 5 ms — but the other N-1 workers continue serving other queries.
- If the pool is saturated, the I/O thread applies back-pressure: it stops reading from client sockets (io_uring RECV is not re-submitted), TCP flow control kicks in, and clients experience latency — which is the correct behavior under overload.
- Optional: a per-query CPU budget (checked at loop iteration points in graph traversal) that yields the worker back to the pool and resumes the query later. This is partial preemption — fiber-like semantics within a thread pool.
### Lifecycle of a Subscription
Subscriptions are long-lived — they span many commits. The runtime handles them without dedicated threads or fibers:
```
1. Client sends SUBSCRIBE frame
→ I/O thread registers (predicate, client_fd) in subscription registry
2. A commit happens on a worker thread
→ Worker evaluates subscription registry against the commit’s delta
→ For each match: serialize DELTA frame, push to that client’s send queue
3. I/O thread drains send queues
→ Submits io_uring SEND for each queued DELTA
4. Client disconnects (CQE reports EPOLLHUP-equivalent)
→ I/O thread removes all subscriptions for that fd
```
No thread or fiber per subscription. No polling. The cost of a subscription is: one entry in the registry (a few hundred bytes) + O(1) evaluation per commit (if keyed) or O(subs) per commit (if arbitrary predicate). Thousands of active subscriptions add microseconds to each commit, not threads.
### Comparison With Real Database Runtimes
| Database | Language | Runtime model | I/O | Query scheduling |
| --- | --- | --- | --- | --- |
| **Postgres** | C | Process-per-connection | epoll + blocking I/O on worker processes | One process per query |
| **MySQL** | C++ | Thread-per-connection or thread pool | epoll | One thread per query |
| **SurrealDB** | Rust | tokio (M:N async) | epoll (tokio) + Rayon for CPU | async tasks + Rayon parallel iterators |
| **ScyllaDB** | C++ | Thread-per-core (Seastar) | io_uring / epoll / SPDK | Futures on per-core reactor |
| **TigerBeetle** | Zig | Single-threaded io_uring loop | io_uring | Deterministic, single-threaded |
| **CockroachDB** | Go | M:N goroutines | epoll (Go netpoller) | Goroutine per query |
| **DuckDB** | C++ | Thread pool | Blocking I/O | Morsel-driven parallelism |
| **`.wo`** (this design) | Rust/C++ | **Hybrid: io_uring I/O thread + pinned worker pool + dedicated WAL thread** | io_uring | Worker per query, lock-free shared state |
### What the Runtime Does NOT Do
Keeping the scope honest:
- **No work-stealing**. Workers are pinned, queries are assigned round-robin or by shard affinity. Work-stealing adds scheduling complexity for marginal throughput gains when the data is in RAM and queries are sub-millisecond.
- **No async/await in the engine codebase**. Workers run synchronous code against in-memory structures. The only async code is the I/O thread’s io_uring event loop.
- **No fiber stacks**. No `swapcontext`, no stack allocation per query. Each worker has one OS stack, runs one query at a time.
- **No M:N scheduling**. N queries map to N worker invocations, but each invocation is a plain function call on a fixed thread — not a scheduled task on a shared executor.
The runtime is deliberately simpler than tokio, Go’s scheduler, or Seastar. The bet is that with all data in RAM, query execution is fast enough that a fixed-size thread pool with lock-free shared state is sufficient — and far easier to debug, profile, and reason about.
### Reference Implementations
- **Seastar** — <https://github.com/scylladb/seastar>. The thread-per-core framework. Read `seastar/core/reactor.cc` for the io_uring event loop and `seastar/core/smp.cc` for inter-core messaging.
- **Glommio** — <https://github.com/DataDog/glommio>. Rust thread-per-core runtime built on io_uring. Closest Rust analogue to Seastar.
- **tokio** — <https://github.com/tokio-rs/tokio>. Work-stealing async runtime. Read `tokio/src/runtime/scheduler/` for the work-stealing logic.
- **TigerBeetle** — <https://github.com/tigerbeetle/tigerbeetle>. Deterministic single-threaded io_uring loop. Read `src/io.zig` for the I/O ring and `src/state_machine.zig` for deterministic processing.
- **crossbeam** — <https://github.com/crossbeam-rs/crossbeam>. Lock-free data structures for Rust: channels, skiplist, deque, epoch-based GC.
- **LMAX Disruptor** — <https://github.com/LMAX-Exchange/disruptor>. Lock-free ring buffer for inter-thread communication. The intellectual ancestor of the WAL batch queue.
- **Readings**:
- *The Seastar Tutorial* — thread-per-core explained from first principles.
- *Fibers under the magnifying glass* (Vyukov) — M:N scheduling in practice.
- *io_uring and networking in 2023* (Axboe) — io_uring for both storage and networking.
- *LMAX Architecture* (Fowler) — single-writer, mechanical sympathy, lock-free coordination.
### Where This Fits in the Database Series
This runtime design is the missing piece between [Phase 2 (ACID engine)](./database/02-wo-language.md) and [Phase 3 (in-memory storage)](./database/03-inmemory-engine.md). Phase 2 described *what* the engine does (ACID transactions across three paradigms). Phase 3 described *where* data lives (RAM, with io_uring WAL). This section describes *how* work is scheduled — the thread architecture that connects client I/O, query execution, WAL durability, and subscription push into a single coherent runtime.

249
docs/runtime/wo-language.md Normal file
View file

@ -0,0 +1,249 @@
# writeonce — the `.wo` Language and Runtime
> A declarative programming language with database and subscription-native HTTP in its standard runtime. Like `go run`, you write `.wo` files and execute them — but your program is a full-stack application.
---
## What writeonce is
`writeonce` is a programming language, a standard runtime, and a toolchain. Three layers of one product:
1. **The language** — `.wo` source files. Declarative by default (`type`, `service`, `policy`, `on <event>`) with a hybrid SQL+Cypher query sublanguage for the imperative parts. Types, queries, transactions, subscriptions, policies, triggers, HTTP endpoints, and UI screens are all first-class language constructs.
2. **The runtime** — an ACID multi-paradigm database (relational + document + graph), an HTTP server, a subscription engine, and a scheduler. All of it links into a single binary with your program. No external Postgres, no external Redis, no separate Node process.
3. **The toolchain** — the `wo` command: `wo run`, `wo build`, `wo test`, `wo fmt`, `wo mod`, `wo gen`. Modelled directly on the Go toolchain. One binary per project; no runtime to install on the target host.
The one-line pitch: **Go + Postgres + `net/http` + Phoenix LiveView, folded into one language and one binary.**
## Hello, world
> Full example projects:
> - [`docs/examples/blog/`](../examples/blog/) — a blog (~200 lines): articles, authors, tags, comments, live subscriptions, row-level policies, typed Go client.
> - [`docs/examples/ecommerce/`](../examples/ecommerce/) — an e-commerce store (~300 lines): cross-paradigm ACID checkout, link types with properties, tagged unions, a **live order-ops table** that delta-updates in place.
A complete `.wo` program that creates a database table, exposes six REST endpoints with live subscriptions, and emits a typed Go client:
```wo
-- article.wo
type Article {
id: Id
title: Text
body: Markdown
author: Text
created_at: Timestamp = now()
service rest "/api/articles"
expose list, get, create, update, delete, subscribe
}
```
Run it:
```bash
$ wo run
[wo] compiling ./article.wo
[wo] schema: 1 type, 0 migrations needed
[wo] listening on :8080
GET /api/articles list
GET /api/articles/:id get
POST /api/articles create
PATCH /api/articles/:id update
DELETE /api/articles/:id delete
WS /api/articles/live subscribe
```
Use it:
```bash
$ curl -X POST localhost:8080/api/articles \
-H "Content-Type: application/json" \
-d '{"title":"Hello","body":"# First post","author":"me"}'
{"id":1,"title":"Hello","body":"# First post","author":"me","created_at":"2026-04-17T..."}
$ curl localhost:8080/api/articles
[{"id":1,"title":"Hello","...":"..."}]
```
Generate a typed client:
```bash
$ wo gen sdk --lang go --out ./client
# produces ./client/sdk.go with typed Article struct and Subscribe helper
```
Subscribe from the client — deltas push on every commit, no polling:
```go
import "myapp.example.com/client"
c, _ := client.Connect("wo://localhost:8080")
sub, _ := c.Articles.Subscribe(ctx, client.Where{Author: "me"})
for delta := range sub.C {
fmt.Printf("%s: %+v\n", delta.Kind, delta.Row)
}
```
Three files, five commands, zero infrastructure. Compare the same thing in Go+Postgres+React: one SQL schema, one migration tool, one ORM, one HTTP router, one subscription layer (polling or Redis pub-sub), one hand-written client, one React hook — roughly 2000 lines before you write any business logic.
## The toolchain
Go-literal. Every command maps to a Go equivalent so the mental model transfers:
| Command | Go equivalent | Purpose |
| --- | --- | --- |
| `wo init <name>` | `go mod init` | scaffold a new project |
| `wo run` | `go run ./...` | compile and execute |
| `wo build` | `go build` | emit a static binary |
| `wo test` | `go test` | run `.wo` tests |
| `wo fmt` | `gofmt` | canonical formatter |
| `wo vet` | `go vet` | lint + type-check without running |
| `wo mod <cmd>` | `go mod` | dependencies |
| `wo doc <sym>` | `go doc` | render docs for a type |
| `wo gen sdk --lang <L>` | `go generate` (codegen) | emit a client SDK |
| `wo migrate [--plan\|--apply]` | no direct equivalent | schema evolution |
| `wo dev` | no direct equivalent | hot-reload dev server |
**`wo run` vs `wo build`.** Same as Go: `wo run` compiles to a temp binary and executes it; `wo build` writes a named binary. No interpreter mode — `.wo` is compiled, always.
**`wo dev` is the one non-Go addition.** Edit a `.wo` file, the runtime hot-swaps the affected module without restarting. Live subscriptions survive the reload. This is the Phoenix LiveView influence.
## Program structure
```
myapp/
├── wo.toml # like go.mod — name, version, dependencies
├── main.wo # optional entry point
├── types/ # `type` declarations (one file per domain concept)
│ ├── article.wo
│ └── user.wo
├── ui/ # ##ui screens (optional)
├── tests/ # *_test.wo files
└── wo.lock # locked dependency graph (like go.sum)
```
Minimum project is one `.wo` file with one `type` declaration. The compiler generates:
- the database schema (relational row, document structures, graph edges) from the type's fields
- HTTP handlers from type-attached `service` blocks
- transactional triggers from `on <event>` blocks
- row-level policies from `policy` blocks
- typed client SDKs from the same type, on demand
No `main()` is required for a pure type-and-service app. The runtime starts the HTTP server, loads the database, and dispatches. If you need procedural entry logic (CLI args, graceful shutdown hooks, cron jobs), add `main.wo` with a `main { ... }` block.
## The runtime — what's in the standard library
Every `wo build` links these in. They're not external packages you import — they're the language.
| Component | Responsibility | Mapped to phase |
| --- | --- | --- |
| **Database** | In-RAM ACID multi-paradigm (relational + doc + graph) with WAL durability | [Phase 2](./database/02-wo-language.md) + [Phase 3](./database/03-inmemory-engine.md) |
| **Transaction coordinator** | MVCC, snapshot isolation, cross-paradigm `RETURNING` alias table | [Phase 2](./database/02-wo-language.md) |
| **HTTP server** | REST + GraphQL dispatch generated from `service` blocks | [Phase 4](./database/04-client-api.md) + `crates/http` |
| **Subscription engine** | `LIVE` queries push deltas on commit, zero polling | [Phase 4](./database/04-client-api.md) |
| **Wire protocol** | Native binary codec for typed clients | [Phase 4](./database/04-client-api.md) |
| **Codegen** | `wo gen sdk` — Go, TypeScript, Rust, Python clients from `type` declarations | [Phase 5](./database/05-go-sdk.md) |
| **UI renderer** | `##ui` screens → SSR HTML + client runtime | [Phase 6](./database/06-lowcode-fullstack.md) |
| **Authorization** | `policy` blocks compiled into planner rewrite rules | [Phase 6](./database/06-lowcode-fullstack.md) |
| **Scheduler** | Single-threaded event loop over io_uring; one core per process (shard to scale) | [async.md](./async.md) + [Phase 2 concurrency](./database/02-wo-language.md#concurrency-model) |
Comparison to Go's stdlib:
| Need | Go | writeonce |
| --- | --- | --- |
| HTTP server | `net/http` | built-in `service rest` |
| Database | none (use `database/sql` + driver + Postgres) | **built-in** |
| Template rendering | `html/template` | `##ui` blocks |
| Concurrency | goroutines + channels | single-threaded event loop (Redis-style); shard to scale past one core |
| Testing | `testing` | `wo test` + `.wo` test syntax |
| Formatting | `gofmt` | `wo fmt` |
| Modules | `go.mod` + `go.sum` | `wo.toml` + `wo.lock` |
## Clients — who consumes your program
The same `.wo` type declarations that define the database also define the wire format. `wo gen sdk` emits:
- **Go** — typed structs, `*Client`, `TypedSubscription[T]` generics over a channel
- **TypeScript / browser** — types + `fetch` + WebSocket subscriptions
- **Rust** — structs, `tokio` async client, `impl Stream<Item = Delta>`
- **Python** — dataclasses, `async for delta in sub`
- **curl / raw REST** — documented via auto-generated OpenAPI spec at `/openapi.json`
- **GraphQL clients** — SDL auto-generated at `/graphql/schema.graphql`
**Raw `.wo` DML is a first-class escape hatch.** Every client SDK exposes a single method — `client.Wo(ctx, src, params)` in Go, equivalents in TypeScript/Rust/Python — that accepts any `.wo` source the server would accept: mixed SQL + Cypher, `BEGIN … COMMIT` blocks with `RETURNING` aliases threading across statements, ad-hoc MATCH-then-SELECT queries that cross multiple generated types. The typed methods are sugar; the engine speaks `.wo` on the wire. A Go program can send a cross-paradigm transaction as a single string and the server parses + executes it exactly like `wo run` would — see [Phase 5: Go Client SDK](./database/05-go-sdk.md) for the full API.
One schema, every protocol. A browser app, a mobile client, and a background worker can all subscribe to the same live query and receive the same delta stream.
## What this is, and isn't
**Is.** A declarative, full-stack, single-binary language for building CRUD apps with live data. A replacement for the "Go backend + Postgres + Redis + React + Prisma + GraphQL server" stack.
**Isn't.**
- Not a general-purpose language like Rust or Go. You can't write a kernel module or a video codec in `.wo`. The scope is data-shaped applications.
- Not a JavaScript meta-framework. No Node, no React. The UI layer (`##ui`) is declarative and compiles to SSR HTML with a small vanilla-JS client.
- Not a DSL that transpiles to another language. `.wo` has its own lexer, parser, analyzer, and bytecode. The [`prototypes/`db`/`](../../prototypes/`db`/) C++ prototype and the planned Rust crates implement the runtime natively.
- Not a hosted service. Your binary owns its own DB file. No managed cloud offering is required.
## How this maps to the design series
This overview is the user-facing frame. The underlying engineering plan is the 7-phase series linked from [database.md](./database.md):
- **[Phase 2](./database/02-wo-language.md)** designs the language and the transaction coordinator.
- **[Phase 3](./database/03-inmemory-engine.md)** builds the storage engine.
- **[Phase 4](./database/04-client-api.md)** builds the wire protocol and subscription engine.
- **[Phase 5](./database/05-go-sdk.md)** builds the first typed client (Go) and `wo gen`.
- **[Phase 6](./database/06-lowcode-fullstack.md)** adds `##ui` and the application-level blocks.
- **[Phase 7](./database/07-wo-seg-migration.md)** migrates writeonce-the-blog from `wo-seg` onto this runtime.
Phase 1 (evaluation) and the case studies in [surreal-case-study.md](./surreal-case-study.md) argue *why* the language exists at all. Read those first if you're skeptical; read the phase docs if you're implementing; read this page if you want to know what it feels like to use.
## Reference points
The design absorbs lessons from several systems. In order of influence:
- **Go** — toolchain shape, single-binary deployment, "the language is the build system"
- **Phoenix LiveView** — subscription-native UI, hot-reloading dev server
- **SAP CDS** — declarative entity/service language, admin UI generation
- **SurrealDB** — multi-paradigm query language, `LIVE` subscriptions over wire
- **PocketBase** — single-binary CRUD backend (the proof of concept that this is shippable)
- **EdgeDB** — unified type system above storage paradigms
- **Elixir / Erlang / OTP** — hot code loading, supervision, subscription semantics
- **Django** — admin UI as a built-in, not a bolt-on
None of these give you all of: a language, a database, a subscription engine, a UI toolkit, a client codegen, and a single-binary output. writeonce is the attempt to fuse the best of each into one thing.
## Minimal "hello, world" as a full program
If you want a pure procedural test, without the server:
```wo
-- hello.wo
main {
print("hello, world")
}
```
```bash
$ wo run hello.wo
hello, world
```
If you want the database without HTTP:
```wo
type Counter {
name: Text @unique
value: Int = 0
}
main {
insert Counter { name: "visits" };
update Counter{ name == "visits" }.value += 1;
let c = select Counter{ name == "visits" };
print(c.value);
}
```
If you want the full app — database, HTTP, subscriptions, clients — it's the article example at the top of this page.
Three progressive shapes, one language, one command to run each.

140
prototypes/wo-db/README.md Normal file
View file

@ -0,0 +1,140 @@
# wo-db — `.wo` language prototype (C++)
Phase 2 milestone 1 from [docs/runtime/database.md](../../docs/runtime/database.md):
**parser + analyzer + in-memory executor, single-user, no durability.**
Implements a cut-down `.wo` dialect spanning all three paradigms described in
[02-wo-language.md](../../docs/runtime/database/02-wo-language.md): relational,
document, and graph — all in one grammar, one process, one in-RAM store.
## Build & run
```bash
make # builds build/wo-db
make test # runs tests/smoke.wo + tests/checkout.wo
./build/wo-db # interactive REPL
./build/wo-db < file.wo # batch mode
```
Requires `g++` with C++20 (tested on 13.3). No external dependencies.
## Supported grammar
### Schema
```wo
##sql
#users
id int
name string
email string
##doc
#article_meta
id int
title string
##graph
#recommendations
(user)-[:PURCHASED]->(product)
```
Paradigms: `sql`, `doc`, `graph`. Types are parsed but not enforced (prototype).
Graph schema blocks are documentation only — nodes and edges are created at
runtime by `CREATE`.
### Queries
```wo
-- relational
INSERT INTO users (id, name) VALUES (1, 'Alice');
SELECT id, name FROM users WHERE id > 1 AND name != 'Carol';
UPDATE users SET email = 'a@b.com' WHERE id = 1;
DELETE FROM users WHERE id = 3;
-- document (same syntax as sql; dotted columns write nested objects)
INSERT INTO article_meta (id, title) VALUES (100, 'Intro');
SELECT title FROM article_meta;
-- graph
CREATE (u:user {id: 1, name: 'Alice'});
CREATE (u:user {id: 1})-[:PURCHASED {qty: 2}]->(p:product {id: 10});
MATCH (u:user {id: 1})-[:PURCHASED]->(p:product) RETURN u, p;
```
### Fixed-glue: parameters, RETURNING, transactions, LIVE
The five things from the two-layer design ([docs/runtime/database/02-wo-language.md](../../docs/runtime/database/02-wo-language.md)) that tie the three grammars together:
```wo
-- $name parameters work everywhere (SQL, Cypher, expressions)
INSERT INTO users (email) VALUES ('a@b.com') RETURNING id AS uid;
SELECT * FROM users WHERE id = $uid;
-- cross-paradigm RETURNING threads ids from SQL into Cypher
BEGIN SNAPSHOT;
INSERT INTO orders (user_id, status) VALUES ($uid, 'pending') RETURNING id AS oid;
CREATE (u:user {id: $uid})-[:PURCHASED {order_id: $oid}]->(p:product {id: $pid});
COMMIT;
-- SAVEPOINT / ROLLBACK TO are parsed but not yet enforced in the prototype
BEGIN;
SAVEPOINT s1;
INSERT INTO users (email) VALUES ('typo@example.com');
ROLLBACK TO s1;
COMMIT;
-- LIVE prefix reserved — inner query runs now, subscription activation in Phase 3
LIVE SELECT id, email FROM users;
LIVE MATCH (u:user)-[:PURCHASED]->(p:product) RETURN u, p;
```
Auto-populated `id` columns: when a table declares an `id int` column and `INSERT` omits it, the engine mints a fresh id from a per-table counter and fills it in. Combined with `RETURNING id AS alias`, this is how ids thread from SQL into Cypher inside a transaction without `LAST_INSERT_ID()`.
### Expressions
Literals (`int`, `string`, `true`/`false`/`null`, `[array]`, `{object}`),
dotted paths (`meta.title`), comparisons (`= != < <= > >=`), boolean
(`AND OR NOT`).
### REPL meta-commands
```
.tables list sql/doc tables
.schema dump schema
.exit
```
## What's intentionally missing
This is a Phase 2 *prototype*, not a database. It does not implement:
- **real** transactions — `BEGIN/COMMIT/SAVEPOINT/ROLLBACK` parse and are acknowledged, but no atomic rollback, no MVCC, no WAL
- **real** `LIVE` subscriptions — the inner query runs; no delta frames, no push (Phase 3/4)
- durability, crash recovery (Phase 3)
- real document operations (array push/splice, deep path updates beyond simple dotted SET)
- joins, aggregations, subqueries
- the LSM document engine, B+ tree relational pages, native graph adjacency (Phase 3)
- wire protocol, auth (Phase 4+)
See [docs/runtime/database.md](../../docs/runtime/database.md) for the full phase plan
and [docs/runtime/database/02-wo-language.md](../../docs/runtime/database/02-wo-language.md)
for the target language spec.
## Source layout
```
src/
value.hpp tagged Value (null/bool/int/string/array/object)
lexer.{hpp,cpp} tokenizer
ast.hpp AST node types
parser.{hpp,cpp} recursive-descent parser
storage.{hpp,cpp} in-memory Database (sql tables + doc tables + graph)
executor.{hpp,cpp} walks AST, returns ResultSet
main.cpp REPL + batch driver
tests/
smoke.wo core statement coverage
checkout.wo cross-paradigm RETURNING + BEGIN/SAVEPOINT/COMMIT + LIVE
```
Roughly 1.1k lines of C++ in the `wo` namespace.

35
reference/README.md Normal file
View file

@ -0,0 +1,35 @@
# `reference/` — archived for reference
Source material preserved alongside the active codebase. Nothing here is on the build path of the root workspace.
## `reference/crates/` — v1 writeonce blog
The original writeonce blog engine: 13 Rust crates implementing a flat-file content store, three sidecar indexes, a minimal HTTP server, and an `.htmlx` template layer.
| Crate | Responsibility |
| --- | --- |
| `wo-model` | `Article` domain type + content loader |
| `wo-seg` | Append-only `.seg` binary file format (~475 LOC) |
| `wo-index` | `title.idx`, `date.idx`, `tags.idx` sidecar indexes |
| `wo-store` | Composes seg + indexes, exposes the query API |
| `wo-watch` | inotify-driven content ingest |
| `wo-event` | Domain event type |
| `wo-sub` | Subscription layer (pre-engine-native `LIVE`) |
| `wo-rt` | v1 runtime glue |
| `wo-http`, `wo-route`, `wo-serve` | HTTP server |
| `wo-htmlx` | `.htmlx` template engine |
| `wo-md` | Markdown rendering |
It is a nested workspace. Build or test it standalone:
```bash
cd reference/crates
cargo build
cargo test
```
See [docs/runtime/database/07-wo-seg-migration.md](../docs/runtime/database/07-wo-seg-migration.md) for the plan to replace these crates with the new `.wo` runtime at `crates/wo-rt/` via phased coexistence. The v1 codebase is the source that migration reads from.
## `reference/writeonce-api/` and `reference/writeonce-app/`
Earlier exploration snapshots that predate the v1 crates. Self-contained; no active build wiring.

92
reference/rest/README.md Normal file
View file

@ -0,0 +1,92 @@
# `reference/rest/` — HTTP test files for the `.wo` runtime
`.rest` (or `.http`) is the plain-text HTTP-request format supported by the two main editor HTTP clients:
- **VS Code** — install [REST Client](https://marketplace.visualstudio.com/items?itemName=humao.rest-client) (`humao.rest-client`) and click "Send Request" above any block.
- **JetBrains IDEs** (IntelliJ, WebStorm, RustRover, Goland) — built-in HTTP Client recognises `.rest` and `.http` natively.
Both clients understand:
- `### ...` block separators
- `@var = value` document-level variables referenced as `{{var}}`
- `# @name foo` on a request, whose response fields are later reachable as `{{foo.response.body.id}}` — useful for threading auto-generated ids from `create` responses into later `get`/`patch`/`delete` calls
## Files
| File | Against | What it exercises |
| --- | --- | --- |
| [`blog.rest`](./blog.rest) | `docs/examples/blog/` | Full CRUD on Article/Comment, read-only on Author/Tag (matches the sample's `expose` lists). End-to-end flow: create article → list → get by id → PATCH title → PATCH embedded doc → DELETE draft → verify final state. |
| [`ecommerce.rest`](./ecommerce.rest) | `docs/examples/ecommerce/` | What Stage 2 currently serves for the ecommerce sample: read-only Product/Order/Customer lists, 405s for non-exposed create endpoints, 501s for Stage 3 stubs. Documents the shape of Stage 3/4 endpoints (`fn checkout`, LIVE subscribe, `/me`) even though they're not wired yet. |
## Running
```bash
# one terminal — start the runtime
cargo run --bin wo -- run docs/examples/blog
# [wo] listening on http://127.0.0.1:8080
# another terminal — or just open the .rest file in VS Code/JetBrains and click
```
Override the port via `WO_LISTEN`:
```bash
WO_LISTEN=127.0.0.1:9000 cargo run --bin wo -- run docs/examples/blog
```
…and update the `@host` line at the top of the `.rest` file to match.
## Without an editor (just `curl`)
Each `.rest` block maps directly to `curl`. Some examples:
```bash
# Runtime info
curl http://127.0.0.1:8080/
# List
curl http://127.0.0.1:8080/api/articles
# Create — server assigns `id` automatically
curl -X POST http://127.0.0.1:8080/api/articles \
-H "Content-Type: application/json" \
-d '{
"slug": "hello-writeonce",
"title": "Hello, writeonce",
"author": 1,
"published": true,
"meta": { "excerpt": "first post", "body_md": "# hi" }
}'
# Get by id
curl http://127.0.0.1:8080/api/articles/1
# Partial update
curl -X PATCH http://127.0.0.1:8080/api/articles/1 \
-H "Content-Type: application/json" \
-d '{ "title": "Updated" }'
# Delete
curl -X DELETE http://127.0.0.1:8080/api/articles/1 # 204 on success
# Stage 3 stub
curl -i http://127.0.0.1:8080/api/articles/live # 501 Not Implemented
```
For a scripted smoke run against the blog sample, the top-to-bottom `curl` sequence that exactly mirrors `blog.rest` is in [`docs/examples/blog/README.md`](../../docs/examples/blog/README.md).
## Expected-status cheat sheet
Every block in the `.rest` files ends its description with the expected HTTP status. Quick legend:
| Status | Meaning in this prototype |
| --- | --- |
| `200` | OK — list / get / update succeeded |
| `201` | Created — new row, `id` in the response body |
| `204` | No Content — delete succeeded |
| `400` | Bad JSON body |
| `404` | No such row, OR no method at all is attached to the path (e.g. `/api/customers` when the `service rest` block doesn't `expose` any collection-root op) |
| `405` | Method not allowed — the path is registered for a *different* method (e.g. POST against `/api/products` when only `list` is exposed, so GET is attached but POST isn't) |
| `501` | Not Implemented — Stage 3+ feature (LIVE subscriptions, `/me`, transactional fns) |
A `405` is a *feature* of the sample — it confirms the `expose` list in the `.wo` file is being honoured. A `501` is a Stage marker — the runtime acknowledges the shape but hasn't wired the handler yet.