diff --git a/docs/examples/employee-list/README.md b/docs/examples/employee-list/README.md index e4511d0..1bdd3e8 100644 --- a/docs/examples/employee-list/README.md +++ b/docs/examples/employee-list/README.md @@ -2,8 +2,8 @@ > **Status: target workload — does not compile on today's toolchain.** > Written ahead of iterations -> [20 (cross-program tables)](../../stories/language-runtime-database/20-cross-program-tables.md) -> and [21 (keypair attach auth)](../../stories/language-runtime-database/21-keypair-attach-auth.md), +> [9 (cross-program tables)](../../stories/databasev2/09-cross-program-tables.md) +> and [10 (keypair attach auth)](../../stories/databasev2/10-keypair-attach-auth.md), > the way every acceptance sample here precedes its features. It also leans > on 9/9b (the [employee sample](../employee/) it attaches to must run > first). diff --git a/docs/plan/discarded.md b/docs/plan/discarded.md index edd5467..1b9e83a 100644 --- a/docs/plan/discarded.md +++ b/docs/plan/discarded.md @@ -72,7 +72,7 @@ goes to find what took each one's place. | --- | --- | | `09-concurrency-scaleout.md` | [`08-shard-actor-runtime.md`](../stories/language-runtime-database/08-shard-actor-runtime.md) + [`11-fibers.md`](../stories/language-runtime-database/11-fibers.md) — the arc, landed 2026-08-21 | | `10-storage-foundations.md`, `11-wal-and-recovery.md` | [`09-database-engine.md`](../stories/language-runtime-database/09-database-engine.md) (typed WAL + replay) and [`22-durability-throughput-scale.md`](../stories/language-runtime-database/22-durability-throughput-scale.md) (the measurements) | -| `12-engine-disk-cutover.md` | Nothing — RAM stays authoritative by doctrine (principle 7). The disk story is the WAL; reclamation is [`32-wal-checkpoint.md`](../stories/language-runtime-database/32-wal-checkpoint.md) | +| `12-engine-disk-cutover.md` | Nothing — RAM stays authoritative by doctrine (principle 7). The disk story is the WAL; reclamation is [`databasev2 3, WAL checkpoint`](../stories/databasev2/03-wal-checkpoint.md) | | `13-class-model-live-pricing.md` | [`09b-table-relations-query.md`](../stories/language-runtime-database/09b-table-relations-query.md) — `@table`, `ref`/`backlink`, the compiler-checked query surface | | `07-inotify-content-watcher.md` | [`07-logwatcher-proof.md`](../stories/language-runtime-database/07-logwatcher-proof.md) — the log-watcher sample polls via `fs.stat`; inotify was never surfaced as a builtin | | `08-sendfile-static-assets.md` | Nothing. `sendfile` is not exposed; static assets are served as `Text` through `net.write` | diff --git a/docs/plan/exploration/postgresql/00-postgresql.md b/docs/plan/exploration/postgresql/00-postgresql.md index f71e299..7d2780b 100644 --- a/docs/plan/exploration/postgresql/00-postgresql.md +++ b/docs/plan/exploration/postgresql/00-postgresql.md @@ -46,7 +46,7 @@ The implementation phases that lean on this material: - The Rust-era consumers (plans 10/11/12: storage foundations, WAL and recovery, disk cutover) were removed with that track 2026-08-18; their ideas shipped in `database/src/` (typed WAL + replay) and the rest wait - on [iteration 32](../../../stories/language-runtime-database/32-wal-checkpoint.md) + on [databasev2 3](../../../stories/databasev2/03-wal-checkpoint.md) (checkpoint) — the wal/buffer cards are its entry material. - Current consumers: [constraints-and-grammar](constraints-and-grammar.md) (the `@table` PK/FK grammar direction) and diff --git a/docs/stories/00-status.md b/docs/stories/00-status.md index f5809f9..8ca6f94 100644 --- a/docs/stories/00-status.md +++ b/docs/stories/00-status.md @@ -10,11 +10,14 @@ The single place to learn where this project stands. Organised in six buckets: folders** — a doc stays where it was authored, and only its frontmatter, its banner and this board change. -**Two tracks** (2026-08-26): [`language-runtime-database/`](language-runtime-database/00-story.md) -— the language, runtime and database — and [`porch/`](porch/00-story.md), the web -framework written in it. Each numbers its iterations from 1, so a porch 3 is not -a language 3; porch stories carry `track: porch` in frontmatter to keep queries -honest. Track folders are fine; **status** folders are not. +**Three tracks** (2026-08-26): +[`language-runtime-database/`](language-runtime-database/00-story.md) — the +language and runtime; [`porch/`](porch/00-story.md) — the web framework written +in it; [`databasev2/`](databasev2/00-story.md) — the database beyond RAM. Each +numbers its iterations from 1, so a porch 3 is not a language 3; every non-language +story carries `track:` in frontmatter, and moved ones keep +`was_language_iteration:` so a search for the old number still finds them. Track +folders are fine; **status** folders are not. **Status lives in frontmatter, nowhere else** (directive 2026-08-26). Every story iteration file sits flat in its track folder and @@ -348,19 +351,19 @@ that sequences its tasks. Read one, approve, then the next starts. | 22 | [Durability, throughput, scale](language-runtime-database/22-durability-throughput-scale.md) | ✅ **landed 2026-08-21** — db-bench + baseline.json (74 metrics) + restart/kill -9 proofs both shard counts; durable 4.5k vs ram 297k inserts/s, reads O(table), msgrate 13.4M/2.45M | | 31 | [Actor lifecycle](language-runtime-database/31-actor-lifecycle.md) | 🔄 **absorbed into 24** (directive 2026-08-23) and half landed there: `call` request/response with a typed scalar reply (`WO_B_CALL = 88`, WO-E226), bounded mailboxes (`WO_MAILBOX`, cap 1024, catchable `WO_T_ACTOR`), and actor death that traps callers instead of hanging them. Still open: `monitor` and `time.after` — ids **89 and 90 are reserved holes** in `wob.h`, which is the machine-checkable proof of what is left. Supervision trees stay out of v1 | | 24 | [chat: WebSocket workload](language-runtime-database/24-chat-websocket-workload.md) | 🔄 **the live slice** (absorbing 31 + 34, directive 2026-08-23) — branch `chat-ws-lifecycle`, 5/10 tasks landed: crypto, bounded mailboxes, WS upgrade, frame codec, `call`/reply + actor death. Pending: `monitor`, `time.after`, the chat sample, its gate, closeout. State lives in [the marker](../active-slice-2026-08-23-chat-ws-lifecycle.md) | -| 23 | [io_uring group-commit](language-runtime-database/23-io-uring-commit.md) | ⬜ fifth in chain, after stage 3 + 22 | -| 32 | [WAL checkpoint](language-runtime-database/32-wal-checkpoint.md) | ⬜ last in chain, after 23 — disk reclamation + bounded replay (story written 2026-08-21) | -| 33 | [Single-file store](language-runtime-database/33-single-file-db.md) | ⬜ off-chain, small — `WO_DATA=.db` file form; driver-only (story written 2026-08-22) | +| 23 | [io_uring group-commit](databasev2/04-io-uring-commit.md) | ⬜ fifth in chain, after stage 3 + 22 | +| 32 | [WAL checkpoint](databasev2/03-wal-checkpoint.md) | ⬜ last in chain, after 23 — disk reclamation + bounded replay (story written 2026-08-21) | +| 33 | [Single-file store](databasev2/07-single-file-db.md) | ⬜ off-chain, small — `WO_DATA=.db` file form; driver-only (story written 2026-08-22) | | 34 | [Crypto builtins](language-runtime-database/34-crypto-builtins.md) | 🔄 **code landed** as 24's T1 (`d14fa9f`): `sha1`/`sha256`/`hmac_sha256`, ids 85–87 in `wob.h`, `runtime/src/crypto.c`, RFC/FIPS vectors 18/0, corpus pin. The 24 gate that once needed it is cleared. Frontmatter keeps `status: refine` only until 24's T10 closeout sets it to `done` | | 38 | [Content platform capabilities](language-runtime-database/38-content-platform-capabilities.md) | ⬜ off-chain, needs a spec — the two capability families no iteration owns, confirmed against `runtime/src/wob.h`: `fs` mutation verbs (six fs builtins, ids 40–45; `append` creates-if-absent, so nothing is ever replaced, truncated, deleted or renamed) and `net.connect` (ids 51–55 + 91–95, no connect, and no `connect()` anywhere in `runtime/src/` — so no OIDC/SMTP/object-store/webhook/federation). Driven by a `docs/examples/vault` content-collaboration workload, in 28's mould. New builtins from 96 (89/90 reserved for 31); no `.wob` bump (`WOB_VERSION 6u`, last moved by 36). Story written 2026-08-26 from the "can it build a Nextcloud?" ask | | 39 | [Web framework parity](language-runtime-database/39-web-framework-parity.md) | ⬜ off-chain, needs a spec — from [the Fiber v3.5.0 study](../plan/exploration/fiber/00-fiber-parity.md) (all 32 of its middleware read against `porch`; **nine already have a counterpart**). Leads with a **random-bytes builtin**: the framework ledger claimed CSRF/sessions were unblocked by iteration 34's HMAC, but HMAC authenticates a token and cannot mint one — there is no RNG anywhere in the runtime. Then cookies (absent both ways; `Resp.headers` being a map cannot carry two `Set-Cookie` lines), then limiter/idempotency (cheapest wins — `@table` + `time.ticks`, nothing new), sessions, CSRF, and the routing/response sugar. Streaming/SSE/compression, `@derive` binding, TTL cache, `proxy` and metrics all excluded with owners named | | 37 | [wo-html components](language-runtime-database/37-wo-html-components.md) | ✅ off-chain — LANDED 2026-08-25. Raw text literal (backtick, margin stripped at lex time, `{{ }}` auto-escapes) + the component layer: `Component`/`render_all`/`Layout` in wo-html, `ok_html` moved into the framework, site and shop both migrated | | 35 | [net runtime seams](language-runtime-database/35-net-runtime-seams.md) | ⬜ off-chain — fd deadlines on the park plane, Unix sockets, peer address; owns the ledger's three 🔧 rows (story written 2026-08-22) | -| 20 | [Cross-program tables](language-runtime-database/20-cross-program-tables.md) | ⏸ hold (2026-08-21); channel done (branch ipc-attach keeps its manifest) | -| 21 | [Keypair attach auth](language-runtime-database/21-keypair-attach-auth.md) | ⏸ hold (2026-08-21); crypto+handshake done (branch keypair-auth keeps its manifest) | +| 20 | [Cross-program tables](databasev2/09-cross-program-tables.md) | ⏸ hold (2026-08-21); channel done (branch ipc-attach keeps its manifest) | +| 21 | [Keypair attach auth](databasev2/10-keypair-attach-auth.md) | ⏸ hold (2026-08-21); crypto+handshake done (branch keypair-auth keeps its manifest) | | 25 | [HTTP service layer](../superpowers/plans/2026-08-01-http-service-layer.md) | ⏸ hold (2026-08-21) — story file removed; the plan doc remains | | 26 | [Blue-green deploy](language-runtime-database/26-blue-green-deploy.md) | ⏸ hold (2026-08-21) | -| 27 | [Query grammar corpus](language-runtime-database/27-query-grammar-corpus.md) | ⏸ hold (2026-08-21) | +| 27 | [Query grammar corpus](databasev2/08-query-grammar-corpus.md) | ⏸ hold (2026-08-21) | | 28 | [skillhost host workload](language-runtime-database/28-skillhost-host-workload.md) | ⏸ hold (2026-08-21); gaps recorded (branch query-grammar found skillhost needs no new query grammar) | | 29 | [Compile-time metaprogramming](language-runtime-database/29-compile-time-metaprogramming.md) | ⏸ hold (2026-08-21) | | 15 | [deps: `wo.toml [deps]`](language-runtime-database/15-deps-package-manager.md) | ✅ **landed 2026-08-18** (branch web-framework): [deps] inline tables, git-binary fetch, wo.lock pinning, offline-when-locked, --update-deps, WO-E106/E107; `just deps-accept` 8/0 | @@ -591,7 +594,7 @@ precedence notes for resumption. 5. **23** — io_uring group-commit; the WAL's WRITE+FSYNC chains ride the arc's per-shard ring (T4); after 22's baseline — the payoff, measured. 6. **32** — WAL checkpoint - ([story](language-runtime-database/32-wal-checkpoint.md), + ([story](databasev2/03-wal-checkpoint.md), written 2026-08-21): the WAL is append-only forever — snapshot + truncate reclaims disk and bounds replay; after 23 (composes with group-commit), policy set by 22's aged-store numbers. @@ -599,6 +602,36 @@ precedence notes for resumption. **30** — observability, CI, fuzz: named 2026-08-20, still row-only (no story file); slots in when scheduled — nothing in the chain depends on it. +### ▸ databasev2 — the database beyond RAM + +New 2026-08-26. **The problem:** RAM is authoritative (principle 7) and nothing +declares a budget. Rows live in `malloc`'d slabs whose addresses are stable +forever; there is no eviction, spill or paging anywhere in `database/src/`; the +WAL never checkpoints so boot replays all history; and durability is one +process-global `WO_DATA`, so no table can say it matters more than another. An +allocation failure *is* a clean catchable `WO_T_OOM` — but swap thrash arrives +first and carries no error signal at all. + +**The lever** is per-table storage modes, which is why this track has a grammar +iteration. Six pending iterations moved here from the language track (their old +ids in the rows below); four are new. Done database work — 9, 9b, 22 — stays in +the language arc as v1 history. + +| # | Iteration | State | +| --- | --- | --- | +| 1 | [RAM ceiling: measure the breaking point](databasev2/01-ram-ceiling-measurement.md) | ⬜ **first, and startable today** — nobody here can say what happens at 90% RAM. Curve not cliff: swap onset, latency departure, the three exits (checked trap / swap thrash / OOM killer), and `kill -9` durability *at exhaustion*. Output is `perf-targets.md` + baseline rows, not prose | +| 2 | [`@table` storage modes](databasev2/02-table-storage-modes.md) | ⬜ **the language enrichment** — `mode: ram \| durable \| cold` per table, replacing the global switch. `durable` defaults so nothing changes silently; the compiler refuses a `durable` row holding a `ref` into a `ram` table. `.wob` format change. Grammar is small (`Ast.table_cfg` gains a key); semantics are the iteration | +| 3 | [WAL checkpoint](databasev2/03-wal-checkpoint.md) *(was 32)* | ⬜ snapshot + truncate: disk reclaimed, replay bounded | +| 4 | [io_uring group commit](databasev2/04-io-uring-commit.md) *(was 23)* | ⬜ close the 66× gap iteration 22 measured (durable 4.5k vs ram 297k inserts/s) | +| 5 | [Bounded tables and eviction](databasev2/05-bounded-tables-eviction.md) | ⬜ a declared capacity + refuse/evict/back-pressure, and a process-level pressure signal that sheds **before** the allocator or OS gets involved — turning the invisible failure into a managed one | +| 6 | [Cold tiering](databasev2/06-cold-tiering.md) | ⬜ the iteration that raises the ceiling, and the riskiest. Mostly forks: which shape, whether the index itself fits, whether the *language* surfaces the fault cost, and whether `@unique` on a cold table is refused outright. A paged B-tree stays rejected — if tiering needs one, reject tiering | +| 7 | [Single-file store](databasev2/07-single-file-db.md) *(was 33)* | ⬜ `WO_DATA=.db`; driver-only, independent | +| 8 | [Query grammar from corpora](databasev2/08-query-grammar-corpus.md) *(was 27)* | ⬜ whole-query `count`, `exists`; independent | +| 9 | [Cross-program tables](databasev2/09-cross-program-tables.md) *(was 20)* | ⏸ hold — attach to a running program's database over local IPC | +| 10 | [Keypair attach auth](databasev2/10-keypair-attach-auth.md) *(was 21)* | ⏸ hold — program identity as a keypair; needs 9 | + +--- + ### ▸ porch — the web framework track New 2026-08-26, from [the Fiber v3.5.0 parity study](../plan/exploration/fiber/00-fiber-parity.md). diff --git a/docs/stories/board-views.md b/docs/stories/board-views.md index 0a3369f..bda9e01 100644 --- a/docs/stories/board-views.md +++ b/docs/stories/board-views.md @@ -28,12 +28,13 @@ standup narrative; these queries are the live views over the same facts. Adjust the `FROM` path to your vault root (queries below assume the vault opens at the repo root). -Two tracks now carry iterations, each numbered from 1: -`language-runtime-database/` (the language, runtime and database) and `porch/` -(the web framework, added 2026-08-26). Iteration ids therefore repeat across -tracks — a porch 3 is not a language 3 — so every query below is scoped by -`FROM` path, and porch stories carry `track: porch` so a combined query can -still tell them apart. +Three tracks now carry iterations, each numbered from 1: +`language-runtime-database/` (the language and runtime), `porch/` (the web +framework) and `databasev2/` (the database beyond RAM) — the latter two added +2026-08-26. Iteration ids therefore repeat across tracks, so every query below is +scoped by `FROM` path; non-language stories carry `track:`, and iterations moved +between tracks keep `was_language_iteration:` so the old number stays +searchable. ## Everything not done, chain order first @@ -79,6 +80,15 @@ the place status is edited.** A status change is one edit to one `status:` key; a Kanban card drag that only rewrites the Kanban file is a lie the next query won't see. +## The databasev2 track + +```dataview +TABLE iteration, status, was_language_iteration AS "was" +FROM "docs/stories/databasev2" +WHERE status != "done" +SORT iteration ASC +``` + ## The porch track ```dataview diff --git a/docs/stories/databasev2/00-story.md b/docs/stories/databasev2/00-story.md new file mode 100644 index 0000000..a4be9ab --- /dev/null +++ b/docs/stories/databasev2/00-story.md @@ -0,0 +1,152 @@ +# Story — databasev2: the database beyond RAM + +The third track. `language-runtime-database/` built the engine; +[`porch/`](../porch/00-story.md) is the framework on top; this track answers the +question v1 deliberately deferred: **what happens when the data does not fit in +memory.** + +Numbering restarts at 1, local to this track. Frontmatter carries +`track: databasev2`, and iterations moved here keep their old id in +`was_language_iteration:` so a search for "iteration 32" still finds the WAL +checkpoint. Status stays where it belongs — the `status:` key, never a directory. + +## The problem, stated honestly + +Principle 7 says **RAM is authoritative; the WAL makes it durable.** That is a +real design, not a shortcut: reads never touch disk, so latency is predictable, +and durability is a sequential append rather than a storage engine bolted to the +side. Iteration 22 measured what it buys — reads at 1.3M ops/s after the index +probe landed, p50 1µs. + +The bill comes due at the ceiling. Read from the engine as it stands: + +- **Rows live in `malloc`'d slabs of 256, and their addresses are stable + forever** (`database/src/table.c`, `DB_SLAB_ROWS`). Slabs are allocated as a + table grows and freed only when the table is destroyed. The free-slot list + recycles removed slots, so a delete-heavy table plateaus — but a growing table + only grows. +- **The ceiling is process RSS, not a configured number.** `WO_HEAP_MB` (default + 64 MiB) bounds the VM object arena; table storage is separate `malloc`, so + nothing in the system declares a maximum dataset size. There is no knob that + says "this database may use at most N". +- **There is no eviction, no spill, no paging, no LRU.** Grep + `database/src/` for any of them and nothing comes back. Every row ever + inserted and not deleted is resident. +- **The WAL is append-only with no checkpoint.** Boot replays every record ever + written, so startup time is O(all writes in the file's history) and disk grows + without bound. That is databasev2 [3](03-wal-checkpoint.md). +- **Durability is process-global.** `WO_DATA` is one environment variable that + turns on one `shard-0.wal` for the whole process (`runtime/src/main.c`). There + is no way to say "this table matters, that one is scratch". + +### What actually breaks first + +Worth being precise, because the failure mode determines the fix — and the good +news is that the engine's own behaviour is clean: + +**An allocation failure is a catchable trap, not a crash.** Every `malloc` in +the row encoder is checked and jumps to an `oom` label; `DB_ERR_OOM` maps to +`WO_T_OOM`, which a program can `try`/`catch`. So a writeonce program that runs +out of memory *refuses the insert* rather than corrupting or dying. That is a +much better starting position than most engines have. + +**But the trap is almost never what a real deployment hits first.** Long before +`malloc` returns NULL, the box starts swapping, and a RAM-authoritative database +on swap is the worst of both worlds: it has paid for in-memory data structures +and is now serving them from disk with no read path designed for that. On a +cgroup-limited host the OOM killer arrives instead, and an external `SIGKILL` is +the one shutdown path that skips every guarantee the WAL was written to provide — +though ack-after-fsync means acked writes still survive; iteration 22's `kill -9` +battery proves that much. + +So the honest problem statement is not "malloc fails". It is: **there is no +declared budget, no back-pressure as the budget is approached, and no way to +distinguish data that must be resident from data that merely is.** Iteration +[1](01-ram-ceiling-measurement.md) exists to replace this paragraph with +numbers before anything is designed on top of it. + +## The lever: per-table storage modes + +The developer's ask, and the reason this track has a grammar iteration. + +Today every `@table` is identical: resident, and durable if and only if +`WO_DATA` is set for the whole process. Real applications are not uniform — +a session table, a rate-limit counter and a page cache want *resident and +disposable*; an orders table wants *resident and durable*; an audit log wants +*durable and rarely read*. One global switch cannot express that, so it forces +either "everything is precious" or "nothing is". + +Extending `@table` with a storage mode moves the decision into the language, +where the compiler can act on it: + +- **`ram`** — resident, never WAL-logged, gone on restart. The compiler knows + no durability code is needed; the engine knows these rows are the first + candidates to shed under pressure; and — the part that matters — a program + that expects a `ram` table to survive a restart is now stating something the + compiler can refuse. +- **`durable`** — today's behaviour: resident and WAL-logged, ack after fsync. +- **`cold`** — durable, and *not* required to be resident. This is the mode that + actually raises the ceiling, and it is the one with real design work behind it + (iteration [6](06-cold-tiering.md)). + +The grammar change is small and the surface is already the right shape: +`Ast.table_cfg` is `{ table_name; indexes }`, the parser's argument match +already rejects unknown keys with a catalogued diagnostic +(`unknown @table argument ... (supported: name, index)`), and adding one more +key follows the path `index` already cut. The *semantics* are the work, not the +syntax — which is exactly why it gets its own iteration +([2](02-table-storage-modes.md)) and why it comes after the measurement. + +This is also the honest answer to "does this break principle 7?" It does not. +RAM stays authoritative **for the tables that say so**. `cold` is a declared +exception a developer opts into per table, with the trade written at the +declaration site rather than buried in an operations runbook. + +## The sequence + +| # | Iteration | Delivers | Needs | +| --- | --- | --- | --- | +| 1 | [RAM ceiling: measure the breaking point](01-ram-ceiling-measurement.md) | what actually happens from 50% RAM to OOM — swap onset, latency cliff, trap behaviour, `kill -9` survival | nothing; extends iteration 22's harness | +| 2 | [`@table` storage modes](02-table-storage-modes.md) | the grammar: `mode: ram \| durable \| cold`, per table, replacing the global `WO_DATA` all-or-nothing | 1 for its defaults | +| 3 | [WAL checkpoint](03-wal-checkpoint.md) *(was language 32)* | snapshot + truncate: disk reclaimed, replay bounded | 4 composes | +| 4 | [io_uring group commit](04-io-uring-commit.md) *(was language 23)* | close the 66× durable/RAM write gap (4.5k vs 297k inserts/s) | the arc (landed) | +| 5 | [Bounded tables and eviction](05-bounded-tables-eviction.md) | a capacity a `ram` table may not exceed, and what happens when it does | 2 | +| 6 | [Cold tiering](06-cold-tiering.md) | rows that leave RAM and come back — the iteration that raises the ceiling | 2, 3, 5 | +| 7 | [Single-file store](07-single-file-db.md) *(was language 33)* | `WO_DATA=.db` — a file path IS the store | independent | +| 8 | [Query grammar from corpora](08-query-grammar-corpus.md) *(was language 27)* | whole-query `count`, `exists` | independent | +| 9 | [Cross-program tables](09-cross-program-tables.md) *(was language 20)* | attach to a running program's database over local IPC | independent | +| 10 | [Keypair attach auth](10-keypair-attach-auth.md) *(was language 21)* | program identity as a keypair; mutual challenge–response | 9 | + +``` +1 ──▶ 2 ──▶ 5 ──▶ 6 + │ ▲ + 3 ──▶ 4 ─────┘ +7, 8 independent +9 ──▶ 10 +``` + +Order rationale: **1 before 2** because a mode's default should follow from a +measurement, not a guess. **3 and 4 before 6** because tiering onto a log that +never truncates would make the disk problem worse, not better. **5 before 6** +because eviction from a bounded resident table is the simpler half of the same +mechanism, and getting the policy right there de-risks the hard half. + +## What this track does NOT own + +| Not databasev2's | Owner | +| --- | --- | +| `transaction { }` and `@table` feature flags | language [iteration 18](../language-runtime-database/18-memory-db-features.md) — approved spec, left whole on purpose | +| the TTL cache middleware | also language 18 (and [porch 1](../porch/01-store-backed-middleware.md) points there) | +| typed binding of rows into app classes | language [iteration 29 `@derive`](../language-runtime-database/29-compile-time-metaprogramming.md) | +| `fs` mutation verbs, outbound sockets | language [iteration 38](../language-runtime-database/38-content-platform-capabilities.md) | +| benchmark harness and CI | iteration 22 (landed) built the harness; per-change CI is language iteration 30 | +| a paged B-tree storage engine | **nobody, deliberately.** Recorded as rejected in [`discarded.md`](../../plan/discarded.md): the disk story is the WAL. `cold` tiering is not a licence to build SQLite. | + +## Review protocol + +The language track's, unchanged: one iteration read and approved before the next +starts; every iteration an unsplittable slice with phases, per-phase tasks, +Given/When/Then criteria and an out-of-scope list. Every engine change is gated +by `just employee`, `just db-actor` and `just db-bench` against +`bench/baseline.json` — and any iteration that claims a performance change must +move a number in that baseline, or it did not happen. diff --git a/docs/stories/databasev2/01-ram-ceiling-measurement.md b/docs/stories/databasev2/01-ram-ceiling-measurement.md new file mode 100644 index 0000000..4234194 --- /dev/null +++ b/docs/stories/databasev2/01-ram-ceiling-measurement.md @@ -0,0 +1,148 @@ +--- +track: databasev2 +iteration: "1" +status: refine +--- + +# databasev2 1 — the RAM ceiling: measure the breaking point before designing for it + +> Part of [Story — databasev2: the database beyond RAM](00-story.md). +> +> **First because the repo's own doctrine says so.** "Always inspect crashsites. +> Always measure. Never assume." Every later iteration in this track — the +> storage modes' defaults, the eviction policy, the tiering threshold — is a +> decision that should follow from a number. Right now nobody in this project +> can say what happens to a writeonce program at 90% of RAM, and designing +> tiering without that is guessing with extra steps. + +## Goals + +- **Find the curve, not the cliff.** Not "does it die" — it dies, everything + does. What matters is the shape on the way down: at what fraction of RAM does + p99 read latency leave its 1µs baseline, what does insert throughput do as + slabs stop coming from a warm allocator, and how much warning is there between + "fine" and "unusable". +- **Characterise all three exits.** The engine can leave the happy path three + ways and they are not equally survivable: a checked `malloc` failure + (`DB_ERR_OOM` → `WO_T_OOM`, a catchable trap — the clean one), swap thrash + (no trap, no error, just latency collapse — the dangerous one because nothing + reports it), and the external OOM killer (`SIGKILL`, skipping every shutdown + path). Establish which arrives first under realistic limits, because the + answer determines whether the fix is back-pressure or eviction. +- **Prove the durability floor holds at the ceiling.** Iteration 22's `kill -9` + battery proved acked writes survive under load. Re-run it *at memory + exhaustion*, which is a different and nastier state — an allocation failure + mid-commit is exactly where an ack-before-durable bug would hide. +- **Publish numbers others can build on.** The output is a section in + `perf-targets.md` and rows in `bench/baseline.json`, not a paragraph of + prose. A measurement that only printed once is not a measurement. + +## Phases + +### Phase A — a workload that can actually reach the ceiling + +- Extend `docs/examples/db-bench` with a growth mode: insert until a target RSS + fraction, holding row shape and index count constant so the variable is size + alone. +- Run it under an explicit memory limit (a cgroup or `ulimit`) rather than on a + big box — "it survived on a 64 GB workstation" measures the workstation. +- Record RSS against row count so the per-row overhead is known: slab headroom, + the id hash, the secondary-index multimaps and the per-row engine-owned values + (`db_text`, `db_rec`, `db_multi`, `db_map` are each their own allocation). +- Verify: RSS growth is linear and its slope is written down; the run is + reproducible twice within the tolerance policy iteration 22 established. + +### Phase B — the latency and throughput curve + +- Sample read p50/p99, query p99 and insert throughput at fixed fractions of the + limit, so the result is a curve rather than two endpoints. +- Separate the two effects deliberately: allocator pressure (still resident) and + swap (no longer resident). They have different fixes and conflating them would + send iteration 6 after the wrong one. +- Include the DB-actor path, since a cross-shard statement's reply materialises + a copy — memory pressure and the actor RPC interact and nobody has looked. +- Verify: the curve is recorded per metric class with iteration 22's per-class + tolerances; the swap onset point is identified, not interpolated. + +### Phase C — the three exits, deliberately triggered + +- Drive a checked allocation failure and confirm `WO_T_OOM` is catchable, the + insert is refused whole, no partial row or index entry is left, and the + process continues serving. +- Drive swap thrash and record what a client sees. This is the case with no + error signal at all, and naming it is most of the value of this iteration. +- Drive the OOM killer under a cgroup limit and confirm what survives: replay + the WAL and check every acked write is present. +- Verify: the trap path leaves no torn state (row count and index agree after a + refused insert); replay after `SIGKILL` at exhaustion loses no acked write. + +### Phase D — write it down where decisions get made + +- A `perf-targets.md` section with the curve, the swap onset, the per-row + overhead and the exit characterisation. +- Baseline rows for the growth metrics so a regression is caught by the existing + gate rather than by a person remembering. +- A short statement of what the numbers *imply* for iterations 2, 5 and 6 — + which is the point of going first. +- Verify: `just db-bench` green against the extended baseline; the gate bites + when a growth metric is doctored. + +## Acceptance Criteria + +- **Given** the growth workload under a fixed memory limit, **when** it runs + twice, **then** RSS-per-row agrees within the tolerance policy and the slope + is recorded in `perf-targets.md`. +- **Given** the workload at rising RAM fractions, **when** latency is sampled, + **then** the fraction at which read p99 first leaves its baseline is + identified as a measured point, not an estimate. +- **Given** a deliberately induced allocation failure, **when** an insert is + attempted, **then** it traps `WO_T_OOM` catchably, the table's row count is + unchanged, every index agrees with the slab contents, and the process keeps + serving subsequent requests. +- **Given** swap thrash, **when** a client issues reads, **then** the observed + degradation is quantified and the fact that **no error is surfaced** is + recorded explicitly as a finding. +- **Given** a cgroup limit and a workload that exceeds it, **when** the OOM + killer fires, **then** replaying the WAL shows every acked write present — + ack-after-fsync holding in the one shutdown path that skips all cleanup. +- **Given** the extended baseline, **when** a growth metric is doctored, **then** + `just db-bench` fails on exactly that metric. + +## Out Of Scope + +- **Any fix.** This iteration measures. Eviction is + [5](05-bounded-tables-eviction.md), tiering is [6](06-cold-tiering.md), + declared budgets are [2](02-table-storage-modes.md). Shipping a fix inside the + measurement slice would remove the ability to tell whether it helped. +- **Changing the OOM behaviour.** The checked-`malloc`-to-catchable-trap path is + good and should not be touched; if the measurement finds a hole in it, that is + a bug fix, reported separately. +- **A memory profiler or allocator instrumentation.** Observability is language + iteration 30. RSS from the OS and the existing `time.ticks` are enough for a + curve. +- **Multi-machine or sharded-across-hosts scaling.** One binary owns its data; + cross-process is [9](09-cross-program-tables.md). +- **Comparing against SQLite at the ceiling.** `bench/compare/go-sqlite` exists + and the comparison would be interesting, but SQLite's whole architecture is + the paged design this project rejected — the numbers would not inform any + decision here. + +## Info + +Forks the spec must settle: + +1. **What is the limit mechanism for the harness?** A cgroup v2 `memory.max` is + the closest thing to how this would actually be deployed; `ulimit -v` is + simpler but bounds address space rather than resident set, which for an engine + that `malloc`s slabs is a materially different constraint. Leaning cgroup, and + the campaign already runs off the fast path so the setup cost is acceptable. +2. **Which table shape is the reference?** Per-row overhead depends heavily on + whether fields are scalars or heap values — a `Text` column is a separate + `db_text` allocation per row, so a text-heavy table and an Int-only table + will produce very different slopes. Probably both, reported separately, + because "bytes per row" is meaningless without saying which row. +3. **Is swap even in scope for the target deployment?** If the intended answer + is "run with swap off and let the OOM killer decide", the swap curve is + informational rather than load-bearing — but that stance should be stated in + the doctrine, not assumed. It also changes which exit iteration 5's + back-pressure is defending against. diff --git a/docs/stories/databasev2/02-table-storage-modes.md b/docs/stories/databasev2/02-table-storage-modes.md new file mode 100644 index 0000000..f2ad592 --- /dev/null +++ b/docs/stories/databasev2/02-table-storage-modes.md @@ -0,0 +1,188 @@ +--- +track: databasev2 +iteration: "2" +status: refine +--- + +# databasev2 2 — `@table` storage modes: durability becomes a language decision + +> Part of [Story — databasev2: the database beyond RAM](00-story.md). +> Needs [1](01-ram-ceiling-measurement.md) — a mode's default should follow from +> a measurement, not a preference. +> +> **The developer's ask, and the language enrichment this track exists for.** +> Today durability is one environment variable for a whole process: `WO_DATA` is +> set and every `@table` is WAL-logged, or it is not and none are +> (`runtime/src/main.c`). Real applications are not uniform. A session table, a +> rate-limit counter and a page cache are resident and disposable; an orders +> table is resident and precious; an audit log is precious and rarely read. One +> global switch forces "everything is precious" or "nothing is", and the +> developer pays for the wrong one either way. + +## Goals + +- **Move the storage decision to the declaration site.** `@table(mode: ...)`, + chosen per table, in the source, where the person who knows what the data is + worth is already writing. An operations runbook is the wrong place for a fact + the compiler could hold. +- **Three modes, each earning its existence.** `ram` — resident, never logged, + gone on restart. `durable` — today's behaviour, resident and ack-after-fsync, + and the **default** so every existing program is byte-identical. `cold` — + durable and not required to be resident, declared here and *implemented* in + [6](06-cold-tiering.md), because a mode with no engine behind it is a promise. +- **Make the compiler enforce what the mode means.** This is the part that makes + it a language feature rather than a config key. A program that inserts into a + `ram` table and expects the row after a restart is stating a contradiction, and + the compiler is the right place to say so — at minimum for the cases it can + see statically, with the diagnostic catalogued like every other. +- **Let the engine act on it.** `ram` tables skip the WAL write entirely, which + is not merely a saving — it is the 66× gap iteration 22 measured (4.5k durable + vs 297k RAM inserts/s) becoming available per table instead of per process. + They also become the first candidates to shed under pressure, which is what + [5](05-bounded-tables-eviction.md) builds on. +- **Keep principle 7 intact and say why.** RAM stays authoritative for every + table that says so. `cold` is a declared, per-table exception a developer opts + into with the trade visible at the declaration. + +## Phases + +### Phase A — the grammar + +- Extend the `@table` argument parser with `mode:`. The path is already cut: + the argument loop matches `name` and `index` and rejects anything else with a + catalogued diagnostic (`unknown @table argument ... (supported: name, index)`), + so this is one more arm plus an updated message. +- Extend `Ast.table_cfg` — today `{ table_name; indexes }` — with the mode, and + give it the default so every existing `@table` keeps its meaning. +- Reject the incoherent cases at parse time: `mode` given twice, an unknown mode + name. Both belong in the same diagnostic family as the existing `@table` + errors, and both go in the error catalog **in this change**, not later — that + catalog went stale once by exactly that omission. +- Verify: golden AST fixtures for each mode; `woc-test` green; every existing + `@table` in every sample parses unchanged. + +### Phase B — the mode reaches the image and the engine + +- Carry the mode through the class descriptor into the `.wob` image so the + runtime knows it without re-deriving anything. This is a format change, so it + moves `WOB_VERSION` and the format contract in the same commit as the code. +- The engine consults it at the one choke point that already exists: the + `INDEX HOOK` / WAL staging site in `wo_row_insert` / `wo_row_remove`, which + `database/src/CODE-LOGIC.md` names as the only place storage may be mutated. + A `ram` table stages nothing. +- Replay must skip records for tables that are now `ram` — a WAL written when a + table was `durable` and replayed after the source changed is a real + migration case, and silently resurrecting rows into a `ram` table would be + worse than refusing. +- Verify: a `ram` table's inserts produce no WAL growth (measured, not assumed); + a `durable` table is byte-identical to today; a mode change across a restart + is handled explicitly rather than by accident. + +### Phase C — the compiler's enforcement + +- Decide how far static checking goes (fork 2) and implement that much. The + floor: `WO_DATA` set with every table `ram` is a program that asked for a data + directory it will never write to — worth a warning at least. +- The interesting case is a `ram` table participating in a `ref`/`backlink` + relation with a `durable` one. A durable row holding a foreign key into a + table that evaporates on restart is a dangling reference by construction, and + FK-restrict cannot save it. This is the check most worth having, and it is + statically visible from the class table. +- Verify: corpus `compile-fail` fixtures for each refusal; each carries the + exact `WO-E###` the catalog now documents. + +### Phase D — prove it on a real workload + +- Give `docs/examples/employee` or the `db-bench` sample a mixed schema — at + least one `ram` table and one `durable` — and gate the distinction: after a + restart, the durable rows are present and the ram rows are gone. That single + assertion is the whole feature. +- Extend the durability legs of `just db-bench` so the per-table write-path + saving appears as a baseline number, not a claim. +- Verify: `just employee`, `just db-actor`, `just db-bench`, `oop-accept` green. + +### Phase E — document the contract + +- The `@table` mode surface in the language-surface guide and the db-binding + contract; the new diagnostics in the error catalog; the mode's effect on + replay in `database/src/CODE-LOGIC.md`. +- Verify: `just linkcheck` clean; the language-surface guide's `@table` row + matches what the parser actually accepts. + +## Acceptance Criteria + +- **Given** every existing `@table` declaration in the repository, **when** it is + compiled after this change, **then** behaviour is byte-identical — `durable` + is the default and nothing opts in silently. +- **Given** a table declared `mode: ram`, **when** rows are inserted with + `WO_DATA` set, **then** the WAL does not grow, and after a restart the table is + empty while `durable` tables in the same program replay intact. +- **Given** `mode: ram` and a measured insert workload, **when** it runs against + the same shape as a `durable` table, **then** the write-path saving is visible + in `bench/baseline.json` — the per-table half of iteration 22's 66× gap. +- **Given** an unknown mode name or `mode:` given twice, **when** it is compiled, + **then** it fails with the catalogued diagnostic naming the legal modes. +- **Given** a `durable` table holding a `ref` into a `ram` table, **when** it is + compiled, **then** the compiler refuses (or warns, per fork 2) — a persistent + row cannot reference one that evaporates. +- **Given** a WAL containing records for a table whose source now says `ram`, + **when** the program starts, **then** the situation is handled explicitly + (refuse, or skip and report) and never by silently loading rows into a table + declared not to have any. +- **Given** the `.wob` format change, **when** an image from the previous version + is loaded, **then** the loader refuses it clearly on the version rather than + misreading a descriptor. + +## Out Of Scope + +- **Implementing `cold`.** Declared here so the mode set is settled and the + format carries it; the engine behaviour is [6](06-cold-tiering.md). Until then + a `cold` declaration must be refused rather than silently treated as + `durable` — accepting a mode that does nothing is how a feature becomes a lie. +- **Per-table capacity limits and eviction** — [5](05-bounded-tables-eviction.md). + This iteration says what a table *is*; that one says how much of it there may + be. +- **`@table` feature flags and `transaction { }`** — language + [iteration 18](../language-runtime-database/18-memory-db-features.md), whose + spec is approved and deliberately left whole. +- **Per-table WAL files.** One log, one writer, shard 0 — the invariant stage 3 + established and [7](07-single-file-db.md) depends on. Modes decide *whether* a + table logs, never *where*. +- **Migrating an existing dataset between modes.** A schema-change story, and + the repo already records destructive migrations as a recorded future. +- **Encryption at rest, compression of the WAL.** Neither has a consumer. + +## Info + +The grammar surface this touches, read from the source: the parser's `@table` +argument loop and its `unknown @table argument` failure; `Ast.table_cfg` as +`{ table_name : string option; indexes : string list list }`; and the class +descriptor in `runtime/src/wob.h` that the loader validates. Adding a key is +genuinely small — the semantics are the iteration. + +Forks the spec must settle: + +1. **What are the modes called?** `ram` / `durable` / `cold` is descriptive of + mechanism. `scratch` / `persistent` / `archived` is descriptive of intent and + is what a developer reasons about. The names are the API and are hard to + change later; leaning the intent-shaped set for the first two if a + short-enough pair can be found, since a developer choosing a mode is thinking + about what the data is *for*, not about where it sits. +2. **How hard does the compiler push?** Three levels: warn on the suspicious + cases; refuse the provably-broken ones (a `durable`→`ram` `ref`); or a full + dataflow check that an insert into a `ram` table is never expected to persist. + The third is not statically decidable in general. Leaning: refuse the + relation case (provable, high value), warn on the `WO_DATA`-with-no-durable- + table case, and stop there. +3. **What is the default, and does it depend on iteration 1?** `durable` keeps + every existing program identical, which is nearly decisive. But if iteration + 1's numbers show the WAL write dominating a workload nobody wanted durable, + there is an argument for making the choice mandatory — no default, every + `@table` states its mode. That is a bigger source change and a better + language; the fork is whether the churn is worth it now or at 1.0. +4. **Does `mode: ram` imply anything about the actor/DB-actor path?** Stage 3 + marshals worker-shard statements to shard 0 because the owner shard holds the + store and the WAL. A `ram` table has no WAL — so does it still need to live on + the owner? A per-shard `ram` table would be dramatically faster and a + different consistency story. Tempting, out of scope here, and worth recording + as a candidate rather than deciding in passing. diff --git a/docs/stories/language-runtime-database/32-wal-checkpoint.md b/docs/stories/databasev2/03-wal-checkpoint.md similarity index 91% rename from docs/stories/language-runtime-database/32-wal-checkpoint.md rename to docs/stories/databasev2/03-wal-checkpoint.md index 981da48..de775e1 100644 --- a/docs/stories/language-runtime-database/32-wal-checkpoint.md +++ b/docs/stories/databasev2/03-wal-checkpoint.md @@ -1,13 +1,19 @@ --- -iteration: "32" +track: databasev2 +iteration: "3" +was_language_iteration: "32" status: refine chain: 6 --- -# Iteration 32 — WAL checkpoint: disk space reclamation and bounded replay +# databasev2 3 — WAL checkpoint: disk space reclamation and bounded replay + +> **Moved 2026-08-26** from the language track, where this was iteration 32. +> Part of [Story — the database beyond RAM](../language-runtime-database/00-story.md). Content unchanged by +> the move; its dependencies are restated in that track index. > Format: `product/story-iteration-template`. Part of -> [Story — one language, one runtime, one database, one binary](00-story.md). +> [Story — one language, one runtime, one database, one binary](../language-runtime-database/00-story.md). > > **Inserted 2026-08-21** (stage-3 guarantee refinement found the hole): > the WAL is append-only FOREVER — no checkpoint, no truncation exists diff --git a/docs/stories/language-runtime-database/23-io-uring-commit.md b/docs/stories/databasev2/04-io-uring-commit.md similarity index 95% rename from docs/stories/language-runtime-database/23-io-uring-commit.md rename to docs/stories/databasev2/04-io-uring-commit.md index bb659df..0dec19f 100644 --- a/docs/stories/language-runtime-database/23-io-uring-commit.md +++ b/docs/stories/databasev2/04-io-uring-commit.md @@ -1,13 +1,19 @@ --- -iteration: "23" +track: databasev2 +iteration: "4" +was_language_iteration: "23" status: refine chain: 5 --- -# Iteration 23 — io_uring group-commit write path +# databasev2 4 — io_uring group-commit write path + +> **Moved 2026-08-26** from the language track, where this was iteration 23. +> Part of [Story — the database beyond RAM](../language-runtime-database/00-story.md). Content unchanged by +> the move; its dependencies are restated in that track index. > Format: `product/story-iteration-template`. Part of -> [Story — one language, one runtime, one database, one binary](00-story.md). +> [Story — one language, one runtime, one database, one binary](../language-runtime-database/00-story.md). > > **Inserted 2026-08-15.** The write-path optimization, and deliberately the > LAST database performance iteration: it only earns its complexity once diff --git a/docs/stories/databasev2/05-bounded-tables-eviction.md b/docs/stories/databasev2/05-bounded-tables-eviction.md new file mode 100644 index 0000000..dc5010a --- /dev/null +++ b/docs/stories/databasev2/05-bounded-tables-eviction.md @@ -0,0 +1,169 @@ +--- +track: databasev2 +iteration: "5" +status: refine +--- + +# databasev2 5 — bounded tables and eviction: a declared budget, and back-pressure before the cliff + +> Part of [Story — databasev2: the database beyond RAM](00-story.md). +> Needs [2](02-table-storage-modes.md) for the mode a bound attaches to, and +> [1](01-ram-ceiling-measurement.md) for the numbers that set a sane default. +> +> **The simpler half of the hard problem, done first on purpose.** Evicting from +> a bounded resident table and evicting to disk are the same policy question with +> different destinations. Getting the policy right where the answer is "drop it" +> de-risks [6](06-cold-tiering.md), where the answer is "write it somewhere and +> be able to find it again". + +## Goals + +- **A table may declare a maximum.** Rows, bytes, or both — a `ram` table that + is a cache or a session store has a size the application is willing to spend, + and today it has no way to say so. Unbounded growth in a table nobody intended + to be large is the most common route to the ceiling iteration 1 measured. +- **Something defined happens at the bound.** Today the answer is "grow until + the process dies". The candidates are eviction (drop the least valuable row), + refusal (trap, let the caller decide), and back-pressure (make the writer + wait). Each is right for a different table, which argues for the policy being + declared rather than chosen for the developer. +- **Back-pressure before the cliff, not at it.** The dangerous exit iteration 1 + characterises is swap thrash, which arrives with **no error signal at all**. + A budget that is enforced at 100% has already lost; the value is in acting at + a threshold, while there is still headroom to act. +- **Eviction that respects the engine's actual invariants.** Rows have stable + addresses forever, the free-slot list recycles slots, ids are never reused, and + every secondary index and unique shadow must stay consistent with the slab. + Eviction is `wo_row_remove` with a policy in front — it must go through the + same choke point, not around it. + +## Phases + +### Phase A — declaring the bound + +- Extend the `@table` surface from [2](02-table-storage-modes.md) with a + capacity and a policy. One new grammar arm, the same catalogued-diagnostic + discipline, defaults that keep every existing table unbounded so nothing + changes silently. +- Decide whether a bound is legal on a `durable` table (fork 1) — evicting a row + that was acked as durable is a promise being broken, and the answer is + probably "only with an explicit, differently-named policy". +- Verify: golden fixtures per policy; existing tables unchanged; illegal + combinations refused at compile time with catalogued codes. + +### Phase B — the accounting + +- Track per-table size cheaply. Row count is free; bytes are not — per-row + footprint includes the slab slot plus each heap value's own allocation + (`db_text`, `db_rec`, `db_multi`, `db_map`), which iteration 1 will have + quantified. Decide what is counted and be honest that it is an estimate of RSS, + not RSS. +- Expose it, because a budget nobody can observe is a budget nobody can tune. + How it is exposed is fork 3. +- Verify: accounting tracks a known workload within a stated error bound; + deleting rows returns the accounting to its prior value (the free-slot list + already makes this true of slots). + +### Phase C — the policies + +- **Refuse**: the bound is a hard ceiling, an insert past it traps catchably. + The simplest correct behaviour and the right default for anything precious. +- **Evict**: drop the least valuable row via `wo_row_remove` so indexes, unique + shadows and the free-slot list all stay honest. The recency metadata this needs + is fork 2 — and note the engine currently stores no per-row access time, so + true LRU is not free. +- **Back-pressure**: at a threshold below the bound, slow or park the writer. + This composes with the fiber model (a parked writer blocks nobody) and is the + only policy that addresses the swap-thrash exit rather than the allocation + exit. +- Verify: each policy behaves at the bound; after eviction every index agrees + with the slab; a parked writer resumes and does not deadlock the shard. + +### Phase D — the pressure signal + +- A process-level threshold, not just per-table: when total resident size + crosses a configured fraction, tables with an eviction policy start shedding + **before** the allocator or the OS gets involved. This is the iteration's real + contribution — turning an invisible failure into a managed one. +- Decide precedence when several tables could shed (fork 4). +- Verify: under the iteration-1 growth workload with the signal enabled, the + process holds a steady state instead of walking into swap; the latency curve + stays inside its baseline. + +### Phase E — gate it against the measurement + +- Re-run iteration 1's growth workload with bounds and policies configured. The + proof is a before/after on the same harness: previously the curve degraded and + the process died; now it plateaus. +- Baseline rows for steady-state throughput under pressure. +- Verify: `just employee`, `just db-actor`, `just db-bench` green; the gate bites + on a doctored pressure metric. + +## Acceptance Criteria + +- **Given** a table bounded at N rows with policy `refuse`, **when** the N+1st + insert is attempted, **then** it traps catchably, the row count stays N, and + every index agrees with the slab. +- **Given** a table bounded at N rows with policy `evict`, **when** the N+1st + insert arrives, **then** exactly one row is evicted, the new row is present, + the count is N, and no index or unique shadow references the evicted row. +- **Given** an evicted row's id, **when** it is looked up, **then** it is absent + — and its id is never reused by a later insert, preserving the invariant the + id hash's tombstone sentinel depends on. +- **Given** a bound expressed in bytes, **when** rows of a known shape are + inserted, **then** the bound is honoured within the stated accounting error, + and that error is documented rather than implied. +- **Given** back-pressure configured at a threshold, **when** the threshold is + crossed, **then** writers are slowed or parked, reads are unaffected, and no + shard deadlocks. +- **Given** the process-level pressure signal and the iteration-1 growth + workload, **when** it runs to what previously exhausted memory, **then** the + process reaches a steady state and read p99 stays within its baseline — the + before/after that justifies the iteration. +- **Given** a `durable` table, **when** an eviction policy is applied to it, + **then** either it is refused at compile time or it is a distinctly named + policy that says out loud it discards acked data. + +## Out Of Scope + +- **Writing evicted rows anywhere** — that is [6](06-cold-tiering.md). Here + eviction means the row is gone. Keeping the two apart is what makes the policy + work reviewable on its own. +- **True LRU if it costs a write per read.** Touching per-row metadata on every + read would turn the 1µs read path into a write path — the same trap iteration + 3's session touch has. An approximation (insertion order, a coarse clock, a + sampled counter) is very likely the right answer and fork 2 should say so + explicitly rather than defaulting to textbook LRU. +- **The TTL cache middleware** — language + [iteration 18](../language-runtime-database/18-memory-db-features.md). Expiry + by *time* is that; bounding by *size* is this. They compose. +- **Query-level result limits.** `take n` already exists in the query surface. +- **Shrinking slabs back to the allocator.** Slab addresses are stable forever + by design and that invariant is load-bearing; reclaiming a slab whose rows were + all evicted is a separate, delicate change with its own iteration if anyone + wants it. + +## Info + +Forks the spec must settle: + +1. **May a `durable` table be bounded?** Evicting an acked row contradicts the + durability promise. But an audit table that must not grow forever is a real + need, and the honest form of it is probably archival (iteration 6) rather than + eviction. Leaning: bounds on `durable` are refused, and the need is redirected + to 6. +2. **What is "least valuable"?** No per-row access time exists today, so LRU + costs a write per read. Candidates: insertion order (free — ids are already + monotonic per table), a coarse epoch stamped on write only, or sampled + approximation. Insertion order is FIFO not LRU, which is wrong for a cache + and fine for a queue — so the policy name should say which it is rather than + claiming "LRU" and delivering FIFO. +3. **How is size observed?** Without observability (language iteration 30) there + is no metrics endpoint to publish it on. Options: a builtin returning a + table's current size, a `@table`-derived query, or stderr on threshold + crossing. A builtin is the smallest thing that makes the feature tunable by + the program that owns the budget. +4. **Precedence when several tables can shed.** Largest first is simple; the + application's own priority order is more correct and needs a way to express + it. Proportional shedding is fairest and hardest to reason about. This + decides whether the pressure signal is predictable enough to trust. diff --git a/docs/stories/databasev2/06-cold-tiering.md b/docs/stories/databasev2/06-cold-tiering.md new file mode 100644 index 0000000..0ac6994 --- /dev/null +++ b/docs/stories/databasev2/06-cold-tiering.md @@ -0,0 +1,183 @@ +--- +track: databasev2 +iteration: "6" +status: refine +--- + +# databasev2 6 — cold tiering: rows that leave RAM and come back + +> Part of [Story — databasev2: the database beyond RAM](00-story.md). +> Needs [2](02-table-storage-modes.md) for the `cold` mode declaration, +> [3](03-wal-checkpoint.md) so the log this builds on does not grow forever, +> and [5](05-bounded-tables-eviction.md) for the policy machinery. +> +> **The iteration that actually raises the ceiling, and the one most likely to +> go wrong.** Everything before it makes the limit visible, declared and +> managed. This one removes it — for tables that opt in — and in doing so +> touches the project's most load-bearing principle. It should be approached +> with more suspicion than enthusiasm. + +## Goals + +- **A `cold` table may hold more rows than fit in memory.** Recently-used rows + are resident; the rest live on disk and are faulted back on access. This is + the whole feature and every other goal is a constraint on it. +- **Do it without becoming a paged storage engine.** `discarded.md` records that + the disk story is the WAL and that a paged B-tree engine was rejected. `cold` + must not be a licence to rebuild SQLite inside `database/src`. The design that + respects the doctrine reuses the log that already exists plus an index into + it — a log-structured read path, not a page cache. +- **Keep the resident path exactly as fast as it is.** A `durable` or `ram` + table must not pay one instruction for a feature it does not use. Iteration + 22's 1.3M ops/s read baseline is the regression gate, and a measurable read + regression on non-`cold` tables is grounds to reject the design, not to tune + it. +- **Be honest in the query surface about what a fault costs.** A scan over a + `cold` table can touch disk per row. The engine currently answers every read + from memory at p50 1µs; a `cold` scan is a different animal and the language + should not pretend otherwise — see fork 3, which is the most important + question in this iteration. +- **Never lose an acked write.** Every guarantee iterations 9 and 22 established + holds byte-for-byte: ack-after-fsync, whole-or-nothing replay, torn tails + dropped by CRC. A tiering layer that weakens any of those is a regression + disguised as a feature. + +## Phases + +### Phase A — settle the design before writing any of it + +- This iteration needs a spec more than any other in the track. The candidate + shapes are genuinely different: (a) the WAL becomes the primary store with an + in-memory id→offset index and a resident row cache; (b) a separate + append-only row file per cold table, checkpointed by 3's machinery; (c) + eviction to disk with a free-space map, which is the paged engine wearing a + hat. +- Whichever wins must state its read amplification, its recovery story, and what + happens when the index itself does not fit — an id→offset map for a billion + rows is not free either, and a design that only moves the ceiling is worth + knowing about before it is built. +- Verify: the spec names the shape, the amplification, and the failure modes. + No code in this phase. + +### Phase B — the resident/cold boundary + +- Which rows are resident: reuse [5](05-bounded-tables-eviction.md)'s policy and + accounting rather than inventing a second notion of "least valuable". +- Eviction becomes write-then-drop instead of drop, and it must be atomic with + respect to a concurrent reader — a row that is being written out must not be + briefly unreachable. +- The fault path: a lookup that misses resident memory reads from disk, + materialises the row, and admits it under the resident policy. +- Verify: a table larger than the resident bound serves correct rows for every + id; a row evicted and faulted back is byte-identical, including every heap + value (`db_text`, `db_rec`, `db_multi`, `db_map` each round-trip). + +### Phase C — indexes and constraints across the boundary + +- **The hard part, and the reason this is late in the track.** A secondary index + over a `cold` table either stays fully resident (bounding the table by index + size rather than row size — which may be the honest answer) or is itself + tiered. A `@unique` constraint must hold across rows nobody has in memory: the + shadow check cannot scan a slab that is not there. +- Foreign-key restrict must also hold — a delete has to know whether any cold + row references it. +- Verify: `@unique` refuses a duplicate whose only conflicting row is cold; FK + restrict refuses a delete whose only referrer is cold. These two criteria are + the correctness core of the iteration. + +### Phase D — recovery + +- Crash mid-eviction, crash mid-fault, crash mid-checkpoint-of-a-cold-table. + Each must recover to a consistent state with no acked write lost and no row + visible twice. +- Interaction with [3](03-wal-checkpoint.md)'s snapshot: a cold table's on-disk + rows are part of the durable state a checkpoint must account for, not + something it can truncate past. +- Verify: `kill -9` at each of the three points, replayed, with every acked write + present and the resident/cold split re-derived correctly. + +### Phase E — measure it, then decide whether to keep it + +- Read/write throughput and p99 for a `cold` table at several resident ratios, + and a **regression check that non-`cold` tables did not move**. +- Publish the amplification honestly in `perf-targets.md`: how much slower a + cold fault is than a resident read, as a number. +- Verify: `just db-bench` green with new cold-path rows; the resident baseline + unchanged; `just employee`, `just db-actor`, `oop-accept` green; ASan and TSan + clean on the fault path. + +## Acceptance Criteria + +- **Given** a `cold` table with more rows than the resident bound, **when** any + row is looked up by id, **then** it is returned correctly whether resident or + faulted, byte-identical including every heap-valued column. +- **Given** a `cold` table under a read workload, **when** the resident set is + smaller than the working set, **then** the process holds steady state without + approaching the RAM ceiling iteration 1 measured. +- **Given** a `@unique` column on a `cold` table, **when** a duplicate is + inserted whose conflicting row is **not resident**, **then** the insert is + refused — the constraint holds across the boundary or it does not hold. +- **Given** a `ref` into a `cold` table, **when** the referenced row's owner is + deleted and the only referrer is cold, **then** FK restrict refuses the delete. +- **Given** `kill -9` during an eviction, a fault, and a checkpoint, **when** the + program restarts, **then** every acked write is present, no row appears twice, + and the resident/cold split is re-derived correctly. +- **Given** a `durable` or `ram` table, **when** the read benchmark runs after + this iteration, **then** its throughput and p99 are inside the existing + baseline tolerance — no cost for a feature not used. +- **Given** a cold fault, **when** its latency is measured, **then** the + amplification versus a resident read is recorded in `perf-targets.md` as a + number a developer can plan around. + +## Out Of Scope + +- **A paged B-tree storage engine.** Explicitly rejected in + [`discarded.md`](../../plan/discarded.md) and not reopened by this iteration. + If the spec phase concludes that tiering *requires* one, the correct outcome is + to reject tiering and say so — not to quietly build the thing the project + decided against. +- **Making `cold` the default, or applying it to a table that did not ask.** + Opt-in per table, forever. +- **Tiering to anything but the local filesystem.** Object storage needs + outbound sockets (language + [iteration 38](../language-runtime-database/38-content-platform-capabilities.md)) + and would change the latency story by orders of magnitude. +- **Compression of cold rows.** Composes with + [porch 7](../porch/07-sse-and-compression.md)'s codec if that lands first; + not a dependency either way and not this slice. +- **Cross-shard cold tables.** The owner shard owns the store and the WAL; a + cold table is more of the same. Per-shard storage is a separate architectural + question noted in [2](02-table-storage-modes.md)'s forks. +- **Tiering the query planner's behaviour.** If a scan over a cold table is + expensive, the answer for now is that it is expensive and documented — not a + cost-based planner. + +## Info + +Forks the spec must settle — this iteration is mostly forks, which is why phase +A produces no code: + +1. **Which shape?** WAL-as-primary-store with an id→offset index and a row + cache reuses machinery that exists and keeps the doctrine ("the disk story is + the WAL") literally true. A separate per-table row file is cleaner to reason + about and duplicates the log. Eviction with a free-space map is the rejected + paged design. Leaning (a), with the caveat in fork 2. +2. **What if the index does not fit either?** An id→offset entry per row is far + smaller than a row, so this moves the ceiling by a large constant — but it + does not remove it. Say so plainly in the spec: `cold` buys an order of + magnitude, not infinity. A design sold as unlimited will be deployed as if it + were. +3. **Does the language surface the cost?** Three positions. Silent — a cold + table reads like any other and the developer discovers the latency in + production. Annotated — the mode is at the declaration, so an attentive + reader knows, which is the status quo of this design. Or *explicit at the use + site*, where a query over a cold table must acknowledge it somehow. The third + is most in keeping with a language whose whole thesis is that the compiler + tells you the truth — and it is also the most intrusive. This is the fork with + the largest effect on what writeonce *is*, and it deserves the brainstorm more + than any implementation detail here. +4. **Is `@unique` on a cold table simply refused?** Keeping a unique index fully + resident is a bound on the table by index size, which is honest and simple. + Refusing `@unique` on `cold` outright is even simpler and might be right for + a first version — a constraint that silently only checks resident rows would + be a correctness hole, and that is the one outcome that must not ship. diff --git a/docs/stories/language-runtime-database/33-single-file-db.md b/docs/stories/databasev2/07-single-file-db.md similarity index 86% rename from docs/stories/language-runtime-database/33-single-file-db.md rename to docs/stories/databasev2/07-single-file-db.md index a42b80d..7eeb598 100644 --- a/docs/stories/language-runtime-database/33-single-file-db.md +++ b/docs/stories/databasev2/07-single-file-db.md @@ -1,12 +1,18 @@ --- -iteration: "33" +track: databasev2 +iteration: "7" +was_language_iteration: "33" status: refine --- -# Iteration 33 — `WO_DATA=.db`: the persistent store as one file +# databasev2 7 — `WO_DATA=.db`: the persistent store as one file + +> **Moved 2026-08-26** from the language track, where this was iteration 33. +> Part of [Story — the database beyond RAM](../language-runtime-database/00-story.md). Content unchanged by +> the move; its dependencies are restated in that track index. > Format: `product/story-iteration-template`. Part of -> [Story — one language, one runtime, one database, one binary](00-story.md). +> [Story — one language, one runtime, one database, one binary](../language-runtime-database/00-story.md). > > **Inserted 2026-08-22** (developer ask: "can the persistent db be in > file.db form?"). The truth is already almost there: `WO_DATA=` @@ -46,7 +52,7 @@ status: refine - A paged database file (SQLite's shape) — RAM is authoritative; the disk story is the WAL, full stop. -- Checkpoint/compaction — [iteration 32](32-wal-checkpoint.md)'s; its +- Checkpoint/compaction — [iteration 32](03-wal-checkpoint.md)'s; its rename-swap (write snapshot+tail to a NEW file, fsync, `rename()` over the old) is exactly what keeps the single-file promise crash-safe when it lands. 33 before or after 32 works; landing 33 first means diff --git a/docs/stories/language-runtime-database/27-query-grammar-corpus.md b/docs/stories/databasev2/08-query-grammar-corpus.md similarity index 95% rename from docs/stories/language-runtime-database/27-query-grammar-corpus.md rename to docs/stories/databasev2/08-query-grammar-corpus.md index 155f394..8113af6 100644 --- a/docs/stories/language-runtime-database/27-query-grammar-corpus.md +++ b/docs/stories/databasev2/08-query-grammar-corpus.md @@ -1,12 +1,18 @@ --- -iteration: "27" +track: databasev2 +iteration: "8" +was_language_iteration: "27" status: hold --- -# Iteration 27 — query grammar, driven by real embedded-DB corpora +# databasev2 8 — query grammar, driven by real embedded-DB corpora + +> **Moved 2026-08-26** from the language track, where this was iteration 27. +> Part of [Story — the database beyond RAM](../language-runtime-database/00-story.md). Content unchanged by +> the move; its dependencies are restated in that track index. > Format: `product/story-iteration-template`. Part of -> [Story — one language, one runtime, one database, one binary](00-story.md). +> [Story — one language, one runtime, one database, one binary](../language-runtime-database/00-story.md). > > **Inserted 2026-08-16.** A query-surface iteration in the 9b family: the > language-integrated query grows to cover the grammar that *real diff --git a/docs/stories/language-runtime-database/20-cross-program-tables.md b/docs/stories/databasev2/09-cross-program-tables.md similarity index 95% rename from docs/stories/language-runtime-database/20-cross-program-tables.md rename to docs/stories/databasev2/09-cross-program-tables.md index 5e38cfe..86b9bff 100644 --- a/docs/stories/language-runtime-database/20-cross-program-tables.md +++ b/docs/stories/databasev2/09-cross-program-tables.md @@ -1,12 +1,18 @@ --- -iteration: "20" +track: databasev2 +iteration: "9" +was_language_iteration: "20" status: hold --- -# Iteration 20 — cross-program tables: attach to a running program's database +# databasev2 9 — cross-program tables: attach to a running program's database + +> **Moved 2026-08-26** from the language track, where this was iteration 20. +> Part of [Story — the database beyond RAM](../language-runtime-database/00-story.md). Content unchanged by +> the move; its dependencies are restated in that track index. > Format: `product/story-iteration-template`. Part of -> [Story — one language, one runtime, one database, one binary](00-story.md). +> [Story — one language, one runtime, one database, one binary](../language-runtime-database/00-story.md). > > **Inserted 2026-08-15**, hence `20`. It follows 9b because a program > attaching to another's tables wants the same typed statements and queries @@ -141,7 +147,7 @@ read or read+write (per-table refinement deferred until a workload needs it), and the registration is A's manifest so a grant is a config change + restart, not an API. **Superseded as the end state (2026-08-15):** identity is a keypair and grants name public keys — iteration -[21](21-keypair-attach-auth.md) owns that; the uid check is only this +[21](10-keypair-attach-auth.md) owns that; the uid check is only this iteration's bootstrap and must be flagged pre-21 wherever it ships. **4. What does B's statement actually block on?** B's insert crosses the diff --git a/docs/stories/language-runtime-database/21-keypair-attach-auth.md b/docs/stories/databasev2/10-keypair-attach-auth.md similarity index 94% rename from docs/stories/language-runtime-database/21-keypair-attach-auth.md rename to docs/stories/databasev2/10-keypair-attach-auth.md index 90f0d7b..b3bdc78 100644 --- a/docs/stories/language-runtime-database/21-keypair-attach-auth.md +++ b/docs/stories/databasev2/10-keypair-attach-auth.md @@ -1,12 +1,18 @@ --- -iteration: "21" +track: databasev2 +iteration: "10" +was_language_iteration: "21" status: hold --- -# Iteration 21 — keypair authentication for cross-program attach +# databasev2 10 — keypair authentication for cross-program attach + +> **Moved 2026-08-26** from the language track, where this was iteration 21. +> Part of [Story — the database beyond RAM](../language-runtime-database/00-story.md). Content unchanged by +> the move; its dependencies are restated in that track index. > Format: `product/story-iteration-template`. Part of -> [Story — one language, one runtime, one database, one binary](00-story.md). +> [Story — one language, one runtime, one database, one binary](../language-runtime-database/00-story.md). > > **Inserted 2026-08-15.** Promotes iteration 20's identity fork (Info, > fork 3) to its own iteration: the name + unix-uid lean is the milestone diff --git a/docs/stories/language-runtime-database/00-story.md b/docs/stories/language-runtime-database/00-story.md index 56d55a4..383e493 100644 --- a/docs/stories/language-runtime-database/00-story.md +++ b/docs/stories/language-runtime-database/00-story.md @@ -105,22 +105,17 @@ still pending IS the runtime-concurrency chain; order: | 16 | 19 | [Float + Bytes](19-missing-scalar-types.md) | **LANDED 2026-08-20** — `.wob` v5; the full stack: IEEE-quiet f64 through literals/VM/@table/WAL/json + Bytes as the binary carrier, no implicit mixing, total-order indexes. Unblocks 24 (WS frames) and the crypto fork (digests). *(was 20)* | | 17 | 31 | [Actor lifecycle](31-actor-lifecycle.md) | request/response (today `send` is one-way and callers `sleep` to await), bounded mailboxes with backpressure (today the FIFO just grows), actor death/supervision, and timers beyond `time.sleep`. 24 cannot be written honestly without these. *(story written 2026-08-21)* | | 18 | 24 | [chat: WebSocket workload](24-chat-websocket-workload.md) | the arc's acceptance: WS upgrade + frames (SHA-1 via crypto fork, Bytes via 19), rooms/broadcast, 1k clients, drain-clean. *(was 19)* | -| 19 | 23 | [io_uring group-commit](23-io-uring-commit.md) | WAL WRITE+FSYNC chains on the arc's per-shard rings; fsync fallback kept (after 22 + the arc). *(was 9f)* | -| 20 | 32 | [WAL checkpoint](32-wal-checkpoint.md) | **NEW 2026-08-21** (stage-3 guarantee refinement found the hole) — the WAL is append-only forever: snapshot + truncate reclaims disk and bounds replay time; every durability guarantee byte-identical; crash mid-checkpoint recovers from the previous snapshot + full tail. After 23 (composes with group-commit); RAM slot-reuse already contracted in `04-db-binding.md`. | -| 21 | 33 | [Single-file store](33-single-file-db.md) | **NEW 2026-08-22** — `WO_DATA=.db`: a file path IS the wal (the store already lives in exactly one file; this makes the surface say so). Driver-only, independent of the chain; composes with 32's rename-swap. | | 22 | 34 | [Crypto builtins](34-crypto-builtins.md) | **NEW 2026-08-22** — SHA-1/SHA-256/HMAC-SHA256 as C builtins over Bytes (no bitwise ops in the language, hand-rolled per doctrine, vector-verified). GATES 24's WS handshake; digest floor for held 21 and the ETag row. | | 23 | 37 | [wo-html components](37-wo-html-components.md) | **NEW 2026-08-23** — an MVC-shaped view layer in the wo-html LIBRARY (framework stays micro): structural `Component` interface (`render() -> Text`), layout components with slots, the site sample migrated as acceptance. Angular's component FORMAT studied and translated to server-rendered no-JS `.wo`; DI/bindings rejected. **LANDED 2026-08-25.** Raw text literal 2026-08-24 (backtick, verbatim content, margin stripped at lex time, `${ }` raw / `{{ }}` auto-escaping; WO-E004/WO-E005) — lexer plus one parser desugar, nothing downstream. Component layer 2026-08-25: `Component`/`render_all`/`Layout` in wo-html, `ok_html` moved into the framework, site and shop both migrated. | | 23 | 35 | [net runtime seams](35-net-runtime-seams.md) | **NEW 2026-08-22** — the ledger's three 🔧 rows owned: fd deadlines composing with the park plane, Unix-socket listeners, peer address (trusted-proxy check). Framework knobs stay framework slices; pairs naturally with 24 (dead-client eviction). | | 23 | 36 | [operator parity](36-operator-parity.md) | **NEW 2026-08-22** — `not`, the five bitwise operators (`& \| ^ << >>`, Int-only, riding the additive/multiplicative rungs Go-style), hex/binary/`_` literals, and compound assigns wired to the written-out form. **Code landed 2026-08-22** on branch `operator-parity`: `.wob` v6, opcodes 42–46 with the 0..63 shift trap (WO-E223), all gates green; reference project `.dev/reference/go` drove the operator-precedence design. Awaiting the developer's MANUAL pass over `docs/examples/operators/` (no test fixtures, by directive). Unblocks story 34's pure-`.wo` HMAC question. | | 24 | 25 | [HTTP service layer](../../superpowers/plans/2026-08-01-http-service-layer.md) | `service` blocks lower onto the framework (after 9b + 20 by their own precedence notes). **HELD 2026-08-21** — story file removed; the plan doc remains. *(was 10)* | | 25 | 18 | [framework v2: memory-rich features](18-memory-db-features.md) | spec+plan approved: TTL cache, @table flags, durable job queue, `transaction { }` over the WAL's staged batch. **Demoted from seq 14**: more surface on a framework with one consumer, and the cache still stores `Text` because there are no generics | -| 26 | 27 | [Query grammar corpus](27-query-grammar-corpus.md) | grow the query grammar from real corpora; likely collapses to "confirm `len(query)` + add `exists`"; precedes 28. *(was 9g)* | | 27 | 26 | [Blue-green deploy](26-blue-green-deploy.md) | two VM slots, in-runtime compile, atomic switch, resident rollback (plan authored after 9 + 25). *(was 12)* | -| 28 | 20 | [Cross-program tables](20-cross-program-tables.md) | attach to a running program's database over local IPC; owner stays the single writer (channel half-built). **Demoted from seq 16**: new distribution surface while there is no TLS, no crypto, and the multi-shard DB still traps. *(was 9c)* | -| 29 | 21 | [Keypair attach auth](21-keypair-attach-auth.md) | program identity is a keypair; mutual challenge–response at attach (crypto half-built; plan folds into 20's). **Demoted with 20** — and it needs crypto primitives that do not exist. *(was 9d)* | | 30 | 28 | [skillhost host workload](28-skillhost-host-workload.md) | host-shaped driving workload naming runtime gaps — demoted with the framework goal. *(was 14)* | | 31 | 29 | [Compile-time metaprogramming](29-compile-time-metaprogramming.md) | `@derive(...)` from class-table metadata; held with the parked drain by the 2026-08-08 scope directive. *(was 13)* | | 32 | 38 | [Content platform capabilities](38-content-platform-capabilities.md) | **NEW 2026-08-26** — the two capability families nothing owns, read off `wob.h`: `fs` mutation (`write`/`remove`/`rename`/`mkdir` — the table holds exactly six fs builtins, ids 40–45, where `append` creates-if-absent and grows, so a file is never replaced, truncated, deleted or renamed) and `net.connect` (ids 51–55 + 91–95, no connect; no `connect()` call in `runtime/src/` at all — which is every identity/notification/federation story at once; 35 deferred connect-side timeouts as "no workload asks yet" — this is the ask). Driven by a `docs/examples/vault` content-collaboration workload in 28's mould; WebDAV, CRDT editing, previews and FTS each get a written verdict instead of an implication. New builtins start at 96 (89/90 are 31's reserved holes); no `.wob` bump. Off-chain, wants a spec. | +| — | — | **[▸ the `databasev2` track](../databasev2/00-story.md)** | **Six pending database iterations moved out 2026-08-26** — WAL checkpoint *(was 32)*, io_uring group commit *(was 23)*, single-file store *(was 33)*, query grammar *(was 27)*, cross-program tables *(was 20)*, keypair attach auth *(was 21)* — renumbered 1–10 in that track alongside four new ones: the RAM-ceiling measurement, `@table` storage modes (the grammar that makes durability per-table instead of one global `WO_DATA`), bounded tables with eviction, and cold tiering. **Done database work stays here as v1 history:** 9 (engine), 9b (`@table`/relations/query) and 22 (durability baseline) are rows above and did not move. | | 33 | 39 | [Web framework parity](39-web-framework-parity.md) | **NEW 2026-08-26** — gofiber/fiber v3.5.0 added as the web-framework reference and read end to end ([the study](../../plan/exploration/fiber/00-fiber-parity.md)); nine of its 32 middleware already have a `porch` counterpart, so the gaps are breadth, not foundations — with one exception. The study's sharpest finding: the framework's own ledger said CSRF and sessions were UNBLOCKED because iteration 34 landed HMAC, but **there is no source of randomness in the runtime at all**, and an HMAC over a guessable session id is a signed guess. So 39 leads with a random-bytes builtin, then cookies (absent in both directions; `Resp.headers: map` structurally cannot emit two `Set-Cookie` lines), then the store-backed chain — limiter and idempotency first since they need only `@table` + `time.ticks`. Off-chain, wants a spec. | | ✅ | 17 | [library projects + `internal/`](17-library-projects-internal.md) | **LANDED 2026-08-20** — `kind = "library"` + entry-less check mode (retires the `--emit` workaround) and Go's `internal/` rule as WO-E108 at the consumer's `use`; driver-only, VM/GC untouched. `just web-app` 26/0 | diff --git a/docs/stories/language-runtime-database/08-shard-actor-runtime.md b/docs/stories/language-runtime-database/08-shard-actor-runtime.md index e9a3fa5..36d1358 100644 --- a/docs/stories/language-runtime-database/08-shard-actor-runtime.md +++ b/docs/stories/language-runtime-database/08-shard-actor-runtime.md @@ -163,7 +163,7 @@ the slice's marker doc when it landed): | Durability | ✅ fsync-per-commit, ack-after-durable; the ack crosses shards only AFTER the owner's fsync (`just db-actor`'s WAL pair). Power-loss rides fdatasync semantics; 22's kill battery is the scripted proof. | | Crash recovery | ✅ boot replay, torn-tail drop, index rebuild; replay completes on the primary before any worker serves (main.c boots the engine before the shards). 22 scripts the restart proof. | | Concurrency control | ✅ stage 3 — the DB actor serializes every statement; replies are materialized copies, no torn read by construction. Cross-statement snapshots arrive with 18. | -| Space reclamation | RAM ✅ (deleted rows free their slot — ids never reused, slots are); disk ✖ → [story 32](32-wal-checkpoint.md), end of chain. | +| Space reclamation | RAM ✅ (deleted rows free their slot — ids never reused, slots are); disk ✖ → [databasev2 story 3](../databasev2/03-wal-checkpoint.md), end of chain. | - The C proving ground (`docs/plan/exploration/c-runtime/`, phases A–F: epoll loops, eventfd mail) is the substrate this lifts into `wovm`. @@ -174,7 +174,7 @@ the slice's marker doc when it landed): - **Gated by the benchmark:** landing the arc means re-running [22](22-durability-throughput-scale.md) at the concurrency scale it unlocks and recording the before/after delta; it is also - where [23](23-io-uring-commit.md) gets a thread to overlap + where [4](../databasev2/04-io-uring-commit.md) gets a thread to overlap durability against. ## Proposed Solution diff --git a/docs/stories/language-runtime-database/34-crypto-builtins.md b/docs/stories/language-runtime-database/34-crypto-builtins.md index 8aa5b3c..2a8b32a 100644 --- a/docs/stories/language-runtime-database/34-crypto-builtins.md +++ b/docs/stories/language-runtime-database/34-crypto-builtins.md @@ -25,7 +25,7 @@ Four consumers already wait on it, none able to proceed: (`Sec-WebSocket-Accept` = base64(SHA-1(key + GUID)) — SHA-1 specifically, not a choice); the framework's ETag/conditional-request row (wants a content hash); HMAC-signed tokens the auth core can grow; -and held [iteration 21](21-keypair-attach-auth.md), whose +and held [databasev2 10](../databasev2/10-keypair-attach-auth.md), whose challenge–response needs primitives that "do not exist" (its demotion note). Bytes and base64 landed with iteration 19 — the carriers exist, only the digests are missing. diff --git a/docs/stories/language-runtime-database/38-content-platform-capabilities.md b/docs/stories/language-runtime-database/38-content-platform-capabilities.md index 95aa4f7..6222b20 100644 --- a/docs/stories/language-runtime-database/38-content-platform-capabilities.md +++ b/docs/stories/language-runtime-database/38-content-platform-capabilities.md @@ -93,7 +93,7 @@ status: refine *does* depend on 28's killable-subprocess work for previews, which is exactly why previews are deferred here rather than attempted. - **WAL checkpoint and disk reclamation** — - [iteration 32](32-wal-checkpoint.md). A metadata store whose boot replays + [databasev2 3](../databasev2/03-wal-checkpoint.md). A metadata store whose boot replays every write ever made is a real ceiling for this workload, and naming it here is the point; fixing it is 32's. This story's gate should record the replay time it observes so 32 inherits a number. @@ -109,7 +109,7 @@ status: refine route Nextcloud itself takes is unreachable until `connect` lands. - **Full-text search.** The engine indexes equality probes on declared columns; there is no prefix scan or FTS. Query-grammar growth is - [iteration 27](27-query-grammar-corpus.md)'s. + [databasev2 8](../databasev2/08-query-grammar-corpus.md)'s. - **TLS** — proxy-terminated, by doctrine, unchanged. The outbound half therefore speaks plaintext to a local sidecar or a trusted-network peer, and the story says so out loud rather than implying HTTPS clients. @@ -196,9 +196,9 @@ a wish list. Order: brainstorm the five forks (fork 1 and 2 together, fork establish exactly where it blocks, then the fs verbs, then `connect`, each with corpus fixtures and its own gate leg. Big enough to want a spec and a plan document — this is not a bounded slice like -[33](33-single-file-db.md). +[7](../databasev2/07-single-file-db.md). Precedence: independent of the concurrency chain (stage 3 → 22 → 31 → 24 → 23 → 32) and startable beside it, with one caveat — the honest disk story -needs [32](32-wal-checkpoint.md), so if this lands first its gate records +needs [3](../databasev2/03-wal-checkpoint.md), so if this lands first its gate records the replay number rather than claiming the platform is operationally done. diff --git a/docs/stories/porch/01-store-backed-middleware.md b/docs/stories/porch/01-store-backed-middleware.md index c743a31..6e719f3 100644 --- a/docs/stories/porch/01-store-backed-middleware.md +++ b/docs/stories/porch/01-store-backed-middleware.md @@ -115,7 +115,7 @@ status: refine it. - **Distributed limiting across processes.** One program owns its database; cross-program state is language - [iteration 20](../language-runtime-database/20-cross-program-tables.md). + [databasev2 9](../databasev2/09-cross-program-tables.md). - **A background expiry sweeper.** No timer exists (`time.after` is still a reserved builtin id in `wob.h`). Lazy pruning on access, deliberately. - **The TTL cache middleware** — language