docs(databasev2): third track — the database beyond RAM, with per-table storage modes

- docs/stories/databasev2/, numbered from 1. Six PENDING database iterations
  moved from the language track and renumbered, keeping the old id in
  `was_language_iteration:` so a search for "iteration 32" still finds it:
  32 -> 3 WAL checkpoint, 23 -> 4 io_uring commit, 33 -> 7 single-file store,
  27 -> 8 query grammar, 20 -> 9 cross-program, 21 -> 10 keypair auth.
  Done work (9, 9b, 22) stays as v1 history; language 18 left whole
- the problem, read off the engine not guessed: rows are malloc'd slabs with
  addresses stable forever, NO eviction/spill/paging anywhere in database/src,
  the WAL never checkpoints so boot replays all history, and durability is one
  process-global WO_DATA so no table can say it matters more than another.
  An allocation failure IS a clean catchable WO_T_OOM — but swap thrash
  arrives first and carries no error signal at all, which is the real hazard
- four new iterations:
  1 measure the ceiling FIRST (curve not cliff; the three exits; kill -9 at
    exhaustion) — every later default should follow from a number
  2 `@table(mode: ram | durable | cold)` — the grammar ask. Small surface
    (Ast.table_cfg gains a key, the parser already rejects unknown args), big
    semantics: `durable` defaults so nothing changes silently, and the
    compiler refuses a durable row holding a `ref` into a ram table
  5 bounded tables + refuse/evict/back-pressure, shedding BEFORE the OS acts
  6 cold tiering — mostly forks, incl. whether the language surfaces the
    fault cost and whether @unique on cold is refused outright. A paged
    B-tree stays rejected: if tiering needs one, reject tiering
- 39 links repointed, link TEXT renumbered to track-local ids; arc gains one
  pointer row replacing the six moved; board + board-views cover three tracks
- linkcheck 0 broken / 0 anchors; no code blocks in any story

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
shoney.arickathil 2026-08-26 20:52:48 +02:00
parent 01df75245f
commit 746dc2b42b
21 changed files with 970 additions and 56 deletions

View file

@ -2,8 +2,8 @@
> **Status: target workload — does not compile on today's toolchain.**
> Written ahead of iterations
> [20 (cross-program tables)](../../stories/language-runtime-database/20-cross-program-tables.md)
> and [21 (keypair attach auth)](../../stories/language-runtime-database/21-keypair-attach-auth.md),
> [9 (cross-program tables)](../../stories/databasev2/09-cross-program-tables.md)
> and [10 (keypair attach auth)](../../stories/databasev2/10-keypair-attach-auth.md),
> the way every acceptance sample here precedes its features. It also leans
> on 9/9b (the [employee sample](../employee/) it attaches to must run
> first).

View file

@ -72,7 +72,7 @@ goes to find what took each one's place.
| --- | --- |
| `09-concurrency-scaleout.md` | [`08-shard-actor-runtime.md`](../stories/language-runtime-database/08-shard-actor-runtime.md) + [`11-fibers.md`](../stories/language-runtime-database/11-fibers.md) — the arc, landed 2026-08-21 |
| `10-storage-foundations.md`, `11-wal-and-recovery.md` | [`09-database-engine.md`](../stories/language-runtime-database/09-database-engine.md) (typed WAL + replay) and [`22-durability-throughput-scale.md`](../stories/language-runtime-database/22-durability-throughput-scale.md) (the measurements) |
| `12-engine-disk-cutover.md` | Nothing — RAM stays authoritative by doctrine (principle 7). The disk story is the WAL; reclamation is [`32-wal-checkpoint.md`](../stories/language-runtime-database/32-wal-checkpoint.md) |
| `12-engine-disk-cutover.md` | Nothing — RAM stays authoritative by doctrine (principle 7). The disk story is the WAL; reclamation is [`databasev2 3, WAL checkpoint`](../stories/databasev2/03-wal-checkpoint.md) |
| `13-class-model-live-pricing.md` | [`09b-table-relations-query.md`](../stories/language-runtime-database/09b-table-relations-query.md) — `@table`, `ref`/`backlink`, the compiler-checked query surface |
| `07-inotify-content-watcher.md` | [`07-logwatcher-proof.md`](../stories/language-runtime-database/07-logwatcher-proof.md) — the log-watcher sample polls via `fs.stat`; inotify was never surfaced as a builtin |
| `08-sendfile-static-assets.md` | Nothing. `sendfile` is not exposed; static assets are served as `Text` through `net.write` |

View file

@ -46,7 +46,7 @@ The implementation phases that lean on this material:
- The Rust-era consumers (plans 10/11/12: storage foundations, WAL and
recovery, disk cutover) were removed with that track 2026-08-18; their
ideas shipped in `database/src/` (typed WAL + replay) and the rest wait
on [iteration 32](../../../stories/language-runtime-database/32-wal-checkpoint.md)
on [databasev2 3](../../../stories/databasev2/03-wal-checkpoint.md)
(checkpoint) — the wal/buffer cards are its entry material.
- Current consumers: [constraints-and-grammar](constraints-and-grammar.md)
(the `@table` PK/FK grammar direction) and

View file

@ -10,11 +10,14 @@ The single place to learn where this project stands. Organised in six buckets:
folders** — a doc stays where it was authored, and only its frontmatter, its
banner and this board change.
**Two tracks** (2026-08-26): [`language-runtime-database/`](language-runtime-database/00-story.md)
— the language, runtime and database — and [`porch/`](porch/00-story.md), the web
framework written in it. Each numbers its iterations from 1, so a porch 3 is not
a language 3; porch stories carry `track: porch` in frontmatter to keep queries
honest. Track folders are fine; **status** folders are not.
**Three tracks** (2026-08-26):
[`language-runtime-database/`](language-runtime-database/00-story.md) — the
language and runtime; [`porch/`](porch/00-story.md) — the web framework written
in it; [`databasev2/`](databasev2/00-story.md) — the database beyond RAM. Each
numbers its iterations from 1, so a porch 3 is not a language 3; every non-language
story carries `track:` in frontmatter, and moved ones keep
`was_language_iteration:` so a search for the old number still finds them. Track
folders are fine; **status** folders are not.
**Status lives in frontmatter, nowhere else** (directive 2026-08-26). Every
story iteration file sits flat in its track folder and
@ -348,19 +351,19 @@ that sequences its tasks. Read one, approve, then the next starts.
| 22 | [Durability, throughput, scale](language-runtime-database/22-durability-throughput-scale.md) | ✅ **landed 2026-08-21** — db-bench + baseline.json (74 metrics) + restart/kill -9 proofs both shard counts; durable 4.5k vs ram 297k inserts/s, reads O(table), msgrate 13.4M/2.45M |
| 31 | [Actor lifecycle](language-runtime-database/31-actor-lifecycle.md) | 🔄 **absorbed into 24** (directive 2026-08-23) and half landed there: `call` request/response with a typed scalar reply (`WO_B_CALL = 88`, WO-E226), bounded mailboxes (`WO_MAILBOX`, cap 1024, catchable `WO_T_ACTOR`), and actor death that traps callers instead of hanging them. Still open: `monitor` and `time.after` — ids **89 and 90 are reserved holes** in `wob.h`, which is the machine-checkable proof of what is left. Supervision trees stay out of v1 |
| 24 | [chat: WebSocket workload](language-runtime-database/24-chat-websocket-workload.md) | 🔄 **the live slice** (absorbing 31 + 34, directive 2026-08-23) — branch `chat-ws-lifecycle`, 5/10 tasks landed: crypto, bounded mailboxes, WS upgrade, frame codec, `call`/reply + actor death. Pending: `monitor`, `time.after`, the chat sample, its gate, closeout. State lives in [the marker](../active-slice-2026-08-23-chat-ws-lifecycle.md) |
| 23 | [io_uring group-commit](language-runtime-database/23-io-uring-commit.md) | ⬜ fifth in chain, after stage 3 + 22 |
| 32 | [WAL checkpoint](language-runtime-database/32-wal-checkpoint.md) | ⬜ last in chain, after 23 — disk reclamation + bounded replay (story written 2026-08-21) |
| 33 | [Single-file store](language-runtime-database/33-single-file-db.md) | ⬜ off-chain, small — `WO_DATA=<path>.db` file form; driver-only (story written 2026-08-22) |
| 23 | [io_uring group-commit](databasev2/04-io-uring-commit.md) | ⬜ fifth in chain, after stage 3 + 22 |
| 32 | [WAL checkpoint](databasev2/03-wal-checkpoint.md) | ⬜ last in chain, after 23 — disk reclamation + bounded replay (story written 2026-08-21) |
| 33 | [Single-file store](databasev2/07-single-file-db.md) | ⬜ off-chain, small — `WO_DATA=<path>.db` file form; driver-only (story written 2026-08-22) |
| 34 | [Crypto builtins](language-runtime-database/34-crypto-builtins.md) | 🔄 **code landed** as 24's T1 (`d14fa9f`): `sha1`/`sha256`/`hmac_sha256`, ids 85–87 in `wob.h`, `runtime/src/crypto.c`, RFC/FIPS vectors 18/0, corpus pin. The 24 gate that once needed it is cleared. Frontmatter keeps `status: refine` only until 24's T10 closeout sets it to `done` |
| 38 | [Content platform capabilities](language-runtime-database/38-content-platform-capabilities.md) | ⬜ off-chain, needs a spec — the two capability families no iteration owns, confirmed against `runtime/src/wob.h`: `fs` mutation verbs (six fs builtins, ids 40–45; `append` creates-if-absent, so nothing is ever replaced, truncated, deleted or renamed) and `net.connect` (ids 51–55 + 91–95, no connect, and no `connect()` anywhere in `runtime/src/` — so no OIDC/SMTP/object-store/webhook/federation). Driven by a `docs/examples/vault` content-collaboration workload, in 28's mould. New builtins from 96 (89/90 reserved for 31); no `.wob` bump (`WOB_VERSION 6u`, last moved by 36). Story written 2026-08-26 from the "can it build a Nextcloud?" ask |
| 39 | [Web framework parity](language-runtime-database/39-web-framework-parity.md) | ⬜ off-chain, needs a spec — from [the Fiber v3.5.0 study](../plan/exploration/fiber/00-fiber-parity.md) (all 32 of its middleware read against `porch`; **nine already have a counterpart**). Leads with a **random-bytes builtin**: the framework ledger claimed CSRF/sessions were unblocked by iteration 34's HMAC, but HMAC authenticates a token and cannot mint one — there is no RNG anywhere in the runtime. Then cookies (absent both ways; `Resp.headers` being a map cannot carry two `Set-Cookie` lines), then limiter/idempotency (cheapest wins — `@table` + `time.ticks`, nothing new), sessions, CSRF, and the routing/response sugar. Streaming/SSE/compression, `@derive` binding, TTL cache, `proxy` and metrics all excluded with owners named |
| 37 | [wo-html components](language-runtime-database/37-wo-html-components.md) | ✅ off-chain — LANDED 2026-08-25. Raw text literal (backtick, margin stripped at lex time, `{{ }}` auto-escapes) + the component layer: `Component`/`render_all`/`Layout` in wo-html, `ok_html` moved into the framework, site and shop both migrated |
| 35 | [net runtime seams](language-runtime-database/35-net-runtime-seams.md) | ⬜ off-chain — fd deadlines on the park plane, Unix sockets, peer address; owns the ledger's three 🔧 rows (story written 2026-08-22) |
| 20 | [Cross-program tables](language-runtime-database/20-cross-program-tables.md) | ⏸ hold (2026-08-21); channel done (branch ipc-attach keeps its manifest) |
| 21 | [Keypair attach auth](language-runtime-database/21-keypair-attach-auth.md) | ⏸ hold (2026-08-21); crypto+handshake done (branch keypair-auth keeps its manifest) |
| 20 | [Cross-program tables](databasev2/09-cross-program-tables.md) | ⏸ hold (2026-08-21); channel done (branch ipc-attach keeps its manifest) |
| 21 | [Keypair attach auth](databasev2/10-keypair-attach-auth.md) | ⏸ hold (2026-08-21); crypto+handshake done (branch keypair-auth keeps its manifest) |
| 25 | [HTTP service layer](../superpowers/plans/2026-08-01-http-service-layer.md) | ⏸ hold (2026-08-21) — story file removed; the plan doc remains |
| 26 | [Blue-green deploy](language-runtime-database/26-blue-green-deploy.md) | ⏸ hold (2026-08-21) |
| 27 | [Query grammar corpus](language-runtime-database/27-query-grammar-corpus.md) | ⏸ hold (2026-08-21) |
| 27 | [Query grammar corpus](databasev2/08-query-grammar-corpus.md) | ⏸ hold (2026-08-21) |
| 28 | [skillhost host workload](language-runtime-database/28-skillhost-host-workload.md) | ⏸ hold (2026-08-21); gaps recorded (branch query-grammar found skillhost needs no new query grammar) |
| 29 | [Compile-time metaprogramming](language-runtime-database/29-compile-time-metaprogramming.md) | ⏸ hold (2026-08-21) |
| 15 | [deps: `wo.toml [deps]`](language-runtime-database/15-deps-package-manager.md) | ✅ **landed 2026-08-18** (branch web-framework): [deps] inline tables, git-binary fetch, wo.lock pinning, offline-when-locked, --update-deps, WO-E106/E107; `just deps-accept` 8/0 |
@ -591,7 +594,7 @@ precedence notes for resumption.
5. **23** — io_uring group-commit; the WAL's WRITE+FSYNC chains ride the
arc's per-shard ring (T4); after 22's baseline — the payoff, measured.
6. **32** — WAL checkpoint
([story](language-runtime-database/32-wal-checkpoint.md),
([story](databasev2/03-wal-checkpoint.md),
written 2026-08-21): the WAL is append-only forever — snapshot +
truncate reclaims disk and bounds replay; after 23 (composes with
group-commit), policy set by 22's aged-store numbers.
@ -599,6 +602,36 @@ precedence notes for resumption.
**30** — observability, CI, fuzz: named 2026-08-20, still row-only (no
story file); slots in when scheduled — nothing in the chain depends on it.
### ▸ databasev2 — the database beyond RAM
New 2026-08-26. **The problem:** RAM is authoritative (principle 7) and nothing
declares a budget. Rows live in `malloc`'d slabs whose addresses are stable
forever; there is no eviction, spill or paging anywhere in `database/src/`; the
WAL never checkpoints so boot replays all history; and durability is one
process-global `WO_DATA`, so no table can say it matters more than another. An
allocation failure *is* a clean catchable `WO_T_OOM` — but swap thrash arrives
first and carries no error signal at all.
**The lever** is per-table storage modes, which is why this track has a grammar
iteration. Six pending iterations moved here from the language track (their old
ids in the rows below); four are new. Done database work — 9, 9b, 22 — stays in
the language arc as v1 history.
| # | Iteration | State |
| --- | --- | --- |
| 1 | [RAM ceiling: measure the breaking point](databasev2/01-ram-ceiling-measurement.md) | ⬜ **first, and startable today** — nobody here can say what happens at 90% RAM. Curve not cliff: swap onset, latency departure, the three exits (checked trap / swap thrash / OOM killer), and `kill -9` durability *at exhaustion*. Output is `perf-targets.md` + baseline rows, not prose |
| 2 | [`@table` storage modes](databasev2/02-table-storage-modes.md) | ⬜ **the language enrichment** — `mode: ram \| durable \| cold` per table, replacing the global switch. `durable` defaults so nothing changes silently; the compiler refuses a `durable` row holding a `ref` into a `ram` table. `.wob` format change. Grammar is small (`Ast.table_cfg` gains a key); semantics are the iteration |
| 3 | [WAL checkpoint](databasev2/03-wal-checkpoint.md) *(was 32)* | ⬜ snapshot + truncate: disk reclaimed, replay bounded |
| 4 | [io_uring group commit](databasev2/04-io-uring-commit.md) *(was 23)* | ⬜ close the 66× gap iteration 22 measured (durable 4.5k vs ram 297k inserts/s) |
| 5 | [Bounded tables and eviction](databasev2/05-bounded-tables-eviction.md) | ⬜ a declared capacity + refuse/evict/back-pressure, and a process-level pressure signal that sheds **before** the allocator or OS gets involved — turning the invisible failure into a managed one |
| 6 | [Cold tiering](databasev2/06-cold-tiering.md) | ⬜ the iteration that raises the ceiling, and the riskiest. Mostly forks: which shape, whether the index itself fits, whether the *language* surfaces the fault cost, and whether `@unique` on a cold table is refused outright. A paged B-tree stays rejected — if tiering needs one, reject tiering |
| 7 | [Single-file store](databasev2/07-single-file-db.md) *(was 33)* | ⬜ `WO_DATA=<path>.db`; driver-only, independent |
| 8 | [Query grammar from corpora](databasev2/08-query-grammar-corpus.md) *(was 27)* | ⬜ whole-query `count`, `exists`; independent |
| 9 | [Cross-program tables](databasev2/09-cross-program-tables.md) *(was 20)* | ⏸ hold — attach to a running program's database over local IPC |
| 10 | [Keypair attach auth](databasev2/10-keypair-attach-auth.md) *(was 21)* | ⏸ hold — program identity as a keypair; needs 9 |
---
### ▸ porch — the web framework track
New 2026-08-26, from [the Fiber v3.5.0 parity study](../plan/exploration/fiber/00-fiber-parity.md).

View file

@ -28,12 +28,13 @@ standup narrative; these queries are the live views over the same facts.
Adjust the `FROM` path to your vault root (queries below assume the
vault opens at the repo root).
Two tracks now carry iterations, each numbered from 1:
`language-runtime-database/` (the language, runtime and database) and `porch/`
(the web framework, added 2026-08-26). Iteration ids therefore repeat across
tracks — a porch 3 is not a language 3 — so every query below is scoped by
`FROM` path, and porch stories carry `track: porch` so a combined query can
still tell them apart.
Three tracks now carry iterations, each numbered from 1:
`language-runtime-database/` (the language and runtime), `porch/` (the web
framework) and `databasev2/` (the database beyond RAM) — the latter two added
2026-08-26. Iteration ids therefore repeat across tracks, so every query below is
scoped by `FROM` path; non-language stories carry `track:`, and iterations moved
between tracks keep `was_language_iteration:` so the old number stays
searchable.
## Everything not done, chain order first
@ -79,6 +80,15 @@ the place status is edited.** A status change is one edit to one
`status:` key; a Kanban card drag that only rewrites the Kanban file is a
lie the next query won't see.
## The databasev2 track
```dataview
TABLE iteration, status, was_language_iteration AS "was"
FROM "docs/stories/databasev2"
WHERE status != "done"
SORT iteration ASC
```
## The porch track
```dataview

View file

@ -0,0 +1,152 @@
# Story — databasev2: the database beyond RAM
The third track. `language-runtime-database/` built the engine;
[`porch/`](../porch/00-story.md) is the framework on top; this track answers the
question v1 deliberately deferred: **what happens when the data does not fit in
memory.**
Numbering restarts at 1, local to this track. Frontmatter carries
`track: databasev2`, and iterations moved here keep their old id in
`was_language_iteration:` so a search for "iteration 32" still finds the WAL
checkpoint. Status stays where it belongs — the `status:` key, never a directory.
## The problem, stated honestly
Principle 7 says **RAM is authoritative; the WAL makes it durable.** That is a
real design, not a shortcut: reads never touch disk, so latency is predictable,
and durability is a sequential append rather than a storage engine bolted to the
side. Iteration 22 measured what it buys — reads at 1.3M ops/s after the index
probe landed, p50 1µs.
The bill comes due at the ceiling. Read from the engine as it stands:
- **Rows live in `malloc`'d slabs of 256, and their addresses are stable
forever** (`database/src/table.c`, `DB_SLAB_ROWS`). Slabs are allocated as a
table grows and freed only when the table is destroyed. The free-slot list
recycles removed slots, so a delete-heavy table plateaus — but a growing table
only grows.
- **The ceiling is process RSS, not a configured number.** `WO_HEAP_MB` (default
64 MiB) bounds the VM object arena; table storage is separate `malloc`, so
nothing in the system declares a maximum dataset size. There is no knob that
says "this database may use at most N".
- **There is no eviction, no spill, no paging, no LRU.** Grep
`database/src/` for any of them and nothing comes back. Every row ever
inserted and not deleted is resident.
- **The WAL is append-only with no checkpoint.** Boot replays every record ever
written, so startup time is O(all writes in the file's history) and disk grows
without bound. That is databasev2 [3](03-wal-checkpoint.md).
- **Durability is process-global.** `WO_DATA` is one environment variable that
turns on one `shard-0.wal` for the whole process (`runtime/src/main.c`). There
is no way to say "this table matters, that one is scratch".
### What actually breaks first
Worth being precise, because the failure mode determines the fix — and the good
news is that the engine's own behaviour is clean:
**An allocation failure is a catchable trap, not a crash.** Every `malloc` in
the row encoder is checked and jumps to an `oom` label; `DB_ERR_OOM` maps to
`WO_T_OOM`, which a program can `try`/`catch`. So a writeonce program that runs
out of memory *refuses the insert* rather than corrupting or dying. That is a
much better starting position than most engines have.
**But the trap is almost never what a real deployment hits first.** Long before
`malloc` returns NULL, the box starts swapping, and a RAM-authoritative database
on swap is the worst of both worlds: it has paid for in-memory data structures
and is now serving them from disk with no read path designed for that. On a
cgroup-limited host the OOM killer arrives instead, and an external `SIGKILL` is
the one shutdown path that skips every guarantee the WAL was written to provide —
though ack-after-fsync means acked writes still survive; iteration 22's `kill -9`
battery proves that much.
So the honest problem statement is not "malloc fails". It is: **there is no
declared budget, no back-pressure as the budget is approached, and no way to
distinguish data that must be resident from data that merely is.** Iteration
[1](01-ram-ceiling-measurement.md) exists to replace this paragraph with
numbers before anything is designed on top of it.
## The lever: per-table storage modes
The developer's ask, and the reason this track has a grammar iteration.
Today every `@table` is identical: resident, and durable if and only if
`WO_DATA` is set for the whole process. Real applications are not uniform —
a session table, a rate-limit counter and a page cache want *resident and
disposable*; an orders table wants *resident and durable*; an audit log wants
*durable and rarely read*. One global switch cannot express that, so it forces
either "everything is precious" or "nothing is".
Extending `@table` with a storage mode moves the decision into the language,
where the compiler can act on it:
- **`ram`** — resident, never WAL-logged, gone on restart. The compiler knows
no durability code is needed; the engine knows these rows are the first
candidates to shed under pressure; and — the part that matters — a program
that expects a `ram` table to survive a restart is now stating something the
compiler can refuse.
- **`durable`** — today's behaviour: resident and WAL-logged, ack after fsync.
- **`cold`** — durable, and *not* required to be resident. This is the mode that
actually raises the ceiling, and it is the one with real design work behind it
(iteration [6](06-cold-tiering.md)).
The grammar change is small and the surface is already the right shape:
`Ast.table_cfg` is `{ table_name; indexes }`, the parser's argument match
already rejects unknown keys with a catalogued diagnostic
(`unknown @table argument ... (supported: name, index)`), and adding one more
key follows the path `index` already cut. The *semantics* are the work, not the
syntax — which is exactly why it gets its own iteration
([2](02-table-storage-modes.md)) and why it comes after the measurement.
This is also the honest answer to "does this break principle 7?" It does not.
RAM stays authoritative **for the tables that say so**. `cold` is a declared
exception a developer opts into per table, with the trade written at the
declaration site rather than buried in an operations runbook.
## The sequence
| # | Iteration | Delivers | Needs |
| --- | --- | --- | --- |
| 1 | [RAM ceiling: measure the breaking point](01-ram-ceiling-measurement.md) | what actually happens from 50% RAM to OOM — swap onset, latency cliff, trap behaviour, `kill -9` survival | nothing; extends iteration 22's harness |
| 2 | [`@table` storage modes](02-table-storage-modes.md) | the grammar: `mode: ram \| durable \| cold`, per table, replacing the global `WO_DATA` all-or-nothing | 1 for its defaults |
| 3 | [WAL checkpoint](03-wal-checkpoint.md) *(was language 32)* | snapshot + truncate: disk reclaimed, replay bounded | 4 composes |
| 4 | [io_uring group commit](04-io-uring-commit.md) *(was language 23)* | close the 66× durable/RAM write gap (4.5k vs 297k inserts/s) | the arc (landed) |
| 5 | [Bounded tables and eviction](05-bounded-tables-eviction.md) | a capacity a `ram` table may not exceed, and what happens when it does | 2 |
| 6 | [Cold tiering](06-cold-tiering.md) | rows that leave RAM and come back — the iteration that raises the ceiling | 2, 3, 5 |
| 7 | [Single-file store](07-single-file-db.md) *(was language 33)* | `WO_DATA=<path>.db` — a file path IS the store | independent |
| 8 | [Query grammar from corpora](08-query-grammar-corpus.md) *(was language 27)* | whole-query `count`, `exists` | independent |
| 9 | [Cross-program tables](09-cross-program-tables.md) *(was language 20)* | attach to a running program's database over local IPC | independent |
| 10 | [Keypair attach auth](10-keypair-attach-auth.md) *(was language 21)* | program identity as a keypair; mutual challenge–response | 9 |
```
1 ──▶ 2 ──▶ 5 ──▶ 6
│ ▲
3 ──▶ 4 ─────┘
7, 8 independent
9 ──▶ 10
```
Order rationale: **1 before 2** because a mode's default should follow from a
measurement, not a guess. **3 and 4 before 6** because tiering onto a log that
never truncates would make the disk problem worse, not better. **5 before 6**
because eviction from a bounded resident table is the simpler half of the same
mechanism, and getting the policy right there de-risks the hard half.
## What this track does NOT own
| Not databasev2's | Owner |
| --- | --- |
| `transaction { }` and `@table` feature flags | language [iteration 18](../language-runtime-database/18-memory-db-features.md) — approved spec, left whole on purpose |
| the TTL cache middleware | also language 18 (and [porch 1](../porch/01-store-backed-middleware.md) points there) |
| typed binding of rows into app classes | language [iteration 29 `@derive`](../language-runtime-database/29-compile-time-metaprogramming.md) |
| `fs` mutation verbs, outbound sockets | language [iteration 38](../language-runtime-database/38-content-platform-capabilities.md) |
| benchmark harness and CI | iteration 22 (landed) built the harness; per-change CI is language iteration 30 |
| a paged B-tree storage engine | **nobody, deliberately.** Recorded as rejected in [`discarded.md`](../../plan/discarded.md): the disk story is the WAL. `cold` tiering is not a licence to build SQLite. |
## Review protocol
The language track's, unchanged: one iteration read and approved before the next
starts; every iteration an unsplittable slice with phases, per-phase tasks,
Given/When/Then criteria and an out-of-scope list. Every engine change is gated
by `just employee`, `just db-actor` and `just db-bench` against
`bench/baseline.json` — and any iteration that claims a performance change must
move a number in that baseline, or it did not happen.

View file

@ -0,0 +1,148 @@
---
track: databasev2
iteration: "1"
status: refine
---
# databasev2 1 — the RAM ceiling: measure the breaking point before designing for it
> Part of [Story — databasev2: the database beyond RAM](00-story.md).
>
> **First because the repo's own doctrine says so.** "Always inspect crashsites.
> Always measure. Never assume." Every later iteration in this track — the
> storage modes' defaults, the eviction policy, the tiering threshold — is a
> decision that should follow from a number. Right now nobody in this project
> can say what happens to a writeonce program at 90% of RAM, and designing
> tiering without that is guessing with extra steps.
## Goals
- **Find the curve, not the cliff.** Not "does it die" — it dies, everything
does. What matters is the shape on the way down: at what fraction of RAM does
p99 read latency leave its 1µs baseline, what does insert throughput do as
slabs stop coming from a warm allocator, and how much warning is there between
"fine" and "unusable".
- **Characterise all three exits.** The engine can leave the happy path three
ways and they are not equally survivable: a checked `malloc` failure
(`DB_ERR_OOM` → `WO_T_OOM`, a catchable trap — the clean one), swap thrash
(no trap, no error, just latency collapse — the dangerous one because nothing
reports it), and the external OOM killer (`SIGKILL`, skipping every shutdown
path). Establish which arrives first under realistic limits, because the
answer determines whether the fix is back-pressure or eviction.
- **Prove the durability floor holds at the ceiling.** Iteration 22's `kill -9`
battery proved acked writes survive under load. Re-run it *at memory
exhaustion*, which is a different and nastier state — an allocation failure
mid-commit is exactly where an ack-before-durable bug would hide.
- **Publish numbers others can build on.** The output is a section in
`perf-targets.md` and rows in `bench/baseline.json`, not a paragraph of
prose. A measurement that only printed once is not a measurement.
## Phases
### Phase A — a workload that can actually reach the ceiling
- Extend `docs/examples/db-bench` with a growth mode: insert until a target RSS
fraction, holding row shape and index count constant so the variable is size
alone.
- Run it under an explicit memory limit (a cgroup or `ulimit`) rather than on a
big box — "it survived on a 64 GB workstation" measures the workstation.
- Record RSS against row count so the per-row overhead is known: slab headroom,
the id hash, the secondary-index multimaps and the per-row engine-owned values
(`db_text`, `db_rec`, `db_multi`, `db_map` are each their own allocation).
- Verify: RSS growth is linear and its slope is written down; the run is
reproducible twice within the tolerance policy iteration 22 established.
### Phase B — the latency and throughput curve
- Sample read p50/p99, query p99 and insert throughput at fixed fractions of the
limit, so the result is a curve rather than two endpoints.
- Separate the two effects deliberately: allocator pressure (still resident) and
swap (no longer resident). They have different fixes and conflating them would
send iteration 6 after the wrong one.
- Include the DB-actor path, since a cross-shard statement's reply materialises
a copy — memory pressure and the actor RPC interact and nobody has looked.
- Verify: the curve is recorded per metric class with iteration 22's per-class
tolerances; the swap onset point is identified, not interpolated.
### Phase C — the three exits, deliberately triggered
- Drive a checked allocation failure and confirm `WO_T_OOM` is catchable, the
insert is refused whole, no partial row or index entry is left, and the
process continues serving.
- Drive swap thrash and record what a client sees. This is the case with no
error signal at all, and naming it is most of the value of this iteration.
- Drive the OOM killer under a cgroup limit and confirm what survives: replay
the WAL and check every acked write is present.
- Verify: the trap path leaves no torn state (row count and index agree after a
refused insert); replay after `SIGKILL` at exhaustion loses no acked write.
### Phase D — write it down where decisions get made
- A `perf-targets.md` section with the curve, the swap onset, the per-row
overhead and the exit characterisation.
- Baseline rows for the growth metrics so a regression is caught by the existing
gate rather than by a person remembering.
- A short statement of what the numbers *imply* for iterations 2, 5 and 6 —
which is the point of going first.
- Verify: `just db-bench` green against the extended baseline; the gate bites
when a growth metric is doctored.
## Acceptance Criteria
- **Given** the growth workload under a fixed memory limit, **when** it runs
twice, **then** RSS-per-row agrees within the tolerance policy and the slope
is recorded in `perf-targets.md`.
- **Given** the workload at rising RAM fractions, **when** latency is sampled,
**then** the fraction at which read p99 first leaves its baseline is
identified as a measured point, not an estimate.
- **Given** a deliberately induced allocation failure, **when** an insert is
attempted, **then** it traps `WO_T_OOM` catchably, the table's row count is
unchanged, every index agrees with the slab contents, and the process keeps
serving subsequent requests.
- **Given** swap thrash, **when** a client issues reads, **then** the observed
degradation is quantified and the fact that **no error is surfaced** is
recorded explicitly as a finding.
- **Given** a cgroup limit and a workload that exceeds it, **when** the OOM
killer fires, **then** replaying the WAL shows every acked write present —
ack-after-fsync holding in the one shutdown path that skips all cleanup.
- **Given** the extended baseline, **when** a growth metric is doctored, **then**
`just db-bench` fails on exactly that metric.
## Out Of Scope
- **Any fix.** This iteration measures. Eviction is
[5](05-bounded-tables-eviction.md), tiering is [6](06-cold-tiering.md),
declared budgets are [2](02-table-storage-modes.md). Shipping a fix inside the
measurement slice would remove the ability to tell whether it helped.
- **Changing the OOM behaviour.** The checked-`malloc`-to-catchable-trap path is
good and should not be touched; if the measurement finds a hole in it, that is
a bug fix, reported separately.
- **A memory profiler or allocator instrumentation.** Observability is language
iteration 30. RSS from the OS and the existing `time.ticks` are enough for a
curve.
- **Multi-machine or sharded-across-hosts scaling.** One binary owns its data;
cross-process is [9](09-cross-program-tables.md).
- **Comparing against SQLite at the ceiling.** `bench/compare/go-sqlite` exists
and the comparison would be interesting, but SQLite's whole architecture is
the paged design this project rejected — the numbers would not inform any
decision here.
## Info
Forks the spec must settle:
1. **What is the limit mechanism for the harness?** A cgroup v2 `memory.max` is
the closest thing to how this would actually be deployed; `ulimit -v` is
simpler but bounds address space rather than resident set, which for an engine
that `malloc`s slabs is a materially different constraint. Leaning cgroup, and
the campaign already runs off the fast path so the setup cost is acceptable.
2. **Which table shape is the reference?** Per-row overhead depends heavily on
whether fields are scalars or heap values — a `Text` column is a separate
`db_text` allocation per row, so a text-heavy table and an Int-only table
will produce very different slopes. Probably both, reported separately,
because "bytes per row" is meaningless without saying which row.
3. **Is swap even in scope for the target deployment?** If the intended answer
is "run with swap off and let the OOM killer decide", the swap curve is
informational rather than load-bearing — but that stance should be stated in
the doctrine, not assumed. It also changes which exit iteration 5's
back-pressure is defending against.

View file

@ -0,0 +1,188 @@
---
track: databasev2
iteration: "2"
status: refine
---
# databasev2 2 — `@table` storage modes: durability becomes a language decision
> Part of [Story — databasev2: the database beyond RAM](00-story.md).
> Needs [1](01-ram-ceiling-measurement.md) — a mode's default should follow from
> a measurement, not a preference.
>
> **The developer's ask, and the language enrichment this track exists for.**
> Today durability is one environment variable for a whole process: `WO_DATA` is
> set and every `@table` is WAL-logged, or it is not and none are
> (`runtime/src/main.c`). Real applications are not uniform. A session table, a
> rate-limit counter and a page cache are resident and disposable; an orders
> table is resident and precious; an audit log is precious and rarely read. One
> global switch forces "everything is precious" or "nothing is", and the
> developer pays for the wrong one either way.
## Goals
- **Move the storage decision to the declaration site.** `@table(mode: ...)`,
chosen per table, in the source, where the person who knows what the data is
worth is already writing. An operations runbook is the wrong place for a fact
the compiler could hold.
- **Three modes, each earning its existence.** `ram` — resident, never logged,
gone on restart. `durable` — today's behaviour, resident and ack-after-fsync,
and the **default** so every existing program is byte-identical. `cold` —
durable and not required to be resident, declared here and *implemented* in
[6](06-cold-tiering.md), because a mode with no engine behind it is a promise.
- **Make the compiler enforce what the mode means.** This is the part that makes
it a language feature rather than a config key. A program that inserts into a
`ram` table and expects the row after a restart is stating a contradiction, and
the compiler is the right place to say so — at minimum for the cases it can
see statically, with the diagnostic catalogued like every other.
- **Let the engine act on it.** `ram` tables skip the WAL write entirely, which
is not merely a saving — it is the 66× gap iteration 22 measured (4.5k durable
vs 297k RAM inserts/s) becoming available per table instead of per process.
They also become the first candidates to shed under pressure, which is what
[5](05-bounded-tables-eviction.md) builds on.
- **Keep principle 7 intact and say why.** RAM stays authoritative for every
table that says so. `cold` is a declared, per-table exception a developer opts
into with the trade visible at the declaration.
## Phases
### Phase A — the grammar
- Extend the `@table` argument parser with `mode:`. The path is already cut:
the argument loop matches `name` and `index` and rejects anything else with a
catalogued diagnostic (`unknown @table argument ... (supported: name, index)`),
so this is one more arm plus an updated message.
- Extend `Ast.table_cfg` — today `{ table_name; indexes }` — with the mode, and
give it the default so every existing `@table` keeps its meaning.
- Reject the incoherent cases at parse time: `mode` given twice, an unknown mode
name. Both belong in the same diagnostic family as the existing `@table`
errors, and both go in the error catalog **in this change**, not later — that
catalog went stale once by exactly that omission.
- Verify: golden AST fixtures for each mode; `woc-test` green; every existing
`@table` in every sample parses unchanged.
### Phase B — the mode reaches the image and the engine
- Carry the mode through the class descriptor into the `.wob` image so the
runtime knows it without re-deriving anything. This is a format change, so it
moves `WOB_VERSION` and the format contract in the same commit as the code.
- The engine consults it at the one choke point that already exists: the
`INDEX HOOK` / WAL staging site in `wo_row_insert` / `wo_row_remove`, which
`database/src/CODE-LOGIC.md` names as the only place storage may be mutated.
A `ram` table stages nothing.
- Replay must skip records for tables that are now `ram` — a WAL written when a
table was `durable` and replayed after the source changed is a real
migration case, and silently resurrecting rows into a `ram` table would be
worse than refusing.
- Verify: a `ram` table's inserts produce no WAL growth (measured, not assumed);
a `durable` table is byte-identical to today; a mode change across a restart
is handled explicitly rather than by accident.
### Phase C — the compiler's enforcement
- Decide how far static checking goes (fork 2) and implement that much. The
floor: `WO_DATA` set with every table `ram` is a program that asked for a data
directory it will never write to — worth a warning at least.
- The interesting case is a `ram` table participating in a `ref`/`backlink`
relation with a `durable` one. A durable row holding a foreign key into a
table that evaporates on restart is a dangling reference by construction, and
FK-restrict cannot save it. This is the check most worth having, and it is
statically visible from the class table.
- Verify: corpus `compile-fail` fixtures for each refusal; each carries the
exact `WO-E###` the catalog now documents.
### Phase D — prove it on a real workload
- Give `docs/examples/employee` or the `db-bench` sample a mixed schema — at
least one `ram` table and one `durable` — and gate the distinction: after a
restart, the durable rows are present and the ram rows are gone. That single
assertion is the whole feature.
- Extend the durability legs of `just db-bench` so the per-table write-path
saving appears as a baseline number, not a claim.
- Verify: `just employee`, `just db-actor`, `just db-bench`, `oop-accept` green.
### Phase E — document the contract
- The `@table` mode surface in the language-surface guide and the db-binding
contract; the new diagnostics in the error catalog; the mode's effect on
replay in `database/src/CODE-LOGIC.md`.
- Verify: `just linkcheck` clean; the language-surface guide's `@table` row
matches what the parser actually accepts.
## Acceptance Criteria
- **Given** every existing `@table` declaration in the repository, **when** it is
compiled after this change, **then** behaviour is byte-identical — `durable`
is the default and nothing opts in silently.
- **Given** a table declared `mode: ram`, **when** rows are inserted with
`WO_DATA` set, **then** the WAL does not grow, and after a restart the table is
empty while `durable` tables in the same program replay intact.
- **Given** `mode: ram` and a measured insert workload, **when** it runs against
the same shape as a `durable` table, **then** the write-path saving is visible
in `bench/baseline.json` — the per-table half of iteration 22's 66× gap.
- **Given** an unknown mode name or `mode:` given twice, **when** it is compiled,
**then** it fails with the catalogued diagnostic naming the legal modes.
- **Given** a `durable` table holding a `ref` into a `ram` table, **when** it is
compiled, **then** the compiler refuses (or warns, per fork 2) — a persistent
row cannot reference one that evaporates.
- **Given** a WAL containing records for a table whose source now says `ram`,
**when** the program starts, **then** the situation is handled explicitly
(refuse, or skip and report) and never by silently loading rows into a table
declared not to have any.
- **Given** the `.wob` format change, **when** an image from the previous version
is loaded, **then** the loader refuses it clearly on the version rather than
misreading a descriptor.
## Out Of Scope
- **Implementing `cold`.** Declared here so the mode set is settled and the
format carries it; the engine behaviour is [6](06-cold-tiering.md). Until then
a `cold` declaration must be refused rather than silently treated as
`durable` — accepting a mode that does nothing is how a feature becomes a lie.
- **Per-table capacity limits and eviction** — [5](05-bounded-tables-eviction.md).
This iteration says what a table *is*; that one says how much of it there may
be.
- **`@table` feature flags and `transaction { }`** — language
[iteration 18](../language-runtime-database/18-memory-db-features.md), whose
spec is approved and deliberately left whole.
- **Per-table WAL files.** One log, one writer, shard 0 — the invariant stage 3
established and [7](07-single-file-db.md) depends on. Modes decide *whether* a
table logs, never *where*.
- **Migrating an existing dataset between modes.** A schema-change story, and
the repo already records destructive migrations as a recorded future.
- **Encryption at rest, compression of the WAL.** Neither has a consumer.
## Info
The grammar surface this touches, read from the source: the parser's `@table`
argument loop and its `unknown @table argument` failure; `Ast.table_cfg` as
`{ table_name : string option; indexes : string list list }`; and the class
descriptor in `runtime/src/wob.h` that the loader validates. Adding a key is
genuinely small — the semantics are the iteration.
Forks the spec must settle:
1. **What are the modes called?** `ram` / `durable` / `cold` is descriptive of
mechanism. `scratch` / `persistent` / `archived` is descriptive of intent and
is what a developer reasons about. The names are the API and are hard to
change later; leaning the intent-shaped set for the first two if a
short-enough pair can be found, since a developer choosing a mode is thinking
about what the data is *for*, not about where it sits.
2. **How hard does the compiler push?** Three levels: warn on the suspicious
cases; refuse the provably-broken ones (a `durable`→`ram` `ref`); or a full
dataflow check that an insert into a `ram` table is never expected to persist.
The third is not statically decidable in general. Leaning: refuse the
relation case (provable, high value), warn on the `WO_DATA`-with-no-durable-
table case, and stop there.
3. **What is the default, and does it depend on iteration 1?** `durable` keeps
every existing program identical, which is nearly decisive. But if iteration
1's numbers show the WAL write dominating a workload nobody wanted durable,
there is an argument for making the choice mandatory — no default, every
`@table` states its mode. That is a bigger source change and a better
language; the fork is whether the churn is worth it now or at 1.0.
4. **Does `mode: ram` imply anything about the actor/DB-actor path?** Stage 3
marshals worker-shard statements to shard 0 because the owner shard holds the
store and the WAL. A `ram` table has no WAL — so does it still need to live on
the owner? A per-shard `ram` table would be dramatically faster and a
different consistency story. Tempting, out of scope here, and worth recording
as a candidate rather than deciding in passing.

View file

@ -1,13 +1,19 @@
---
iteration: "32"
track: databasev2
iteration: "3"
was_language_iteration: "32"
status: refine
chain: 6
---
# Iteration 32 — WAL checkpoint: disk space reclamation and bounded replay
# databasev2 3 — WAL checkpoint: disk space reclamation and bounded replay
> **Moved 2026-08-26** from the language track, where this was iteration 32.
> Part of [Story — the database beyond RAM](../language-runtime-database/00-story.md). Content unchanged by
> the move; its dependencies are restated in that track index.
> Format: `product/story-iteration-template`. Part of
> [Story — one language, one runtime, one database, one binary](00-story.md).
> [Story — one language, one runtime, one database, one binary](../language-runtime-database/00-story.md).
>
> **Inserted 2026-08-21** (stage-3 guarantee refinement found the hole):
> the WAL is append-only FOREVER — no checkpoint, no truncation exists

View file

@ -1,13 +1,19 @@
---
iteration: "23"
track: databasev2
iteration: "4"
was_language_iteration: "23"
status: refine
chain: 5
---
# Iteration 23 — io_uring group-commit write path
# databasev2 4 — io_uring group-commit write path
> **Moved 2026-08-26** from the language track, where this was iteration 23.
> Part of [Story — the database beyond RAM](../language-runtime-database/00-story.md). Content unchanged by
> the move; its dependencies are restated in that track index.
> Format: `product/story-iteration-template`. Part of
> [Story — one language, one runtime, one database, one binary](00-story.md).
> [Story — one language, one runtime, one database, one binary](../language-runtime-database/00-story.md).
>
> **Inserted 2026-08-15.** The write-path optimization, and deliberately the
> LAST database performance iteration: it only earns its complexity once

View file

@ -0,0 +1,169 @@
---
track: databasev2
iteration: "5"
status: refine
---
# databasev2 5 — bounded tables and eviction: a declared budget, and back-pressure before the cliff
> Part of [Story — databasev2: the database beyond RAM](00-story.md).
> Needs [2](02-table-storage-modes.md) for the mode a bound attaches to, and
> [1](01-ram-ceiling-measurement.md) for the numbers that set a sane default.
>
> **The simpler half of the hard problem, done first on purpose.** Evicting from
> a bounded resident table and evicting to disk are the same policy question with
> different destinations. Getting the policy right where the answer is "drop it"
> de-risks [6](06-cold-tiering.md), where the answer is "write it somewhere and
> be able to find it again".
## Goals
- **A table may declare a maximum.** Rows, bytes, or both — a `ram` table that
is a cache or a session store has a size the application is willing to spend,
and today it has no way to say so. Unbounded growth in a table nobody intended
to be large is the most common route to the ceiling iteration 1 measured.
- **Something defined happens at the bound.** Today the answer is "grow until
the process dies". The candidates are eviction (drop the least valuable row),
refusal (trap, let the caller decide), and back-pressure (make the writer
wait). Each is right for a different table, which argues for the policy being
declared rather than chosen for the developer.
- **Back-pressure before the cliff, not at it.** The dangerous exit iteration 1
characterises is swap thrash, which arrives with **no error signal at all**.
A budget that is enforced at 100% has already lost; the value is in acting at
a threshold, while there is still headroom to act.
- **Eviction that respects the engine's actual invariants.** Rows have stable
addresses forever, the free-slot list recycles slots, ids are never reused, and
every secondary index and unique shadow must stay consistent with the slab.
Eviction is `wo_row_remove` with a policy in front — it must go through the
same choke point, not around it.
## Phases
### Phase A — declaring the bound
- Extend the `@table` surface from [2](02-table-storage-modes.md) with a
capacity and a policy. One new grammar arm, the same catalogued-diagnostic
discipline, defaults that keep every existing table unbounded so nothing
changes silently.
- Decide whether a bound is legal on a `durable` table (fork 1) — evicting a row
that was acked as durable is a promise being broken, and the answer is
probably "only with an explicit, differently-named policy".
- Verify: golden fixtures per policy; existing tables unchanged; illegal
combinations refused at compile time with catalogued codes.
### Phase B — the accounting
- Track per-table size cheaply. Row count is free; bytes are not — per-row
footprint includes the slab slot plus each heap value's own allocation
(`db_text`, `db_rec`, `db_multi`, `db_map`), which iteration 1 will have
quantified. Decide what is counted and be honest that it is an estimate of RSS,
not RSS.
- Expose it, because a budget nobody can observe is a budget nobody can tune.
How it is exposed is fork 3.
- Verify: accounting tracks a known workload within a stated error bound;
deleting rows returns the accounting to its prior value (the free-slot list
already makes this true of slots).
### Phase C — the policies
- **Refuse**: the bound is a hard ceiling, an insert past it traps catchably.
The simplest correct behaviour and the right default for anything precious.
- **Evict**: drop the least valuable row via `wo_row_remove` so indexes, unique
shadows and the free-slot list all stay honest. The recency metadata this needs
is fork 2 — and note the engine currently stores no per-row access time, so
true LRU is not free.
- **Back-pressure**: at a threshold below the bound, slow or park the writer.
This composes with the fiber model (a parked writer blocks nobody) and is the
only policy that addresses the swap-thrash exit rather than the allocation
exit.
- Verify: each policy behaves at the bound; after eviction every index agrees
with the slab; a parked writer resumes and does not deadlock the shard.
### Phase D — the pressure signal
- A process-level threshold, not just per-table: when total resident size
crosses a configured fraction, tables with an eviction policy start shedding
**before** the allocator or the OS gets involved. This is the iteration's real
contribution — turning an invisible failure into a managed one.
- Decide precedence when several tables could shed (fork 4).
- Verify: under the iteration-1 growth workload with the signal enabled, the
process holds a steady state instead of walking into swap; the latency curve
stays inside its baseline.
### Phase E — gate it against the measurement
- Re-run iteration 1's growth workload with bounds and policies configured. The
proof is a before/after on the same harness: previously the curve degraded and
the process died; now it plateaus.
- Baseline rows for steady-state throughput under pressure.
- Verify: `just employee`, `just db-actor`, `just db-bench` green; the gate bites
on a doctored pressure metric.
## Acceptance Criteria
- **Given** a table bounded at N rows with policy `refuse`, **when** the N+1st
insert is attempted, **then** it traps catchably, the row count stays N, and
every index agrees with the slab.
- **Given** a table bounded at N rows with policy `evict`, **when** the N+1st
insert arrives, **then** exactly one row is evicted, the new row is present,
the count is N, and no index or unique shadow references the evicted row.
- **Given** an evicted row's id, **when** it is looked up, **then** it is absent
— and its id is never reused by a later insert, preserving the invariant the
id hash's tombstone sentinel depends on.
- **Given** a bound expressed in bytes, **when** rows of a known shape are
inserted, **then** the bound is honoured within the stated accounting error,
and that error is documented rather than implied.
- **Given** back-pressure configured at a threshold, **when** the threshold is
crossed, **then** writers are slowed or parked, reads are unaffected, and no
shard deadlocks.
- **Given** the process-level pressure signal and the iteration-1 growth
workload, **when** it runs to what previously exhausted memory, **then** the
process reaches a steady state and read p99 stays within its baseline — the
before/after that justifies the iteration.
- **Given** a `durable` table, **when** an eviction policy is applied to it,
**then** either it is refused at compile time or it is a distinctly named
policy that says out loud it discards acked data.
## Out Of Scope
- **Writing evicted rows anywhere** — that is [6](06-cold-tiering.md). Here
eviction means the row is gone. Keeping the two apart is what makes the policy
work reviewable on its own.
- **True LRU if it costs a write per read.** Touching per-row metadata on every
read would turn the 1µs read path into a write path — the same trap iteration
3's session touch has. An approximation (insertion order, a coarse clock, a
sampled counter) is very likely the right answer and fork 2 should say so
explicitly rather than defaulting to textbook LRU.
- **The TTL cache middleware** — language
[iteration 18](../language-runtime-database/18-memory-db-features.md). Expiry
by *time* is that; bounding by *size* is this. They compose.
- **Query-level result limits.** `take n` already exists in the query surface.
- **Shrinking slabs back to the allocator.** Slab addresses are stable forever
by design and that invariant is load-bearing; reclaiming a slab whose rows were
all evicted is a separate, delicate change with its own iteration if anyone
wants it.
## Info
Forks the spec must settle:
1. **May a `durable` table be bounded?** Evicting an acked row contradicts the
durability promise. But an audit table that must not grow forever is a real
need, and the honest form of it is probably archival (iteration 6) rather than
eviction. Leaning: bounds on `durable` are refused, and the need is redirected
to 6.
2. **What is "least valuable"?** No per-row access time exists today, so LRU
costs a write per read. Candidates: insertion order (free — ids are already
monotonic per table), a coarse epoch stamped on write only, or sampled
approximation. Insertion order is FIFO not LRU, which is wrong for a cache
and fine for a queue — so the policy name should say which it is rather than
claiming "LRU" and delivering FIFO.
3. **How is size observed?** Without observability (language iteration 30) there
is no metrics endpoint to publish it on. Options: a builtin returning a
table's current size, a `@table`-derived query, or stderr on threshold
crossing. A builtin is the smallest thing that makes the feature tunable by
the program that owns the budget.
4. **Precedence when several tables can shed.** Largest first is simple; the
application's own priority order is more correct and needs a way to express
it. Proportional shedding is fairest and hardest to reason about. This
decides whether the pressure signal is predictable enough to trust.

View file

@ -0,0 +1,183 @@
---
track: databasev2
iteration: "6"
status: refine
---
# databasev2 6 — cold tiering: rows that leave RAM and come back
> Part of [Story — databasev2: the database beyond RAM](00-story.md).
> Needs [2](02-table-storage-modes.md) for the `cold` mode declaration,
> [3](03-wal-checkpoint.md) so the log this builds on does not grow forever,
> and [5](05-bounded-tables-eviction.md) for the policy machinery.
>
> **The iteration that actually raises the ceiling, and the one most likely to
> go wrong.** Everything before it makes the limit visible, declared and
> managed. This one removes it — for tables that opt in — and in doing so
> touches the project's most load-bearing principle. It should be approached
> with more suspicion than enthusiasm.
## Goals
- **A `cold` table may hold more rows than fit in memory.** Recently-used rows
are resident; the rest live on disk and are faulted back on access. This is
the whole feature and every other goal is a constraint on it.
- **Do it without becoming a paged storage engine.** `discarded.md` records that
the disk story is the WAL and that a paged B-tree engine was rejected. `cold`
must not be a licence to rebuild SQLite inside `database/src`. The design that
respects the doctrine reuses the log that already exists plus an index into
it — a log-structured read path, not a page cache.
- **Keep the resident path exactly as fast as it is.** A `durable` or `ram`
table must not pay one instruction for a feature it does not use. Iteration
22's 1.3M ops/s read baseline is the regression gate, and a measurable read
regression on non-`cold` tables is grounds to reject the design, not to tune
it.
- **Be honest in the query surface about what a fault costs.** A scan over a
`cold` table can touch disk per row. The engine currently answers every read
from memory at p50 1µs; a `cold` scan is a different animal and the language
should not pretend otherwise — see fork 3, which is the most important
question in this iteration.
- **Never lose an acked write.** Every guarantee iterations 9 and 22 established
holds byte-for-byte: ack-after-fsync, whole-or-nothing replay, torn tails
dropped by CRC. A tiering layer that weakens any of those is a regression
disguised as a feature.
## Phases
### Phase A — settle the design before writing any of it
- This iteration needs a spec more than any other in the track. The candidate
shapes are genuinely different: (a) the WAL becomes the primary store with an
in-memory id→offset index and a resident row cache; (b) a separate
append-only row file per cold table, checkpointed by 3's machinery; (c)
eviction to disk with a free-space map, which is the paged engine wearing a
hat.
- Whichever wins must state its read amplification, its recovery story, and what
happens when the index itself does not fit — an id→offset map for a billion
rows is not free either, and a design that only moves the ceiling is worth
knowing about before it is built.
- Verify: the spec names the shape, the amplification, and the failure modes.
No code in this phase.
### Phase B — the resident/cold boundary
- Which rows are resident: reuse [5](05-bounded-tables-eviction.md)'s policy and
accounting rather than inventing a second notion of "least valuable".
- Eviction becomes write-then-drop instead of drop, and it must be atomic with
respect to a concurrent reader — a row that is being written out must not be
briefly unreachable.
- The fault path: a lookup that misses resident memory reads from disk,
materialises the row, and admits it under the resident policy.
- Verify: a table larger than the resident bound serves correct rows for every
id; a row evicted and faulted back is byte-identical, including every heap
value (`db_text`, `db_rec`, `db_multi`, `db_map` each round-trip).
### Phase C — indexes and constraints across the boundary
- **The hard part, and the reason this is late in the track.** A secondary index
over a `cold` table either stays fully resident (bounding the table by index
size rather than row size — which may be the honest answer) or is itself
tiered. A `@unique` constraint must hold across rows nobody has in memory: the
shadow check cannot scan a slab that is not there.
- Foreign-key restrict must also hold — a delete has to know whether any cold
row references it.
- Verify: `@unique` refuses a duplicate whose only conflicting row is cold; FK
restrict refuses a delete whose only referrer is cold. These two criteria are
the correctness core of the iteration.
### Phase D — recovery
- Crash mid-eviction, crash mid-fault, crash mid-checkpoint-of-a-cold-table.
Each must recover to a consistent state with no acked write lost and no row
visible twice.
- Interaction with [3](03-wal-checkpoint.md)'s snapshot: a cold table's on-disk
rows are part of the durable state a checkpoint must account for, not
something it can truncate past.
- Verify: `kill -9` at each of the three points, replayed, with every acked write
present and the resident/cold split re-derived correctly.
### Phase E — measure it, then decide whether to keep it
- Read/write throughput and p99 for a `cold` table at several resident ratios,
and a **regression check that non-`cold` tables did not move**.
- Publish the amplification honestly in `perf-targets.md`: how much slower a
cold fault is than a resident read, as a number.
- Verify: `just db-bench` green with new cold-path rows; the resident baseline
unchanged; `just employee`, `just db-actor`, `oop-accept` green; ASan and TSan
clean on the fault path.
## Acceptance Criteria
- **Given** a `cold` table with more rows than the resident bound, **when** any
row is looked up by id, **then** it is returned correctly whether resident or
faulted, byte-identical including every heap-valued column.
- **Given** a `cold` table under a read workload, **when** the resident set is
smaller than the working set, **then** the process holds steady state without
approaching the RAM ceiling iteration 1 measured.
- **Given** a `@unique` column on a `cold` table, **when** a duplicate is
inserted whose conflicting row is **not resident**, **then** the insert is
refused — the constraint holds across the boundary or it does not hold.
- **Given** a `ref` into a `cold` table, **when** the referenced row's owner is
deleted and the only referrer is cold, **then** FK restrict refuses the delete.
- **Given** `kill -9` during an eviction, a fault, and a checkpoint, **when** the
program restarts, **then** every acked write is present, no row appears twice,
and the resident/cold split is re-derived correctly.
- **Given** a `durable` or `ram` table, **when** the read benchmark runs after
this iteration, **then** its throughput and p99 are inside the existing
baseline tolerance — no cost for a feature not used.
- **Given** a cold fault, **when** its latency is measured, **then** the
amplification versus a resident read is recorded in `perf-targets.md` as a
number a developer can plan around.
## Out Of Scope
- **A paged B-tree storage engine.** Explicitly rejected in
[`discarded.md`](../../plan/discarded.md) and not reopened by this iteration.
If the spec phase concludes that tiering *requires* one, the correct outcome is
to reject tiering and say so — not to quietly build the thing the project
decided against.
- **Making `cold` the default, or applying it to a table that did not ask.**
Opt-in per table, forever.
- **Tiering to anything but the local filesystem.** Object storage needs
outbound sockets (language
[iteration 38](../language-runtime-database/38-content-platform-capabilities.md))
and would change the latency story by orders of magnitude.
- **Compression of cold rows.** Composes with
[porch 7](../porch/07-sse-and-compression.md)'s codec if that lands first;
not a dependency either way and not this slice.
- **Cross-shard cold tables.** The owner shard owns the store and the WAL; a
cold table is more of the same. Per-shard storage is a separate architectural
question noted in [2](02-table-storage-modes.md)'s forks.
- **Tiering the query planner's behaviour.** If a scan over a cold table is
expensive, the answer for now is that it is expensive and documented — not a
cost-based planner.
## Info
Forks the spec must settle — this iteration is mostly forks, which is why phase
A produces no code:
1. **Which shape?** WAL-as-primary-store with an id→offset index and a row
cache reuses machinery that exists and keeps the doctrine ("the disk story is
the WAL") literally true. A separate per-table row file is cleaner to reason
about and duplicates the log. Eviction with a free-space map is the rejected
paged design. Leaning (a), with the caveat in fork 2.
2. **What if the index does not fit either?** An id→offset entry per row is far
smaller than a row, so this moves the ceiling by a large constant — but it
does not remove it. Say so plainly in the spec: `cold` buys an order of
magnitude, not infinity. A design sold as unlimited will be deployed as if it
were.
3. **Does the language surface the cost?** Three positions. Silent — a cold
table reads like any other and the developer discovers the latency in
production. Annotated — the mode is at the declaration, so an attentive
reader knows, which is the status quo of this design. Or *explicit at the use
site*, where a query over a cold table must acknowledge it somehow. The third
is most in keeping with a language whose whole thesis is that the compiler
tells you the truth — and it is also the most intrusive. This is the fork with
the largest effect on what writeonce *is*, and it deserves the brainstorm more
than any implementation detail here.
4. **Is `@unique` on a cold table simply refused?** Keeping a unique index fully
resident is a bound on the table by index size, which is honest and simple.
Refusing `@unique` on `cold` outright is even simpler and might be right for
a first version — a constraint that silently only checks resident rows would
be a correctness hole, and that is the one outcome that must not ship.

View file

@ -1,12 +1,18 @@
---
iteration: "33"
track: databasev2
iteration: "7"
was_language_iteration: "33"
status: refine
---
# Iteration 33 — `WO_DATA=<path>.db`: the persistent store as one file
# databasev2 7 — `WO_DATA=<path>.db`: the persistent store as one file
> **Moved 2026-08-26** from the language track, where this was iteration 33.
> Part of [Story — the database beyond RAM](../language-runtime-database/00-story.md). Content unchanged by
> the move; its dependencies are restated in that track index.
> Format: `product/story-iteration-template`. Part of
> [Story — one language, one runtime, one database, one binary](00-story.md).
> [Story — one language, one runtime, one database, one binary](../language-runtime-database/00-story.md).
>
> **Inserted 2026-08-22** (developer ask: "can the persistent db be in
> file.db form?"). The truth is already almost there: `WO_DATA=<dir>`
@ -46,7 +52,7 @@ status: refine
- A paged database file (SQLite's shape) — RAM is authoritative; the
disk story is the WAL, full stop.
- Checkpoint/compaction — [iteration 32](32-wal-checkpoint.md)'s; its
- Checkpoint/compaction — [iteration 32](03-wal-checkpoint.md)'s; its
rename-swap (write snapshot+tail to a NEW file, fsync, `rename()`
over the old) is exactly what keeps the single-file promise crash-safe
when it lands. 33 before or after 32 works; landing 33 first means

View file

@ -1,12 +1,18 @@
---
iteration: "27"
track: databasev2
iteration: "8"
was_language_iteration: "27"
status: hold
---
# Iteration 27 — query grammar, driven by real embedded-DB corpora
# databasev2 8 — query grammar, driven by real embedded-DB corpora
> **Moved 2026-08-26** from the language track, where this was iteration 27.
> Part of [Story — the database beyond RAM](../language-runtime-database/00-story.md). Content unchanged by
> the move; its dependencies are restated in that track index.
> Format: `product/story-iteration-template`. Part of
> [Story — one language, one runtime, one database, one binary](00-story.md).
> [Story — one language, one runtime, one database, one binary](../language-runtime-database/00-story.md).
>
> **Inserted 2026-08-16.** A query-surface iteration in the 9b family: the
> language-integrated query grows to cover the grammar that *real

View file

@ -1,12 +1,18 @@
---
iteration: "20"
track: databasev2
iteration: "9"
was_language_iteration: "20"
status: hold
---
# Iteration 20 — cross-program tables: attach to a running program's database
# databasev2 9 — cross-program tables: attach to a running program's database
> **Moved 2026-08-26** from the language track, where this was iteration 20.
> Part of [Story — the database beyond RAM](../language-runtime-database/00-story.md). Content unchanged by
> the move; its dependencies are restated in that track index.
> Format: `product/story-iteration-template`. Part of
> [Story — one language, one runtime, one database, one binary](00-story.md).
> [Story — one language, one runtime, one database, one binary](../language-runtime-database/00-story.md).
>
> **Inserted 2026-08-15**, hence `20`. It follows 9b because a program
> attaching to another's tables wants the same typed statements and queries
@ -141,7 +147,7 @@ read or read+write (per-table refinement deferred until a workload needs
it), and the registration is A's manifest so a grant is a config change +
restart, not an API. **Superseded as the end state (2026-08-15):**
identity is a keypair and grants name public keys — iteration
[21](21-keypair-attach-auth.md) owns that; the uid check is only this
[21](10-keypair-attach-auth.md) owns that; the uid check is only this
iteration's bootstrap and must be flagged pre-21 wherever it ships.
**4. What does B's statement actually block on?** B's insert crosses the

View file

@ -1,12 +1,18 @@
---
iteration: "21"
track: databasev2
iteration: "10"
was_language_iteration: "21"
status: hold
---
# Iteration 21 — keypair authentication for cross-program attach
# databasev2 10 — keypair authentication for cross-program attach
> **Moved 2026-08-26** from the language track, where this was iteration 21.
> Part of [Story — the database beyond RAM](../language-runtime-database/00-story.md). Content unchanged by
> the move; its dependencies are restated in that track index.
> Format: `product/story-iteration-template`. Part of
> [Story — one language, one runtime, one database, one binary](00-story.md).
> [Story — one language, one runtime, one database, one binary](../language-runtime-database/00-story.md).
>
> **Inserted 2026-08-15.** Promotes iteration 20's identity fork (Info,
> fork 3) to its own iteration: the name + unix-uid lean is the milestone

View file

@ -105,22 +105,17 @@ still pending IS the runtime-concurrency chain; order:
| 16 | 19 | [Float + Bytes](19-missing-scalar-types.md) | **LANDED 2026-08-20** — `.wob` v5; the full stack: IEEE-quiet f64 through literals/VM/@table/WAL/json + Bytes as the binary carrier, no implicit mixing, total-order indexes. Unblocks 24 (WS frames) and the crypto fork (digests). *(was 20)* |
| 17 | 31 | [Actor lifecycle](31-actor-lifecycle.md) | request/response (today `send` is one-way and callers `sleep` to await), bounded mailboxes with backpressure (today the FIFO just grows), actor death/supervision, and timers beyond `time.sleep`. 24 cannot be written honestly without these. *(story written 2026-08-21)* |
| 18 | 24 | [chat: WebSocket workload](24-chat-websocket-workload.md) | the arc's acceptance: WS upgrade + frames (SHA-1 via crypto fork, Bytes via 19), rooms/broadcast, 1k clients, drain-clean. *(was 19)* |
| 19 | 23 | [io_uring group-commit](23-io-uring-commit.md) | WAL WRITE+FSYNC chains on the arc's per-shard rings; fsync fallback kept (after 22 + the arc). *(was 9f)* |
| 20 | 32 | [WAL checkpoint](32-wal-checkpoint.md) | **NEW 2026-08-21** (stage-3 guarantee refinement found the hole) — the WAL is append-only forever: snapshot + truncate reclaims disk and bounds replay time; every durability guarantee byte-identical; crash mid-checkpoint recovers from the previous snapshot + full tail. After 23 (composes with group-commit); RAM slot-reuse already contracted in `04-db-binding.md`. |
| 21 | 33 | [Single-file store](33-single-file-db.md) | **NEW 2026-08-22** — `WO_DATA=<path>.db`: a file path IS the wal (the store already lives in exactly one file; this makes the surface say so). Driver-only, independent of the chain; composes with 32's rename-swap. |
| 22 | 34 | [Crypto builtins](34-crypto-builtins.md) | **NEW 2026-08-22** — SHA-1/SHA-256/HMAC-SHA256 as C builtins over Bytes (no bitwise ops in the language, hand-rolled per doctrine, vector-verified). GATES 24's WS handshake; digest floor for held 21 and the ETag row. |
| 23 | 37 | [wo-html components](37-wo-html-components.md) | **NEW 2026-08-23** — an MVC-shaped view layer in the wo-html LIBRARY (framework stays micro): structural `Component` interface (`render() -> Text`), layout components with slots, the site sample migrated as acceptance. Angular's component FORMAT studied and translated to server-rendered no-JS `.wo`; DI/bindings rejected. **LANDED 2026-08-25.** Raw text literal 2026-08-24 (backtick, verbatim content, margin stripped at lex time, `${ }` raw / `{{ }}` auto-escaping; WO-E004/WO-E005) — lexer plus one parser desugar, nothing downstream. Component layer 2026-08-25: `Component`/`render_all`/`Layout` in wo-html, `ok_html` moved into the framework, site and shop both migrated. |
| 23 | 35 | [net runtime seams](35-net-runtime-seams.md) | **NEW 2026-08-22** — the ledger's three 🔧 rows owned: fd deadlines composing with the park plane, Unix-socket listeners, peer address (trusted-proxy check). Framework knobs stay framework slices; pairs naturally with 24 (dead-client eviction). |
| 23 | 36 | [operator parity](36-operator-parity.md) | **NEW 2026-08-22** — `not`, the five bitwise operators (`& \| ^ << >>`, Int-only, riding the additive/multiplicative rungs Go-style), hex/binary/`_` literals, and compound assigns wired to the written-out form. **Code landed 2026-08-22** on branch `operator-parity`: `.wob` v6, opcodes 42–46 with the 0..63 shift trap (WO-E223), all gates green; reference project `.dev/reference/go` drove the operator-precedence design. Awaiting the developer's MANUAL pass over `docs/examples/operators/` (no test fixtures, by directive). Unblocks story 34's pure-`.wo` HMAC question. |
| 24 | 25 | [HTTP service layer](../../superpowers/plans/2026-08-01-http-service-layer.md) | `service` blocks lower onto the framework (after 9b + 20 by their own precedence notes). **HELD 2026-08-21** — story file removed; the plan doc remains. *(was 10)* |
| 25 | 18 | [framework v2: memory-rich features](18-memory-db-features.md) | spec+plan approved: TTL cache, @table flags, durable job queue, `transaction { }` over the WAL's staged batch. **Demoted from seq 14**: more surface on a framework with one consumer, and the cache still stores `Text` because there are no generics |
| 26 | 27 | [Query grammar corpus](27-query-grammar-corpus.md) | grow the query grammar from real corpora; likely collapses to "confirm `len(query)` + add `exists`"; precedes 28. *(was 9g)* |
| 27 | 26 | [Blue-green deploy](26-blue-green-deploy.md) | two VM slots, in-runtime compile, atomic switch, resident rollback (plan authored after 9 + 25). *(was 12)* |
| 28 | 20 | [Cross-program tables](20-cross-program-tables.md) | attach to a running program's database over local IPC; owner stays the single writer (channel half-built). **Demoted from seq 16**: new distribution surface while there is no TLS, no crypto, and the multi-shard DB still traps. *(was 9c)* |
| 29 | 21 | [Keypair attach auth](21-keypair-attach-auth.md) | program identity is a keypair; mutual challenge–response at attach (crypto half-built; plan folds into 20's). **Demoted with 20** — and it needs crypto primitives that do not exist. *(was 9d)* |
| 30 | 28 | [skillhost host workload](28-skillhost-host-workload.md) | host-shaped driving workload naming runtime gaps — demoted with the framework goal. *(was 14)* |
| 31 | 29 | [Compile-time metaprogramming](29-compile-time-metaprogramming.md) | `@derive(...)` from class-table metadata; held with the parked drain by the 2026-08-08 scope directive. *(was 13)* |
| 32 | 38 | [Content platform capabilities](38-content-platform-capabilities.md) | **NEW 2026-08-26** — the two capability families nothing owns, read off `wob.h`: `fs` mutation (`write`/`remove`/`rename`/`mkdir` — the table holds exactly six fs builtins, ids 40–45, where `append` creates-if-absent and grows, so a file is never replaced, truncated, deleted or renamed) and `net.connect` (ids 51–55 + 91–95, no connect; no `connect()` call in `runtime/src/` at all — which is every identity/notification/federation story at once; 35 deferred connect-side timeouts as "no workload asks yet" — this is the ask). Driven by a `docs/examples/vault` content-collaboration workload in 28's mould; WebDAV, CRDT editing, previews and FTS each get a written verdict instead of an implication. New builtins start at 96 (89/90 are 31's reserved holes); no `.wob` bump. Off-chain, wants a spec. |
| — | — | **[▸ the `databasev2` track](../databasev2/00-story.md)** | **Six pending database iterations moved out 2026-08-26** — WAL checkpoint *(was 32)*, io_uring group commit *(was 23)*, single-file store *(was 33)*, query grammar *(was 27)*, cross-program tables *(was 20)*, keypair attach auth *(was 21)* — renumbered 1–10 in that track alongside four new ones: the RAM-ceiling measurement, `@table` storage modes (the grammar that makes durability per-table instead of one global `WO_DATA`), bounded tables with eviction, and cold tiering. **Done database work stays here as v1 history:** 9 (engine), 9b (`@table`/relations/query) and 22 (durability baseline) are rows above and did not move. |
| 33 | 39 | [Web framework parity](39-web-framework-parity.md) | **NEW 2026-08-26** — gofiber/fiber v3.5.0 added as the web-framework reference and read end to end ([the study](../../plan/exploration/fiber/00-fiber-parity.md)); nine of its 32 middleware already have a `porch` counterpart, so the gaps are breadth, not foundations — with one exception. The study's sharpest finding: the framework's own ledger said CSRF and sessions were UNBLOCKED because iteration 34 landed HMAC, but **there is no source of randomness in the runtime at all**, and an HMAC over a guessable session id is a signed guess. So 39 leads with a random-bytes builtin, then cookies (absent in both directions; `Resp.headers: map<Text,Text>` structurally cannot emit two `Set-Cookie` lines), then the store-backed chain — limiter and idempotency first since they need only `@table` + `time.ticks`. Off-chain, wants a spec. |
| ✅ | 17 | [library projects + `internal/`](17-library-projects-internal.md) | **LANDED 2026-08-20** — `kind = "library"` + entry-less check mode (retires the `--emit` workaround) and Go's `internal/` rule as WO-E108 at the consumer's `use`; driver-only, VM/GC untouched. `just web-app` 26/0 |

View file

@ -163,7 +163,7 @@ the slice's marker doc when it landed):
| Durability | ✅ fsync-per-commit, ack-after-durable; the ack crosses shards only AFTER the owner's fsync (`just db-actor`'s WAL pair). Power-loss rides fdatasync semantics; 22's kill battery is the scripted proof. |
| Crash recovery | ✅ boot replay, torn-tail drop, index rebuild; replay completes on the primary before any worker serves (main.c boots the engine before the shards). 22 scripts the restart proof. |
| Concurrency control | ✅ stage 3 — the DB actor serializes every statement; replies are materialized copies, no torn read by construction. Cross-statement snapshots arrive with 18. |
| Space reclamation | RAM ✅ (deleted rows free their slot — ids never reused, slots are); disk ✖ → [story 32](32-wal-checkpoint.md), end of chain. |
| Space reclamation | RAM ✅ (deleted rows free their slot — ids never reused, slots are); disk ✖ → [databasev2 story 3](../databasev2/03-wal-checkpoint.md), end of chain. |
- The C proving ground (`docs/plan/exploration/c-runtime/`, phases A–F:
epoll loops, eventfd mail) is the substrate this lifts into `wovm`.
@ -174,7 +174,7 @@ the slice's marker doc when it landed):
- **Gated by the benchmark:** landing the arc means re-running
[22](22-durability-throughput-scale.md) at the concurrency
scale it unlocks and recording the before/after delta; it is also
where [23](23-io-uring-commit.md) gets a thread to overlap
where [4](../databasev2/04-io-uring-commit.md) gets a thread to overlap
durability against.
## Proposed Solution

View file

@ -25,7 +25,7 @@ Four consumers already wait on it, none able to proceed:
(`Sec-WebSocket-Accept` = base64(SHA-1(key + GUID)) — SHA-1
specifically, not a choice); the framework's ETag/conditional-request
row (wants a content hash); HMAC-signed tokens the auth core can grow;
and held [iteration 21](21-keypair-attach-auth.md), whose
and held [databasev2 10](../databasev2/10-keypair-attach-auth.md), whose
challenge–response needs primitives that "do not exist" (its demotion
note). Bytes and base64 landed with iteration 19 — the carriers exist,
only the digests are missing.

View file

@ -93,7 +93,7 @@ status: refine
*does* depend on 28's killable-subprocess work for previews, which is
exactly why previews are deferred here rather than attempted.
- **WAL checkpoint and disk reclamation** —
[iteration 32](32-wal-checkpoint.md). A metadata store whose boot replays
[databasev2 3](../databasev2/03-wal-checkpoint.md). A metadata store whose boot replays
every write ever made is a real ceiling for this workload, and naming it
here is the point; fixing it is 32's. This story's gate should record the
replay time it observes so 32 inherits a number.
@ -109,7 +109,7 @@ status: refine
route Nextcloud itself takes is unreachable until `connect` lands.
- **Full-text search.** The engine indexes equality probes on declared
columns; there is no prefix scan or FTS. Query-grammar growth is
[iteration 27](27-query-grammar-corpus.md)'s.
[databasev2 8](../databasev2/08-query-grammar-corpus.md)'s.
- **TLS** — proxy-terminated, by doctrine, unchanged. The outbound half
therefore speaks plaintext to a local sidecar or a trusted-network peer,
and the story says so out loud rather than implying HTTPS clients.
@ -196,9 +196,9 @@ a wish list. Order: brainstorm the five forks (fork 1 and 2 together, fork
establish exactly where it blocks, then the fs verbs, then `connect`, each
with corpus fixtures and its own gate leg. Big enough to want a spec and a
plan document — this is not a bounded slice like
[33](33-single-file-db.md).
[7](../databasev2/07-single-file-db.md).
Precedence: independent of the concurrency chain (stage 3 → 22 → 31 → 24 →
23 → 32) and startable beside it, with one caveat — the honest disk story
needs [32](32-wal-checkpoint.md), so if this lands first its gate records
needs [3](../databasev2/03-wal-checkpoint.md), so if this lands first its gate records
the replay number rather than claiming the platform is operationally done.

View file

@ -115,7 +115,7 @@ status: refine
it.
- **Distributed limiting across processes.** One program owns its database;
cross-program state is language
[iteration 20](../language-runtime-database/20-cross-program-tables.md).
[databasev2 9](../databasev2/09-cross-program-tables.md).
- **A background expiry sweeper.** No timer exists (`time.after` is still a
reserved builtin id in `wob.h`). Lazy pruning on access, deliberately.
- **The TTL cache middleware** — language