- `Wide` text-heavy reference shape beside Int-only `Item` - `growth N int|text`: per-decile RSS read from own /proc/self/status - `growth-verify`: survivor of a crash must be a contiguous intact prefix - four footprint legs under a rootless cgroup v2 cap, swap on/off - `ceiling` leg: die at the cap, then replay must come back intact - footprint read as median-of-marginals; doublings a separate metric - 121 checks, 0 failures; footprint gated ±10%, kill-timing ±100% Measured, and it inverted two of the iteration's own predictions: - footprint 96.5-100 B/row Int vs 320.6-324 B/row text = 3.3x, NOT the "order of magnitude" three docs asserted - table storage has NO checked ceiling: SIGKILL signal 9, not a catchable WO_T_OOM. overcommit lets malloc succeed; kernel kills on page touch - swap is NOT latency collapse: 900k rows 148s capped-with-swap vs 150s uncapped. Append-mostly never re-touches cold pages - ack-after-fsync survives an OOM kill: ~40k rows, no holes, no corruption - iteration 2's budget dependency is REMOVED not satisfied — there is no "swap onset" to derive a fraction from - fix: subprocess returncode -9 was labelled a "checked refusal"; 137 is the shell spelling of the same signal Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
11 KiB
Story — databasev2: the database beyond RAM
The third track. language-runtime-database/ built the engine;
porch/ is the framework on top; this track answers the
question v1 deliberately deferred: what happens when the data does not fit in
memory.
Numbering restarts at 1, local to this track. Frontmatter carries
track: databasev2, and iterations moved here keep their old id in
was_language_iteration: so a search for "iteration 32" still finds the WAL
checkpoint. Status stays where it belongs — the status: key, never a directory.
The problem, stated honestly
Principle 7 says RAM is authoritative; the WAL makes it durable. That is a real design, not a shortcut: reads never touch disk, so latency is predictable, and durability is a sequential append rather than a storage engine bolted to the side. Iteration 22 measured what it buys — reads at 1.3M ops/s after the index probe landed, p50 1µs.
The bill comes due at the ceiling. Read from the engine as it stands:
- Rows live in
malloc'd slabs of 256, and their addresses are stable forever (database/src/table.c,DB_SLAB_ROWS). Slabs are allocated as a table grows and freed only when the table is destroyed. The free-slot list recycles removed slots, so a delete-heavy table plateaus — but a growing table only grows. - The ceiling is process RSS, not a configured number.
WO_HEAP_MB(default 64 MiB) bounds the VM object arena; table storage is separatemalloc, so nothing in the system declares a maximum dataset size. There is no knob that says "this database may use at most N". - There is no eviction, no spill, no paging, no LRU. Grep
database/src/for any of them and nothing comes back. Every row ever inserted and not deleted is resident. - The WAL is append-only with no checkpoint. Boot replays every record ever written, so startup time is O(all writes in the file's history) and disk grows without bound. That is databasev2 3.
- Durability is process-global.
WO_DATAis one environment variable that turns on oneshard-0.walfor the whole process (runtime/src/main.c). There is no way to say "this table matters, that one is scratch".
What actually breaks first
Worth being precise, because the failure mode determines the fix — and the good news is that the engine's own behaviour is clean:
Corrected 2026-08-27 by measurement. This section used to open "an
allocation failure is a catchable trap, not a crash", and that is true only of
the VM arena. Table storage has no ceiling, and with vm.overcommit_memory = 0
its malloc never fails — the process is SIGKILLed (rc=137, measured at
360 000 rows under a 64 MiB cap). The checked path below is real, but it is the
arena's, not the store's. See iteration 1.
Every malloc in
the row encoder is checked and jumps to an oom label; DB_ERR_OOM maps to
WO_T_OOM, which a program can try/catch. So a writeonce program that runs
out of memory refuses the insert rather than corrupting or dying. That is a
much better starting position than most engines have.
But the trap is almost never what a real deployment hits first. Long before
malloc returns NULL, the box starts swapping, and a RAM-authoritative database
on swap is the worst of both worlds: it has paid for in-memory data structures
and is now serving them from disk with no read path designed for that. On a
cgroup-limited host the OOM killer arrives instead, and an external SIGKILL is
the one shutdown path that skips every guarantee the WAL was written to provide —
though ack-after-fsync means acked writes still survive; iteration 22's kill -9
battery proves that much.
So the honest problem statement is not "malloc fails". It is: there is no declared budget, no back-pressure as the budget is approached, and no way to distinguish data that must be resident from data that merely is.
Iteration 1 has now measured this
(2026-08-27), and it strengthened the statement rather than softening it. A row
costs 96.5–100 B Int-only and 320.6–324 B text-heavy (3.3× apart, so no
single per-row number can bound RAM). At the ceiling the engine has exactly two
behaviours and neither one tells anybody: without swap the process is
SIGKILLed on signal 9 — table storage has no checked ceiling, and under
vm.overcommit_memory = 0 its malloc succeeds and the kernel kills on page
touch — and with swap it keeps returning 0 while serving from disk, finishing
900 000 rows in 148 s against 150 s uncapped. Durability is the one thing that
does hold: acked writes came back as an intact prefix across an OOM kill.
That is why "back-pressure at exhaustion" is not a design option. Exhaustion either kills without warning or never arrives. Only a declared threshold can speak in time.
The lever: per-table storage modes
The developer's ask, and the reason this track has a grammar iteration.
Today every @table is identical: resident, and durable if and only if
WO_DATA is set for the whole process. Real applications are not uniform —
a session table, a rate-limit counter and a page cache want resident and
disposable; an orders table wants resident and durable; an audit log wants
durable and rarely read. One global switch cannot express that, so it forces
either "everything is precious" or "nothing is".
Extending @table moves the decision into the language, where the compiler can
act on it. Two keys, not one enum — the developer is answering two
independent questions, and an enum would need a name for every combination:
durable: true | false(defaulttrue).falseskips the WAL append entirely: no record, no fsync, ack from RAM, table empty after restart. The compiler can then refuse a program that stores a durablerefinto such a table, because that id would dangle across a restart (WO-E224).resident: all | keys(defaultall).keyskeeps the id map, the secondary indexes and the unique shadows resident and reads rows back from the log by offset. This is the key that raises the ceiling — and the arithmetic is why it works: 240M rows × 16 B of index ≈ 3.8 GB resident for a 120 GB table.
durable: false with resident: keys is refused: rows would be neither logged
nor resident, so there would be nowhere to read them from.
The grammar change was small, as predicted — Ast.table_cfg gained two fields
and the parser's argument match two arms. The semantics were the work, which
is why iteration 2 is 7 tasks rather than one.
Does this break principle 7? It amends it, deliberately, and the amendment is applied: the log is authoritative and residency is a declared per-table policy. Durability is untouched and unconditional — ack after fsync, replay whole-or-nothing, torn tails dropped by CRC. What stays rejected is a second engine: a paged B-tree with its own buffer pool. Reading rows from the log we already write is not that.
An earlier draft of this section proposed a three-valued mode: enum including
cold. That name conflated durability with residency and could not be defined
before its mechanism existed; the history is in
iteration 2.
The sequence
| # | Iteration | Delivers | Needs |
|---|---|---|---|
| 1 | RAM ceiling: measure the breaking point | 🔄 measured 2026-08-27: footprint per shape (3.3× apart), the two silent exits (SIGKILL vs swap-serving-from-disk at ~uncapped speed), and ack-after-fsync surviving an OOM kill. Outstanding: the random-read-over-cap collapse, and a replay baseline | nothing; extends iteration 22's harness |
| 2 | per-table storage | the grammar: durable: true|false and resident: all|keys, per table, replacing the global WO_DATA all-or-nothing. In progress — the durable half is done |
1 for the budget default |
| 3 | WAL checkpoint (was language 32) | snapshot + truncate: disk reclaimed, replay bounded | 4 composes |
| 4 | io_uring group commit (was language 23) | close the 66× durable/RAM write gap (4.5k vs 297k inserts/s) | the arc (landed) |
| 5 | Bounded tables and eviction | a capacity a ram table may not exceed, and what happens when it does |
2 |
| 6 | Cold tiering | ⚠ largely superseded by 2 — resident: keys is the ceiling-raiser. Its premise (a user-space resident working set) was rejected in favour of the kernel page cache. Revisit only with a measurement showing the page cache insufficient |
— |
| 7 | Single-file store (was language 33) | WO_DATA=<path>.db — a file path IS the store |
independent |
| 8 | Query grammar from corpora (was language 27) | whole-query count, exists |
independent |
| 9 | Cross-program tables (was language 20) | attach to a running program's database over local IPC | independent |
| 10 | Keypair attach auth (was language 21) | program identity as a keypair; mutual challenge–response | 9 |
1 ──▶ 2 ──▶ 5 ──▶ 6
│ ▲
3 ──▶ 4 ─────┘
7, 8 independent
9 ──▶ 10
Order rationale: 1 before 2 because the budget default should follow from a
measurement, not a guess. 3 and 4 matter to 2 for the same reason tiering
onto a never-truncating log would have: resident: keys rebuilds its offset map
by scanning the whole log at boot until 3's snapshot persists it.
Amended 2026-08-27: the original rationale sequenced 6 as the ceiling-raiser
after 3, 4 and 5. resident: keys took that role into iteration 2, so 6 is
largely superseded and 5 is no longer a prerequisite for anything on the
critical path.
What this track does NOT own
| Not databasev2's | Owner |
|---|---|
transaction { } and @table feature flags |
language iteration 18 — approved spec, left whole on purpose |
| the TTL cache middleware | also language 18 (and porch 1 points there) |
| typed binding of rows into app classes | language iteration 29 @derive |
fs mutation verbs, outbound sockets |
language iteration 38 |
| benchmark harness and CI | iteration 22 (landed) built the harness; per-change CI is language iteration 30 |
| a paged B-tree storage engine | nobody, deliberately. Recorded as rejected in discarded.md: the disk story is the WAL. cold tiering is not a licence to build SQLite. |
Review protocol
The language track's, unchanged: one iteration read and approved before the next
starts; every iteration an unsplittable slice with phases, per-phase tasks,
Given/When/Then criteria and an out-of-scope list. Every engine change is gated
by just employee, just db-actor and just db-bench against
bench/baseline.json — and any iteration that claims a performance change must
move a number in that baseline, or it did not happen.