- two tables identical except the annotation, 200k rows, 40k reads in one key order, WAL on ext4 (not /tmp, which is tmpfs here and would have put the log in RAM), rootless cgroup v2 cap - WIDE shape, 2.55x smaller resident set: 34.4 MB vs 87.5 MB. That is the real win and the thing the mode was built for - under a 48 MB cap (between the two resident sets): keys 19635 ops/s vs all 12854 — only 1.53x faster than letting the kernel swap - degradation is far gentler though: all collapses 105x from its own uncapped throughput, keys 16x - costs 4.2x read throughput when memory is not tight, and writes are markedly slower — the keys fill did not finish in 2 min where the resident fill plus 40k reads did. No design doc had costed writes - THE UNANTICIPATED FINDING: cgroup limits charge the PAGE CACHE, so moving rows to a file does not escape a container memory limit. WAL 37 MB + RSS 34 MB cannot both live under a 48 MB cap, so every pread reaches disk. The premise "the page cache will hold the hot rows" fails in exactly the deployment this targets - first attempt used Int-only rows and showed parity; recorded, because drop_payload frees a field's VALUE and an Int's value is its inline slot word, so that shape cannot benefit and would have condemned the feature for the wrong reason - verdict: keep it, to fit ~2.5x more data in given RAM — not to make an over-capacity table fast Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> (cherry picked from commit 7cba9b1174b0bf581314b3147e25cc49e6f49464)
21 KiB
| track | iteration | status | readiness |
|---|---|---|---|
| databasev2 | 2 | in-progress | ready |
databasev2 2 — per-table storage: durable and resident
Part of Story — databasev2: the database beyond RAM. Spec:
2026-08-26-table-residency-design.md· plan:2026-08-26-table-residency.md· runnable example:docs/examples/residency, gated byjust residency.The language enrichment this track exists for. Before this, durability was one environment variable for a whole process:
WO_DATAset and every@tableis WAL-logged, or unset and none are (runtime/src/main.c, anddb.cguards each append on the WAL pointer). Real applications are not uniform — a session table and a rate-limit counter are disposable, an orders table is precious, a 120 GB audit table does not fit in RAM at all. One global switch forces "everything is precious" or "nothing is", and the developer pays for the wrong one either way.Rewritten 2026-08-27 to match what was designed and built. Two earlier drafts of this file described a three-valued
mode:enum includingcold; that design was replaced during the brainstorm and the history is at the bottom.
The design, as built
Two optional @table arguments, because the developer is answering two
independent questions — do I need this after a restart? and does it fit in
RAM? A single enum would have forced a name for every combination, which is
what made the third value unwriteable before its mechanism existed.
| Argument | Values | Default | Meaning |
|---|---|---|---|
durable |
true, false |
true |
false skips the WAL append entirely: no record, no fsync, ack from RAM, table empty after restart |
resident |
all, keys |
all |
keys keeps the id map, secondary indexes and unique shadows resident; rows are read back from the log by offset |
Both default to the pre-existing behaviour, which is why all 28 @table
declarations in the repository compiled unchanged and no golden moved.
durable: false with resident: keys is refused — rows would be neither
logged nor resident, so there would be nowhere to read them from.
The engine stays one log-structured store. The WAL already held every row;
this iteration stops discarding the payload. No second engine, no user-space row
cache — the kernel page cache is the hot copy, which is the position
exploration/postgresql/buffer-and-checkpoint.md already argued and the reason
the engine avoids O_DIRECT.
Principle 7 was amended for this: the log is authoritative, residency is a declared per-table policy. Durability is untouched and unconditional.
Progress
| # | Task | State |
|---|---|---|
| 1 | grammar: both arguments, defaults preserve behaviour | ✅ 69b7ce2 |
| 2 | WO-E224: refuse a durable ref into a volatile table |
✅ 753e6c4 |
| 3 | .wob v7: the class descriptor carries both properties |
✅ 7e68c99 |
| 4 | durable: false skips the WAL append and replay |
✅ dd67e31 |
| 5a | wo_wal_next_offset — exact record offsets |
✅ ac7d8af |
| 5b | wo_wal_read_row_at — a row from a log offset |
✅ d0c370c |
| 5c | shared borrow/release accessor, then id→offset storage | ✅ 2e347de (accessor, pure refactor, db-bench --quick 85/0), 18ce4d5 (offset storage), f9c36ef (insert + boot wiring) |
| 5d | rewire the readers: remaining wo_row_ptr sites, slab scans, FK restrict, @unique across the boundary |
✅ 11a92df (db.c), + this commit (table.c, wal.c, compaction). Updates refused, not rewired — see below |
| 6 | the two runtime refusals (no-WO_DATA, the byte budget) |
⬜ |
| 7 | measure, gate, document, close out | 🔄 measured 2026-08-30 (below); gate + closeout outstanding |
The durable half is complete and usable. A volatile table is a full table
in-process — same indexes, same @unique, same FK restrict, same query surface
— and is simply empty after a restart. That is what
porch 1–3 need for sessions,
rate-limit counters and idempotency keys.
The resident: keys half is fully wired for CRUD. Storage, reads, scans,
@unique, deletes and updates (a WAL delta record, folded back to a value on
every read) all work, and survive both a restart and a WAL checkpoint. Task
6's two runtime refusals and task 7's measurement are what remain — see
Outstanding below.
Acceptance Criteria
Met:
-
Given every existing
@tabledeclaration, when compiled, then behaviour is byte-identical. ✅ verified asgit diffovercompiler/test/golden/being empty after aWOC_BLESSrun — a green test run alone proves nothing, since blessing rewrites every golden. -
Given
durable: falsewithWO_DATAset, when rows are inserted, then the WAL does not grow and the table is empty after a restart while durable siblings replay. ✅ measured: 50 inserts wrote 1500 bytes durable and 0 volatile. Measured against the file's non-zero prefix, because the file isfallocate'd to 1 MiB and its size proves nothing. -
Given an unknown value, a repeated argument, a retired design word, or the refused combination, when compiled, then WO-E102 with a message naming what to write instead. ✅
-
Given a
durabletable holding arefinto a volatile one, when compiled, then WO-E224 naming both classes and both escapes. ✅ The reverse direction and everybacklinkshape stay legal, pinned by a run fixture so the check cannot grow over-broad. -
Given a WAL holding records for a class the source now declares volatile, when the program starts, then it refuses, exits 2, names the class, and is not reported as corruption. ✅
-
Given a v6 image, when loaded, then refused on version rather than misread. ✅
-
Given a
resident: keystable, when rows are read by id and scanned, then every row is byte-identical including heap-valued columns. ✅ 5d. Every read path goes throughwo_row_borrow/wo_row_release, and the scans go throughwo_row_next_id— deliberately the id map for a keys table and the bitmap for a resident one, since hash order would reorder every unordered query. -
Given
@uniqueon aresident: keystable, when a duplicate arrives whose conflicting row is not resident, then it is refused. ✅ 5d. The shadow probe borrows each bucket candidate, so the check costs onepreadper candidate — bounded by the bucket, not the table — and never silently narrows to the resident subset. -
Given a
resident: keystable and a WAL checkpoint, when the log is compacted, then every such row survives and still reads correctly. ✅ 5d, and this is the obligation databasev2 3 left behind. Two independent ways to fail it, both pinned bytest_keys_resident_survives_compaction: compaction walked the bitmap, which a keys row has no bit in, so every one of them would have been dropped from the new log; and the id map would still have named offsets into the replaced file. Rows are rewritten in hash order, so offsets genuinely move and a missing re-point cannot pass by luck. -
Given a
deleteof a row on aresident: keystable, when it runs, then the row is gone and nothing else is touched. ✅ fixed 2026-08-29, and it was memory corruption before the fix.wo_row_removeread the id map's value as a slot, but on a keys table that value is a LOG OFFSET, andslot_rowdoes no bounds check — so a delete indexed the slab array with a byte offset and then freed whatever it landed on. Pinned bytest_keys_resident_delete, which SEGVs against the old code. The same latent trap inwo_row_ptris closed too: it now returns NULL rather than a wild pointer when the value is an offset.This is why the loader refusal earns its keep. The gap was not one missing operation but a second one that corrupted memory silently, found only by auditing every reader of the id map.
-
Given a logged
deleteon aresident: keystable, when the process restarts, then the tombstone replays. ✅ fixed 2026-08-30 — and it was broken by the delete fix itself.wo_row_remove's keys arm borrows the row out of the log to find its index entries, and a borrow reads throughdb->rt->wal. At boot that pointer is not wired yet —main.creplays first and assignsrt.walafterwards — so the borrow found no log, the remove failed, and replay reported a valid tombstone as CORRUPTION. Replay now lends the runtime a read-only view over the fd it already has open. Pinned bytest_keys_resident_delete_then_replay, which fails against the unfixed code.Found by asking whether the read-modify-append plan was ready, not by a gate — it is unreachable today only because the loader refuses the annotation.
-
Given an
updateto a row on aresident: keystable, when it runs, then it is applied. ✅ lifted 2026-08-30. Read-modify-append: a WAL delta record chains off the row's previous offset, andwo_wal_fold_row_at— the ONE fold every reader, replay and compaction call — walks the chain back to a value. Verified four ways: the fold itself, on a chain built by hand (databasev2 2 tasks); the request path stages the delta under group commit and defers the id-map re-point to the post-barrier flush, so a hot row costs one fsync per DRAIN, not per update; replay and compaction fold delta chains the same way an ordinary read does; and the oracle test (test_oracle_all_vs_keys_same_update_sequence,test_wal.c) drives the SAME sequence of updates against aresident: alltable and aresident: keystable and asserts the rows read byte-identical at every step — the strongest available check that the fold agrees with ordinary storage, since the resident table IS the oracle.docs/examples/residency'sProducttable is genuinelyresident: keysnow;place_order's stock decrement survives a restart, gated end-to-end byscripts/residency-accept.sh.A second gap surfaced auditing the request path before lifting the refusal — the same audit class that caught the
deletememory corruption below.idx_hash,idx_cols_equalandwo_idx_probe(table.c) read a TEXT column's slot as an enginedb_text*, but the keys-resident fold was handing back VM-decodedwo_str*— a different struct layout. Reproduced as a genuine ASan heap-buffer-overflow, not merely wrong values, and present too indb.c'sGET_FIELDandPROBEarms (inline and request-path alike) — nobody had audited those against a keys-resident row because nothing could reach one while the annotation was refused. Fixed at the root rather than patched at each reader: a keys-resident borrow now hands back engine values, exactlywo_row_ptr's contract forresident: all(table.h's own "a row stores NO VM pointer" doctrine) — no index function needed to change. Pinned bytest_keys_resident_update_indexed_text, which reproduces the heap-buffer-overflow against the pre-fix code.Three limitations shipped, not fixed — documented, not papered over:
-
Mid-drain stale reads. A request reading a row inside the same uncommitted drain, while an earlier request in that drain has an in-flight update to it, may see the last durable value — read-your-writes holds within a request, not across requests in one drain. Closing it needs the fold to consult the WAL staging buffer generally, which is materially bigger.
-
Replay is O(N²) in a row's delta-chain length. Each replayed delta re-folds the whole chain back to its base record, so boot cost for one long chain is quadratic.
-
Compaction cannot see chain length.
wo_wal_should_compacttriggers on a byte ratio only, with no per-row delta-count trigger, so one hot row taking many small updates — a single popular SKU, this feature's own motivating workload — can grow a long chain without moving the aggregate ratio enough to fire a checkpoint. The delta-updates design's decision not to cap chain length rests on compaction bounding it instead; for this shape it does not.Answered by iteration 11 (spec written 2026-08-30): the update path already folds the row and the fold already walks hop by hop, so it reports the depth for free — past a fixed K the update writes a full row instead of a delta, and the chain resets. Read cost becomes at most K+1 reads and replay O(K²) per row, independent of when a checkpoint fires. Limitations 2 and 3 above both fall to it.
-
Task 7 — measured 2026-08-30, and the answer is qualified
The question, in the words this file has carried since the iteration was
written: pread through the page cache should beat the 273× collapse
iteration 1 measured for demand-paged anonymous memory, "and the whole value of
resident: keys rests on how much better."
Method. Two tables identical except the annotation, so any difference is the
storage mode's doing: 200 000 rows, 40 000 reads in the same Weyl key order,
WO_SHARDS=1, WAL on ext4 (not /tmp, which is tmpfs here and would have
put the "log" in RAM), memory capped with a rootless cgroup v2 scope.
First attempt measured the wrong thing, and is worth recording. With
Int-only rows the two modes were indistinguishable — 6 061 vs 5 599 ops/s, RSS
15.0 MB vs 14.3 MB. The cause is structural: wo_row_drop_payload frees each
field's value and returns the slot to a free list, but never releases the
slab, and an Int's value IS its inline slot word. So dropping an Int-only
row frees nothing at all. The mode cannot help that shape, and a benchmark built
on it would have condemned the feature for the wrong reason.
The wide shape (one Int, three Text) is where the mode can act.
| 200k rows, 40k reads | ops/s | p50 | p99 | RSS |
|---|---|---|---|---|
resident: all, no pressure (256 MB) |
1 354 554 | 1 µs | 2 µs | 87.5 MB |
resident: keys, no pressure (256 MB) |
320 053 | 3 µs | 5 µs | 34.4 MB |
resident: all, 48 MB cap |
12 854 | 67 µs | 231 µs | 48.1 MB |
resident: keys, 48 MB cap |
19 635 | 65 µs | 227 µs | 34.4 MB |
The 48 MB cap is chosen to sit between the two resident sets: resident: all
needs 87 MB and must page, resident: keys needs 34 MB and fits.
What it buys.
- 2.55× smaller resident set — 34.4 MB against 87.5 MB. This is the real, unambiguous win, and it is the thing the mode was built for.
- A far gentler degradation curve: under the cap
resident: allcollapses 105× from its own uncapped throughput,resident: keysonly 16×. - 1.53× faster than swapping at the same cap — 19 635 vs 12 854 ops/s.
What it costs.
- 4.2× slower reads when memory is not tight (320k vs 1.35M ops/s). A
preadand a fold per row against a pointer dereference. - Writes are markedly slower, uncosted by any design document so far: the
keys fill of 200 000 rows did not finish inside two minutes where the resident
fill plus 40 000 reads did. The per-insert drop-and-re-point work is the
difference; both tables are
durable: true, so the WAL is not.
The finding that matters most, and it was not anticipated. resident: keys
is only 1.53× faster than swapping under the cap, not the order of magnitude the
design implies — because cgroup memory limits charge the page cache. The WAL
here is 37 MB; the resident set is 34 MB; a 48 MB cap cannot hold both, so the
log's pages are evicted and every pread reaches the disk. Moving rows out of
the heap and into a file does not escape a container memory limit — the
cache the design leans on is charged to the same cgroup. The mode's premise,
"the kernel's page cache will hold the hot rows", fails in precisely the
containerised deployment it targets.
Verdict. The feature is worth keeping, but for a narrower reason than claimed: it lets a given amount of RAM hold ~2.5× more data, and degrades far more gracefully than swapping. It is not a way to make an over-capacity table fast — under a hard memory cap it is within 1.5× of simply letting the kernel swap. The honest guidance is "use it to fit more, not to go faster", and the docs should say so.
Still outstanding for task 7: wire these legs into scripts/db-bench.py
with tolerances and a baseline entry, and re-measure the resident: all read
baseline to confirm no cost for a feature not used.
Outstanding:
- Given
durable: trueand noWO_DATA, when the program starts, then it refuses. (task 6 — today this combination silently discards every write) - Given the resident footprint crossing the budget, when it does, then a refusal naming the table and the annotation. (task 6)
- Given the
resident: allread baseline, when re-measured, then inside tolerance — no cost for a feature not used. (task 7) - Given a
resident: keystable larger than RAM, when read randomly, then its read cost is measured against the resident baseline on its own read path, not inherited from databasev2 1's swap figure. (task 7) — that figure is 273× for demand-paged anonymous memory (1);preadthrough the page cache should do better, and the whole value ofresident: keysrests on how much better. If it is not materially better than swapping, the design buys nothing that the kernel was not already doing.
Out Of Scope
- Checkpoint and compaction — 3. Boot rebuilds the offset map by scanning the log until that lands, which is O(all history); 3's snapshot should persist the map.
- Eviction and a resident row cache — 5. This iteration's tables are either fully resident or keys-only.
- io_uring on the read path — a real question that only exists after this; noted in 4, deliberately not folded in.
transaction { }and@tablefeature flags — language iteration 18, approved spec, left whole.- Per-shard residency for volatile tables — a volatile table has no WAL, so it arguably need not live on the owner shard at all. Faster, and a different consistency story. Recorded as a candidate, not decided.
- Converting an existing dataset between settings. Refuse on mismatch, do not convert — implemented in task 4.
Info — the forks, settled
- Two keys, not one enum. An enum needs a name per combination, and the brainstorm demonstrated the third name is unwriteable before its mechanism is decided.
keys, notindex.index:is already an argument key, so@table(index: [c], resident: index)read badly.all/keysalso put both values on one axis — what row data stays resident.resident: nonewas rejected as overclaiming, since the indexes are very much resident.- Optional with
durabledefaulting true, not mandatory. Mandatory would have touched 28 declarations, 13 corpus fixtures and 3 goldens; the README already says nothing is API-stable, so making it mandatory at 1.0 stays available. @uniqueonresident: keysis allowed, with its index unconditionally resident. Roughly doubles the resident index; stated at the declaration so the cost is visible.- The budget is bytes, not rows — a text-heavy row and an Int-only row differ by 3.3× (measured, databasev2 1), so a row count cannot bound RAM.
History — two corrections worth keeping
The three-mode design was replaced. Earlier drafts had
mode: ram | durable | cold. cold conflated two independent properties and
could not be named honestly before its mechanism existed, and the developer's
120 GB-on-32 GB case showed the real axis was residency. Replaced by two keys,
and principle 7 amended rather than worked around.
The "one real rewrite" was fiction. The spec claimed the on-disk record was
pointer-bearing and that re-encoding it was this iteration's substantive
engineering. That came from reading table.c's db_val_encode — which builds
the in-memory slot — and inferring the file format from it. wal.c's
enc_val has been flat since iteration 9. The task was deleted, not reduced.
The opposite half then turned out to be genuinely deep. With the format
fine, the plan's storage steps still read as plumbing. Measured instead:
wo_row_ptr returns a db_row * into a slab and has 11 call sites, table.c
has 37 slab references, db.c:105-181 walks slabs for scans, enc_val
serialises from the slab, and no operation exists that drops a row's payload
while keeping its index entries. Hence the 5a–5d split. 5c and 5d need their
own write-ups, and the two open design questions for 5c are whether the id hash
stores offsets in place of slot indices or gains a parallel map, and what the
new operation does about the unique shadows, which currently point at slots.