writeonce/docs/stories/databasev2/02-table-storage-modes.md
shoney.arickathil 841f6c8f3f feat(db2-keys): gate the residency measurement, close out task 7
- new `residency` leg in db-bench.py driving docs/examples/residency-bench:
  two tables identical except the annotation, control cap + binding cap
- ITS OWN PROGRAM, not a db-bench mode: declaring a resident: keys table
  is a WHOLE-PROGRAM constraint, so the no-WO_DATA refusal fires for
  every mode in the module. Putting those classes in db-bench's shared
  types made growth/ceiling/randread — which run without WO_DATA —
  refuse to start. Caught by running the leg, not by reading it
- gates the RATIOS, waives the absolutes: ops/sec under a cap is swap
  and disk I/O and belongs to the box. Same split randread makes
- rss_ratio 2.55 floor 2.0 tol 10% (structural, like bytes_per_row);
  overcap_vs_swap_x 1.53 floor 1.0; in_ram_cost_x 4.23 ceiling 8.0;
  all_collapse_x 105.4 floor 2.0
- all_collapse_x exists because the leg's FIRST run silently measured
  nothing: at QUICK's 40k rows a 48 MiB cap binds neither mode, so the
  "over-cap" half was not over cap. The cap now scales with N and the
  leg asserts it binds
- verified the gate bites: rss_ratio 1.4, overcap_vs_swap_x 0.6 and
  in_ram_cost_x 12.0 are all rejected
- task 7 closed: both criteria moved to Met with how each was verified

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
(cherry picked from commit a310496664982c372f51b43113465eb8ad9e9fb5)
2026-08-30 20:38:03 +02:00

23 KiB
Raw Blame History

track iteration status readiness
databasev2 2 in-progress ready

databasev2 2 — per-table storage: durable and resident

Part of Story — databasev2: the database beyond RAM. Spec: 2026-08-26-table-residency-design.md · plan: 2026-08-26-table-residency.md · runnable example: docs/examples/residency, gated by just residency.

The language enrichment this track exists for. Before this, durability was one environment variable for a whole process: WO_DATA set and every @table is WAL-logged, or unset and none are (runtime/src/main.c, and db.c guards each append on the WAL pointer). Real applications are not uniform — a session table and a rate-limit counter are disposable, an orders table is precious, a 120 GB audit table does not fit in RAM at all. One global switch forces "everything is precious" or "nothing is", and the developer pays for the wrong one either way.

Rewritten 2026-08-27 to match what was designed and built. Two earlier drafts of this file described a three-valued mode: enum including cold; that design was replaced during the brainstorm and the history is at the bottom.

The design, as built

Two optional @table arguments, because the developer is answering two independent questions — do I need this after a restart? and does it fit in RAM? A single enum would have forced a name for every combination, which is what made the third value unwriteable before its mechanism existed.

Argument Values Default Meaning
durable true, false true false skips the WAL append entirely: no record, no fsync, ack from RAM, table empty after restart
resident all, keys all keys keeps the id map, secondary indexes and unique shadows resident; rows are read back from the log by offset

Both default to the pre-existing behaviour, which is why all 28 @table declarations in the repository compiled unchanged and no golden moved. durable: false with resident: keys is refused — rows would be neither logged nor resident, so there would be nowhere to read them from.

The engine stays one log-structured store. The WAL already held every row; this iteration stops discarding the payload. No second engine, no user-space row cache — the kernel page cache is the hot copy, which is the position exploration/postgresql/buffer-and-checkpoint.md already argued and the reason the engine avoids O_DIRECT.

Principle 7 was amended for this: the log is authoritative, residency is a declared per-table policy. Durability is untouched and unconditional.

Progress

# Task State
1 grammar: both arguments, defaults preserve behaviour ✅ 69b7ce2
2 WO-E224: refuse a durable ref into a volatile table ✅ 753e6c4
3 .wob v7: the class descriptor carries both properties ✅ 7e68c99
4 durable: false skips the WAL append and replay ✅ dd67e31
5a wo_wal_next_offset — exact record offsets ✅ ac7d8af
5b wo_wal_read_row_at — a row from a log offset ✅ d0c370c
5c shared borrow/release accessor, then id→offset storage ✅ 2e347de (accessor, pure refactor, db-bench --quick 85/0), 18ce4d5 (offset storage), f9c36ef (insert + boot wiring)
5d rewire the readers: remaining wo_row_ptr sites, slab scans, FK restrict, @unique across the boundary ✅ 11a92df (db.c), + this commit (table.c, wal.c, compaction). Updates refused, not rewired — see below
6 the two runtime refusals (no-WO_DATA, the byte budget) ⬜
7 measure, gate, document, close out ✅ measured, gated and documented 2026-08-30

The durable half is complete and usable. A volatile table is a full table in-process — same indexes, same @unique, same FK restrict, same query surface — and is simply empty after a restart. That is what porch 1–3 need for sessions, rate-limit counters and idempotency keys.

The resident: keys half is fully wired for CRUD. Storage, reads, scans, @unique, deletes and updates (a WAL delta record, folded back to a value on every read) all work, and survive both a restart and a WAL checkpoint. Task 6's two runtime refusals and task 7's measurement are what remain — see Outstanding below.

Acceptance Criteria

Met:

  • Given every existing @table declaration, when compiled, then behaviour is byte-identical. ✅ verified as git diff over compiler/test/golden/ being empty after a WOC_BLESS run — a green test run alone proves nothing, since blessing rewrites every golden.

  • Given durable: false with WO_DATA set, when rows are inserted, then the WAL does not grow and the table is empty after a restart while durable siblings replay. ✅ measured: 50 inserts wrote 1500 bytes durable and 0 volatile. Measured against the file's non-zero prefix, because the file is fallocate'd to 1 MiB and its size proves nothing.

  • Given an unknown value, a repeated argument, a retired design word, or the refused combination, when compiled, then WO-E102 with a message naming what to write instead. ✅

  • Given a durable table holding a ref into a volatile one, when compiled, then WO-E224 naming both classes and both escapes. ✅ The reverse direction and every backlink shape stay legal, pinned by a run fixture so the check cannot grow over-broad.

  • Given a WAL holding records for a class the source now declares volatile, when the program starts, then it refuses, exits 2, names the class, and is not reported as corruption. ✅

  • Given a v6 image, when loaded, then refused on version rather than misread. ✅

  • Given a resident: keys table, when rows are read by id and scanned, then every row is byte-identical including heap-valued columns. ✅ 5d. Every read path goes through wo_row_borrow/wo_row_release, and the scans go through wo_row_next_id — deliberately the id map for a keys table and the bitmap for a resident one, since hash order would reorder every unordered query.

  • Given @unique on a resident: keys table, when a duplicate arrives whose conflicting row is not resident, then it is refused. ✅ 5d. The shadow probe borrows each bucket candidate, so the check costs one pread per candidate — bounded by the bucket, not the table — and never silently narrows to the resident subset.

  • Given a resident: keys table and a WAL checkpoint, when the log is compacted, then every such row survives and still reads correctly. ✅ 5d, and this is the obligation databasev2 3 left behind. Two independent ways to fail it, both pinned by test_keys_resident_survives_compaction: compaction walked the bitmap, which a keys row has no bit in, so every one of them would have been dropped from the new log; and the id map would still have named offsets into the replaced file. Rows are rewritten in hash order, so offsets genuinely move and a missing re-point cannot pass by luck.

  • Given a delete of a row on a resident: keys table, when it runs, then the row is gone and nothing else is touched. ✅ fixed 2026-08-29, and it was memory corruption before the fix. wo_row_remove read the id map's value as a slot, but on a keys table that value is a LOG OFFSET, and slot_row does no bounds check — so a delete indexed the slab array with a byte offset and then freed whatever it landed on. Pinned by test_keys_resident_delete, which SEGVs against the old code. The same latent trap in wo_row_ptr is closed too: it now returns NULL rather than a wild pointer when the value is an offset.

    This is why the loader refusal earns its keep. The gap was not one missing operation but a second one that corrupted memory silently, found only by auditing every reader of the id map.

  • Given a logged delete on a resident: keys table, when the process restarts, then the tombstone replays. ✅ fixed 2026-08-30 — and it was broken by the delete fix itself. wo_row_remove's keys arm borrows the row out of the log to find its index entries, and a borrow reads through db->rt->wal. At boot that pointer is not wired yet — main.c replays first and assigns rt.wal afterwards — so the borrow found no log, the remove failed, and replay reported a valid tombstone as CORRUPTION. Replay now lends the runtime a read-only view over the fd it already has open. Pinned by test_keys_resident_delete_then_replay, which fails against the unfixed code.

    Found by asking whether the read-modify-append plan was ready, not by a gate — it is unreachable today only because the loader refuses the annotation.

  • Given an update to a row on a resident: keys table, when it runs, then it is applied. ✅ lifted 2026-08-30. Read-modify-append: a WAL delta record chains off the row's previous offset, and wo_wal_fold_row_at — the ONE fold every reader, replay and compaction call — walks the chain back to a value. Verified four ways: the fold itself, on a chain built by hand (databasev2 2 tasks); the request path stages the delta under group commit and defers the id-map re-point to the post-barrier flush, so a hot row costs one fsync per DRAIN, not per update; replay and compaction fold delta chains the same way an ordinary read does; and the oracle test (test_oracle_all_vs_keys_same_update_sequence, test_wal.c) drives the SAME sequence of updates against a resident: all table and a resident: keys table and asserts the rows read byte-identical at every step — the strongest available check that the fold agrees with ordinary storage, since the resident table IS the oracle. docs/examples/residency's Product table is genuinely resident: keys now; place_order's stock decrement survives a restart, gated end-to-end by scripts/residency-accept.sh.

    A second gap surfaced auditing the request path before lifting the refusal — the same audit class that caught the delete memory corruption below. idx_hash, idx_cols_equal and wo_idx_probe (table.c) read a TEXT column's slot as an engine db_text*, but the keys-resident fold was handing back VM-decoded wo_str* — a different struct layout. Reproduced as a genuine ASan heap-buffer-overflow, not merely wrong values, and present too in db.c's GET_FIELD and PROBE arms (inline and request-path alike) — nobody had audited those against a keys-resident row because nothing could reach one while the annotation was refused. Fixed at the root rather than patched at each reader: a keys-resident borrow now hands back engine values, exactly wo_row_ptr's contract for resident: all (table.h's own "a row stores NO VM pointer" doctrine) — no index function needed to change. Pinned by test_keys_resident_update_indexed_text, which reproduces the heap-buffer-overflow against the pre-fix code.

    Three limitations shipped, not fixed — documented, not papered over:

    1. Mid-drain stale reads. A request reading a row inside the same uncommitted drain, while an earlier request in that drain has an in-flight update to it, may see the last durable value — read-your-writes holds within a request, not across requests in one drain. Closing it needs the fold to consult the WAL staging buffer generally, which is materially bigger.

    2. Replay is O(N²) in a row's delta-chain length. Each replayed delta re-folds the whole chain back to its base record, so boot cost for one long chain is quadratic.

    3. Compaction cannot see chain length. wo_wal_should_compact triggers on a byte ratio only, with no per-row delta-count trigger, so one hot row taking many small updates — a single popular SKU, this feature's own motivating workload — can grow a long chain without moving the aggregate ratio enough to fire a checkpoint. The delta-updates design's decision not to cap chain length rests on compaction bounding it instead; for this shape it does not.

      Answered by iteration 11 (spec written 2026-08-30): the update path already folds the row and the fold already walks hop by hop, so it reports the depth for free — past a fixed K the update writes a full row instead of a delta, and the chain resets. Read cost becomes at most K+1 reads and replay O(K²) per row, independent of when a checkpoint fires. Limitations 2 and 3 above both fall to it.

Task 7 — measured 2026-08-30, and the answer is qualified

The question, in the words this file has carried since the iteration was written: pread through the page cache should beat the 273× collapse iteration 1 measured for demand-paged anonymous memory, "and the whole value of resident: keys rests on how much better."

Method. Two tables identical except the annotation, so any difference is the storage mode's doing: 200 000 rows, 40 000 reads in the same Weyl key order, WO_SHARDS=1, WAL on ext4 (not /tmp, which is tmpfs here and would have put the "log" in RAM), memory capped with a rootless cgroup v2 scope.

First attempt measured the wrong thing, and is worth recording. With Int-only rows the two modes were indistinguishable — 6 061 vs 5 599 ops/s, RSS 15.0 MB vs 14.3 MB. The cause is structural: wo_row_drop_payload frees each field's value and returns the slot to a free list, but never releases the slab, and an Int's value IS its inline slot word. So dropping an Int-only row frees nothing at all. The mode cannot help that shape, and a benchmark built on it would have condemned the feature for the wrong reason.

The wide shape (one Int, three Text) is where the mode can act.

200k rows, 40k reads ops/s p50 p99 RSS
resident: all, no pressure (256 MB) 1 354 554 1 µs 2 µs 87.5 MB
resident: keys, no pressure (256 MB) 320 053 3 µs 5 µs 34.4 MB
resident: all, 48 MB cap 12 854 67 µs 231 µs 48.1 MB
resident: keys, 48 MB cap 19 635 65 µs 227 µs 34.4 MB

The 48 MB cap is chosen to sit between the two resident sets: resident: all needs 87 MB and must page, resident: keys needs 34 MB and fits.

What it buys.

  • 2.55× smaller resident set — 34.4 MB against 87.5 MB. This is the real, unambiguous win, and it is the thing the mode was built for.
  • A far gentler degradation curve: under the cap resident: all collapses 105× from its own uncapped throughput, resident: keys only 16×.
  • 1.53× faster than swapping at the same cap — 19 635 vs 12 854 ops/s.

What it costs.

  • 4.2× slower reads when memory is not tight (320k vs 1.35M ops/s). A pread and a fold per row against a pointer dereference.
  • Writes are markedly slower, uncosted by any design document so far: the keys fill of 200 000 rows did not finish inside two minutes where the resident fill plus 40 000 reads did. The per-insert drop-and-re-point work is the difference; both tables are durable: true, so the WAL is not.

The finding that matters most, and it was not anticipated. resident: keys is only 1.53× faster than swapping under the cap, not the order of magnitude the design implies — because cgroup memory limits charge the page cache. The WAL here is 37 MB; the resident set is 34 MB; a 48 MB cap cannot hold both, so the log's pages are evicted and every pread reaches the disk. Moving rows out of the heap and into a file does not escape a container memory limit — the cache the design leans on is charged to the same cgroup. The mode's premise, "the kernel's page cache will hold the hot rows", fails in precisely the containerised deployment it targets.

Verdict. The feature is worth keeping, but for a narrower reason than claimed: it lets a given amount of RAM hold ~2.5× more data, and degrades far more gracefully than swapping. It is not a way to make an over-capacity table fast — under a hard memory cap it is within 1.5× of simply letting the kernel swap. The honest guidance is "use it to fit more, not to go faster", and the docs should say so.

Gated 2026-08-30. scripts/db-bench.py grew a residency leg driving docs/examples/residency-bench — its own program, because declaring a resident: keys table is a WHOLE-PROGRAM constraint: the runtime refuses to start without WO_DATA, for every mode in the module. Putting those classes in db-bench's shared types made growth, ceiling and randread — which deliberately run without WO_DATA — refuse to start. That regression was caught by running the leg, not by reading it.

What is gated, and what deliberately is not, follows randread's existing split: the absolute ops/sec under a cap is swap and disk I/O and belongs to the box, so it is recorded and waived; the RATIOS are the engine's property.

metric baseline floor tolerance
residency.rss_ratio 2.55 2.0 10%
residency.overcap_vs_swap_x 1.53 1.0 100%
residency.in_ram_cost_x 4.23 8.0 (ceiling) 50%
residency.all_collapse_x 105.4 2.0 100%

rss_ratio carries the tight tolerance because footprint is structural — the same class of number as bytes_per_row. The two throughput ratios are guarded by their FLOORS rather than their bands, which is this harness's established answer to a metric whose absolute value belongs to the disk. all_collapse_x exists only to assert the cap actually binds; a leg whose "over-cap" half is not over cap silently measures nothing, which is exactly what the first run of this leg did.

Verified by feeding the gate a breaching run: rss_ratio 1.4, overcap_vs_swap_x 0.6 and in_ram_cost_x 12.0 are all rejected.

  • Given the resident: all read baseline, when re-measured, then inside tolerance — no cost for a feature not used. ✅ Verified as a by-product of the residency leg: resident: all is unchanged at 1 354 554 reads/sec uncapped, and every other db-bench leg still runs, which the keys-resident classes had briefly broken by forcing WO_DATA module-wide.
  • Given a resident: keys table larger than RAM, when read randomly, then its read cost is measured against the resident baseline on its own read path. ✅ Measured 2026-08-30 and the answer is qualified: 1.53× faster than letting the kernel swap under a cap that binds one and not the other — real, but nowhere near iteration 1's 273× swap figure would suggest, because cgroup limits charge the page cache, so moving rows to a file does not escape a container memory limit. The unambiguous win is footprint: 2.55× smaller resident set. Full method, numbers and the failed first attempt above; gated by residency.*.

Outstanding:

  • Given durable: true and no WO_DATA, when the program starts, then it refuses. (task 6 — today this combination silently discards every write)
  • Given the resident footprint crossing the budget, when it does, then a refusal naming the table and the annotation. (task 6)

Out Of Scope

  • Checkpoint and compaction — 3. Boot rebuilds the offset map by scanning the log until that lands, which is O(all history); 3's snapshot should persist the map.
  • Eviction and a resident row cache — 5. This iteration's tables are either fully resident or keys-only.
  • io_uring on the read path — a real question that only exists after this; noted in 4, deliberately not folded in.
  • transaction { } and @table feature flags — language iteration 18, approved spec, left whole.
  • Per-shard residency for volatile tables — a volatile table has no WAL, so it arguably need not live on the owner shard at all. Faster, and a different consistency story. Recorded as a candidate, not decided.
  • Converting an existing dataset between settings. Refuse on mismatch, do not convert — implemented in task 4.

Info — the forks, settled

  1. Two keys, not one enum. An enum needs a name per combination, and the brainstorm demonstrated the third name is unwriteable before its mechanism is decided.
  2. keys, not index. index: is already an argument key, so @table(index: [c], resident: index) read badly. all/keys also put both values on one axis — what row data stays resident. resident: none was rejected as overclaiming, since the indexes are very much resident.
  3. Optional with durable defaulting true, not mandatory. Mandatory would have touched 28 declarations, 13 corpus fixtures and 3 goldens; the README already says nothing is API-stable, so making it mandatory at 1.0 stays available.
  4. @unique on resident: keys is allowed, with its index unconditionally resident. Roughly doubles the resident index; stated at the declaration so the cost is visible.
  5. The budget is bytes, not rows — a text-heavy row and an Int-only row differ by 3.3× (measured, databasev2 1), so a row count cannot bound RAM.

History — two corrections worth keeping

The three-mode design was replaced. Earlier drafts had mode: ram | durable | cold. cold conflated two independent properties and could not be named honestly before its mechanism existed, and the developer's 120 GB-on-32 GB case showed the real axis was residency. Replaced by two keys, and principle 7 amended rather than worked around.

The "one real rewrite" was fiction. The spec claimed the on-disk record was pointer-bearing and that re-encoding it was this iteration's substantive engineering. That came from reading table.c's db_val_encode — which builds the in-memory slot — and inferring the file format from it. wal.c's enc_val has been flat since iteration 9. The task was deleted, not reduced.

The opposite half then turned out to be genuinely deep. With the format fine, the plan's storage steps still read as plumbing. Measured instead: wo_row_ptr returns a db_row * into a slab and has 11 call sites, table.c has 37 slab references, db.c:105-181 walks slabs for scans, enc_val serialises from the slab, and no operation exists that drops a row's payload while keeping its index entries. Hence the 5a–5d split. 5c and 5d need their own write-ups, and the two open design questions for 5c are whether the id hash stores offsets in place of slot indices or gains a parallel map, and what the new operation does about the unique shadows, which currently point at slots.