writeonce/docs/stories/databasev2/01-ram-ceiling-measurement.md
shoney.arickathil 746dc2b42b docs(databasev2): third track — the database beyond RAM, with per-table storage modes
- docs/stories/databasev2/, numbered from 1. Six PENDING database iterations
  moved from the language track and renumbered, keeping the old id in
  `was_language_iteration:` so a search for "iteration 32" still finds it:
  32 -> 3 WAL checkpoint, 23 -> 4 io_uring commit, 33 -> 7 single-file store,
  27 -> 8 query grammar, 20 -> 9 cross-program, 21 -> 10 keypair auth.
  Done work (9, 9b, 22) stays as v1 history; language 18 left whole
- the problem, read off the engine not guessed: rows are malloc'd slabs with
  addresses stable forever, NO eviction/spill/paging anywhere in database/src,
  the WAL never checkpoints so boot replays all history, and durability is one
  process-global WO_DATA so no table can say it matters more than another.
  An allocation failure IS a clean catchable WO_T_OOM — but swap thrash
  arrives first and carries no error signal at all, which is the real hazard
- four new iterations:
  1 measure the ceiling FIRST (curve not cliff; the three exits; kill -9 at
    exhaustion) — every later default should follow from a number
  2 `@table(mode: ram | durable | cold)` — the grammar ask. Small surface
    (Ast.table_cfg gains a key, the parser already rejects unknown args), big
    semantics: `durable` defaults so nothing changes silently, and the
    compiler refuses a durable row holding a `ref` into a ram table
  5 bounded tables + refuse/evict/back-pressure, shedding BEFORE the OS acts
  6 cold tiering — mostly forks, incl. whether the language surfaces the
    fault cost and whether @unique on cold is refused outright. A paged
    B-tree stays rejected: if tiering needs one, reject tiering
- 39 links repointed, link TEXT renumbered to track-local ids; arc gains one
  pointer row replacing the six moved; board + board-views cover three tracks
- linkcheck 0 broken / 0 anchors; no code blocks in any story

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-26 20:52:48 +02:00

8 KiB

track iteration status
databasev2 1 refine

databasev2 1 — the RAM ceiling: measure the breaking point before designing for it

Part of Story — databasev2: the database beyond RAM.

First because the repo's own doctrine says so. "Always inspect crashsites. Always measure. Never assume." Every later iteration in this track — the storage modes' defaults, the eviction policy, the tiering threshold — is a decision that should follow from a number. Right now nobody in this project can say what happens to a writeonce program at 90% of RAM, and designing tiering without that is guessing with extra steps.

Goals

  • Find the curve, not the cliff. Not "does it die" — it dies, everything does. What matters is the shape on the way down: at what fraction of RAM does p99 read latency leave its 1µs baseline, what does insert throughput do as slabs stop coming from a warm allocator, and how much warning is there between "fine" and "unusable".
  • Characterise all three exits. The engine can leave the happy path three ways and they are not equally survivable: a checked malloc failure (DB_ERR_OOM → WO_T_OOM, a catchable trap — the clean one), swap thrash (no trap, no error, just latency collapse — the dangerous one because nothing reports it), and the external OOM killer (SIGKILL, skipping every shutdown path). Establish which arrives first under realistic limits, because the answer determines whether the fix is back-pressure or eviction.
  • Prove the durability floor holds at the ceiling. Iteration 22's kill -9 battery proved acked writes survive under load. Re-run it at memory exhaustion, which is a different and nastier state — an allocation failure mid-commit is exactly where an ack-before-durable bug would hide.
  • Publish numbers others can build on. The output is a section in perf-targets.md and rows in bench/baseline.json, not a paragraph of prose. A measurement that only printed once is not a measurement.

Phases

Phase A — a workload that can actually reach the ceiling

  • Extend docs/examples/db-bench with a growth mode: insert until a target RSS fraction, holding row shape and index count constant so the variable is size alone.
  • Run it under an explicit memory limit (a cgroup or ulimit) rather than on a big box — "it survived on a 64 GB workstation" measures the workstation.
  • Record RSS against row count so the per-row overhead is known: slab headroom, the id hash, the secondary-index multimaps and the per-row engine-owned values (db_text, db_rec, db_multi, db_map are each their own allocation).
  • Verify: RSS growth is linear and its slope is written down; the run is reproducible twice within the tolerance policy iteration 22 established.

Phase B — the latency and throughput curve

  • Sample read p50/p99, query p99 and insert throughput at fixed fractions of the limit, so the result is a curve rather than two endpoints.
  • Separate the two effects deliberately: allocator pressure (still resident) and swap (no longer resident). They have different fixes and conflating them would send iteration 6 after the wrong one.
  • Include the DB-actor path, since a cross-shard statement's reply materialises a copy — memory pressure and the actor RPC interact and nobody has looked.
  • Verify: the curve is recorded per metric class with iteration 22's per-class tolerances; the swap onset point is identified, not interpolated.

Phase C — the three exits, deliberately triggered

  • Drive a checked allocation failure and confirm WO_T_OOM is catchable, the insert is refused whole, no partial row or index entry is left, and the process continues serving.
  • Drive swap thrash and record what a client sees. This is the case with no error signal at all, and naming it is most of the value of this iteration.
  • Drive the OOM killer under a cgroup limit and confirm what survives: replay the WAL and check every acked write is present.
  • Verify: the trap path leaves no torn state (row count and index agree after a refused insert); replay after SIGKILL at exhaustion loses no acked write.

Phase D — write it down where decisions get made

  • A perf-targets.md section with the curve, the swap onset, the per-row overhead and the exit characterisation.
  • Baseline rows for the growth metrics so a regression is caught by the existing gate rather than by a person remembering.
  • A short statement of what the numbers imply for iterations 2, 5 and 6 — which is the point of going first.
  • Verify: just db-bench green against the extended baseline; the gate bites when a growth metric is doctored.

Acceptance Criteria

  • Given the growth workload under a fixed memory limit, when it runs twice, then RSS-per-row agrees within the tolerance policy and the slope is recorded in perf-targets.md.
  • Given the workload at rising RAM fractions, when latency is sampled, then the fraction at which read p99 first leaves its baseline is identified as a measured point, not an estimate.
  • Given a deliberately induced allocation failure, when an insert is attempted, then it traps WO_T_OOM catchably, the table's row count is unchanged, every index agrees with the slab contents, and the process keeps serving subsequent requests.
  • Given swap thrash, when a client issues reads, then the observed degradation is quantified and the fact that no error is surfaced is recorded explicitly as a finding.
  • Given a cgroup limit and a workload that exceeds it, when the OOM killer fires, then replaying the WAL shows every acked write present — ack-after-fsync holding in the one shutdown path that skips all cleanup.
  • Given the extended baseline, when a growth metric is doctored, then just db-bench fails on exactly that metric.

Out Of Scope

  • Any fix. This iteration measures. Eviction is 5, tiering is 6, declared budgets are 2. Shipping a fix inside the measurement slice would remove the ability to tell whether it helped.
  • Changing the OOM behaviour. The checked-malloc-to-catchable-trap path is good and should not be touched; if the measurement finds a hole in it, that is a bug fix, reported separately.
  • A memory profiler or allocator instrumentation. Observability is language iteration 30. RSS from the OS and the existing time.ticks are enough for a curve.
  • Multi-machine or sharded-across-hosts scaling. One binary owns its data; cross-process is 9.
  • Comparing against SQLite at the ceiling. bench/compare/go-sqlite exists and the comparison would be interesting, but SQLite's whole architecture is the paged design this project rejected — the numbers would not inform any decision here.

Info

Forks the spec must settle:

  1. What is the limit mechanism for the harness? A cgroup v2 memory.max is the closest thing to how this would actually be deployed; ulimit -v is simpler but bounds address space rather than resident set, which for an engine that mallocs slabs is a materially different constraint. Leaning cgroup, and the campaign already runs off the fast path so the setup cost is acceptable.
  2. Which table shape is the reference? Per-row overhead depends heavily on whether fields are scalars or heap values — a Text column is a separate db_text allocation per row, so a text-heavy table and an Int-only table will produce very different slopes. Probably both, reported separately, because "bytes per row" is meaningless without saying which row.
  3. Is swap even in scope for the target deployment? If the intended answer is "run with swap off and let the OOM killer decide", the swap curve is informational rather than load-bearing — but that stance should be stated in the doctrine, not assumed. It also changes which exit iteration 5's back-pressure is defending against.