writeonce/docs/stories/databasev2/05-bounded-tables-eviction.md
shoney.arickathil 746dc2b42b docs(databasev2): third track — the database beyond RAM, with per-table storage modes
- docs/stories/databasev2/, numbered from 1. Six PENDING database iterations
  moved from the language track and renumbered, keeping the old id in
  `was_language_iteration:` so a search for "iteration 32" still finds it:
  32 -> 3 WAL checkpoint, 23 -> 4 io_uring commit, 33 -> 7 single-file store,
  27 -> 8 query grammar, 20 -> 9 cross-program, 21 -> 10 keypair auth.
  Done work (9, 9b, 22) stays as v1 history; language 18 left whole
- the problem, read off the engine not guessed: rows are malloc'd slabs with
  addresses stable forever, NO eviction/spill/paging anywhere in database/src,
  the WAL never checkpoints so boot replays all history, and durability is one
  process-global WO_DATA so no table can say it matters more than another.
  An allocation failure IS a clean catchable WO_T_OOM — but swap thrash
  arrives first and carries no error signal at all, which is the real hazard
- four new iterations:
  1 measure the ceiling FIRST (curve not cliff; the three exits; kill -9 at
    exhaustion) — every later default should follow from a number
  2 `@table(mode: ram | durable | cold)` — the grammar ask. Small surface
    (Ast.table_cfg gains a key, the parser already rejects unknown args), big
    semantics: `durable` defaults so nothing changes silently, and the
    compiler refuses a durable row holding a `ref` into a ram table
  5 bounded tables + refuse/evict/back-pressure, shedding BEFORE the OS acts
  6 cold tiering — mostly forks, incl. whether the language surfaces the
    fault cost and whether @unique on cold is refused outright. A paged
    B-tree stays rejected: if tiering needs one, reject tiering
- 39 links repointed, link TEXT renumbered to track-local ids; arc gains one
  pointer row replacing the six moved; board + board-views cover three tracks
- linkcheck 0 broken / 0 anchors; no code blocks in any story

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-26 20:52:48 +02:00

9.1 KiB

track iteration status
databasev2 5 refine

databasev2 5 — bounded tables and eviction: a declared budget, and back-pressure before the cliff

Part of Story — databasev2: the database beyond RAM. Needs 2 for the mode a bound attaches to, and 1 for the numbers that set a sane default.

The simpler half of the hard problem, done first on purpose. Evicting from a bounded resident table and evicting to disk are the same policy question with different destinations. Getting the policy right where the answer is "drop it" de-risks 6, where the answer is "write it somewhere and be able to find it again".

Goals

  • A table may declare a maximum. Rows, bytes, or both — a ram table that is a cache or a session store has a size the application is willing to spend, and today it has no way to say so. Unbounded growth in a table nobody intended to be large is the most common route to the ceiling iteration 1 measured.
  • Something defined happens at the bound. Today the answer is "grow until the process dies". The candidates are eviction (drop the least valuable row), refusal (trap, let the caller decide), and back-pressure (make the writer wait). Each is right for a different table, which argues for the policy being declared rather than chosen for the developer.
  • Back-pressure before the cliff, not at it. The dangerous exit iteration 1 characterises is swap thrash, which arrives with no error signal at all. A budget that is enforced at 100% has already lost; the value is in acting at a threshold, while there is still headroom to act.
  • Eviction that respects the engine's actual invariants. Rows have stable addresses forever, the free-slot list recycles slots, ids are never reused, and every secondary index and unique shadow must stay consistent with the slab. Eviction is wo_row_remove with a policy in front — it must go through the same choke point, not around it.

Phases

Phase A — declaring the bound

  • Extend the @table surface from 2 with a capacity and a policy. One new grammar arm, the same catalogued-diagnostic discipline, defaults that keep every existing table unbounded so nothing changes silently.
  • Decide whether a bound is legal on a durable table (fork 1) — evicting a row that was acked as durable is a promise being broken, and the answer is probably "only with an explicit, differently-named policy".
  • Verify: golden fixtures per policy; existing tables unchanged; illegal combinations refused at compile time with catalogued codes.

Phase B — the accounting

  • Track per-table size cheaply. Row count is free; bytes are not — per-row footprint includes the slab slot plus each heap value's own allocation (db_text, db_rec, db_multi, db_map), which iteration 1 will have quantified. Decide what is counted and be honest that it is an estimate of RSS, not RSS.
  • Expose it, because a budget nobody can observe is a budget nobody can tune. How it is exposed is fork 3.
  • Verify: accounting tracks a known workload within a stated error bound; deleting rows returns the accounting to its prior value (the free-slot list already makes this true of slots).

Phase C — the policies

  • Refuse: the bound is a hard ceiling, an insert past it traps catchably. The simplest correct behaviour and the right default for anything precious.
  • Evict: drop the least valuable row via wo_row_remove so indexes, unique shadows and the free-slot list all stay honest. The recency metadata this needs is fork 2 — and note the engine currently stores no per-row access time, so true LRU is not free.
  • Back-pressure: at a threshold below the bound, slow or park the writer. This composes with the fiber model (a parked writer blocks nobody) and is the only policy that addresses the swap-thrash exit rather than the allocation exit.
  • Verify: each policy behaves at the bound; after eviction every index agrees with the slab; a parked writer resumes and does not deadlock the shard.

Phase D — the pressure signal

  • A process-level threshold, not just per-table: when total resident size crosses a configured fraction, tables with an eviction policy start shedding before the allocator or the OS gets involved. This is the iteration's real contribution — turning an invisible failure into a managed one.
  • Decide precedence when several tables could shed (fork 4).
  • Verify: under the iteration-1 growth workload with the signal enabled, the process holds a steady state instead of walking into swap; the latency curve stays inside its baseline.

Phase E — gate it against the measurement

  • Re-run iteration 1's growth workload with bounds and policies configured. The proof is a before/after on the same harness: previously the curve degraded and the process died; now it plateaus.
  • Baseline rows for steady-state throughput under pressure.
  • Verify: just employee, just db-actor, just db-bench green; the gate bites on a doctored pressure metric.

Acceptance Criteria

  • Given a table bounded at N rows with policy refuse, when the N+1st insert is attempted, then it traps catchably, the row count stays N, and every index agrees with the slab.
  • Given a table bounded at N rows with policy evict, when the N+1st insert arrives, then exactly one row is evicted, the new row is present, the count is N, and no index or unique shadow references the evicted row.
  • Given an evicted row's id, when it is looked up, then it is absent — and its id is never reused by a later insert, preserving the invariant the id hash's tombstone sentinel depends on.
  • Given a bound expressed in bytes, when rows of a known shape are inserted, then the bound is honoured within the stated accounting error, and that error is documented rather than implied.
  • Given back-pressure configured at a threshold, when the threshold is crossed, then writers are slowed or parked, reads are unaffected, and no shard deadlocks.
  • Given the process-level pressure signal and the iteration-1 growth workload, when it runs to what previously exhausted memory, then the process reaches a steady state and read p99 stays within its baseline — the before/after that justifies the iteration.
  • Given a durable table, when an eviction policy is applied to it, then either it is refused at compile time or it is a distinctly named policy that says out loud it discards acked data.

Out Of Scope

  • Writing evicted rows anywhere — that is 6. Here eviction means the row is gone. Keeping the two apart is what makes the policy work reviewable on its own.
  • True LRU if it costs a write per read. Touching per-row metadata on every read would turn the 1µs read path into a write path — the same trap iteration 3's session touch has. An approximation (insertion order, a coarse clock, a sampled counter) is very likely the right answer and fork 2 should say so explicitly rather than defaulting to textbook LRU.
  • The TTL cache middleware — language iteration 18. Expiry by time is that; bounding by size is this. They compose.
  • Query-level result limits. take n already exists in the query surface.
  • Shrinking slabs back to the allocator. Slab addresses are stable forever by design and that invariant is load-bearing; reclaiming a slab whose rows were all evicted is a separate, delicate change with its own iteration if anyone wants it.

Info

Forks the spec must settle:

  1. May a durable table be bounded? Evicting an acked row contradicts the durability promise. But an audit table that must not grow forever is a real need, and the honest form of it is probably archival (iteration 6) rather than eviction. Leaning: bounds on durable are refused, and the need is redirected to 6.
  2. What is "least valuable"? No per-row access time exists today, so LRU costs a write per read. Candidates: insertion order (free — ids are already monotonic per table), a coarse epoch stamped on write only, or sampled approximation. Insertion order is FIFO not LRU, which is wrong for a cache and fine for a queue — so the policy name should say which it is rather than claiming "LRU" and delivering FIFO.
  3. How is size observed? Without observability (language iteration 30) there is no metrics endpoint to publish it on. Options: a builtin returning a table's current size, a @table-derived query, or stderr on threshold crossing. A builtin is the smallest thing that makes the feature tunable by the program that owns the budget.
  4. Precedence when several tables can shed. Largest first is simple; the application's own priority order is more correct and needs a way to express it. Proportional shedding is fairest and hardest to reason about. This decides whether the pressure signal is predictable enough to trust.