- docs/stories/databasev2/, numbered from 1. Six PENDING database iterations
moved from the language track and renumbered, keeping the old id in
`was_language_iteration:` so a search for "iteration 32" still finds it:
32 -> 3 WAL checkpoint, 23 -> 4 io_uring commit, 33 -> 7 single-file store,
27 -> 8 query grammar, 20 -> 9 cross-program, 21 -> 10 keypair auth.
Done work (9, 9b, 22) stays as v1 history; language 18 left whole
- the problem, read off the engine not guessed: rows are malloc'd slabs with
addresses stable forever, NO eviction/spill/paging anywhere in database/src,
the WAL never checkpoints so boot replays all history, and durability is one
process-global WO_DATA so no table can say it matters more than another.
An allocation failure IS a clean catchable WO_T_OOM — but swap thrash
arrives first and carries no error signal at all, which is the real hazard
- four new iterations:
1 measure the ceiling FIRST (curve not cliff; the three exits; kill -9 at
exhaustion) — every later default should follow from a number
2 `@table(mode: ram | durable | cold)` — the grammar ask. Small surface
(Ast.table_cfg gains a key, the parser already rejects unknown args), big
semantics: `durable` defaults so nothing changes silently, and the
compiler refuses a durable row holding a `ref` into a ram table
5 bounded tables + refuse/evict/back-pressure, shedding BEFORE the OS acts
6 cold tiering — mostly forks, incl. whether the language surfaces the
fault cost and whether @unique on cold is refused outright. A paged
B-tree stays rejected: if tiering needs one, reject tiering
- 39 links repointed, link TEXT renumbered to track-local ids; arc gains one
pointer row replacing the six moved; board + board-views cover three tracks
- linkcheck 0 broken / 0 anchors; no code blocks in any story
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
9.1 KiB
9.1 KiB
| track | iteration | status |
|---|---|---|
| databasev2 | 5 | refine |
databasev2 5 — bounded tables and eviction: a declared budget, and back-pressure before the cliff
Part of Story — databasev2: the database beyond RAM. Needs 2 for the mode a bound attaches to, and 1 for the numbers that set a sane default.
The simpler half of the hard problem, done first on purpose. Evicting from a bounded resident table and evicting to disk are the same policy question with different destinations. Getting the policy right where the answer is "drop it" de-risks 6, where the answer is "write it somewhere and be able to find it again".
Goals
- A table may declare a maximum. Rows, bytes, or both — a
ramtable that is a cache or a session store has a size the application is willing to spend, and today it has no way to say so. Unbounded growth in a table nobody intended to be large is the most common route to the ceiling iteration 1 measured. - Something defined happens at the bound. Today the answer is "grow until the process dies". The candidates are eviction (drop the least valuable row), refusal (trap, let the caller decide), and back-pressure (make the writer wait). Each is right for a different table, which argues for the policy being declared rather than chosen for the developer.
- Back-pressure before the cliff, not at it. The dangerous exit iteration 1 characterises is swap thrash, which arrives with no error signal at all. A budget that is enforced at 100% has already lost; the value is in acting at a threshold, while there is still headroom to act.
- Eviction that respects the engine's actual invariants. Rows have stable
addresses forever, the free-slot list recycles slots, ids are never reused, and
every secondary index and unique shadow must stay consistent with the slab.
Eviction is
wo_row_removewith a policy in front — it must go through the same choke point, not around it.
Phases
Phase A — declaring the bound
- Extend the
@tablesurface from 2 with a capacity and a policy. One new grammar arm, the same catalogued-diagnostic discipline, defaults that keep every existing table unbounded so nothing changes silently. - Decide whether a bound is legal on a
durabletable (fork 1) — evicting a row that was acked as durable is a promise being broken, and the answer is probably "only with an explicit, differently-named policy". - Verify: golden fixtures per policy; existing tables unchanged; illegal combinations refused at compile time with catalogued codes.
Phase B — the accounting
- Track per-table size cheaply. Row count is free; bytes are not — per-row
footprint includes the slab slot plus each heap value's own allocation
(
db_text,db_rec,db_multi,db_map), which iteration 1 will have quantified. Decide what is counted and be honest that it is an estimate of RSS, not RSS. - Expose it, because a budget nobody can observe is a budget nobody can tune. How it is exposed is fork 3.
- Verify: accounting tracks a known workload within a stated error bound; deleting rows returns the accounting to its prior value (the free-slot list already makes this true of slots).
Phase C — the policies
- Refuse: the bound is a hard ceiling, an insert past it traps catchably. The simplest correct behaviour and the right default for anything precious.
- Evict: drop the least valuable row via
wo_row_removeso indexes, unique shadows and the free-slot list all stay honest. The recency metadata this needs is fork 2 — and note the engine currently stores no per-row access time, so true LRU is not free. - Back-pressure: at a threshold below the bound, slow or park the writer. This composes with the fiber model (a parked writer blocks nobody) and is the only policy that addresses the swap-thrash exit rather than the allocation exit.
- Verify: each policy behaves at the bound; after eviction every index agrees with the slab; a parked writer resumes and does not deadlock the shard.
Phase D — the pressure signal
- A process-level threshold, not just per-table: when total resident size crosses a configured fraction, tables with an eviction policy start shedding before the allocator or the OS gets involved. This is the iteration's real contribution — turning an invisible failure into a managed one.
- Decide precedence when several tables could shed (fork 4).
- Verify: under the iteration-1 growth workload with the signal enabled, the process holds a steady state instead of walking into swap; the latency curve stays inside its baseline.
Phase E — gate it against the measurement
- Re-run iteration 1's growth workload with bounds and policies configured. The proof is a before/after on the same harness: previously the curve degraded and the process died; now it plateaus.
- Baseline rows for steady-state throughput under pressure.
- Verify:
just employee,just db-actor,just db-benchgreen; the gate bites on a doctored pressure metric.
Acceptance Criteria
- Given a table bounded at N rows with policy
refuse, when the N+1st insert is attempted, then it traps catchably, the row count stays N, and every index agrees with the slab. - Given a table bounded at N rows with policy
evict, when the N+1st insert arrives, then exactly one row is evicted, the new row is present, the count is N, and no index or unique shadow references the evicted row. - Given an evicted row's id, when it is looked up, then it is absent — and its id is never reused by a later insert, preserving the invariant the id hash's tombstone sentinel depends on.
- Given a bound expressed in bytes, when rows of a known shape are inserted, then the bound is honoured within the stated accounting error, and that error is documented rather than implied.
- Given back-pressure configured at a threshold, when the threshold is crossed, then writers are slowed or parked, reads are unaffected, and no shard deadlocks.
- Given the process-level pressure signal and the iteration-1 growth workload, when it runs to what previously exhausted memory, then the process reaches a steady state and read p99 stays within its baseline — the before/after that justifies the iteration.
- Given a
durabletable, when an eviction policy is applied to it, then either it is refused at compile time or it is a distinctly named policy that says out loud it discards acked data.
Out Of Scope
- Writing evicted rows anywhere — that is 6. Here eviction means the row is gone. Keeping the two apart is what makes the policy work reviewable on its own.
- True LRU if it costs a write per read. Touching per-row metadata on every read would turn the 1µs read path into a write path — the same trap iteration 3's session touch has. An approximation (insertion order, a coarse clock, a sampled counter) is very likely the right answer and fork 2 should say so explicitly rather than defaulting to textbook LRU.
- The TTL cache middleware — language iteration 18. Expiry by time is that; bounding by size is this. They compose.
- Query-level result limits.
take nalready exists in the query surface. - Shrinking slabs back to the allocator. Slab addresses are stable forever by design and that invariant is load-bearing; reclaiming a slab whose rows were all evicted is a separate, delicate change with its own iteration if anyone wants it.
Info
Forks the spec must settle:
- May a
durabletable be bounded? Evicting an acked row contradicts the durability promise. But an audit table that must not grow forever is a real need, and the honest form of it is probably archival (iteration 6) rather than eviction. Leaning: bounds ondurableare refused, and the need is redirected to 6. - What is "least valuable"? No per-row access time exists today, so LRU costs a write per read. Candidates: insertion order (free — ids are already monotonic per table), a coarse epoch stamped on write only, or sampled approximation. Insertion order is FIFO not LRU, which is wrong for a cache and fine for a queue — so the policy name should say which it is rather than claiming "LRU" and delivering FIFO.
- How is size observed? Without observability (language iteration 30) there
is no metrics endpoint to publish it on. Options: a builtin returning a
table's current size, a
@table-derived query, or stderr on threshold crossing. A builtin is the smallest thing that makes the feature tunable by the program that owns the budget. - Precedence when several tables can shed. Largest first is simple; the application's own priority order is more correct and needs a way to express it. Proportional shedding is fairest and hardest to reason about. This decides whether the pressure signal is predictable enough to trust.