Audit found 4 of 10 iterations citing it1 and 4 carrying stale claims the
measurement contradicts.
- 05: framing was contradicted, not merely incomplete. Its goal expected a
gradient to detect ("back-pressure before the cliff"); there is no cliff
— SIGKILL with swap off, exit 0 with swap on, and read latency STEPS
(1us -> 487us) rather than departing. Heading and goal rewritten; the
measurement makes the goal stronger, not weaker
- 05: budget must be bytes — 3.3x footprint spread — with headroom for
index doublings, else it fires during a rehash
- 05: new goal — eviction policy QUALITY is decisive, since getting the
resident set wrong costs 273x, not a few percent
- 06: its revival question now has a reference point. 273x is the KERNEL
SWAP path; `resident: keys` preads via page cache and must beat it. This
file revives only if 5c/5d lands near 273x rather than well below
- 04: write path is not where pressure bites (append ~1%, read 273x), so
the io_uring question that matters is iteration 2's deferred read-path
one, not group-commit
- 00-story: problem statement asserted the store "refuses the insert
rather than dying". Corrected in place — a banner above it was not
enough, a skimmer never reaches it
- residency spec: "swap thrash and the OOM killer" named exits that were
not measured; replaced with silence-or-a-corpse
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
11 KiB
| track | iteration | status | readiness |
|---|---|---|---|
| databasev2 | 5 | pending | refine |
databasev2 5 — bounded tables and eviction: a declared budget, and back-pressure at a declared threshold
Part of Story — databasev2: the database beyond RAM. Needs 2 for the mode a bound attaches to, and 1, whose numbers landed 2026-08-27 and corrected this iteration's framing — see the third goal.
The simpler half of the hard problem, done first on purpose. Evicting from a bounded resident table and evicting to disk are the same policy question with different destinations. Getting the policy right where the answer is "drop it" de-risks 6, where the answer is "write it somewhere and be able to find it again".
Goals
-
A table may declare a maximum. Rows, bytes, or both — a
ramtable that is a cache or a session store has a size the application is willing to spend, and today it has no way to say so. Unbounded growth in a table nobody intended to be large is the most common route to the ceiling iteration 1 measured. -
Something defined happens at the bound. Today the answer is "grow until the process dies". The candidates are eviction (drop the least valuable row), refusal (trap, let the caller decide), and back-pressure (make the writer wait). Each is right for a different table, which argues for the policy being declared rather than chosen for the developer.
-
Back-pressure at a declared threshold — because there is no cliff to be before. This goal was written expecting a gradient to detect. Iteration 1 measured (2026-08-27) that no such gradient exists, which makes the goal stronger, not weaker:
- Exceeding RAM without swap is SIGKILL, signal 9 — no trap, no
diagnostic. Table storage has no checked ceiling, and under
vm.overcommit_memory = 0itsmallocsucceeds and the kernel kills on page touch, so the checked path never runs. - Exceeding RAM with swap returns exit 0 and keeps serving from disk. An append-mostly workload pays ~1% (148 s vs 150 s uncapped for 900k rows), so "swap thrash" — which this goal previously named as the dangerous exit — is not what happens on the write path at all.
- Read latency does not depart, it steps: 1 µs resident to 487 µs over-cap with nothing in between.
So there is no early-warning signal anywhere to react to — not an error, not a latency knee. A budget enforced at 100% has not merely "already lost"; it can never fire, because the process is dead or silently fine. Only a declared threshold can speak, and it must be declared in bytes — footprint is 96.5–100 B/row Int-only against 320.6–324 B/row text-heavy, 3.3× apart, so a row count cannot bound RAM. Leave headroom for index doublings, which are transient RSS steps (measured at ~24k and ~48k rows): a budget without headroom fires during a rehash instead of at a real threshold.
- Exceeding RAM without swap is SIGKILL, signal 9 — no trap, no
diagnostic. Table storage has no checked ceiling, and under
-
Eviction policy QUALITY is decisive, not incidental. Iteration 1 measured random reads over an oversized table at 273× slower than resident (1 851 166 vs 6 771 reads/s; p99 1 µs vs 487 µs). That is the cost of getting the resident set wrong, so the gap between a good policy and a careless one is not a few percent — it is the difference between a working system and an unusable one. Whatever policy ships must be measured against that spread, not merely shown to be correct.
-
Eviction that respects the engine's actual invariants. Rows have stable addresses forever, the free-slot list recycles slots, ids are never reused, and every secondary index and unique shadow must stay consistent with the slab. Eviction is
wo_row_removewith a policy in front — it must go through the same choke point, not around it.
Phases
Phase A — declaring the bound
- Extend the
@tablesurface from 2 with a capacity and a policy. One new grammar arm, the same catalogued-diagnostic discipline, defaults that keep every existing table unbounded so nothing changes silently. - Decide whether a bound is legal on a
durabletable (fork 1) — evicting a row that was acked as durable is a promise being broken, and the answer is probably "only with an explicit, differently-named policy". - Verify: golden fixtures per policy; existing tables unchanged; illegal combinations refused at compile time with catalogued codes.
Phase B — the accounting
- Track per-table size cheaply. Row count is free; bytes are not — per-row
footprint includes the slab slot plus each heap value's own allocation
(
db_text,db_rec,db_multi,db_map), which iteration 1 will have quantified. Decide what is counted and be honest that it is an estimate of RSS, not RSS. - Expose it, because a budget nobody can observe is a budget nobody can tune. How it is exposed is fork 3.
- Verify: accounting tracks a known workload within a stated error bound; deleting rows returns the accounting to its prior value (the free-slot list already makes this true of slots).
Phase C — the policies
- Refuse: the bound is a hard ceiling, an insert past it traps catchably. The simplest correct behaviour and the right default for anything precious.
- Evict: drop the least valuable row via
wo_row_removeso indexes, unique shadows and the free-slot list all stay honest. The recency metadata this needs is fork 2 — and note the engine currently stores no per-row access time, so true LRU is not free. - Back-pressure: at a threshold below the bound, slow or park the writer. This composes with the fiber model (a parked writer blocks nobody) and is the only policy that addresses the swap-thrash exit rather than the allocation exit.
- Verify: each policy behaves at the bound; after eviction every index agrees with the slab; a parked writer resumes and does not deadlock the shard.
Phase D — the pressure signal
- A process-level threshold, not just per-table: when total resident size crosses a configured fraction, tables with an eviction policy start shedding before the allocator or the OS gets involved. This is the iteration's real contribution — turning an invisible failure into a managed one.
- Decide precedence when several tables could shed (fork 4).
- Verify: under the iteration-1 growth workload with the signal enabled, the process holds a steady state instead of walking into swap; the latency curve stays inside its baseline.
Phase E — gate it against the measurement
- Re-run iteration 1's growth workload with bounds and policies configured. The proof is a before/after on the same harness: previously the curve degraded and the process died; now it plateaus.
- Baseline rows for steady-state throughput under pressure.
- Verify:
just employee,just db-actor,just db-benchgreen; the gate bites on a doctored pressure metric.
Acceptance Criteria
- Given a table bounded at N rows with policy
refuse, when the N+1st insert is attempted, then it traps catchably, the row count stays N, and every index agrees with the slab. - Given a table bounded at N rows with policy
evict, when the N+1st insert arrives, then exactly one row is evicted, the new row is present, the count is N, and no index or unique shadow references the evicted row. - Given an evicted row's id, when it is looked up, then it is absent — and its id is never reused by a later insert, preserving the invariant the id hash's tombstone sentinel depends on.
- Given a bound expressed in bytes, when rows of a known shape are inserted, then the bound is honoured within the stated accounting error, and that error is documented rather than implied.
- Given back-pressure configured at a threshold, when the threshold is crossed, then writers are slowed or parked, reads are unaffected, and no shard deadlocks.
- Given the process-level pressure signal and the iteration-1 growth workload, when it runs to what previously exhausted memory, then the process reaches a steady state and read p99 stays within its baseline — the before/after that justifies the iteration.
- Given a
durabletable, when an eviction policy is applied to it, then either it is refused at compile time or it is a distinctly named policy that says out loud it discards acked data.
Out Of Scope
- Writing evicted rows anywhere — that is 6. Here eviction means the row is gone. Keeping the two apart is what makes the policy work reviewable on its own.
- True LRU if it costs a write per read. Touching per-row metadata on every read would turn the 1µs read path into a write path — the same trap iteration 3's session touch has. An approximation (insertion order, a coarse clock, a sampled counter) is very likely the right answer and fork 2 should say so explicitly rather than defaulting to textbook LRU.
- The TTL cache middleware — language iteration 18. Expiry by time is that; bounding by size is this. They compose.
- Query-level result limits.
take nalready exists in the query surface. - Shrinking slabs back to the allocator. Slab addresses are stable forever by design and that invariant is load-bearing; reclaiming a slab whose rows were all evicted is a separate, delicate change with its own iteration if anyone wants it.
Info
Forks the spec must settle:
- May a
durabletable be bounded? Evicting an acked row contradicts the durability promise. But an audit table that must not grow forever is a real need, and the honest form of it is probably archival (iteration 6) rather than eviction. Leaning: bounds ondurableare refused, and the need is redirected to 6. - What is "least valuable"? No per-row access time exists today, so LRU costs a write per read. Candidates: insertion order (free — ids are already monotonic per table), a coarse epoch stamped on write only, or sampled approximation. Insertion order is FIFO not LRU, which is wrong for a cache and fine for a queue — so the policy name should say which it is rather than claiming "LRU" and delivering FIFO.
- How is size observed? Without observability (language iteration 30) there
is no metrics endpoint to publish it on. Options: a builtin returning a
table's current size, a
@table-derived query, or stderr on threshold crossing. A builtin is the smallest thing that makes the feature tunable by the program that owns the budget. - Precedence when several tables can shed. Largest first is simple; the application's own priority order is more correct and needs a way to express it. Proportional shedding is fairest and hardest to reason about. This decides whether the pressure signal is predictable enough to trust.