Audit found 4 of 10 iterations citing it1 and 4 carrying stale claims the
measurement contradicts.
- 05: framing was contradicted, not merely incomplete. Its goal expected a
gradient to detect ("back-pressure before the cliff"); there is no cliff
— SIGKILL with swap off, exit 0 with swap on, and read latency STEPS
(1us -> 487us) rather than departing. Heading and goal rewritten; the
measurement makes the goal stronger, not weaker
- 05: budget must be bytes — 3.3x footprint spread — with headroom for
index doublings, else it fires during a rehash
- 05: new goal — eviction policy QUALITY is decisive, since getting the
resident set wrong costs 273x, not a few percent
- 06: its revival question now has a reference point. 273x is the KERNEL
SWAP path; `resident: keys` preads via page cache and must beat it. This
file revives only if 5c/5d lands near 273x rather than well below
- 04: write path is not where pressure bites (append ~1%, read 273x), so
the io_uring question that matters is iteration 2's deferred read-path
one, not group-commit
- 00-story: problem statement asserted the store "refuses the insert
rather than dying". Corrected in place — a banner above it was not
enough, a skimmer never reaches it
- residency spec: "swap thrash and the OOM killer" named exits that were
not measured; replaced with silence-or-a-corpse
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
197 lines
11 KiB
Markdown
197 lines
11 KiB
Markdown
---
|
||
track: databasev2
|
||
iteration: "5"
|
||
status: pending
|
||
readiness: refine
|
||
---
|
||
|
||
# databasev2 5 — bounded tables and eviction: a declared budget, and back-pressure at a declared threshold
|
||
|
||
> Part of [Story — databasev2: the database beyond RAM](00-story.md).
|
||
> Needs [2](02-table-storage-modes.md) for the mode a bound attaches to, and
|
||
> [1](01-ram-ceiling-measurement.md), whose numbers landed 2026-08-27 and
|
||
> **corrected this iteration's framing** — see the third goal.
|
||
>
|
||
> **The simpler half of the hard problem, done first on purpose.** Evicting from
|
||
> a bounded resident table and evicting to disk are the same policy question with
|
||
> different destinations. Getting the policy right where the answer is "drop it"
|
||
> de-risks [6](06-cold-tiering.md), where the answer is "write it somewhere and
|
||
> be able to find it again".
|
||
|
||
## Goals
|
||
|
||
- **A table may declare a maximum.** Rows, bytes, or both — a `ram` table that
|
||
is a cache or a session store has a size the application is willing to spend,
|
||
and today it has no way to say so. Unbounded growth in a table nobody intended
|
||
to be large is the most common route to the ceiling iteration 1 measured.
|
||
- **Something defined happens at the bound.** Today the answer is "grow until
|
||
the process dies". The candidates are eviction (drop the least valuable row),
|
||
refusal (trap, let the caller decide), and back-pressure (make the writer
|
||
wait). Each is right for a different table, which argues for the policy being
|
||
declared rather than chosen for the developer.
|
||
- **Back-pressure at a declared threshold — because there is no cliff to be
|
||
before.** This goal was written expecting a gradient to detect. Iteration 1
|
||
measured (2026-08-27) that no such gradient exists, which makes the goal
|
||
*stronger*, not weaker:
|
||
- Exceeding RAM **without** swap is **SIGKILL, signal 9** — no trap, no
|
||
diagnostic. Table storage has no checked ceiling, and under
|
||
`vm.overcommit_memory = 0` its `malloc` succeeds and the kernel kills on
|
||
page touch, so the checked path never runs.
|
||
- Exceeding RAM **with** swap returns **exit 0** and keeps serving from disk.
|
||
An append-mostly workload pays **~1%** (148 s vs 150 s uncapped for 900k
|
||
rows), so "swap thrash" — which this goal previously named as the dangerous
|
||
exit — is not what happens on the write path at all.
|
||
- Read latency does not *depart*, it **steps**: 1 µs resident to 487 µs
|
||
over-cap with nothing in between.
|
||
|
||
So there is no early-warning signal anywhere to react to — not an error, not a
|
||
latency knee. A budget enforced at 100% has not merely "already lost"; it can
|
||
never fire, because the process is dead or silently fine. **Only a declared
|
||
threshold can speak, and it must be declared in bytes** — footprint is
|
||
96.5–100 B/row Int-only against 320.6–324 B/row text-heavy, **3.3× apart**, so
|
||
a row count cannot bound RAM. Leave headroom for index doublings, which are
|
||
transient RSS steps (measured at ~24k and ~48k rows): a budget without headroom
|
||
fires during a rehash instead of at a real threshold.
|
||
- **Eviction policy QUALITY is decisive, not incidental.** Iteration 1 measured
|
||
random reads over an oversized table at **273× slower** than resident
|
||
(1 851 166 vs 6 771 reads/s; p99 1 µs vs 487 µs). That is the cost of getting
|
||
the resident set wrong, so the gap between a good policy and a careless one is
|
||
not a few percent — it is the difference between a working system and an
|
||
unusable one. Whatever policy ships must be measured against that spread, not
|
||
merely shown to be correct.
|
||
- **Eviction that respects the engine's actual invariants.** Rows have stable
|
||
addresses forever, the free-slot list recycles slots, ids are never reused, and
|
||
every secondary index and unique shadow must stay consistent with the slab.
|
||
Eviction is `wo_row_remove` with a policy in front — it must go through the
|
||
same choke point, not around it.
|
||
|
||
## Phases
|
||
|
||
### Phase A — declaring the bound
|
||
|
||
- Extend the `@table` surface from [2](02-table-storage-modes.md) with a
|
||
capacity and a policy. One new grammar arm, the same catalogued-diagnostic
|
||
discipline, defaults that keep every existing table unbounded so nothing
|
||
changes silently.
|
||
- Decide whether a bound is legal on a `durable` table (fork 1) — evicting a row
|
||
that was acked as durable is a promise being broken, and the answer is
|
||
probably "only with an explicit, differently-named policy".
|
||
- Verify: golden fixtures per policy; existing tables unchanged; illegal
|
||
combinations refused at compile time with catalogued codes.
|
||
|
||
### Phase B — the accounting
|
||
|
||
- Track per-table size cheaply. Row count is free; bytes are not — per-row
|
||
footprint includes the slab slot plus each heap value's own allocation
|
||
(`db_text`, `db_rec`, `db_multi`, `db_map`), which iteration 1 will have
|
||
quantified. Decide what is counted and be honest that it is an estimate of RSS,
|
||
not RSS.
|
||
- Expose it, because a budget nobody can observe is a budget nobody can tune.
|
||
How it is exposed is fork 3.
|
||
- Verify: accounting tracks a known workload within a stated error bound;
|
||
deleting rows returns the accounting to its prior value (the free-slot list
|
||
already makes this true of slots).
|
||
|
||
### Phase C — the policies
|
||
|
||
- **Refuse**: the bound is a hard ceiling, an insert past it traps catchably.
|
||
The simplest correct behaviour and the right default for anything precious.
|
||
- **Evict**: drop the least valuable row via `wo_row_remove` so indexes, unique
|
||
shadows and the free-slot list all stay honest. The recency metadata this needs
|
||
is fork 2 — and note the engine currently stores no per-row access time, so
|
||
true LRU is not free.
|
||
- **Back-pressure**: at a threshold below the bound, slow or park the writer.
|
||
This composes with the fiber model (a parked writer blocks nobody) and is the
|
||
only policy that addresses the swap-thrash exit rather than the allocation
|
||
exit.
|
||
- Verify: each policy behaves at the bound; after eviction every index agrees
|
||
with the slab; a parked writer resumes and does not deadlock the shard.
|
||
|
||
### Phase D — the pressure signal
|
||
|
||
- A process-level threshold, not just per-table: when total resident size
|
||
crosses a configured fraction, tables with an eviction policy start shedding
|
||
**before** the allocator or the OS gets involved. This is the iteration's real
|
||
contribution — turning an invisible failure into a managed one.
|
||
- Decide precedence when several tables could shed (fork 4).
|
||
- Verify: under the iteration-1 growth workload with the signal enabled, the
|
||
process holds a steady state instead of walking into swap; the latency curve
|
||
stays inside its baseline.
|
||
|
||
### Phase E — gate it against the measurement
|
||
|
||
- Re-run iteration 1's growth workload with bounds and policies configured. The
|
||
proof is a before/after on the same harness: previously the curve degraded and
|
||
the process died; now it plateaus.
|
||
- Baseline rows for steady-state throughput under pressure.
|
||
- Verify: `just employee`, `just db-actor`, `just db-bench` green; the gate bites
|
||
on a doctored pressure metric.
|
||
|
||
## Acceptance Criteria
|
||
|
||
- **Given** a table bounded at N rows with policy `refuse`, **when** the N+1st
|
||
insert is attempted, **then** it traps catchably, the row count stays N, and
|
||
every index agrees with the slab.
|
||
- **Given** a table bounded at N rows with policy `evict`, **when** the N+1st
|
||
insert arrives, **then** exactly one row is evicted, the new row is present,
|
||
the count is N, and no index or unique shadow references the evicted row.
|
||
- **Given** an evicted row's id, **when** it is looked up, **then** it is absent
|
||
— and its id is never reused by a later insert, preserving the invariant the
|
||
id hash's tombstone sentinel depends on.
|
||
- **Given** a bound expressed in bytes, **when** rows of a known shape are
|
||
inserted, **then** the bound is honoured within the stated accounting error,
|
||
and that error is documented rather than implied.
|
||
- **Given** back-pressure configured at a threshold, **when** the threshold is
|
||
crossed, **then** writers are slowed or parked, reads are unaffected, and no
|
||
shard deadlocks.
|
||
- **Given** the process-level pressure signal and the iteration-1 growth
|
||
workload, **when** it runs to what previously exhausted memory, **then** the
|
||
process reaches a steady state and read p99 stays within its baseline — the
|
||
before/after that justifies the iteration.
|
||
- **Given** a `durable` table, **when** an eviction policy is applied to it,
|
||
**then** either it is refused at compile time or it is a distinctly named
|
||
policy that says out loud it discards acked data.
|
||
|
||
## Out Of Scope
|
||
|
||
- **Writing evicted rows anywhere** — that is [6](06-cold-tiering.md). Here
|
||
eviction means the row is gone. Keeping the two apart is what makes the policy
|
||
work reviewable on its own.
|
||
- **True LRU if it costs a write per read.** Touching per-row metadata on every
|
||
read would turn the 1µs read path into a write path — the same trap iteration
|
||
3's session touch has. An approximation (insertion order, a coarse clock, a
|
||
sampled counter) is very likely the right answer and fork 2 should say so
|
||
explicitly rather than defaulting to textbook LRU.
|
||
- **The TTL cache middleware** — language
|
||
[iteration 18](../language-runtime-database/18-memory-db-features.md). Expiry
|
||
by *time* is that; bounding by *size* is this. They compose.
|
||
- **Query-level result limits.** `take n` already exists in the query surface.
|
||
- **Shrinking slabs back to the allocator.** Slab addresses are stable forever
|
||
by design and that invariant is load-bearing; reclaiming a slab whose rows were
|
||
all evicted is a separate, delicate change with its own iteration if anyone
|
||
wants it.
|
||
|
||
## Info
|
||
|
||
Forks the spec must settle:
|
||
|
||
1. **May a `durable` table be bounded?** Evicting an acked row contradicts the
|
||
durability promise. But an audit table that must not grow forever is a real
|
||
need, and the honest form of it is probably archival (iteration 6) rather than
|
||
eviction. Leaning: bounds on `durable` are refused, and the need is redirected
|
||
to 6.
|
||
2. **What is "least valuable"?** No per-row access time exists today, so LRU
|
||
costs a write per read. Candidates: insertion order (free — ids are already
|
||
monotonic per table), a coarse epoch stamped on write only, or sampled
|
||
approximation. Insertion order is FIFO not LRU, which is wrong for a cache
|
||
and fine for a queue — so the policy name should say which it is rather than
|
||
claiming "LRU" and delivering FIFO.
|
||
3. **How is size observed?** Without observability (language iteration 30) there
|
||
is no metrics endpoint to publish it on. Options: a builtin returning a
|
||
table's current size, a `@table`-derived query, or stderr on threshold
|
||
crossing. A builtin is the smallest thing that makes the feature tunable by
|
||
the program that owns the budget.
|
||
4. **Precedence when several tables can shed.** Largest first is simple; the
|
||
application's own priority order is more correct and needs a way to express
|
||
it. Proportional shedding is fairest and hardest to reason about. This
|
||
decides whether the pressure signal is predictable enough to trust.
|