docs(databasev2): propagate iteration 1's findings to every consumer

Audit found 4 of 10 iterations citing it1 and 4 carrying stale claims the
measurement contradicts.

- 05: framing was contradicted, not merely incomplete. Its goal expected a
  gradient to detect ("back-pressure before the cliff"); there is no cliff
  — SIGKILL with swap off, exit 0 with swap on, and read latency STEPS
  (1us -> 487us) rather than departing. Heading and goal rewritten; the
  measurement makes the goal stronger, not weaker
- 05: budget must be bytes — 3.3x footprint spread — with headroom for
  index doublings, else it fires during a rehash
- 05: new goal — eviction policy QUALITY is decisive, since getting the
  resident set wrong costs 273x, not a few percent
- 06: its revival question now has a reference point. 273x is the KERNEL
  SWAP path; `resident: keys` preads via page cache and must beat it. This
  file revives only if 5c/5d lands near 273x rather than well below
- 04: write path is not where pressure bites (append ~1%, read 273x), so
  the io_uring question that matters is iteration 2's deferred read-path
  one, not group-commit
- 00-story: problem statement asserted the store "refuses the insert
  rather than dying". Corrected in place — a banner above it was not
  enough, a skimmer never reaches it
- residency spec: "swap thrash and the OOM killer" named exits that were
  not measured; replaced with silence-or-a-corpse

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
shoney.arickathil 2026-08-27 22:47:37 +02:00
parent 5b1a8c96a1
commit 3507aafea3
5 changed files with 86 additions and 26 deletions

View file

@ -45,26 +45,36 @@ Worth being precise, because the failure mode determines the fix — and the goo
news is that the engine's own behaviour is clean:
**Corrected 2026-08-27 by measurement.** This section used to open "an
allocation failure is a catchable trap, not a crash", and that is true only of
the VM arena. Table storage has no ceiling, and with `vm.overcommit_memory = 0`
its `malloc` never fails — the process is **SIGKILLed** (rc=137, measured at
360 000 rows under a 64 MiB cap). The checked path below is real, but it is the
arena's, not the store's. See [iteration 1](01-ram-ceiling-measurement.md).
allocation failure is a catchable trap, not a crash". That is true only of the VM
object arena, whose `WO_HEAP_MB` ceiling is checked and does trap
(`trap 4 … out of memory`, exit 1). **Table storage has no ceiling at all** —
details below, measured. See [iteration 1](01-ram-ceiling-measurement.md).
Every `malloc` in
the row encoder is checked and jumps to an `oom` label; `DB_ERR_OOM` maps to
`WO_T_OOM`, which a program can `try`/`catch`. So a writeonce program that runs
out of memory *refuses the insert* rather than corrupting or dying. That is a
much better starting position than most engines have.
`WO_T_OOM`, which a program can `try`/`catch`. On paper a writeonce program that
runs out of memory *refuses the insert* rather than corrupting or dying.
**But the trap is almost never what a real deployment hits first.** Long before
`malloc` returns NULL, the box starts swapping, and a RAM-authoritative database
on swap is the worst of both worlds: it has paid for in-memory data structures
and is now serving them from disk with no read path designed for that. On a
cgroup-limited host the OOM killer arrives instead, and an external `SIGKILL` is
the one shutdown path that skips every guarantee the WAL was written to provide —
though ack-after-fsync means acked writes still survive; iteration 22's `kill -9`
battery proves that much.
**Measured 2026-08-27: that code does not run.** Under
`vm.overcommit_memory = 0` — the Linux default — `malloc` **succeeds** and the
kernel kills the process when it later *touches* the pages. So the checked path
never gets a NULL to check. It is not dead code in principle, just unreachable in
the configuration everything actually runs in. What a deployment gets instead,
both exits measured under a cgroup cap:
- **swap off: `SIGKILL`, signal 9.** No trap, no message. An external `SIGKILL`
is the one shutdown path that skips every guarantee the WAL was written to
provide — though ack-after-fsync holds: ~40 000 rows came back as a contiguous
intact prefix, no holes, not read as corruption. Iteration 22's `kill -9`
battery proved this for an external kill; iteration 1 proved it for the OOM
killer.
- **swap on: exit 0.** The process finishes, returns success, and serves from
disk. The price depends entirely on access pattern: appending pays **~1%**
(148 s vs 150 s uncapped for 900 000 rows) because cold pages are written once
and never re-read, while random reads across the table pay **273×** (1 851 166
vs 6 771 reads/s; p99 1 µs vs 487 µs). "A RAM-authoritative database on swap is
the worst of both worlds" is therefore true of the **read** path specifically,
not of writes.
So the honest problem statement is not "malloc fails". It is: **there is no
declared budget, no back-pressure as the budget is approached, and no way to

View file

@ -24,6 +24,14 @@ chain: 5
> would optimize a number nobody had measured, against a runtime that
> couldn't use it.
>
> **Supporting evidence for staying last (iteration 1, 2026-08-27):** the write
> path is *not* where memory pressure bites. Inserting 900 000 rows inside a
> 64 MiB cap with swap cost **~1%** (148 s vs 150 s uncapped), because appending
> never re-touches its cold pages. Random *reads* over the same oversized table
> cost **273×**. So the pressure is on the read path, and the io_uring question
> that may actually matter is the one iteration 2 deferred here — io_uring for
> `resident: keys` row reads — not group-commit for writes.
>
> **No spec exists yet.** ~~The forks in *Info* are genuine decisions.~~
>
> **REFINED 2026-08-20: the four forks are SETTLED as their recorded

View file

@ -5,11 +5,12 @@ status: pending
readiness: refine
---
# databasev2 5 — bounded tables and eviction: a declared budget, and back-pressure before the cliff
# databasev2 5 — bounded tables and eviction: a declared budget, and back-pressure at a declared threshold
> Part of [Story — databasev2: the database beyond RAM](00-story.md).
> Needs [2](02-table-storage-modes.md) for the mode a bound attaches to, and
> [1](01-ram-ceiling-measurement.md) for the numbers that set a sane default.
> [1](01-ram-ceiling-measurement.md), whose numbers landed 2026-08-27 and
> **corrected this iteration's framing** — see the third goal.
>
> **The simpler half of the hard problem, done first on purpose.** Evicting from
> a bounded resident table and evicting to disk are the same policy question with
@ -28,10 +29,36 @@ readiness: refine
refusal (trap, let the caller decide), and back-pressure (make the writer
wait). Each is right for a different table, which argues for the policy being
declared rather than chosen for the developer.
- **Back-pressure before the cliff, not at it.** The dangerous exit iteration 1
characterises is swap thrash, which arrives with **no error signal at all**.
A budget that is enforced at 100% has already lost; the value is in acting at
a threshold, while there is still headroom to act.
- **Back-pressure at a declared threshold — because there is no cliff to be
before.** This goal was written expecting a gradient to detect. Iteration 1
measured (2026-08-27) that no such gradient exists, which makes the goal
*stronger*, not weaker:
- Exceeding RAM **without** swap is **SIGKILL, signal 9** — no trap, no
diagnostic. Table storage has no checked ceiling, and under
`vm.overcommit_memory = 0` its `malloc` succeeds and the kernel kills on
page touch, so the checked path never runs.
- Exceeding RAM **with** swap returns **exit 0** and keeps serving from disk.
An append-mostly workload pays **~1%** (148 s vs 150 s uncapped for 900k
rows), so "swap thrash" — which this goal previously named as the dangerous
exit — is not what happens on the write path at all.
- Read latency does not *depart*, it **steps**: 1 µs resident to 487 µs
over-cap with nothing in between.
So there is no early-warning signal anywhere to react to — not an error, not a
latency knee. A budget enforced at 100% has not merely "already lost"; it can
never fire, because the process is dead or silently fine. **Only a declared
threshold can speak, and it must be declared in bytes** — footprint is
96.5–100 B/row Int-only against 320.6–324 B/row text-heavy, **3.3× apart**, so
a row count cannot bound RAM. Leave headroom for index doublings, which are
transient RSS steps (measured at ~24k and ~48k rows): a budget without headroom
fires during a rehash instead of at a real threshold.
- **Eviction policy QUALITY is decisive, not incidental.** Iteration 1 measured
random reads over an oversized table at **273× slower** than resident
(1 851 166 vs 6 771 reads/s; p99 1 µs vs 487 µs). That is the cost of getting
the resident set wrong, so the gap between a good policy and a careless one is
not a few percent — it is the difference between a working system and an
unusable one. Whatever policy ships must be measured against that spread, not
merely shown to be correct.
- **Eviction that respects the engine's actual invariants.** Rows have stable
addresses forever, the free-slot list recycles slots, ids are never reused, and
every secondary index and unique shadow must stay consistent with the slab.

View file

@ -29,7 +29,16 @@ readiness: refine
> **What may still be left:** if measurement after 5c/5d shows the page cache
> insufficient for some workload, a user-space working set becomes arguable
> again — but only with that number in hand, which is the opposite of how this
> file was written. Until then treat the design questions below as answered
> file was written.
>
> **That question now has a reference point (iteration 1, 2026-08-27).** Random
> reads over a table larger than RAM measured **273× slower** than resident
> (1 851 166 vs 6 771 reads/s; p99 1 µs vs 487 µs) — but that is the *kernel's
> swap* path: demand-paged anonymous memory, 4 KiB per fault, no readahead. It is
> the number `resident: keys` must **beat**, since it `pread`s through the page
> cache, which gets readahead and a shared cache. So this file revives if and
> only if 5c/5d measures the page-cache path landing near 273× rather than well
> below it. Until that measurement exists, neither outcome is assumed. Until then treat the design questions below as answered
> elsewhere and the phases as void. Its genuinely durable contribution is its
> fork list, especially "does the language surface the fault cost at the *use*
> site" — still open, and still the largest question about what writeonce is.

View file

@ -249,9 +249,15 @@ commit as the code. An older image is refused on version rather than misread.
takes the same path it takes today. The 1.3M ops/s read baseline is the
regression gate, and a measurable regression there is grounds to reject the
implementation rather than tune it.
- **The failure mode becomes a diagnostic.** Two of the three exits
characterised in databasev2 1 — swap thrash and the OOM killer — are replaced
by a refusal that names the fix.
- **The failure mode becomes a diagnostic.** Both exits **as measured** in
databasev2 1 (2026-08-27) are replaced by a refusal that names the fix — and
the measurement made this argument stronger than the draft that named "swap
thrash and the OOM killer". What actually happens is **SIGKILL signal 9** with
swap off (no trap: overcommit lets `malloc` succeed, the kernel kills on page
touch, so the checked path never runs) or **exit 0 while serving from disk**
with swap on. Neither is a diagnostic; one is silence and the other is a
corpse. A declared budget is the only way this engine can say anything at all
before either.
- **It is declared, not automatic.** No threshold heuristic, no performance
cliff the compiler cannot explain. Consistent with a language whose thesis is
that the compiler tells you the truth.