docs(databasev2): propagate iteration 1's findings to every consumer

Audit found 4 of 10 iterations citing it1 and 4 carrying stale claims the
measurement contradicts.

- 05: framing was contradicted, not merely incomplete. Its goal expected a
  gradient to detect ("back-pressure before the cliff"); there is no cliff
  — SIGKILL with swap off, exit 0 with swap on, and read latency STEPS
  (1us -> 487us) rather than departing. Heading and goal rewritten; the
  measurement makes the goal stronger, not weaker
- 05: budget must be bytes — 3.3x footprint spread — with headroom for
  index doublings, else it fires during a rehash
- 05: new goal — eviction policy QUALITY is decisive, since getting the
  resident set wrong costs 273x, not a few percent
- 06: its revival question now has a reference point. 273x is the KERNEL
  SWAP path; `resident: keys` preads via page cache and must beat it. This
  file revives only if 5c/5d lands near 273x rather than well below
- 04: write path is not where pressure bites (append ~1%, read 273x), so
  the io_uring question that matters is iteration 2's deferred read-path
  one, not group-commit
- 00-story: problem statement asserted the store "refuses the insert
  rather than dying". Corrected in place — a banner above it was not
  enough, a skimmer never reaches it
- residency spec: "swap thrash and the OOM killer" named exits that were
  not measured; replaced with silence-or-a-corpse

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
shoney.arickathil 2026-08-27 22:47:37 +02:00
parent 5b1a8c96a1
commit 3507aafea3
5 changed files with 86 additions and 26 deletions

View file

@ -45,26 +45,36 @@ Worth being precise, because the failure mode determines the fix — and the goo
news is that the engine's own behaviour is clean: news is that the engine's own behaviour is clean:
**Corrected 2026-08-27 by measurement.** This section used to open "an **Corrected 2026-08-27 by measurement.** This section used to open "an
allocation failure is a catchable trap, not a crash", and that is true only of allocation failure is a catchable trap, not a crash". That is true only of the VM
the VM arena. Table storage has no ceiling, and with `vm.overcommit_memory = 0` object arena, whose `WO_HEAP_MB` ceiling is checked and does trap
its `malloc` never fails — the process is **SIGKILLed** (rc=137, measured at (`trap 4 … out of memory`, exit 1). **Table storage has no ceiling at all** —
360 000 rows under a 64 MiB cap). The checked path below is real, but it is the details below, measured. See [iteration 1](01-ram-ceiling-measurement.md).
arena's, not the store's. See [iteration 1](01-ram-ceiling-measurement.md).
Every `malloc` in Every `malloc` in
the row encoder is checked and jumps to an `oom` label; `DB_ERR_OOM` maps to the row encoder is checked and jumps to an `oom` label; `DB_ERR_OOM` maps to
`WO_T_OOM`, which a program can `try`/`catch`. So a writeonce program that runs `WO_T_OOM`, which a program can `try`/`catch`. On paper a writeonce program that
out of memory *refuses the insert* rather than corrupting or dying. That is a runs out of memory *refuses the insert* rather than corrupting or dying.
much better starting position than most engines have.
**But the trap is almost never what a real deployment hits first.** Long before **Measured 2026-08-27: that code does not run.** Under
`malloc` returns NULL, the box starts swapping, and a RAM-authoritative database `vm.overcommit_memory = 0` — the Linux default — `malloc` **succeeds** and the
on swap is the worst of both worlds: it has paid for in-memory data structures kernel kills the process when it later *touches* the pages. So the checked path
and is now serving them from disk with no read path designed for that. On a never gets a NULL to check. It is not dead code in principle, just unreachable in
cgroup-limited host the OOM killer arrives instead, and an external `SIGKILL` is the configuration everything actually runs in. What a deployment gets instead,
the one shutdown path that skips every guarantee the WAL was written to provide — both exits measured under a cgroup cap:
though ack-after-fsync means acked writes still survive; iteration 22's `kill -9`
battery proves that much. - **swap off: `SIGKILL`, signal 9.** No trap, no message. An external `SIGKILL`
is the one shutdown path that skips every guarantee the WAL was written to
provide — though ack-after-fsync holds: ~40 000 rows came back as a contiguous
intact prefix, no holes, not read as corruption. Iteration 22's `kill -9`
battery proved this for an external kill; iteration 1 proved it for the OOM
killer.
- **swap on: exit 0.** The process finishes, returns success, and serves from
disk. The price depends entirely on access pattern: appending pays **~1%**
(148 s vs 150 s uncapped for 900 000 rows) because cold pages are written once
and never re-read, while random reads across the table pay **273×** (1 851 166
vs 6 771 reads/s; p99 1 µs vs 487 µs). "A RAM-authoritative database on swap is
the worst of both worlds" is therefore true of the **read** path specifically,
not of writes.
So the honest problem statement is not "malloc fails". It is: **there is no So the honest problem statement is not "malloc fails". It is: **there is no
declared budget, no back-pressure as the budget is approached, and no way to declared budget, no back-pressure as the budget is approached, and no way to

View file

@ -24,6 +24,14 @@ chain: 5
> would optimize a number nobody had measured, against a runtime that > would optimize a number nobody had measured, against a runtime that
> couldn't use it. > couldn't use it.
> >
> **Supporting evidence for staying last (iteration 1, 2026-08-27):** the write
> path is *not* where memory pressure bites. Inserting 900 000 rows inside a
> 64 MiB cap with swap cost **~1%** (148 s vs 150 s uncapped), because appending
> never re-touches its cold pages. Random *reads* over the same oversized table
> cost **273×**. So the pressure is on the read path, and the io_uring question
> that may actually matter is the one iteration 2 deferred here — io_uring for
> `resident: keys` row reads — not group-commit for writes.
>
> **No spec exists yet.** ~~The forks in *Info* are genuine decisions.~~ > **No spec exists yet.** ~~The forks in *Info* are genuine decisions.~~
> >
> **REFINED 2026-08-20: the four forks are SETTLED as their recorded > **REFINED 2026-08-20: the four forks are SETTLED as their recorded

View file

@ -5,11 +5,12 @@ status: pending
readiness: refine readiness: refine
--- ---
# databasev2 5 — bounded tables and eviction: a declared budget, and back-pressure before the cliff # databasev2 5 — bounded tables and eviction: a declared budget, and back-pressure at a declared threshold
> Part of [Story — databasev2: the database beyond RAM](00-story.md). > Part of [Story — databasev2: the database beyond RAM](00-story.md).
> Needs [2](02-table-storage-modes.md) for the mode a bound attaches to, and > Needs [2](02-table-storage-modes.md) for the mode a bound attaches to, and
> [1](01-ram-ceiling-measurement.md) for the numbers that set a sane default. > [1](01-ram-ceiling-measurement.md), whose numbers landed 2026-08-27 and
> **corrected this iteration's framing** — see the third goal.
> >
> **The simpler half of the hard problem, done first on purpose.** Evicting from > **The simpler half of the hard problem, done first on purpose.** Evicting from
> a bounded resident table and evicting to disk are the same policy question with > a bounded resident table and evicting to disk are the same policy question with
@ -28,10 +29,36 @@ readiness: refine
refusal (trap, let the caller decide), and back-pressure (make the writer refusal (trap, let the caller decide), and back-pressure (make the writer
wait). Each is right for a different table, which argues for the policy being wait). Each is right for a different table, which argues for the policy being
declared rather than chosen for the developer. declared rather than chosen for the developer.
- **Back-pressure before the cliff, not at it.** The dangerous exit iteration 1 - **Back-pressure at a declared threshold — because there is no cliff to be
characterises is swap thrash, which arrives with **no error signal at all**. before.** This goal was written expecting a gradient to detect. Iteration 1
A budget that is enforced at 100% has already lost; the value is in acting at measured (2026-08-27) that no such gradient exists, which makes the goal
a threshold, while there is still headroom to act. *stronger*, not weaker:
- Exceeding RAM **without** swap is **SIGKILL, signal 9** — no trap, no
diagnostic. Table storage has no checked ceiling, and under
`vm.overcommit_memory = 0` its `malloc` succeeds and the kernel kills on
page touch, so the checked path never runs.
- Exceeding RAM **with** swap returns **exit 0** and keeps serving from disk.
An append-mostly workload pays **~1%** (148 s vs 150 s uncapped for 900k
rows), so "swap thrash" — which this goal previously named as the dangerous
exit — is not what happens on the write path at all.
- Read latency does not *depart*, it **steps**: 1 µs resident to 487 µs
over-cap with nothing in between.
So there is no early-warning signal anywhere to react to — not an error, not a
latency knee. A budget enforced at 100% has not merely "already lost"; it can
never fire, because the process is dead or silently fine. **Only a declared
threshold can speak, and it must be declared in bytes** — footprint is
96.5–100 B/row Int-only against 320.6–324 B/row text-heavy, **3.3× apart**, so
a row count cannot bound RAM. Leave headroom for index doublings, which are
transient RSS steps (measured at ~24k and ~48k rows): a budget without headroom
fires during a rehash instead of at a real threshold.
- **Eviction policy QUALITY is decisive, not incidental.** Iteration 1 measured
random reads over an oversized table at **273× slower** than resident
(1 851 166 vs 6 771 reads/s; p99 1 µs vs 487 µs). That is the cost of getting
the resident set wrong, so the gap between a good policy and a careless one is
not a few percent — it is the difference between a working system and an
unusable one. Whatever policy ships must be measured against that spread, not
merely shown to be correct.
- **Eviction that respects the engine's actual invariants.** Rows have stable - **Eviction that respects the engine's actual invariants.** Rows have stable
addresses forever, the free-slot list recycles slots, ids are never reused, and addresses forever, the free-slot list recycles slots, ids are never reused, and
every secondary index and unique shadow must stay consistent with the slab. every secondary index and unique shadow must stay consistent with the slab.

View file

@ -29,7 +29,16 @@ readiness: refine
> **What may still be left:** if measurement after 5c/5d shows the page cache > **What may still be left:** if measurement after 5c/5d shows the page cache
> insufficient for some workload, a user-space working set becomes arguable > insufficient for some workload, a user-space working set becomes arguable
> again — but only with that number in hand, which is the opposite of how this > again — but only with that number in hand, which is the opposite of how this
> file was written. Until then treat the design questions below as answered > file was written.
>
> **That question now has a reference point (iteration 1, 2026-08-27).** Random
> reads over a table larger than RAM measured **273× slower** than resident
> (1 851 166 vs 6 771 reads/s; p99 1 µs vs 487 µs) — but that is the *kernel's
> swap* path: demand-paged anonymous memory, 4 KiB per fault, no readahead. It is
> the number `resident: keys` must **beat**, since it `pread`s through the page
> cache, which gets readahead and a shared cache. So this file revives if and
> only if 5c/5d measures the page-cache path landing near 273× rather than well
> below it. Until that measurement exists, neither outcome is assumed. Until then treat the design questions below as answered
> elsewhere and the phases as void. Its genuinely durable contribution is its > elsewhere and the phases as void. Its genuinely durable contribution is its
> fork list, especially "does the language surface the fault cost at the *use* > fork list, especially "does the language surface the fault cost at the *use*
> site" — still open, and still the largest question about what writeonce is. > site" — still open, and still the largest question about what writeonce is.

View file

@ -249,9 +249,15 @@ commit as the code. An older image is refused on version rather than misread.
takes the same path it takes today. The 1.3M ops/s read baseline is the takes the same path it takes today. The 1.3M ops/s read baseline is the
regression gate, and a measurable regression there is grounds to reject the regression gate, and a measurable regression there is grounds to reject the
implementation rather than tune it. implementation rather than tune it.
- **The failure mode becomes a diagnostic.** Two of the three exits - **The failure mode becomes a diagnostic.** Both exits **as measured** in
characterised in databasev2 1 — swap thrash and the OOM killer — are replaced databasev2 1 (2026-08-27) are replaced by a refusal that names the fix — and
by a refusal that names the fix. the measurement made this argument stronger than the draft that named "swap
thrash and the OOM killer". What actually happens is **SIGKILL signal 9** with
swap off (no trap: overcommit lets `malloc` succeed, the kernel kills on page
touch, so the checked path never runs) or **exit 0 while serving from disk**
with swap on. Neither is a diagnostic; one is silence and the other is a
corpse. A declared budget is the only way this engine can say anything at all
before either.
- **It is declared, not automatic.** No threshold heuristic, no performance - **It is declared, not automatic.** No threshold heuristic, no performance
cliff the compiler cannot explain. Consistent with a language whose thesis is cliff the compiler cannot explain. Consistent with a language whose thesis is
that the compiler tells you the truth. that the compiler tells you the truth.