diff --git a/docs/stories/databasev2/00-story.md b/docs/stories/databasev2/00-story.md index 278eb17..bae4cb5 100644 --- a/docs/stories/databasev2/00-story.md +++ b/docs/stories/databasev2/00-story.md @@ -45,26 +45,36 @@ Worth being precise, because the failure mode determines the fix — and the goo news is that the engine's own behaviour is clean: **Corrected 2026-08-27 by measurement.** This section used to open "an -allocation failure is a catchable trap, not a crash", and that is true only of -the VM arena. Table storage has no ceiling, and with `vm.overcommit_memory = 0` -its `malloc` never fails — the process is **SIGKILLed** (rc=137, measured at -360 000 rows under a 64 MiB cap). The checked path below is real, but it is the -arena's, not the store's. See [iteration 1](01-ram-ceiling-measurement.md). +allocation failure is a catchable trap, not a crash". That is true only of the VM +object arena, whose `WO_HEAP_MB` ceiling is checked and does trap +(`trap 4 … out of memory`, exit 1). **Table storage has no ceiling at all** — +details below, measured. See [iteration 1](01-ram-ceiling-measurement.md). Every `malloc` in the row encoder is checked and jumps to an `oom` label; `DB_ERR_OOM` maps to -`WO_T_OOM`, which a program can `try`/`catch`. So a writeonce program that runs -out of memory *refuses the insert* rather than corrupting or dying. That is a -much better starting position than most engines have. +`WO_T_OOM`, which a program can `try`/`catch`. On paper a writeonce program that +runs out of memory *refuses the insert* rather than corrupting or dying. -**But the trap is almost never what a real deployment hits first.** Long before -`malloc` returns NULL, the box starts swapping, and a RAM-authoritative database -on swap is the worst of both worlds: it has paid for in-memory data structures -and is now serving them from disk with no read path designed for that. On a -cgroup-limited host the OOM killer arrives instead, and an external `SIGKILL` is -the one shutdown path that skips every guarantee the WAL was written to provide — -though ack-after-fsync means acked writes still survive; iteration 22's `kill -9` -battery proves that much. +**Measured 2026-08-27: that code does not run.** Under +`vm.overcommit_memory = 0` — the Linux default — `malloc` **succeeds** and the +kernel kills the process when it later *touches* the pages. So the checked path +never gets a NULL to check. It is not dead code in principle, just unreachable in +the configuration everything actually runs in. What a deployment gets instead, +both exits measured under a cgroup cap: + +- **swap off: `SIGKILL`, signal 9.** No trap, no message. An external `SIGKILL` + is the one shutdown path that skips every guarantee the WAL was written to + provide — though ack-after-fsync holds: ~40 000 rows came back as a contiguous + intact prefix, no holes, not read as corruption. Iteration 22's `kill -9` + battery proved this for an external kill; iteration 1 proved it for the OOM + killer. +- **swap on: exit 0.** The process finishes, returns success, and serves from + disk. The price depends entirely on access pattern: appending pays **~1%** + (148 s vs 150 s uncapped for 900 000 rows) because cold pages are written once + and never re-read, while random reads across the table pay **273×** (1 851 166 + vs 6 771 reads/s; p99 1 µs vs 487 µs). "A RAM-authoritative database on swap is + the worst of both worlds" is therefore true of the **read** path specifically, + not of writes. So the honest problem statement is not "malloc fails". It is: **there is no declared budget, no back-pressure as the budget is approached, and no way to diff --git a/docs/stories/databasev2/04-io-uring-commit.md b/docs/stories/databasev2/04-io-uring-commit.md index a0fe26f..dd5c255 100644 --- a/docs/stories/databasev2/04-io-uring-commit.md +++ b/docs/stories/databasev2/04-io-uring-commit.md @@ -24,6 +24,14 @@ chain: 5 > would optimize a number nobody had measured, against a runtime that > couldn't use it. > +> **Supporting evidence for staying last (iteration 1, 2026-08-27):** the write +> path is *not* where memory pressure bites. Inserting 900 000 rows inside a +> 64 MiB cap with swap cost **~1%** (148 s vs 150 s uncapped), because appending +> never re-touches its cold pages. Random *reads* over the same oversized table +> cost **273×**. So the pressure is on the read path, and the io_uring question +> that may actually matter is the one iteration 2 deferred here — io_uring for +> `resident: keys` row reads — not group-commit for writes. +> > **No spec exists yet.** ~~The forks in *Info* are genuine decisions.~~ > > **REFINED 2026-08-20: the four forks are SETTLED as their recorded diff --git a/docs/stories/databasev2/05-bounded-tables-eviction.md b/docs/stories/databasev2/05-bounded-tables-eviction.md index 22b6c7d..c4cdaac 100644 --- a/docs/stories/databasev2/05-bounded-tables-eviction.md +++ b/docs/stories/databasev2/05-bounded-tables-eviction.md @@ -5,11 +5,12 @@ status: pending readiness: refine --- -# databasev2 5 — bounded tables and eviction: a declared budget, and back-pressure before the cliff +# databasev2 5 — bounded tables and eviction: a declared budget, and back-pressure at a declared threshold > Part of [Story — databasev2: the database beyond RAM](00-story.md). > Needs [2](02-table-storage-modes.md) for the mode a bound attaches to, and -> [1](01-ram-ceiling-measurement.md) for the numbers that set a sane default. +> [1](01-ram-ceiling-measurement.md), whose numbers landed 2026-08-27 and +> **corrected this iteration's framing** — see the third goal. > > **The simpler half of the hard problem, done first on purpose.** Evicting from > a bounded resident table and evicting to disk are the same policy question with @@ -28,10 +29,36 @@ readiness: refine refusal (trap, let the caller decide), and back-pressure (make the writer wait). Each is right for a different table, which argues for the policy being declared rather than chosen for the developer. -- **Back-pressure before the cliff, not at it.** The dangerous exit iteration 1 - characterises is swap thrash, which arrives with **no error signal at all**. - A budget that is enforced at 100% has already lost; the value is in acting at - a threshold, while there is still headroom to act. +- **Back-pressure at a declared threshold — because there is no cliff to be + before.** This goal was written expecting a gradient to detect. Iteration 1 + measured (2026-08-27) that no such gradient exists, which makes the goal + *stronger*, not weaker: + - Exceeding RAM **without** swap is **SIGKILL, signal 9** — no trap, no + diagnostic. Table storage has no checked ceiling, and under + `vm.overcommit_memory = 0` its `malloc` succeeds and the kernel kills on + page touch, so the checked path never runs. + - Exceeding RAM **with** swap returns **exit 0** and keeps serving from disk. + An append-mostly workload pays **~1%** (148 s vs 150 s uncapped for 900k + rows), so "swap thrash" — which this goal previously named as the dangerous + exit — is not what happens on the write path at all. + - Read latency does not *depart*, it **steps**: 1 µs resident to 487 µs + over-cap with nothing in between. + + So there is no early-warning signal anywhere to react to — not an error, not a + latency knee. A budget enforced at 100% has not merely "already lost"; it can + never fire, because the process is dead or silently fine. **Only a declared + threshold can speak, and it must be declared in bytes** — footprint is + 96.5–100 B/row Int-only against 320.6–324 B/row text-heavy, **3.3× apart**, so + a row count cannot bound RAM. Leave headroom for index doublings, which are + transient RSS steps (measured at ~24k and ~48k rows): a budget without headroom + fires during a rehash instead of at a real threshold. +- **Eviction policy QUALITY is decisive, not incidental.** Iteration 1 measured + random reads over an oversized table at **273× slower** than resident + (1 851 166 vs 6 771 reads/s; p99 1 µs vs 487 µs). That is the cost of getting + the resident set wrong, so the gap between a good policy and a careless one is + not a few percent — it is the difference between a working system and an + unusable one. Whatever policy ships must be measured against that spread, not + merely shown to be correct. - **Eviction that respects the engine's actual invariants.** Rows have stable addresses forever, the free-slot list recycles slots, ids are never reused, and every secondary index and unique shadow must stay consistent with the slab. diff --git a/docs/stories/databasev2/06-cold-tiering.md b/docs/stories/databasev2/06-cold-tiering.md index 3b21194..e45e7f0 100644 --- a/docs/stories/databasev2/06-cold-tiering.md +++ b/docs/stories/databasev2/06-cold-tiering.md @@ -29,7 +29,16 @@ readiness: refine > **What may still be left:** if measurement after 5c/5d shows the page cache > insufficient for some workload, a user-space working set becomes arguable > again — but only with that number in hand, which is the opposite of how this -> file was written. Until then treat the design questions below as answered +> file was written. +> +> **That question now has a reference point (iteration 1, 2026-08-27).** Random +> reads over a table larger than RAM measured **273× slower** than resident +> (1 851 166 vs 6 771 reads/s; p99 1 µs vs 487 µs) — but that is the *kernel's +> swap* path: demand-paged anonymous memory, 4 KiB per fault, no readahead. It is +> the number `resident: keys` must **beat**, since it `pread`s through the page +> cache, which gets readahead and a shared cache. So this file revives if and +> only if 5c/5d measures the page-cache path landing near 273× rather than well +> below it. Until that measurement exists, neither outcome is assumed. Until then treat the design questions below as answered > elsewhere and the phases as void. Its genuinely durable contribution is its > fork list, especially "does the language surface the fault cost at the *use* > site" — still open, and still the largest question about what writeonce is. diff --git a/docs/superpowers/specs/2026-08-26-table-residency-design.md b/docs/superpowers/specs/2026-08-26-table-residency-design.md index efc6d65..98b16e5 100644 --- a/docs/superpowers/specs/2026-08-26-table-residency-design.md +++ b/docs/superpowers/specs/2026-08-26-table-residency-design.md @@ -249,9 +249,15 @@ commit as the code. An older image is refused on version rather than misread. takes the same path it takes today. The 1.3M ops/s read baseline is the regression gate, and a measurable regression there is grounds to reject the implementation rather than tune it. -- **The failure mode becomes a diagnostic.** Two of the three exits - characterised in databasev2 1 — swap thrash and the OOM killer — are replaced - by a refusal that names the fix. +- **The failure mode becomes a diagnostic.** Both exits **as measured** in + databasev2 1 (2026-08-27) are replaced by a refusal that names the fix — and + the measurement made this argument stronger than the draft that named "swap + thrash and the OOM killer". What actually happens is **SIGKILL signal 9** with + swap off (no trap: overcommit lets `malloc` succeed, the kernel kills on page + touch, so the checked path never runs) or **exit 0 while serving from disk** + with swap on. Neither is a diagnostic; one is silence and the other is a + corpse. A declared budget is the only way this engine can say anything at all + before either. - **It is declared, not automatic.** No threshold heuristic, no performance cliff the compiler cannot explain. Consistent with a language whose thesis is that the compiler tells you the truth.