docs(databasev2): propagate iteration 1's findings to every consumer
Audit found 4 of 10 iterations citing it1 and 4 carrying stale claims the
measurement contradicts.
- 05: framing was contradicted, not merely incomplete. Its goal expected a
gradient to detect ("back-pressure before the cliff"); there is no cliff
— SIGKILL with swap off, exit 0 with swap on, and read latency STEPS
(1us -> 487us) rather than departing. Heading and goal rewritten; the
measurement makes the goal stronger, not weaker
- 05: budget must be bytes — 3.3x footprint spread — with headroom for
index doublings, else it fires during a rehash
- 05: new goal — eviction policy QUALITY is decisive, since getting the
resident set wrong costs 273x, not a few percent
- 06: its revival question now has a reference point. 273x is the KERNEL
SWAP path; `resident: keys` preads via page cache and must beat it. This
file revives only if 5c/5d lands near 273x rather than well below
- 04: write path is not where pressure bites (append ~1%, read 273x), so
the io_uring question that matters is iteration 2's deferred read-path
one, not group-commit
- 00-story: problem statement asserted the store "refuses the insert
rather than dying". Corrected in place — a banner above it was not
enough, a skimmer never reaches it
- residency spec: "swap thrash and the OOM killer" named exits that were
not measured; replaced with silence-or-a-corpse
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
parent
5b1a8c96a1
commit
3507aafea3
5 changed files with 86 additions and 26 deletions
|
|
@ -45,26 +45,36 @@ Worth being precise, because the failure mode determines the fix — and the goo
|
||||||
news is that the engine's own behaviour is clean:
|
news is that the engine's own behaviour is clean:
|
||||||
|
|
||||||
**Corrected 2026-08-27 by measurement.** This section used to open "an
|
**Corrected 2026-08-27 by measurement.** This section used to open "an
|
||||||
allocation failure is a catchable trap, not a crash", and that is true only of
|
allocation failure is a catchable trap, not a crash". That is true only of the VM
|
||||||
the VM arena. Table storage has no ceiling, and with `vm.overcommit_memory = 0`
|
object arena, whose `WO_HEAP_MB` ceiling is checked and does trap
|
||||||
its `malloc` never fails — the process is **SIGKILLed** (rc=137, measured at
|
(`trap 4 … out of memory`, exit 1). **Table storage has no ceiling at all** —
|
||||||
360 000 rows under a 64 MiB cap). The checked path below is real, but it is the
|
details below, measured. See [iteration 1](01-ram-ceiling-measurement.md).
|
||||||
arena's, not the store's. See [iteration 1](01-ram-ceiling-measurement.md).
|
|
||||||
|
|
||||||
Every `malloc` in
|
Every `malloc` in
|
||||||
the row encoder is checked and jumps to an `oom` label; `DB_ERR_OOM` maps to
|
the row encoder is checked and jumps to an `oom` label; `DB_ERR_OOM` maps to
|
||||||
`WO_T_OOM`, which a program can `try`/`catch`. So a writeonce program that runs
|
`WO_T_OOM`, which a program can `try`/`catch`. On paper a writeonce program that
|
||||||
out of memory *refuses the insert* rather than corrupting or dying. That is a
|
runs out of memory *refuses the insert* rather than corrupting or dying.
|
||||||
much better starting position than most engines have.
|
|
||||||
|
|
||||||
**But the trap is almost never what a real deployment hits first.** Long before
|
**Measured 2026-08-27: that code does not run.** Under
|
||||||
`malloc` returns NULL, the box starts swapping, and a RAM-authoritative database
|
`vm.overcommit_memory = 0` — the Linux default — `malloc` **succeeds** and the
|
||||||
on swap is the worst of both worlds: it has paid for in-memory data structures
|
kernel kills the process when it later *touches* the pages. So the checked path
|
||||||
and is now serving them from disk with no read path designed for that. On a
|
never gets a NULL to check. It is not dead code in principle, just unreachable in
|
||||||
cgroup-limited host the OOM killer arrives instead, and an external `SIGKILL` is
|
the configuration everything actually runs in. What a deployment gets instead,
|
||||||
the one shutdown path that skips every guarantee the WAL was written to provide —
|
both exits measured under a cgroup cap:
|
||||||
though ack-after-fsync means acked writes still survive; iteration 22's `kill -9`
|
|
||||||
battery proves that much.
|
- **swap off: `SIGKILL`, signal 9.** No trap, no message. An external `SIGKILL`
|
||||||
|
is the one shutdown path that skips every guarantee the WAL was written to
|
||||||
|
provide — though ack-after-fsync holds: ~40 000 rows came back as a contiguous
|
||||||
|
intact prefix, no holes, not read as corruption. Iteration 22's `kill -9`
|
||||||
|
battery proved this for an external kill; iteration 1 proved it for the OOM
|
||||||
|
killer.
|
||||||
|
- **swap on: exit 0.** The process finishes, returns success, and serves from
|
||||||
|
disk. The price depends entirely on access pattern: appending pays **~1%**
|
||||||
|
(148 s vs 150 s uncapped for 900 000 rows) because cold pages are written once
|
||||||
|
and never re-read, while random reads across the table pay **273×** (1 851 166
|
||||||
|
vs 6 771 reads/s; p99 1 µs vs 487 µs). "A RAM-authoritative database on swap is
|
||||||
|
the worst of both worlds" is therefore true of the **read** path specifically,
|
||||||
|
not of writes.
|
||||||
|
|
||||||
So the honest problem statement is not "malloc fails". It is: **there is no
|
So the honest problem statement is not "malloc fails". It is: **there is no
|
||||||
declared budget, no back-pressure as the budget is approached, and no way to
|
declared budget, no back-pressure as the budget is approached, and no way to
|
||||||
|
|
|
||||||
|
|
@ -24,6 +24,14 @@ chain: 5
|
||||||
> would optimize a number nobody had measured, against a runtime that
|
> would optimize a number nobody had measured, against a runtime that
|
||||||
> couldn't use it.
|
> couldn't use it.
|
||||||
>
|
>
|
||||||
|
> **Supporting evidence for staying last (iteration 1, 2026-08-27):** the write
|
||||||
|
> path is *not* where memory pressure bites. Inserting 900 000 rows inside a
|
||||||
|
> 64 MiB cap with swap cost **~1%** (148 s vs 150 s uncapped), because appending
|
||||||
|
> never re-touches its cold pages. Random *reads* over the same oversized table
|
||||||
|
> cost **273×**. So the pressure is on the read path, and the io_uring question
|
||||||
|
> that may actually matter is the one iteration 2 deferred here — io_uring for
|
||||||
|
> `resident: keys` row reads — not group-commit for writes.
|
||||||
|
>
|
||||||
> **No spec exists yet.** ~~The forks in *Info* are genuine decisions.~~
|
> **No spec exists yet.** ~~The forks in *Info* are genuine decisions.~~
|
||||||
>
|
>
|
||||||
> **REFINED 2026-08-20: the four forks are SETTLED as their recorded
|
> **REFINED 2026-08-20: the four forks are SETTLED as their recorded
|
||||||
|
|
|
||||||
|
|
@ -5,11 +5,12 @@ status: pending
|
||||||
readiness: refine
|
readiness: refine
|
||||||
---
|
---
|
||||||
|
|
||||||
# databasev2 5 — bounded tables and eviction: a declared budget, and back-pressure before the cliff
|
# databasev2 5 — bounded tables and eviction: a declared budget, and back-pressure at a declared threshold
|
||||||
|
|
||||||
> Part of [Story — databasev2: the database beyond RAM](00-story.md).
|
> Part of [Story — databasev2: the database beyond RAM](00-story.md).
|
||||||
> Needs [2](02-table-storage-modes.md) for the mode a bound attaches to, and
|
> Needs [2](02-table-storage-modes.md) for the mode a bound attaches to, and
|
||||||
> [1](01-ram-ceiling-measurement.md) for the numbers that set a sane default.
|
> [1](01-ram-ceiling-measurement.md), whose numbers landed 2026-08-27 and
|
||||||
|
> **corrected this iteration's framing** — see the third goal.
|
||||||
>
|
>
|
||||||
> **The simpler half of the hard problem, done first on purpose.** Evicting from
|
> **The simpler half of the hard problem, done first on purpose.** Evicting from
|
||||||
> a bounded resident table and evicting to disk are the same policy question with
|
> a bounded resident table and evicting to disk are the same policy question with
|
||||||
|
|
@ -28,10 +29,36 @@ readiness: refine
|
||||||
refusal (trap, let the caller decide), and back-pressure (make the writer
|
refusal (trap, let the caller decide), and back-pressure (make the writer
|
||||||
wait). Each is right for a different table, which argues for the policy being
|
wait). Each is right for a different table, which argues for the policy being
|
||||||
declared rather than chosen for the developer.
|
declared rather than chosen for the developer.
|
||||||
- **Back-pressure before the cliff, not at it.** The dangerous exit iteration 1
|
- **Back-pressure at a declared threshold — because there is no cliff to be
|
||||||
characterises is swap thrash, which arrives with **no error signal at all**.
|
before.** This goal was written expecting a gradient to detect. Iteration 1
|
||||||
A budget that is enforced at 100% has already lost; the value is in acting at
|
measured (2026-08-27) that no such gradient exists, which makes the goal
|
||||||
a threshold, while there is still headroom to act.
|
*stronger*, not weaker:
|
||||||
|
- Exceeding RAM **without** swap is **SIGKILL, signal 9** — no trap, no
|
||||||
|
diagnostic. Table storage has no checked ceiling, and under
|
||||||
|
`vm.overcommit_memory = 0` its `malloc` succeeds and the kernel kills on
|
||||||
|
page touch, so the checked path never runs.
|
||||||
|
- Exceeding RAM **with** swap returns **exit 0** and keeps serving from disk.
|
||||||
|
An append-mostly workload pays **~1%** (148 s vs 150 s uncapped for 900k
|
||||||
|
rows), so "swap thrash" — which this goal previously named as the dangerous
|
||||||
|
exit — is not what happens on the write path at all.
|
||||||
|
- Read latency does not *depart*, it **steps**: 1 µs resident to 487 µs
|
||||||
|
over-cap with nothing in between.
|
||||||
|
|
||||||
|
So there is no early-warning signal anywhere to react to — not an error, not a
|
||||||
|
latency knee. A budget enforced at 100% has not merely "already lost"; it can
|
||||||
|
never fire, because the process is dead or silently fine. **Only a declared
|
||||||
|
threshold can speak, and it must be declared in bytes** — footprint is
|
||||||
|
96.5–100 B/row Int-only against 320.6–324 B/row text-heavy, **3.3× apart**, so
|
||||||
|
a row count cannot bound RAM. Leave headroom for index doublings, which are
|
||||||
|
transient RSS steps (measured at ~24k and ~48k rows): a budget without headroom
|
||||||
|
fires during a rehash instead of at a real threshold.
|
||||||
|
- **Eviction policy QUALITY is decisive, not incidental.** Iteration 1 measured
|
||||||
|
random reads over an oversized table at **273× slower** than resident
|
||||||
|
(1 851 166 vs 6 771 reads/s; p99 1 µs vs 487 µs). That is the cost of getting
|
||||||
|
the resident set wrong, so the gap between a good policy and a careless one is
|
||||||
|
not a few percent — it is the difference between a working system and an
|
||||||
|
unusable one. Whatever policy ships must be measured against that spread, not
|
||||||
|
merely shown to be correct.
|
||||||
- **Eviction that respects the engine's actual invariants.** Rows have stable
|
- **Eviction that respects the engine's actual invariants.** Rows have stable
|
||||||
addresses forever, the free-slot list recycles slots, ids are never reused, and
|
addresses forever, the free-slot list recycles slots, ids are never reused, and
|
||||||
every secondary index and unique shadow must stay consistent with the slab.
|
every secondary index and unique shadow must stay consistent with the slab.
|
||||||
|
|
|
||||||
|
|
@ -29,7 +29,16 @@ readiness: refine
|
||||||
> **What may still be left:** if measurement after 5c/5d shows the page cache
|
> **What may still be left:** if measurement after 5c/5d shows the page cache
|
||||||
> insufficient for some workload, a user-space working set becomes arguable
|
> insufficient for some workload, a user-space working set becomes arguable
|
||||||
> again — but only with that number in hand, which is the opposite of how this
|
> again — but only with that number in hand, which is the opposite of how this
|
||||||
> file was written. Until then treat the design questions below as answered
|
> file was written.
|
||||||
|
>
|
||||||
|
> **That question now has a reference point (iteration 1, 2026-08-27).** Random
|
||||||
|
> reads over a table larger than RAM measured **273× slower** than resident
|
||||||
|
> (1 851 166 vs 6 771 reads/s; p99 1 µs vs 487 µs) — but that is the *kernel's
|
||||||
|
> swap* path: demand-paged anonymous memory, 4 KiB per fault, no readahead. It is
|
||||||
|
> the number `resident: keys` must **beat**, since it `pread`s through the page
|
||||||
|
> cache, which gets readahead and a shared cache. So this file revives if and
|
||||||
|
> only if 5c/5d measures the page-cache path landing near 273× rather than well
|
||||||
|
> below it. Until that measurement exists, neither outcome is assumed. Until then treat the design questions below as answered
|
||||||
> elsewhere and the phases as void. Its genuinely durable contribution is its
|
> elsewhere and the phases as void. Its genuinely durable contribution is its
|
||||||
> fork list, especially "does the language surface the fault cost at the *use*
|
> fork list, especially "does the language surface the fault cost at the *use*
|
||||||
> site" — still open, and still the largest question about what writeonce is.
|
> site" — still open, and still the largest question about what writeonce is.
|
||||||
|
|
|
||||||
|
|
@ -249,9 +249,15 @@ commit as the code. An older image is refused on version rather than misread.
|
||||||
takes the same path it takes today. The 1.3M ops/s read baseline is the
|
takes the same path it takes today. The 1.3M ops/s read baseline is the
|
||||||
regression gate, and a measurable regression there is grounds to reject the
|
regression gate, and a measurable regression there is grounds to reject the
|
||||||
implementation rather than tune it.
|
implementation rather than tune it.
|
||||||
- **The failure mode becomes a diagnostic.** Two of the three exits
|
- **The failure mode becomes a diagnostic.** Both exits **as measured** in
|
||||||
characterised in databasev2 1 — swap thrash and the OOM killer — are replaced
|
databasev2 1 (2026-08-27) are replaced by a refusal that names the fix — and
|
||||||
by a refusal that names the fix.
|
the measurement made this argument stronger than the draft that named "swap
|
||||||
|
thrash and the OOM killer". What actually happens is **SIGKILL signal 9** with
|
||||||
|
swap off (no trap: overcommit lets `malloc` succeed, the kernel kills on page
|
||||||
|
touch, so the checked path never runs) or **exit 0 while serving from disk**
|
||||||
|
with swap on. Neither is a diagnostic; one is silence and the other is a
|
||||||
|
corpse. A declared budget is the only way this engine can say anything at all
|
||||||
|
before either.
|
||||||
- **It is declared, not automatic.** No threshold heuristic, no performance
|
- **It is declared, not automatic.** No threshold heuristic, no performance
|
||||||
cliff the compiler cannot explain. Consistent with a language whose thesis is
|
cliff the compiler cannot explain. Consistent with a language whose thesis is
|
||||||
that the compiler tells you the truth.
|
that the compiler tells you the truth.
|
||||||
|
|
|
||||||
Loading…
Reference in a new issue