Closes the last gap in databasev2 1; gives databasev2 3 its "before". - `boot` mode: does NOTHING. WO_DATA replay runs before main, so a mode with no work measures replay plus a fixed startup - `replayseed N M`: N inserts + M updates — same live rows, longer log - `replay` leg: empty-store startup floor measured and SUBTRACTED, then two shapes timed, median of 3 boots each - premise check: updates must actually append WAL records, else the two shapes are one measurement and the penalty means nothing - WAL bytes = non-zero prefix, never file size (fallocate'd to 1 MiB) - per-record cost stored in NANOseconds: as us it rounded 5.5 and 5.3 to 6 and 5, too coarse for the number a checkpoint exists to improve - 148 checks, 0 failures; gate bites on a doctored ns_per_record Measured — same 20 000 live rows, different history: - 20 000 records: 980 035 B WAL, 110 ms replay, 5.5 us/record - 40 000 records: 1 960 035 B WAL, 211 ms replay, 5.3 us/record - 1.9x boot cost for an IDENTICAL dataset; per-record cost flat, so replay is linear in records not rows - extrapolated: 10M records ~55 s of boot, 100M ~9 min - databasev2 3 correction: it planned to use "22's aged-store replay numbers", which never existed — 22 proved restart correctness, never timed it - databasev2 3 hazard recorded: compaction rewrites the log and moves every record, so it invalidates every `resident: keys` offset — an arbitrary byte in a rewritten file, not stale-but-readable - databasev2 1 -> status: done Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
11 KiB
Performance targets — measured, named, waiting
A register like discarded.md/learnings.md:
optimization candidates that exist because a NUMBER says so, not a
hunch. Every row cites its measurement (the db-bench campaign,
bench/baseline.json, or the go-sqlite comparison)
and names an owner iteration when one exists. A target leaves this file
by landing (delta recorded in the baseline) or by being rejected into
discarded.md with its reason.
1. The write path — update-through-query re-probes, insert re-encodes
Measured 2026-08-22 (go-sqlite comparison, N=20k, same machine, ext4): ram mixed writes 195,465 ops/s vs SQLite's 380,069 (×1.9 behind); durable mixed writes 2,324 vs 3,257 (×1.4 behind) — while writeonce WINS durable seed ×1.4 and reads ×2.6–6.4. The write gap is specifically the UPDATE half of the mix.
Suspected costs, in probable order (attribute before optimizing — the C-API microbench the 22 spec reserves exists for exactly this):
- Update-through-query runs a whole query statement per update:
probe (now O(1)) + materialize an id
multi(arena alloc) +DB_GET_FIELD/DB_UPDATE_FIELDbuiltin round-trips per touched field. SQLite's equivalent is one page write inside one statement. wo_row_update_fieldwalks every index three times (shadow unique check, old-entry removal, new-entry add — threetouches-loops overt->indexesper update; seedatabase/src/table.c).- Insert encodes per field with a malloc per text/owned value
(
db_val_encode) — visible as ram seed ×1.2 behind SQLite (245k vs 297k) even though the durable flavor wins. - A WAL update record re-encodes the whole row
(
wo_wal_append_updatewrites the row image, not a delta).
Owner: none yet. Sequence note: iteration 23 (io_uring group-commit) rewrites the durable write path's syscall story anyway — re-measure after 23 lands, then decide whether the RAM-side costs (1–3) earn their own slice. Acceptance shape: ram write ops/s closes on SQLite's number with reads unharmed; baseline refreshed with the delta recorded.
2. Cross-shard DB RPC halves concurrent read throughput
Measured 2026-08-22 (db-bench campaign): ram mixread 89,538 ops/s single-shard vs 44,918 at default cores; durable 9,211 vs 4,324. The DB actor serializes every statement on shard 0 and each op pays an envelope + park/unpark round-trip.
Owner: by design, priced deliberately (story 8's settled decision 1 — rejected alternatives: engine lock, partitioned tables "wait for a measured need"). THIS is the measured need's first data point; the recorded escalation path is partitioned/replicated read state, only if a real workload (iteration 24's chat) hurts. Not actionable before 24.
3. The mutex inbox costs ~6× on cross-shard message rate
Measured 2026-08-22: 16.7M msgs/s same-heap vs 2.85M cross-shard
(msgrate). Owner: iteration 31 (mailbox/backpressure decisions
consume this number) and stage-2 deviation 4 (lock-free rings arrive
only if the mutex is the measured bottleneck — at 2.85M msgs/s it is
not the limiting factor for any current workload).
4. Durable write throughput is fsync-bound at ~4.5k/s
Measured 2026-08-21: durable seed 4,460 inserts/s vs ram 245k — the ~55× gap is one fdatasync per statement (~220µs each). Owner: iteration 23 (io_uring group-commit) — its acceptance is literally this number moving while the crash battery stays green.
5. The RAM ceiling: footprint, and how the engine actually dies
Measured 2026-08-27 (databasev2 1), rootless cgroup v2 via
systemd-run --user --scope -p MemoryMax=N -p MemorySwapMax=M, dev box.
Per-row resident footprint, by shape
| Shape | Columns | Steady-state | Doubling steps |
|---|---|---|---|
Int-only (Item) |
2× Int + 1 ref | 96.5 B/row | at ~24k and ~48k rows |
Text-heavy (Wide) |
1× Int + 3× Text | 320.6 B/row | at ~24k and ~48k rows |
3.3×, not the "order of magnitude" an earlier doc asserted. Two shapes are
published, never one number: a Text column is a separate db_text allocation
per row on top of the slab slot, so a row count cannot bound RAM.
Read the steady-state figure as the median of per-interval marginals, not a two-point slope. The id hash and index buckets are open-addressing pow2 and double periodically; a two-point slope lands arbitrarily on or off a doubling and swings 2× (96 vs 205 B/row measured for the same shape). The doublings are reported separately because a transient RSS step is exactly what a resident-footprint budget must leave headroom for — a budget without it fires during a rehash rather than at a steady-state threshold. Direct input to databasev2 2's budget design.
How it dies — and it is not the way the docs claimed
| Allocator | Ceiling | Failure mode |
|---|---|---|
| VM object arena | WO_HEAP_MB, checked |
trap 4 … out of memory, rc=1, reportable. Verified at 4 and 16 MiB |
| table storage (slabs + heap values) | none | SIGKILL, signal 9 (shell rc 137). Verified at 360 000 rows / 57 188 KiB under a 64 MiB cap |
Three docs asserted that an allocation failure surfaces as a catchable
WO_T_OOM. For table storage it does not: vm.overcommit_memory = 0 means
malloc succeeds and the kernel kills the process when it touches the pages,
so the checked-malloc code never runs. The trap path is real, but it is the
arena's.
Consequence, and the strongest available argument for databasev2 2's byte
budget: a declared budget is the only way table storage can acquire a
checked ceiling, because malloc under default overcommit will never report a
problem. Owner: databasev2 2.
Swap: the ceiling that does not announce itself
| Leg | 900 000 Int rows, 64 MiB cap | Wall | Final RSS |
|---|---|---|---|
swap OFF (MemorySwapMax=0) |
SIGKILL at 360 000 rows | — | 57 188 KiB |
| swap ON (256 MiB) | completed, exit 0 | 148 s | 62 264 KiB (rest paged out) |
| uncapped | completed, exit 0 | 150 s | 169 416 KiB |
Swap cost ~1%. A prior draft predicted "latency collapse"; the prediction had
the wrong sign. Inserting is append-mostly, so cold pages are written once and
never re-read — paging is sequential and off the critical path. The swap device
is a real disk file (/swap.img; no zram, zswap disabled), so this is genuine
disk paging.
Do not generalise this to "swap is fine". It measures an append-mostly workload — and the opposite pattern was then measured too, below.
The operational consequence is that the RAM ceiling has two shapes and neither reports itself: without swap the process vanishes on signal 9, with swap it keeps returning 0 while serving from disk. A budget that fires at a declared threshold is the only one that can speak before either happens.
Durability across the ceiling
60 000 Int rows, 8 MiB cap, swap off, WO_DATA set — the process is OOM-killed
mid-insert, then replayed:
| Claim | Result |
|---|---|
| the survivor is a contiguous prefix | ✅ ~40 000 rows, rows 1..M all present |
| every surviving row's payload is correct | ✅ every v matches item_v(i) |
| the truncated tail is not read as corruption | ✅ replay exits 0 |
Ack-after-fsync holds through an OOM kill — the one shutdown path that skips
every cleanup handler. Gated as db-bench's ceiling leg, which asserts the
shape of the survivor rather than its size: where the SIGKILL lands is the
scheduler's business, so rows_recovered carries ±100% tolerance. The leg
asserts the exit but never records it as a metric, so that when databasev2 2's
byte budget turns the kill into a checked refusal, the gate does not fail on the
improvement.
Random reads over an oversized table: 273×
60 000 Int rows, both legs reading the same Weyl key order
(i*2654435761 mod n), differing only in the cap:
| Leg | Cap | Throughput | p50 | p99 |
|---|---|---|---|---|
| all resident | 256 MiB | 1 851 166 reads/s | 0 µs | 1 µs |
| over-cap, swap on | 6 MiB | 6 771 reads/s | 128 µs | 487 µs |
All 20 000 reads resolved in both legs, so this is the cost of faulting pages back, not of failed lookups. Swap-off is not an option in this configuration — it is SIGKILLed.
The two access patterns are ~270× apart under identical memory pressure:
| Pattern | Cost of exceeding RAM |
|---|---|
| append-mostly insert | ~1% (cold pages written once, never re-read) |
| random read across the table | 273× |
Departure is a step, not a curve. 1 µs to 487 µs with nothing in between —
p99_departure_decile looks for a gentle knee that does not exist. Residency is
close to binary, which is why a budget must fire at a declared threshold: there
is no early warning in the latency signal to react to.
Mechanism caveat, and it is a design input for databasev2 2. This is
demand-paging of anonymous slab memory through swap — 4 KiB per fault, no
readahead. resident: keys instead preads rows from the WAL, through the
page cache: same physical constraint, different mechanism, plausibly a better
constant because file reads get readahead and a shared cache. That is a
hypothesis. 273× bounds what swapping costs; iteration 2 must measure its own
read path rather than inherit this figure.
Gated as db-bench's randread leg, which gates the ratio — the absolute
reads/sec of the over-cap half is the box's swap device, while the factor between
two runs differing only in their cap is the engine's.
Replay: boot cost tracks history, not data
Two stores with the same 20 000 live rows and different history lengths. Process startup (3.5 ms, empty store) is subtracted, so these are replay:
| Shape | Records | WAL used | Replay | Per record |
|---|---|---|---|---|
| N inserts | 20 000 | 980 035 B | 110 ms | 5.5 µs |
| N inserts + N updates | 40 000 | 1 960 035 B | 211 ms | 5.3 µs |
1.9× the boot cost for an identical dataset. Per-record cost is flat, so replay is linear in records, not rows. An update appends a record and nothing ever collapses it, so a row updated a thousand times costs a thousand records at every boot, forever.
Extrapolated at 5.5 µs/record: 10M records ≈ 55 s of boot, 100M ≈ 9 minutes.
This is the "before" databasev2 3 lacked — bench/baseline.json carried no
replay, restart, boot or recovery metric at all, because iteration 22 proved
restart correctness and never timed it. Gated as db-bench's replay leg:
replay.inserts.*, replay.history.*, replay.history_penalty_x. Per-record
cost is stored in nanoseconds — as µs it rounded 5.5 and 5.3 to 6 and 5,
which is too coarse for the one number a checkpoint is meant to improve.
WAL bytes are measured as the file's non-zero prefix, never its size: shard
WALs are fallocate'd to 1 MiB, so an empty store reports 1048576.