The branch was 17 ahead / 25 behind with 11 conflicting files, and drifting
further: db.c had been rewritten twice on master since (group commit, then
compaction). Resolved rather than rebased so both histories stay legible.
Conflicts, and how each was settled:
- db.c: BOTH semantics kept. Master's fatal path and compaction check now sit
behind the branch's `table_is_durable` predicate, in all three inline arms —
a volatile table reaches neither the barrier nor the compaction check
- db-bench sample: every mode from both sides (growth, growth-verify, randread,
replayseed, wmix) and ONE `boot` mode, which both sides had added
independently
- db-bench.py: all six legs kept. Both sides had also grown the same
WAL-size helper under different names; collapsed into one
- perf-targets: the branch's §5 (RAM ceiling) then master's §6/§7 — master's
numbering had already assumed a §5 it did not have
- story frontmatter: master's `status` (the landing truth) plus the branch's
`readiness` axis. 03 would have read `done` + `refine`, which is a
contradiction — it was brainstormed and landed on master, so `ready`
- board: both standup blocks newest-first; master's chain rows (a superset);
the branch's databasev2 1-2 rows with master's 3-4. Fixed a stray `|` in
master's row 3
- baseline: master's, then REGENERATED from a full campaign — 143 metrics,
132 checks, 0 failures with both sides' legs present
TWO HALF-EXPOSED FEATURES FIXED, because the merge rule is that master gets
no feature that is honoured in name only:
- `resident: keys` PARSED, set a .wob flag, and did nothing: rows stayed fully
resident. A developer could declare a 120 GB table keys-resident, watch it
compile, and be OOM-killed. The loader now REFUSES it with a message naming
what to write instead, until tasks 5c/5d land. The compiler still parses it
and its AST golden still passes, so the grammar work stays tested
- `durable: false` was honoured ONLY on the inline path. wo_db_exec_req had no
guard at all, so a volatile table written from an actor on a worker shard
would still be logged — precisely porch's session-table case, and precisely
what iteration 2 exists to provide. All three request-path arms now carry the
same predicate. Found by reading the merged code, not by a test: the obvious
probe runs main() on the primary and therefore only exercises the inline path
Verified on the merged tree: wovm-test 0, woc-test 0, oop-e2e 122/0,
residency-accept 8/0, db-bench 132/0, linkcheck clean.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
databasev2 3, task 5.
Full campaign, same workload twice, differing only in whether
checkpointing may fire:
- WAL used 1962358 -> 907094 bytes (2.16x reclaimed)
- boot 114 -> 64 ms (1.78x), median of 3
- stop-the-world pause max 2651us against a STATED 50ms budget
The budget is asserted, not assumed: 50ms is a stall a serving process
can absorb without a client seeing a timeout, and the leg fails if it is
exceeded. The pause is O(live rows) — at ~181 MB/s a 1GB live set implies
~5.5s, which is the number an incremental design must be bought against.
The spec deliberately did not buy it in advance.
FOUND BY MEASURING: the dump was 8x slower than it needed to be. It
flushed through wo_wal_commit, which fdatasyncs, so it paid one barrier
per 256 records. Intermediate durability there is worthless — the temp is
not authoritative until the rename and is fsynced once immediately before
it. With a single final barrier:
- ~107KB live: 23948us -> 2903us
- ~500KB live: 36361us -> 7526us
- ~1.98MB live: 107649us -> 13212us
- marginal ~22 MB/s -> ~181 MB/s, sync-bound to bandwidth-bound
Correctness re-proven after that change: wovm-test 36 suites 0 fail,
test_wal 760 pass including the 40-round kill-during-compaction battery.
Two measurement defects of my own, fixed rather than reported:
- boot measured through the driver's run() helper reported 251ms both
with and without checkpointing — run() samples RSS on a 250ms poll, so
every timing floors at the quantum. Measured directly instead, median
of 3
- ckpt.reclaim_x was recorded as lower-is-better by the default detector,
which would have PASSED "reclaimed nothing" and FAILED an improvement:
the feature's central claim, gated backwards. Now higher-is-better,
gated at 15% while the wall-clock metrics stay wide — waiving them all
would have left the leg ungated, part A's task 4 mistake
- sample gains a `boot` mode that does nothing, so boot time is boot time
- walstats now reports compactions, pause max/total and compacted bytes
- baseline refreshed from the FULL campaign (N=20000, crash_reps=3), and
a fresh full run passes 116 checks 0 failures
- gate bites: reclaim_x doctored to 1.0 -> FAIL on exactly that metric
One flake seen and checked, not papered over: durable.sN.query.ops_sec
failed once at 53% below baseline. It is a read-only metric that touches
no WAL code, and a re-run passed 116/0 with the box at load 1.85.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
databasev2 4 part A, task 4. Scope extended with developer approval: the
plan authorised touching the sample only for observability, but no
existing leg has enough concurrent durable writes to exercise group
commit at all, so the payoff was unevaluable either way.
The finding that forced it:
- `mix` writes on one op in ten with C=4 (all_mode calls mix_mode(n/10,
4); Mixer writes on i % 10 == 9), so the quick run performs 20 writes
total. Measured mean batch 1.01 over 3112 barriers, peak 3
- that is a property of the WORKLOAD, not the mechanism: peak 3 of a
possible 4 shows batches form whenever writes actually coincide
- `wmix N C` added: every op a durable write, C at once. Updates rather
than inserts, so it is comparable to mixwrite and the row count stays
flat. Histogram kind 2 — a replayed store still holds the seeding run's
kind-0/1 Hist rows and merging those would report someone else's
latencies
- WO_WAL_STATS=1 prints one line at exit: batches, records, peak_batch,
peak_staged. Opt-in, because it would otherwise pollute every durable
program's output. Counters live in wo_wal; no builtin, the numbers are
diagnostic and not part of the language
Measured, and it scales with concurrency exactly as designed:
- C = 4 / 16 / 64 -> mean batch 1.13 / 1.76 / 5.35, peak 3 / 10 / 39
- the gate's own legs: durable.s1 5412 records over 5412 barriers (mean
1.0, peak 1 — the inline path, one barrier per statement BY DESIGN),
durable.sN 7757 over 2296 (mean 3.38, peak 28) at 2x the throughput
- peak staged 1372 B settles the no-cap decision with a number: the batch
is tiny, so the upstream mailbox bound is sufficient
- mean_batch/peak_batch are higher-is-better (the default detector would
have called bigger batches worse)
- only the batch SHAPE metrics are waived to 100%; wmix throughput and
latency keep real tolerances (15% s1, 50% sN) — a blanket waiver would
have left the entire new leg ungated
- the live assertion `mean > 1.0` on the sN leg is what catches inertness
- gate bites: sN wmix ops_sec -60% -> FAIL on exactly that metric, 1 of 86
Verified: db-bench-quick 89 checks 0 failures; baseline refreshed (86
metrics).
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Closes the last gap in databasev2 1; gives databasev2 3 its "before".
- `boot` mode: does NOTHING. WO_DATA replay runs before main, so a mode
with no work measures replay plus a fixed startup
- `replayseed N M`: N inserts + M updates — same live rows, longer log
- `replay` leg: empty-store startup floor measured and SUBTRACTED, then
two shapes timed, median of 3 boots each
- premise check: updates must actually append WAL records, else the two
shapes are one measurement and the penalty means nothing
- WAL bytes = non-zero prefix, never file size (fallocate'd to 1 MiB)
- per-record cost stored in NANOseconds: as us it rounded 5.5 and 5.3 to
6 and 5, too coarse for the number a checkpoint exists to improve
- 148 checks, 0 failures; gate bites on a doctored ns_per_record
Measured — same 20 000 live rows, different history:
- 20 000 records: 980 035 B WAL, 110 ms replay, 5.5 us/record
- 40 000 records: 1 960 035 B WAL, 211 ms replay, 5.3 us/record
- 1.9x boot cost for an IDENTICAL dataset; per-record cost flat, so
replay is linear in records not rows
- extrapolated: 10M records ~55 s of boot, 100M ~9 min
- databasev2 3 correction: it planned to use "22's aged-store replay
numbers", which never existed — 22 proved restart correctness, never
timed it
- databasev2 3 hazard recorded: compaction rewrites the log and moves
every record, so it invalidates every `resident: keys` offset — an
arbitrary byte in a rewritten file, not stale-but-readable
- databasev2 1 -> status: done
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
- `randread N R` in the sample: fill N rows, read R across the WHOLE range
- Weyl order `i*2654435761 mod n` — no RNG in the language, none needed;
both legs read the SAME key order so residency is the only variable
- `randread` driver leg: control (256 MiB, does not bind) vs over-cap
(6 MiB + swap), sizes kept modest — quick resolves it in ~5s
- gates the RATIO, not the absolutes: over-cap reads/sec belongs to the
box's swap device, the factor between two runs belongs to the engine
- reads must all resolve (hits == R) or the leg fails; a collapse measured
over unresolved reads is noise
- 133 checks, 0 failures; gate bites on a doctored collapse_x
Measured — this closes the gap the swap leg left:
- resident 1 851 166 reads/sec, p50 0us p99 1us
- over-cap 6 771 reads/sec, p50 128us p99 487us
- 273x throughput, ~480x p99, all 20 000 reads resolving in both
- so the two access patterns sit ~270x apart under identical pressure:
append-mostly insert ~1%, random read 273x
- departure is a STEP not a curve (1us -> 487us, nothing between), which
is why p99_departure_decile finds no knee — there is none
- caveat recorded, NOT inherited: this is demand-paged anonymous memory
through swap (4 KiB/fault, no readahead). `resident: keys` preads via
the page cache — should be better, but databasev2 2 task 7 must measure
its own read path. New criterion added there
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
- `Wide` text-heavy reference shape beside Int-only `Item`
- `growth N int|text`: per-decile RSS read from own /proc/self/status
- `growth-verify`: survivor of a crash must be a contiguous intact prefix
- four footprint legs under a rootless cgroup v2 cap, swap on/off
- `ceiling` leg: die at the cap, then replay must come back intact
- footprint read as median-of-marginals; doublings a separate metric
- 121 checks, 0 failures; footprint gated ±10%, kill-timing ±100%
Measured, and it inverted two of the iteration's own predictions:
- footprint 96.5-100 B/row Int vs 320.6-324 B/row text = 3.3x, NOT the
"order of magnitude" three docs asserted
- table storage has NO checked ceiling: SIGKILL signal 9, not a catchable
WO_T_OOM. overcommit lets malloc succeed; kernel kills on page touch
- swap is NOT latency collapse: 900k rows 148s capped-with-swap vs 150s
uncapped. Append-mostly never re-touches cold pages
- ack-after-fsync survives an OOM kill: ~40k rows, no holes, no corruption
- iteration 2's budget dependency is REMOVED not satisfied — there is no
"swap onset" to derive a fraction from
- fix: subprocess returncode -9 was labelled a "checked refusal"; 137 is
the shell spelling of the same signal
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
- Mixer actors: 90/10 read/write, per-actor histograms merged through
the store itself (Hist rows) — exact aggregate percentiles
- msgrate: one-way flood at a worker-placed sink; measured 15.3M
msgs/s same-heap vs 2.06M cross-shard — the mutex-inbox number
- finding: point lookups are O(table) (probe walks all slabs), so
read-heavy mix is quadratic in store size — all-mode calibrated to
N/10 mix ops; the number 22 exists to publish
- TSan clean both shard counts (setarch -R, fibers-gate pattern)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
- seed/read/query/write/wal/verify/verify-acked + all (one-process
campaign: RAM store dies with the process)
- per-op time.ticks, 1us-bucket histogram percentiles (reservoir
deviation: no element-write/sort in language; better tail anyway)
- Meta expectation rows ride the same WAL verify checks
- finding: hand-built multi<TableClass> SEGVs on drop (elems classed
OWNED, refs are scalar ids) — worked around, recorded
- finding: reads ~1.6k/s p50 595us vs 287k/s inserts — probe walks
all slabs; the number 22 exists to surface
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>