Plan for the approved spec. Code-free per the repo convention
(docs/plan/discarded.md:54); the executor writes the code.
- T1 a failed barrier is detected and fatal — one entry point that names
the operation, errno, WAL path and batch size, then exits. The abort
path itself stays unexercised and the task says so rather than buying
coverage with a fault-injection switch
- T2 the barrier moves to the drain point and replies are held; the
request path stops committing per append. Riskiest task, and its risk
is one place: the crash legs. Plan says STOP if they fail, do not
adjust the test
- T3 the inline path takes the same fatal rule but keeps its own barrier,
with a comment explaining the asymmetry so the next reader does not
"fix" it. Looks like a no-op; without it the two paths disagree, which
is the unevenness the spec exists to remove
- T4 prove batches actually form BEFORE measuring the payoff — otherwise
a win gets attributed to the wrong cause. Also records peak staged
bytes, settling the no-cap decision with a number
- T5 measure, gate, write it down. If the payoff is absent, say so and
stop: part B must not start on an unproven premise
- T6 closeout, including the error catalogue — WO_T_IO leaving the write
path is language-visible and must be written down
Spec corrected while planning: it pointed at durable.s1.seed as the
payoff. Wrong, structurally — worker shards hold no WAL, so a queue only
exists when other shards write, and a serial writer has nothing to batch
with. The real target is durable.sN.mixwrite: 480 ops/s at p99 5888us
against s1's 1023 at p99 664, so adding shards currently makes durable
writing WORSE. That inversion is a better argument for the iteration than
the one the story recorded.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Brainstormed 2026-08-28. The iteration is split: part A batches, part B
(io_uring submission) is deferred until A's measurement says whether the
blocking boundary still dominates.
The story's premise needed correcting first:
- it says "replace fsync-per-commit with io_uring group-commit", but the
engine commits per STATEMENT — db.c calls wo_wal_commit right after
every append, all six sites, so each row change is one pwrite + one
fdatasync
- so two independent wins were being carried as one, and only the second
needs io_uring. The staging buffer already holds any number of records;
today it never holds more than one. Part A is mostly deleting calls
- iteration 22's numbers say A is where the payoff is: durable writes
4460 ops/s, mixwrite 1023 ops/s p99 664us, against 1.28M ops/s reads
Forks settled:
- batch boundary is QUEUE-DRAIN, not the tick this story had recorded: a
tick adds latency to a lone writer, taxing an idle system to serve a
busy one. Queue-drain self-tunes and needs no knob
- shard 0 holds each reply envelope instead of sending it, commits once
when the queue empties, then releases all — so a writer is acked after
the barrier carrying ITS record, which today is true only because
every batch has one member
- a failure between "RAM mutated" and "record durable" is a FATAL,
diagnosed abort. This replaces uneven behaviour that already exists:
insert rolls back, update and delete do not and say so in a comment
("RAM ahead of disk"). Batching would have multiplied that
- consequence stated, not slipped in: WO_T_IO leaves the write path
- no batch cap initially; peak staged bytes is measured so the question
is settled by a number
One gap disclosed rather than hidden: forcing a real fdatasync failure
needs mount privileges, so the unit test proves the error is DETECTED and
the abort itself stays covered by inspection.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
- docs/stories/databasev2/, numbered from 1. Six PENDING database iterations
moved from the language track and renumbered, keeping the old id in
`was_language_iteration:` so a search for "iteration 32" still finds it:
32 -> 3 WAL checkpoint, 23 -> 4 io_uring commit, 33 -> 7 single-file store,
27 -> 8 query grammar, 20 -> 9 cross-program, 21 -> 10 keypair auth.
Done work (9, 9b, 22) stays as v1 history; language 18 left whole
- the problem, read off the engine not guessed: rows are malloc'd slabs with
addresses stable forever, NO eviction/spill/paging anywhere in database/src,
the WAL never checkpoints so boot replays all history, and durability is one
process-global WO_DATA so no table can say it matters more than another.
An allocation failure IS a clean catchable WO_T_OOM — but swap thrash
arrives first and carries no error signal at all, which is the real hazard
- four new iterations:
1 measure the ceiling FIRST (curve not cliff; the three exits; kill -9 at
exhaustion) — every later default should follow from a number
2 `@table(mode: ram | durable | cold)` — the grammar ask. Small surface
(Ast.table_cfg gains a key, the parser already rejects unknown args), big
semantics: `durable` defaults so nothing changes silently, and the
compiler refuses a durable row holding a `ref` into a ram table
5 bounded tables + refuse/evict/back-pressure, shedding BEFORE the OS acts
6 cold tiering — mostly forks, incl. whether the language surfaces the
fault cost and whether @unique on cold is refused outright. A paged
B-tree stays rejected: if tiering needs one, reject tiering
- 39 links repointed, link TEXT renumbered to track-local ids; arc gains one
pointer row replacing the six moved; board + board-views cover three tracks
- linkcheck 0 broken / 0 anchors; no code blocks in any story
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>