- retitled "group commit, and the async barrier"; the 2026-08-15/20/28 banners compressed into a trail; `readiness: refine`, `review_pending` (forks 1–5 and 8–10 decided under autonomy 2026-09-10 by codd-shoney; 6 and 7 keep readiness at refine) - what part B is FOR: the read tail on shard 0 while a barrier blocks — and only that; mechanism: the drain pwrites as today, then submits ONE bare IORING_OP_FSYNC and keeps working; held replies released by the completion; the epoll fallback is part A unchanged; ordering with compaction and the deferred drops/re-points; the inline durable write on shard 0 unified through its own inbox while a barrier is in flight; completion delivery on a busy shard 0; shutdown reaps an in-flight barrier before wo_wal_close - fork 6 (kernel floor and raw-syscall shape) goes to lintor; fork 7 (go/no-go) is settled by one measurement — tmpfs vs ext4 `mixread.p99` — then one developer answer; both answers already sit in .dev/zack/databasev2-4b.md, the fold into this story is pending - progress table B1–B9 with sizes and owners (cyril B1/B8, lintor B2, the runtime agent B3, codd + pm B9); no new knob, no new dependency Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> (cherry picked from commit 7ceb7b8805da41e67661744746bf2d0b7879cc50)
34 KiB
| track | iteration | was_language_iteration | status | readiness | review_pending | chain |
|---|---|---|---|---|---|---|
| databasev2 | 4 | 23 | in-progress | refine | part B forks 1–5 and 8–10 decided under autonomy 2026-09-10 by codd-shoney; forks 6 (lintor) and 7 (cyril's tmpfs ceiling) are open and keep readiness at refine | 5 |
databasev2 4 — group commit, and the async barrier (was: io_uring group-commit write path)
Moved 2026-08-26 from the language track, where this was iteration 23. Part of Story — databasev2: the database beyond RAM. Content unchanged by the move; its dependencies are restated in that track index.
Format:
product/story-iteration-template. Part of Story — one language, one runtime, one database, one binary — the track this iteration was authored in before the 2026-08-26 move.Inserted 2026-08-15. The write-path optimization, and deliberately the LAST database performance iteration: it only earns its complexity once there is a measured fsync-per-commit baseline to beat (iteration 22) and a multithreaded runtime to overlap against (iteration 8). Doing it earlier would optimize a number nobody had measured, against a runtime that couldn't use it.
Supporting evidence for staying last (iteration 1, 2026-08-27): the write path is not where memory pressure bites. Inserting 900 000 rows inside a 64 MiB cap with swap cost ~1% (148 s vs 150 s uncapped), because appending never re-touches its cold pages. Random reads over the same oversized table cost 273×. So the pressure is on the read path, and the io_uring question that may actually matter is the one iteration 2 deferred here — io_uring for
resident: keysrow reads — not group-commit for writes.REFINED 2026-08-20 (developer confirmation, no code): the four original forks were settled as their leanings — drop-in behind
wo_wal_commitfirst; rawio_uring_setup/io_uring_enter, libc only; the batch boundary the shard tick; a startup probe plus the arc-wideWO_IO=uring|epolloverride. Position re-sequenced 2026-08-21 to fifth in the concurrency chain; the per-shard ring already exists (arc T4,WO_IO=uring|epoll) and the developer's io_uring-first directive of 2026-08-20 says the WAL's durability ops ride the SAME per-shard ring the fibers park on — one event loop per shard. Those settlements are recorded in Info below under "the 2026-08-15 forks, for the record"; the tick boundary was later replaced by queue-drain (part A) and "WRITE+FSYNC chains" by a single FSYNC (part B, fork 2).
BRAINSTORMED 2026-08-28 — and SPLIT IN TWO. Spec for part A:
2026-08-28-wal-group-commit-design.md· plan:2026-08-28-wal-group-commit.md(6 tasks).The premise needed correcting. This story said "replace fsync-per-commit with io_uring group-commit", but the engine committed per statement:
db.ccalledwo_wal_commitimmediately after every append, at all six sites. That split the goal into two independent wins, and only the second needs io_uring:
- Part A — batching. Let many statements share one barrier. Landed 2026-08-28, ≈2.9× concurrent durable write throughput.
- Part B — the async barrier. The shard submits the barrier and keeps working instead of blocking in
fdatasync. Re-brainstormed 2026-09-10 on the corrected premise below.Forks settled in the 2026-08-28 brainstorm: batch boundary is queue-drain (not the tick — a tick taxes an idle system to serve a busy one); a failure between "RAM mutated" and "record durable" is a fatal, diagnosed abort, which removed
WO_T_IOfrom the write path — a language-visible change, recorded here deliberately.RE-BRAINSTORMED 2026-09-10 (part B, codd-shoney, autonomy). The board's original target for part B — "close the 66× gap" — was wrong, and the story's own part-A closeout said so. What part A left behind is a latency tail on shard 0: everything queued behind a barrier waits for the device, and
durable.sN.mixread.p99measured 1043 / 2318 / 4147 µs across three identical runs. Part B is now FOR that tail, and nothing else. Ten forks, eight settled from code evidence, two open with the measurement and the kernel question that settle them named — see Info — the forks, settled.readiness: refineuntil those two return.
Progress — part A landed 2026-08-28
| # | Task | State |
|---|---|---|
| 1 | a failed barrier is detected, and fatal | ✅ d3ff03e |
| 2 | one barrier per drain; replies held | ✅ b9b8a45 |
| 3 | the inline path takes the fatal rule, asymmetry documented | ✅ a6ccdbe |
| 4 | prove batches form — the wmix write-concurrent leg |
✅ 40d029c |
| 5 | measure the payoff, gate it, record it | ✅ d52ea8a |
| 6 | closeout | ✅ this change |
Progress — part B, the async barrier (brainstormed 2026-09-10)
Sizes: S ≤ half a day, M ≤ two days, L longer. B1 and B2 gate everything after them; nothing below B2 starts until both have reported.
| # | Task | Size | State |
|---|---|---|---|
| B1 | the ceiling measurement (codd-cyril): mix at default shards, same build, WO_DATA on tmpfs vs on bench/ (ext4), three runs each; report mixread ops/s, p50, p99 — settles fork 7 |
S | ⬜ |
| B2 | the kernel questions (lintor): floor and shape of a bare IORING_OP_FSYNC on the 5.4 ring, io-wq behaviour on the floor kernel, teardown with a barrier in flight, how to detect the op at boot — settles fork 6 |
S | ⬜ |
| B3 | the park-plane seam (owner: the runtime agent, not codd): submit one FSYNC SQE with a WAL sentinel, a non-blocking reap, a completion hook; the reap loop matches the sentinel before its fiber cast | S | ⬜ |
| B4 | wal.c: split the write from the barrier — the drain's pwrite advances off as today, the barrier becomes submit-or-queue with an in-flight mark and a synced-bytes counter; boot self-test of the op with the loud fallback notice; stats line names the path |
M | ⬜ |
| B5 | vm.c drain: pwrite, then submit or queue behind the in-flight barrier; the two held-reply lists move from drain locals into shard state; the completion handler releases ITS list, then runs the compaction check; the hot path polls completions where it polls the inbox |
M | ⬜ |
| B6 | the inline unification: a durable-table write on shard 0 takes the request path through its own inbox; reads and volatile writes stay inline; guarded by durable.s1.seed.p50us staying inside its 15% tolerance |
M | ⬜ |
| B7 | shutdown: shard 0 reaps an in-flight barrier to completion (fatal rule applies) before wo_wal_close; never closes the fd with a barrier outstanding |
S | ⬜ |
| B8 | gates and numbers (codd-cyril): durable legs and the crash battery under BOTH WO_IO=uring and WO_IO=epoll; a test_wal case that submits FSYNC on a bad fd and sees the failure detected; baseline re-run; perf-targets.md §6 addendum with the before/after against the B1 ceiling; re-tighten the 300% tolerance if the tail stabilises |
M | ⬜ |
| B9 | contracts and closeout: CODE-LOGIC.md "Group commit" section extended with the async barrier, 04-db-binding.md commit-contract paragraph, story/board/graph (codd + codd-pm) |
S | ⬜ |
The payoff, measured two ways
| Measurement | Before | After |
|---|---|---|
controlled (same build, only db.c/vm.c swapped; wmix 4000 32) |
2213 · 2177 ops/s, p50 7183 · 7251 µs | 6216 · 6525 ops/s, p50 3458 · 3444 µs |
committed baseline: s1 inline vs sN batched |
1467 ops/s, mean batch 1.0 | 5117 ops/s, mean batch 5.43, peak 57 |
≈2.9× throughput, ≈2.1× lower p50, and the two methods agree (2.9× and 3.5×). Batching scales with contention: mean batch 1.13 / 1.76 / 5.35 at C = 4 / 16 / 64.
The cost side, and a bug the battery caught
Reads were being held behind the barrier. The drain first held every DB
reply until the commit — including reads, which stage nothing. mixread p99 rose
from ~1043 µs to 4057 µs until only staging statements had their replies
held. Caught by the gate, not by review.
What remains is inherent: a barrier blocks the owner shard longer (more
records per fsync) though less often, so anything queued behind one waits. Three
full runs of the same build gave durable.sN.mixread.p99 of 1043 / 2318 /
4147 µs — a 2–4× spread near idle. So part A buys ~3× write throughput at the
cost of a longer, noisier tail on the owner shard. durable.sN.*.p99us was
re-baselined at 300% tolerance for that reason, with the floor as the real guard
(mixread's came within 25 µs of tripping).
This is the strongest argument for part B — submitting the barrier and continuing to serve is exactly what removes this cost.
What did NOT improve — and it was predicted
durable.sN.mixwrite: 480 → 492 ops/s, i.e. unchanged. This was the spec's original payoff metric, and correcting it was part of the brainstorm:mixwrites on one op in ten with C=4, so a quick run performs 20 writes and measured mean batch 1.01. A workload that never has two writes in flight cannot be helped by batching them.durable.*.seed: unchanged. A serial single writer has nothing to batch with, under any scheme.- This board's stated target was mis-stated. It read "close the 66× gap
iteration 22 measured (durable 4.5k vs ram 297k inserts/s)". Part A does not
close that gap and structurally cannot:
seedis serial, and one writer waiting on one barrier is a latency problem, not a batching one. Recorded rather than quietly renumbered. - The before-p99 is not a measurement.
hist_addclamps at 20000 µs and both before-runs pinned exactly there, so the true value is ≥20 ms and unknown. The gain is at least 2.3×.
Goals
- One barrier per drain, not per statement (part A, met). Many statements
share one
pwrite+fdatasync; each writer is acknowledged only after the barrier that carried its record. The batch boundary is the queue going empty — no tick, no timer, nothing to tune. - Shard 0 never blocks in the barrier (part B, rewritten 2026-09-10). The
drain writes its batch into the page cache and submits the durability barrier
to the shard's own ring, then goes back to serving. Reads arriving while the
device flushes are answered before the flush completes; the next drain's
writes are written and queued behind the in-flight barrier. The metric is the
read tail on the owner shard:
durable.sN.mixread.p99us(baseline 4050 µs; measured 1043 / 2318 / 4147 µs) anddurable.sN.mixread.ops_sec(4733), against the no-barrier referenceram.sN.mixread(p99 81 µs, 44 874 ops/s) and the tmpfs ceiling B1 measures. Write throughput may rise as a side effect of the disk never idling between barriers (durable.sN.wmix.ops_sec, 6017); it is watched, not targeted. - One write path, not two. A durable-table write on shard 0 takes the same
request path a worker's does, so the inline path's private barrier — the
asymmetry part A documented as a standing hazard — goes away, and
WO_SHARDS=1gets batching and the async barrier with it. - Keep the durability promise byte-for-byte. Every guarantee iterations 9 and 22 proved — replay-whole-or-not-at-all, torn-tail drop, no acknowledged write ever lost, durable-or-process-death — holds identically. Part B changes WHEN shard 0 learns the barrier finished, never WHETHER an ack means durable.
Acceptance Criteria
Met:
- Given the batched write path under iteration 22's crash battery, when
it runs, then every acknowledged write is present after replay. ✅ —
crash.sN(the batched path) recovered every acked row afterkill -9,crash.s1likewise, and both restart legs replay byte-true. This was the one thing batching could break. - Given the durable write benchmark before and after, then throughput is
materially higher and p99 lower, recorded. ✅ ~2.9× and ~2.1× (p50); see
perf-targets.md§6. Scoped honestly: on a write-concurrent workload only, and p99's "before" is at the histogram ceiling. - Given batching, when it runs, then it is proven to engage rather than assumed. ✅ mean batch 5.43, peak 57 on the gated leg, and the live assertion fails the suite if the mean drops to 1.
- Given a durability failure, when it happens, then the engine does not continue with RAM ahead of disk. ✅ fatal, diagnosed, exit 74 — replacing three behaviours that disagreed.
Outstanding (part B; rewritten 2026-09-10):
- Given the ceiling measurement B1, when
mixruns on tmpfs (a free barrier) against ext4 on the same build, then the gap between the twomixreadfigures is recorded and is the go/no-go for everything below. If tmpfsmixread.p99is not materially below the ext4 figure, the tail is not the barrier and part B closes as "measured, not built" (fork 7). - Given a barrier in flight on shard 0, when worker reads arrive,
then they are answered before it completes:
durable.sN.mixread.p99uswithin 2× of the B1 tmpfs figure across three runs, anddurable.sN.mixread.ops_secat least half of it;durable.*.seedand everydurable.s1.*leg inside their 15% tolerance. - Given the async path, when a writer's reply is released, then the
barrier that carried its record has completed: the crash battery recovers
every acked row after
kill -9underWO_IO=uring, and again underWO_IO=epoll. - Given
WO_IO=epoll, or a ring that cannot run the op, when the runtime starts, then the drain commits synchronously exactly as part A does, one stderr line says so when the fall-back was not forced, andWO_WAL_STATS=1names the path a run took. - Given
WO_SHARDS=1with concurrent writers, whenwmixruns, then batches form (durable.s1.wmix.mean_batchabove 1.0, today exactly 1.0) — the inline unification's proof. - Given a completion that reports failure, when the handler sees it,
then the process ends with the same diagnostic and exit 74 that part A
gives a failed
fdatasync. The unit test proves the failure is detected (an FSYNC submitted on a bad descriptor completes with an error); the exit stays covered by inspection, disclosed exactly as part A disclosed it. - Given the process exits with a barrier in flight, when shard 0 tears down, then it waits for the completion first; the stats line's submitted and completed barrier counts agree at exit.
Part B — its premise changed (2026-08-28), and what replaced it (2026-09-10)
Part B was justified by "close the 66× durable gap". Part A showed that framing
was wrong: the gap is two problems. Concurrent write fan-in was a batching
problem and is now ~3× better. What remains for a serial writer is a device
flush it must wait for before its ack — no submission model shortens that, and
acknowledging before the flush is the one thing principle 7 forbids. The tail
that IS addressable is shard 0's: while it sits in fdatasync, every read and
every write queued in its inbox waits for the device. The forks below settle
part B against that, and only that.
Out Of Scope
- io_uring for the network/accept path. This iteration is the WAL write path only; the socket side is the shard-actor runtime's and the network layer's concern.
- io_uring for reads. For a fully-resident table reads never touch a
descriptor. A table declaring
resident: keys(databasev2 2) doespreadrows from the log, and accelerating that is a real, separate question — not this iteration's, whose acceptance is a durability tail. - Async commit — acknowledging a writer before its barrier completes,
PostgreSQL's
synchronous_commit = off/XLogSetAsyncXactLSN. Refused: principle 7, "an ack means the commit reached disk … none of it is negotiable". It is the only thing that would help a serial writer's latency, and it is not on offer. - A linked WRITE→FSYNC SQE chain, registered buffers, fixed files, SQPOLL,
sync_file_range. Fork 2 explains why the chain buys nothing here; the rest are ring features to add only when a measurement names one. - Replacing the WAL format or the commit contract. The bytes on disk and the meaning of an ack are iteration 9's; this changes when shard 0 learns the barrier finished, not what the log contains.
transaction { }(language 18, hold). A transaction is a staged batch; under part B it is one drain's write and one barrier. The two compose without either knowing the other.
Info — the forks, settled
Part B — ten forks, 2026-09-10 (codd-shoney, under autonomy; the
frontmatter's review_pending asks for the developer's second look at the eight
decided ones, and names the two that keep readiness at refine).
-
What part B is FOR: the read tail on shard 0 — and only that. Three candidates were on the table. (a) The tail: reads and writes queued in shard 0's inbox wait while the drain blocks in
fdatasync(runtime/src/vm.c233–241 callswo_wal_commit_fatal, which ispwrite+fdatasyncatdatabase/src/wal.c965–981); measureddurable.sN.mixread.p99us1043 / 2318 / 4147 µs, baseline 4050, where the same read path with no barrier in the way (ram.sN.mixread) sits at 81 µs and 44 874 ops/s against 4733. (b) Serial-writer latency (durable.s1.seedp50 212 µs): refused — one writer needs one completed device flush before its ack, and the only mechanism that shortens the wait is acknowledging early, which principle 7 forbids; io_uring moves the wait off the thread, it does not shorten it. (c) Throughput: part A already took the batching win; pipelining (the next drain writes while the previous barrier flushes) may raisedurable.sN.wmix.ops_secfrom 6017 because the disk stops idling during the drain's execution phase — recorded as a watched side effect, not the goal, because a goal needs one number and (a) has it. Decision: (a). The metric isdurable.sN.mixread.p99usand.ops_sec; the bar is fork 7's. Evidence that the tail is shard 0's and not the requester's: reads are already released before the barrier (vm.c207–214), so a read waits only when it ARRIVES during one — which is exactly what an async barrier ends. -
Mechanism: the drain
pwrites as today, then submits ONE bareIORING_OP_FSYNC(IORING_FSYNC_DATASYNC) on shard 0's existing ring; at most one barrier in flight; a drain that finds one in flight writes its bytes and queues behind it; the completion submits the queued one. Three options were weighed. (i) A linkedIORING_OP_WRITE→IORING_OP_FSYNCchain — the 2026-08-20 wording and the linux card's advice (docs/plan/exploration/linux/07-io_uring.md, "essential for WAL"). It needs the staging buffer to stay stable until the WRITE completes (a second buffer, and every reader of[off, off+len)—scan_record_stagedatwal.c311–330,wo_wal_next_offsetatwal.h222–242 — would have to learn a second not-yet-visible region), two CQEs to check with a short-write rule (io_uring/rw.c550–560 fails the request on a short write;io_uring/io_uring.c1841–1848 cancels the link), and it buys no asynchrony the plain write lacks: a buffered write to a regular file is punted to an io-wq worker anyway unless the filesystem advertisesFOP_BUFFER_WASYNC(rw.c1146–1155). The card was written for a Rust-era buffer model that no longer exists; the "same ring" half of the directive stands, the "chain" half does not. (ii) A dedicated fsync fiber or thread of ours — a thread plus a second wake mechanism, against codd's "no locks, single-threaded by contract"; and redundant, because the kernel already runs the op on a worker thread it owns:io_fsync_prepsetsREQ_F_FORCE_ASYNCandio_fsyncwarns if ever called non-blocking (io_uring/sync.c69, 79). (iii) Moving the barrier off shard 0 entirely — there is no other owner of the WAL (principle 5; workers assert they hold neitherdbnorwal,vm.c2362). Decision: the bare FSYNC.pwritereturning means the bytes are in the page cache, so every existing offset invariant holds unchanged —offadvances at the write exactly as it does atwal.c978, the staged region[off, off+len)exists only inside a drain,preadof any written offset returns data — and the WAL stays the only truth. What is new is bookkeeping about the file: an in-flight mark and a synced-bytes counter, both used only by compaction (fork 5) and the stats line. Cost: ~30 lines inpark.c(a submit with a WAL sentinel —uring_submitis static atpark.c136; a non-blocking reap; a completion hook; and the reap loop'selsebranch atpark.c392–397 casts any unknownuser_datato a fiber pointer, so the sentinel MUST be matched before it) — that seam is the park plane's, not codd's, and its owner is named in B3. PostgreSQL's shape, ported as behaviour: a backend that finds the flush lock held waits for the holder and re-checks whether its record was flushed for it (xlog.c2875–2889) — our "queue behind the in-flight barrier" is that, with the drain as the only backend. PG'scommit_delay(xlog.c2902–2906) is a timer knob and stays rejected for the reason part A rejected the tick. -
The ack contract across an async completion. Today the held replies are drain locals — "nothing here needs to outlive the batch it describes" (
vm.c98–101). Under part B they must: two FIFO lists in shard state, in-flight and next; each list is released by ITS barrier's completion, never earlier. A completion withresbelow zero takeswal_diewith "fdatasync" and-resas the errno — the same text and exit 74 (WO_EXIT_DURABILITY,wal.h304) part A gives a failedfdatasync; apwritefailure stays synchronous at the drain and takes the existing "pwrite" arm (wal.c1016). "Durable or process death" is unchanged. What widens is the WINDOW between RAM-mutated and durable, and it is the same exposure part A already accepted: a read arriving after a write in the same drain sees the write's effect and is answered before the barrier (vm.c207–214), so a process death loses nothing acknowledged and may have shown a reader an unacknowledged state — today's behaviour, longer. Rejected: a per-batch sequence number stamped on each reply (a third piece of state to keep consistent; FIFO-in-order completion makes two lists sufficient). -
The epoll fallback is part A, unchanged: synchronous
pwrite+fdatasyncin the drain. No helper thread — a thread of ours doingfdatasyncand signalling an eventfd would be a second mechanism to keep correct for the portability path, and principle 9 says the portability path is not where sharpness goes. The contract is identical on both paths; only where shard 0 waits differs. Two consequences, both gated: (a)scripts/db-bench.pynever setsWO_IOtoday — its durable legs and the crash battery (lines 257–274) run whatever the probe picks — so cyril addsWO_IO=uringandWO_IO=epolllegs to the restart proof and the crash battery, the patternscripts/db-actor-accept.sh49–50 already uses; (b) theWO_WAL_STATS=1line (runtime/src/main.c175–179) gains the barrier path and the submitted/completed counts, so no measurement can be misattributed. The baseline keeps the probe's default (uring on this box). -
Ordering with compaction, and the deferred drops and re-points. Compaction renames a new file over the log and re-opens the descriptor (
wal.c1317–1327); with an FSYNC in flight on the OLD descriptor, its completion would certify an unlinked inode. Decision: compaction runs only when staging is empty AND no barrier is in flight — the refusal atwal.c1214 gains the second condition as its backstop, and the drain's compaction check (vm.c250–267) moves into the completion handler, after that batch's replies are released, keeping today's order. Compaction itself stays synchronous; it is rare and measured at 2.7 ms. The keys-resident drops and re-points were deferred to "after the commit" because "the record is still in the staging buffer, so its offset would pread zeros" (wal.h96–98,wo_db_flush_dropsatwal.c451–461); the reason is the staging buffer, not durability, and the drain'spwriteends it. Decision: flush drops and re-points right after the drain's write, keeping one pending list and no per-batch split; a later barrier failure kills the process, so nothing observable depends on the difference. zack rewords thewal.h339 comment ("Call ONLY after a commit has succeeded") to name the real precondition — written, not synced. -
Kernel floor and raw-syscall shape — OPEN, for the
lintoragent, not answered from memory. The reference tree is Linux 7.0 (.dev/reference/linux/Makefile); the ring's floor is 5.4 (park.h13, "ops restricted to the TIMEOUT floor"). Questions lintor settles, each with the kernel path it read: (a) the floor ofIORING_OP_FSYNCwithIORING_FSYNC_DATASYNC(include/uapi/linux/io_uring.h257, 340) and whether a bare FSYNC SQE needs anyIOSQE_*flag; (b) whetherpark.c's hand-mirroredio_uring_sqealready covers thefsync_flagsunion member (it mirrorspoll32_eventsat the same offset) or the mirror must grow; (c) how the 5.4-era io-wq executes a forced-async op — a kernel thread per ring, visible where, anyRLIMIT_NPROCor cgroup consequence — versusIORING_FEAT_NATIVE_WORKERSkernels; (d) what happens to an in-flight FSYNC at ring teardown and atclose(fd)(does exit wait for it; does the op hold its own file reference); (e) how to know at boot that the op is supported withoutIORING_REGISTER_PROBE(5.6) — the candidate is a self-test: one real FSYNC on the freshly opened, preallocated WAL at first use, judged by itsres; (f) confirm thatfdatasyncfrom a worker while the issuing threadpwrites the same file has no ordering hazard beyond "a later write may or may not be covered", which fork 2 already assumes. What is settled here regardless of the answers: forcedWO_IO=uringwith the op unavailable refuses at boot (mirrorspark.c187–189, forced and absent is fatal); unforced falls back to the synchronous path with one stderr notice, never silently. -
Go/no-go — OPEN, settled by one measurement, then one developer answer. The measurement (B1, codd-cyril): the
mixleg at default shards on the same build,WO_DATAon tmpfs (wherefdatasyncis free — the board recorded 195 000 vs 2200 ops/s forwmix,00-status.md1010–1014) and onbench/(ext4), three runs each. tmpfs is the ceiling for part B: a barrier that costs shard 0 no wall time. Reading it: if tmpfsmixread.p99lands nearram.sN.mixread's 81 µs and ops/s near its 44 874, the barrier is the whole tail and the bar is set from that figure (within 2× on p99, at least half on ops/s, three ext4 runs,seedands1legs inside 15%). If tmpfsmixread.p99stays in the thousands of microseconds, the tail is not the barrier — the drain's execution phase or scheduling — and part B closes as "measured, not built", its Progress table recording the ceiling. The developer answer, after the number: is a 1–4 ms read p99 behind a write, with the 2–4× run-to-run spread that forced the gate to 300%, acceptable for the north-star app (a porch service that fits RAM)? If yes, the do-nothing option is taken with the measurement on record and the gate tolerance kept. Nothing after B2 starts before both answers exist. -
The inline path while a barrier is in flight: unify it. A statement on shard 0 stages and commits its own barrier (
database/src/db.c56–84, 100–115, 133–137; routed atruntime/src/builtin.c188–189), and themixleg puts one of its four mixers on shard 0 (round-robin placement,vm.c1101), so its writes block the drain today and would still block it under an async drain. Options: (a) keep the inline barrier synchronous — a secondfdatasyncbeside the in-flight one is harmless, but the tail comes back for every shard-0 write andWO_SHARDS=1gets nothing; (b) route a durable-table write on shard 0 through the request path —wo_db_rpcpushes to shard 0's own inbox (inbox_push_to,vm.c59–73; the primary inbox exists at every shard count,vm.c574–577 andmain.c234), the fiber parks onWO_PARK_INBOX, the drain executes it, holds its reply, and the completion releases it — one write path, andWO_SHARDS=1gets batching and the async barrier for free; (c) stage inline and park the fiber with a done-flag of its own — the request path minus the marshal, a third path. Decision: (b). It deletes the asymmetryCODE-LOGIC.mdcalls a standing hazard ("if that ever stops holding, the inline path would make another statement's record durable early and acknowledge it to the wrong writer") and themaybe_compactcopy atdb.c32. Reads and@table(durable: false)writes stay inline —ram.s1.seedis 4 µs per op and two inbox hops would double it. Guard: the hops costdurable.s1.seed.p50us(212 µs, 15% tolerance) at most a few microseconds against a ~200 µs flush; if the gate says otherwise, (a) is the fallback and this fork reopens. The route condition needstable_is_durablevisible tobuiltin.c; the codd doctrine that traps stay byte-identical between the two paths is what makes the unification safe to do. -
Completion delivery on a busy shard 0. The reap loop lives in
wo_io_wait, which runs only when nothing is runnable (park.c330–400); a completed barrier's replies would otherwise wait for shard 0 to go idle. Decision: poll the completion queue, non-blocking, at the two points the inbox is already polled —NEXT_RUNNABLE(vm.c1697) and the reduction-slice check (vm.c1734–1736) — plus the sentinel branch inwo_io_wait. Two acquire-loads of the mmap'd head and tail per slice; the bound is one reduction budget, the same bound the DB RPC already carries ("a busy shard adopts its inbox once per reduction slice",runtime/src/CODE-LOGIC.md, "The transparent DB actor"). Rejected: routing the completion through the wake eventfd — the CQE already is the event. -
Shutdown with a barrier in flight.
wo_engine_stop(vm.c713–770) joins workers and discards queued envelopes; it knows nothing of the WAL, andwo_wal_close(wal.c463–474) closes the descriptor unconditionally. Decision: beforewo_wal_close, shard 0 reaps an outstanding barrier to completion with a blockingio_uring_enter(GETEVENTS), applies the fatal rule to its result, then releases the held replies into inboxes the teardown discards. One wait, no new mechanism; the invariant it buys is that exit 0 means every submitted barrier completed, which the stats line's two counts make checkable. A parked requester woken by STOP re-executeswo_db_rpcand re-sends (park.c341–352,vm.c309–357) — pre-existing behaviour for a request in flight at stop, noted, not part B's to change.
Knob, dependency, mode check, applied to all ten: no new environment
variable (WO_IO is the arc's), no library, no thread of ours, no rollback
path, no second copy of any record; the one behavioural difference between
backends is where shard 0 waits, and the stats line and the boot notice name
it.
The 2026-08-15 forks, for the record (superseded, kept so the trail reads):
(1) drop-in behind wo_wal_commit first, async only if the boundary proved the
bottleneck — part A was the drop-in; the boundary proved to be the tail above.
(2) liburing or raw syscalls — raw, and already landed for the ring in
park.c (arc T4); part B adds one op to it. (3) the batch boundary — the tick
was replaced by queue-drain in the 2026-08-28 brainstorm. (4) fallback and CI —
the startup probe plus WO_IO=uring|epoll, landed with the arc; fork 4 above
says what part B owes it.
Proposed Solution
Sequence: B1 and B2 first, in parallel, both S; a GO on B1 and lintor's answers
on B2 unlock the rest. Then B3 (the park-plane seam, its owner assigned by the
developer since the plane is not codd's), B4 and B5 together (the WAL split and
the drain rewrite are one change to reason about), B6 (the inline unification,
measured against its guard), B7, then B8 and B9. Expected shape: no new file —
wal.c gains the submit-or-queue and the in-flight state, vm.c's drain and
two poll sites change, park.c gains three small exports, builtin.c one route
condition. The contract paragraphs in 04-db-binding.md and the "Group commit"
section of CODE-LOGIC.md say what an ack means under the async barrier — the
same thing, later observed — and the crash battery proves it under both
backends.
History
2026-09-10 — part B re-brainstormed by codd-shoney (autonomy). The board's
"close the 66× gap" target was retired for good and the story re-aimed at the
shard-0 read tail. Ten forks enumerated; eight decided from code, kernel and
PostgreSQL evidence (file:line above); two left open — the kernel floor to the
lintor agent, the go/no-go to one tmpfs-vs-ext4 measurement and one developer
answer. Two findings worth keeping regardless of the outcome: the linux card's
"link WRITE→FSYNC, essential" advice does not apply to this engine's buffer
model (the kernel punts both ops to a worker anyway, and pwrite already keeps
every offset invariant); and the deferred keys-resident drops and re-points were
waiting on the staging buffer, not on durability. One inconsistency to hand to
codd: CODE-LOGIC.md's "Group commit" section says the durability abort "exits
3" where wal.h 304 and this story say 74.