docs(db2-4b): part B re-brainstormed — the async barrier on the corrected premise, ten forks

- retitled "group commit, and the async barrier"; the 2026-08-15/20/28
  banners compressed into a trail; `readiness: refine`, `review_pending`
  (forks 1–5 and 8–10 decided under autonomy 2026-09-10 by codd-shoney;
  6 and 7 keep readiness at refine)
- what part B is FOR: the read tail on shard 0 while a barrier blocks — and
  only that; mechanism: the drain pwrites as today, then submits ONE bare
  IORING_OP_FSYNC and keeps working; held replies released by the
  completion; the epoll fallback is part A unchanged; ordering with
  compaction and the deferred drops/re-points; the inline durable write on
  shard 0 unified through its own inbox while a barrier is in flight;
  completion delivery on a busy shard 0; shutdown reaps an in-flight
  barrier before wo_wal_close
- fork 6 (kernel floor and raw-syscall shape) goes to lintor; fork 7
  (go/no-go) is settled by one measurement — tmpfs vs ext4 `mixread.p99` —
  then one developer answer; both answers already sit in
  .dev/zack/databasev2-4b.md, the fold into this story is pending
- progress table B1–B9 with sizes and owners (cyril B1/B8, lintor B2, the
  runtime agent B3, codd + pm B9); no new knob, no new dependency

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
(cherry picked from commit 7ceb7b8805da41e67661744746bf2d0b7879cc50)
This commit is contained in:
shoney.arickathil 2026-09-15 01:04:54 +02:00
parent 54b160070f
commit 15b15dbe1c

View file

@ -3,11 +3,12 @@ track: databasev2
iteration: "4" iteration: "4"
was_language_iteration: "23" was_language_iteration: "23"
status: in-progress status: in-progress
readiness: ready readiness: refine
review_pending: "part B forks 1–5 and 8–10 decided under autonomy 2026-09-10 by codd-shoney; forks 6 (lintor) and 7 (cyril's tmpfs ceiling) are open and keep readiness at refine"
chain: 5 chain: 5
--- ---
# databasev2 4 — io_uring group-commit write path # databasev2 4 — group commit, and the async barrier (was: io_uring group-commit write path)
> **Moved 2026-08-26** from the language track, where this was iteration 23. > **Moved 2026-08-26** from the language track, where this was iteration 23.
> Part of [Story — databasev2: the database beyond RAM](00-story.md). Content unchanged by > Part of [Story — databasev2: the database beyond RAM](00-story.md). Content unchanged by
@ -32,78 +33,49 @@ chain: 5
> that may actually matter is the one iteration 2 deferred here — io_uring for > that may actually matter is the one iteration 2 deferred here — io_uring for
> `resident: keys` row reads — not group-commit for writes. > `resident: keys` row reads — not group-commit for writes.
> >
> **No spec exists yet.** ~~The forks in *Info* are genuine decisions.~~ > **REFINED 2026-08-20** (developer confirmation, no code): the four original
> > forks were settled as their leanings — drop-in behind `wo_wal_commit` first;
> **REFINED 2026-08-20: the four forks are SETTLED as their recorded > raw `io_uring_setup`/`io_uring_enter`, libc only; the batch boundary the shard
> leanings** (developer confirmation, no code): (1) drop-in behind > tick; a startup probe plus the arc-wide `WO_IO=uring|epoll` override. Position
> `wo_wal_commit` first, an async variant only if the arc's scheduler > re-sequenced 2026-08-21 to fifth in the concurrency chain; the per-shard ring
> proves the blocking boundary is the bottleneck; (2) raw > already exists (arc T4, `WO_IO=uring|epoll`) and the developer's io_uring-first
> `io_uring_setup`/`io_uring_enter` syscalls — libc-only doctrine holds, > directive of 2026-08-20 says the WAL's durability ops ride the SAME per-shard
> ring layout documented normatively; (3) the batch boundary is the shard > ring the fibers park on — one event loop per shard. Those settlements are
> tick (the 8+11 arc's quantum), single-writer fallback batches whatever > recorded in *Info* below under "the 2026-08-15 forks, for the record"; the
> accumulated; (4) startup auto-probe + an env override so CI proves both > tick boundary was later replaced by queue-drain (part A) and "WRITE+FSYNC
> paths on one kernel — AMENDED: the override is the arc-wide > chains" by a single FSYNC (part B, fork 2).
> `WO_IO=uring|epoll` (the arc's T4 owns the probe and the per-shard
> ring; `WO_WAL_MODE` is subsumed). Position — RE-SEQUENCED 2026-08-21:
> FIFTH in the concurrency chain (32, WAL checkpoint, follows it —
> added 2026-08-21), **stage 3 → 22 → 31 → 24 → 23 → 32**
> (supersedes the 2026-08-20 old-id ordering "9e → 8+11 → 9f"); the
> per-shard ring already exists (arc T4 landed 2026-08-20,
> `WO_IO=uring|epoll`) — this iteration adds the WAL's WRITE+FSYNC
> chains to it. AMENDED 2026-08-20 (io_uring-first
> directive): the WAL's WRITE+FSYNC chains ride the SAME per-shard ring
> T4 creates for fiber parking — one event loop per shard, readiness ops
> and durability ops together, exactly the linux reference project's
> "single event loop" card. One composition
> note added since iteration 18: a `transaction { }` already IS a staged
> batch — under io_uring it becomes exactly one submission, so the two
> features compose without either knowing the other.
> **BRAINSTORMED 2026-08-28 — and SPLIT IN TWO.** Spec for part A: > **BRAINSTORMED 2026-08-28 — and SPLIT IN TWO.** Spec for part A:
> [`2026-08-28-wal-group-commit-design.md`](../../superpowers/specs/2026-08-28-wal-group-commit-design.md) > [`2026-08-28-wal-group-commit-design.md`](../../superpowers/specs/2026-08-28-wal-group-commit-design.md)
> · plan: [`2026-08-28-wal-group-commit.md`](../../superpowers/plans/2026-08-28-wal-group-commit.md) > · plan: [`2026-08-28-wal-group-commit.md`](../../superpowers/plans/2026-08-28-wal-group-commit.md)
> (6 tasks). > (6 tasks).
> >
> **The payoff metric is `durable.sN.mixwrite`, not the s1 numbers.** Worker > **The premise needed correcting.** This story said "replace fsync-per-commit
> shards hold no WAL — the runtime asserts it — so every statement on a worker > with io_uring group-commit", but the engine committed per **statement**: `db.c`
> marshals to shard 0 and parks, while a statement already on shard 0 runs > called `wo_wal_commit` immediately after every append, at all six sites. That
> inline. Batches form only where there is a queue, so concurrent multi-shard > split the goal into two independent wins, and only the second needs io_uring:
> writes batch and a single-shard or serial workload does not. The baseline
> shows why that is the right target anyway: **multi-shard concurrent writes are
> 480 ops/s at p99 5888 µs against single-shard's 1023 at p99 664 — adding
> shards makes durable writing WORSE today**, because every marshaled statement
> still buys its own barrier on the owner.
> >
> **The premise below needed correcting.** This story says "replace > - **Part A — batching.** Let many statements share one barrier. Landed
> fsync-per-commit with io_uring group-commit", but the engine does not commit > 2026-08-28, ≈2.9× concurrent durable write throughput.
> per commit — it commits per **statement**: `db.c` calls `wo_wal_commit` > - **Part B — the async barrier.** The shard submits the barrier and keeps
> immediately after every append, at all six sites, so every row change is one > working instead of blocking in `fdatasync`. **Re-brainstormed 2026-09-10 on
> `pwrite` plus one `fdatasync`. That splits the goal into two independent > the corrected premise below.**
> wins, and only the second needs io_uring:
> >
> - **Part A — batching.** Let many statements share one barrier. The staging > **Forks settled in the 2026-08-28 brainstorm:** batch boundary is
> buffer already holds any number of records; today it never holds more than > **queue-drain** (not the tick — a tick taxes an idle system to serve a busy
> one because the caller commits immediately. Mostly a deletion of calls.
> - **Part B — async submission.** The shard submits and keeps working instead
> of blocking in `fdatasync`. Deferred until A's measurement says whether the
> blocking boundary is still the bottleneck.
>
> **A is where most of the number lives.** Iteration 22 measured durable writes
> at 4460 ops/s and mixed writes at 1023 ops/s (p99 664 µs) against 1.28M ops/s
> for durable reads — ~290× apart, essentially all of it the per-statement
> barrier.
>
> **Forks settled in the brainstorm:** batch boundary is **queue-drain** (not
> the tick this story recorded — a tick taxes an idle system to serve a busy
> one); a failure between "RAM mutated" and "record durable" is a **fatal, > one); a failure between "RAM mutated" and "record durable" is a **fatal,
> diagnosed abort**, replacing today's uneven rollback where `insert` undoes > diagnosed abort**, which **removed `WO_T_IO` from the write path** — a
> itself and `update`/`delete` admit in a comment that they leave RAM ahead of > language-visible change, recorded here deliberately.
> disk. **That removes `WO_T_IO` from the write path** — a language-visible
> change, recorded here deliberately.
> >
> `status: in-progress` because the brainstorm is done and the spec is > **RE-BRAINSTORMED 2026-09-10 (part B, codd-shoney, autonomy).** The board's
> approved; the plan is next. (The `readiness` axis that would say this > original target for part B — "close the 66× gap" — was wrong, and the story's
> precisely lives on the unmerged `db-residency-doctrine`.) > own part-A closeout said so. What part A left behind is a **latency tail on
> shard 0**: everything queued behind a barrier waits for the device, and
> `durable.sN.mixread.p99` measured **1043 / 2318 / 4147 µs** across three
> identical runs. Part B is now FOR that tail, and nothing else. Ten forks,
> eight settled from code evidence, two open with the measurement and the
> kernel question that settle them named — see *Info — the forks, settled*.
> `readiness: refine` until those two return.
## Progress — part A landed 2026-08-28 ## Progress — part A landed 2026-08-28
@ -115,7 +87,23 @@ chain: 5
| 4 | prove batches form — the `wmix` write-concurrent leg | ✅ `40d029c` | | 4 | prove batches form — the `wmix` write-concurrent leg | ✅ `40d029c` |
| 5 | measure the payoff, gate it, record it | ✅ `d52ea8a` | | 5 | measure the payoff, gate it, record it | ✅ `d52ea8a` |
| 6 | closeout | ✅ this change | | 6 | closeout | ✅ this change |
| — | **part B — io_uring submission** | ⬜ **not started; its premise changed, see below** |
## Progress — part B, the async barrier (brainstormed 2026-09-10)
Sizes: S ≤ half a day, M ≤ two days, L longer. B1 and B2 gate everything after
them; nothing below B2 starts until both have reported.
| # | Task | Size | State |
| --- | --- | --- | --- |
| B1 | **the ceiling measurement** (codd-cyril): `mix` at default shards, same build, `WO_DATA` on tmpfs vs on `bench/` (ext4), three runs each; report `mixread` ops/s, p50, p99 — settles fork 7 | S | ⬜ |
| B2 | **the kernel questions** (lintor): floor and shape of a bare `IORING_OP_FSYNC` on the 5.4 ring, io-wq behaviour on the floor kernel, teardown with a barrier in flight, how to detect the op at boot — settles fork 6 | S | ⬜ |
| B3 | the park-plane seam (owner: the runtime agent, not codd): submit one FSYNC SQE with a WAL sentinel, a non-blocking reap, a completion hook; the reap loop matches the sentinel before its fiber cast | S | ⬜ |
| B4 | `wal.c`: split the write from the barrier — the drain's `pwrite` advances `off` as today, the barrier becomes submit-or-queue with an in-flight mark and a synced-bytes counter; boot self-test of the op with the loud fallback notice; stats line names the path | M | ⬜ |
| B5 | `vm.c` drain: pwrite, then submit or queue behind the in-flight barrier; the two held-reply lists move from drain locals into shard state; the completion handler releases ITS list, then runs the compaction check; the hot path polls completions where it polls the inbox | M | ⬜ |
| B6 | the inline unification: a durable-table write on shard 0 takes the request path through its own inbox; reads and volatile writes stay inline; guarded by `durable.s1.seed.p50us` staying inside its 15% tolerance | M | ⬜ |
| B7 | shutdown: shard 0 reaps an in-flight barrier to completion (fatal rule applies) before `wo_wal_close`; never closes the fd with a barrier outstanding | S | ⬜ |
| B8 | gates and numbers (codd-cyril): durable legs and the crash battery under BOTH `WO_IO=uring` and `WO_IO=epoll`; a `test_wal` case that submits FSYNC on a bad fd and sees the failure detected; baseline re-run; `perf-targets.md` §6 addendum with the before/after against the B1 ceiling; re-tighten the 300% tolerance if the tail stabilises | M | ⬜ |
| B9 | contracts and closeout: `CODE-LOGIC.md` "Group commit" section extended with the async barrier, `04-db-binding.md` commit-contract paragraph, story/board/graph (codd + codd-pm) | S | ⬜ |
### The payoff, measured two ways ### The payoff, measured two ways
@ -140,7 +128,7 @@ records per fsync) though less often, so anything queued behind one waits. Three
full runs of the same build gave `durable.sN.mixread.p99` of **1043 / 2318 / full runs of the same build gave `durable.sN.mixread.p99` of **1043 / 2318 /
4147 µs** — a 2–4× spread near idle. So part A buys ~3× write throughput at the 4147 µs** — a 2–4× spread near idle. So part A buys ~3× write throughput at the
cost of a longer, noisier tail on the owner shard. `durable.sN.*.p99us` was cost of a longer, noisier tail on the owner shard. `durable.sN.*.p99us` was
re-baselined at 100% tolerance for that reason, with the floor as the real guard re-baselined at 300% tolerance for that reason, with the floor as the real guard
(`mixread`'s came within 25 µs of tripping). (`mixread`'s came within 25 µs of tripping).
**This is the strongest argument for part B** — submitting the barrier and **This is the strongest argument for part B** — submitting the barrier and
@ -166,29 +154,39 @@ continuing to serve is exactly what removes this cost.
## Goals ## Goals
- **Replace fsync-per-commit with io_uring group-commit** on the WAL write - **One barrier per drain, not per statement** (part A, met). Many statements
path: batch a tick's committed records into one submission, let the kernel share one `pwrite` + `fdatasync`; each writer is acknowledged only after the
overlap the write and the durability barrier, and acknowledge each writer barrier that carried its record. The batch boundary is the queue going
only after the barrier its record rode has completed — the same empty — no tick, no timer, nothing to tune.
ack-after-durable contract, at a fraction of the syscall cost. - **Shard 0 never blocks in the barrier** (part B, rewritten 2026-09-10). The
- **Overlap durability with work.** With the shard-actor runtime drain writes its batch into the page cache and submits the durability barrier
(iteration 8) the shard thread submits its batch and keeps executing ready to the shard's own ring, then goes back to serving. Reads arriving while the
statements while the ring drains, instead of blocking one thread on one device flushes are answered before the flush completes; the next drain's
fdatasync — the multithreading the throughput number has been waiting for. writes are written and queued behind the in-flight barrier. The metric is the
read tail on the owner shard: `durable.sN.mixread.p99us` (baseline 4050 µs;
measured 1043 / 2318 / 4147 µs) and `durable.sN.mixread.ops_sec` (4733),
against the no-barrier reference `ram.sN.mixread` (p99 81 µs, 44 874 ops/s)
and the tmpfs ceiling B1 measures. Write throughput may rise as a side effect
of the disk never idling between barriers (`durable.sN.wmix.ops_sec`, 6017);
it is watched, not targeted.
- **One write path, not two.** A durable-table write on shard 0 takes the same
request path a worker's does, so the inline path's private barrier — the
asymmetry part A documented as a standing hazard — goes away, and
`WO_SHARDS=1` gets batching and the async barrier with it.
- **Keep the durability promise byte-for-byte.** Every guarantee iterations 9 - **Keep the durability promise byte-for-byte.** Every guarantee iterations 9
and 22 proved — replay-whole-or-not-at-all, torn-tail drop, no and 22 proved — replay-whole-or-not-at-all, torn-tail drop, no acknowledged
acknowledged write ever lost — holds identically; io_uring changes HOW the write ever lost, durable-or-process-death — holds identically. Part B changes
bytes reach the platter, never WHETHER an ack means durable. WHEN shard 0 learns the barrier finished, never WHETHER an ack means durable.
## Acceptance Criteria ## Acceptance Criteria
Met: Met:
- **Given** the io_uring write path under iteration 22's crash battery, **when** - **Given** the batched write path under iteration 22's crash battery, **when**
it runs, **then** every acknowledged write is present after replay. ✅ — the it runs, **then** every acknowledged write is present after replay. ✅ —
criterion applies unchanged to part A's batching. `crash.sN` (the batched `crash.sN` (the batched path) recovered every acked row after `kill -9`,
path) recovered every acked row after `kill -9`, `crash.s1` likewise, and both `crash.s1` likewise, and both restart legs replay byte-true. This was the one
restart legs replay byte-true. This was the one thing batching could break. thing batching could break.
- **Given** the durable write benchmark before and after, **then** throughput is - **Given** the durable write benchmark before and after, **then** throughput is
materially higher and p99 lower, recorded. ✅ ~2.9× and ~2.1× (p50); see materially higher and p99 lower, recorded. ✅ ~2.9× and ~2.1× (p50); see
`perf-targets.md` §6. **Scoped honestly:** on a write-concurrent workload `perf-targets.md` §6. **Scoped honestly:** on a write-concurrent workload
@ -200,32 +198,48 @@ Met:
not continue with RAM ahead of disk. ✅ fatal, diagnosed, exit 74 — replacing not continue with RAM ahead of disk. ✅ fatal, diagnosed, exit 74 — replacing
three behaviours that disagreed. three behaviours that disagreed.
Outstanding: Outstanding (part B; rewritten 2026-09-10):
- **Given** a kernel without io_uring, **when** the runtime starts, **then** it - **Given** the ceiling measurement B1, **when** `mix` runs on tmpfs (a free
falls back automatically. *(part B — part A adds no syscall interface, so barrier) against ext4 on the same build, **then** the gap between the two
nothing to fall back from yet.)* `mixread` figures is recorded and is the go/no-go for everything below. If
- **Single-shard concurrent batching.** A statement on shard 0 commits inline tmpfs `mixread.p99` is not materially below the ext4 figure, the tail is not
and cannot batch; doing so needs the inline path to park its fiber on the the barrier and part B closes as "measured, not built" (fork 7).
barrier — the same machinery part B needs. So `WO_SHARDS=1` gets no batching - **Given** a barrier in flight on shard 0, **when** worker reads arrive,
at all, by design and measured (mean batch 1.0). **then** they are answered before it completes: `durable.sN.mixread.p99us`
- **The abort path is not exercised.** Forcing a real `fdatasync` failure needs a within 2× of the B1 tmpfs figure across three runs, and
full or read-only filesystem, which the gate cannot arrange without mount `durable.sN.mixread.ops_sec` at least half of it; `durable.*.seed` and every
privileges. The unit test proves the error is *detected*; the exit three lines `durable.s1.*` leg inside their 15% tolerance.
later is covered by inspection. Disclosed rather than papered over — iteration - **Given** the async path, **when** a writer's reply is released, **then** the
40 was exactly a fatal path nothing exercised. barrier that carried its record has completed: the crash battery recovers
every acked row after `kill -9` under `WO_IO=uring`, and again under
`WO_IO=epoll`.
- **Given** `WO_IO=epoll`, or a ring that cannot run the op, **when** the
runtime starts, **then** the drain commits synchronously exactly as part A
does, one stderr line says so when the fall-back was not forced, and
`WO_WAL_STATS=1` names the path a run took.
- **Given** `WO_SHARDS=1` with concurrent writers, **when** `wmix` runs,
**then** batches form (`durable.s1.wmix.mean_batch` above 1.0, today exactly
1.0) — the inline unification's proof.
- **Given** a completion that reports failure, **when** the handler sees it,
**then** the process ends with the same diagnostic and exit 74 that part A
gives a failed `fdatasync`. The unit test proves the failure is *detected*
(an FSYNC submitted on a bad descriptor completes with an error); the exit
stays covered by inspection, disclosed exactly as part A disclosed it.
- **Given** the process exits with a barrier in flight, **when** shard 0 tears
down, **then** it waits for the completion first; the stats line's
submitted and completed barrier counts agree at exit.
## Part B — its premise changed ## Part B — its premise changed (2026-08-28), and what replaced it (2026-09-10)
Part B was justified by "close the 66× durable gap". Part A shows that framing Part B was justified by "close the 66× durable gap". Part A showed that framing
was wrong: the gap is **two** problems. Concurrent write fan-in was a batching was wrong: the gap is **two** problems. Concurrent write fan-in was a batching
problem and is now ~3× better. What remains is a **serial** writer waiting on a problem and is now ~3× better. What remains for a **serial** writer is a device
single barrier, which no amount of batching can help — and io_uring does not flush it must wait for before its ack — no submission model shortens that, and
obviously help it either, since one writer still needs one durable barrier acknowledging before the flush is the one thing principle 7 forbids. The tail
before its ack. Part B's real candidates are overlapping the barrier with other that IS addressable is shard 0's: while it sits in `fdatasync`, every read and
work on the shard, and the inline-path park that single-shard batching also every write queued in its inbox waits for the device. The forks below settle
needs. **It should be re-brainstormed against that, not started on the old part B against that, and only that.
premise.**
## Out Of Scope ## Out Of Scope
@ -233,65 +247,277 @@ premise.**
path only; the socket side is the shard-actor runtime's and the network path only; the socket side is the shard-actor runtime's and the network
layer's concern. layer's concern.
- **io_uring for reads.** For a fully-resident table reads never touch a - **io_uring for reads.** For a fully-resident table reads never touch a
descriptor, so there is nothing to accelerate on the read path. This is a descriptor. A table declaring `resident: keys` ([databasev2 2](02-table-storage-modes.md))
write-durability optimization, full stop. **Note (2026-08-26):** principle 7's *does* `pread` rows from the log, and accelerating that is a real, separate
residency half was amended, so a table declaring `resident: keys` question — not this iteration's, whose acceptance is a durability tail.
([databasev2 2](02-table-storage-modes.md)) *does* `pread` rows from the log — - **Async commit** — acknowledging a writer before its barrier completes,
and accelerating that read path with io_uring becomes a real, separate PostgreSQL's `synchronous_commit = off` / `XLogSetAsyncXactLSN`. Refused:
question. It is not this iteration's, and it should not be folded in: this principle 7, "an ack means the commit reached disk … none of it is
one is about the commit path and its acceptance is a durability number. negotiable". It is the only thing that would help a serial writer's latency,
- **Registered buffers / fixed files / SQPOLL tuning** beyond what the and it is not on offer.
benchmark shows is worth it. Start with the plain submit/complete model; - **A linked WRITE→FSYNC SQE chain, registered buffers, fixed files, SQPOLL,
add ring features only when 22's number says a specific one pays. `sync_file_range`.** Fork 2 explains why the chain buys nothing here; the
rest are ring features to add only when a measurement names one.
- **Replacing the WAL format or the commit contract.** The bytes on disk and - **Replacing the WAL format or the commit contract.** The bytes on disk and
the meaning of an ack are iteration 9's; this changes the syscall, not the the meaning of an ack are iteration 9's; this changes when shard 0 learns the
format. barrier finished, not what the log contains.
- **`transaction { }`** (language 18, hold). A transaction is a staged batch;
under part B it is one drain's write and one barrier. The two compose without
either knowing the other.
## Info ## Info — the forks, settled
Forks the spec must settle: **Part B — ten forks, 2026-09-10** (codd-shoney, under autonomy; the
frontmatter's `review_pending` asks for the developer's second look at the eight
decided ones, and names the two that keep readiness at `refine`).
**1. How much of the ring model, and behind what abstraction?** The write 1. **What part B is FOR: the read tail on shard 0 — and only that.** Three
path today is `pwrite` + `fdatasync` in `database/src/wal.c`; io_uring adds a candidates were on the table. (a) The tail: reads and writes queued in shard
submission/completion queue and a durability barrier op 0's inbox wait while the drain blocks in `fdatasync` (`runtime/src/vm.c`
(`IORING_OP_FSYNC`/`IORING_FSYNC_DATASYNC` or `O_DSYNC` writes). The fork: 233–241 calls `wo_wal_commit_fatal`, which is `pwrite` + `fdatasync` at
wrap it behind the existing `wo_wal_commit` boundary (drop-in, the engine `database/src/wal.c` 965–981); measured `durable.sN.mixread.p99us` 1043 /
never learns) or expose an async-commit primitive the shard scheduler drives 2318 / 4147 µs, baseline 4050, where the same read path with no barrier in
(faster overlap, but couples the WAL to iteration 8's loop). Leaning: the way (`ram.sN.mixread`) sits at 81 µs and 44 874 ops/s against 4733. (b)
drop-in behind `wo_wal_commit` first — it is the correctness-preserving Serial-writer latency (`durable.s1.seed` p50 212 µs): refused — one writer
step and 22 can measure it standalone — then an async variant only if 8's needs one completed device flush before its ack, and the only mechanism
scheduler shows the blocking boundary is the remaining bottleneck. that shortens the wait is acknowledging early, which principle 7 forbids;
io_uring moves the wait off the thread, it does not shorten it. (c)
Throughput: part A already took the batching win; pipelining (the next
drain writes while the previous barrier flushes) may raise
`durable.sN.wmix.ops_sec` from 6017 because the disk stops idling during
the drain's execution phase — recorded as a watched side effect, not the
goal, because a goal needs one number and (a) has it. **Decision: (a).** The
metric is `durable.sN.mixread.p99us` and `.ops_sec`; the bar is fork 7's.
Evidence that the tail is shard 0's and not the requester's: reads are
already released before the barrier (`vm.c` 207–214), so a read waits only
when it ARRIVES during one — which is exactly what an async barrier ends.
**2. liburing or raw syscalls?** liburing is the ergonomic wrapper but is a 2. **Mechanism: the drain `pwrite`s as today, then submits ONE bare
new external dependency, against the libc-only doctrine; the raw `IORING_OP_FSYNC` (`IORING_FSYNC_DATASYNC`) on shard 0's existing ring; at
`io_uring_setup`/`io_uring_enter` syscalls are a few hundred lines and keep most one barrier in flight; a drain that finds one in flight writes its bytes
the doctrine. Leaning: raw syscalls (the doctrine is load-bearing and this is and queues behind it; the completion submits the queued one.** Three options
a bounded surface), with the mmap'd ring setup written down in the binding were weighed. (i) A linked `IORING_OP_WRITE` → `IORING_OP_FSYNC` chain — the
doc the way the WAL format is — normative, versioned. 2026-08-20 wording and the linux card's advice
(`docs/plan/exploration/linux/07-io_uring.md`, "essential for WAL"). It needs
the staging buffer to stay stable until the WRITE completes (a second
buffer, and every reader of `[off, off+len)` — `scan_record_staged` at
`wal.c` 311–330, `wo_wal_next_offset` at `wal.h` 222–242 — would have to
learn a second not-yet-visible region), two CQEs to check with a
short-write rule (`io_uring/rw.c` 550–560 fails the request on a short
write; `io_uring/io_uring.c` 1841–1848 cancels the link), and it buys no
asynchrony the plain write lacks: a buffered write to a regular file is
punted to an io-wq worker anyway unless the filesystem advertises
`FOP_BUFFER_WASYNC` (`rw.c` 1146–1155). The card was written for a Rust-era
buffer model that no longer exists; the "same ring" half of the directive
stands, the "chain" half does not. (ii) A dedicated fsync fiber or thread of
ours — a thread plus a second wake mechanism, against codd's "no locks,
single-threaded by contract"; and redundant, because the kernel already runs
the op on a worker thread it owns: `io_fsync_prep` sets `REQ_F_FORCE_ASYNC`
and `io_fsync` warns if ever called non-blocking (`io_uring/sync.c` 69, 79).
(iii) Moving the barrier off shard 0 entirely — there is no other owner of
the WAL (principle 5; workers assert they hold neither `db` nor `wal`,
`vm.c` 2362). **Decision: the bare FSYNC.** `pwrite` returning means the
bytes are in the page cache, so every existing offset invariant holds
unchanged — `off` advances at the write exactly as it does at `wal.c` 978,
the staged region `[off, off+len)` exists only inside a drain, `pread` of any
written offset returns data — and the WAL stays the only truth. What is
new is bookkeeping about the file: an in-flight mark and a synced-bytes
counter, both used only by compaction (fork 5) and the stats line. **Cost:**
~30 lines in `park.c` (a submit with a WAL sentinel — `uring_submit` is
static at `park.c` 136; a non-blocking reap; a completion hook; and the reap
loop's `else` branch at `park.c` 392–397 casts any unknown `user_data` to a
fiber pointer, so the sentinel MUST be matched before it) — that seam is the
park plane's, not codd's, and its owner is named in B3. PostgreSQL's shape,
ported as behaviour: a backend that finds the flush lock held waits for the
holder and re-checks whether its record was flushed for it (`xlog.c`
2875–2889) — our "queue behind the in-flight barrier" is that, with the drain
as the only backend. PG's `commit_delay` (`xlog.c` 2902–2906) is a timer
knob and stays rejected for the reason part A rejected the tick.
**3. What is the batch boundary?** Per-statement commit (today) is the 3. **The ack contract across an async completion.** Today the held replies are
simplest correct thing and the slowest; a group commit needs a boundary — a drain locals — "nothing here needs to outlive the batch it describes"
tick (iteration 8's scheduler quantum), a count, or a short time window. (`vm.c` 98–101). Under part B they must: **two FIFO lists in shard state,
Leaning: the shard tick once iteration 8 lands (a batch is "everything in-flight and next; each list is released by ITS barrier's completion, never
committed this tick"), with a single-writer fallback that batches whatever earlier.** A completion with `res` below zero takes `wal_die` with
accumulated between one `wo_wal_commit` call and the ring draining. "fdatasync" and `-res` as the errno — the same text and exit 74
(`WO_EXIT_DURABILITY`, `wal.h` 304) part A gives a failed `fdatasync`; a
`pwrite` failure stays synchronous at the drain and takes the existing
"pwrite" arm (`wal.c` 1016). "Durable or process death" is unchanged. What
widens is the WINDOW between RAM-mutated and durable, and it is the same
exposure part A already accepted: a read arriving after a write in the same
drain sees the write's effect and is answered before the barrier (`vm.c`
207–214), so a process death loses nothing acknowledged and may have shown
a reader an unacknowledged state — today's behaviour, longer. Rejected: a
per-batch sequence number stamped on each reply (a third piece of state to
keep consistent; FIFO-in-order completion makes two lists sufficient).
**4. How is the fallback chosen and tested?** A kernel probe at startup 4. **The epoll fallback is part A, unchanged: synchronous `pwrite` +
(attempt `io_uring_setup`, fall back on ENOSYS/EPERM) is the mechanism; the `fdatasync` in the drain.** No helper thread — a thread of ours doing
question is how CI proves BOTH paths without two kernels. Leaning: an `fdatasync` and signalling an eventfd would be a second mechanism to keep
environment override (`WO_WAL_MODE=fsync|uring`) so the test matrix runs the correct for the portability path, and principle 9 says the portability path
crash battery and the benchmark on both on any capable machine, and the is not where sharpness goes. The contract is identical on both paths; only
auto-probe is what production uses. where shard 0 waits differs. **Two consequences, both gated:** (a)
`scripts/db-bench.py` never sets `WO_IO` today — its durable legs and the
crash battery (lines 257–274) run whatever the probe picks — so cyril adds
`WO_IO=uring` and `WO_IO=epoll` legs to the restart proof and the crash
battery, the pattern `scripts/db-actor-accept.sh` 49–50 already uses; (b)
the `WO_WAL_STATS=1` line (`runtime/src/main.c` 175–179) gains the barrier
path and the submitted/completed counts, so no measurement can be
misattributed. The baseline keeps the probe's default (uring on this box).
5. **Ordering with compaction, and the deferred drops and re-points.**
Compaction renames a new file over the log and re-opens the descriptor
(`wal.c` 1317–1327); with an FSYNC in flight on the OLD descriptor, its
completion would certify an unlinked inode. **Decision: compaction runs only
when staging is empty AND no barrier is in flight** — the refusal at `wal.c`
1214 gains the second condition as its backstop, and the drain's compaction
check (`vm.c` 250–267) moves into the completion handler, after that batch's
replies are released, keeping today's order. Compaction itself stays
synchronous; it is rare and measured at 2.7 ms. The keys-resident drops and
re-points were deferred to "after the commit" because "the record is still
in the staging buffer, so its offset would pread zeros" (`wal.h` 96–98,
`wo_db_flush_drops` at `wal.c` 451–461); the reason is the staging buffer,
not durability, and the drain's `pwrite` ends it. **Decision: flush drops
and re-points right after the drain's write**, keeping one pending list and
no per-batch split; a later barrier failure kills the process, so nothing
observable depends on the difference. zack rewords the `wal.h` 339 comment
("Call ONLY after a commit has succeeded") to name the real precondition —
written, not synced.
6. **Kernel floor and raw-syscall shape — OPEN, for the `lintor` agent, not
answered from memory.** The reference tree is Linux 7.0
(`.dev/reference/linux/Makefile`); the ring's floor is 5.4 (`park.h` 13,
"ops restricted to the TIMEOUT floor"). Questions lintor settles, each with
the kernel path it read: (a) the floor of `IORING_OP_FSYNC` with
`IORING_FSYNC_DATASYNC` (`include/uapi/linux/io_uring.h` 257, 340) and
whether a bare FSYNC SQE needs any `IOSQE_*` flag; (b) whether `park.c`'s
hand-mirrored `io_uring_sqe` already covers the `fsync_flags` union member
(it mirrors `poll32_events` at the same offset) or the mirror must grow;
(c) how the 5.4-era io-wq executes a forced-async op — a kernel thread per
ring, visible where, any `RLIMIT_NPROC` or cgroup consequence — versus
`IORING_FEAT_NATIVE_WORKERS` kernels; (d) what happens to an in-flight FSYNC
at ring teardown and at `close(fd)` (does exit wait for it; does the op
hold its own file reference); (e) how to know at boot that the op is
supported without `IORING_REGISTER_PROBE` (5.6) — the candidate is a
self-test: one real FSYNC on the freshly opened, preallocated WAL at first
use, judged by its `res`; (f) confirm that `fdatasync` from a worker while
the issuing thread `pwrite`s the same file has no ordering hazard beyond
"a later write may or may not be covered", which fork 2 already assumes.
What is settled here regardless of the answers: forced `WO_IO=uring` with
the op unavailable refuses at boot (mirrors `park.c` 187–189, forced and
absent is fatal); unforced falls back to the synchronous path with one
stderr notice, never silently.
7. **Go/no-go — OPEN, settled by one measurement, then one developer answer.**
The measurement (B1, codd-cyril): the `mix` leg at default shards on the
same build, `WO_DATA` on tmpfs (where `fdatasync` is free — the board
recorded 195 000 vs 2200 ops/s for `wmix`, `00-status.md` 1010–1014) and on
`bench/` (ext4), three runs each. tmpfs is the ceiling for part B: a barrier
that costs shard 0 no wall time. Reading it: if tmpfs `mixread.p99` lands
near `ram.sN.mixread`'s 81 µs and ops/s near its 44 874, the barrier is the
whole tail and the bar is set from that figure (within 2× on p99, at least
half on ops/s, three ext4 runs, `seed` and `s1` legs inside 15%). If tmpfs
`mixread.p99` stays in the thousands of microseconds, the tail is not the
barrier — the drain's execution phase or scheduling — and part B closes as
"measured, not built", its Progress table recording the ceiling. The
developer answer, after the number: is a 1–4 ms read p99 behind a write, with
the 2–4× run-to-run spread that forced the gate to 300%, acceptable for the
north-star app (a porch service that fits RAM)? If yes, the do-nothing
option is taken with the measurement on record and the gate tolerance kept.
Nothing after B2 starts before both answers exist.
8. **The inline path while a barrier is in flight: unify it.** A statement on
shard 0 stages and commits its own barrier (`database/src/db.c` 56–84,
100–115, 133–137; routed at `runtime/src/builtin.c` 188–189), and the
`mix` leg puts one of its four mixers on shard 0 (round-robin placement,
`vm.c` 1101), so its writes block the drain today and would still block it
under an async drain. Options: (a) keep the inline barrier synchronous — a
second `fdatasync` beside the in-flight one is harmless, but the tail comes
back for every shard-0 write and `WO_SHARDS=1` gets nothing; (b) route a
durable-table write on shard 0 through the request path — `wo_db_rpc` pushes
to shard 0's own inbox (`inbox_push_to`, `vm.c` 59–73; the primary inbox
exists at every shard count, `vm.c` 574–577 and `main.c` 234), the fiber
parks on `WO_PARK_INBOX`, the drain executes it, holds its reply, and the
completion releases it — one write path, and `WO_SHARDS=1` gets batching
and the async barrier for free; (c) stage inline and park the fiber with a
done-flag of its own — the request path minus the marshal, a third path.
**Decision: (b).** It deletes the asymmetry `CODE-LOGIC.md` calls a standing
hazard ("if that ever stops holding, the inline path would make another
statement's record durable early and acknowledge it to the wrong writer")
and the `maybe_compact` copy at `db.c` 32. Reads and `@table(durable:
false)` writes stay inline — `ram.s1.seed` is 4 µs per op and two inbox
hops would double it. **Guard:** the hops cost `durable.s1.seed.p50us`
(212 µs, 15% tolerance) at most a few microseconds against a ~200 µs flush;
if the gate says otherwise, (a) is the fallback and this fork reopens. The
route condition needs `table_is_durable` visible to `builtin.c`; the codd
doctrine that traps stay byte-identical between the two paths is what makes
the unification safe to do.
9. **Completion delivery on a busy shard 0.** The reap loop lives in
`wo_io_wait`, which runs only when nothing is runnable (`park.c` 330–400);
a completed barrier's replies would otherwise wait for shard 0 to go idle.
**Decision: poll the completion queue, non-blocking, at the two points the
inbox is already polled** — `NEXT_RUNNABLE` (`vm.c` 1697) and the
reduction-slice check (`vm.c` 1734–1736) — plus the sentinel branch in
`wo_io_wait`. Two acquire-loads of the mmap'd head and tail per slice; the
bound is one reduction budget, the same bound the DB RPC already carries
("a busy shard adopts its inbox once per reduction slice",
`runtime/src/CODE-LOGIC.md`, "The transparent DB actor"). Rejected: routing
the completion through the wake eventfd — the CQE already is the event.
10. **Shutdown with a barrier in flight.** `wo_engine_stop` (`vm.c` 713–770)
joins workers and discards queued envelopes; it knows nothing of the WAL,
and `wo_wal_close` (`wal.c` 463–474) closes the descriptor unconditionally.
**Decision: before `wo_wal_close`, shard 0 reaps an outstanding barrier to
completion with a blocking `io_uring_enter(GETEVENTS)`, applies the fatal
rule to its result, then releases the held replies into inboxes the
teardown discards.** One wait, no new mechanism; the invariant it buys is
that exit 0 means every submitted barrier completed, which the stats
line's two counts make checkable. A parked requester woken by STOP
re-executes `wo_db_rpc` and re-sends (`park.c` 341–352, `vm.c` 309–357) —
pre-existing behaviour for a request in flight at stop, noted, not part
B's to change.
**Knob, dependency, mode check, applied to all ten:** no new environment
variable (`WO_IO` is the arc's), no library, no thread of ours, no rollback
path, no second copy of any record; the one behavioural difference between
backends is where shard 0 waits, and the stats line and the boot notice name
it.
**The 2026-08-15 forks, for the record** (superseded, kept so the trail reads):
(1) drop-in behind `wo_wal_commit` first, async only if the boundary proved the
bottleneck — part A was the drop-in; the boundary proved to be the tail above.
(2) liburing or raw syscalls — raw, and already landed for the ring in
`park.c` (arc T4); part B adds one op to it. (3) the batch boundary — the tick
was replaced by queue-drain in the 2026-08-28 brainstorm. (4) fallback and CI —
the startup probe plus `WO_IO=uring|epoll`, landed with the arc; fork 4 above
says what part B owes it.
## Proposed Solution ## Proposed Solution
- **Brainstorm the spec** after iterations 8 and 22 exist — this iteration is Sequence: B1 and B2 first, in parallel, both S; a GO on B1 and lintor's answers
meaningless without a multithreaded runtime to overlap against and a on B2 unlock the rest. Then B3 (the park-plane seam, its owner assigned by the
measured baseline to beat, and its plan's acceptance is literally "22's developer since the plane is not codd's), B4 and B5 together (the WAL split and
durable number improved, 22's crash battery still green, fsync fallback the drain rewrite are one change to reason about), B6 (the inline unification,
still correct". measured against its guard), B7, then B8 and B9. Expected shape: no new file —
- Expected shape: a `wo_wal` write-mode switch (fsync vs uring), the raw ring `wal.c` gains the submit-or-queue and the in-flight state, `vm.c`'s drain and
setup + submit/complete in `database/src/wal.c` (or a `wal_uring.c` two poll sites change, `park.c` gains three small exports, `builtin.c` one route
beside it), the startup probe + `WO_WAL_MODE` override, the binding doc's condition. The contract paragraphs in `04-db-binding.md` and the "Group commit"
WAL section extended with the ring layout, and iteration 22 re-run on both section of `CODE-LOGIC.md` say what an ack means under the async barrier — the
paths with the delta committed. same thing, later observed — and the crash battery proves it under both
backends.
## History
**2026-09-10 — part B re-brainstormed by codd-shoney (autonomy).** The board's
"close the 66× gap" target was retired for good and the story re-aimed at the
shard-0 read tail. Ten forks enumerated; eight decided from code, kernel and
PostgreSQL evidence (file:line above); two left open — the kernel floor to the
`lintor` agent, the go/no-go to one tmpfs-vs-ext4 measurement and one developer
answer. Two findings worth keeping regardless of the outcome: the linux card's
"link WRITE→FSYNC, essential" advice does not apply to this engine's buffer
model (the kernel punts both ops to a worker anyway, and `pwrite` already keeps
every offset invariant); and the deferred keys-resident drops and re-points were
waiting on the staging buffer, not on durability. One inconsistency to hand to
codd: `CODE-LOGIC.md`'s "Group commit" section says the durability abort "exits
3" where `wal.h` 304 and this story say 74.