writeonce/docs/stories/databasev2/04-io-uring-commit.md
shoney.arickathil 75aedf1216 docs(spec): WAL group commit — databasev2 4 part A
Brainstormed 2026-08-28. The iteration is split: part A batches, part B
(io_uring submission) is deferred until A's measurement says whether the
blocking boundary still dominates.

The story's premise needed correcting first:

- it says "replace fsync-per-commit with io_uring group-commit", but the
  engine commits per STATEMENT — db.c calls wo_wal_commit right after
  every append, all six sites, so each row change is one pwrite + one
  fdatasync
- so two independent wins were being carried as one, and only the second
  needs io_uring. The staging buffer already holds any number of records;
  today it never holds more than one. Part A is mostly deleting calls
- iteration 22's numbers say A is where the payoff is: durable writes
  4460 ops/s, mixwrite 1023 ops/s p99 664us, against 1.28M ops/s reads

Forks settled:

- batch boundary is QUEUE-DRAIN, not the tick this story had recorded: a
  tick adds latency to a lone writer, taxing an idle system to serve a
  busy one. Queue-drain self-tunes and needs no knob
- shard 0 holds each reply envelope instead of sending it, commits once
  when the queue empties, then releases all — so a writer is acked after
  the barrier carrying ITS record, which today is true only because
  every batch has one member
- a failure between "RAM mutated" and "record durable" is a FATAL,
  diagnosed abort. This replaces uneven behaviour that already exists:
  insert rolls back, update and delete do not and say so in a comment
  ("RAM ahead of disk"). Batching would have multiplied that
- consequence stated, not slipped in: WO_T_IO leaves the write path
- no batch cap initially; peak staged bytes is measured so the question
  is settled by a number

One gap disclosed rather than hidden: forcing a real fdatasync failure
needs mount privileges, so the unit test proves the error is DETECTED and
the abort itself stays covered by inspection.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-28 07:48:11 +02:00

10 KiB
Raw Blame History

track iteration was_language_iteration status chain
databasev2 4 23 in-progress 5

databasev2 4 — io_uring group-commit write path

Moved 2026-08-26 from the language track, where this was iteration 23. Part of Story — the database beyond RAM. Content unchanged by the move; its dependencies are restated in that track index.

Format: product/story-iteration-template. Part of Story — one language, one runtime, one database, one binary.

Inserted 2026-08-15. The write-path optimization, and deliberately the LAST database performance iteration: it only earns its complexity once there is a measured fsync-per-commit baseline to beat (iteration 22) and a multithreaded runtime to overlap against (iteration 8). Doing it earlier would optimize a number nobody had measured, against a runtime that couldn't use it.

No spec exists yet. The forks in Info are genuine decisions.

REFINED 2026-08-20: the four forks are SETTLED as their recorded leanings (developer confirmation, no code): (1) drop-in behind wo_wal_commit first, an async variant only if the arc's scheduler proves the blocking boundary is the bottleneck; (2) raw io_uring_setup/io_uring_enter syscalls — libc-only doctrine holds, ring layout documented normatively; (3) the batch boundary is the shard tick (the 8+11 arc's quantum), single-writer fallback batches whatever accumulated; (4) startup auto-probe + an env override so CI proves both paths on one kernel — AMENDED: the override is the arc-wide WO_IO=uring|epoll (the arc's T4 owns the probe and the per-shard ring; WO_WAL_MODE is subsumed). Position — RE-SEQUENCED 2026-08-21: FIFTH in the concurrency chain (32, WAL checkpoint, follows it — added 2026-08-21), stage 3 → 22 → 31 → 24 → 23 → 32 (supersedes the 2026-08-20 old-id ordering "9e → 8+11 → 9f"); the per-shard ring already exists (arc T4 landed 2026-08-20, WO_IO=uring|epoll) — this iteration adds the WAL's WRITE+FSYNC chains to it. AMENDED 2026-08-20 (io_uring-first directive): the WAL's WRITE+FSYNC chains ride the SAME per-shard ring T4 creates for fiber parking — one event loop per shard, readiness ops and durability ops together, exactly the linux reference project's "single event loop" card. One composition note added since iteration 18: a transaction { } already IS a staged batch — under io_uring it becomes exactly one submission, so the two features compose without either knowing the other.

BRAINSTORMED 2026-08-28 — and SPLIT IN TWO. Spec for part A: 2026-08-28-wal-group-commit-design.md.

The premise below needed correcting. This story says "replace fsync-per-commit with io_uring group-commit", but the engine does not commit per commit — it commits per statement: db.c calls wo_wal_commit immediately after every append, at all six sites, so every row change is one pwrite plus one fdatasync. That splits the goal into two independent wins, and only the second needs io_uring:

  • Part A — batching. Let many statements share one barrier. The staging buffer already holds any number of records; today it never holds more than one because the caller commits immediately. Mostly a deletion of calls.
  • Part B — async submission. The shard submits and keeps working instead of blocking in fdatasync. Deferred until A's measurement says whether the blocking boundary is still the bottleneck.

A is where most of the number lives. Iteration 22 measured durable writes at 4460 ops/s and mixed writes at 1023 ops/s (p99 664 µs) against 1.28M ops/s for durable reads — ~290× apart, essentially all of it the per-statement barrier.

Forks settled in the brainstorm: batch boundary is queue-drain (not the tick this story recorded — a tick taxes an idle system to serve a busy one); a failure between "RAM mutated" and "record durable" is a fatal, diagnosed abort, replacing today's uneven rollback where insert undoes itself and update/delete admit in a comment that they leave RAM ahead of disk. That removes WO_T_IO from the write path — a language-visible change, recorded here deliberately.

status: in-progress because the brainstorm is done and the spec is approved; the plan is next. (The readiness axis that would say this precisely lives on the unmerged db-residency-doctrine.)

Goals

  • Replace fsync-per-commit with io_uring group-commit on the WAL write path: batch a tick's committed records into one submission, let the kernel overlap the write and the durability barrier, and acknowledge each writer only after the barrier its record rode has completed — the same ack-after-durable contract, at a fraction of the syscall cost.
  • Overlap durability with work. With the shard-actor runtime (iteration 8) the shard thread submits its batch and keeps executing ready statements while the ring drains, instead of blocking one thread on one fdatasync — the multithreading the throughput number has been waiting for.
  • Keep the durability promise byte-for-byte. Every guarantee iterations 9 and 22 proved — replay-whole-or-not-at-all, torn-tail drop, no acknowledged write ever lost — holds identically; io_uring changes HOW the bytes reach the platter, never WHETHER an ack means durable.

Acceptance Criteria

  • What to achieve?
    • Given the io_uring write path under the iteration-22 crash battery (concurrent writers, kill -9 mid-stream, reboot, replay),
    • when it runs,
    • then every acknowledged write is present after replay and no unacknowledged partial write is ever visible — the exact result the fsync path gives, so durability is provably unchanged.
  • What to achieve?
    • Given the iteration-22 durable write benchmark,
    • when it is run on the fsync-per-commit path and then the io_uring group-commit path on the same machine,
    • then the io_uring path's write throughput is materially higher and its p99 commit latency lower, with the before/after numbers recorded — the payoff, measured, not asserted.
  • What to achieve?
    • Given a kernel without io_uring (old, or restricted by seccomp),
    • when the runtime starts,
    • then it falls back to the pwrite + fdatasync path automatically and correctly — io_uring is an accelerator, never a hard dependency, and a binary that runs everywhere is the whole project's premise.

Out Of Scope

  • io_uring for the network/accept path. This iteration is the WAL write path only; the socket side is the shard-actor runtime's and the network layer's concern.
  • io_uring for reads. RAM is authoritative — reads never touch a descriptor (phase-B doctrine), so there is nothing to accelerate on the read path. This is a write-durability optimization, full stop.
  • Registered buffers / fixed files / SQPOLL tuning beyond what the benchmark shows is worth it. Start with the plain submit/complete model; add ring features only when 22's number says a specific one pays.
  • Replacing the WAL format or the commit contract. The bytes on disk and the meaning of an ack are iteration 9's; this changes the syscall, not the format.

Info

Forks the spec must settle:

1. How much of the ring model, and behind what abstraction? The write path today is pwrite + fdatasync in database/src/wal.c; io_uring adds a submission/completion queue and a durability barrier op (IORING_OP_FSYNC/IORING_FSYNC_DATASYNC or O_DSYNC writes). The fork: wrap it behind the existing wo_wal_commit boundary (drop-in, the engine never learns) or expose an async-commit primitive the shard scheduler drives (faster overlap, but couples the WAL to iteration 8's loop). Leaning: drop-in behind wo_wal_commit first — it is the correctness-preserving step and 22 can measure it standalone — then an async variant only if 8's scheduler shows the blocking boundary is the remaining bottleneck.

2. liburing or raw syscalls? liburing is the ergonomic wrapper but is a new external dependency, against the libc-only doctrine; the raw io_uring_setup/io_uring_enter syscalls are a few hundred lines and keep the doctrine. Leaning: raw syscalls (the doctrine is load-bearing and this is a bounded surface), with the mmap'd ring setup written down in the binding doc the way the WAL format is — normative, versioned.

3. What is the batch boundary? Per-statement commit (today) is the simplest correct thing and the slowest; a group commit needs a boundary — a tick (iteration 8's scheduler quantum), a count, or a short time window. Leaning: the shard tick once iteration 8 lands (a batch is "everything committed this tick"), with a single-writer fallback that batches whatever accumulated between one wo_wal_commit call and the ring draining.

4. How is the fallback chosen and tested? A kernel probe at startup (attempt io_uring_setup, fall back on ENOSYS/EPERM) is the mechanism; the question is how CI proves BOTH paths without two kernels. Leaning: an environment override (WO_WAL_MODE=fsync|uring) so the test matrix runs the crash battery and the benchmark on both on any capable machine, and the auto-probe is what production uses.

Proposed Solution

  • Brainstorm the spec after iterations 8 and 22 exist — this iteration is meaningless without a multithreaded runtime to overlap against and a measured baseline to beat, and its plan's acceptance is literally "22's durable number improved, 22's crash battery still green, fsync fallback still correct".
  • Expected shape: a wo_wal write-mode switch (fsync vs uring), the raw ring setup + submit/complete in database/src/wal.c (or a wal_uring.c beside it), the startup probe + WO_WAL_MODE override, the binding doc's WAL section extended with the ring layout, and iteration 22 re-run on both paths with the delta committed.