Commit graph

3 commits

Author SHA1 Message Date
026919762b docs(plan): WAL group commit — 6 tasks, databasev2 4 part A
Plan for the approved spec. Code-free per the repo convention
(docs/plan/discarded.md:54); the executor writes the code.

- T1 a failed barrier is detected and fatal — one entry point that names
  the operation, errno, WAL path and batch size, then exits. The abort
  path itself stays unexercised and the task says so rather than buying
  coverage with a fault-injection switch
- T2 the barrier moves to the drain point and replies are held; the
  request path stops committing per append. Riskiest task, and its risk
  is one place: the crash legs. Plan says STOP if they fail, do not
  adjust the test
- T3 the inline path takes the same fatal rule but keeps its own barrier,
  with a comment explaining the asymmetry so the next reader does not
  "fix" it. Looks like a no-op; without it the two paths disagree, which
  is the unevenness the spec exists to remove
- T4 prove batches actually form BEFORE measuring the payoff — otherwise
  a win gets attributed to the wrong cause. Also records peak staged
  bytes, settling the no-cap decision with a number
- T5 measure, gate, write it down. If the payoff is absent, say so and
  stop: part B must not start on an unproven premise
- T6 closeout, including the error catalogue — WO_T_IO leaving the write
  path is language-visible and must be written down

Spec corrected while planning: it pointed at durable.s1.seed as the
payoff. Wrong, structurally — worker shards hold no WAL, so a queue only
exists when other shards write, and a serial writer has nothing to batch
with. The real target is durable.sN.mixwrite: 480 ops/s at p99 5888us
against s1's 1023 at p99 664, so adding shards currently makes durable
writing WORSE. That inversion is a better argument for the iteration than
the one the story recorded.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-28 09:06:09 +02:00
75aedf1216 docs(spec): WAL group commit — databasev2 4 part A
Brainstormed 2026-08-28. The iteration is split: part A batches, part B
(io_uring submission) is deferred until A's measurement says whether the
blocking boundary still dominates.

The story's premise needed correcting first:

- it says "replace fsync-per-commit with io_uring group-commit", but the
  engine commits per STATEMENT — db.c calls wo_wal_commit right after
  every append, all six sites, so each row change is one pwrite + one
  fdatasync
- so two independent wins were being carried as one, and only the second
  needs io_uring. The staging buffer already holds any number of records;
  today it never holds more than one. Part A is mostly deleting calls
- iteration 22's numbers say A is where the payoff is: durable writes
  4460 ops/s, mixwrite 1023 ops/s p99 664us, against 1.28M ops/s reads

Forks settled:

- batch boundary is QUEUE-DRAIN, not the tick this story had recorded: a
  tick adds latency to a lone writer, taxing an idle system to serve a
  busy one. Queue-drain self-tunes and needs no knob
- shard 0 holds each reply envelope instead of sending it, commits once
  when the queue empties, then releases all — so a writer is acked after
  the barrier carrying ITS record, which today is true only because
  every batch has one member
- a failure between "RAM mutated" and "record durable" is a FATAL,
  diagnosed abort. This replaces uneven behaviour that already exists:
  insert rolls back, update and delete do not and say so in a comment
  ("RAM ahead of disk"). Batching would have multiplied that
- consequence stated, not slipped in: WO_T_IO leaves the write path
- no batch cap initially; peak staged bytes is measured so the question
  is settled by a number

One gap disclosed rather than hidden: forcing a real fdatasync failure
needs mount privileges, so the unit test proves the error is DETECTED and
the abort itself stays covered by inspection.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-28 07:48:11 +02:00
746dc2b42b docs(databasev2): third track — the database beyond RAM, with per-table storage modes
- docs/stories/databasev2/, numbered from 1. Six PENDING database iterations
  moved from the language track and renumbered, keeping the old id in
  `was_language_iteration:` so a search for "iteration 32" still finds it:
  32 -> 3 WAL checkpoint, 23 -> 4 io_uring commit, 33 -> 7 single-file store,
  27 -> 8 query grammar, 20 -> 9 cross-program, 21 -> 10 keypair auth.
  Done work (9, 9b, 22) stays as v1 history; language 18 left whole
- the problem, read off the engine not guessed: rows are malloc'd slabs with
  addresses stable forever, NO eviction/spill/paging anywhere in database/src,
  the WAL never checkpoints so boot replays all history, and durability is one
  process-global WO_DATA so no table can say it matters more than another.
  An allocation failure IS a clean catchable WO_T_OOM — but swap thrash
  arrives first and carries no error signal at all, which is the real hazard
- four new iterations:
  1 measure the ceiling FIRST (curve not cliff; the three exits; kill -9 at
    exhaustion) — every later default should follow from a number
  2 `@table(mode: ram | durable | cold)` — the grammar ask. Small surface
    (Ast.table_cfg gains a key, the parser already rejects unknown args), big
    semantics: `durable` defaults so nothing changes silently, and the
    compiler refuses a durable row holding a `ref` into a ram table
  5 bounded tables + refuse/evict/back-pressure, shedding BEFORE the OS acts
  6 cold tiering — mostly forks, incl. whether the language surfaces the
    fault cost and whether @unique on cold is refused outright. A paged
    B-tree stays rejected: if tiering needs one, reject tiering
- 39 links repointed, link TEXT renumbered to track-local ids; arc gains one
  pointer row replacing the six moved; board + board-views cover three tracks
- linkcheck 0 broken / 0 anchors; no code blocks in any story

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-26 20:52:48 +02:00
Renamed from docs/stories/language-runtime-database/23-io-uring-commit.md (Browse further)