writeonce/docs/stories/databasev2/03-wal-checkpoint.md
shoney.arickathil 69c34c9a89 docs(spec): WAL checkpoint — compact by rewrite + atomic rename
databasev2 3, chain 6. Brainstormed 2026-08-28 after databasev2 4 part A
landed.

Design: compact the log by rewriting it as one record per live row into a
temp file, fsync, rename over the live WAL, fsync the parent dir, reopen.
Recovery is COMPLETELY UNCHANGED — boot still opens one file and replays
it — and the crash criterion ("the same store as if the checkpoint had
never started") is satisfied by rename, not by code we must get right.

Read .dev/reference/postgresql for this. The finding is that PG's design
is UNAVAILABLE to us, which is what makes the simpler option legitimate:

- PG never compacts its WAL; segments before the redo point are recycled
  by rename or unlinked. Its records are page deltas, so a compacted redo
  log is not a store — hence heap files, a control file, a redo pointer,
  a second recovery source and a separate process
- ours are FULL ROW IMAGES (apply_record implements UPDATE as
  remove-then-recreate), so a compacted log IS a complete store. That one
  difference deletes all of the above from the design
- what IS worth porting is the ordering discipline: publish the new
  "recovery starts here" atomically and LAST, so a crash falls back. PG
  needs a start-of-checkpoint redo pointer plus an end-of-checkpoint
  control file update; we get the same property from one rename, because
  we can swap the whole data set atomically and PG cannot

Forks settled:

- no snapshot format — the compacted log is the snapshot, existing grammar,
  so no new encoder or decoder and the dump reuses wo_wal_append_insert
- one source, not two
- volume-only trigger, as a ratio against the LAST compaction's measured
  output (the denominator is known exactly; estimating the live set would
  mean estimating Text) with an absolute floor. NO TIMER — PG's exists to
  bound loss from unflushed buffers and we have none; an idle log does not
  grow. Copying the mechanism without the reason was the trap
- stop-the-world, with the pause measured against a stated budget rather
  than assumed acceptable; alternatives are bought against a number
- compaction may run ONLY where nothing is staged (right after a barrier),
  or a staged record lands in a file about to be replaced. Normative

Recorded before it can be found late: compaction invalidates every WAL
offset iteration 2's `resident: keys` stores, so the compactor rebuilds the
offset map as it writes. Nothing breaks today because that storage half is
unimplemented — it would break later, looking like corruption.

Also corrected exploration/postgresql/buffer-and-checkpoint.md, which was
wrong on two counts: PG does NOT update its control file by rename (in-place
full-block write + CRC32C), and its checkpoint sketch assumes writeonce has
segment files, which it does not and deliberately will not.

Grounding measured on master: seed 20000 leaves a 986614-byte log; 20000
updates take it to 2590262 bytes with the SAME live rows, and boot+verify on
that store is 155ms.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-28 17:39:38 +02:00

7.1 KiB
Raw Blame History

track iteration was_language_iteration status chain
databasev2 3 32 in-progress 6

databasev2 3 — WAL checkpoint: disk space reclamation and bounded replay

Moved 2026-08-26 from the language track, where this was iteration 32. Part of Story — the database beyond RAM. Content unchanged by the move; its dependencies are restated in that track index.

Format: product/story-iteration-template. Part of Story — one language, one runtime, one database, one binary.

Inserted 2026-08-21 (stage-3 guarantee refinement found the hole): the WAL is append-only FOREVER — no checkpoint, no truncation exists in the engine or anywhere on the roadmap. Disk grows without bound and replay time grows with history, so restart cost rises with every write the program ever made. RAM reclamation already exists (deleted rows free their slot — 04-db-binding.md: "Ids are never reused; slots are"); this iteration is the DISK half. LAST in the concurrency chain: stage 3 → 22 → 31 → 24 → 23 → 32 — it wants 22's measured replay/restart numbers to justify its policy and must compose with 23's group-commit write path.

BRAINSTORMED 2026-08-28. Spec: 2026-08-28-wal-checkpoint-design.md. Read .dev/reference/postgresql for this — and the conclusion was that Postgres' design is unavailable to us, which is what makes the simpler one legitimate.

The design in one sentence: compact the log by rewriting it as one record per live row into a temp file, then rename it over the live WAL. Recovery is completely unchanged — boot still opens one file and replays it — and the crash criterion is satisfied by the filesystem rather than by code we must get right.

Why one file works here and not in Postgres. Postgres never compacts its WAL: its records are page deltas, so a compacted redo log is not a store, and it must keep heap files, a control file, a redo pointer and a second recovery source. Ours are full row images — apply_record implements UPDATE as remove-then-recreate — so a compacted log is a complete store. That one difference deletes the control file, the redo pointer, the cutoff offset and the separate process from the design.

Forks settled: no snapshot format (the compacted log is the snapshot); one source, not two; volume-only trigger as a ratio against the last compaction's own measured output, with an absolute floor — no timer, because Postgres' timer exists to bound loss from unflushed buffers and we have none; stop-the-world, with the pause measured against a stated budget rather than assumed acceptable.

The coupling that would otherwise be found late: compaction moves every record, so it invalidates every WAL offset iteration 2's resident: keys stores. The compactor rebuilds the offset map as it writes. Recorded now because iteration 2's storage half is unimplemented, so nothing breaks today — it would break later, looking like corruption rather than a design gap.

Measured on master 2026-08-28, grounding the whole iteration: seed 20000 leaves a 986 614-byte log; 20 000 updates take it to 2 590 262 bytes with the same live rows (2.6× history for no data), and boot+verify on that store is 155 ms.

Goals

  • Disk space is reclaimed. A checkpoint writes the live store as a snapshot and truncates the WAL behind it; deleted rows and overwritten versions stop occupying disk forever.
  • Replay is bounded. Startup replays snapshot + WAL tail, not the program's whole write history — restart time becomes a function of store size, not store age.
  • Every existing guarantee holds byte-for-byte. Ack-after-durable, replay-whole-or-not-at-all, torn-tail drop, ids never reused — a checkpoint changes where bytes live, never what an ack means. A crash DURING checkpoint recovers from the previous snapshot + full tail: the old WAL is not truncated until the new snapshot is durable.

Acceptance Criteria (draft — the spec refines)

  • Given a store with N rows after many writes and deletes, when a checkpoint completes, then disk usage reflects the live rows (plus the WAL tail), and a restart replays snapshot + tail to the byte-identical store.
  • Given kill -9 at ANY instant during a checkpoint, when the process restarts, then recovery produces the same consistent store as if the checkpoint had never started — no acknowledged write lost, no partial snapshot ever read.
  • Given the iteration-22 restart benchmark re-run after checkpoint lands, when replay time is measured on an aged store, then the bounded-replay improvement is recorded as a before/after delta.
  • Given writes arriving while a checkpoint runs (the DB actor serializes statements; the checkpoint must not stall them beyond the stated budget), when the mixed load completes, then every ack held its durability contract and the tail contains exactly the post-snapshot writes.

Out Of Scope

  • MVCC / multi-version reads — the store is update-in-place RAM; "old versions" exist only as WAL history, which is exactly what truncation reclaims.
  • Incremental/streaming backup, point-in-time recovery — a snapshot is a recovery artifact here, not a backup product.
  • Cross-shard checkpoint coordination — the WAL is owner-shard-only (stage 3's rule); one shard, one checkpoint.
  • Compression, dedup, tiering — measure first (22), add only what a number justifies.

Info

Forks the spec must settle:

  1. Snapshot format — a row-image dump of the live store (simple, O(live rows)) vs a rewritten-compacted WAL (reuses replay machinery, O(live rows) too but stays in one format). Leaning: row-image dump in the WAL's existing record grammar, so replay needs no second decoder.
  2. Trigger policy — size threshold (WAL bytes vs snapshot bytes ratio), boot-time compaction, explicit call, or some mix. Leaning: ratio threshold checked at commit, plus manual trigger for tests; decided against 22's numbers.
  3. Write availability during checkpoint — stop-the-world dump (simplest; the DB actor just runs one long "statement") vs fork-and-dump vs incremental copy. Leaning: measure the stop-the-world pause on the 1M-row store first (22); complexity only if the pause breaks a stated budget.
  4. Composition with 23 — the snapshot's durability barrier rides the same per-shard ring (WRITE+FSYNC chain, then the truncate); ordering vs in-flight group commits must be stated normatively in 04-db-binding.md's WAL section.

Proposed Solution

Brainstorm → spec → plan after 23 lands (the write path it composes with) using 22's aged-store replay numbers as the policy input; extend 04-db-binding.md's WAL section with the snapshot format the way the record grammar is documented today.