databasev2 3, chain 6. Brainstormed 2026-08-28 after databasev2 4 part A
landed.
Design: compact the log by rewriting it as one record per live row into a
temp file, fsync, rename over the live WAL, fsync the parent dir, reopen.
Recovery is COMPLETELY UNCHANGED — boot still opens one file and replays
it — and the crash criterion ("the same store as if the checkpoint had
never started") is satisfied by rename, not by code we must get right.
Read .dev/reference/postgresql for this. The finding is that PG's design
is UNAVAILABLE to us, which is what makes the simpler option legitimate:
- PG never compacts its WAL; segments before the redo point are recycled
by rename or unlinked. Its records are page deltas, so a compacted redo
log is not a store — hence heap files, a control file, a redo pointer,
a second recovery source and a separate process
- ours are FULL ROW IMAGES (apply_record implements UPDATE as
remove-then-recreate), so a compacted log IS a complete store. That one
difference deletes all of the above from the design
- what IS worth porting is the ordering discipline: publish the new
"recovery starts here" atomically and LAST, so a crash falls back. PG
needs a start-of-checkpoint redo pointer plus an end-of-checkpoint
control file update; we get the same property from one rename, because
we can swap the whole data set atomically and PG cannot
Forks settled:
- no snapshot format — the compacted log is the snapshot, existing grammar,
so no new encoder or decoder and the dump reuses wo_wal_append_insert
- one source, not two
- volume-only trigger, as a ratio against the LAST compaction's measured
output (the denominator is known exactly; estimating the live set would
mean estimating Text) with an absolute floor. NO TIMER — PG's exists to
bound loss from unflushed buffers and we have none; an idle log does not
grow. Copying the mechanism without the reason was the trap
- stop-the-world, with the pause measured against a stated budget rather
than assumed acceptable; alternatives are bought against a number
- compaction may run ONLY where nothing is staged (right after a barrier),
or a staged record lands in a file about to be replaced. Normative
Recorded before it can be found late: compaction invalidates every WAL
offset iteration 2's `resident: keys` stores, so the compactor rebuilds the
offset map as it writes. Nothing breaks today because that storage half is
unimplemented — it would break later, looking like corruption.
Also corrected exploration/postgresql/buffer-and-checkpoint.md, which was
wrong on two counts: PG does NOT update its control file by rename (in-place
full-block write + CRC32C), and its checkpoint sketch assumes writeonce has
segment files, which it does not and deliberately will not.
Grounding measured on master: seed 20000 leaves a 986614-byte log; 20000
updates take it to 2590262 bytes with the SAME live rows, and boot+verify on
that store is 155ms.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
- docs/stories/databasev2/, numbered from 1. Six PENDING database iterations
moved from the language track and renumbered, keeping the old id in
`was_language_iteration:` so a search for "iteration 32" still finds it:
32 -> 3 WAL checkpoint, 23 -> 4 io_uring commit, 33 -> 7 single-file store,
27 -> 8 query grammar, 20 -> 9 cross-program, 21 -> 10 keypair auth.
Done work (9, 9b, 22) stays as v1 history; language 18 left whole
- the problem, read off the engine not guessed: rows are malloc'd slabs with
addresses stable forever, NO eviction/spill/paging anywhere in database/src,
the WAL never checkpoints so boot replays all history, and durability is one
process-global WO_DATA so no table can say it matters more than another.
An allocation failure IS a clean catchable WO_T_OOM — but swap thrash
arrives first and carries no error signal at all, which is the real hazard
- four new iterations:
1 measure the ceiling FIRST (curve not cliff; the three exits; kill -9 at
exhaustion) — every later default should follow from a number
2 `@table(mode: ram | durable | cold)` — the grammar ask. Small surface
(Ast.table_cfg gains a key, the parser already rejects unknown args), big
semantics: `durable` defaults so nothing changes silently, and the
compiler refuses a durable row holding a `ref` into a ram table
5 bounded tables + refuse/evict/back-pressure, shedding BEFORE the OS acts
6 cold tiering — mostly forks, incl. whether the language surfaces the
fault cost and whether @unique on cold is refused outright. A paged
B-tree stays rejected: if tiering needs one, reject tiering
- 39 links repointed, link TEXT renumbered to track-local ids; arc gains one
pointer row replacing the six moved; board + board-views cover three tracks
- linkcheck 0 broken / 0 anchors; no code blocks in any story
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-26 20:52:48 +02:00
Renamed from docs/stories/language-runtime-database/32-wal-checkpoint.md (Browse further)