Commit graph

3 commits

Author SHA1 Message Date
74399ffc68 docs(plan): WAL checkpoint — 6 tasks, databasev2 3
Plan for the approved spec. Code-free per the repo convention
(docs/plan/discarded.md:54); the executor writes the code.

- T1 wo_wal_compact: walk live rows via the bitmap, append one INSERT
  each through the EXISTING append path, fsync, rename over the live log,
  fsync the parent dir, reopen the descriptor. Test asserts BOTH that the
  log shrank AND that a replay reproduces the same rows/ids/values —
  shorter alone is worthless, a truncating bug also passes that
- T2 a stale temp file is removed at open and never read. The test uses
  PLAUSIBLE records, not garbage: garbage would be rejected anyway and
  would prove nothing
- T3 the trigger as a PURE decision (used bytes, last compaction's
  measured output, floor) so it is unit-testable without a store; env
  knobs for floor and ratio, which is what makes the policy testable at
  all. No timer, with the reason. The check is called only where nothing
  is staged, asserted by a test that stages and expects deferral
- T4 kill -9 DURING compaction, extending the existing fork-based crash
  battery. Asserts the PROPERTY — the store equals the pre- or the
  post-compaction content, never a mixture, and every acked id survives.
  Run repeatedly and state the count: it is a race, one green run proves
  little
- T5 measure space reclaimed, boot before/after, and the stop-the-world
  PAUSE against a stated budget. If the pause exceeds it, stop and report
  — the alternatives are bought against that number, not before it
- T6 closeout, including the normative ordering rule in 04-db-binding.md

Constraints carried from the spec into every task:

- recovery must NOT change; a task editing the replay path should stop
- the dump must FLUSH PERIODICALLY. stage() grows the staging buffer by
  doubling, so dumping a whole store through one buffer would hold the
  entire store in RAM — the unbounded growth databasev2 1 identified as
  how this engine dies
- a FAILED compaction is a missed optimisation, not a durability event,
  so it must not take databasev2 4's fatal path
- gate tolerances must not be waived wholesale (part A's T4 made that
  mistake), and the baseline is full-mode — writing a quick-mode baseline
  over it is a regression part A also made

Deliberately NOT a task: rebuilding the `resident: keys` offset map. It
cannot be implemented against a feature that does not exist yet, so T6
records it as an obligation at the compactor and in the story instead of
a stub nobody can test.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-28 17:44:18 +02:00
69c34c9a89 docs(spec): WAL checkpoint — compact by rewrite + atomic rename
databasev2 3, chain 6. Brainstormed 2026-08-28 after databasev2 4 part A
landed.

Design: compact the log by rewriting it as one record per live row into a
temp file, fsync, rename over the live WAL, fsync the parent dir, reopen.
Recovery is COMPLETELY UNCHANGED — boot still opens one file and replays
it — and the crash criterion ("the same store as if the checkpoint had
never started") is satisfied by rename, not by code we must get right.

Read .dev/reference/postgresql for this. The finding is that PG's design
is UNAVAILABLE to us, which is what makes the simpler option legitimate:

- PG never compacts its WAL; segments before the redo point are recycled
  by rename or unlinked. Its records are page deltas, so a compacted redo
  log is not a store — hence heap files, a control file, a redo pointer,
  a second recovery source and a separate process
- ours are FULL ROW IMAGES (apply_record implements UPDATE as
  remove-then-recreate), so a compacted log IS a complete store. That one
  difference deletes all of the above from the design
- what IS worth porting is the ordering discipline: publish the new
  "recovery starts here" atomically and LAST, so a crash falls back. PG
  needs a start-of-checkpoint redo pointer plus an end-of-checkpoint
  control file update; we get the same property from one rename, because
  we can swap the whole data set atomically and PG cannot

Forks settled:

- no snapshot format — the compacted log is the snapshot, existing grammar,
  so no new encoder or decoder and the dump reuses wo_wal_append_insert
- one source, not two
- volume-only trigger, as a ratio against the LAST compaction's measured
  output (the denominator is known exactly; estimating the live set would
  mean estimating Text) with an absolute floor. NO TIMER — PG's exists to
  bound loss from unflushed buffers and we have none; an idle log does not
  grow. Copying the mechanism without the reason was the trap
- stop-the-world, with the pause measured against a stated budget rather
  than assumed acceptable; alternatives are bought against a number
- compaction may run ONLY where nothing is staged (right after a barrier),
  or a staged record lands in a file about to be replaced. Normative

Recorded before it can be found late: compaction invalidates every WAL
offset iteration 2's `resident: keys` stores, so the compactor rebuilds the
offset map as it writes. Nothing breaks today because that storage half is
unimplemented — it would break later, looking like corruption.

Also corrected exploration/postgresql/buffer-and-checkpoint.md, which was
wrong on two counts: PG does NOT update its control file by rename (in-place
full-block write + CRC32C), and its checkpoint sketch assumes writeonce has
segment files, which it does not and deliberately will not.

Grounding measured on master: seed 20000 leaves a 986614-byte log; 20000
updates take it to 2590262 bytes with the SAME live rows, and boot+verify on
that store is 155ms.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-28 17:39:38 +02:00
746dc2b42b docs(databasev2): third track — the database beyond RAM, with per-table storage modes
- docs/stories/databasev2/, numbered from 1. Six PENDING database iterations
  moved from the language track and renumbered, keeping the old id in
  `was_language_iteration:` so a search for "iteration 32" still finds it:
  32 -> 3 WAL checkpoint, 23 -> 4 io_uring commit, 33 -> 7 single-file store,
  27 -> 8 query grammar, 20 -> 9 cross-program, 21 -> 10 keypair auth.
  Done work (9, 9b, 22) stays as v1 history; language 18 left whole
- the problem, read off the engine not guessed: rows are malloc'd slabs with
  addresses stable forever, NO eviction/spill/paging anywhere in database/src,
  the WAL never checkpoints so boot replays all history, and durability is one
  process-global WO_DATA so no table can say it matters more than another.
  An allocation failure IS a clean catchable WO_T_OOM — but swap thrash
  arrives first and carries no error signal at all, which is the real hazard
- four new iterations:
  1 measure the ceiling FIRST (curve not cliff; the three exits; kill -9 at
    exhaustion) — every later default should follow from a number
  2 `@table(mode: ram | durable | cold)` — the grammar ask. Small surface
    (Ast.table_cfg gains a key, the parser already rejects unknown args), big
    semantics: `durable` defaults so nothing changes silently, and the
    compiler refuses a durable row holding a `ref` into a ram table
  5 bounded tables + refuse/evict/back-pressure, shedding BEFORE the OS acts
  6 cold tiering — mostly forks, incl. whether the language surfaces the
    fault cost and whether @unique on cold is refused outright. A paged
    B-tree stays rejected: if tiering needs one, reject tiering
- 39 links repointed, link TEXT renumbered to track-local ids; arc gains one
  pointer row replacing the six moved; board + board-views cover three tracks
- linkcheck 0 broken / 0 anchors; no code blocks in any story

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-26 20:52:48 +02:00
Renamed from docs/stories/language-runtime-database/32-wal-checkpoint.md (Browse further)