Plan for the approved spec. Code-free per the repo convention
(docs/plan/discarded.md:54); the executor writes the code.
- T1 wo_wal_compact: walk live rows via the bitmap, append one INSERT
each through the EXISTING append path, fsync, rename over the live log,
fsync the parent dir, reopen the descriptor. Test asserts BOTH that the
log shrank AND that a replay reproduces the same rows/ids/values —
shorter alone is worthless, a truncating bug also passes that
- T2 a stale temp file is removed at open and never read. The test uses
PLAUSIBLE records, not garbage: garbage would be rejected anyway and
would prove nothing
- T3 the trigger as a PURE decision (used bytes, last compaction's
measured output, floor) so it is unit-testable without a store; env
knobs for floor and ratio, which is what makes the policy testable at
all. No timer, with the reason. The check is called only where nothing
is staged, asserted by a test that stages and expects deferral
- T4 kill -9 DURING compaction, extending the existing fork-based crash
battery. Asserts the PROPERTY — the store equals the pre- or the
post-compaction content, never a mixture, and every acked id survives.
Run repeatedly and state the count: it is a race, one green run proves
little
- T5 measure space reclaimed, boot before/after, and the stop-the-world
PAUSE against a stated budget. If the pause exceeds it, stop and report
— the alternatives are bought against that number, not before it
- T6 closeout, including the normative ordering rule in 04-db-binding.md
Constraints carried from the spec into every task:
- recovery must NOT change; a task editing the replay path should stop
- the dump must FLUSH PERIODICALLY. stage() grows the staging buffer by
doubling, so dumping a whole store through one buffer would hold the
entire store in RAM — the unbounded growth databasev2 1 identified as
how this engine dies
- a FAILED compaction is a missed optimisation, not a durability event,
so it must not take databasev2 4's fatal path
- gate tolerances must not be waived wholesale (part A's T4 made that
mistake), and the baseline is full-mode — writing a quick-mode baseline
over it is a regression part A also made
Deliberately NOT a task: rebuilding the `resident: keys` offset map. It
cannot be implemented against a feature that does not exist yet, so T6
records it as an obligation at the compactor and in the story instead of
a stub nobody can test.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>