9 commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
| 8311330531 |
docs(db2-keys): reconcile databasev2 and porch markdown with the code
- loader's resident:keys refusal said "rows are still fully resident" and "until tasks 5c/5d land". Both false since f606fc9. Corrected to name the real blocker: UPDATE needs read-modify-append - databasev2 00-story: the sequence graph drew 2->3->4, which reads as 3 needing 2 and 4 needing 3. Both backwards, and it still drew the 2->5->6 path the 2026-08-27 amendment retired. Redrawn stating only real dependencies, with 4 and 3 shown as composing rather than ordered, and the execution order that actually happened - databasev2 03: the hazard and its Outstanding entry both claimed nothing fails "because iteration 2's storage half is unimplemented". Marked discharged, and recorded that the hazard named only half the danger — the bitmap walk would have dropped keys rows outright - databasev2 06: pending -> hold (largely superseded, revisit only on a measurement); dated its 5c/5d references - porch 01: rewritten to the settled shape. readiness ready, status in-progress, phases B and C marked superseded with why - porch 01 claimed time.after "is still a reserved builtin id". False — builtin 90, implemented. That claim is what made the iteration look cheaper than it is - porch README gains honest ledger rows for both features (partial, being rebuilt), not shipped - skill-catalog README pointed at a story path that moved tracks; linkcheck now 0 broken Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> (cherry picked from commit b3d8c403e1d19ac27ec966de85cb293e0765795c) |
|||
| 02b4b13a52 |
Merge master into db-residency-doctrine — and close the two half-exposed features
The branch was 17 ahead / 25 behind with 11 conflicting files, and drifting further: db.c had been rewritten twice on master since (group commit, then compaction). Resolved rather than rebased so both histories stay legible. Conflicts, and how each was settled: - db.c: BOTH semantics kept. Master's fatal path and compaction check now sit behind the branch's `table_is_durable` predicate, in all three inline arms — a volatile table reaches neither the barrier nor the compaction check - db-bench sample: every mode from both sides (growth, growth-verify, randread, replayseed, wmix) and ONE `boot` mode, which both sides had added independently - db-bench.py: all six legs kept. Both sides had also grown the same WAL-size helper under different names; collapsed into one - perf-targets: the branch's §5 (RAM ceiling) then master's §6/§7 — master's numbering had already assumed a §5 it did not have - story frontmatter: master's `status` (the landing truth) plus the branch's `readiness` axis. 03 would have read `done` + `refine`, which is a contradiction — it was brainstormed and landed on master, so `ready` - board: both standup blocks newest-first; master's chain rows (a superset); the branch's databasev2 1-2 rows with master's 3-4. Fixed a stray `|` in master's row 3 - baseline: master's, then REGENERATED from a full campaign — 143 metrics, 132 checks, 0 failures with both sides' legs present TWO HALF-EXPOSED FEATURES FIXED, because the merge rule is that master gets no feature that is honoured in name only: - `resident: keys` PARSED, set a .wob flag, and did nothing: rows stayed fully resident. A developer could declare a 120 GB table keys-resident, watch it compile, and be OOM-killed. The loader now REFUSES it with a message naming what to write instead, until tasks 5c/5d land. The compiler still parses it and its AST golden still passes, so the grammar work stays tested - `durable: false` was honoured ONLY on the inline path. wo_db_exec_req had no guard at all, so a volatile table written from an actor on a worker shard would still be logged — precisely porch's session-table case, and precisely what iteration 2 exists to provide. All three request-path arms now carry the same predicate. Found by reading the merged code, not by a test: the obvious probe runs main() on the primary and therefore only exercises the inline path Verified on the merged tree: wovm-test 0, woc-test 0, oop-e2e 122/0, residency-accept 8/0, db-bench 132/0, linkcheck clean. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
|||
| 8b29eb492c |
docs(db): T6 closeout — checkpoint documented, chain's last link lands
databasev2 3, task 6. Documentation, plus three gate-tolerance corrections that are justified rather than silent. - 04-db-binding.md: the NORMATIVE rule — compaction may run only where nothing is staged (a correctness requirement, not scheduling), recovery is unchanged, and a failed compaction is a missed optimisation rather than a durability event - database/src/CODE-LOGIC.md: why one file and not snapshot-plus-tail (Postgres CANNOT compact — page deltas; ours are full row images, so a compacted log IS a store), why rename is the whole crash-safety story, why the dump flushes but does NOT fsync when it does, why the replacement is preallocated, and where the trigger is checked - README: the checkpoint knobs, the extended walstats line, the boot mode - story -> status: done, with criteria split met/outstanding - board: standup entry in the six-question shape, both rows rewritten THE OBLIGATION IS AT THE COMPACTOR, not only in a spec: compaction moves every record, so it invalidates every WAL offset iteration 2's `resident: keys` stores, and the loop that knows each record's new position must rebuild that map. Nothing fails today because that storage half is unimplemented — it would fail later, looking like corruption. Board claim corrected before it shipped: I wrote that the concurrency chain is "complete". It is not — chain 5 stays in-progress because databasev2 4's part B was never done and its premise was invalidated by part A. Every link has landed its PLANNED work; that is a different statement. Gate tolerances, each with the measurement that justifies it: - ckpt.pause_us_max is no longer gated relatively. The raw pause scales with the live set and this workload's live set is not fixed (wmix's hist_dump inserts a row per latency bucket), so gating it gates the box. Added ckpt.pause_us_per_mb — the engine's own rate, gated for real, and the metric that would have caught the 8x dump regression — with the absolute 50ms budget still guarding the raw pause - ram.*.msgrate 15% -> 70%. PRE-EXISTING, and measured: 10.7M-17.9M msgs/sec across ten full runs, several predating this work — a 1.67x spread against a 15% gate - durable.sN.*.p99us 100% -> 300%, with more evidence than the first widening: mixread 1043/2318/4147us, mixwrite 1623/4446us on the same build. Floors stay the real guard and are not slack Battery: wovm-test 36 suites 0 fail, woc-test, oop-e2e 119/0, db-bench 117 checks 0 failures, linkcheck clean. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
|||
| 74399ffc68 |
docs(plan): WAL checkpoint — 6 tasks, databasev2 3
Plan for the approved spec. Code-free per the repo convention (docs/plan/discarded.md:54); the executor writes the code. - T1 wo_wal_compact: walk live rows via the bitmap, append one INSERT each through the EXISTING append path, fsync, rename over the live log, fsync the parent dir, reopen the descriptor. Test asserts BOTH that the log shrank AND that a replay reproduces the same rows/ids/values — shorter alone is worthless, a truncating bug also passes that - T2 a stale temp file is removed at open and never read. The test uses PLAUSIBLE records, not garbage: garbage would be rejected anyway and would prove nothing - T3 the trigger as a PURE decision (used bytes, last compaction's measured output, floor) so it is unit-testable without a store; env knobs for floor and ratio, which is what makes the policy testable at all. No timer, with the reason. The check is called only where nothing is staged, asserted by a test that stages and expects deferral - T4 kill -9 DURING compaction, extending the existing fork-based crash battery. Asserts the PROPERTY — the store equals the pre- or the post-compaction content, never a mixture, and every acked id survives. Run repeatedly and state the count: it is a race, one green run proves little - T5 measure space reclaimed, boot before/after, and the stop-the-world PAUSE against a stated budget. If the pause exceeds it, stop and report — the alternatives are bought against that number, not before it - T6 closeout, including the normative ordering rule in 04-db-binding.md Constraints carried from the spec into every task: - recovery must NOT change; a task editing the replay path should stop - the dump must FLUSH PERIODICALLY. stage() grows the staging buffer by doubling, so dumping a whole store through one buffer would hold the entire store in RAM — the unbounded growth databasev2 1 identified as how this engine dies - a FAILED compaction is a missed optimisation, not a durability event, so it must not take databasev2 4's fatal path - gate tolerances must not be waived wholesale (part A's T4 made that mistake), and the baseline is full-mode — writing a quick-mode baseline over it is a regression part A also made Deliberately NOT a task: rebuilding the `resident: keys` offset map. It cannot be implemented against a feature that does not exist yet, so T6 records it as an obligation at the compactor and in the story instead of a stub nobody can test. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
|||
| 69c34c9a89 |
docs(spec): WAL checkpoint — compact by rewrite + atomic rename
databasev2 3, chain 6. Brainstormed 2026-08-28 after databasev2 4 part A
landed.
Design: compact the log by rewriting it as one record per live row into a
temp file, fsync, rename over the live WAL, fsync the parent dir, reopen.
Recovery is COMPLETELY UNCHANGED — boot still opens one file and replays
it — and the crash criterion ("the same store as if the checkpoint had
never started") is satisfied by rename, not by code we must get right.
Read .dev/reference/postgresql for this. The finding is that PG's design
is UNAVAILABLE to us, which is what makes the simpler option legitimate:
- PG never compacts its WAL; segments before the redo point are recycled
by rename or unlinked. Its records are page deltas, so a compacted redo
log is not a store — hence heap files, a control file, a redo pointer,
a second recovery source and a separate process
- ours are FULL ROW IMAGES (apply_record implements UPDATE as
remove-then-recreate), so a compacted log IS a complete store. That one
difference deletes all of the above from the design
- what IS worth porting is the ordering discipline: publish the new
"recovery starts here" atomically and LAST, so a crash falls back. PG
needs a start-of-checkpoint redo pointer plus an end-of-checkpoint
control file update; we get the same property from one rename, because
we can swap the whole data set atomically and PG cannot
Forks settled:
- no snapshot format — the compacted log is the snapshot, existing grammar,
so no new encoder or decoder and the dump reuses wo_wal_append_insert
- one source, not two
- volume-only trigger, as a ratio against the LAST compaction's measured
output (the denominator is known exactly; estimating the live set would
mean estimating Text) with an absolute floor. NO TIMER — PG's exists to
bound loss from unflushed buffers and we have none; an idle log does not
grow. Copying the mechanism without the reason was the trap
- stop-the-world, with the pause measured against a stated budget rather
than assumed acceptable; alternatives are bought against a number
- compaction may run ONLY where nothing is staged (right after a barrier),
or a staged record lands in a file about to be replaced. Normative
Recorded before it can be found late: compaction invalidates every WAL
offset iteration 2's `resident: keys` stores, so the compactor rebuilds the
offset map as it writes. Nothing breaks today because that storage half is
unimplemented — it would break later, looking like corruption.
Also corrected exploration/postgresql/buffer-and-checkpoint.md, which was
wrong on two counts: PG does NOT update its control file by rename (in-place
full-block write + CRC32C), and its checkpoint sketch assumes writeonce has
segment files, which it does not and deliberately will not.
Grounding measured on master: seed 20000 leaves a 986614-byte log; 20000
updates take it to 2590262 bytes with the SAME live rows, and boot+verify on
that store is 155ms.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
|||
| 5b1a8c96a1 |
feat(db-bench): replay baseline — boot cost tracks history, not data
Closes the last gap in databasev2 1; gives databasev2 3 its "before". - `boot` mode: does NOTHING. WO_DATA replay runs before main, so a mode with no work measures replay plus a fixed startup - `replayseed N M`: N inserts + M updates — same live rows, longer log - `replay` leg: empty-store startup floor measured and SUBTRACTED, then two shapes timed, median of 3 boots each - premise check: updates must actually append WAL records, else the two shapes are one measurement and the penalty means nothing - WAL bytes = non-zero prefix, never file size (fallocate'd to 1 MiB) - per-record cost stored in NANOseconds: as us it rounded 5.5 and 5.3 to 6 and 5, too coarse for the number a checkpoint exists to improve - 148 checks, 0 failures; gate bites on a doctored ns_per_record Measured — same 20 000 live rows, different history: - 20 000 records: 980 035 B WAL, 110 ms replay, 5.5 us/record - 40 000 records: 1 960 035 B WAL, 211 ms replay, 5.3 us/record - 1.9x boot cost for an IDENTICAL dataset; per-record cost flat, so replay is linear in records not rows - extrapolated: 10M records ~55 s of boot, 100M ~9 min - databasev2 3 correction: it planned to use "22's aged-store replay numbers", which never existed — 22 proved restart correctness, never timed it - databasev2 3 hazard recorded: compaction rewrites the log and moves every record, so it invalidates every `resident: keys` offset — an arbitrary byte in a rewritten file, not stale-but-readable - databasev2 1 -> status: done Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
|||
| 1fe808b7a4 |
docs(stories): add readiness, retire status: refine, sweep all 47 iterations
- `readiness: ready | refine` is a SECOND axis, orthogonal to status.
`ready` = the brainstorm is complete and the decisions are LOCKED (a spec
approved, or the forks explicitly confirmed). `refine` = open forks remain
and it cannot be planned yet
- `status: refine` RETIRED because it carried both meanings at once, so a held
iteration with an approved spec (language 18, 26) was indistinguishable from
one nobody had thought about. status is now purely where the WORK is:
done | in-progress | pending | hold — `pending` was already the board's own
rendering word, so nothing new was invented
- all 47 iterations classified from EVIDENCE in their own text, not by guess:
"the four forks are SETTLED" / "spec + plan approved" / "Approved spec:" for
ready; "Forks the spec must settle" / "no spec exists yet" for refine. Every
shipped iteration is ready by definition. 19 done, 5 in-progress, 15
pending, 8 hold; 27 ready, 20 refine
- two iterations moved refine -> in-progress rather than -> pending: language
31 and 34 are absorbed into 24 and work on them is literally happening, which
the board already showed as 🔄 while their frontmatter said otherwise. That
disagreement is now gone
- board legend, board-views' frontmatter contract, and two new Dataview
queries updated — the useful one being `readiness: ready AND status:
pending`, the startable set
WHAT THE NEW AXIS IMMEDIATELY SURFACED: of 15 pending iterations, exactly ONE
is startable — databasev2 4, io_uring group-commit, whose forks were confirmed
settled 2026-08-20. Everything else pending needs a brainstorm first. That was
invisible while one key carried both meanings, and it is now on the board.
Also caught by the sweep, unrelated to readiness but found by cross-checking
frontmatter against the board: SIX duplicate rows. Every iteration moved into
databasev2 was still listed in the LANGUAGE pending table under its retired id
(23, 32, 33, 20, 21, 27) as well as its new one. Stale copies removed. And two
databasev2 rows made claims the sweep contradicts — iteration 1 was billed
"startable today" while its forks are open, and 6 still called itself the
ceiling-raiser after 2 took that role.
Docs only. linkcheck 0 broken / 0 anchors.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
|||
| 566559cf70 |
docs: name the residency value keys, not index (review change)
- `resident: all | keys` replaces `resident: all | index`. Two reasons beyond taste: it kills the collision with the `index:` argument (`@table(index: [customer], resident: index)` read badly), and it puts both values on ONE axis — each now answers "what row data stays resident", where `all`/`index` mixed a quantity with a structure name - accurate as well as clearer: what stays resident is the id->offset map, the secondary indexes and the unique shadows — all key structures; row payloads are exactly what leaves. `resident: none` was rejected as overclaiming, since the indexes very much are resident - checked for collisions: neither `all` nor `keys` is a keyword or a builtin (`key_at`/`val_at` exist, bare `keys` does not) - the spec's wart note became a recorded decision; the rejected spelling is kept quoted so the rationale still reads - fixes a bug I introduced in the 2026-08-26 track move: all six moved iterations carried a banner reading "Part of [Story — the database beyond RAM]" whose link pointed at the LANGUAGE arc — correct target, lying text, the exact failure mode the link audit warned about. Banners now point at the databasev2 story, and the original "Part of" line says plainly which track the iteration was authored in before the move - linkcheck 0 broken / 0 anchors Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
|||
| 746dc2b42b |
docs(databasev2): third track — the database beyond RAM, with per-table storage modes
- docs/stories/databasev2/, numbered from 1. Six PENDING database iterations
moved from the language track and renumbered, keeping the old id in
`was_language_iteration:` so a search for "iteration 32" still finds it:
32 -> 3 WAL checkpoint, 23 -> 4 io_uring commit, 33 -> 7 single-file store,
27 -> 8 query grammar, 20 -> 9 cross-program, 21 -> 10 keypair auth.
Done work (9, 9b, 22) stays as v1 history; language 18 left whole
- the problem, read off the engine not guessed: rows are malloc'd slabs with
addresses stable forever, NO eviction/spill/paging anywhere in database/src,
the WAL never checkpoints so boot replays all history, and durability is one
process-global WO_DATA so no table can say it matters more than another.
An allocation failure IS a clean catchable WO_T_OOM — but swap thrash
arrives first and carries no error signal at all, which is the real hazard
- four new iterations:
1 measure the ceiling FIRST (curve not cliff; the three exits; kill -9 at
exhaustion) — every later default should follow from a number
2 `@table(mode: ram | durable | cold)` — the grammar ask. Small surface
(Ast.table_cfg gains a key, the parser already rejects unknown args), big
semantics: `durable` defaults so nothing changes silently, and the
compiler refuses a durable row holding a `ref` into a ram table
5 bounded tables + refuse/evict/back-pressure, shedding BEFORE the OS acts
6 cold tiering — mostly forks, incl. whether the language surfaces the
fault cost and whether @unique on cold is refused outright. A paged
B-tree stays rejected: if tiering needs one, reject tiering
- 39 links repointed, link TEXT renumbered to track-local ids; arc gains one
pointer row replacing the six moved; board + board-views cover three tracks
- linkcheck 0 broken / 0 anchors; no code blocks in any story
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
Renamed from docs/stories/language-runtime-database/32-wal-checkpoint.md (Browse further)