The branch was 17 ahead / 25 behind with 11 conflicting files, and drifting
further: db.c had been rewritten twice on master since (group commit, then
compaction). Resolved rather than rebased so both histories stay legible.
Conflicts, and how each was settled:
- db.c: BOTH semantics kept. Master's fatal path and compaction check now sit
behind the branch's `table_is_durable` predicate, in all three inline arms —
a volatile table reaches neither the barrier nor the compaction check
- db-bench sample: every mode from both sides (growth, growth-verify, randread,
replayseed, wmix) and ONE `boot` mode, which both sides had added
independently
- db-bench.py: all six legs kept. Both sides had also grown the same
WAL-size helper under different names; collapsed into one
- perf-targets: the branch's §5 (RAM ceiling) then master's §6/§7 — master's
numbering had already assumed a §5 it did not have
- story frontmatter: master's `status` (the landing truth) plus the branch's
`readiness` axis. 03 would have read `done` + `refine`, which is a
contradiction — it was brainstormed and landed on master, so `ready`
- board: both standup blocks newest-first; master's chain rows (a superset);
the branch's databasev2 1-2 rows with master's 3-4. Fixed a stray `|` in
master's row 3
- baseline: master's, then REGENERATED from a full campaign — 143 metrics,
132 checks, 0 failures with both sides' legs present
TWO HALF-EXPOSED FEATURES FIXED, because the merge rule is that master gets
no feature that is honoured in name only:
- `resident: keys` PARSED, set a .wob flag, and did nothing: rows stayed fully
resident. A developer could declare a 120 GB table keys-resident, watch it
compile, and be OOM-killed. The loader now REFUSES it with a message naming
what to write instead, until tasks 5c/5d land. The compiler still parses it
and its AST golden still passes, so the grammar work stays tested
- `durable: false` was honoured ONLY on the inline path. wo_db_exec_req had no
guard at all, so a volatile table written from an actor on a worker shard
would still be logged — precisely porch's session-table case, and precisely
what iteration 2 exists to provide. All three request-path arms now carry the
same predicate. Found by reading the merged code, not by a test: the obvious
probe runs main() on the primary and therefore only exercises the inline path
Verified on the merged tree: wovm-test 0, woc-test 0, oop-e2e 122/0,
residency-accept 8/0, db-bench 132/0, linkcheck clean.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
databasev2 3, task 6. Documentation, plus three gate-tolerance
corrections that are justified rather than silent.
- 04-db-binding.md: the NORMATIVE rule — compaction may run only where
nothing is staged (a correctness requirement, not scheduling), recovery
is unchanged, and a failed compaction is a missed optimisation rather
than a durability event
- database/src/CODE-LOGIC.md: why one file and not snapshot-plus-tail
(Postgres CANNOT compact — page deltas; ours are full row images, so a
compacted log IS a store), why rename is the whole crash-safety story,
why the dump flushes but does NOT fsync when it does, why the
replacement is preallocated, and where the trigger is checked
- README: the checkpoint knobs, the extended walstats line, the boot mode
- story -> status: done, with criteria split met/outstanding
- board: standup entry in the six-question shape, both rows rewritten
THE OBLIGATION IS AT THE COMPACTOR, not only in a spec: compaction moves
every record, so it invalidates every WAL offset iteration 2's
`resident: keys` stores, and the loop that knows each record's new
position must rebuild that map. Nothing fails today because that storage
half is unimplemented — it would fail later, looking like corruption.
Board claim corrected before it shipped: I wrote that the concurrency
chain is "complete". It is not — chain 5 stays in-progress because
databasev2 4's part B was never done and its premise was invalidated by
part A. Every link has landed its PLANNED work; that is a different
statement.
Gate tolerances, each with the measurement that justifies it:
- ckpt.pause_us_max is no longer gated relatively. The raw pause scales
with the live set and this workload's live set is not fixed (wmix's
hist_dump inserts a row per latency bucket), so gating it gates the
box. Added ckpt.pause_us_per_mb — the engine's own rate, gated for
real, and the metric that would have caught the 8x dump regression —
with the absolute 50ms budget still guarding the raw pause
- ram.*.msgrate 15% -> 70%. PRE-EXISTING, and measured: 10.7M-17.9M
msgs/sec across ten full runs, several predating this work — a 1.67x
spread against a 15% gate
- durable.sN.*.p99us 100% -> 300%, with more evidence than the first
widening: mixread 1043/2318/4147us, mixwrite 1623/4446us on the same
build. Floors stay the real guard and are not slack
Battery: wovm-test 36 suites 0 fail, woc-test, oop-e2e 119/0,
db-bench 117 checks 0 failures, linkcheck clean.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Plan for the approved spec. Code-free per the repo convention
(docs/plan/discarded.md:54); the executor writes the code.
- T1 wo_wal_compact: walk live rows via the bitmap, append one INSERT
each through the EXISTING append path, fsync, rename over the live log,
fsync the parent dir, reopen the descriptor. Test asserts BOTH that the
log shrank AND that a replay reproduces the same rows/ids/values —
shorter alone is worthless, a truncating bug also passes that
- T2 a stale temp file is removed at open and never read. The test uses
PLAUSIBLE records, not garbage: garbage would be rejected anyway and
would prove nothing
- T3 the trigger as a PURE decision (used bytes, last compaction's
measured output, floor) so it is unit-testable without a store; env
knobs for floor and ratio, which is what makes the policy testable at
all. No timer, with the reason. The check is called only where nothing
is staged, asserted by a test that stages and expects deferral
- T4 kill -9 DURING compaction, extending the existing fork-based crash
battery. Asserts the PROPERTY — the store equals the pre- or the
post-compaction content, never a mixture, and every acked id survives.
Run repeatedly and state the count: it is a race, one green run proves
little
- T5 measure space reclaimed, boot before/after, and the stop-the-world
PAUSE against a stated budget. If the pause exceeds it, stop and report
— the alternatives are bought against that number, not before it
- T6 closeout, including the normative ordering rule in 04-db-binding.md
Constraints carried from the spec into every task:
- recovery must NOT change; a task editing the replay path should stop
- the dump must FLUSH PERIODICALLY. stage() grows the staging buffer by
doubling, so dumping a whole store through one buffer would hold the
entire store in RAM — the unbounded growth databasev2 1 identified as
how this engine dies
- a FAILED compaction is a missed optimisation, not a durability event,
so it must not take databasev2 4's fatal path
- gate tolerances must not be waived wholesale (part A's T4 made that
mistake), and the baseline is full-mode — writing a quick-mode baseline
over it is a regression part A also made
Deliberately NOT a task: rebuilding the `resident: keys` offset map. It
cannot be implemented against a feature that does not exist yet, so T6
records it as an obligation at the compactor and in the story instead of
a stub nobody can test.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
databasev2 3, chain 6. Brainstormed 2026-08-28 after databasev2 4 part A
landed.
Design: compact the log by rewriting it as one record per live row into a
temp file, fsync, rename over the live WAL, fsync the parent dir, reopen.
Recovery is COMPLETELY UNCHANGED — boot still opens one file and replays
it — and the crash criterion ("the same store as if the checkpoint had
never started") is satisfied by rename, not by code we must get right.
Read .dev/reference/postgresql for this. The finding is that PG's design
is UNAVAILABLE to us, which is what makes the simpler option legitimate:
- PG never compacts its WAL; segments before the redo point are recycled
by rename or unlinked. Its records are page deltas, so a compacted redo
log is not a store — hence heap files, a control file, a redo pointer,
a second recovery source and a separate process
- ours are FULL ROW IMAGES (apply_record implements UPDATE as
remove-then-recreate), so a compacted log IS a complete store. That one
difference deletes all of the above from the design
- what IS worth porting is the ordering discipline: publish the new
"recovery starts here" atomically and LAST, so a crash falls back. PG
needs a start-of-checkpoint redo pointer plus an end-of-checkpoint
control file update; we get the same property from one rename, because
we can swap the whole data set atomically and PG cannot
Forks settled:
- no snapshot format — the compacted log is the snapshot, existing grammar,
so no new encoder or decoder and the dump reuses wo_wal_append_insert
- one source, not two
- volume-only trigger, as a ratio against the LAST compaction's measured
output (the denominator is known exactly; estimating the live set would
mean estimating Text) with an absolute floor. NO TIMER — PG's exists to
bound loss from unflushed buffers and we have none; an idle log does not
grow. Copying the mechanism without the reason was the trap
- stop-the-world, with the pause measured against a stated budget rather
than assumed acceptable; alternatives are bought against a number
- compaction may run ONLY where nothing is staged (right after a barrier),
or a staged record lands in a file about to be replaced. Normative
Recorded before it can be found late: compaction invalidates every WAL
offset iteration 2's `resident: keys` stores, so the compactor rebuilds the
offset map as it writes. Nothing breaks today because that storage half is
unimplemented — it would break later, looking like corruption.
Also corrected exploration/postgresql/buffer-and-checkpoint.md, which was
wrong on two counts: PG does NOT update its control file by rename (in-place
full-block write + CRC32C), and its checkpoint sketch assumes writeonce has
segment files, which it does not and deliberately will not.
Grounding measured on master: seed 20000 leaves a 986614-byte log; 20000
updates take it to 2590262 bytes with the SAME live rows, and boot+verify on
that store is 155ms.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
databasev2 4 part A, task 6. Mostly documentation, plus one real fix the
full battery caught.
THE FIX. The drain held EVERY DB reply until the barrier — including
reads, which stage nothing and have no stake in durability. That parked
readers behind an fsync for no reason: durable.sN.mixread.p99 rose from
~1043us to 4057us. Only a statement that actually staged a record now has
its reply held. Caught by the gate, not by review.
THE TRADE, recorded rather than smoothed over. What remains is inherent: a
barrier blocks the owner shard LONGER (more records per fsync) though LESS
OFTEN, so anything queued behind one waits. Three full runs of the same
build gave durable.sN.mixread.p99 of 1043 / 2318 / 4147us and wmix.p99 of
8758 / 20000us — a 2-4x spread with the box near idle. So part A buys ~3x
write throughput at the cost of a longer, noisier tail on the owner shard,
and that is the strongest argument for part B (submit and keep serving).
- durable.sN.*.p99us tolerance widened to 100% WITH the reason in the
code: a 2-4x-variable tail gated at 50% gates the disk, not the engine.
The floor is the real guard and is not slack — mixread's (4172us) came
within 25us of tripping on the worst run. Baseline refreshed; a fresh
full run then passed 106 checks 0 failures
EXIT STATUS MOVED 3 -> 74 (sysexits EX_IOERR). 3 and 4 are already used by
SAMPLES for their own meanings — db-bench's own `verify` exits 3 on a
checksum mismatch, and it is the gate that exercises durability, so a
durability abort exiting 3 would have been indistinguishable from the
mismatch it should help diagnose. The low range belongs to programs.
Docs:
- story: progress, the payoff measured two ways, the cost side, criteria
split met/outstanding, and a "part B — its premise changed" section:
it was justified by "close the 66x gap", but that gap is two problems
and only the concurrent one was a batching problem
- board: standup entry in the six-question shape; both databasev2 4 rows
rewritten. They had said "close the 66x gap" — recorded as MIS-STATED
rather than quietly renumbered
- 00-wob-format.md and 04-db-binding.md: the normative failure contract
("a failed WAL commit traps WO_T_IO after un-applying the row") was
false; corrected, along with the tick-scoped group commit that never
happened
- database/src/CODE-LOGIC.md: where the barrier runs and why there, why
replies are held, why the inline path is asymmetric, the one failure
rule, and how to measure it
- db-bench README: the wmix mode, the env knobs, and the tmpfs warning
Battery: wovm-test 36 suites 0 fail, woc-test, oop-e2e 119/0,
db-bench 106/0, linkcheck clean.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Plan for the approved spec. Code-free per the repo convention
(docs/plan/discarded.md:54); the executor writes the code.
- T1 a failed barrier is detected and fatal — one entry point that names
the operation, errno, WAL path and batch size, then exits. The abort
path itself stays unexercised and the task says so rather than buying
coverage with a fault-injection switch
- T2 the barrier moves to the drain point and replies are held; the
request path stops committing per append. Riskiest task, and its risk
is one place: the crash legs. Plan says STOP if they fail, do not
adjust the test
- T3 the inline path takes the same fatal rule but keeps its own barrier,
with a comment explaining the asymmetry so the next reader does not
"fix" it. Looks like a no-op; without it the two paths disagree, which
is the unevenness the spec exists to remove
- T4 prove batches actually form BEFORE measuring the payoff — otherwise
a win gets attributed to the wrong cause. Also records peak staged
bytes, settling the no-cap decision with a number
- T5 measure, gate, write it down. If the payoff is absent, say so and
stop: part B must not start on an unproven premise
- T6 closeout, including the error catalogue — WO_T_IO leaving the write
path is language-visible and must be written down
Spec corrected while planning: it pointed at durable.s1.seed as the
payoff. Wrong, structurally — worker shards hold no WAL, so a queue only
exists when other shards write, and a serial writer has nothing to batch
with. The real target is durable.sN.mixwrite: 480 ops/s at p99 5888us
against s1's 1023 at p99 664, so adding shards currently makes durable
writing WORSE. That inversion is a better argument for the iteration than
the one the story recorded.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Brainstormed 2026-08-28. The iteration is split: part A batches, part B
(io_uring submission) is deferred until A's measurement says whether the
blocking boundary still dominates.
The story's premise needed correcting first:
- it says "replace fsync-per-commit with io_uring group-commit", but the
engine commits per STATEMENT — db.c calls wo_wal_commit right after
every append, all six sites, so each row change is one pwrite + one
fdatasync
- so two independent wins were being carried as one, and only the second
needs io_uring. The staging buffer already holds any number of records;
today it never holds more than one. Part A is mostly deleting calls
- iteration 22's numbers say A is where the payoff is: durable writes
4460 ops/s, mixwrite 1023 ops/s p99 664us, against 1.28M ops/s reads
Forks settled:
- batch boundary is QUEUE-DRAIN, not the tick this story had recorded: a
tick adds latency to a lone writer, taxing an idle system to serve a
busy one. Queue-drain self-tunes and needs no knob
- shard 0 holds each reply envelope instead of sending it, commits once
when the queue empties, then releases all — so a writer is acked after
the barrier carrying ITS record, which today is true only because
every batch has one member
- a failure between "RAM mutated" and "record durable" is a FATAL,
diagnosed abort. This replaces uneven behaviour that already exists:
insert rolls back, update and delete do not and say so in a comment
("RAM ahead of disk"). Batching would have multiplied that
- consequence stated, not slipped in: WO_T_IO leaves the write path
- no batch cap initially; peak staged bytes is measured so the question
is settled by a number
One gap disclosed rather than hidden: forcing a real fdatasync failure
needs mount privileges, so the unit test proves the error is DETECTED and
the abort itself stays covered by inspection.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Audit found 4 of 10 iterations citing it1 and 4 carrying stale claims the
measurement contradicts.
- 05: framing was contradicted, not merely incomplete. Its goal expected a
gradient to detect ("back-pressure before the cliff"); there is no cliff
— SIGKILL with swap off, exit 0 with swap on, and read latency STEPS
(1us -> 487us) rather than departing. Heading and goal rewritten; the
measurement makes the goal stronger, not weaker
- 05: budget must be bytes — 3.3x footprint spread — with headroom for
index doublings, else it fires during a rehash
- 05: new goal — eviction policy QUALITY is decisive, since getting the
resident set wrong costs 273x, not a few percent
- 06: its revival question now has a reference point. 273x is the KERNEL
SWAP path; `resident: keys` preads via page cache and must beat it. This
file revives only if 5c/5d lands near 273x rather than well below
- 04: write path is not where pressure bites (append ~1%, read 273x), so
the io_uring question that matters is iteration 2's deferred read-path
one, not group-commit
- 00-story: problem statement asserted the store "refuses the insert
rather than dying". Corrected in place — a banner above it was not
enough, a skimmer never reaches it
- residency spec: "swap thrash and the OOM killer" named exits that were
not measured; replaced with silence-or-a-corpse
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Closes the last gap in databasev2 1; gives databasev2 3 its "before".
- `boot` mode: does NOTHING. WO_DATA replay runs before main, so a mode
with no work measures replay plus a fixed startup
- `replayseed N M`: N inserts + M updates — same live rows, longer log
- `replay` leg: empty-store startup floor measured and SUBTRACTED, then
two shapes timed, median of 3 boots each
- premise check: updates must actually append WAL records, else the two
shapes are one measurement and the penalty means nothing
- WAL bytes = non-zero prefix, never file size (fallocate'd to 1 MiB)
- per-record cost stored in NANOseconds: as us it rounded 5.5 and 5.3 to
6 and 5, too coarse for the number a checkpoint exists to improve
- 148 checks, 0 failures; gate bites on a doctored ns_per_record
Measured — same 20 000 live rows, different history:
- 20 000 records: 980 035 B WAL, 110 ms replay, 5.5 us/record
- 40 000 records: 1 960 035 B WAL, 211 ms replay, 5.3 us/record
- 1.9x boot cost for an IDENTICAL dataset; per-record cost flat, so
replay is linear in records not rows
- extrapolated: 10M records ~55 s of boot, 100M ~9 min
- databasev2 3 correction: it planned to use "22's aged-store replay
numbers", which never existed — 22 proved restart correctness, never
timed it
- databasev2 3 hazard recorded: compaction rewrites the log and moves
every record, so it invalidates every `resident: keys` offset — an
arbitrary byte in a rewritten file, not stale-but-readable
- databasev2 1 -> status: done
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
- `randread N R` in the sample: fill N rows, read R across the WHOLE range
- Weyl order `i*2654435761 mod n` — no RNG in the language, none needed;
both legs read the SAME key order so residency is the only variable
- `randread` driver leg: control (256 MiB, does not bind) vs over-cap
(6 MiB + swap), sizes kept modest — quick resolves it in ~5s
- gates the RATIO, not the absolutes: over-cap reads/sec belongs to the
box's swap device, the factor between two runs belongs to the engine
- reads must all resolve (hits == R) or the leg fails; a collapse measured
over unresolved reads is noise
- 133 checks, 0 failures; gate bites on a doctored collapse_x
Measured — this closes the gap the swap leg left:
- resident 1 851 166 reads/sec, p50 0us p99 1us
- over-cap 6 771 reads/sec, p50 128us p99 487us
- 273x throughput, ~480x p99, all 20 000 reads resolving in both
- so the two access patterns sit ~270x apart under identical pressure:
append-mostly insert ~1%, random read 273x
- departure is a STEP not a curve (1us -> 487us, nothing between), which
is why p99_departure_decile finds no knee — there is none
- caveat recorded, NOT inherited: this is demand-paged anonymous memory
through swap (4 KiB/fault, no readahead). `resident: keys` preads via
the page cache — should be better, but databasev2 2 task 7 must measure
its own read path. New criterion added there
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
- `Wide` text-heavy reference shape beside Int-only `Item`
- `growth N int|text`: per-decile RSS read from own /proc/self/status
- `growth-verify`: survivor of a crash must be a contiguous intact prefix
- four footprint legs under a rootless cgroup v2 cap, swap on/off
- `ceiling` leg: die at the cap, then replay must come back intact
- footprint read as median-of-marginals; doublings a separate metric
- 121 checks, 0 failures; footprint gated ±10%, kill-timing ±100%
Measured, and it inverted two of the iteration's own predictions:
- footprint 96.5-100 B/row Int vs 320.6-324 B/row text = 3.3x, NOT the
"order of magnitude" three docs asserted
- table storage has NO checked ceiling: SIGKILL signal 9, not a catchable
WO_T_OOM. overcommit lets malloc succeed; kernel kills on page touch
- swap is NOT latency collapse: 900k rows 148s capped-with-swap vs 150s
uncapped. Append-mostly never re-touches cold pages
- ack-after-fsync survives an OOM kill: ~40k rows, no holes, no corruption
- iteration 2's budget dependency is REMOVED not satisfied — there is no
"swap onset" to derive a fraction from
- fix: subprocess returncode -9 was labelled a "checked refusal"; 137 is
the shell spelling of the same signal
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
- `readiness: ready | refine` is a SECOND axis, orthogonal to status.
`ready` = the brainstorm is complete and the decisions are LOCKED (a spec
approved, or the forks explicitly confirmed). `refine` = open forks remain
and it cannot be planned yet
- `status: refine` RETIRED because it carried both meanings at once, so a held
iteration with an approved spec (language 18, 26) was indistinguishable from
one nobody had thought about. status is now purely where the WORK is:
done | in-progress | pending | hold — `pending` was already the board's own
rendering word, so nothing new was invented
- all 47 iterations classified from EVIDENCE in their own text, not by guess:
"the four forks are SETTLED" / "spec + plan approved" / "Approved spec:" for
ready; "Forks the spec must settle" / "no spec exists yet" for refine. Every
shipped iteration is ready by definition. 19 done, 5 in-progress, 15
pending, 8 hold; 27 ready, 20 refine
- two iterations moved refine -> in-progress rather than -> pending: language
31 and 34 are absorbed into 24 and work on them is literally happening, which
the board already showed as 🔄 while their frontmatter said otherwise. That
disagreement is now gone
- board legend, board-views' frontmatter contract, and two new Dataview
queries updated — the useful one being `readiness: ready AND status:
pending`, the startable set
WHAT THE NEW AXIS IMMEDIATELY SURFACED: of 15 pending iterations, exactly ONE
is startable — databasev2 4, io_uring group-commit, whose forks were confirmed
settled 2026-08-20. Everything else pending needs a brainstorm first. That was
invisible while one key carried both meanings, and it is now on the board.
Also caught by the sweep, unrelated to readiness but found by cross-checking
frontmatter against the board: SIX duplicate rows. Every iteration moved into
databasev2 was still listed in the LANGUAGE pending table under its retired id
(23, 32, 33, 20, 21, 27) as well as its new one. Stale copies removed. And two
databasev2 rows made claims the sweep contradicts — iteration 1 was billed
"startable today" while its forks are open, and 6 still called itself the
ceiling-raiser after 2 took that role.
Docs only. linkcheck 0 broken / 0 anchors.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Flagged by the developer: the iteration still described the pre-brainstorm
three-mode design behind a "superseded in part" banner while four tasks had
landed against it.
- iteration 2 rewritten around the shape as built: two keys
(`durable: true|false`, `resident: all|keys`), not a `mode:` enum with
`cold`. status refine -> in-progress
- added a task-by-task progress table with commit hashes, and split the
acceptance criteria into MET (each with how it was verified, not just that
it passed — e.g. the goldens-unchanged claim is `git diff` over golden/
being empty after a WOC_BLESS run, since blessing rewrites all of them) and
OUTSTANDING with the task that owns each
- kept the history rather than deleting it: the three-mode replacement, the
"one real rewrite" that was fiction, and the opposite half that turned out
genuinely deep. An iteration file is where that record belongs
- board row rewritten to agree; the track index's "the lever" section, its
principle-7 paragraph and its sequence rationale all still taught the dead
three-mode design
ITERATION 6 IS NOW LARGELY SUPERSEDED, and bannered as such rather than
quietly gutted. `resident: keys` is the ceiling-raiser and it lives in
iteration 2 (tasks 5c/5d). More than relocated: 6's premise — a user-space
resident working set with faulting and 5's eviction policy — was specifically
REJECTED by the spec in favour of the kernel page cache, since a pread against
a cached page is a memcpy. What may still be left is recorded honestly: revisit
only with a measurement showing the page cache insufficient. Its fork list
survives, especially "does the language surface the fault cost at the use
site", which is still open and still the largest question about what writeonce
is. The sequence rationale is amended too — it had 6 as the ceiling-raiser and
5 as a prerequisite on the critical path; neither holds.
Docs only. linkcheck 0 broken / 0 anchors.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
- `resident: all | keys` replaces `resident: all | index`. Two reasons beyond
taste: it kills the collision with the `index:` argument
(`@table(index: [customer], resident: index)` read badly), and it puts both
values on ONE axis — each now answers "what row data stays resident",
where `all`/`index` mixed a quantity with a structure name
- accurate as well as clearer: what stays resident is the id->offset map, the
secondary indexes and the unique shadows — all key structures; row payloads
are exactly what leaves. `resident: none` was rejected as overclaiming,
since the indexes very much are resident
- checked for collisions: neither `all` nor `keys` is a keyword or a builtin
(`key_at`/`val_at` exist, bare `keys` does not)
- the spec's wart note became a recorded decision; the rejected spelling is
kept quoted so the rationale still reads
- fixes a bug I introduced in the 2026-08-26 track move: all six moved
iterations carried a banner reading "Part of [Story — the database beyond
RAM]" whose link pointed at the LANGUAGE arc — correct target, lying text,
the exact failure mode the link audit warned about. Banners now point at
the databasev2 story, and the original "Part of" line says plainly which
track the iteration was authored in before the move
- linkcheck 0 broken / 0 anchors
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
- driving case: a 120 GB order table on a 32 GB host. Not a tuning problem;
no eviction policy fixes it. Developer accepted reconsidering the principle
- principle 7 rewritten: durability half UNCHANGED and unconditional
(WAL-logged, fsync before ack, CRC-dropped torn tail); residency half
demoted from law to per-table declaration. Old wording quoted in place so
the amendment is legible, with the reason: a doctrine a real workload
cannot satisfy gets ignored, and the failure it produced was an OOM kill
- spec: docs/superpowers/specs/2026-08-26-table-residency-design.md
One log-structured engine — the WAL already holds every row, so keep an
in-RAM id->offset map and pread rows back. No second engine, no user-space
row cache (the kernel page cache is the hot copy, which is already this
repo's stated position and why it avoids O_DIRECT)
- arithmetic that makes it work: 240M rows x 16 B of index = ~3.8 GB
resident in 32 GB. Indexes stay resident, rows do not. Buys ~2 orders of
magnitude, not infinity — stated plainly in the spec
- grammar: two optional keys, `durable: true|false` and `resident: all|index`,
both defaulting to today's behaviour, so all 28 existing @table
declarations compile untouched and no golden is reblessed
- rejected, with reasons recorded: mmap (rows are pointer-bearing —
table.c returns (uintptr_t)t as the slot word), buffer pool (the Rust-era
phase-12 design that died with that track), paged B-tree (stays rejected),
a three-valued enum, automatic spill, disk-backed-by-default
- self-review caught the budget defaulting to "none" while promising the ERP
developer a diagnostic instead of the OOM killer — contradiction fixed:
the budget defaults to a fraction of host memory, and its value comes from
databasev2 1's swap-onset measurement
- live docs that contradicted the amendment updated (subagent doctrine,
its guide, discarded.md's two rows, iteration 04's read claim, 07, 38);
dated specs/plans left as records. linkcheck 0 broken / 0 anchors
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
- docs/stories/databasev2/, numbered from 1. Six PENDING database iterations
moved from the language track and renumbered, keeping the old id in
`was_language_iteration:` so a search for "iteration 32" still finds it:
32 -> 3 WAL checkpoint, 23 -> 4 io_uring commit, 33 -> 7 single-file store,
27 -> 8 query grammar, 20 -> 9 cross-program, 21 -> 10 keypair auth.
Done work (9, 9b, 22) stays as v1 history; language 18 left whole
- the problem, read off the engine not guessed: rows are malloc'd slabs with
addresses stable forever, NO eviction/spill/paging anywhere in database/src,
the WAL never checkpoints so boot replays all history, and durability is one
process-global WO_DATA so no table can say it matters more than another.
An allocation failure IS a clean catchable WO_T_OOM — but swap thrash
arrives first and carries no error signal at all, which is the real hazard
- four new iterations:
1 measure the ceiling FIRST (curve not cliff; the three exits; kill -9 at
exhaustion) — every later default should follow from a number
2 `@table(mode: ram | durable | cold)` — the grammar ask. Small surface
(Ast.table_cfg gains a key, the parser already rejects unknown args), big
semantics: `durable` defaults so nothing changes silently, and the
compiler refuses a durable row holding a `ref` into a ram table
5 bounded tables + refuse/evict/back-pressure, shedding BEFORE the OS acts
6 cold tiering — mostly forks, incl. whether the language surfaces the
fault cost and whether @unique on cold is refused outright. A paged
B-tree stays rejected: if tiering needs one, reject tiering
- 39 links repointed, link TEXT renumbered to track-local ids; arc gains one
pointer row replacing the six moved; board + board-views cover three tracks
- linkcheck 0 broken / 0 anchors; no code blocks in any story
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>