Commit graph

7 commits

Author SHA1 Message Date
04017aea26 docs(db2-keys): spec — delta records for keys-resident updates
- updates append a DELTA (class, id, field, value, back-pointer), not a
  full row. The workload decides it: a product catalogue changes one
  narrow field of a wide row on every order, so a full-row append would
  rewrite every field to move one integer on a shop's hottest path
- the back-pointer keeps the id map at one slot per row, which is the
  mode's whole premise; a map growing per update would defeat it
- ONE fold function, three callers (read, replay, compaction). Three
  implementations of one rule is how they drift, and a fold that differs
  between reading and replaying is a database that changes its mind at
  boot. Named as the design's principal risk
- indexed columns MAY change: price is exactly what a catalogue indexes,
  so forbidding it would be a restriction users meet immediately
- no chain cap, deliberately. Compaction already rewrites live rows, so
  every checkpoint resets every chain, and deltas grow the log which
  pulls the next checkpoint forward — the workload that lengthens chains
  triggers the fold that shortens them
- the risk that accepts: one hot SKU under an otherwise quiet write
  rate. Task 7's benchmark must include it
- supersedes the 2026-08-26 spec's one-line full-row Update sketch,
  marked in place rather than left as a second design in the tree

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
(cherry picked from commit c9a88b05c4c62a5df13253686c990aaccc3f9cbb)
2026-08-30 20:37:27 +02:00
3507aafea3 docs(databasev2): propagate iteration 1's findings to every consumer
Audit found 4 of 10 iterations citing it1 and 4 carrying stale claims the
measurement contradicts.

- 05: framing was contradicted, not merely incomplete. Its goal expected a
  gradient to detect ("back-pressure before the cliff"); there is no cliff
  — SIGKILL with swap off, exit 0 with swap on, and read latency STEPS
  (1us -> 487us) rather than departing. Heading and goal rewritten; the
  measurement makes the goal stronger, not weaker
- 05: budget must be bytes — 3.3x footprint spread — with headroom for
  index doublings, else it fires during a rehash
- 05: new goal — eviction policy QUALITY is decisive, since getting the
  resident set wrong costs 273x, not a few percent
- 06: its revival question now has a reference point. 273x is the KERNEL
  SWAP path; `resident: keys` preads via page cache and must beat it. This
  file revives only if 5c/5d lands near 273x rather than well below
- 04: write path is not where pressure bites (append ~1%, read 273x), so
  the io_uring question that matters is iteration 2's deferred read-path
  one, not group-commit
- 00-story: problem statement asserted the store "refuses the insert
  rather than dying". Corrected in place — a banner above it was not
  enough, a skimmer never reaches it
- residency spec: "swap thrash and the OOM killer" named exits that were
  not measured; replaced with silence-or-a-corpse

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-27 22:47:37 +02:00
a873cf7331 feat(db-bench): random-read-over-cap leg — the 273x collapse
- `randread N R` in the sample: fill N rows, read R across the WHOLE range
- Weyl order `i*2654435761 mod n` — no RNG in the language, none needed;
  both legs read the SAME key order so residency is the only variable
- `randread` driver leg: control (256 MiB, does not bind) vs over-cap
  (6 MiB + swap), sizes kept modest — quick resolves it in ~5s
- gates the RATIO, not the absolutes: over-cap reads/sec belongs to the
  box's swap device, the factor between two runs belongs to the engine
- reads must all resolve (hits == R) or the leg fails; a collapse measured
  over unresolved reads is noise
- 133 checks, 0 failures; gate bites on a doctored collapse_x

Measured — this closes the gap the swap leg left:

- resident 1 851 166 reads/sec, p50 0us p99 1us
- over-cap    6 771 reads/sec, p50 128us p99 487us
- 273x throughput, ~480x p99, all 20 000 reads resolving in both
- so the two access patterns sit ~270x apart under identical pressure:
  append-mostly insert ~1%, random read 273x
- departure is a STEP not a curve (1us -> 487us, nothing between), which
  is why p99_departure_decile finds no knee — there is none

- caveat recorded, NOT inherited: this is demand-paged anonymous memory
  through swap (4 KiB/fault, no readahead). `resident: keys` preads via
  the page cache — should be better, but databasev2 2 task 7 must measure
  its own read path. New criterion added there

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-27 20:56:19 +02:00
0c9b2c45d8 feat(db-bench): measure the RAM ceiling — databasev2 1
- `Wide` text-heavy reference shape beside Int-only `Item`
- `growth N int|text`: per-decile RSS read from own /proc/self/status
- `growth-verify`: survivor of a crash must be a contiguous intact prefix
- four footprint legs under a rootless cgroup v2 cap, swap on/off
- `ceiling` leg: die at the cap, then replay must come back intact
- footprint read as median-of-marginals; doublings a separate metric
- 121 checks, 0 failures; footprint gated ±10%, kill-timing ±100%

Measured, and it inverted two of the iteration's own predictions:

- footprint 96.5-100 B/row Int vs 320.6-324 B/row text = 3.3x, NOT the
  "order of magnitude" three docs asserted
- table storage has NO checked ceiling: SIGKILL signal 9, not a catchable
  WO_T_OOM. overcommit lets malloc succeed; kernel kills on page touch
- swap is NOT latency collapse: 900k rows 148s capped-with-swap vs 150s
  uncapped. Append-mostly never re-touches cold pages
- ack-after-fsync survives an OOM kill: ~40k rows, no holes, no corruption
- iteration 2's budget dependency is REMOVED not satisfied — there is no
  "swap onset" to derive a fraction from

- fix: subprocess returncode -9 was labelled a "checked refusal"; 137 is
  the shell spelling of the same signal

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-27 20:18:47 +02:00
1ee7cce597 docs: retract the row-encoding rewrite — the flat record format already exists
- found during the pre-execution review of the plan, before any code
- the claim was wrong in both spec and plan: table.c's db_val_encode builds
  the IN-MEMORY slot; the FILE record is a separate encoding in wal.c and has
  been flat since iteration 9. enc_val inlines every kind recursively with no
  pointer anywhere; dec_val reads it back; a record is
  `WO_WAL_INSERT | class_id | id | <value per field>` in the
  len|crc|payload|mark frame; scan_record already preads and CRC-verifies a
  record at an arbitrary offset
- so the row encoding needs NO change, and Task 5 (a "self-contained,
  offset-based" rewrite billed as the iteration's substantive engineering) is
  DELETED, not reduced. 8 tasks -> 7, and the highest-risk task is gone
- the real difficulty is where the spec never looked: wo_wal_append_insert
  stages into a 1 MiB buffer, so a record's final offset is unknown until
  flush. Threading an accurate offset back through a buffered writer —
  correct across partial flush, failed commit and torn tail — is now Task 5's
  first two steps, with a unit test that straddles a buffer boundary and a
  case asserting no offset is published for a record that never reached disk
- dependent claims corrected: the mmap alternative's premise, the read-path
  bullet (now names scan_record/dec_val), and the self-review coverage table,
  which records the retraction rather than quietly dropping the row
- root cause worth noting: reading one layer and inferring another. Second
  time this iteration — the first was assuming WO_HEAP_MB bounded table
  storage when it bounds the VM arena
- no code written yet; linkcheck 0 broken / 0 anchors

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-26 23:00:18 +02:00
566559cf70 docs: name the residency value keys, not index (review change)
- `resident: all | keys` replaces `resident: all | index`. Two reasons beyond
  taste: it kills the collision with the `index:` argument
  (`@table(index: [customer], resident: index)` read badly), and it puts both
  values on ONE axis — each now answers "what row data stays resident",
  where `all`/`index` mixed a quantity with a structure name
- accurate as well as clearer: what stays resident is the id->offset map, the
  secondary indexes and the unique shadows — all key structures; row payloads
  are exactly what leaves. `resident: none` was rejected as overclaiming,
  since the indexes very much are resident
- checked for collisions: neither `all` nor `keys` is a keyword or a builtin
  (`key_at`/`val_at` exist, bare `keys` does not)
- the spec's wart note became a recorded decision; the rejected spelling is
  kept quoted so the rationale still reads
- fixes a bug I introduced in the 2026-08-26 track move: all six moved
  iterations carried a banner reading "Part of [Story — the database beyond
  RAM]" whose link pointed at the LANGUAGE arc — correct target, lying text,
  the exact failure mode the link audit warned about. Banners now point at
  the databasev2 story, and the original "Part of" line says plainly which
  track the iteration was authored in before the move
- linkcheck 0 broken / 0 anchors

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-26 22:47:24 +02:00
da30aa6527 docs: amend principle 7 — the log is authoritative, residency is declared
- driving case: a 120 GB order table on a 32 GB host. Not a tuning problem;
  no eviction policy fixes it. Developer accepted reconsidering the principle
- principle 7 rewritten: durability half UNCHANGED and unconditional
  (WAL-logged, fsync before ack, CRC-dropped torn tail); residency half
  demoted from law to per-table declaration. Old wording quoted in place so
  the amendment is legible, with the reason: a doctrine a real workload
  cannot satisfy gets ignored, and the failure it produced was an OOM kill
- spec: docs/superpowers/specs/2026-08-26-table-residency-design.md
  One log-structured engine — the WAL already holds every row, so keep an
  in-RAM id->offset map and pread rows back. No second engine, no user-space
  row cache (the kernel page cache is the hot copy, which is already this
  repo's stated position and why it avoids O_DIRECT)
- arithmetic that makes it work: 240M rows x 16 B of index = ~3.8 GB
  resident in 32 GB. Indexes stay resident, rows do not. Buys ~2 orders of
  magnitude, not infinity — stated plainly in the spec
- grammar: two optional keys, `durable: true|false` and `resident: all|index`,
  both defaulting to today's behaviour, so all 28 existing @table
  declarations compile untouched and no golden is reblessed
- rejected, with reasons recorded: mmap (rows are pointer-bearing —
  table.c returns (uintptr_t)t as the slot word), buffer pool (the Rust-era
  phase-12 design that died with that track), paged B-tree (stays rejected),
  a three-valued enum, automatic spill, disk-backed-by-default
- self-review caught the budget defaulting to "none" while promising the ERP
  developer a diagnostic instead of the OOM killer — contradiction fixed:
  the budget defaults to a fraction of host memory, and its value comes from
  databasev2 1's swap-onset measurement
- live docs that contradicted the amendment updated (subagent doctrine,
  its guide, discarded.md's two rows, iteration 04's read claim, 07, 38);
  dated specs/plans left as records. linkcheck 0 broken / 0 anchors

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-26 22:42:27 +02:00