Commit graph

11 commits

Author SHA1 Message Date
4ad24d6381 docs(db2-migrate): close out iteration 12
- crash-before-rename test: a COMPLETE valid migrated temp beside the
  untouched original is discarded and the boot re-migrates — the
  sharpest point on the crash timeline, deterministic, no fault
  injection needed
- story: all six tasks done with commit hashes, all eight criteria met
  with the test that proves each, plus the three deviations from the
  plan and why (transcode over replay, lazy head, poison forces
  transcode)
- CODE-LOGIC: migration section; also corrected limitation 3, which
  still claimed unbounded hot-row chains — iteration 11 closed that
- status board row 12; deploy guide's rollback section gets its real
  answer (rolling back across a migration is a migration backwards:
  expect the refusal, restore the .bak)
- test_wal 5966 pass, 0 fail

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
(cherry picked from commit 4bb6ece2531e2123eb958c91d9bef4a6528eab3b)
2026-08-31 21:54:27 +02:00
e08c26309a feat(db2-delta): lift the resident:keys refusal, prove it end to end
- loader.c: delete the INCOMPLETE-update BAIL; durable:false +
  resident:keys stays refused (nowhere to read from)
- table.c: root-cause fix for the Text-index gap — a keys-resident
  borrow now holds ENGINE values, matching wo_row_ptr's contract
  (table.h's "no VM pointer" doctrine), not a VM-decoded row. Fixes
  idx_hash/idx_cols_equal/wo_idx_probe AND db.c's GET_FIELD/PROBE
  arms with one change; reproduced pre-fix as an ASan
  heap-buffer-overflow
- docs/examples/residency: Product is genuinely resident:keys;
  residency-accept.sh's refusal leg replaced by proving the program
  runs and stock survives a restart (11/0)
- test_wal.c: oracle test drives resident:all and resident:keys
  through the same update sequence and asserts identical rows;
  Text-indexed-update test catches the representation bug; five
  pre-existing tests corrected to the fixed contract (4746/0)
- story, README, status board, CODE-LOGIC.md updated; three known
  limitations documented: mid-drain stale reads, O(N^2) replay in
  chain length, compaction blind to per-row chain length

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
(cherry picked from commit b87c68f950f01aa5e572fbb86a0f374adc83d813)
2026-08-30 20:37:49 +02:00
02b4b13a52 Merge master into db-residency-doctrine — and close the two half-exposed features
The branch was 17 ahead / 25 behind with 11 conflicting files, and drifting
further: db.c had been rewritten twice on master since (group commit, then
compaction). Resolved rather than rebased so both histories stay legible.

Conflicts, and how each was settled:

- db.c: BOTH semantics kept. Master's fatal path and compaction check now sit
  behind the branch's `table_is_durable` predicate, in all three inline arms —
  a volatile table reaches neither the barrier nor the compaction check
- db-bench sample: every mode from both sides (growth, growth-verify, randread,
  replayseed, wmix) and ONE `boot` mode, which both sides had added
  independently
- db-bench.py: all six legs kept. Both sides had also grown the same
  WAL-size helper under different names; collapsed into one
- perf-targets: the branch's §5 (RAM ceiling) then master's §6/§7 — master's
  numbering had already assumed a §5 it did not have
- story frontmatter: master's `status` (the landing truth) plus the branch's
  `readiness` axis. 03 would have read `done` + `refine`, which is a
  contradiction — it was brainstormed and landed on master, so `ready`
- board: both standup blocks newest-first; master's chain rows (a superset);
  the branch's databasev2 1-2 rows with master's 3-4. Fixed a stray `|` in
  master's row 3
- baseline: master's, then REGENERATED from a full campaign — 143 metrics,
  132 checks, 0 failures with both sides' legs present

TWO HALF-EXPOSED FEATURES FIXED, because the merge rule is that master gets
no feature that is honoured in name only:

- `resident: keys` PARSED, set a .wob flag, and did nothing: rows stayed fully
  resident. A developer could declare a 120 GB table keys-resident, watch it
  compile, and be OOM-killed. The loader now REFUSES it with a message naming
  what to write instead, until tasks 5c/5d land. The compiler still parses it
  and its AST golden still passes, so the grammar work stays tested
- `durable: false` was honoured ONLY on the inline path. wo_db_exec_req had no
  guard at all, so a volatile table written from an actor on a worker shard
  would still be logged — precisely porch's session-table case, and precisely
  what iteration 2 exists to provide. All three request-path arms now carry the
  same predicate. Found by reading the merged code, not by a test: the obvious
  probe runs main() on the primary and therefore only exercises the inline path

Verified on the merged tree: wovm-test 0, woc-test 0, oop-e2e 122/0,
residency-accept 8/0, db-bench 132/0, linkcheck clean.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-29 10:14:25 +02:00
8b29eb492c docs(db): T6 closeout — checkpoint documented, chain's last link lands
databasev2 3, task 6. Documentation, plus three gate-tolerance
corrections that are justified rather than silent.

- 04-db-binding.md: the NORMATIVE rule — compaction may run only where
  nothing is staged (a correctness requirement, not scheduling), recovery
  is unchanged, and a failed compaction is a missed optimisation rather
  than a durability event
- database/src/CODE-LOGIC.md: why one file and not snapshot-plus-tail
  (Postgres CANNOT compact — page deltas; ours are full row images, so a
  compacted log IS a store), why rename is the whole crash-safety story,
  why the dump flushes but does NOT fsync when it does, why the
  replacement is preallocated, and where the trigger is checked
- README: the checkpoint knobs, the extended walstats line, the boot mode
- story -> status: done, with criteria split met/outstanding
- board: standup entry in the six-question shape, both rows rewritten

THE OBLIGATION IS AT THE COMPACTOR, not only in a spec: compaction moves
every record, so it invalidates every WAL offset iteration 2's
`resident: keys` stores, and the loop that knows each record's new
position must rebuild that map. Nothing fails today because that storage
half is unimplemented — it would fail later, looking like corruption.

Board claim corrected before it shipped: I wrote that the concurrency
chain is "complete". It is not — chain 5 stays in-progress because
databasev2 4's part B was never done and its premise was invalidated by
part A. Every link has landed its PLANNED work; that is a different
statement.

Gate tolerances, each with the measurement that justifies it:

- ckpt.pause_us_max is no longer gated relatively. The raw pause scales
  with the live set and this workload's live set is not fixed (wmix's
  hist_dump inserts a row per latency bucket), so gating it gates the
  box. Added ckpt.pause_us_per_mb — the engine's own rate, gated for
  real, and the metric that would have caught the 8x dump regression —
  with the absolute 50ms budget still guarding the raw pause
- ram.*.msgrate 15% -> 70%. PRE-EXISTING, and measured: 10.7M-17.9M
  msgs/sec across ten full runs, several predating this work — a 1.67x
  spread against a 15% gate
- durable.sN.*.p99us 100% -> 300%, with more evidence than the first
  widening: mixread 1043/2318/4147us, mixwrite 1623/4446us on the same
  build. Floors stay the real guard and are not slack

Battery: wovm-test 36 suites 0 fail, woc-test, oop-e2e 119/0,
db-bench 117 checks 0 failures, linkcheck clean.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-29 06:48:34 +02:00
0b618ace19 docs+fix(db): T6 closeout — and reads no longer wait for the barrier
databasev2 4 part A, task 6. Mostly documentation, plus one real fix the
full battery caught.

THE FIX. The drain held EVERY DB reply until the barrier — including
reads, which stage nothing and have no stake in durability. That parked
readers behind an fsync for no reason: durable.sN.mixread.p99 rose from
~1043us to 4057us. Only a statement that actually staged a record now has
its reply held. Caught by the gate, not by review.

THE TRADE, recorded rather than smoothed over. What remains is inherent: a
barrier blocks the owner shard LONGER (more records per fsync) though LESS
OFTEN, so anything queued behind one waits. Three full runs of the same
build gave durable.sN.mixread.p99 of 1043 / 2318 / 4147us and wmix.p99 of
8758 / 20000us — a 2-4x spread with the box near idle. So part A buys ~3x
write throughput at the cost of a longer, noisier tail on the owner shard,
and that is the strongest argument for part B (submit and keep serving).

- durable.sN.*.p99us tolerance widened to 100% WITH the reason in the
  code: a 2-4x-variable tail gated at 50% gates the disk, not the engine.
  The floor is the real guard and is not slack — mixread's (4172us) came
  within 25us of tripping on the worst run. Baseline refreshed; a fresh
  full run then passed 106 checks 0 failures

EXIT STATUS MOVED 3 -> 74 (sysexits EX_IOERR). 3 and 4 are already used by
SAMPLES for their own meanings — db-bench's own `verify` exits 3 on a
checksum mismatch, and it is the gate that exercises durability, so a
durability abort exiting 3 would have been indistinguishable from the
mismatch it should help diagnose. The low range belongs to programs.

Docs:

- story: progress, the payoff measured two ways, the cost side, criteria
  split met/outstanding, and a "part B — its premise changed" section:
  it was justified by "close the 66x gap", but that gap is two problems
  and only the concurrent one was a batching problem
- board: standup entry in the six-question shape; both databasev2 4 rows
  rewritten. They had said "close the 66x gap" — recorded as MIS-STATED
  rather than quietly renumbered
- 00-wob-format.md and 04-db-binding.md: the normative failure contract
  ("a failed WAL commit traps WO_T_IO after un-applying the row") was
  false; corrected, along with the tick-scoped group commit that never
  happened
- database/src/CODE-LOGIC.md: where the barrier runs and why there, why
  replies are held, why the inline path is asymmetric, the one failure
  rule, and how to measure it
- db-bench README: the wmix mode, the env knobs, and the tmpfs warning

Battery: wovm-test 36 suites 0 fail, woc-test, oop-e2e 119/0,
db-bench 106/0, linkcheck clean.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-28 16:48:23 +02:00
b74e13d21e feat(db): durable:false skips the WAL append and replay
Task 4 of docs/superpowers/plans/2026-08-26-table-residency.md — the first
behavioural change in the iteration.

- db.c: one predicate, `table_is_durable`, gating the three EXISTING mutation
  sites. Kept as a function rather than an inlined condition so
  database/src/CODE-LOGIC.md's "nothing else may mutate storage" claim keeps
  holding — the choke points stayed three
- the ack contract is untouched for durable tables: RAM applied, record
  staged, one commit before the ack, and a failed commit still removes the row
- replay: a log holding records for a class the image now declares volatile is
  a real migration case, not corruption. apply_record returns -2 (distinct
  from -1), wo_wal_replay_ex reports the class id, and main.c names it and
  exits 2. `wo_wal_replay` stays as the NULL wrapper, so all 156 WAL unit
  checks are untouched
- measured, not asserted: 50 inserts wrote 1500 WAL bytes into a durable
  table and ZERO into a volatile one. The file's SIZE proves nothing (it is
  fallocate'd to 1 MiB up front), so the gate measures the non-zero prefix

BUG I INTRODUCED AND CAUGHT: the mismatch message first printed the class name
with %s, but wo_str.data is `char data[]` with NO NUL terminator (obj.h) — a
buffer over-read. Now %.*s with the explicit length, and re-verified under
ASan.

New gate `just residency` (8 checks), because everything above was otherwise
a one-off manual measurement: restart behaviour, the zero-byte write path, the
mismatch refusal (exit 2, names the class, NOT reported as corruption), and
both compile-time refusals. Its own first run failed two checks for a bug in
the script rather than the feature — `woc | grep` under `set -o pipefail`
returns woc's exit 1 even when grep matches, since woc exits 1 whenever it
reports diagnostics. Captures first now, with the reason noted inline.

Also new: corpus run/table-volatile-inprocess pins that a volatile table is a
FULL table in-process — same @unique enforcement, same index probe, same query
surface. Only survival differs, and that is unobservable from inside one
process.

Gates: woc-test 557/0, 18 runtime suites 0 fail, cli_smoke OK, oop-e2e 119/0
(was 118), residency 8/0, employee 8/0, db-actor 8/0, site 21/0, ASan clean on
the new replay path.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-27 08:23:49 +02:00
881b90e9c7 feat: O(1) read path — index probe wired end to end
- engine: wo_idx_probe answers single-column equality from the index
  hash buckets (idx_hash_key1 reproduces idx_hash bit for bit; verify
  compares exactly as the slab walk did, so results identical);
  composite indexes keep the walk; both executors wired (local + DB
  actor RPC)
- compiler: probe_key_of_where lowers "var.col == key" on an indexed
  column to DB_PROBE; all where guards still run (guard stays the
  final arbiter); keys = ident/int-literal only; Float/Bytes excluded
  (engine raw-eq narrower than VM float-eq)
- measured: reads 1.3k -> 1.3M ops/s, p50 600us -> 1us (~x850);
  query x830; mixread 1.3k -> 89k s1, 21 -> ~1.9k sN
- gate policy moved into the driver (tolerance_for: refresh-proof);
  latency floors max(4x,100us); quick mode skips poll-bound mix
  floors; both tolerance classes proven to bite
- proof: test_table wo_idx_probe suite (RED first), corpus
  query-index-probe 105/0, full battery green, TSan clean, two
  campaigns pass the refreshed baseline

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-22 16:44:56 +02:00
5e51a151c0 docs: CODE-LOGIC — DB-actor RPC + ring-params fix, slot surface
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-21 13:16:36 +02:00
30ef8d8533 feat: insert executes (iteration 9, Task 3) -- DB_STUB retires for insert
- compiler: `insert Class { ... }` is a typed Ast.Insert in statement
  AND expression position, sharing the ctor literal's field grammar;
  typechecked with the ctor's omittable rule; result = the row id (Int)
- owner pass: the engine copies at the row API, so an insert BORROWS
  its field values -- no transfer, no E304; node is trap-capable and
  carries a live-mask drop entry like DbStub did
- emit: builtin 61 window = class-id const + one slot per DECLARED
  field in declaration order; omitted defaults emitted, omitted ?scalar
  gets WO_NIL_SCALAR, other omitted optionals the zero word; fresh
  argument values reaped after (the push/set copy semantics)
- runtime: database/src/db.c executes via the choke-point row API;
  rt.db/rt.wal opaque handles on wo_rt; WO_DATA=<dir> = replay
  <dir>/shard-0.wal at boot + commit-before-ack per statement (the
  builtin's return IS the ack until iteration 8 ticks); failed commit
  un-applies the row and traps WO_T_IO; loader validates the class-id
  slot (variable window documented in wob.h + format doc)
- the promised diff: trap/pricing-set-price-db-stub is now
  run/pricing-set-price-insert printing engine-allocated ids;
  durability smoke prints 1,2 then 3,4 across two WO_DATA runs
- old "bare insert is an Ident" unit test rewritten to the new
  contract; runner's loader mirror accepts id 61; goldens re-blessed
- gates: oop-accept ALL CRITERIA MET, oop-e2e 71/0, woc-test 566/0,
  wovm-test green, log-watcher 7/0

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-15 11:15:06 +02:00
c3741262cb feat(database): typed WAL + boot replay (iteration 9, Task 2)
- database/src/wal.{c,h}: framed records len|crc32|payload|mark
  ("WOL1" written last -- no mark, no record), typed-row payloads
  walking the class-table kinds (nested records, containers, nil
  encodings), little-endian like the loader
- commit order verbatim from the shipped phase-D pattern: RAM apply,
  stage, ONE pwrite + ONE fdatasync for the batch, ack after -- group
  commit is everything staged riding one sync
- replay decodes straight into engine-owned values (no VM at boot)
  and re-enters rows through the choke-point row API, so Task 4's
  indexes will rebuild for free; next_id advances past replayed ids
  this shard owns (wo_row_create_raw)
- torn tail = short/CRC-fail/no-mark/zero-len: intact prefix applies,
  tear dropped whole, wo_wal_open positions AT the tear so the next
  commit overwrites it; CRC-valid-but-undecodable = corruption, loud
- wo_wal_check: offline oracle, no engine needed -- the crash
  battery's verifier
- test_wal 90/0 ASan+UBSan incl. five crash-battery rounds (fork,
  insert/commit/ack-over-pipe, SIGKILL mid-stream: zero acked-but-
  missing, zero acked-but-wrong); all runtime suites green, oop-e2e
  71/0; binding doc WAL section + CODE-LOGIC + plan Task 2 checked

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-15 11:01:41 +02:00
936bd14bff feat(database): class-shaped row storage (iteration 9, Task 1)
- database/src/table.{c,h}: per-shard per-class slabs (256 rows,
  malloc'd, never moved -- row addresses stable for 9b's row views),
  occupancy bitmap, LIFO slot reuse, open-addressing id hash with
  tombstones (ids never 0, never reused)
- field encoding walks the same .wob class-table kinds the VM walks:
  scalars raw (WO_NIL_SCALAR passes through), Texts copied to db_text,
  owned objects flattened recursively to db_rec, containers
  element-wise; GCREF refused at encode (the GC bulkhead, defensively)
- two one-way copy gates: insert copies VM values in, read allocates
  fresh VM values out -- no VM pointer in a slab, no slab pointer in
  the VM, proven by mutating originals after insert
- id discipline: per table per shard, S+1 step N; owner = (id-1) % N;
  N-parametric, runs at N=1 until iteration 8, tested at N=3
- choke points: wo_row_insert/wo_row_remove carry the INDEX HOOK
  sites Task 4 attaches to; nothing else mutates storage
- runtime/Makefile links database/src into every wovm + test binary
- test_table 827/0 ASan+UBSan; oop-e2e 71/0; log-watcher 7/0;
  binding doc docs/plan/oop-vm/04-db-binding.md; CODE-LOGIC.md beside
  the code; plan Task 1 checked off

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-15 10:54:24 +02:00