Commit graph

10 commits

Author SHA1 Message Date
e08c26309a feat(db2-delta): lift the resident:keys refusal, prove it end to end
- loader.c: delete the INCOMPLETE-update BAIL; durable:false +
  resident:keys stays refused (nowhere to read from)
- table.c: root-cause fix for the Text-index gap — a keys-resident
  borrow now holds ENGINE values, matching wo_row_ptr's contract
  (table.h's "no VM pointer" doctrine), not a VM-decoded row. Fixes
  idx_hash/idx_cols_equal/wo_idx_probe AND db.c's GET_FIELD/PROBE
  arms with one change; reproduced pre-fix as an ASan
  heap-buffer-overflow
- docs/examples/residency: Product is genuinely resident:keys;
  residency-accept.sh's refusal leg replaced by proving the program
  runs and stock survives a restart (11/0)
- test_wal.c: oracle test drives resident:all and resident:keys
  through the same update sequence and asserts identical rows;
  Text-indexed-update test catches the representation bug; five
  pre-existing tests corrected to the fixed contract (4746/0)
- story, README, status board, CODE-LOGIC.md updated; three known
  limitations documented: mid-drain stale reads, O(N^2) replay in
  chain length, compaction blind to per-row chain length

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
(cherry picked from commit b87c68f950f01aa5e572fbb86a0f374adc83d813)
2026-08-30 20:37:49 +02:00
02b4b13a52 Merge master into db-residency-doctrine — and close the two half-exposed features
The branch was 17 ahead / 25 behind with 11 conflicting files, and drifting
further: db.c had been rewritten twice on master since (group commit, then
compaction). Resolved rather than rebased so both histories stay legible.

Conflicts, and how each was settled:

- db.c: BOTH semantics kept. Master's fatal path and compaction check now sit
  behind the branch's `table_is_durable` predicate, in all three inline arms —
  a volatile table reaches neither the barrier nor the compaction check
- db-bench sample: every mode from both sides (growth, growth-verify, randread,
  replayseed, wmix) and ONE `boot` mode, which both sides had added
  independently
- db-bench.py: all six legs kept. Both sides had also grown the same
  WAL-size helper under different names; collapsed into one
- perf-targets: the branch's §5 (RAM ceiling) then master's §6/§7 — master's
  numbering had already assumed a §5 it did not have
- story frontmatter: master's `status` (the landing truth) plus the branch's
  `readiness` axis. 03 would have read `done` + `refine`, which is a
  contradiction — it was brainstormed and landed on master, so `ready`
- board: both standup blocks newest-first; master's chain rows (a superset);
  the branch's databasev2 1-2 rows with master's 3-4. Fixed a stray `|` in
  master's row 3
- baseline: master's, then REGENERATED from a full campaign — 143 metrics,
  132 checks, 0 failures with both sides' legs present

TWO HALF-EXPOSED FEATURES FIXED, because the merge rule is that master gets
no feature that is honoured in name only:

- `resident: keys` PARSED, set a .wob flag, and did nothing: rows stayed fully
  resident. A developer could declare a 120 GB table keys-resident, watch it
  compile, and be OOM-killed. The loader now REFUSES it with a message naming
  what to write instead, until tasks 5c/5d land. The compiler still parses it
  and its AST golden still passes, so the grammar work stays tested
- `durable: false` was honoured ONLY on the inline path. wo_db_exec_req had no
  guard at all, so a volatile table written from an actor on a worker shard
  would still be logged — precisely porch's session-table case, and precisely
  what iteration 2 exists to provide. All three request-path arms now carry the
  same predicate. Found by reading the merged code, not by a test: the obvious
  probe runs main() on the primary and therefore only exercises the inline path

Verified on the merged tree: wovm-test 0, woc-test 0, oop-e2e 122/0,
residency-accept 8/0, db-bench 132/0, linkcheck clean.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-29 10:14:25 +02:00
8b29eb492c docs(db): T6 closeout — checkpoint documented, chain's last link lands
databasev2 3, task 6. Documentation, plus three gate-tolerance
corrections that are justified rather than silent.

- 04-db-binding.md: the NORMATIVE rule — compaction may run only where
  nothing is staged (a correctness requirement, not scheduling), recovery
  is unchanged, and a failed compaction is a missed optimisation rather
  than a durability event
- database/src/CODE-LOGIC.md: why one file and not snapshot-plus-tail
  (Postgres CANNOT compact — page deltas; ours are full row images, so a
  compacted log IS a store), why rename is the whole crash-safety story,
  why the dump flushes but does NOT fsync when it does, why the
  replacement is preallocated, and where the trigger is checked
- README: the checkpoint knobs, the extended walstats line, the boot mode
- story -> status: done, with criteria split met/outstanding
- board: standup entry in the six-question shape, both rows rewritten

THE OBLIGATION IS AT THE COMPACTOR, not only in a spec: compaction moves
every record, so it invalidates every WAL offset iteration 2's
`resident: keys` stores, and the loop that knows each record's new
position must rebuild that map. Nothing fails today because that storage
half is unimplemented — it would fail later, looking like corruption.

Board claim corrected before it shipped: I wrote that the concurrency
chain is "complete". It is not — chain 5 stays in-progress because
databasev2 4's part B was never done and its premise was invalidated by
part A. Every link has landed its PLANNED work; that is a different
statement.

Gate tolerances, each with the measurement that justifies it:

- ckpt.pause_us_max is no longer gated relatively. The raw pause scales
  with the live set and this workload's live set is not fixed (wmix's
  hist_dump inserts a row per latency bucket), so gating it gates the
  box. Added ckpt.pause_us_per_mb — the engine's own rate, gated for
  real, and the metric that would have caught the 8x dump regression —
  with the absolute 50ms budget still guarding the raw pause
- ram.*.msgrate 15% -> 70%. PRE-EXISTING, and measured: 10.7M-17.9M
  msgs/sec across ten full runs, several predating this work — a 1.67x
  spread against a 15% gate
- durable.sN.*.p99us 100% -> 300%, with more evidence than the first
  widening: mixread 1043/2318/4147us, mixwrite 1623/4446us on the same
  build. Floors stay the real guard and are not slack

Battery: wovm-test 36 suites 0 fail, woc-test, oop-e2e 119/0,
db-bench 117 checks 0 failures, linkcheck clean.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-29 06:48:34 +02:00
0b618ace19 docs+fix(db): T6 closeout — and reads no longer wait for the barrier
databasev2 4 part A, task 6. Mostly documentation, plus one real fix the
full battery caught.

THE FIX. The drain held EVERY DB reply until the barrier — including
reads, which stage nothing and have no stake in durability. That parked
readers behind an fsync for no reason: durable.sN.mixread.p99 rose from
~1043us to 4057us. Only a statement that actually staged a record now has
its reply held. Caught by the gate, not by review.

THE TRADE, recorded rather than smoothed over. What remains is inherent: a
barrier blocks the owner shard LONGER (more records per fsync) though LESS
OFTEN, so anything queued behind one waits. Three full runs of the same
build gave durable.sN.mixread.p99 of 1043 / 2318 / 4147us and wmix.p99 of
8758 / 20000us — a 2-4x spread with the box near idle. So part A buys ~3x
write throughput at the cost of a longer, noisier tail on the owner shard,
and that is the strongest argument for part B (submit and keep serving).

- durable.sN.*.p99us tolerance widened to 100% WITH the reason in the
  code: a 2-4x-variable tail gated at 50% gates the disk, not the engine.
  The floor is the real guard and is not slack — mixread's (4172us) came
  within 25us of tripping on the worst run. Baseline refreshed; a fresh
  full run then passed 106 checks 0 failures

EXIT STATUS MOVED 3 -> 74 (sysexits EX_IOERR). 3 and 4 are already used by
SAMPLES for their own meanings — db-bench's own `verify` exits 3 on a
checksum mismatch, and it is the gate that exercises durability, so a
durability abort exiting 3 would have been indistinguishable from the
mismatch it should help diagnose. The low range belongs to programs.

Docs:

- story: progress, the payoff measured two ways, the cost side, criteria
  split met/outstanding, and a "part B — its premise changed" section:
  it was justified by "close the 66x gap", but that gap is two problems
  and only the concurrent one was a batching problem
- board: standup entry in the six-question shape; both databasev2 4 rows
  rewritten. They had said "close the 66x gap" — recorded as MIS-STATED
  rather than quietly renumbered
- 00-wob-format.md and 04-db-binding.md: the normative failure contract
  ("a failed WAL commit traps WO_T_IO after un-applying the row") was
  false; corrected, along with the tick-scoped group commit that never
  happened
- database/src/CODE-LOGIC.md: where the barrier runs and why there, why
  replies are held, why the inline path is asymmetric, the one failure
  rule, and how to measure it
- db-bench README: the wmix mode, the env knobs, and the tmpfs warning

Battery: wovm-test 36 suites 0 fail, woc-test, oop-e2e 119/0,
db-bench 106/0, linkcheck clean.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-28 16:48:23 +02:00
b74e13d21e feat(db): durable:false skips the WAL append and replay
Task 4 of docs/superpowers/plans/2026-08-26-table-residency.md — the first
behavioural change in the iteration.

- db.c: one predicate, `table_is_durable`, gating the three EXISTING mutation
  sites. Kept as a function rather than an inlined condition so
  database/src/CODE-LOGIC.md's "nothing else may mutate storage" claim keeps
  holding — the choke points stayed three
- the ack contract is untouched for durable tables: RAM applied, record
  staged, one commit before the ack, and a failed commit still removes the row
- replay: a log holding records for a class the image now declares volatile is
  a real migration case, not corruption. apply_record returns -2 (distinct
  from -1), wo_wal_replay_ex reports the class id, and main.c names it and
  exits 2. `wo_wal_replay` stays as the NULL wrapper, so all 156 WAL unit
  checks are untouched
- measured, not asserted: 50 inserts wrote 1500 WAL bytes into a durable
  table and ZERO into a volatile one. The file's SIZE proves nothing (it is
  fallocate'd to 1 MiB up front), so the gate measures the non-zero prefix

BUG I INTRODUCED AND CAUGHT: the mismatch message first printed the class name
with %s, but wo_str.data is `char data[]` with NO NUL terminator (obj.h) — a
buffer over-read. Now %.*s with the explicit length, and re-verified under
ASan.

New gate `just residency` (8 checks), because everything above was otherwise
a one-off manual measurement: restart behaviour, the zero-byte write path, the
mismatch refusal (exit 2, names the class, NOT reported as corruption), and
both compile-time refusals. Its own first run failed two checks for a bug in
the script rather than the feature — `woc | grep` under `set -o pipefail`
returns woc's exit 1 even when grep matches, since woc exits 1 whenever it
reports diagnostics. Captures first now, with the reason noted inline.

Also new: corpus run/table-volatile-inprocess pins that a volatile table is a
FULL table in-process — same @unique enforcement, same index probe, same query
surface. Only survival differs, and that is unobservable from inside one
process.

Gates: woc-test 557/0, 18 runtime suites 0 fail, cli_smoke OK, oop-e2e 119/0
(was 118), residency 8/0, employee 8/0, db-actor 8/0, site 21/0, ASan clean on
the new replay path.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-27 08:23:49 +02:00
881b90e9c7 feat: O(1) read path — index probe wired end to end
- engine: wo_idx_probe answers single-column equality from the index
  hash buckets (idx_hash_key1 reproduces idx_hash bit for bit; verify
  compares exactly as the slab walk did, so results identical);
  composite indexes keep the walk; both executors wired (local + DB
  actor RPC)
- compiler: probe_key_of_where lowers "var.col == key" on an indexed
  column to DB_PROBE; all where guards still run (guard stays the
  final arbiter); keys = ident/int-literal only; Float/Bytes excluded
  (engine raw-eq narrower than VM float-eq)
- measured: reads 1.3k -> 1.3M ops/s, p50 600us -> 1us (~x850);
  query x830; mixread 1.3k -> 89k s1, 21 -> ~1.9k sN
- gate policy moved into the driver (tolerance_for: refresh-proof);
  latency floors max(4x,100us); quick mode skips poll-bound mix
  floors; both tolerance classes proven to bite
- proof: test_table wo_idx_probe suite (RED first), corpus
  query-index-probe 105/0, full battery green, TSan clean, two
  campaigns pass the refreshed baseline

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-22 16:44:56 +02:00
5e51a151c0 docs: CODE-LOGIC — DB-actor RPC + ring-params fix, slot surface
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-21 13:16:36 +02:00
30ef8d8533 feat: insert executes (iteration 9, Task 3) -- DB_STUB retires for insert
- compiler: `insert Class { ... }` is a typed Ast.Insert in statement
  AND expression position, sharing the ctor literal's field grammar;
  typechecked with the ctor's omittable rule; result = the row id (Int)
- owner pass: the engine copies at the row API, so an insert BORROWS
  its field values -- no transfer, no E304; node is trap-capable and
  carries a live-mask drop entry like DbStub did
- emit: builtin 61 window = class-id const + one slot per DECLARED
  field in declaration order; omitted defaults emitted, omitted ?scalar
  gets WO_NIL_SCALAR, other omitted optionals the zero word; fresh
  argument values reaped after (the push/set copy semantics)
- runtime: database/src/db.c executes via the choke-point row API;
  rt.db/rt.wal opaque handles on wo_rt; WO_DATA=<dir> = replay
  <dir>/shard-0.wal at boot + commit-before-ack per statement (the
  builtin's return IS the ack until iteration 8 ticks); failed commit
  un-applies the row and traps WO_T_IO; loader validates the class-id
  slot (variable window documented in wob.h + format doc)
- the promised diff: trap/pricing-set-price-db-stub is now
  run/pricing-set-price-insert printing engine-allocated ids;
  durability smoke prints 1,2 then 3,4 across two WO_DATA runs
- old "bare insert is an Ident" unit test rewritten to the new
  contract; runner's loader mirror accepts id 61; goldens re-blessed
- gates: oop-accept ALL CRITERIA MET, oop-e2e 71/0, woc-test 566/0,
  wovm-test green, log-watcher 7/0

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-15 11:15:06 +02:00
c3741262cb feat(database): typed WAL + boot replay (iteration 9, Task 2)
- database/src/wal.{c,h}: framed records len|crc32|payload|mark
  ("WOL1" written last -- no mark, no record), typed-row payloads
  walking the class-table kinds (nested records, containers, nil
  encodings), little-endian like the loader
- commit order verbatim from the shipped phase-D pattern: RAM apply,
  stage, ONE pwrite + ONE fdatasync for the batch, ack after -- group
  commit is everything staged riding one sync
- replay decodes straight into engine-owned values (no VM at boot)
  and re-enters rows through the choke-point row API, so Task 4's
  indexes will rebuild for free; next_id advances past replayed ids
  this shard owns (wo_row_create_raw)
- torn tail = short/CRC-fail/no-mark/zero-len: intact prefix applies,
  tear dropped whole, wo_wal_open positions AT the tear so the next
  commit overwrites it; CRC-valid-but-undecodable = corruption, loud
- wo_wal_check: offline oracle, no engine needed -- the crash
  battery's verifier
- test_wal 90/0 ASan+UBSan incl. five crash-battery rounds (fork,
  insert/commit/ack-over-pipe, SIGKILL mid-stream: zero acked-but-
  missing, zero acked-but-wrong); all runtime suites green, oop-e2e
  71/0; binding doc WAL section + CODE-LOGIC + plan Task 2 checked

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-15 11:01:41 +02:00
936bd14bff feat(database): class-shaped row storage (iteration 9, Task 1)
- database/src/table.{c,h}: per-shard per-class slabs (256 rows,
  malloc'd, never moved -- row addresses stable for 9b's row views),
  occupancy bitmap, LIFO slot reuse, open-addressing id hash with
  tombstones (ids never 0, never reused)
- field encoding walks the same .wob class-table kinds the VM walks:
  scalars raw (WO_NIL_SCALAR passes through), Texts copied to db_text,
  owned objects flattened recursively to db_rec, containers
  element-wise; GCREF refused at encode (the GC bulkhead, defensively)
- two one-way copy gates: insert copies VM values in, read allocates
  fresh VM values out -- no VM pointer in a slab, no slab pointer in
  the VM, proven by mutating originals after insert
- id discipline: per table per shard, S+1 step N; owner = (id-1) % N;
  N-parametric, runs at N=1 until iteration 8, tested at N=3
- choke points: wo_row_insert/wo_row_remove carry the INDEX HOOK
  sites Task 4 attaches to; nothing else mutates storage
- runtime/Makefile links database/src into every wovm + test binary
- test_table 827/0 ASan+UBSan; oop-e2e 71/0; log-watcher 7/0;
  binding doc docs/plan/oop-vm/04-db-binding.md; CODE-LOGIC.md beside
  the code; plan Task 1 checked off

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-15 10:54:24 +02:00