Commit graph

14 commits

Author SHA1 Message Date
7ad52937b2 fix(db2-keys): delete on a keys-resident table was memory corruption
- wo_row_remove read the id map's value as a slot, but on a keys table
  that value is a LOG OFFSET (hput(t, id, wal_off + 1)). slot_row does
  no bounds check, so a delete indexed t->slabs[] with a byte offset and
  then called db_val_free on whatever it landed on — arbitrary frees,
  not a wrong answer
- keys tables now take their own arm: no slab slot, no bitmap bit, no
  free-list entry to return. The index hook needs the row's values, so
  the row is borrowed from the log for exactly that long
- wo_row_ptr carried the same trap and is public. It cannot refuse keys
  tables outright (insert legitimately calls it while the map still
  holds a slot), so it now detects the offset case — index past the
  slabs, or bitmap bit clear — and returns NULL. Callers all handle NULL
- test_keys_resident_delete pins it; it SEGVs against the old code,
  verified by reverting the fix rather than assumed
- found while auditing every hget() reader before narrowing the loader
  refusal to allow benchmarking. The refusal was justified in the docs
  by "updates are unimplemented" while actually standing in front of
  this too: a guard whose stated reason is narrower than its real one
  gets removed by someone who believes the stated reason

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
(cherry picked from commit 76b8fd944af9ed062467bdf9ab93c2e96dd198cf)
2026-08-30 20:37:27 +02:00
533fc5294f feat(db2-keys): rewire remaining readers, survive compaction
- wo_row_read and the @unique shadow probe go through borrow/release;
  release runs before every exit, including wo_row_read's early return
- updates on a keys table refused explicitly in wo_row_update_field and
  the slot variant: no slab slot to mutate, and writing the borrow's
  scratch would discard the write silently. Needs read-modify-append
- compaction walked the bitmap, which a keys row has no bit in — every
  such row would have been dropped from the new log. Now walks
  wo_row_next_id and re-points each row to where it lands
- moves records byte-for-byte (copy_record) rather than decoding: a
  borrowed row holds VM values, enc_val expects engine values, and ASan
  caught that mismatch as a 4294967292-byte memcpy
- wo_row_set_offset updates a value in place and never rehashes, so a
  wo_row_next_id cursor stays valid while compaction re-points
- a compaction that fails after moving rows is fatal: the map would name
  an unlinked temp file, and the intact log replays correctly
- test_keys_resident_survives_compaction pins both failure modes; rows
  rewrite in hash order so offsets really move
- loader still refuses resident: keys — updates are not implemented

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
(cherry picked from commit f606fc9b76d983cac2b348f03f7d4f01433cd905)
2026-08-30 20:36:31 +02:00
636f36b0f6 feat(db2-keys): the query paths read through the iterator and borrow
databasev2 2, task 5d. Every reader in db.c now works for both backings.

The measured problem: a keys-resident table's bitmap is EMPTY by construction
(its payloads live in the log), so all four bitmap walks would have silently
returned no rows — a query over such a table would find nothing, with no error.

- wo_row_next_id: one iterator, two backings. Keys tables walk the id map;
  resident tables keep walking the BITMAP deliberately, because the id map
  holds the same set in hash order and switching would reorder the results of
  every unordered query in the repo. No behaviour change where none was needed
- the three id-collecting scans move onto it. They only ever collected ids
  (the 9b cursor-stability rule materialises the list up front), so they needed
  no row access at all — which is why this was far smaller than the plan feared
- the two filtered scans borrow, compare, and RELEASE BEFORE any exit. The
  scratch is per-table, so a borrow leaked past a `return` or `break` would
  make the next borrow on that table fail as a nested one. That is a real
  hazard, not a hypothetical: the request-path GET_FIELD borrowed and then
  `break`ed without releasing until this commit
- point reads decode or clone BEFORE releasing, because a keys-resident row's
  slots point into the scratch that release frees

Verified: just wovm-test — 36 suites 0 fail.

Still to do in 5d: table.c's unique shadow and its three remaining wo_row_ptr
sites, wal.c's append encode, and compaction's own walk — which is where the
recorded `resident: keys` offset obligation has to be honoured. The loader
refusal stays until all of it lands.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
(cherry picked from commit 0c97fa48d3e31ea9eea3287d9c23ed29f0260488)
2026-08-30 20:36:31 +02:00
7cc80405dd feat(db2-keys): storage — drop the payload, read it back from the log
databasev2 2, task 5c step 2. The storage half the accessor was left waiting
for. Not yet wired into insert (that and 5d remain), and the loader still
refuses `resident: keys`, so nothing is exposed to a program yet.

- wo_db gains an `rt` back-pointer, set in main.c beside VM.rt.db. wo_rt
  already carries `db` and `wal` as opaque handles, so this closes the loop
  and a borrow can reach the log WITHOUT threading a wal pointer through
  eleven call sites — which is the whole reason 5c is one accessor
- wo_row_drop_payload: the operation the plan recorded as MISSING. Frees the
  slot and its engine-owned values, then re-points the id map at the record's
  log offset (off + 1, reusing the same 0-is-empty trick as slot + 1). It
  deliberately does NOT touch the secondary indexes (they store row ids, so
  they stay correct), does NOT decrement count (the row is still live, only
  its backing moved), and does NOT remove the id (that is how it is found)
- wo_row_borrow materialises for a keys table: reads the offset from the id
  map, calls 5b's wo_wal_read_row_at into the per-table scratch, and checks
  the record actually holds the expected class and id — a compaction that
  moved records without rebuilding the map lands exactly there, which is the
  obligation recorded at wo_wal_compact
- fully-resident tables keep today's path and pay one predicate

A REAL BUG, exposed the first time the path was used: wo_row_release freed the
materialised values with the ENGINE's allocator. They are VM values —
wo_wal_read_row_at is the out-gate and always copies — so ASan reported a
bad-free immediately. It now drops them through the runtime. That stub was
written in 5c step 1 for a path that did not exist yet.

Recorded while implementing: wo_wal_next_offset's contract says to trust an
offset "only after the matching commit returns 0". Group commit (databasev2 4)
defers that barrier to the drain, so db.c can no longer check inline — but part
A also made a failed commit FATAL, so no execution can record an offset whose
record never became durable. Same guarantee, different mechanism.

Test: a heap-valued row is inserted, committed, has its payload dropped, and is
read back out of the log with its Text intact; count is unchanged (still live);
and a second borrow succeeds, which fails if release did not clear the scratch.

Verified: just wovm-test — 36 suites 0 fail, test_wal 4273 pass, cli_smoke OK.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
(cherry picked from commit 125bd09218d616b2a16b140de770d3f38b45f0ac)
2026-08-30 20:36:31 +02:00
f72b3310a8 refactor(db): wo_row_borrow/wo_row_release — one read path for both residencies
Task 5c step 1 of docs/superpowers/plans/2026-08-26-table-residency.md, as a
PURE REFACTOR: no storage change, no keys-table anywhere. Borrow is wo_row_ptr
plus a seam, every release is a no-op. Provable on its own before the storage
change it exists to enable.

DESIGN SETTLED BY READING THE STRUCTURES, and both answers make 5c smaller:

- the id hash needs NO new storage. `hvals` is already uint64 holding
  slot+1 with 0 = empty (table.h), so offset+1 fits the same field, and the
  interpretation is per-table because a table is wholly `all` or wholly
  `keys`. No parallel map
- secondary indexes need NO change. `db_ibucket.ids` stores row IDS, not slot
  indices, and table.c resolves them through the id hash. I had told the
  developer these pointed at slab slots — that was wrong, and it is why this
  is one shared accessor rather than 11 rewrites
- the real coupling is the unique shadow: idx_add_row and
  row_apply_field_slot both FETCH the conflicting row and compare columns.
  Both now borrow/release, so a keys-table's non-resident conflict will be
  found rather than silently skipped — a unique check that only examines
  resident rows is a correctness hole, not a limitation

The scratch lives on `db_table`, not on the stack and not per call. Per call
would allocate once per candidate inside a bucket loop, turning an O(1) probe
into an allocation storm; a stack buffer is unsafe because the loader bounds
field_cnt at 65535 (loader.c:189), so the worst case is ~512 KB. It is safe
per-table because the store is single-writer, and a `busy` flag is there to
catch a nested borrow rather than let it alias silently. Freed in
table_destroy.

Gates: all 18 runtime suites 0 fail under ASan+UBSan (test_table 856/0,
test_wal 3654/0), oop-e2e 119/0, residency 8/0, employee 8/0, db-actor 8/0.
And the pure-refactor proof the plan asked for: `db-bench --quick` 85/0, every
resident read/query/seed/write floor held — a refactor that moves a number is
not a refactor.

Remaining wo_row_ptr sites for 5d: 6 in table.c, 2 in db.c, 2 in wal.c.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-27 16:33:36 +02:00
881b90e9c7 feat: O(1) read path — index probe wired end to end
- engine: wo_idx_probe answers single-column equality from the index
  hash buckets (idx_hash_key1 reproduces idx_hash bit for bit; verify
  compares exactly as the slab walk did, so results identical);
  composite indexes keep the walk; both executors wired (local + DB
  actor RPC)
- compiler: probe_key_of_where lowers "var.col == key" on an indexed
  column to DB_PROBE; all where guards still run (guard stays the
  final arbiter); keys = ident/int-literal only; Float/Bytes excluded
  (engine raw-eq narrower than VM float-eq)
- measured: reads 1.3k -> 1.3M ops/s, p50 600us -> 1us (~x850);
  query x830; mixread 1.3k -> 89k s1, 21 -> ~1.9k sN
- gate policy moved into the driver (tolerance_for: refresh-proof);
  latency floors max(4x,100us); quick mode skips poll-bound mix
  floors; both tolerance classes proven to bite
- proof: test_table wo_idx_probe suite (RED first), corpus
  query-index-probe 105/0, full battery green, TSan clean, two
  campaigns pass the refreshed baseline

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-22 16:44:56 +02:00
07c1f7b257 feat: arc stage 3 T7 — transparent DB actor (WO_T_DB hole closed)
- worker DB builtins marshal to shard 0: requester-side slot encode
  (VM heaps never read cross-shard), owner executes serialized in
  adopt, reply unparks via new WO_PARK_INBOX park + envelope 3/4
- engine gains thread-agnostic slot entry points (insert_slots,
  update_field_slot, val_encode/clone, wo_db_exec_req); traps and
  messages byte-identical to the local path
- main.c: engine + replay boot BEFORE shards spawn; workers assert
  rt.db/rt.wal NULL; busy shard adopts inbox once per slice
- latent stage-1 bug fixed: shared io_uring params static raced by
  lazy worker init lost park wakes (~1/20 hangs); params per-vm,
  short submit now fails loud
- new sample docs/examples/db-actor + just db-actor gate 8/0 (multi
  x3, uring/epoll forced, single byte-exact, WAL replay pair);
  ASan+TSan 6/6; full battery green

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-21 13:10:26 +02:00
d24c705860 feat: iterations 19 + 17 — Float/Bytes scalars (.wob v5), library kind + internal/
- Float full stack: literals (fraction/exponent; `0..10` still a range), f64
  opcodes 34-41, @table column, WAL bit-exact replay, json fractions in and
  shortest-round-trip out. IEEE-quiet — FDIV never traps where DIV does.
- Bytes: a wo_str with its own class id, so alloc/free/copy are shared but no
  Text builtin accepts one; len/at/slice/eq/concat, base64 both ways, json
  boundary as base64; TEXT_COPY preserves the kind.
- No implicit Int/Float mixing (WO-E201 in the typechecker, not the emitter,
  which picks the opcode from one side and would misread the other).
- One IEEE deviation: float_cmp total order (NaN last, -0.0 == +0.0) for
  indexes and order-by, keys canonicalized to match. `?Float` nil is a
  reserved quiet NaN — the zero word is +0.0, WO_NIL_SCALAR's bits are -2.0.
- Renderer prefers fixed over exponential in 1e-6..1e21: pure shortest makes
  a price of 900.0 read `9e+02`. One renderer for interp/json/float_to_text.
- Fixed en route: lexer double-counted the leading digit; is_scalar_shaped
  took Float/Bytes as Int-shaped; Bytes ownership needed a shared heap-scalar
  predicate or temps never dropped; order-by bit-compared negatives backwards.
- Iteration 17: `kind = "library"` (absent = program; bad value = WO-E109),
  entry-less check mode retiring the `--emit` workaround, Go's `internal/` as
  WO-E108 at the consumer's `use`. Driver-only; VM/.wob/GC untouched.
- Framework reorg: internal/{parse,serve}.wo; http/form.wo split out to keep
  media_type/form_values public (parse.wo had grown public surface).
- Docs: link audit (97 -> 88 broken, conflict markers resolved, 2 duplicate
  stories removed), 00-code-review verified 26/27, iterations re-sequenced.
- Also carries the pre-staged pub(read)/using/#if work from the index.
- Gates: corpus 103/0, test_wal 156/0, web-app 26/0, oop-accept ALL MET.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-20 19:24:15 +02:00
2677c1457e feat: FK restrict on delete + employee sample runs; group-by parked (9b)
- FK restrict: deleting a row a non-nullable `ref` still points at traps
  WO_T_FK (11), catchable. The compiler now records a `ref` field's
  target class in the class-table field_class metadata; the engine
  (wo_row_has_referrers) scans referencing scalar columns before a
  delete. Correctness-first full scan; the backlink-index optimization
  is recorded for later
- docs/examples/employee now COMPILES AND RUNS all six modes against a
  WAL-durable database: seed (+@unique trap across restart), report
  (per-dept aggregates + payroll), staff (unique probe + backlink +
  ref nav), raise (update-through-row), drop (FK restrict), and
  persistence via replay
- group-by SYNTAX parked to a future iteration (user decision): the
  report mode is hand-rolled from the shipped primitives meanwhile
  (same numbers). "table relations and FK" is complete
- scripts/employee-accept.sh (8 checks) + a `just employee` module;
  manifest parser tolerates iteration 9c's [share]/[[share.clients]]
  sections so `woc .` builds the sample on this branch
- fixtures trap/db-fk-restrict (code 11) + run/db-fk-restrict-catch;
  oop-e2e 79/0, woc-test 566/0, 15 runtime suites, log-watcher 7/0,
  employee-accept 8/0
- 9b story + status board updated

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-16 18:37:23 +02:00
048c242a01 feat(database): query-read engine builtins (iteration 9b foundation)
- DB_SCAN(64): class -> multi<Int> of every row id, materialized up
  front (the 9b cursor-stability rule: the loop body point-reads, so a
  row updated mid-loop cannot disturb iteration)
- DB_GET_FIELD(65): class,id,field -> the field decoded to a fresh VM
  value (the out-gate copy); a table-class value IS its row id at
  runtime, so this is how a compiled query reads a column, and a ref
  field decodes to the target id for navigation
- DB_PROBE(66): class,index,key -> multi<Int> of ids whose first
  indexed column equals key (backlink + indexed where)
- wo_val_decode_vm wrapper exposed; dispatch range 61..66, loader
  arities, runner mirror updated
- 15 runtime suites green, oop-e2e 73/0

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-15 22:19:13 +02:00
44c77ff5bf feat: update-point + delete engine half (iteration 9, Task 5 engine)
- wo_row_update_field: encode new value, unique re-check against a
  shadow BEFORE any mutation (violating update leaves the row
  untouched, DB_ERR_UNIQUE), index entries moved old-hash -> new-hash,
  old engine value freed; proven by test_table (unique refusal keeps
  the row, released key becomes insertable)
- WAL UPDATE record: full-row re-log, replay = replace (remove +
  re-create same id); prefix/suffix delta recorded as later
  optimization; test_wal replays insert+update to the updated state
- builtins 62 DB_UPDATE_FIELD (cid,id,field,value) and 63 DB_DELETE
  (cid,id), commit-before-ack like insert, WO_T_UNIQUE/WO_T_DB/WO_T_IO
  mapping; dispatch range 61..63; loader arities; runner mirror
- plan Task 5 marked superseded-in-part with the recorded deviation:
  the language surface (reads, queries, row views, delete statement)
  is 9b's, where the comprehension design put it -- no interim brace-
  select grammar to retire later
- gates: test_table 839/0, test_wal 102/0, 15 suites, oop-e2e 73/0

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-15 13:01:05 +02:00
6f2f9b6f0a feat: secondary indexes + @unique trap (iteration 9, Task 4; wob v3)
- .wob v3: class records carry an index tail (flags bit0 = unique,
  col_cnt, columns) -- @table(index:[a,b]) entries plus one unique
  single-column entry per @unique field; loader validates columns in
  range and scalar/Text-kinded; emitter validates the declarations
  (unknown column, un-indexable kind => diagnostic)
- engine: db_index hash multimap per table, built from the class
  table at first touch, maintained ONLY inside wo_row_insert/
  wo_row_remove; unique checks re-compare actual column values (a
  hash is a hint); replay re-indexes via wo_row_raw_commit AFTER
  slots are filled, so recovered tables carry their indexes
- WO_T_UNIQUE = 10; a violating insert is un-applied whole (bitmap,
  hash, count, and the never-observable id reclaimed) and traps
  catchably -- the employee SEED-DUP pattern
- wo_row_insert gains err_kind so db.c maps UNIQUE/OOM/other to the
  right trap; test images and the runner's loader mirror speak v3
- fixtures: trap/db-unique-violation (code 10 exact) and
  run/db-unique-catch (catchable dup, composite index accepts
  duplicates, next id dense after a refusal)
- gates: oop-e2e 73/0, all 15 runtime suites, woc-test green,
  log-watcher 7/0

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-15 12:57:32 +02:00
c3741262cb feat(database): typed WAL + boot replay (iteration 9, Task 2)
- database/src/wal.{c,h}: framed records len|crc32|payload|mark
  ("WOL1" written last -- no mark, no record), typed-row payloads
  walking the class-table kinds (nested records, containers, nil
  encodings), little-endian like the loader
- commit order verbatim from the shipped phase-D pattern: RAM apply,
  stage, ONE pwrite + ONE fdatasync for the batch, ack after -- group
  commit is everything staged riding one sync
- replay decodes straight into engine-owned values (no VM at boot)
  and re-enters rows through the choke-point row API, so Task 4's
  indexes will rebuild for free; next_id advances past replayed ids
  this shard owns (wo_row_create_raw)
- torn tail = short/CRC-fail/no-mark/zero-len: intact prefix applies,
  tear dropped whole, wo_wal_open positions AT the tear so the next
  commit overwrites it; CRC-valid-but-undecodable = corruption, loud
- wo_wal_check: offline oracle, no engine needed -- the crash
  battery's verifier
- test_wal 90/0 ASan+UBSan incl. five crash-battery rounds (fork,
  insert/commit/ack-over-pipe, SIGKILL mid-stream: zero acked-but-
  missing, zero acked-but-wrong); all runtime suites green, oop-e2e
  71/0; binding doc WAL section + CODE-LOGIC + plan Task 2 checked

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-15 11:01:41 +02:00
936bd14bff feat(database): class-shaped row storage (iteration 9, Task 1)
- database/src/table.{c,h}: per-shard per-class slabs (256 rows,
  malloc'd, never moved -- row addresses stable for 9b's row views),
  occupancy bitmap, LIFO slot reuse, open-addressing id hash with
  tombstones (ids never 0, never reused)
- field encoding walks the same .wob class-table kinds the VM walks:
  scalars raw (WO_NIL_SCALAR passes through), Texts copied to db_text,
  owned objects flattened recursively to db_rec, containers
  element-wise; GCREF refused at encode (the GC bulkhead, defensively)
- two one-way copy gates: insert copies VM values in, read allocates
  fresh VM values out -- no VM pointer in a slab, no slab pointer in
  the VM, proven by mutating originals after insert
- id discipline: per table per shard, S+1 step N; owner = (id-1) % N;
  N-parametric, runs at N=1 until iteration 8, tested at N=3
- choke points: wo_row_insert/wo_row_remove carry the INDEX HOOK
  sites Task 4 attaches to; nothing else mutates storage
- runtime/Makefile links database/src into every wovm + test binary
- test_table 827/0 ASan+UBSan; oop-e2e 71/0; log-watcher 7/0;
  binding doc docs/plan/oop-vm/04-db-binding.md; CODE-LOGIC.md beside
  the code; plan Task 1 checked off

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-15 10:54:24 +02:00