Commit graph

4 commits

Author SHA1 Message Date
41f48bba3f test(db2-ephemeral): gates opt into WO_EPHEMERAL=1 where a durable table is declared; residency section 7
- residency-accept.sh section 7, six checks: the refusal names the class
  and all three ways forward (exit 2); WO_EPHEMERAL=1 runs from RAM with
  the boot notice and a write round-trips; WO_EPHEMERAL with WO_DATA
  refuses; WO_EPHEMERAL=2 refuses naming the accepted value; keys-resident
  still refuses under the hatch; a plain class (the corpus `methods`
  fixture) runs with no WO_DATA, rc 0, nothing on stderr
- blast radius measured gate by gate — each run without the export first,
  kept only where the program refused: oop-e2e (fixtures declare tables);
  db-bench.py's ram/msgrate/growth/randread legs (the durable legs drop it,
  so a WO_DATA in the caller's shell now refuses loudly instead of silently
  turning a RAM leg durable); db-actor per run (its restart pair sets
  WO_DATA); chat (porch's store declares RateLimitCounter default-durable —
  a library's table binds the consumer); wmux client legs (same image as
  the server, no WO_DATA; servers and the WO_DATA-carrying r11cli `env -u`)
- byte-exact compares (db-actor single-shard, wmux client) drop the one
  notice line; fibers, subprocess, log-watcher declare no table — untouched
- db-bench.py ceiling note: the checked refusal is databasev2 5's now
- residency 32/0, oop-e2e 131/0 with this tree

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
(cherry picked from commit 4553ca15da235c155b7ad31bbb077c3ad8e88fee)
2026-09-15 01:16:24 +02:00
00214bd68e fix(vm): marshal cross-shard actor messages (language 41) — the double free
Root cause (decision 1): cross-shard send/call/monitor pointer-shared the
message into the receiver's shard (e->payload = msg_val), so a worker read and
eventually dropped an object living in the sender's arena — a double free, then
a class-0 forge, then a modulo self-route livelock, all downstream of that one
broken invariant ("VM heaps are never read cross-shard", which wo_db_rpc keeps).

- actor_marshal: the sender encodes the message into an arena-independent neutral
  form (wo_db_val_encode, the same marshal wo_db_rpc uses) and drops its own
  original — no pointer crosses an arena boundary, so the double-free class is
  gone by construction. actor_unmarshal rebuilds it in the receiver's arena
  (wo_val_decode_vm) and frees the neutral. Applied to the 4 cross-shard
  producers (send x2, call, monitor) + the 3 consumers (kinds 0/5/7). Same-shard
  paths untouched (the WO_SHARDS=1 fast path never failed). Call replies are
  scalars by contract, so kind 6 needs no marshal.
- eng_settle_inboxes: undrained kind-0/5/7 payloads at teardown are the neutral
  form now — free with wo_db_val_free, not wo_drop_obj (caught by ASan mid-fix).
- decision 2: wo_route_free traps a shard_id >= nshards header (a corrupt/freed
  block) instead of self-routing it into the settle livelock.
- proof: tests/regress/lang-41/cross-shard-marshal.wo (a multi<Text> sent +
  called cross-shard, both sides drop) — clean 12x/5x under WO_SHARDS=4 + ASan;
  shard-settle repro still clean 8x; full runtime suite 0 fail (same-shard
  byte-unchanged). `just db-actor` extended with the new fixture.
- unblocks porch 9. Follow-ups: poison-on-free (decision 3), corpus fixture (4).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
(cherry picked from commit 63065ff75799f7f43b2bce6de61e77856799566f)
2026-09-15 01:16:13 +02:00
b9ce271b3d fix(lang41): an unadopted shard must not impersonate shard 0
- root cause: a worker's runtime is initialised lazily on first fiber
  adoption, and rt.shard_id is stamped only there — but INBOX_READY[i]
  is set at thread creation. A shard that never adopts is still settled
  at shutdown, carrying rt.shard_id 0 from the memset
- it then impersonated shard 0: wo_drop_obj saw 0 == 0 for anything the
  primary allocated, took the "we are home" branch instead of routing,
  and called class_free against rt->classes, which lazy init never
  filled. &rt->classes[class_id] off a NULL base is the faulting read
- fix: stamp the runtime's real identity at thread creation. An
  uninitialised shard owns nothing, so its true id makes every payload
  correctly foreign and routes it to an owner that can free it
- ASan could not name this: the arena is one hand-managed malloc block,
  so intra-arena reuse is invisible and it surfaces as a bare SEGV
- pinned by tests/regress/lang-41, driven from db-actor-accept. Needs
  multiple shards (the corpus runner pins WO_SHARDS=1) and the ASan
  build. SEGVs twice per run unfixed, clean fixed
- the HANG is a separate defect and is NOT fixed: with this in place the
  harness stops losing whole sections, but idempotent-stop-2 still
  fires ~1 run in 6. The story records where to look

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
(cherry picked from commit 9dca0b4b4727b976d326b29cb4c6522b62d48a73)
2026-09-15 01:15:30 +02:00
07c1f7b257 feat: arc stage 3 T7 — transparent DB actor (WO_T_DB hole closed)
- worker DB builtins marshal to shard 0: requester-side slot encode
  (VM heaps never read cross-shard), owner executes serialized in
  adopt, reply unparks via new WO_PARK_INBOX park + envelope 3/4
- engine gains thread-agnostic slot entry points (insert_slots,
  update_field_slot, val_encode/clone, wo_db_exec_req); traps and
  messages byte-identical to the local path
- main.c: engine + replay boot BEFORE shards spawn; workers assert
  rt.db/rt.wal NULL; busy shard adopts inbox once per slice
- latent stage-1 bug fixed: shared io_uring params static raced by
  lazy worker init lost park wakes (~1/20 hangs); params per-vm,
  short submit now fails loud
- new sample docs/examples/db-actor + just db-actor gate 8/0 (multi
  x3, uring/epoll forced, single byte-exact, WAL replay pair);
  ASan+TSan 6/6; full battery green

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-21 13:10:26 +02:00