- residency-accept.sh section 7, six checks: the refusal names the class
and all three ways forward (exit 2); WO_EPHEMERAL=1 runs from RAM with
the boot notice and a write round-trips; WO_EPHEMERAL with WO_DATA
refuses; WO_EPHEMERAL=2 refuses naming the accepted value; keys-resident
still refuses under the hatch; a plain class (the corpus `methods`
fixture) runs with no WO_DATA, rc 0, nothing on stderr
- blast radius measured gate by gate — each run without the export first,
kept only where the program refused: oop-e2e (fixtures declare tables);
db-bench.py's ram/msgrate/growth/randread legs (the durable legs drop it,
so a WO_DATA in the caller's shell now refuses loudly instead of silently
turning a RAM leg durable); db-actor per run (its restart pair sets
WO_DATA); chat (porch's store declares RateLimitCounter default-durable —
a library's table binds the consumer); wmux client legs (same image as
the server, no WO_DATA; servers and the WO_DATA-carrying r11cli `env -u`)
- byte-exact compares (db-actor single-shard, wmux client) drop the one
notice line; fibers, subprocess, log-watcher declare no table — untouched
- db-bench.py ceiling note: the checked refusal is databasev2 5's now
- residency 32/0, oop-e2e 131/0 with this tree
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
(cherry picked from commit 4553ca15da235c155b7ad31bbb077c3ad8e88fee)
Root cause (decision 1): cross-shard send/call/monitor pointer-shared the
message into the receiver's shard (e->payload = msg_val), so a worker read and
eventually dropped an object living in the sender's arena — a double free, then
a class-0 forge, then a modulo self-route livelock, all downstream of that one
broken invariant ("VM heaps are never read cross-shard", which wo_db_rpc keeps).
- actor_marshal: the sender encodes the message into an arena-independent neutral
form (wo_db_val_encode, the same marshal wo_db_rpc uses) and drops its own
original — no pointer crosses an arena boundary, so the double-free class is
gone by construction. actor_unmarshal rebuilds it in the receiver's arena
(wo_val_decode_vm) and frees the neutral. Applied to the 4 cross-shard
producers (send x2, call, monitor) + the 3 consumers (kinds 0/5/7). Same-shard
paths untouched (the WO_SHARDS=1 fast path never failed). Call replies are
scalars by contract, so kind 6 needs no marshal.
- eng_settle_inboxes: undrained kind-0/5/7 payloads at teardown are the neutral
form now — free with wo_db_val_free, not wo_drop_obj (caught by ASan mid-fix).
- decision 2: wo_route_free traps a shard_id >= nshards header (a corrupt/freed
block) instead of self-routing it into the settle livelock.
- proof: tests/regress/lang-41/cross-shard-marshal.wo (a multi<Text> sent +
called cross-shard, both sides drop) — clean 12x/5x under WO_SHARDS=4 + ASan;
shard-settle repro still clean 8x; full runtime suite 0 fail (same-shard
byte-unchanged). `just db-actor` extended with the new fixture.
- unblocks porch 9. Follow-ups: poison-on-free (decision 3), corpus fixture (4).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
(cherry picked from commit 63065ff75799f7f43b2bce6de61e77856799566f)
- root cause: a worker's runtime is initialised lazily on first fiber
adoption, and rt.shard_id is stamped only there — but INBOX_READY[i]
is set at thread creation. A shard that never adopts is still settled
at shutdown, carrying rt.shard_id 0 from the memset
- it then impersonated shard 0: wo_drop_obj saw 0 == 0 for anything the
primary allocated, took the "we are home" branch instead of routing,
and called class_free against rt->classes, which lazy init never
filled. &rt->classes[class_id] off a NULL base is the faulting read
- fix: stamp the runtime's real identity at thread creation. An
uninitialised shard owns nothing, so its true id makes every payload
correctly foreign and routes it to an owner that can free it
- ASan could not name this: the arena is one hand-managed malloc block,
so intra-arena reuse is invisible and it surfaces as a bare SEGV
- pinned by tests/regress/lang-41, driven from db-actor-accept. Needs
multiple shards (the corpus runner pins WO_SHARDS=1) and the ASan
build. SEGVs twice per run unfixed, clean fixed
- the HANG is a separate defect and is NOT fixed: with this in place the
harness stops losing whole sections, but idempotent-stop-2 still
fires ~1 run in 6. The story records where to look
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
(cherry picked from commit 9dca0b4b4727b976d326b29cb4c6522b62d48a73)
- worker DB builtins marshal to shard 0: requester-side slot encode
(VM heaps never read cross-shard), owner executes serialized in
adopt, reply unparks via new WO_PARK_INBOX park + envelope 3/4
- engine gains thread-agnostic slot entry points (insert_slots,
update_field_slot, val_encode/clone, wo_db_exec_req); traps and
messages byte-identical to the local path
- main.c: engine + replay boot BEFORE shards spawn; workers assert
rt.db/rt.wal NULL; busy shard adopts inbox once per slice
- latent stage-1 bug fixed: shared io_uring params static raced by
lazy worker init lost park wakes (~1/20 hangs); params per-vm,
short submit now fails loud
- new sample docs/examples/db-actor + just db-actor gate 8/0 (multi
x3, uring/epoll forced, single byte-exact, WAL replay pair);
ASan+TSan 6/6; full battery green
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>