Root cause (decision 1): cross-shard send/call/monitor pointer-shared the
message into the receiver's shard (e->payload = msg_val), so a worker read and
eventually dropped an object living in the sender's arena — a double free, then
a class-0 forge, then a modulo self-route livelock, all downstream of that one
broken invariant ("VM heaps are never read cross-shard", which wo_db_rpc keeps).
- actor_marshal: the sender encodes the message into an arena-independent neutral
form (wo_db_val_encode, the same marshal wo_db_rpc uses) and drops its own
original — no pointer crosses an arena boundary, so the double-free class is
gone by construction. actor_unmarshal rebuilds it in the receiver's arena
(wo_val_decode_vm) and frees the neutral. Applied to the 4 cross-shard
producers (send x2, call, monitor) + the 3 consumers (kinds 0/5/7). Same-shard
paths untouched (the WO_SHARDS=1 fast path never failed). Call replies are
scalars by contract, so kind 6 needs no marshal.
- eng_settle_inboxes: undrained kind-0/5/7 payloads at teardown are the neutral
form now — free with wo_db_val_free, not wo_drop_obj (caught by ASan mid-fix).
- decision 2: wo_route_free traps a shard_id >= nshards header (a corrupt/freed
block) instead of self-routing it into the settle livelock.
- proof: tests/regress/lang-41/cross-shard-marshal.wo (a multi<Text> sent +
called cross-shard, both sides drop) — clean 12x/5x under WO_SHARDS=4 + ASan;
shard-settle repro still clean 8x; full runtime suite 0 fail (same-shard
byte-unchanged). `just db-actor` extended with the new fixture.
- unblocks porch 9. Follow-ups: poison-on-free (decision 3), corpus fixture (4).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
(cherry picked from commit 63065ff75799f7f43b2bce6de61e77856799566f)
- root cause: a worker's runtime is initialised lazily on first fiber
adoption, and rt.shard_id is stamped only there — but INBOX_READY[i]
is set at thread creation. A shard that never adopts is still settled
at shutdown, carrying rt.shard_id 0 from the memset
- it then impersonated shard 0: wo_drop_obj saw 0 == 0 for anything the
primary allocated, took the "we are home" branch instead of routing,
and called class_free against rt->classes, which lazy init never
filled. &rt->classes[class_id] off a NULL base is the faulting read
- fix: stamp the runtime's real identity at thread creation. An
uninitialised shard owns nothing, so its true id makes every payload
correctly foreign and routes it to an owner that can free it
- ASan could not name this: the arena is one hand-managed malloc block,
so intra-arena reuse is invisible and it surfaces as a bare SEGV
- pinned by tests/regress/lang-41, driven from db-actor-accept. Needs
multiple shards (the corpus runner pins WO_SHARDS=1) and the ASan
build. SEGVs twice per run unfixed, clean fixed
- the HANG is a separate defect and is NOT fixed: with this in place the
harness stops losing whole sections, but idempotent-stop-2 still
fires ~1 run in 6. The story records where to look
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
(cherry picked from commit 9dca0b4b4727b976d326b29cb4c6522b62d48a73)