Commit graph

2 commits

Author SHA1 Message Date
00214bd68e fix(vm): marshal cross-shard actor messages (language 41) — the double free
Root cause (decision 1): cross-shard send/call/monitor pointer-shared the
message into the receiver's shard (e->payload = msg_val), so a worker read and
eventually dropped an object living in the sender's arena — a double free, then
a class-0 forge, then a modulo self-route livelock, all downstream of that one
broken invariant ("VM heaps are never read cross-shard", which wo_db_rpc keeps).

- actor_marshal: the sender encodes the message into an arena-independent neutral
  form (wo_db_val_encode, the same marshal wo_db_rpc uses) and drops its own
  original — no pointer crosses an arena boundary, so the double-free class is
  gone by construction. actor_unmarshal rebuilds it in the receiver's arena
  (wo_val_decode_vm) and frees the neutral. Applied to the 4 cross-shard
  producers (send x2, call, monitor) + the 3 consumers (kinds 0/5/7). Same-shard
  paths untouched (the WO_SHARDS=1 fast path never failed). Call replies are
  scalars by contract, so kind 6 needs no marshal.
- eng_settle_inboxes: undrained kind-0/5/7 payloads at teardown are the neutral
  form now — free with wo_db_val_free, not wo_drop_obj (caught by ASan mid-fix).
- decision 2: wo_route_free traps a shard_id >= nshards header (a corrupt/freed
  block) instead of self-routing it into the settle livelock.
- proof: tests/regress/lang-41/cross-shard-marshal.wo (a multi<Text> sent +
  called cross-shard, both sides drop) — clean 12x/5x under WO_SHARDS=4 + ASan;
  shard-settle repro still clean 8x; full runtime suite 0 fail (same-shard
  byte-unchanged). `just db-actor` extended with the new fixture.
- unblocks porch 9. Follow-ups: poison-on-free (decision 3), corpus fixture (4).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
(cherry picked from commit 63065ff75799f7f43b2bce6de61e77856799566f)
2026-09-15 01:16:13 +02:00
b9ce271b3d fix(lang41): an unadopted shard must not impersonate shard 0
- root cause: a worker's runtime is initialised lazily on first fiber
  adoption, and rt.shard_id is stamped only there — but INBOX_READY[i]
  is set at thread creation. A shard that never adopts is still settled
  at shutdown, carrying rt.shard_id 0 from the memset
- it then impersonated shard 0: wo_drop_obj saw 0 == 0 for anything the
  primary allocated, took the "we are home" branch instead of routing,
  and called class_free against rt->classes, which lazy init never
  filled. &rt->classes[class_id] off a NULL base is the faulting read
- fix: stamp the runtime's real identity at thread creation. An
  uninitialised shard owns nothing, so its true id makes every payload
  correctly foreign and routes it to an owner that can free it
- ASan could not name this: the arena is one hand-managed malloc block,
  so intra-arena reuse is invisible and it surfaces as a bare SEGV
- pinned by tests/regress/lang-41, driven from db-actor-accept. Needs
  multiple shards (the corpus runner pins WO_SHARDS=1) and the ASan
  build. SEGVs twice per run unfixed, clean fixed
- the HANG is a separate defect and is NOT fixed: with this in place the
  harness stops losing whole sections, but idempotent-stop-2 still
  fires ~1 run in 6. The story records where to look

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
(cherry picked from commit 9dca0b4b4727b976d326b29cb4c6522b62d48a73)
2026-09-15 01:15:30 +02:00