Gate defects, all measured:
- fd check was core-count dependent: `fds_before + 8` read LAZY per-shard
init as a leak. Shards init on first fiber, each taking one io_uring +
one eventfd, capped at nproc; on 20 cores the first wave legitimately
adds 18. Measured 26 -> 44 after 20 clients, still 44 after 40 more.
Replaced with the invariant the check is for: a second wave must not
raise the count. Core-count independent, and catches a slow leak that
any fixed slack would hide
- a failed leg ORPHANED its server: drain inherited $SRV from the soak
leg, so its python died on int("") and the soak server was never
killed — its listener then broke the next run's soak on the same port.
drain now starts its own server; cleanup kills every server a run
started, matched on the run's unique temp dir
- two legs the plan requires were missing: WO_SHARDS=1 (the single-shard
control) and WO_MAILBOX=8 (drop-slow-member backpressure). Both added,
both green. The mailbox leg shrinks the slow client's SO_RCVBUF so it
needs no sleeps
- chat adopted the porch naming (use porch/..., [deps] key) after the
rename landed on master
Decoupling the legs exposed a REAL drain bug, traced and documented in
docs/2026-08-27-chat-drain-finding.md, NOT fixed here:
- on a FRESH server the SIGTERM drain is flaky: 5 of 16 runs left a
client at EOF with no close frame and no diagnostic
- traced: main -> Registry -> Room -> Writer. Registry runs (diag
confirms), the Room NEVER processes its shutdown message, so the
Writer's close branch never runs. Clients that do get a frame are
saved by their own Reader seeing env.stopping()
- ruled out: the spin budget (a 1s wall-clock deadline still failed 2 of
12 — reverted, it fixed nothing and cost 1s per shutdown),
dummy_writer() spawning during shutdown, and write failure
- the fix is an engine guarantee — a send issued before the stop flag is
delivered — which belongs to the actor lifecycle, not a spin count
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
- docs/examples/chat: registry (call consumer) / room / reader+writer
actor pair per connection over ws_accept + wsframe; presence,
broadcast, cross-room isolation, mailbox-full = drop-from-room;
reader tail sends hardened (a full writer no longer orphans the fd)
- RUNTIME SEMANTICS CHANGE (the drain): SIGTERM no longer kills parked
fibers from outside — the plane WAKES them and each wait RESOLVES
(deadline'd waits answer their timeout result, sleeps return early,
plain waits answer WO_SYS_STOPPED and unwind THAT fiber alone; main's
STOPPED still ends the program). Workers keep adopting their inboxes
after stop until eng_shutdown. This is what lets a program drain:
chat's close frames now reach clients (byte-verified 0x88), then
main returns and the reap runs
- also: SIGPIPE ignored process-wide (EPIPE trap instead of death);
two-phase engine teardown (real drops while arenas+routing live,
settle passes for routed frees) — fixes the registry-map leak and
the drain UAF ASan found
- gate scripts/chat-accept.sh + just chat: handshake independently
verified, functional matrix on BOTH backends, 1k-hot-room soak
(1000/1000 in ~35ms), drain close-frames, SIGTERM exit 0, ASan leg
clean. OPEN: soak-fds check (18 fds settle slower than the window)
+ full battery after the semantics change — NOT yet run
- committed for manual testing at the user's request
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>