Commit graph

2 commits

Author SHA1 Message Date
4af1e8bcdd fix(chat gate): every leg starts its own server — and it found a real bug
Gate defects, all measured:

- fd check was core-count dependent: `fds_before + 8` read LAZY per-shard
  init as a leak. Shards init on first fiber, each taking one io_uring +
  one eventfd, capped at nproc; on 20 cores the first wave legitimately
  adds 18. Measured 26 -> 44 after 20 clients, still 44 after 40 more.
  Replaced with the invariant the check is for: a second wave must not
  raise the count. Core-count independent, and catches a slow leak that
  any fixed slack would hide
- a failed leg ORPHANED its server: drain inherited $SRV from the soak
  leg, so its python died on int("") and the soak server was never
  killed — its listener then broke the next run's soak on the same port.
  drain now starts its own server; cleanup kills every server a run
  started, matched on the run's unique temp dir
- two legs the plan requires were missing: WO_SHARDS=1 (the single-shard
  control) and WO_MAILBOX=8 (drop-slow-member backpressure). Both added,
  both green. The mailbox leg shrinks the slow client's SO_RCVBUF so it
  needs no sleeps
- chat adopted the porch naming (use porch/..., [deps] key) after the
  rename landed on master

Decoupling the legs exposed a REAL drain bug, traced and documented in
docs/2026-08-27-chat-drain-finding.md, NOT fixed here:

- on a FRESH server the SIGTERM drain is flaky: 5 of 16 runs left a
  client at EOF with no close frame and no diagnostic
- traced: main -> Registry -> Room -> Writer. Registry runs (diag
  confirms), the Room NEVER processes its shutdown message, so the
  Writer's close branch never runs. Clients that do get a frame are
  saved by their own Reader seeing env.stopping()
- ruled out: the spin budget (a 1s wall-clock deadline still failed 2 of
  12 — reverted, it fixed nothing and cost 1s per shutdown),
  dummy_writer() spawning during shutdown, and write failure
- the fix is an engine guarantee — a send issued before the stop flag is
  delivered — which belongs to the actor lifecycle, not a spin count

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-27 23:27:46 +02:00
735fd270db feat: chat sample + gate (T8/T9, IN PROGRESS) + stop-drain semantics
- docs/examples/chat: registry (call consumer) / room / reader+writer
  actor pair per connection over ws_accept + wsframe; presence,
  broadcast, cross-room isolation, mailbox-full = drop-from-room;
  reader tail sends hardened (a full writer no longer orphans the fd)
- RUNTIME SEMANTICS CHANGE (the drain): SIGTERM no longer kills parked
  fibers from outside — the plane WAKES them and each wait RESOLVES
  (deadline'd waits answer their timeout result, sleeps return early,
  plain waits answer WO_SYS_STOPPED and unwind THAT fiber alone; main's
  STOPPED still ends the program). Workers keep adopting their inboxes
  after stop until eng_shutdown. This is what lets a program drain:
  chat's close frames now reach clients (byte-verified 0x88), then
  main returns and the reap runs
- also: SIGPIPE ignored process-wide (EPIPE trap instead of death);
  two-phase engine teardown (real drops while arenas+routing live,
  settle passes for routed frees) — fixes the registry-map leak and
  the drain UAF ASan found
- gate scripts/chat-accept.sh + just chat: handshake independently
  verified, functional matrix on BOTH backends, 1k-hot-room soak
  (1000/1000 in ~35ms), drain close-frames, SIGTERM exit 0, ASan leg
  clean. OPEN: soak-fds check (18 fds settle slower than the window)
  + full battery after the semantics change — NOT yet run
- committed for manual testing at the user's request

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-23 09:53:21 +02:00