writeonce/scripts
shoney.arickathil 4af1e8bcdd fix(chat gate): every leg starts its own server — and it found a real bug
Gate defects, all measured:

- fd check was core-count dependent: `fds_before + 8` read LAZY per-shard
  init as a leak. Shards init on first fiber, each taking one io_uring +
  one eventfd, capped at nproc; on 20 cores the first wave legitimately
  adds 18. Measured 26 -> 44 after 20 clients, still 44 after 40 more.
  Replaced with the invariant the check is for: a second wave must not
  raise the count. Core-count independent, and catches a slow leak that
  any fixed slack would hide
- a failed leg ORPHANED its server: drain inherited $SRV from the soak
  leg, so its python died on int("") and the soak server was never
  killed — its listener then broke the next run's soak on the same port.
  drain now starts its own server; cleanup kills every server a run
  started, matched on the run's unique temp dir
- two legs the plan requires were missing: WO_SHARDS=1 (the single-shard
  control) and WO_MAILBOX=8 (drop-slow-member backpressure). Both added,
  both green. The mailbox leg shrinks the slow client's SO_RCVBUF so it
  needs no sleeps
- chat adopted the porch naming (use porch/..., [deps] key) after the
  rename landed on master

Decoupling the legs exposed a REAL drain bug, traced and documented in
docs/2026-08-27-chat-drain-finding.md, NOT fixed here:

- on a FRESH server the SIGTERM drain is flaky: 5 of 16 runs left a
  client at EOF with no close frame and no diagnostic
- traced: main -> Registry -> Room -> Writer. Registry runs (diag
  confirms), the Room NEVER processes its shutdown message, so the
  Writer's close branch never runs. Clients that do get a frame are
  saved by their own Reader seeing env.stopping()
- ruled out: the spin budget (a 1s wall-clock deadline still failed 2 of
  12 — reverted, it fixed nothing and cost 1s per shutdown),
  dummy_writer() spawning during shutdown, and write failure
- the fix is an engine guarantee — a send issued before the stop flag is
  delivered — which belongs to the actor lifecycle, not a spin count

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-27 23:27:46 +02:00
..
chat-accept.sh fix(chat gate): every leg starts its own server — and it found a real bug 2026-08-27 23:27:46 +02:00
db-actor-accept.sh feat: arc stage 3 T7 — transparent DB actor (WO_T_DB hole closed) 2026-08-21 13:10:26 +02:00
db-bench.py feat: bounded mailboxes + WO_T_ACTOR (trap 13); try-arm place-copy fix 2026-08-23 01:05:13 +02:00
deps-accept.sh test: deps-accept — iteration 15's gate (Task 4) 2026-08-19 19:21:32 +02:00
employee-accept.sh feat: FK restrict on delete + employee sample runs; group-by parked (9b) 2026-08-16 18:37:23 +02:00
fibers-accept.sh feat: cross-shard actors — placement, envelopes, home-routed frees, WO-E222 (arc T6) 2026-08-20 12:06:13 +02:00
install-accept.sh feat: installable toolchain — version, wovm self-locate, dist tarball 2026-08-18 01:19:20 +02:00
install-readme.tmpl.md feat: installable toolchain — version, wovm self-locate, dist tarball 2026-08-18 01:19:20 +02:00
linkcheck.py docs: audit all markdown against the code, fix findings, flatten status folders 2026-08-26 19:20:22 +02:00
log-watcher-accept.sh fix: close per request, soak the daemon, kill five soak-found leaks (Tasks 5+6) 2026-08-15 00:32:49 +02:00
mkdist.sh feat: installable toolchain — version, wovm self-locate, dist tarball 2026-08-18 01:19:20 +02:00
oop-e2e.sh feat: cross-shard actors — placement, envelopes, home-routed frees, WO-E222 (arc T6) 2026-08-20 12:06:13 +02:00
single-binary-smoke.sh feat: milestone 1 complete — .wob emitter, conformance corpus, single binary; GC redesign specced 2026-08-11 19:31:26 +02:00
site-accept.sh refactor(porch): name the web framework porch, fix the wo.toml identifier claim 2026-08-26 19:36:34 +02:00
web-app-accept.sh refactor(porch): name the web framework porch, fix the wo.toml identifier claim 2026-08-26 19:36:34 +02:00