From 60414a17544b862edcbd3426d7326cc57c59a8d3 Mon Sep 17 00:00:00 2001 From: "shoney.arickathil" Date: Thu, 27 Aug 2026 23:43:45 +0200 Subject: [PATCH] =?UTF-8?q?feat(runtime):=20the=20shutdown=20drain=20guara?= =?UTF-8?q?ntee=20=E2=80=94=20iteration=2040?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit A message sent before the stop flag is observed must be delivered and run before the engine stops. One rule; a spin count could never express it. - root cause in `shard_main` (runtime/src/vm.c): NEXT_RUNNABLE() already stated the contract — "a WORKER on stop keeps DRAINING ... so queued shutdown messages (close frames!) still run" — but the IDLE branch contradicted it, calling fib_reap_all and breaking on WO_IO_STOP, abandoning its inbox for wo_engine_stop() to free wholesale - an actor between messages is exactly that idle case, which is why a WARM soak server hid it: warm shards held live fibers and took the right path - fix: while the primary's drain window is open, an idle worker adopts its inbox and runs what arrives; sched_yield on an empty poll so a drain cannot burn a core per shard and starve the actors it exists to let run - unreachable at WO_SHARDS=1: wo_engine_stop returns early at nshards <= 1 Measured: - fresh-server SIGTERM drain: 5 of 16 failing before, 20 of 20 clean after - `just chat` at the FULL 1000-client soak: 11 checks, 0 failures, both WO_IO backends, ASan clean with zero leaks - the fd leg settled at scale too: 1000 connections left the count at 44, unchanged after 20 more — lazy per-shard init, not a leak - runtime battery 36 suites (18 x both dispatch flavors) 0 fail; compiler 556 checks 0 fail - story: docs/stories/language-runtime-database/40-shutdown-drain-guarantee.md (chain 3 with 31, status done), board row, slice marker updated - outstanding and named: a pin below the gate needs new multithreaded test infrastructure — nothing in runtime/test/ drives wo_engine_start/stop and no corpus fixture can trigger a stop Co-Authored-By: Claude Opus 5 (1M context) --- ...tive-slice-2026-08-23-chat-ws-lifecycle.md | 22 +-- docs/stories/00-status.md | 1 + .../40-shutdown-drain-guarantee.md | 158 ++++++++++++++++++ runtime/src/vm.c | 27 +++ 4 files changed, 198 insertions(+), 10 deletions(-) create mode 100644 docs/stories/language-runtime-database/40-shutdown-drain-guarantee.md diff --git a/docs/active-slice-2026-08-23-chat-ws-lifecycle.md b/docs/active-slice-2026-08-23-chat-ws-lifecycle.md index a70ed53..8287d04 100644 --- a/docs/active-slice-2026-08-23-chat-ws-lifecycle.md +++ b/docs/active-slice-2026-08-23-chat-ws-lifecycle.md @@ -52,7 +52,7 @@ surfaces as "no listener". `runtime/build/wovm_asan` bit the same way. **Rebuild both after any branch switch** (`just woc-build`, `make -C runtime wovm-asan`) before believing a gate failure. -**The remaining failure is a real bug and is NOT fixed** — +**The remaining failure was a real bug and is now FIXED** (iteration 40) — [`2026-08-27-chat-drain-finding.md`](2026-08-27-chat-drain-finding.md). On a *fresh* server the SIGTERM drain leaves a client at EOF with no close frame in 5 of 16 runs. Traced: main → Registry → Room → Writer; the Registry runs but @@ -63,15 +63,17 @@ gate had been hiding it by draining a server the soak had already warmed. ## Pending -- 🔴 **The drain guarantee — the one blocker.** A `send` issued before the - stop flag must be delivered before the engine stops. `main` cannot park - after the flag (a park unwinds), so it spins, and **spinning is not a - barrier** — the evidence says the Room's shard never adopts its inbox, not - that it adopts it late. This is a semantic guarantee belonging to the actor - lifecycle (31), not a tuning parameter: it wants a stated rule in the - runtime lifecycle docs and a corpus fixture, not a bigger spin count. - **Nothing else in the slice should land before this**, because the drain is - half of what "actor lifecycle" means. +- ✅ **The drain guarantee — FIXED, and split into its own iteration** + ([40](stories/language-runtime-database/40-shutdown-drain-guarantee.md), + chain 3 with 31). It was a runtime semantic, not a task in a sample's gate. + Root cause: `NEXT_RUNNABLE()` already stated the contract — "a WORKER on stop + keeps DRAINING … so queued shutdown messages (close frames!) still run" — but + `shard_main`'s IDLE branch contradicted it, reaping and breaking on + `WO_IO_STOP` and abandoning its inbox for teardown to free. An actor between + messages is exactly that idle case, which is why a warm soak server hid it. + One branch now honours the primary's drain window, yielding on an empty poll + so the drain cannot starve the actors it exists to let run. **20 of 20 fresh + server drains clean, from 5 in 16 failing.** - ⬜ **T9 remainder** — the 1k soak has only been run trimmed (`CHAT_SOAK=20`); run it at the default 1000 once the drain is fixed. - ⬜ **T10 closeout** — stories 24/31/34 → `status: done` with banners (note the diff --git a/docs/stories/00-status.md b/docs/stories/00-status.md index 8ca6f94..f8ce9b6 100644 --- a/docs/stories/00-status.md +++ b/docs/stories/00-status.md @@ -357,6 +357,7 @@ that sequences its tasks. Read one, approve, then the next starts. | 34 | [Crypto builtins](language-runtime-database/34-crypto-builtins.md) | 🔄 **code landed** as 24's T1 (`d14fa9f`): `sha1`/`sha256`/`hmac_sha256`, ids 85–87 in `wob.h`, `runtime/src/crypto.c`, RFC/FIPS vectors 18/0, corpus pin. The 24 gate that once needed it is cleared. Frontmatter keeps `status: refine` only until 24's T10 closeout sets it to `done` | | 38 | [Content platform capabilities](language-runtime-database/38-content-platform-capabilities.md) | ⬜ off-chain, needs a spec — the two capability families no iteration owns, confirmed against `runtime/src/wob.h`: `fs` mutation verbs (six fs builtins, ids 40–45; `append` creates-if-absent, so nothing is ever replaced, truncated, deleted or renamed) and `net.connect` (ids 51–55 + 91–95, no connect, and no `connect()` anywhere in `runtime/src/` — so no OIDC/SMTP/object-store/webhook/federation). Driven by a `docs/examples/vault` content-collaboration workload, in 28's mould. New builtins from 96 (89/90 reserved for 31); no `.wob` bump (`WOB_VERSION 6u`, last moved by 36). Story written 2026-08-26 from the "can it build a Nextcloud?" ask | | 39 | [Web framework parity](language-runtime-database/39-web-framework-parity.md) | ⬜ off-chain, needs a spec — from [the Fiber v3.5.0 study](../plan/exploration/fiber/00-fiber-parity.md) (all 32 of its middleware read against `porch`; **nine already have a counterpart**). Leads with a **random-bytes builtin**: the framework ledger claimed CSRF/sessions were unblocked by iteration 34's HMAC, but HMAC authenticates a token and cannot mint one — there is no RNG anywhere in the runtime. Then cookies (absent both ways; `Resp.headers` being a map cannot carry two `Set-Cookie` lines), then limiter/idempotency (cheapest wins — `@table` + `time.ticks`, nothing new), sessions, CSRF, and the routing/response sugar. Streaming/SSE/compression, `@derive` binding, TTL cache, `proxy` and metrics all excluded with owners named | +| 40 | [Shutdown drain guarantee](language-runtime-database/40-shutdown-drain-guarantee.md) | ✅ **LANDED 2026-08-27 — chain 3, with 31; split out of 24.** One rule: **a message sent before the stop flag is observed must be delivered and run before the engine stops.** Found by measurement, not review: making the chat gate's drain leg start its OWN (cold) server exposed that **5 of 16** fresh-server SIGTERM drains left a WebSocket client at EOF with no close frame and no diagnostic. Traced to `shard_main` — `NEXT_RUNNABLE()` already stated the contract ("a WORKER on stop keeps DRAINING … close frames!") but the IDLE branch reaped and broke, abandoning its inbox for teardown to free. An actor between messages is exactly that idle case, which is why a WARM soak server hid it for so long. Fix is one branch honouring the primary's drain window, yielding on an empty poll. **20 of 20 clean after**; `just chat` 11 checks 0 failures at the full 1000-client soak (which also settled the fd question: 1000 connections left the count at 44); runtime battery 36 suites 0 fail, compiler 556 checks 0 fail. Ruled out: a bigger spin (a 1 s wall-clock deadline still failed 2 of 12) and spawn-during-shutdown. Outstanding: a pin below the gate — nothing in `runtime/test/` drives the engine start/stop and no corpus fixture can trigger a stop | | 37 | [wo-html components](language-runtime-database/37-wo-html-components.md) | ✅ off-chain — LANDED 2026-08-25. Raw text literal (backtick, margin stripped at lex time, `{{ }}` auto-escapes) + the component layer: `Component`/`render_all`/`Layout` in wo-html, `ok_html` moved into the framework, site and shop both migrated | | 35 | [net runtime seams](language-runtime-database/35-net-runtime-seams.md) | ⬜ off-chain — fd deadlines on the park plane, Unix sockets, peer address; owns the ledger's three 🔧 rows (story written 2026-08-22) | | 20 | [Cross-program tables](databasev2/09-cross-program-tables.md) | ⏸ hold (2026-08-21); channel done (branch ipc-attach keeps its manifest) | diff --git a/docs/stories/language-runtime-database/40-shutdown-drain-guarantee.md b/docs/stories/language-runtime-database/40-shutdown-drain-guarantee.md new file mode 100644 index 0000000..6da3eb0 --- /dev/null +++ b/docs/stories/language-runtime-database/40-shutdown-drain-guarantee.md @@ -0,0 +1,158 @@ +--- +iteration: "40" +status: done +chain: 3 +--- + +# iteration 40 — the shutdown drain guarantee: a send before the stop flag is delivered + +> Part of [Story — one language, one runtime, one database, one binary](00-story.md). +> +> **Split out of [24](24-chat-websocket-workload.md) on 2026-08-27** because it +> is a runtime *semantic*, not a task in a sample's gate. It belongs to the +> actor lifecycle ([31](31-actor-lifecycle.md), absorbed into 24) and it is +> the half of "lifecycle" that nothing had stated: 31 gave actors a death +> notice, this gives the program a shutdown that does not lose mail. +> +> **Found by measurement, not review.** The chat gate's drain leg had been +> passing only because it drained a server the 1k soak had already warmed. +> Making every leg start its own server exposed it: +> [`2026-08-27-chat-drain-finding.md`](../../2026-08-27-chat-drain-finding.md). + +## The rule + +**A message sent before the stop flag is observed must be delivered and run +before the engine stops.** One sentence, and it is the whole iteration. It is a +guarantee, not a tuning parameter — which is why a spin count could never +express it. + +What it does *not* promise: that a message sent *after* the flag is delivered, +that a parked fiber is resumed, or that an actor gets unbounded time. The drain +window is the primary's, and it closes when the primary returns. + +## The bug, as measured + +Fresh server, two WebSocket clients, `SIGTERM`, both must receive a close frame: + +| Sample | Result | +| --- | --- | +| 5 fresh servers | 1 failure (`eof\|close`) | +| 12 fresh servers | 3 failures, one `eof\|eof` | +| 16 fresh servers | 5 failures | + +The failing client's socket reaches EOF with **no close frame and no +diagnostic** — the process exits and the kernel closes the fd. + +Traced with instrumentation on the sample's actors: `main` → Registry → Room → +Writer. The Registry runs and sees its room. The **Room never processes the +shutdown message**, so the Writer's close branch never runs. Clients that did +get a frame were saved by their own Reader noticing `env.stopping()`, not by the +room broadcast. + +## The design, as built + +`runtime/src/vm.c` already encoded the correct contract in `NEXT_RUNNABLE()`: +a worker that takes a stop while it has a live fiber returns 2 and **keeps +draining its inbox** until the primary sets `eng_shutdown`. Its comment says so +in as many words — "queued shutdown messages (close frames!) still run". + +`shard_main`'s own idle branch contradicted it. A worker with an empty run queue +waits in `wo_io_wait`, and on `WO_IO_STOP` it called `fib_reap_all` and +**broke** — abandoning whatever was still in its inbox, which `wo_engine_stop` +then freed wholesale during teardown. + +So the failure needed a shard that was *idle* at `SIGTERM`. A Room actor between +messages is exactly that, which is why the warm soak server hid it: warm shards +had live fibers and took the correct path. + +The fix makes the idle branch obey the same contract: while the primary's drain +window is open, an idle worker adopts its inbox and runs what arrives, yielding +between empty polls so a drain cannot become a hot spin across every core. Only +`eng_shutdown` — set by the primary after `main` returns — ends it. + +One branch, in one place, matching a contract the file already stated. + +## Progress + +| Piece | State | +| --- | --- | +| the idle-worker drain branch in `shard_main` (`runtime/src/vm.c`) | ✅ one branch, matching the contract `NEXT_RUNNABLE()` already stated | +| `sched_yield` on an empty poll so the drain cannot hot-spin | ✅ | +| fresh-server drain, repeated | ✅ **20 of 20**, from 5-in-16 failing | +| chat gate at the default 1k soak | ✅ **11 checks, 0 failures** — 1000/1000 clients, both `WO_IO` backends, ASan clean | +| full runtime battery (this touches every actor program's shard loop) | ✅ **36 suites** (18 × both dispatch flavors), 0 fail, `cli_smoke: OK`; compiler 556 checks 0 fail | +| the regression pin | ✅ the chat gate's drain leg, now that it starts its OWN (cold) server — that decoupling is what caught this. **Not** a corpus fixture or unit test: nothing in `runtime/test/` drives `wo_engine_start`/`wo_engine_stop` today, and no corpus fixture can trigger a stop, so pinning it below the gate means new multithreaded test infrastructure — named as its own cost, not smuggled in here | + +**Measured 2026-08-27.** Before: 5 of 16 fresh-server drains left a client at +EOF. After: **20 of 20 clean.** At the observed failure rate, 20 clean runs by +luck would be about 0.04%, so this is the fix rather than a quieter race. + +## Acceptance Criteria + +Met: + +- **Given** a fresh server with two connected WebSocket clients, **when** it is + sent `SIGTERM`, **then** both clients receive a close frame — **repeatedly**, + not once. The bug reproduced at 5 in 16, so a single green run proves nothing; + the criterion is a run of at least 16 with zero failures. + ✅ **20 of 20**, from 5-in-16 failing. A single run would have proved nothing. +- **Given** an actor whose shard is idle at the moment of the stop, **when** a + message is sent to it before the stop flag is observed, **then** its + `receive` runs before the engine stops. ✅ this is exactly the case that + failed — the Room between messages — and it is what the branch now covers. +- **Given** the drain window, **when** a worker has nothing to adopt, **then** + it does not hot-spin. ✅ `sched_yield()` on an empty poll; the 1k soak's RSS + and timing legs are unchanged (marker reached all 1000 in 28 ms). +- **Given** `just chat`, **when** it runs at the default soak, **then** all + legs pass on both `WO_IO` backends and under the ASan build with zero leaks. + ✅ 11 checks, 0 failures. The fd leg also settled the lazy-init question at + scale: **1000 connections left the count at 44**, unchanged after 20 more. +- **Given** the full runtime battery, **when** it runs, **then** no suite + regresses — this touches the shard loop every actor program uses. ✅ 36 suites + 0 fail, plus the compiler's 556 checks. +- **Given** a program with no worker shards (`WO_SHARDS=1`), **when** it stops, + **then** behaviour is unchanged. ✅ the gate's `WO_SHARDS=1` leg passes, and + the branch is unreachable there — `wo_engine_stop` returns early at + `nshards <= 1`, so a single-shard program never enters a worker loop. + +Outstanding: + +- **A pin below the gate.** The guarantee is currently proven by the chat gate + only. Nothing in `runtime/test/` drives `wo_engine_start`/`wo_engine_stop`, + and no corpus fixture can trigger a stop, so pinning it lower means new + multithreaded test infrastructure. Named as its own cost rather than assumed + cheap. + +## Out Of Scope + +- **Unbounded drain.** The window is the primary's and closes when `main` + returns. A program that wants longer holds the window open itself. +- **Delivering sends issued *after* the stop flag.** Nothing promises that, and + promising it would mean a program could refuse to exit. +- **Resuming parked fibers on stop.** `WO_SYS_STOPPED` unwinds them; that + contract is iteration 24's and stays. +- **A shutdown acknowledgement in the language surface.** The alternative fix + was a barrier the sample builds itself, rejected below. +- **`main` parking after the stop flag.** Still forbidden — a park after the + flag unwinds. `main` still spins; the point is that spinning now works + because the workers cooperate. + +## Info — the forks, settled + +1. **Engine guarantee, not a sample barrier.** The alternative was an + acknowledged drain: rooms confirm back to `main`, which waits. Rejected — + `main` cannot park after the stop flag, so it could only spin on the + acknowledgement anyway, and every future actor program would have to + re-implement the same handshake to avoid losing mail. A guarantee is stated + once; a barrier is re-invented per program. +2. **Not the spin budget.** Replacing the sample's `spin < 20000000` with a 1 s + wall-clock deadline still failed 2 of 12. More time cannot help when the + shard is not scheduled at all, and the reverted attempt cost a fixed second + on every shutdown. Recorded because a bigger spin is the obvious wrong fix. +3. **Not `dummy_writer()`.** Hoisting the shutdown message's placeholder actor + out of the drain path (it spawned during shutdown) left 5 of 16 failing. +4. **Yield rather than spin in the idle drain.** A worker polling an empty + inbox in a tight loop would burn a core per shard during the window and + starve the actors being drained. +5. **Chain position 3**, with [31](31-actor-lifecycle.md): it is lifecycle + semantics, and [24](24-chat-websocket-workload.md)'s gate is what proves it. diff --git a/runtime/src/vm.c b/runtime/src/vm.c index 86244d0..31c0b9e 100644 --- a/runtime/src/vm.c +++ b/runtime/src/vm.c @@ -447,6 +447,33 @@ static void *shard_main(void *arg) { } else { int rc = wo_io_wait(vm); /* parked fibers AND the wake eventfd */ if (rc == WO_IO_STOP) { + /* iteration 40 — THE DRAIN GUARANTEE. A message sent before + * the stop flag is observed must be delivered and run before + * the engine stops. + * + * NEXT_RUNNABLE() already states this contract for a worker + * holding a live fiber: it returns 2 and keeps draining "so + * queued shutdown messages (close frames!) still run". This + * branch — the IDLE worker, empty run queue, waiting on the + * plane — used to reap and break instead, abandoning whatever + * sat in its inbox for wo_engine_stop() to free wholesale. + * + * An actor between messages is exactly that idle case, which + * is why a WARM server hid the bug: warm shards had live + * fibers and took the correct path. Measured 2026-08-27 on a + * fresh server: 5 of 16 SIGTERM drains left a WebSocket + * client at EOF with no close frame and no diagnostic. + * + * The window belongs to the PRIMARY and closes when it sets + * eng_shutdown (after main returns), so honour it here and + * only exit when the primary says so. Yield on an empty poll: + * a tight loop would burn a core per shard and starve the very + * actors the drain exists to let run. */ + if (!eng_shutdown) { + (void)wo_vm_adopt(vm); + if (!vm->qhead) sched_yield(); + continue; + } fib_reap_all(vm); break; }