feat(runtime): the shutdown drain guarantee — iteration 40

A message sent before the stop flag is observed must be delivered and run
before the engine stops. One rule; a spin count could never express it.

- root cause in `shard_main` (runtime/src/vm.c): NEXT_RUNNABLE() already
  stated the contract — "a WORKER on stop keeps DRAINING ... so queued
  shutdown messages (close frames!) still run" — but the IDLE branch
  contradicted it, calling fib_reap_all and breaking on WO_IO_STOP,
  abandoning its inbox for wo_engine_stop() to free wholesale
- an actor between messages is exactly that idle case, which is why a WARM
  soak server hid it: warm shards held live fibers and took the right path
- fix: while the primary's drain window is open, an idle worker adopts its
  inbox and runs what arrives; sched_yield on an empty poll so a drain
  cannot burn a core per shard and starve the actors it exists to let run
- unreachable at WO_SHARDS=1: wo_engine_stop returns early at nshards <= 1

Measured:

- fresh-server SIGTERM drain: 5 of 16 failing before, 20 of 20 clean after
- `just chat` at the FULL 1000-client soak: 11 checks, 0 failures, both
  WO_IO backends, ASan clean with zero leaks
- the fd leg settled at scale too: 1000 connections left the count at 44,
  unchanged after 20 more — lazy per-shard init, not a leak
- runtime battery 36 suites (18 x both dispatch flavors) 0 fail;
  compiler 556 checks 0 fail

- story: docs/stories/language-runtime-database/40-shutdown-drain-guarantee.md
  (chain 3 with 31, status done), board row, slice marker updated
- outstanding and named: a pin below the gate needs new multithreaded test
  infrastructure — nothing in runtime/test/ drives wo_engine_start/stop and
  no corpus fixture can trigger a stop

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
shoney.arickathil 2026-08-27 23:43:45 +02:00
parent 3bc85d85f3
commit 60414a1754
4 changed files with 198 additions and 10 deletions

View file

@ -52,7 +52,7 @@ surfaces as "no listener". `runtime/build/wovm_asan` bit the same way. **Rebuild
both after any branch switch** (`just woc-build`, `make -C runtime wovm-asan`) both after any branch switch** (`just woc-build`, `make -C runtime wovm-asan`)
before believing a gate failure. before believing a gate failure.
**The remaining failure is a real bug and is NOT fixed** — **The remaining failure was a real bug and is now FIXED** (iteration 40) —
[`2026-08-27-chat-drain-finding.md`](2026-08-27-chat-drain-finding.md). On a [`2026-08-27-chat-drain-finding.md`](2026-08-27-chat-drain-finding.md). On a
*fresh* server the SIGTERM drain leaves a client at EOF with no close frame in *fresh* server the SIGTERM drain leaves a client at EOF with no close frame in
5 of 16 runs. Traced: main → Registry → Room → Writer; the Registry runs but 5 of 16 runs. Traced: main → Registry → Room → Writer; the Registry runs but
@ -63,15 +63,17 @@ gate had been hiding it by draining a server the soak had already warmed.
## Pending ## Pending
- 🔴 **The drain guarantee — the one blocker.** A `send` issued before the - ✅ **The drain guarantee — FIXED, and split into its own iteration**
stop flag must be delivered before the engine stops. `main` cannot park ([40](stories/language-runtime-database/40-shutdown-drain-guarantee.md),
after the flag (a park unwinds), so it spins, and **spinning is not a chain 3 with 31). It was a runtime semantic, not a task in a sample's gate.
barrier** — the evidence says the Room's shard never adopts its inbox, not Root cause: `NEXT_RUNNABLE()` already stated the contract — "a WORKER on stop
that it adopts it late. This is a semantic guarantee belonging to the actor keeps DRAINING … so queued shutdown messages (close frames!) still run" — but
lifecycle (31), not a tuning parameter: it wants a stated rule in the `shard_main`'s IDLE branch contradicted it, reaping and breaking on
runtime lifecycle docs and a corpus fixture, not a bigger spin count. `WO_IO_STOP` and abandoning its inbox for teardown to free. An actor between
**Nothing else in the slice should land before this**, because the drain is messages is exactly that idle case, which is why a warm soak server hid it.
half of what "actor lifecycle" means. One branch now honours the primary's drain window, yielding on an empty poll
so the drain cannot starve the actors it exists to let run. **20 of 20 fresh
server drains clean, from 5 in 16 failing.**
- ⬜ **T9 remainder** — the 1k soak has only been run trimmed - ⬜ **T9 remainder** — the 1k soak has only been run trimmed
(`CHAT_SOAK=20`); run it at the default 1000 once the drain is fixed. (`CHAT_SOAK=20`); run it at the default 1000 once the drain is fixed.
- ⬜ **T10 closeout** — stories 24/31/34 → `status: done` with banners (note the - ⬜ **T10 closeout** — stories 24/31/34 → `status: done` with banners (note the

View file

@ -357,6 +357,7 @@ that sequences its tasks. Read one, approve, then the next starts.
| 34 | [Crypto builtins](language-runtime-database/34-crypto-builtins.md) | 🔄 **code landed** as 24's T1 (`d14fa9f`): `sha1`/`sha256`/`hmac_sha256`, ids 85–87 in `wob.h`, `runtime/src/crypto.c`, RFC/FIPS vectors 18/0, corpus pin. The 24 gate that once needed it is cleared. Frontmatter keeps `status: refine` only until 24's T10 closeout sets it to `done` | | 34 | [Crypto builtins](language-runtime-database/34-crypto-builtins.md) | 🔄 **code landed** as 24's T1 (`d14fa9f`): `sha1`/`sha256`/`hmac_sha256`, ids 85–87 in `wob.h`, `runtime/src/crypto.c`, RFC/FIPS vectors 18/0, corpus pin. The 24 gate that once needed it is cleared. Frontmatter keeps `status: refine` only until 24's T10 closeout sets it to `done` |
| 38 | [Content platform capabilities](language-runtime-database/38-content-platform-capabilities.md) | ⬜ off-chain, needs a spec — the two capability families no iteration owns, confirmed against `runtime/src/wob.h`: `fs` mutation verbs (six fs builtins, ids 40–45; `append` creates-if-absent, so nothing is ever replaced, truncated, deleted or renamed) and `net.connect` (ids 51–55 + 91–95, no connect, and no `connect()` anywhere in `runtime/src/` — so no OIDC/SMTP/object-store/webhook/federation). Driven by a `docs/examples/vault` content-collaboration workload, in 28's mould. New builtins from 96 (89/90 reserved for 31); no `.wob` bump (`WOB_VERSION 6u`, last moved by 36). Story written 2026-08-26 from the "can it build a Nextcloud?" ask | | 38 | [Content platform capabilities](language-runtime-database/38-content-platform-capabilities.md) | ⬜ off-chain, needs a spec — the two capability families no iteration owns, confirmed against `runtime/src/wob.h`: `fs` mutation verbs (six fs builtins, ids 40–45; `append` creates-if-absent, so nothing is ever replaced, truncated, deleted or renamed) and `net.connect` (ids 51–55 + 91–95, no connect, and no `connect()` anywhere in `runtime/src/` — so no OIDC/SMTP/object-store/webhook/federation). Driven by a `docs/examples/vault` content-collaboration workload, in 28's mould. New builtins from 96 (89/90 reserved for 31); no `.wob` bump (`WOB_VERSION 6u`, last moved by 36). Story written 2026-08-26 from the "can it build a Nextcloud?" ask |
| 39 | [Web framework parity](language-runtime-database/39-web-framework-parity.md) | ⬜ off-chain, needs a spec — from [the Fiber v3.5.0 study](../plan/exploration/fiber/00-fiber-parity.md) (all 32 of its middleware read against `porch`; **nine already have a counterpart**). Leads with a **random-bytes builtin**: the framework ledger claimed CSRF/sessions were unblocked by iteration 34's HMAC, but HMAC authenticates a token and cannot mint one — there is no RNG anywhere in the runtime. Then cookies (absent both ways; `Resp.headers` being a map cannot carry two `Set-Cookie` lines), then limiter/idempotency (cheapest wins — `@table` + `time.ticks`, nothing new), sessions, CSRF, and the routing/response sugar. Streaming/SSE/compression, `@derive` binding, TTL cache, `proxy` and metrics all excluded with owners named | | 39 | [Web framework parity](language-runtime-database/39-web-framework-parity.md) | ⬜ off-chain, needs a spec — from [the Fiber v3.5.0 study](../plan/exploration/fiber/00-fiber-parity.md) (all 32 of its middleware read against `porch`; **nine already have a counterpart**). Leads with a **random-bytes builtin**: the framework ledger claimed CSRF/sessions were unblocked by iteration 34's HMAC, but HMAC authenticates a token and cannot mint one — there is no RNG anywhere in the runtime. Then cookies (absent both ways; `Resp.headers` being a map cannot carry two `Set-Cookie` lines), then limiter/idempotency (cheapest wins — `@table` + `time.ticks`, nothing new), sessions, CSRF, and the routing/response sugar. Streaming/SSE/compression, `@derive` binding, TTL cache, `proxy` and metrics all excluded with owners named |
| 40 | [Shutdown drain guarantee](language-runtime-database/40-shutdown-drain-guarantee.md) | ✅ **LANDED 2026-08-27 — chain 3, with 31; split out of 24.** One rule: **a message sent before the stop flag is observed must be delivered and run before the engine stops.** Found by measurement, not review: making the chat gate's drain leg start its OWN (cold) server exposed that **5 of 16** fresh-server SIGTERM drains left a WebSocket client at EOF with no close frame and no diagnostic. Traced to `shard_main` — `NEXT_RUNNABLE()` already stated the contract ("a WORKER on stop keeps DRAINING … close frames!") but the IDLE branch reaped and broke, abandoning its inbox for teardown to free. An actor between messages is exactly that idle case, which is why a WARM soak server hid it for so long. Fix is one branch honouring the primary's drain window, yielding on an empty poll. **20 of 20 clean after**; `just chat` 11 checks 0 failures at the full 1000-client soak (which also settled the fd question: 1000 connections left the count at 44); runtime battery 36 suites 0 fail, compiler 556 checks 0 fail. Ruled out: a bigger spin (a 1 s wall-clock deadline still failed 2 of 12) and spawn-during-shutdown. Outstanding: a pin below the gate — nothing in `runtime/test/` drives the engine start/stop and no corpus fixture can trigger a stop |
| 37 | [wo-html components](language-runtime-database/37-wo-html-components.md) | ✅ off-chain — LANDED 2026-08-25. Raw text literal (backtick, margin stripped at lex time, `{{ }}` auto-escapes) + the component layer: `Component`/`render_all`/`Layout` in wo-html, `ok_html` moved into the framework, site and shop both migrated | | 37 | [wo-html components](language-runtime-database/37-wo-html-components.md) | ✅ off-chain — LANDED 2026-08-25. Raw text literal (backtick, margin stripped at lex time, `{{ }}` auto-escapes) + the component layer: `Component`/`render_all`/`Layout` in wo-html, `ok_html` moved into the framework, site and shop both migrated |
| 35 | [net runtime seams](language-runtime-database/35-net-runtime-seams.md) | ⬜ off-chain — fd deadlines on the park plane, Unix sockets, peer address; owns the ledger's three 🔧 rows (story written 2026-08-22) | | 35 | [net runtime seams](language-runtime-database/35-net-runtime-seams.md) | ⬜ off-chain — fd deadlines on the park plane, Unix sockets, peer address; owns the ledger's three 🔧 rows (story written 2026-08-22) |
| 20 | [Cross-program tables](databasev2/09-cross-program-tables.md) | ⏸ hold (2026-08-21); channel done (branch ipc-attach keeps its manifest) | | 20 | [Cross-program tables](databasev2/09-cross-program-tables.md) | ⏸ hold (2026-08-21); channel done (branch ipc-attach keeps its manifest) |

View file

@ -0,0 +1,158 @@
---
iteration: "40"
status: done
chain: 3
---
# iteration 40 — the shutdown drain guarantee: a send before the stop flag is delivered
> Part of [Story — one language, one runtime, one database, one binary](00-story.md).
>
> **Split out of [24](24-chat-websocket-workload.md) on 2026-08-27** because it
> is a runtime *semantic*, not a task in a sample's gate. It belongs to the
> actor lifecycle ([31](31-actor-lifecycle.md), absorbed into 24) and it is
> the half of "lifecycle" that nothing had stated: 31 gave actors a death
> notice, this gives the program a shutdown that does not lose mail.
>
> **Found by measurement, not review.** The chat gate's drain leg had been
> passing only because it drained a server the 1k soak had already warmed.
> Making every leg start its own server exposed it:
> [`2026-08-27-chat-drain-finding.md`](../../2026-08-27-chat-drain-finding.md).
## The rule
**A message sent before the stop flag is observed must be delivered and run
before the engine stops.** One sentence, and it is the whole iteration. It is a
guarantee, not a tuning parameter — which is why a spin count could never
express it.
What it does *not* promise: that a message sent *after* the flag is delivered,
that a parked fiber is resumed, or that an actor gets unbounded time. The drain
window is the primary's, and it closes when the primary returns.
## The bug, as measured
Fresh server, two WebSocket clients, `SIGTERM`, both must receive a close frame:
| Sample | Result |
| --- | --- |
| 5 fresh servers | 1 failure (`eof\|close`) |
| 12 fresh servers | 3 failures, one `eof\|eof` |
| 16 fresh servers | 5 failures |
The failing client's socket reaches EOF with **no close frame and no
diagnostic** — the process exits and the kernel closes the fd.
Traced with instrumentation on the sample's actors: `main` → Registry → Room →
Writer. The Registry runs and sees its room. The **Room never processes the
shutdown message**, so the Writer's close branch never runs. Clients that did
get a frame were saved by their own Reader noticing `env.stopping()`, not by the
room broadcast.
## The design, as built
`runtime/src/vm.c` already encoded the correct contract in `NEXT_RUNNABLE()`:
a worker that takes a stop while it has a live fiber returns 2 and **keeps
draining its inbox** until the primary sets `eng_shutdown`. Its comment says so
in as many words — "queued shutdown messages (close frames!) still run".
`shard_main`'s own idle branch contradicted it. A worker with an empty run queue
waits in `wo_io_wait`, and on `WO_IO_STOP` it called `fib_reap_all` and
**broke** — abandoning whatever was still in its inbox, which `wo_engine_stop`
then freed wholesale during teardown.
So the failure needed a shard that was *idle* at `SIGTERM`. A Room actor between
messages is exactly that, which is why the warm soak server hid it: warm shards
had live fibers and took the correct path.
The fix makes the idle branch obey the same contract: while the primary's drain
window is open, an idle worker adopts its inbox and runs what arrives, yielding
between empty polls so a drain cannot become a hot spin across every core. Only
`eng_shutdown` — set by the primary after `main` returns — ends it.
One branch, in one place, matching a contract the file already stated.
## Progress
| Piece | State |
| --- | --- |
| the idle-worker drain branch in `shard_main` (`runtime/src/vm.c`) | ✅ one branch, matching the contract `NEXT_RUNNABLE()` already stated |
| `sched_yield` on an empty poll so the drain cannot hot-spin | ✅ |
| fresh-server drain, repeated | ✅ **20 of 20**, from 5-in-16 failing |
| chat gate at the default 1k soak | ✅ **11 checks, 0 failures** — 1000/1000 clients, both `WO_IO` backends, ASan clean |
| full runtime battery (this touches every actor program's shard loop) | ✅ **36 suites** (18 × both dispatch flavors), 0 fail, `cli_smoke: OK`; compiler 556 checks 0 fail |
| the regression pin | ✅ the chat gate's drain leg, now that it starts its OWN (cold) server — that decoupling is what caught this. **Not** a corpus fixture or unit test: nothing in `runtime/test/` drives `wo_engine_start`/`wo_engine_stop` today, and no corpus fixture can trigger a stop, so pinning it below the gate means new multithreaded test infrastructure — named as its own cost, not smuggled in here |
**Measured 2026-08-27.** Before: 5 of 16 fresh-server drains left a client at
EOF. After: **20 of 20 clean.** At the observed failure rate, 20 clean runs by
luck would be about 0.04%, so this is the fix rather than a quieter race.
## Acceptance Criteria
Met:
- **Given** a fresh server with two connected WebSocket clients, **when** it is
sent `SIGTERM`, **then** both clients receive a close frame — **repeatedly**,
not once. The bug reproduced at 5 in 16, so a single green run proves nothing;
the criterion is a run of at least 16 with zero failures.
✅ **20 of 20**, from 5-in-16 failing. A single run would have proved nothing.
- **Given** an actor whose shard is idle at the moment of the stop, **when** a
message is sent to it before the stop flag is observed, **then** its
`receive` runs before the engine stops. ✅ this is exactly the case that
failed — the Room between messages — and it is what the branch now covers.
- **Given** the drain window, **when** a worker has nothing to adopt, **then**
it does not hot-spin. ✅ `sched_yield()` on an empty poll; the 1k soak's RSS
and timing legs are unchanged (marker reached all 1000 in 28 ms).
- **Given** `just chat`, **when** it runs at the default soak, **then** all
legs pass on both `WO_IO` backends and under the ASan build with zero leaks.
✅ 11 checks, 0 failures. The fd leg also settled the lazy-init question at
scale: **1000 connections left the count at 44**, unchanged after 20 more.
- **Given** the full runtime battery, **when** it runs, **then** no suite
regresses — this touches the shard loop every actor program uses. ✅ 36 suites
0 fail, plus the compiler's 556 checks.
- **Given** a program with no worker shards (`WO_SHARDS=1`), **when** it stops,
**then** behaviour is unchanged. ✅ the gate's `WO_SHARDS=1` leg passes, and
the branch is unreachable there — `wo_engine_stop` returns early at
`nshards <= 1`, so a single-shard program never enters a worker loop.
Outstanding:
- **A pin below the gate.** The guarantee is currently proven by the chat gate
only. Nothing in `runtime/test/` drives `wo_engine_start`/`wo_engine_stop`,
and no corpus fixture can trigger a stop, so pinning it lower means new
multithreaded test infrastructure. Named as its own cost rather than assumed
cheap.
## Out Of Scope
- **Unbounded drain.** The window is the primary's and closes when `main`
returns. A program that wants longer holds the window open itself.
- **Delivering sends issued *after* the stop flag.** Nothing promises that, and
promising it would mean a program could refuse to exit.
- **Resuming parked fibers on stop.** `WO_SYS_STOPPED` unwinds them; that
contract is iteration 24's and stays.
- **A shutdown acknowledgement in the language surface.** The alternative fix
was a barrier the sample builds itself, rejected below.
- **`main` parking after the stop flag.** Still forbidden — a park after the
flag unwinds. `main` still spins; the point is that spinning now works
because the workers cooperate.
## Info — the forks, settled
1. **Engine guarantee, not a sample barrier.** The alternative was an
acknowledged drain: rooms confirm back to `main`, which waits. Rejected —
`main` cannot park after the stop flag, so it could only spin on the
acknowledgement anyway, and every future actor program would have to
re-implement the same handshake to avoid losing mail. A guarantee is stated
once; a barrier is re-invented per program.
2. **Not the spin budget.** Replacing the sample's `spin < 20000000` with a 1 s
wall-clock deadline still failed 2 of 12. More time cannot help when the
shard is not scheduled at all, and the reverted attempt cost a fixed second
on every shutdown. Recorded because a bigger spin is the obvious wrong fix.
3. **Not `dummy_writer()`.** Hoisting the shutdown message's placeholder actor
out of the drain path (it spawned during shutdown) left 5 of 16 failing.
4. **Yield rather than spin in the idle drain.** A worker polling an empty
inbox in a tight loop would burn a core per shard during the window and
starve the actors being drained.
5. **Chain position 3**, with [31](31-actor-lifecycle.md): it is lifecycle
semantics, and [24](24-chat-websocket-workload.md)'s gate is what proves it.

View file

@ -447,6 +447,33 @@ static void *shard_main(void *arg) {
} else { } else {
int rc = wo_io_wait(vm); /* parked fibers AND the wake eventfd */ int rc = wo_io_wait(vm); /* parked fibers AND the wake eventfd */
if (rc == WO_IO_STOP) { if (rc == WO_IO_STOP) {
/* iteration 40 — THE DRAIN GUARANTEE. A message sent before
* the stop flag is observed must be delivered and run before
* the engine stops.
*
* NEXT_RUNNABLE() already states this contract for a worker
* holding a live fiber: it returns 2 and keeps draining "so
* queued shutdown messages (close frames!) still run". This
* branch — the IDLE worker, empty run queue, waiting on the
* plane — used to reap and break instead, abandoning whatever
* sat in its inbox for wo_engine_stop() to free wholesale.
*
* An actor between messages is exactly that idle case, which
* is why a WARM server hid the bug: warm shards had live
* fibers and took the correct path. Measured 2026-08-27 on a
* fresh server: 5 of 16 SIGTERM drains left a WebSocket
* client at EOF with no close frame and no diagnostic.
*
* The window belongs to the PRIMARY and closes when it sets
* eng_shutdown (after main returns), so honour it here and
* only exit when the primary says so. Yield on an empty poll:
* a tight loop would burn a core per shard and starve the very
* actors the drain exists to let run. */
if (!eng_shutdown) {
(void)wo_vm_adopt(vm);
if (!vm->qhead) sched_yield();
continue;
}
fib_reap_all(vm); fib_reap_all(vm);
break; break;
} }