feat(runtime): the shutdown drain guarantee — iteration 40
A message sent before the stop flag is observed must be delivered and run before the engine stops. One rule; a spin count could never express it. - root cause in `shard_main` (runtime/src/vm.c): NEXT_RUNNABLE() already stated the contract — "a WORKER on stop keeps DRAINING ... so queued shutdown messages (close frames!) still run" — but the IDLE branch contradicted it, calling fib_reap_all and breaking on WO_IO_STOP, abandoning its inbox for wo_engine_stop() to free wholesale - an actor between messages is exactly that idle case, which is why a WARM soak server hid it: warm shards held live fibers and took the right path - fix: while the primary's drain window is open, an idle worker adopts its inbox and runs what arrives; sched_yield on an empty poll so a drain cannot burn a core per shard and starve the actors it exists to let run - unreachable at WO_SHARDS=1: wo_engine_stop returns early at nshards <= 1 Measured: - fresh-server SIGTERM drain: 5 of 16 failing before, 20 of 20 clean after - `just chat` at the FULL 1000-client soak: 11 checks, 0 failures, both WO_IO backends, ASan clean with zero leaks - the fd leg settled at scale too: 1000 connections left the count at 44, unchanged after 20 more — lazy per-shard init, not a leak - runtime battery 36 suites (18 x both dispatch flavors) 0 fail; compiler 556 checks 0 fail - story: docs/stories/language-runtime-database/40-shutdown-drain-guarantee.md (chain 3 with 31, status done), board row, slice marker updated - outstanding and named: a pin below the gate needs new multithreaded test infrastructure — nothing in runtime/test/ drives wo_engine_start/stop and no corpus fixture can trigger a stop Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
parent
3bc85d85f3
commit
60414a1754
4 changed files with 198 additions and 10 deletions
|
|
@ -52,7 +52,7 @@ surfaces as "no listener". `runtime/build/wovm_asan` bit the same way. **Rebuild
|
|||
both after any branch switch** (`just woc-build`, `make -C runtime wovm-asan`)
|
||||
before believing a gate failure.
|
||||
|
||||
**The remaining failure is a real bug and is NOT fixed** —
|
||||
**The remaining failure was a real bug and is now FIXED** (iteration 40) —
|
||||
[`2026-08-27-chat-drain-finding.md`](2026-08-27-chat-drain-finding.md). On a
|
||||
*fresh* server the SIGTERM drain leaves a client at EOF with no close frame in
|
||||
5 of 16 runs. Traced: main → Registry → Room → Writer; the Registry runs but
|
||||
|
|
@ -63,15 +63,17 @@ gate had been hiding it by draining a server the soak had already warmed.
|
|||
|
||||
## Pending
|
||||
|
||||
- 🔴 **The drain guarantee — the one blocker.** A `send` issued before the
|
||||
stop flag must be delivered before the engine stops. `main` cannot park
|
||||
after the flag (a park unwinds), so it spins, and **spinning is not a
|
||||
barrier** — the evidence says the Room's shard never adopts its inbox, not
|
||||
that it adopts it late. This is a semantic guarantee belonging to the actor
|
||||
lifecycle (31), not a tuning parameter: it wants a stated rule in the
|
||||
runtime lifecycle docs and a corpus fixture, not a bigger spin count.
|
||||
**Nothing else in the slice should land before this**, because the drain is
|
||||
half of what "actor lifecycle" means.
|
||||
- ✅ **The drain guarantee — FIXED, and split into its own iteration**
|
||||
([40](stories/language-runtime-database/40-shutdown-drain-guarantee.md),
|
||||
chain 3 with 31). It was a runtime semantic, not a task in a sample's gate.
|
||||
Root cause: `NEXT_RUNNABLE()` already stated the contract — "a WORKER on stop
|
||||
keeps DRAINING … so queued shutdown messages (close frames!) still run" — but
|
||||
`shard_main`'s IDLE branch contradicted it, reaping and breaking on
|
||||
`WO_IO_STOP` and abandoning its inbox for teardown to free. An actor between
|
||||
messages is exactly that idle case, which is why a warm soak server hid it.
|
||||
One branch now honours the primary's drain window, yielding on an empty poll
|
||||
so the drain cannot starve the actors it exists to let run. **20 of 20 fresh
|
||||
server drains clean, from 5 in 16 failing.**
|
||||
- ⬜ **T9 remainder** — the 1k soak has only been run trimmed
|
||||
(`CHAT_SOAK=20`); run it at the default 1000 once the drain is fixed.
|
||||
- ⬜ **T10 closeout** — stories 24/31/34 → `status: done` with banners (note the
|
||||
|
|
|
|||
|
|
@ -357,6 +357,7 @@ that sequences its tasks. Read one, approve, then the next starts.
|
|||
| 34 | [Crypto builtins](language-runtime-database/34-crypto-builtins.md) | 🔄 **code landed** as 24's T1 (`d14fa9f`): `sha1`/`sha256`/`hmac_sha256`, ids 85–87 in `wob.h`, `runtime/src/crypto.c`, RFC/FIPS vectors 18/0, corpus pin. The 24 gate that once needed it is cleared. Frontmatter keeps `status: refine` only until 24's T10 closeout sets it to `done` |
|
||||
| 38 | [Content platform capabilities](language-runtime-database/38-content-platform-capabilities.md) | ⬜ off-chain, needs a spec — the two capability families no iteration owns, confirmed against `runtime/src/wob.h`: `fs` mutation verbs (six fs builtins, ids 40–45; `append` creates-if-absent, so nothing is ever replaced, truncated, deleted or renamed) and `net.connect` (ids 51–55 + 91–95, no connect, and no `connect()` anywhere in `runtime/src/` — so no OIDC/SMTP/object-store/webhook/federation). Driven by a `docs/examples/vault` content-collaboration workload, in 28's mould. New builtins from 96 (89/90 reserved for 31); no `.wob` bump (`WOB_VERSION 6u`, last moved by 36). Story written 2026-08-26 from the "can it build a Nextcloud?" ask |
|
||||
| 39 | [Web framework parity](language-runtime-database/39-web-framework-parity.md) | ⬜ off-chain, needs a spec — from [the Fiber v3.5.0 study](../plan/exploration/fiber/00-fiber-parity.md) (all 32 of its middleware read against `porch`; **nine already have a counterpart**). Leads with a **random-bytes builtin**: the framework ledger claimed CSRF/sessions were unblocked by iteration 34's HMAC, but HMAC authenticates a token and cannot mint one — there is no RNG anywhere in the runtime. Then cookies (absent both ways; `Resp.headers` being a map cannot carry two `Set-Cookie` lines), then limiter/idempotency (cheapest wins — `@table` + `time.ticks`, nothing new), sessions, CSRF, and the routing/response sugar. Streaming/SSE/compression, `@derive` binding, TTL cache, `proxy` and metrics all excluded with owners named |
|
||||
| 40 | [Shutdown drain guarantee](language-runtime-database/40-shutdown-drain-guarantee.md) | ✅ **LANDED 2026-08-27 — chain 3, with 31; split out of 24.** One rule: **a message sent before the stop flag is observed must be delivered and run before the engine stops.** Found by measurement, not review: making the chat gate's drain leg start its OWN (cold) server exposed that **5 of 16** fresh-server SIGTERM drains left a WebSocket client at EOF with no close frame and no diagnostic. Traced to `shard_main` — `NEXT_RUNNABLE()` already stated the contract ("a WORKER on stop keeps DRAINING … close frames!") but the IDLE branch reaped and broke, abandoning its inbox for teardown to free. An actor between messages is exactly that idle case, which is why a WARM soak server hid it for so long. Fix is one branch honouring the primary's drain window, yielding on an empty poll. **20 of 20 clean after**; `just chat` 11 checks 0 failures at the full 1000-client soak (which also settled the fd question: 1000 connections left the count at 44); runtime battery 36 suites 0 fail, compiler 556 checks 0 fail. Ruled out: a bigger spin (a 1 s wall-clock deadline still failed 2 of 12) and spawn-during-shutdown. Outstanding: a pin below the gate — nothing in `runtime/test/` drives the engine start/stop and no corpus fixture can trigger a stop |
|
||||
| 37 | [wo-html components](language-runtime-database/37-wo-html-components.md) | ✅ off-chain — LANDED 2026-08-25. Raw text literal (backtick, margin stripped at lex time, `{{ }}` auto-escapes) + the component layer: `Component`/`render_all`/`Layout` in wo-html, `ok_html` moved into the framework, site and shop both migrated |
|
||||
| 35 | [net runtime seams](language-runtime-database/35-net-runtime-seams.md) | ⬜ off-chain — fd deadlines on the park plane, Unix sockets, peer address; owns the ledger's three 🔧 rows (story written 2026-08-22) |
|
||||
| 20 | [Cross-program tables](databasev2/09-cross-program-tables.md) | ⏸ hold (2026-08-21); channel done (branch ipc-attach keeps its manifest) |
|
||||
|
|
|
|||
|
|
@ -0,0 +1,158 @@
|
|||
---
|
||||
iteration: "40"
|
||||
status: done
|
||||
chain: 3
|
||||
---
|
||||
|
||||
# iteration 40 — the shutdown drain guarantee: a send before the stop flag is delivered
|
||||
|
||||
> Part of [Story — one language, one runtime, one database, one binary](00-story.md).
|
||||
>
|
||||
> **Split out of [24](24-chat-websocket-workload.md) on 2026-08-27** because it
|
||||
> is a runtime *semantic*, not a task in a sample's gate. It belongs to the
|
||||
> actor lifecycle ([31](31-actor-lifecycle.md), absorbed into 24) and it is
|
||||
> the half of "lifecycle" that nothing had stated: 31 gave actors a death
|
||||
> notice, this gives the program a shutdown that does not lose mail.
|
||||
>
|
||||
> **Found by measurement, not review.** The chat gate's drain leg had been
|
||||
> passing only because it drained a server the 1k soak had already warmed.
|
||||
> Making every leg start its own server exposed it:
|
||||
> [`2026-08-27-chat-drain-finding.md`](../../2026-08-27-chat-drain-finding.md).
|
||||
|
||||
## The rule
|
||||
|
||||
**A message sent before the stop flag is observed must be delivered and run
|
||||
before the engine stops.** One sentence, and it is the whole iteration. It is a
|
||||
guarantee, not a tuning parameter — which is why a spin count could never
|
||||
express it.
|
||||
|
||||
What it does *not* promise: that a message sent *after* the flag is delivered,
|
||||
that a parked fiber is resumed, or that an actor gets unbounded time. The drain
|
||||
window is the primary's, and it closes when the primary returns.
|
||||
|
||||
## The bug, as measured
|
||||
|
||||
Fresh server, two WebSocket clients, `SIGTERM`, both must receive a close frame:
|
||||
|
||||
| Sample | Result |
|
||||
| --- | --- |
|
||||
| 5 fresh servers | 1 failure (`eof\|close`) |
|
||||
| 12 fresh servers | 3 failures, one `eof\|eof` |
|
||||
| 16 fresh servers | 5 failures |
|
||||
|
||||
The failing client's socket reaches EOF with **no close frame and no
|
||||
diagnostic** — the process exits and the kernel closes the fd.
|
||||
|
||||
Traced with instrumentation on the sample's actors: `main` → Registry → Room →
|
||||
Writer. The Registry runs and sees its room. The **Room never processes the
|
||||
shutdown message**, so the Writer's close branch never runs. Clients that did
|
||||
get a frame were saved by their own Reader noticing `env.stopping()`, not by the
|
||||
room broadcast.
|
||||
|
||||
## The design, as built
|
||||
|
||||
`runtime/src/vm.c` already encoded the correct contract in `NEXT_RUNNABLE()`:
|
||||
a worker that takes a stop while it has a live fiber returns 2 and **keeps
|
||||
draining its inbox** until the primary sets `eng_shutdown`. Its comment says so
|
||||
in as many words — "queued shutdown messages (close frames!) still run".
|
||||
|
||||
`shard_main`'s own idle branch contradicted it. A worker with an empty run queue
|
||||
waits in `wo_io_wait`, and on `WO_IO_STOP` it called `fib_reap_all` and
|
||||
**broke** — abandoning whatever was still in its inbox, which `wo_engine_stop`
|
||||
then freed wholesale during teardown.
|
||||
|
||||
So the failure needed a shard that was *idle* at `SIGTERM`. A Room actor between
|
||||
messages is exactly that, which is why the warm soak server hid it: warm shards
|
||||
had live fibers and took the correct path.
|
||||
|
||||
The fix makes the idle branch obey the same contract: while the primary's drain
|
||||
window is open, an idle worker adopts its inbox and runs what arrives, yielding
|
||||
between empty polls so a drain cannot become a hot spin across every core. Only
|
||||
`eng_shutdown` — set by the primary after `main` returns — ends it.
|
||||
|
||||
One branch, in one place, matching a contract the file already stated.
|
||||
|
||||
## Progress
|
||||
|
||||
| Piece | State |
|
||||
| --- | --- |
|
||||
| the idle-worker drain branch in `shard_main` (`runtime/src/vm.c`) | ✅ one branch, matching the contract `NEXT_RUNNABLE()` already stated |
|
||||
| `sched_yield` on an empty poll so the drain cannot hot-spin | ✅ |
|
||||
| fresh-server drain, repeated | ✅ **20 of 20**, from 5-in-16 failing |
|
||||
| chat gate at the default 1k soak | ✅ **11 checks, 0 failures** — 1000/1000 clients, both `WO_IO` backends, ASan clean |
|
||||
| full runtime battery (this touches every actor program's shard loop) | ✅ **36 suites** (18 × both dispatch flavors), 0 fail, `cli_smoke: OK`; compiler 556 checks 0 fail |
|
||||
| the regression pin | ✅ the chat gate's drain leg, now that it starts its OWN (cold) server — that decoupling is what caught this. **Not** a corpus fixture or unit test: nothing in `runtime/test/` drives `wo_engine_start`/`wo_engine_stop` today, and no corpus fixture can trigger a stop, so pinning it below the gate means new multithreaded test infrastructure — named as its own cost, not smuggled in here |
|
||||
|
||||
**Measured 2026-08-27.** Before: 5 of 16 fresh-server drains left a client at
|
||||
EOF. After: **20 of 20 clean.** At the observed failure rate, 20 clean runs by
|
||||
luck would be about 0.04%, so this is the fix rather than a quieter race.
|
||||
|
||||
## Acceptance Criteria
|
||||
|
||||
Met:
|
||||
|
||||
- **Given** a fresh server with two connected WebSocket clients, **when** it is
|
||||
sent `SIGTERM`, **then** both clients receive a close frame — **repeatedly**,
|
||||
not once. The bug reproduced at 5 in 16, so a single green run proves nothing;
|
||||
the criterion is a run of at least 16 with zero failures.
|
||||
✅ **20 of 20**, from 5-in-16 failing. A single run would have proved nothing.
|
||||
- **Given** an actor whose shard is idle at the moment of the stop, **when** a
|
||||
message is sent to it before the stop flag is observed, **then** its
|
||||
`receive` runs before the engine stops. ✅ this is exactly the case that
|
||||
failed — the Room between messages — and it is what the branch now covers.
|
||||
- **Given** the drain window, **when** a worker has nothing to adopt, **then**
|
||||
it does not hot-spin. ✅ `sched_yield()` on an empty poll; the 1k soak's RSS
|
||||
and timing legs are unchanged (marker reached all 1000 in 28 ms).
|
||||
- **Given** `just chat`, **when** it runs at the default soak, **then** all
|
||||
legs pass on both `WO_IO` backends and under the ASan build with zero leaks.
|
||||
✅ 11 checks, 0 failures. The fd leg also settled the lazy-init question at
|
||||
scale: **1000 connections left the count at 44**, unchanged after 20 more.
|
||||
- **Given** the full runtime battery, **when** it runs, **then** no suite
|
||||
regresses — this touches the shard loop every actor program uses. ✅ 36 suites
|
||||
0 fail, plus the compiler's 556 checks.
|
||||
- **Given** a program with no worker shards (`WO_SHARDS=1`), **when** it stops,
|
||||
**then** behaviour is unchanged. ✅ the gate's `WO_SHARDS=1` leg passes, and
|
||||
the branch is unreachable there — `wo_engine_stop` returns early at
|
||||
`nshards <= 1`, so a single-shard program never enters a worker loop.
|
||||
|
||||
Outstanding:
|
||||
|
||||
- **A pin below the gate.** The guarantee is currently proven by the chat gate
|
||||
only. Nothing in `runtime/test/` drives `wo_engine_start`/`wo_engine_stop`,
|
||||
and no corpus fixture can trigger a stop, so pinning it lower means new
|
||||
multithreaded test infrastructure. Named as its own cost rather than assumed
|
||||
cheap.
|
||||
|
||||
## Out Of Scope
|
||||
|
||||
- **Unbounded drain.** The window is the primary's and closes when `main`
|
||||
returns. A program that wants longer holds the window open itself.
|
||||
- **Delivering sends issued *after* the stop flag.** Nothing promises that, and
|
||||
promising it would mean a program could refuse to exit.
|
||||
- **Resuming parked fibers on stop.** `WO_SYS_STOPPED` unwinds them; that
|
||||
contract is iteration 24's and stays.
|
||||
- **A shutdown acknowledgement in the language surface.** The alternative fix
|
||||
was a barrier the sample builds itself, rejected below.
|
||||
- **`main` parking after the stop flag.** Still forbidden — a park after the
|
||||
flag unwinds. `main` still spins; the point is that spinning now works
|
||||
because the workers cooperate.
|
||||
|
||||
## Info — the forks, settled
|
||||
|
||||
1. **Engine guarantee, not a sample barrier.** The alternative was an
|
||||
acknowledged drain: rooms confirm back to `main`, which waits. Rejected —
|
||||
`main` cannot park after the stop flag, so it could only spin on the
|
||||
acknowledgement anyway, and every future actor program would have to
|
||||
re-implement the same handshake to avoid losing mail. A guarantee is stated
|
||||
once; a barrier is re-invented per program.
|
||||
2. **Not the spin budget.** Replacing the sample's `spin < 20000000` with a 1 s
|
||||
wall-clock deadline still failed 2 of 12. More time cannot help when the
|
||||
shard is not scheduled at all, and the reverted attempt cost a fixed second
|
||||
on every shutdown. Recorded because a bigger spin is the obvious wrong fix.
|
||||
3. **Not `dummy_writer()`.** Hoisting the shutdown message's placeholder actor
|
||||
out of the drain path (it spawned during shutdown) left 5 of 16 failing.
|
||||
4. **Yield rather than spin in the idle drain.** A worker polling an empty
|
||||
inbox in a tight loop would burn a core per shard during the window and
|
||||
starve the actors being drained.
|
||||
5. **Chain position 3**, with [31](31-actor-lifecycle.md): it is lifecycle
|
||||
semantics, and [24](24-chat-websocket-workload.md)'s gate is what proves it.
|
||||
|
|
@ -447,6 +447,33 @@ static void *shard_main(void *arg) {
|
|||
} else {
|
||||
int rc = wo_io_wait(vm); /* parked fibers AND the wake eventfd */
|
||||
if (rc == WO_IO_STOP) {
|
||||
/* iteration 40 — THE DRAIN GUARANTEE. A message sent before
|
||||
* the stop flag is observed must be delivered and run before
|
||||
* the engine stops.
|
||||
*
|
||||
* NEXT_RUNNABLE() already states this contract for a worker
|
||||
* holding a live fiber: it returns 2 and keeps draining "so
|
||||
* queued shutdown messages (close frames!) still run". This
|
||||
* branch — the IDLE worker, empty run queue, waiting on the
|
||||
* plane — used to reap and break instead, abandoning whatever
|
||||
* sat in its inbox for wo_engine_stop() to free wholesale.
|
||||
*
|
||||
* An actor between messages is exactly that idle case, which
|
||||
* is why a WARM server hid the bug: warm shards had live
|
||||
* fibers and took the correct path. Measured 2026-08-27 on a
|
||||
* fresh server: 5 of 16 SIGTERM drains left a WebSocket
|
||||
* client at EOF with no close frame and no diagnostic.
|
||||
*
|
||||
* The window belongs to the PRIMARY and closes when it sets
|
||||
* eng_shutdown (after main returns), so honour it here and
|
||||
* only exit when the primary says so. Yield on an empty poll:
|
||||
* a tight loop would burn a core per shard and starve the very
|
||||
* actors the drain exists to let run. */
|
||||
if (!eng_shutdown) {
|
||||
(void)wo_vm_adopt(vm);
|
||||
if (!vm->qhead) sched_yield();
|
||||
continue;
|
||||
}
|
||||
fib_reap_all(vm);
|
||||
break;
|
||||
}
|
||||
|
|
|
|||
Loading…
Reference in a new issue