docs(24): T10 closeout — stories done, board, graph, ledger, CODE-LOGIC

Iteration 24 closes, absorbing 31 and 34. No code in this commit.

- stories 24, 31, 34 -> `status: done`, each with a landing banner. 24's
  records the gate numbers and BOTH disclosed deviations: monitor takes
  three arguments (the caller may be `main`, which has no mailbox) and a
  v1 `call` reply is a typed scalar (which is what let the agreement be
  checked at compile time, WO-E226). 31's notes it landed INSIDE 24 and
  that a fifth mechanism it never anticipated came out of proving the
  gate — the drain guarantee (40). 34's names the gap it did NOT close:
  still no RNG, so CSRF/sessions stay blocked
- board: in-progress row cleared, marker doc deleted (convention), the
  standup entry in the six-question shape, chain note — next link is
  databasev2 4 (io_uring group-commit, chain 5)
- graph: PUBSUB2 (pub/sub + WebSockets, "rejected until here") -> done
- porch ledger: a WebSocket/pub-sub row added; the cancellation row now
  says what it actually waits on rather than repeating "the arc"; the
  README's "no WebSockets/SSE" limitation was stale — WebSockets are
  supported, SSE and chunked encoding are not
- CODE-LOGIC: runtime/src gains the actor-lifecycle section (call, death,
  the cap counter's sender/home-thread split, the monitor walk, the timer
  list), the drain guarantee, and the digest section; docs/examples/chat
  gains its own — actor topology, WHY two actors per connection, fd
  ownership, and the shutdown choreography

Battery after the doc edits: wovm-test 36 suites 0 fail, woc-test exit 0,
oop-e2e 119/0, chat 11/0, web-app 46/0, linkcheck clean.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
shoney.arickathil 2026-08-28 00:06:31 +02:00
parent 7833dd5740
commit 62d29d6a77
10 changed files with 311 additions and 102 deletions

View file

@ -128,6 +128,7 @@ flowchart TD
classDef rt fill:#8250df,color:#fff,stroke:none
classDef gated fill:#eac54f,color:#000,stroke:none
classDef v2 fill:#0969da,color:#fff,stroke:none
classDef done fill:#1a7f37,color:#fff,stroke:none
I7b2["7b per-shard collector (done — the precondition 8 waited on)"]:::rt
I8x["8 shard-actor runtime: thread-per-core, ownership-move messages"]:::rt
@ -140,7 +141,7 @@ flowchart TD
STREAM2["request body streaming + backpressure"]:::gated
SRESP2["streaming responses + explicit commit point"]:::gated
CANCEL2["per-request cancellation propagation"]:::gated
PUBSUB2["pub/sub + WebSockets (rejected until here)"]:::gated
PUBSUB2["DONE 2026-08-27 — pub/sub + WebSockets (iteration 24: ws_accept + wsframe + room actors)"]:::done
ASYNC9C["20 async attach statements (rejected-for-now alternative)"]:::gated
TIMEOUTS2["idle timeouts become schedulable (net seam still needed)"]:::gated

View file

@ -1,86 +0,0 @@
---
slice: "24" # the story that owns the status; see stories/24-chat-websocket-workload.md
status: in-progress
---
# Active slice — chat + actor lifecycle (iteration 24, absorbing 31 + 34)
Branch `chat-ws-lifecycle`. Spec:
[`superpowers/specs/2026-08-23-chat-websocket-actor-lifecycle-design.md`](superpowers/specs/2026-08-23-chat-websocket-actor-lifecycle-design.md)
· plan:
[`superpowers/plans/2026-08-23-chat-ws-lifecycle.md`](superpowers/plans/2026-08-23-chat-ws-lifecycle.md)
· board: [`stories/00-status.md`](stories/00-status.md).
## Progress (2026-08-23)
- ✅ **T1 crypto** (`d14fa9f`): sha1/sha256/hmac_sha256, ids 85–87, RFC
vectors 18/0, corpus pin. Story 34's C-builtin resolution delivered.
- ✅ **T2 bounded mailboxes** (`92754a8`): cap 1024 + `WO_MAILBOX`,
sender-side atomic reserve, WO_T_ACTOR (trap 13) catchable. Plus a
pre-existing compiler fix: try-arm Text places (bare `e.msg`) now
copy before the arm's scope dies (was ASan use-after-free + SEGV).
- ✅ **T6 WS upgrade** (`79cfa01`): `ws_accept` + accept-key + the
101 hijack sentinel; plain HTTP byte-identical (web-app 26/26).
- ✅ **T7 frame codec** (`7ad2ced`): pure-`.wo` RFC 6455 parse/serialize,
probe-verified against the RFC's own bytes.
- ✅ **T3 call/reply** (`ed69841`): `call` parks + typed scalar reply
(WO-E226 through actor-M erasure); actor DEATH landed with it —
callers never hang (mid-call + to-dead both trap catchably). Fixed
TRAPF's fiber-death leak/dangle en route.
- ✅ **T4 monitor + T5 time.after** (`56fe41a`, ids 89/90): the lifecycle
core. Corpus fixtures monitor-death, timer-delivery, timer-generation.
- 🔄 **T8 chat sample + T9 gate** (`6d729cc`, then `bbe0216`): the sample
and all five gate legs exist and run.
Every landed task: full battery 12/12, fresh-built.
## Verified 2026-08-27 (branch merged up to master)
Merged `master` in (clean; the porch rename means chat now says `use porch/...`
and its `[deps]` key is `porch`). Baseline on this branch: **18 runtime suites
× both dispatch flavors, 0 fail, `cli_smoke: OK`.**
`just chat` at `CHAT_SOAK=20` — **11 of 12 legs green**, including the two the
plan required and the gate was missing (`WO_SHARDS=1`, `WO_MAILBOX=8`).
**Three of the four failures found on 2026-08-27 were stale build artifacts,
not code.** Switching branches leaves `compiler/_build/` and `runtime/build/`
holding the *other* branch's binaries: a `woc` emitting `.wob` v7 against a
runtime expecting v6 reports only `wovm: unsupported version 7`, which the gate
surfaces as "no listener". `runtime/build/wovm_asan` bit the same way. **Rebuild
both after any branch switch** (`just woc-build`, `make -C runtime wovm-asan`)
before believing a gate failure.
**The remaining failure was a real bug and is now FIXED** (iteration 40) —
[`2026-08-27-chat-drain-finding.md`](2026-08-27-chat-drain-finding.md). On a
*fresh* server the SIGTERM drain leaves a client at EOF with no close frame in
5 of 16 runs. Traced: main → Registry → Room → Writer; the Registry runs but
the **Room never processes its shutdown message**, so the Writer's close branch
never runs. Ruled out: the spin budget (a 1 s wall-clock deadline still failed
2 of 12), `dummy_writer()` spawning during shutdown, and write failure. The
gate had been hiding it by draining a server the soak had already warmed.
## Pending
- ✅ **The drain guarantee — FIXED, and split into its own iteration**
([40](stories/language-runtime-database/40-shutdown-drain-guarantee.md),
chain 3 with 31). It was a runtime semantic, not a task in a sample's gate.
Root cause: `NEXT_RUNNABLE()` already stated the contract — "a WORKER on stop
keeps DRAINING … so queued shutdown messages (close frames!) still run" — but
`shard_main`'s IDLE branch contradicted it, reaping and breaking on
`WO_IO_STOP` and abandoning its inbox for teardown to free. An actor between
messages is exactly that idle case, which is why a warm soak server hid it.
One branch now honours the primary's drain window, yielding on an empty poll
so the drain cannot starve the actors it exists to let run. **20 of 20 fresh
server drains clean, from 5 in 16 failing.**
- ⬜ **T9 remainder** — the 1k soak has only been run trimmed
(`CHAT_SOAK=20`); run it at the default 1000 once the drain is fixed.
- ⬜ **T10 closeout** — stories 24/31/34 → `status: done` with banners (note the
scalar-reply v1 narrowing + three-argument monitor deviations), board
standup entry, graph nodes, framework README ledger rows, runtime +
chat CODE-LOGIC sections, delete this marker. Final battery.
This file is deleted when the slice lands (board convention). It lives flat in
`docs/` rather than a status folder — since 2026-08-26 no directory in this repo
encodes state; `status:` above is the only place it is recorded.

View file

@ -0,0 +1,82 @@
# `docs/examples/chat` — how the sample is put together
Iteration 24's acceptance workload: rooms, presence and broadcast over
WebSocket, actors on fibers across shards, one binary, no broker. It exists to
*drive* the actor work, so nearly every shape here is chosen to exercise
something the runtime claims.
Gate: `just chat` (`scripts/chat-accept.sh`), which logs to `/tmp/chat.log` —
`tail -F` it while the gate runs.
## The actors
| Actor | Owns | Answers |
| --- | --- | --- |
| `Registry` | name → room map, a fallback room | a `call` returning the room's address; spawns rooms on demand |
| `Room` | its member list (writer address + name) | join, leave, a text line, shutdown |
| `Reader` | the read half of one connection | nothing — it loops on the fd and sends onward |
| `Writer` | the **fd**, and the write half | text, pong, close |
| `ConnWorker` | one accepted connection | runs the HTTP layer over that fd |
`Registry` is the first honest consumer of `call`: the handler runs on the
connection worker's shard, the registry lives wherever placement put it, and
the reply is a scalar — the room's address. That is the cross-shard `call`
proof the gate asserts, not a contrivance added for it.
## Two actors per connection, not one
One fd, two directions, and they block independently. A single actor would have
to be inside `read` to notice the client, and inside `write` to deliver a
broadcast — it cannot be in both, so a broadcast would stall behind a quiet
client's read. Splitting them buys three things:
1. **The `Writer` is the sole writer of that fd.** Frames can never interleave,
which for a framed protocol is a correctness property and not a nicety.
2. **The `Reader` may block as long as it likes.** It sits in `read_dl` with a
30 s idle deadline and nothing else is waiting on it.
3. **The `Writer`'s mailbox becomes the backpressure point.** A slow client
stops draining its socket, its `Writer` blocks in `write_dl`, its mailbox
fills, and the room's next broadcast to it raises a catchable `WO_T_ACTOR`.
The room catches that and drops the member. **This is the whole reason the
mailbox cap is fail-fast** — the room survives its slowest member, and the
gate's `WO_MAILBOX=8` leg proves the path fires rather than assuming it.
`Room.say` is written around that: it shifts every member, tries the send, and
keeps only the members whose send succeeded — a failed one is sent a close and
dropped. So fan-out and eviction are the same pass.
## Who owns the fd
The `Writer`. It closes it, in every branch: a failed write sets `dead` and
closes; a close message writes the close frame and closes. The `Reader` closes
the fd itself in exactly one case — when its `send_close` to the writer traps,
meaning the writer is unreachable and nobody else will. Without that the fd
would leak on a dead-writer path.
`Writer.dead` guards against a second close, which matters because two
independent paths can decide a connection is finished (the reader seeing EOF,
and the room broadcasting shutdown).
## Shutdown choreography
On `env.stopping()` the accept loop stops and `main` sends one message to the
`Registry`, which fans out to every room; each room shifts its members and
sends each `Writer` a close; each writer writes the close frame and closes the
fd. `main` then spins — it may **not** park, because a park after the stop flag
unwinds — and returns, which is what stops the engine.
Independently, every `Reader` notices `env.stopping()` at its loop head and
runs its tail: leave the room, close the writer.
Both paths exist and that is deliberate: the reader path covers a connection
whose room is already gone, the room path covers a reader parked in a read that
has not come back yet.
**This is where iteration 40 came from.** The room path used to be unreliable:
a `Room` whose shard was idle at `SIGTERM` never adopted the shutdown message,
because an idle worker abandoned its inbox on stop. Clients that still got a
close frame were being saved by the reader path alone — which is why the
failure looked random and why a warmed-up server hid it. The engine now
guarantees that a send issued before the stop flag is delivered, so both paths
work as written. Nothing in this file changed to fix it, and that is the point:
the sample was right and the runtime was not.

View file

@ -56,7 +56,8 @@ place it runs.
feature.
- **Why `main` waits.** `main` is not an actor and has no mailbox, so it sleeps
rather than awaiting — the gap iteration 31's `call` closes for actors and
[24's marker](../../active-slice-2026-08-23-chat-ws-lifecycle.md) tracks.
[iteration 24](../../stories/language-runtime-database/24-chat-websocket-workload.md)
landed 2026-08-27.
Reasoning under the engine side: [`database/src/CODE-LOGIC.md`](../../../database/src/CODE-LOGIC.md).
Contract: [`plan/oop-vm/04-db-binding.md`](../../plan/oop-vm/04-db-binding.md).

View file

@ -75,7 +75,11 @@ porch = { git = "https://github.com/shoneyj/porch", rev = "v0.1.0" }
- **TLS: none, anywhere.** Deploy behind nginx/caddy; the proxy terminates
TLS+ALPN and gives browsers HTTP/2 while this backend speaks HTTP/1.1
keep-alive. See the web-app sample's README for the nginx sketch.
- `Content-Length` bodies only (no chunked encoding), no WebSockets/SSE,
- `Content-Length` bodies only (no chunked encoding); **WebSockets ARE
supported since 2026-08-27** — `ws_accept` (`http/ws.wo`) performs the RFC
6455 handshake and hands back the hijacked `net.Conn`, and `http/wsframe.wo`
is a pure-`.wo` frame codec; `docs/examples/chat` is the worked example and
`just chat` its gate. **SSE is still absent**, and so is chunked encoding.
JSON-first (no templates). Form-encoded bodies parse through
`form_values(req)` (`+` and `%XX` decoded, nil on any other
content-type); multipart/form-data through `multipart_parts(req)`
@ -130,6 +134,7 @@ first (pure `.wo` cannot express it yet).
| Content negotiation | ✅ `media_type(req)` request-side; `accepts(req, mtype)` response-side (exact, type/*, */*; q-values stripped not ranked — ranking waits for an app serving alternates) — slice 2 |
| Trusted-proxy client IP | 🔶 `client_ip(req)` parses X-Forwarded-For; `net.peer(fd)` (iteration 35) exposes the peer — the verify middleware is now a pure-`.wo` candidate slice |
| Status/header setting · redirects | ✅ builders + `set_header` |
| WebSockets · pub/sub | ✅ **2026-08-27 (iteration 24)** — `ws_accept` does the RFC 6455 handshake and hands back the hijacked `net.Conn`; `http/wsframe.wo` is a pure-`.wo` frame codec. Rooms/presence/broadcast are actors in `docs/examples/chat`, gated by `just chat` (11 checks, 1000-client soak, both `WO_IO` backends, ASan clean). No SSE |
| Lazy body streaming + backpressure · streaming responses · explicit commit point | ⏸ UNBLOCKED by the arc (8/11 landed 2026-08-21) — stays parked until its own slice |
| ETag + conditional requests | ✅ `etag_for` (quoted base64 SHA-256) + `with_etag` (If-None-Match → 304) over iteration 34's digest builtins — slice 2 |
@ -140,7 +145,7 @@ first (pure `.wo` cannot express it yet).
| Ordered middleware chain | ✅ registration order, `?Resp` short-circuits |
| Request-scoped context | ✅ `req.ctx` map (slice 2): middleware writes, handlers read; identity stays in `principal` |
| Guaranteed teardown | 🔶 every fd closes on every path (gate-proven); no user teardown hooks yet |
| Cancellation into pending storage ops | ⏸ UNBLOCKED by the arc (8/11 landed 2026-08-21) — stays parked until its own slice |
| Cancellation into pending storage ops | ⏸ **unblocked, not built.** The arc landed 2026-08-21 and iteration 24 (2026-08-27) added the lifecycle a cancellation would ride — `call` with a catchable trap when the callee dies, bounded mailboxes, `monitor`, and `time.after` for a deadline. Nothing here consumes them yet; it stays parked until its own slice |
| Panic recovery | 🔶 trap = 500 and the server survives ✅; "rolls back the transaction" is framework v2 (needs `transaction { }`, iteration 18) |
### Storage integration (the differentiator — framework v2 territory)

View file

@ -50,6 +50,56 @@ behind this board; live Obsidian Dataview views:
## ▶ NEXT PLAN
### Landed 2026-08-27 — iteration 24, chat + actor lifecycle (absorbing 31 + 34)
**Implemented last time (2026-08-27):** the slice closed and merged to master
(`ed5334d`, fast-forward). T4 `monitor` + T5 `time.after` (ids 89/90) had
landed on the branch; this session merged master in (adopting the `porch`
rename), finished T8/T9, fixed the gate, found and fixed a runtime bug, and did
T10. Iterations 31 and 34 land inside it.
**Key findings (measured, not asserted):** finishing the gate mattered more than
finishing the sample. Making **every leg start its own server** — instead of the
drain leg inheriting the soak's warmed one — exposed that **5 of 16**
fresh-server SIGTERM drains left a client at EOF with no close frame and no
diagnostic. Traced to `shard_main`: `NEXT_RUNNABLE()` already stated the
contract ("a WORKER on stop keeps DRAINING … close frames!") but the **idle**
branch reaped and broke, abandoning its inbox. An actor between messages is
exactly that idle case. Split out as
[40](language-runtime-database/40-shutdown-drain-guarantee.md); **20 of 20
clean** after. Also measured: the fd check had been core-count dependent — lazy
per-shard init takes one `io_uring` + one `eventfd` per shard, capped at
`nproc`, so 26 → 44 on a 20-core box read as a leak. **1000 connections left it
at 44**, which settled it.
**Learned:** three of the four gate failures were **stale build artifacts**, not
code. A branch switch leaves `compiler/_build/` and `runtime/build/` holding the
other branch's binaries, and a `woc` emitting `.wob` v7 against a v6 runtime
surfaces only as "no listener" — rebuild both before believing a gate failure.
And a gate that reuses another leg's server is not merely untidy: it hid a real
bug, and when its own leg failed it orphaned a listener that broke the *next*
run. Example apps now log to `/tmp/<app>.log` so a developer can `tail -F` them.
**Dependencies unblocked:** PUBSUB2 (WebSockets + pub/sub, rejected until this
point) is done; the porch ledger's WebSocket rows are ✅ and its cancellation row
is unblocked-not-built. Chain position 4 is complete, so **the chain's next link
is [databasev2 4](databasev2/04-io-uring-commit.md)** (io_uring group-commit).
Still blocked: CSRF and sessions — iteration 34 shipped HMAC but **there is
still no RNG**, and HMAC authenticates a token without being able to mint one,
which is [39](language-runtime-database/39-web-framework-parity.md)'s leading
item.
**Next steps:** databasev2 4, or databasev2 2's outstanding 5c/5d. One debt is
named rather than hidden: iteration 40's guarantee is proven only by the chat
gate — nothing in `runtime/test/` drives `wo_engine_start`/`wo_engine_stop` and
no corpus fixture can trigger a stop, so pinning it lower needs new
multithreaded test infrastructure.
**`.dev/reference` used:** none this slice. The sources were RFC 6455, RFC
3174/4231 for the digest vectors, and the kernel's own interfaces for the drain.
---
### Landed 2026-08-25 — packaging + release pipeline (off-chain, no story)
**Implemented last time (2026-08-25):** the toolchain became installable
@ -148,7 +198,7 @@ soak, runtime battery 36 suites 0 fail, compiler 556 checks 0 fail, corpus
119 checks 0 fail. Only **T10 closeout** remains — which is what still holds
stories 24/31/34 open. Finishing T9 exposed and fixed a real runtime bug,
split out as [40](language-runtime-database/40-shutdown-drain-guarantee.md). Its running state is the marker doc
([`2026-08-23-chat-ws-lifecycle.md`](../active-slice-2026-08-23-chat-ws-lifecycle.md)),
(the marker doc, deleted at closeout per the convention),
which is the file to read for what is done and what is next; stories
[31](language-runtime-database/31-actor-lifecycle.md) and
[34](language-runtime-database/34-crypto-builtins.md) keep
@ -355,8 +405,8 @@ that sequences its tasks. Read one, approve, then the next starts.
| 19 | [Float + Bytes](language-runtime-database/19-missing-scalar-types.md) | ✅ **landed 2026-08-20** — `.wob` v5: Float constant tag, field kinds 6/7, opcodes 34-41 (IEEE-quiet f64), builtins 70-83. Full stack: literals, arithmetic, `@table` column, WAL bit-exact replay, json fractions in / shortest-round-trip out, `?Float` reserved-NaN nil, total-order index (NaN last, `-0.0` == `+0.0`), Bytes + base64. No implicit Int/Float mixing (WO-E201); `float`/`trunc` are the only bridges. Proof: web-app price is a real Float (`{"price":9.99}`), `just web-app` 23/0; corpus 103/0 |
| 11 | [Fibers](language-runtime-database/11-fibers.md) | ✅ **landed 2026-08-21** with the arc (`just fibers` 10/0); fs-park re-scoped out of v1, disclosed in the story |
| 22 | [Durability, throughput, scale](language-runtime-database/22-durability-throughput-scale.md) | ✅ **landed 2026-08-21** — db-bench + baseline.json (74 metrics) + restart/kill -9 proofs both shard counts; durable 4.5k vs ram 297k inserts/s, reads O(table), msgrate 13.4M/2.45M |
| 31 | [Actor lifecycle](language-runtime-database/31-actor-lifecycle.md) | 🔄 **absorbed into 24** (directive 2026-08-23) and half landed there: `call` request/response with a typed scalar reply (`WO_B_CALL = 88`, WO-E226), bounded mailboxes (`WO_MAILBOX`, cap 1024, catchable `WO_T_ACTOR`), and actor death that traps callers instead of hanging them. Still open: `monitor` and `time.after` — ids **89 and 90 are reserved holes** in `wob.h`, which is the machine-checkable proof of what is left. Supervision trees stay out of v1 |
| 24 | [chat: WebSocket workload](language-runtime-database/24-chat-websocket-workload.md) | 🔄 **the live slice** (absorbing 31 + 34, directive 2026-08-23) — branch `chat-ws-lifecycle`, 5/10 tasks landed: crypto, bounded mailboxes, WS upgrade, frame codec, `call`/reply + actor death. Pending: `monitor`, `time.after`, the chat sample, its gate, closeout. State lives in [the marker](../active-slice-2026-08-23-chat-ws-lifecycle.md) |
| 31 | [Actor lifecycle](language-runtime-database/31-actor-lifecycle.md) | ✅ **LANDED 2026-08-27 inside 24** (directive 2026-08-23). All four mechanisms: `call`/reply with a typed scalar reply (`WO_B_CALL = 88`, WO-E226), bounded mailboxes (`WO_MAILBOX`, cap 1024, catchable `WO_T_ACTOR`), actor death that traps callers instead of hanging them, **`monitor` (89)** and **`time.after` (90)** — the reserved holes in `wob.h` are filled. A fifth mechanism it did not anticipate came out of proving the gate: the shutdown drain guarantee, [40](language-runtime-database/40-shutdown-drain-guarantee.md). Supervision trees stay out of v1 |
| 24 | [chat: WebSocket workload](language-runtime-database/24-chat-websocket-workload.md) | ✅ **LANDED 2026-08-27** (absorbing 31 + 34) — all ten tasks; merged to master `ed5334d`. `just chat` **11 checks, 0 failures** at the full 1000-client soak: handshake, functional matrix on both `WO_IO` backends and on one shard, the soak, the fd invariant, the SIGTERM drain, `WO_MAILBOX=8` backpressure, ASan clean. Finishing its gate found a real runtime bug, split out as [40](language-runtime-database/40-shutdown-drain-guarantee.md) |
| 23 | [io_uring group-commit](databasev2/04-io-uring-commit.md) | ⬜ fifth in chain, after stage 3 + 22 |
| 32 | [WAL checkpoint](databasev2/03-wal-checkpoint.md) | ⬜ last in chain, after 23 — disk reclamation + bounded replay (story written 2026-08-21) |
| 33 | [Single-file store](databasev2/07-single-file-db.md) | ⬜ off-chain, small — `WO_DATA=<path>.db` file form; driver-only (story written 2026-08-22) |
@ -387,13 +437,16 @@ that sequences its tasks. Read one, approve, then the next starts.
| Language | 🔄 [iteration 36 — operator parity](language-runtime-database/36-operator-parity.md): `not`, bitwise `& \| ^ << >>`, hex/binary/`_` literals, compound assigns — CODE LANDED 2026-08-22 (branch operator-parity, `.wob` v6, all gates green; reference project `.dev/reference/go` drove the design). Awaiting the developer's MANUAL pass on `docs/examples/operators/` (no test fixtures by directive); unblocks story 34's pure-`.wo` HMAC question | [plan](../superpowers/plans/2026-08-22-operator-parity.md) |
| Language | the framework v1-polish slice landed 2026-08-20 (branch framework-v1, awaiting merge); next per the order: brainstorm 20/21's forks | [order](#implementation-order-re-sequenced-2026-08-21--concurrency-chain) |
| Runtime | ✅ **iteration 35 landed 2026-08-23** (branch `framework-v1b`, with framework v1 slice 2 + the serving slice): net deadlines/unix/peer (ids 91–95), fiber pooling, serve_conn + web-app fiber-per-connection — web-app gate 41/0, both WO_IO backends | [design](../superpowers/specs/2026-08-23-net-seams-park-design.md) |
| Runtime | 🔄 **iteration 24 (absorbing 31 + 34): chat + actor lifecycle** — spec + plan approved 2026-08-23 (24 absorbs 31 by directive; 34 resolved C-builtins); executing on branch `chat-ws-lifecycle` | [marker](../active-slice-2026-08-23-chat-ws-lifecycle.md) · [plan](../superpowers/plans/2026-08-23-chat-ws-lifecycle.md) |
The active slice's marker doc is
[`docs/active-slice-2026-08-23-chat-ws-lifecycle.md`](../active-slice-2026-08-23-chat-ws-lifecycle.md)
— one file, deleted when the slice lands. Everything else pending is the
concurrency chain (see *Pending* below); the held tail is every story
whose frontmatter reads `status: hold`.
**No slice is active.** Iteration 24 landed 2026-08-27 and its marker doc was
deleted per the convention. Everything pending is the concurrency chain (see
*Pending* below) — **the chain's next link is
[databasev2 4](databasev2/04-io-uring-commit.md)** (chain 5, the io_uring
group-commit write path, `was_language_iteration: 23`), which now has iteration
22's fsync-per-commit numbers in hand, plus databasev2 1's finding that the
write path is *not* where memory pressure bites (appending under a cap costs
~1%, random reads 273×). The held tail is every story whose frontmatter reads
`status: hold`.
### Landed 2026-08-14 — the compile-and-run milestone

View file

@ -1,6 +1,6 @@
---
iteration: "24"
status: in-progress
status: done
chain: 4
---
@ -19,6 +19,34 @@ chain: 4
> bounded mailboxes, actor death, timers). Iteration 19 LANDED
> 2026-08-20, so Bytes is available for frame parse/serialize.
> **✅ LANDED 2026-08-27** (branch `chat-ws-lifecycle`, merged to master
> `ed5334d`). Ten tasks: crypto (T1), bounded mailboxes (T2), `call`/reply and
> actor death (T3), `monitor` (T4), `time.after` (T5), the WS upgrade seam
> (T6), the pure-`.wo` frame codec (T7), the chat sample (T8), the gate (T9),
> this closeout (T10). It absorbed [31](31-actor-lifecycle.md) and
> [34](34-crypto-builtins.md), which land with it.
>
> **Gate — `just chat`, 11 checks, 0 failures** at the full 1000-client soak:
> handshake with an independently recomputed accept-key, the functional matrix
> (presence, broadcast, room isolation, leave) on **both** `WO_IO` backends and
> on a single shard, the 1k hot-room soak, the fd invariant, the SIGTERM drain,
> `WO_MAILBOX=8` backpressure, and an ASan run with zero leaks. Battery
> alongside: runtime 36 suites 0 fail, compiler 556 checks, corpus 119 checks.
> The sample logs to `/tmp/chat.log`.
>
> **Two disclosed deviations from the spec.** `monitor` takes **three**
> arguments (`watched, observer, msg`) rather than two, because the caller may
> be `main`, which has no mailbox and cannot be an implicit observer. And a
> `call` reply is a **typed scalar** in v1 — which is what let the agreement be
> checked at compile time (WO-E226) instead of carried as a tagged value.
>
> **What finishing the gate found.** Making every leg start its own server
> exposed a real runtime bug the warmed soak server had been hiding: on a fresh
> server, 5 of 16 SIGTERM drains left a client at EOF with no close frame. It
> was not this sample's fault — the fix is an engine guarantee, split out as
> [40](40-shutdown-drain-guarantee.md). Design notes:
> [`docs/examples/chat/CODE-LOGIC.md`](../../examples/chat/CODE-LOGIC.md).
## Why this iteration exists
Everything the framework ledger parks behind concurrency — WebSockets,

View file

@ -1,6 +1,6 @@
---
iteration: "31"
status: refine
status: done
chain: 3
---
@ -16,6 +16,22 @@ chain: 3
> ([iteration 24](24-chat-websocket-workload.md)) cannot be written
> honestly without these four mechanisms.
> **✅ LANDED 2026-08-27 — INSIDE [24](24-chat-websocket-workload.md)**, per
> the 2026-08-23 directive that absorbed it. All four mechanisms shipped:
> `call`/reply with a typed scalar reply (id 88, WO-E226), **bounded mailboxes**
> (`WO_MAILBOX`, default 1024, fail-fast with a catchable `WO_T_ACTOR`),
> **actor death** that traps callers instead of hanging them, `monitor`
> (id 89) and `time.after` (id 90). Ids 89 and 90 were reserved holes in
> `wob.h`; they are filled.
>
> **A fifth mechanism was added that this story did not anticipate**: the
> shutdown drain guarantee, [40](40-shutdown-drain-guarantee.md). It is
> lifecycle semantics — this story gave actors a death notice, 40 gives the
> program a shutdown that does not lose mail — and it was found by measurement
> while proving 24's gate, not by review.
>
> How each piece works: `runtime/src/CODE-LOGIC.md`, "Actor lifecycle".
## Why this iteration exists
The arc's stages 1+2 shipped `spawn`/`send` mechanism without lifecycle:

View file

@ -1,6 +1,6 @@
---
iteration: "34"
status: refine
status: done
---
# Iteration 34 — crypto builtins: digests and HMAC in the runtime
@ -18,6 +18,18 @@ status: refine
> Off the concurrency chain but **gates chain position 4**: iteration
> 24's WebSocket handshake needs SHA-1 before chat can land.
> **✅ LANDED 2026-08-27 — inside [24](24-chat-websocket-workload.md)** as its
> task 1. The fork resolved to **C builtins**: `sha1` (85), `sha256` (86),
> `hmac_sha256` (87), each over one buffer returning a fresh `Bytes`. Pinned to
> the published vectors — RFC 3174, the SHA-256 vectors, RFC 4231 — in
> `runtime/test/test_crypto.c`, 18 checks, plus a corpus fixture hashing "abc"
> from `.wo`. This unblocked chain position 4: the WebSocket handshake needs
> SHA-1, and `just chat` verifies the accept-key independently.
>
> **The gap it did NOT close:** there is still no RNG in the runtime. HMAC
> authenticates a token and cannot mint one, so CSRF and sessions stay blocked
> — which is why [39](39-web-framework-parity.md) leads with a random-bytes
> builtin rather than treating them as unblocked.
## Why this iteration exists
Four consumers already wait on it, none able to proceed:

View file

@ -257,3 +257,100 @@ layout.
- **`listen_unix` sets O_NONBLOCK on the listener itself** — accept4's
SOCK_NONBLOCK flags the ACCEPTED socket only; a blocking listener
would block the whole shard (found by the seam probe, both backends).
## Actor lifecycle: call, death, monitor, timers (iteration 24, ids 88–90)
Four pieces that together answer "what happens to an actor that is waiting,
that dies, that watches, or that wants to be woken later". All four live in
`vm.c` with their entry points in `builtin.c`; the structures are in `vm.h`.
**`call` (id 88) — a send that waits.** An ordinary `send` returns immediately;
`call` parks the calling fiber and resumes it with the receive's return value.
The reply is a **typed scalar**, which is what let the agreement be checked at
compile time (WO-E226) rather than carried as a tagged value at runtime. The
caller is never left hanging: if the callee dies mid-call, or the address is
already dead, the caller **traps catchably** instead of parking forever. That
is the property worth keeping in mind when reading the code — every path out of
a call either resumes the fiber or traps it.
**Death.** A `receive` that traps uncaught marks the actor dead on its home
thread. From then on sends to it drop silently, calls trap, queued callers are
error-unparked, and its state and mailbox are released. Silent-drop for sends
is deliberate: a sender cannot handle another actor's failure, and making every
`send` fallible would put a `try` on every line.
**The mailbox cap and its counter.** One cap for every mailbox (default 1024,
`WO_MAILBOX` overrides at boot; the chat gate shrinks it to 8 to force the
policy). `pending` counts sent-but-not-delivered. It is incremented by the
**sender**, on any shard, and decremented by the **home thread** at delivery —
so it is touched only through `wo_mbox_reserve`/`wo_mbox_release` and their
`__atomic` builtins. The consequence is disclosed rather than hidden: the cap
can overshoot by at most the number of in-flight sends. Overflow is fail-fast —
the send raises a catchable `WO_T_ACTOR` (trap 13), which is what lets a room
drop a slow member instead of growing without bound.
**`monitor` (id 89) — the death notice.** `wo_monitor` is one registration:
observer, the moved-in notice message, next. The list lives on the **watched**
actor and is owned by its home thread, so the death walk needs no lock — dying
is a home-thread event and the list is right there. The notice is the
observer's own M-typed message, so an observer receives death notices in the
same shape as everything else. Monitoring an already-dead actor fires
immediately rather than silently doing nothing. An observer whose mailbox is
full loses the notice, with a disclosed stderr line — the alternative was
blocking a death walk on a slow observer.
It takes **three arguments** (`watched, observer, msg`), not the two the spec
first proposed, because the caller may be `main`, which has no mailbox and so
cannot be an implicit observer.
**`time.after` (id 90) — one-shot, no cancel.** `wo_timer` is `at` (wall ms),
target, message, next. The list lives on the **arming fiber's shard** and is
scanned by the same deadline machinery that already serves fd-park deadlines,
so timers cost no new wait mechanism. Firing is an ordinary runtime send, which
means it inherits the ordinary rules: a full target drops with a stderr line, a
dead target drops silently. There is no cancel; the idiom is a generation
counter in the message, which the `timer-generation` corpus fixture pins.
**Where to look when a lifecycle thing misbehaves:** `wo_vm_actor_monitor` and
`wo_vm_timer_after` in `vm.c` are the two entry points; `shard_main` and
`NEXT_RUNNABLE()` decide when a shard runs, adopts, or stops. The corpus
fixtures `monitor-death`, `timer-delivery` and `timer-generation` are the
smallest working examples of each.
## The shutdown drain guarantee (iteration 40)
**A message sent before the stop flag is observed is delivered and run before
the engine stops.** Stated because it was once untrue in a way nothing caught.
`wo_engine_stop` sets `eng_shutdown`, wakes every worker, joins them, and only
then tears down — freeing whatever envelopes are still queued. So a worker that
leaves its loop early takes its inbox with it. `NEXT_RUNNABLE()` has always
encoded the right behaviour for a worker holding a live fiber: on a stop it
returns 2 and keeps draining, because "only the PRIMARY's stop ends the
program". `shard_main`'s **idle** branch did the opposite — it reaped and broke
— so a shard whose actors happened to be between messages at `SIGTERM`
abandoned everything still in flight.
It now honours the same contract: while the primary's window is open an idle
worker adopts its inbox and runs what arrives, `sched_yield`ing on an empty
poll so a drain cannot burn a core per shard and starve the actors it exists to
let run. Only `eng_shutdown` — which the primary sets after `main` returns —
ends it.
Two things follow that are easy to get wrong. The window is the **primary's**,
so a program that wants a longer drain holds it open itself; `main` cannot park
after the stop flag, because a park there unwinds. And the whole path is
unreachable at `WO_SHARDS=1`, where `wo_engine_stop` returns at `nshards <= 1`.
## Digests: sha1, sha256, hmac_sha256 (iteration 34, ids 85–87)
`crypto.c` holds SHA-1 and SHA-256 over a single buffer and HMAC-SHA-256 on top
of the latter, each returning a fresh `Bytes`. No streaming API and no other
primitives — these exist because WebSocket's handshake needs SHA-1 and ETags
need SHA-256, and that is the whole of the demand so far.
Correctness is pinned to the published vectors rather than to itself:
RFC 3174 for SHA-1, the FIPS/RFC 6234 vectors for SHA-256, RFC 4231 for HMAC,
in `runtime/test/test_crypto.c` (18 checks). **There is still no RNG anywhere
in the runtime** — HMAC authenticates a token but cannot mint one, which is why
iteration 39 leads with a random-bytes builtin.