add wo-rt-c: C runtime prototype, phases A-F shipped

prototypes/wo-rt-c — single-file C reference of the writeonce runtime
layer, zero deps beyond libc + kernel uapi: thread-per-core io_uring
event loops (raw syscalls, no liburing), SO_REUSEPORT listeners, one
mlock'd mmap arena sharded by address, WAL dual-write with group commit
(HTTP ack only after the fsync CQE), boot-time snapshot + WAL replay
recovery, and a bench harness with a Go net/http comparison server.

Measured on 20 cores: 908k reads/s p99 154us and 643k fsync-acked
commits/s p99 177us on 8 shards, vs Go net/http 495k/355k (no
durability) on 20 cores. Crash-under-load testing found and fixed an
ack-before-fsync race and an fd-reuse ABA hazard in commit-ack parking.

docs/plan/exploration/c-runtime — the phased plan (00, exit evidence
per phase), the one-address architecture trace (01), and the
single-binary end-goal contract (02). justfile carries the demo and
bench recipes; .gitignore covers binaries and data dirs.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
This commit is contained in:
shoney.arickathil 2026-06-12 23:05:13 +02:00
parent c8f76484f4
commit 929b2802c4
14 changed files with 1777 additions and 1 deletions

11
.gitignore vendored
View file

@ -5,6 +5,12 @@
# C++ prototype build output # C++ prototype build output
/prototypes/*/build /prototypes/*/build
# C runtime prototype (prototypes/wo-rt-c): compiled binary + WAL/snapshot data.
# `just rt-c-demo` and the default WO_DATA write ./wo-data at the repo root.
/prototypes/wo-rt-c/wo-rt
/wo-data
/prototypes/wo-rt-c/wo-data
# Runtime data directories for the sample projects. # Runtime data directories for the sample projects.
# `wo.toml` points at `./data` which holds the per-project engine state. # `wo.toml` points at `./data` which holds the per-project engine state.
/data /data
@ -32,4 +38,7 @@
!/.vscode/settings.json.example !/.vscode/settings.json.example
!/.vscode/extensions.json !/.vscode/extensions.json
/prototypes /prototypes/wo-db
# Phase-F bench binaries (sources are committed; builds are not)
/prototypes/wo-rt-c/bench/bench
/prototypes/wo-rt-c/bench/goref/goref

8
.vscode/extensions.json vendored Normal file
View file

@ -0,0 +1,8 @@
{
// Workspace extension recommendations — VS Code offers to install these on open.
"recommendations": [
"bierner.markdown-mermaid", // renders ```mermaid blocks in the markdown preview
"rust-lang.rust-analyzer",
"humao.rest-client" // reference/rest/*.rest files
]
}

View file

@ -33,6 +33,8 @@ The repo holds three cuts of the same project plus one research reference:
There is also **`prototypes/wo-db/`** — a ~2k-line **C++ prototype** of the query-layer engine (SQL + Cypher + document paths, `RETURNING` aliases, `LIVE` stub). It keeps its `wo-db` directory name (C++ project, separate from the Rust crate `db`). It's the reference implementation the Rust port follows; `make test` still passes. There is also **`prototypes/wo-db/`** — a ~2k-line **C++ prototype** of the query-layer engine (SQL + Cypher + document paths, `RETURNING` aliases, `LIVE` stub). It keeps its `wo-db` directory name (C++ project, separate from the Rust crate `db`). It's the reference implementation the Rust port follows; `make test` still passes.
And **`prototypes/wo-rt-c/`** — a single-file **C reference of the runtime layer** (edge-triggered epoll loop, signalfd shutdown, non-blocking listener, in-RAM store over minimal HTTP; libc only). It's the runtime-layer sibling of `wo-db`: each block maps one-to-one to a `crates/rt/src/runtime/` module (table in its README). `make` builds it; `just rt-c-demo` exercises it. Its evolution into a multi-threaded io_uring RAM-database runtime (thread-per-core, mmap arena, WAL dual-write, recovery, ACID) is phased A–F in `docs/plan/exploration/c-runtime/00-plan.md`, with the one-address architecture trace beside it (`01-architecture.md`) — the proving ground for plans 09–12. **Documentation belongs under `docs/`** — prototype/crate directories keep only their orientation README.
## Commands ## Commands
```bash ```bash

View file

@ -116,6 +116,7 @@ If the "single core per process, shard across processes" argument ([Redis Cluste
## Cross-references ## Cross-references
- [`./exploration/c-runtime/00-plan.md`](./exploration/c-runtime/00-plan.md) — the C prototype's phased evolution (threads → arena → io_uring → WAL → recovery); the executable proving ground for 09a's thread-per-core skeleton and 09c's per-shard WAL before the Rust work starts.
- [`./08-sendfile-static-assets.md`](./08-sendfile-static-assets.md) — last prerequisite phase; feature-complete single-threaded runtime. - [`./08-sendfile-static-assets.md`](./08-sendfile-static-assets.md) — last prerequisite phase; feature-complete single-threaded runtime.
- [`./assembly/02-writeonce-stance.md`](./assembly/02-writeonce-stance.md) — updated to reference this phase's thread-per-core model; still no asm. - [`./assembly/02-writeonce-stance.md`](./assembly/02-writeonce-stance.md) — updated to reference this phase's thread-per-core model; still no asm.
- [`../runtime/database/02-wo-language.md#concurrency-model`](../runtime/database/02-wo-language.md#concurrency-model) — the stance this plan refines. - [`../runtime/database/02-wo-language.md#concurrency-model`](../runtime/database/02-wo-language.md#concurrency-model) — the stance this plan refines.

View file

@ -0,0 +1,86 @@
# wo-rt-c roadmap — multi-threaded io_uring RAM database runtime, in C
**Context sources:** [`prototypes/wo-rt-c/wo-rt.c`](../../../../prototypes/wo-rt-c/wo-rt.c) (phase 0 — the single-threaded epoll baseline), [`../../09-concurrency-scaleout.md`](../../09-concurrency-scaleout.md) (the thread-per-core doctrine every phase here miniaturizes), [`../../10-storage-foundations.md`](../../10-storage-foundations.md) / [`11-wal-and-recovery.md`](../../11-wal-and-recovery.md) / [`12-engine-disk-cutover.md`](../../12-engine-disk-cutover.md) (the storage track), kernel reference cards [`../linux/07-io_uring.md`](../linux/07-io_uring.md), [`08-mmap.md`](../linux/08-mmap.md), [`09-fallocate.md`](../linux/09-fallocate.md), [`12-pwrite-fsync.md`](../linux/12-pwrite-fsync.md), [`02-eventfd.md`](../linux/02-eventfd.md).
## Goal
Evolve the [`prototypes/wo-rt-c/`](../../../../prototypes/wo-rt-c/) prototype from a single-threaded epoll reference into a **multi-threaded runtime environment for writeonce applications**: thread-per-core io_uring event loops at million-scale read/write concurrency, the whole database resident in RAM (one mmap arena, addressed per shard — no duplication), **ACID** commits that dual-write RAM-first-then-disk, and a boot path that loads the hard drive's state back into RAM before serving. Still one C file's worth of honesty per concern, still **zero dependencies beyond libc** — raw io_uring syscalls, no liburing.
Each phase is the executable proving ground for the matching Rust plan (09–12): get the syscall sequence right here in a few hundred lines, then port with confidence.
## Design decisions (locked — by plan 09 and by the project owner)
1. **Thread-per-core, shared-nothing.** `WO_THREADS` (default `sysconf(_SC_NPROCESSORS_ONLN)`) pthreads, each pinned with `sched_setaffinity`, each owning its event loop, its listener, its connections, its shard. No work stealing, no connection migration, no locks or atomics on the data path (plan 09 decisions 1, 2, 7).
2. **One RAM arena, sharded by address — no duplication.** A single `mmap` region partitioned into N shard slices. Thread *t* touches only addresses inside slice *t*. Data exists exactly once in memory; ownership, not copying, is the concurrency model.
3. **Physical RAM via kernel primitives.** `MAP_POPULATE` to fault pages in up front, `mlock` to pin them (no swap on the read path), `MAP_HUGETLB` attempted with 4 KB fallback. The read path never touches a file descriptor.
4. **Per-thread io_uring ring, raw syscalls.** `io_uring_setup` + mmap'd SQ/CQ rings + `io_uring_enter`, `IORING_SETUP_SINGLE_ISSUER`. No ring sharing (plan 09 decision 4). No liburing — the zero-dep stance is the point of the prototype.
5. **ACID at per-shard scope.**
- *Atomicity* — WAL records are framed `len | crc32 | payload | COMMIT`; a record replays whole or not at all.
- *Consistency* — invariants checked in-thread before the RAM apply; an aborted op writes nothing.
- *Isolation* — one thread is one serial execution stream: per-shard serializability by construction.
- *Durability* — the HTTP ack is sent only after the fsync completion for the commit's WAL tick.
Cross-shard transactions (2PC) belong to the Rust track ([plan 09e](../../09-concurrency-scaleout.md)) — non-scope here.
6. **Dual-write order.** Commit = apply to the RAM slot → append WAL record (write SQE) → one `fdatasync` SQE per loop tick covering all of that tick's commits (group commit) → ack. Boot = replay per-shard WAL/snapshot from disk into the arena, in parallel, before listeners open.
## Phase sequence
One phase per implementation pass; `make` + the smoke endpoints stay green after every phase.
### Phase A — thread-per-core skeleton — ✅ shipped
*Maps to [plan 09a](../../09-concurrency-scaleout.md); cards: `SO_REUSEPORT`, [`eventfd`](../linux/02-eventfd.md).*
`WO_THREADS` pthreads, each pinned to its core; per-thread epoll loop (io_uring arrives in C), per-thread `SO_REUSEPORT` listener on the same port (kernel balances accepts by 4-tuple), per-thread `conns[]` and request counters. Shutdown: thread 0 owns the `signalfd`; on signal it writes each thread's `eventfd`, every loop exits, `pthread_join` all. Store stays per-thread arrays until B. Responses gain `"shard":t`; `/` reports `"threads":N` and per-shard counters.
**Exit (met):** all endpoints green at `WO_THREADS=4`; `/proc/<pid>/task/*/status` shows each thread pinned to its own core (tid→cpu 0,1,2,3); `/` counters proved 200 concurrent requests spread `[68,36,44,53]` across 4 shards; SIGTERM broadcast joined all shards cleanly; `WO_THREADS=1` behaves like phase 0 (shard 0, sequential ids). Note ids are now interleaved per shard (t+1, t+1+N, …) for coordination-free global uniqueness.
### Phase B — the RAM arena — ✅ shipped
*Maps to [plan 10](../../10-storage-foundations.md); cards: [`mmap`](../linux/08-mmap.md), [`fallocate`](../linux/09-fallocate.md).*
One arena: a header page (magic, version, shard count, slot geometry) + N shard slices of fixed-size row slots + a per-shard allocation bitmap. `mmap(MAP_ANONYMOUS|MAP_PRIVATE [|MAP_HUGETLB], MAP_POPULATE)` then `mlock` (graceful fallback + warning if `RLIMIT_MEMLOCK` refuses). Rows move from arrays into slots; every access is a typed pointer into the owning thread's slice — decision 2 made literal.
**Exit (met):** same API; `/` reports `"arena":{bytes, mapped, hugepages, mlocked, slot geometry}` + `shard_used[]`; `smaps_rollup` showed `Locked: 276 kB` exactly matching the mapping; hugepage attempt fell back to 4 K pages gracefully on a box with no hugepage pool; rows live at `(shard, slot)` addresses (`"slot":n` in create responses) — the stable coordinates phase D's WAL records will carry.
### Phase C — io_uring event loops + keep-alive — ✅ shipped
*Maps to plan 09 decision 4; card: [`io_uring`](../linux/07-io_uring.md).*
Replace each thread's epoll loop with a raw ring: `io_uring_setup`, mmap SQ/CQ, `io_uring_enter`; multishot accept, `recv`/`send` SQEs, `user_data = fd | (op << 32)`, the shutdown `eventfd` watched via a POLL_ADD SQE. Rewrite the connection state machine for **HTTP keep-alive** with a per-connection write queue — connection-per-request caps throughput far below the million-scale target. Largest phase (~400 LOC); the ring-setup block is the C reference the Rust port reads.
**Exit (met):** four requests over one socket (`curl` reported `num_connects: 1, 0, 0, 0`), and a create+list pair on one connection lands on the same shard with both rows visible; `strace -c` over 60 keep-alive requests showed **124 `io_uring_enter` and zero `epoll_wait`/`recvfrom`/`sendto`/`accept`** — the only `read`/`write` calls were the signalfd/eventfd shutdown path; SIGTERM broadcast joined all shards cleanly. Raw ring (`io_uring_setup` + SINGLE_MMAP rings + `io_uring_enter`), multishot accept with `CQE_F_MORE` re-arm, one outstanding SQE per connection, pipelined-tail carry-over.
### Phase D — WAL dual write: RAM first, then the hard drive — ✅ shipped
*Maps to [plan 11](../../11-wal-and-recovery.md) + 09c; cards: [`pwrite/fsync`](../linux/12-pwrite-fsync.md), [`fallocate`](../linux/09-fallocate.md).*
Per-shard `shard-<t>.wal`, preallocated with `fallocate`. The commit path is decision 6 verbatim: RAM apply → framed record append (write SQE at the shard's tail offset) → one `fdatasync` SQE per tick → ack on the fsync CQE. CRC32 hand-rolled. Acks for all commits in a tick ride the same fsync — group commit, exactly the Rust runtime's doctrine.
**Exit (met):** crash test — 60 concurrent POSTs, `kill -9` mid-stream — showed **60/60 acked writes present and CRC-valid in the WALs, zero acked-but-missing** (offline verification via the new `wo-rt wal-check <file>` mode, which is also phase E's replay skeleton); a copy truncated mid-record reported `TORN at byte 2584 — 17 whole records before it`, dropping the partial whole; 200 concurrent durable commits in 102 ms (~1,960 commits/s, curl-process-bound — each tick's fsync covers every commit staged that tick via double-buffered write→fsync `IOSQE_IO_LINK` pairs). Failed fsync closes the batch's connections without acking — a client never sees a 201 for a non-durable write. Phase D boots with `O_TRUNC` (fresh log); replay lands in E.
### Phase E — first load: hard drive → RAM — ✅ shipped
*Maps to plans [11](../../11-wal-and-recovery.md)/[12](../../12-engine-disk-cutover.md).*
Boot, before any listener opens: each thread replays its own WAL into its arena slice — parallel recovery — validating frame CRC + commit marker, truncating at the first torn record. Clean shutdown writes a snapshot (`pwrite` of live slots to `shard-<t>.data`) and truncates the WAL; boot prefers snapshot + WAL tail. Recovery time printed at startup.
**Exit (met):** the full cycle verified — (1) 40 writes + `kill -9` + restart: `shard_used [10,7,15,8]` identical, replayed from WAL alone (`recovered 0 snapshot rows + N wal records in 1 ms` per shard); (2) SIGTERM wrote four snapshots and truncated the WALs; (3) restart loaded the snapshots instantly; (4) 5 more writes + `kill -9` + restart recovered **snapshot + WAL tail combined** (`recovered 7 snapshot rows + 2 wal records`), totals exact at 45/45; (5) a restart with `WO_THREADS=8` against a 4-shard data dir **refuses to boot** (`meta` guard — resharding is 09f, never silent data loss). Recovery 1–3 ms at demo geometry; large-arena timing rides phase F's harness, which can generate volume natively (geometry is `-D` overridable: `SLOTS_PER_SHARD`/`SLOT_SIZE`).
### Phase F — million-scale harness + ACID verification — ✅ shipped
*Maps to [plan 09's verification-targets table](../../09-concurrency-scaleout.md).*
`setrlimit(RLIMIT_NOFILE)` raised at boot. A small C load client under `prototypes/wo-rt-c/bench/` (keep-alive, pipelined GETs, latency timestamps — `wrk` would be an external dep). Measure honestly on the dev box and commit the numbers to the prototype README: aggregate read req/s across cores (goal order 10⁶/s on 8–16 cores), concurrent open connections (goal order 10⁵–10⁶; ~8 KB/conn + fd limits are the ceiling), commits/s under group fsync, p99 read latency under write load. ACID scripts: torn-WAL injection (atomicity), single-shard interleaving probe (isolation), the phase-D crash test under load (durability). A `just rt-c-bench` recipe runs it all.
**Exit (met):** measured on a 20-core box (table in the [prototype README](../../../../prototypes/wo-rt-c/README.md)): **908,916 reads/s p99 154 µs and 643,250 fsync-acked commits/s p99 177 µs** on 8 shards — vs Go `net/http` on 20 cores at 495k/355k with ~8× worse p99 and no durability (.NET unavailable on the box); 10k idle connections, 0 errors; only 2xx counted (the client tracks status codes). **The crash-under-load test found two real durability bugs the phase-D test missed** — an ack-armed-before-fsync race in `conn_continue` (route parks the response *during* `try_process`; the pre-check missed it) and an fd-reuse ABA hazard in batch ack-parking (fixed with per-connection generation stamps). After both fixes, three `kill -9`-mid-bench rounds at ~1–2M commits each showed **WAL records ≥ acked, every round** (one exact). Isolation: 300 concurrent commits → 300 distinct interleaved ids. Geometry scaling via `-DSLOTS_PER_SHARD` (bitmap region generalized to multi-page); 512 MB arena verified mlocked.
## Non-scope
- **No cross-shard transactions.** 2PC is plan 09e on the Rust track; the prototype's rows are independent by design.
- **No liburing, ever.** Raw syscalls are the deliverable — the Rust port needs the sequence, not a wrapper.
- **No multi-node anything.** Single box, threads-as-shards, same stance as plan 09.
- **No TLS, no HTTP/2.** The HTTP layer stays minimal; the runtime is the subject.
- **The prototype directory stays a prototype.** Lessons flow into `crates/rt`; production code never imports from there.
## Cross-references
- [`../../09-concurrency-scaleout.md`](../../09-concurrency-scaleout.md) — the doctrine; this prototype is its executable proving ground (A↔09a, C↔09 decision 4, D↔09c).
- [`../../10-storage-foundations.md`](../../10-storage-foundations.md), [`11-wal-and-recovery.md`](../../11-wal-and-recovery.md), [`12-engine-disk-cutover.md`](../../12-engine-disk-cutover.md) — the storage track phases B/D/E miniaturize.
- [`../../../../prototypes/wo-rt-c/README.md`](../../../../prototypes/wo-rt-c/README.md) — current state and module map (phase 0).
- [`./01-architecture.md`](./01-architecture.md) — the target architecture traced through one memory address at million-connection concurrency, plus improvement proposals (seqlock reads, registered buffers, SEND_ZC, SQPOLL) that slot into phases C/F.
- [`./02-single-binary.md`](./02-single-binary.md) — the end goal: how the `wo build` single binary runs on this runtime environment (Go model, not JVM — the kernel is statically linked into every app; the embedding contract between compiler payload and runtime kernel).
- [`../../../../prototypes/wo-db/`](../../../../prototypes/wo-db/) — the query-layer sibling; one day a phase-G could splice its engine on top of this runtime.

View file

@ -0,0 +1,149 @@
# wo-rt-c architecture — one memory address, two spaces, a million connections
This document defines the runtime's architecture by following **one memory address** through user space, kernel space, and hardware, under a million connections reading and writing it concurrently — then suggests improvements. Companion docs: [`00-plan.md`](./00-plan.md) (the phases that build this), [`README.md`](../../../../prototypes/wo-rt-c/README.md) (phase-0 module map).
## The cast: one address
The database is one `mmap` arena. After the 4 KB header page, shard 0's first row slot sits at:
```
row = arena + 4096 → virtual address 0x7f3a2c001000 (say)
```
Three facts define everything that follows:
1. **User space sees a virtual address.** `0x7f3a2c001000` is an entry in this process's page tables; the kernel resolved it to one physical RAM frame at fault time (`MAP_POPULATE` faults it in at boot, before any request).
2. **The kernel pins the frame.** `mlock` guarantees the physical page is never swapped — a load from this address is always a RAM access, never disk I/O in disguise.
3. **Exactly one thread owns writes to it.** The address lies inside shard 0's slice; thread 0 is the only code in the process that may store to it ([00-plan.md decision 2](./00-plan.md)). Data exists once — ownership, not copying, is the concurrency model.
## The two spaces
User space and kernel space touch the same physical pages in exactly two places — the arena (data) and the io_uring rings (control). Everything else crosses by syscall.
```mermaid
flowchart TB
subgraph US["USER SPACE (N pinned threads, shared-nothing)"]
T0["thread 0<br/>event loop"]
ARENA["mmap arena<br/><b>0x7f3a2c001000</b> = shard-0 slot 0<br/>(mlock'd, MAP_POPULATE)"]
SQ["io_uring SQ/CQ rings<br/>(mmap'd — shared with kernel)"]
CB["per-connection buffers"]
T0 -->|"MOV — plain load/store,<br/>no syscall"| ARENA
T0 -->|"write SQE / read CQE<br/>no syscall"| SQ
T0 --- CB
end
subgraph KS["KERNEL SPACE"]
PT["page tables + TLB<br/>VA→PA for 0x7f3a2c001000"]
URING["io_uring engine"]
SKB["socket buffers (1M sockets,<br/>SO_REUSEPORT spread over N listeners)"]
PC["page cache + block layer<br/>(WAL file shard-0.wal)"]
end
subgraph HW["HARDWARE"]
RAM["RAM frame<br/>(the one physical copy)"]
NIC["NIC"]
SSD["SSD"]
end
SQ <-->|"io_uring_enter — ONE syscall<br/>per loop tick, batched"| URING
URING --> SKB
URING --> PC
ARENA -.->|"page tables map it"| PT
PT -.-> RAM
SKB <--> NIC
PC <--> SSD
```
The arrows worth staring at: the thread's access to the database (`MOV`) and to the I/O queues (ring writes) cross **no** boundary. The only recurring syscall is one batched `io_uring_enter` per loop tick.
## Write path — the address changes
One of the million connections POSTs a new value. Dual-write order per [00-plan.md decision 6](./00-plan.md): RAM first, then the hard drive, ack only after the disk confirms.
```mermaid
sequenceDiagram
participant NIC as NIC (hw)
participant K as kernel
participant T0 as thread 0 (user)
participant ROW as 0x7f3a2c001000 (RAM)
participant SSD as SSD (hw)
NIC->>K: packets → socket buffer
K->>T0: recv CQE (ring memory, no syscall)
T0->>T0: parse request, check invariants [C of ACID]
T0->>ROW: MOV — store new row bytes [RAM applied]
T0->>K: write SQE: WAL record len|crc32|payload|COMMIT
K->>SSD: page cache → block layer
T0->>K: one fdatasync SQE for ALL commits this tick [group commit]
SSD-->>K: flush done
K-->>T0: fsync CQE
T0->>K: send SQE — HTTP 201 ack [D of ACID: ack after fsync]
K->>NIC: response bytes out
```
Boundary crossings per commit: amortized to **one** `io_uring_enter` shared by every commit in the tick — the store to the address itself costs zero. Atomicity lives in the WAL framing (a torn record fails CRC and is dropped whole on replay); isolation is thread 0's serial execution; durability is the ack ordering.
## Read path — the address is observed
```
GET /api/notes/0 → thread 0 formats JSON straight from 0x7f3a2c001000
→ send SQE → socket buffer → NIC
```
The database read is **a memory load**. No file descriptor, no syscall, no kernel involvement until the response leaves. This is what "the whole database resides in RAM" buys: the kernel is in the room for networking and durability, not for reads.
## A million connections against this one address
- The kernel's `SO_REUSEPORT` hash spreads ~1M sockets across N listeners → each thread owns ~1M/N connections outright (accepts never migrate).
- **Memory ceiling:** ~8 KB user-space state per connection (`conns[]` buffer) + kernel sk_buffs → 1M connections ≈ 8 GB user + kernel-tunable socket memory. `RLIMIT_NOFILE` must be raised at boot (00-plan.md phase F).
- **Writes:** every write to the address funnels to thread 0 and serializes — that *is* the ACID isolation story, and group commit keeps the WAL from becoming a per-write fsync storm.
- **Reads — the honest bottleneck:** today a connection that hashed to thread 3 cannot serve the address; only shard 0's thread may touch it. One hot row = one core's worth of read throughput (~the per-core ceiling), while the other N−1 cores idle on that row. The improvements below exist for exactly this.
## Suggested improvements
Ordered by how cleanly each fits the locked doctrine (thread-per-core, no locks on the data path, no duplication, libc only).
### 1. Per-slot seqlock — every core may read the one address
The single highest-leverage change. Give each slot a version counter; the owning thread (still the **only writer**) increments it before and after the store (odd = mid-write). Any thread on any core may then read the address directly:
```c
do { v1 = atomic_load_acquire(&slot->ver); /* spin only while odd */
memcpy(local, slot->bytes, len);
v2 = atomic_load_acquire(&slot->ver);
} while (v1 != v2 || (v1 & 1));
```
A hot row becomes readable by all N cores **with zero duplication — same physical frame, same address** — and writes stay serial, so ACID isolation is untouched. Cost: two atomic increments per write, a retry loop per read (C11 atomics, no library). Doctrine note: this relaxes "only the owner touches the slice" to "only the owner *writes* the slice"; contrast with [plan 13e's hot-row read replicas](../../13-class-model-live-pricing.md), which solve the same bottleneck by *copying* rows per thread — seqlock is the no-duplication answer the replica design isn't.
### 2. Registered buffers and files (`IORING_REGISTER_BUFFERS` / `_FILES`)
The arena and connection buffers are already `mlock`-pinned; registering them lets the kernel skip per-operation page lookup/refcounting, and WAL fds skip the fd-table walk. Pure win, no doctrine impact.
### 3. Zero-copy send (`IORING_OP_SEND_ZC`)
Responses currently copy user → socket buffer. Zero-copy send transmits straight from user memory — strongest when combined with improvement 5, where the hot row's bytes are already response-shaped.
### 4. SQPOLL (`IORING_SETUP_SQPOLL`)
A kernel-side poller consumes the SQ ring; steady state needs **zero** syscalls — even the per-tick `io_uring_enter` disappears. This is precisely the "the only non-userland thread is the kernel-owned io_uring SQPOLL helper" end state already written into the project's concurrency model (CLAUDE.md). Cost: one kernel thread per ring burning a core fraction; enable per-deployment.
### 5. Serialized-row cache beside the slot
Store the rendered JSON next to the row bytes, invalidated by the same seqlock version bump. A hot read becomes `memcpy` from the address — no formatting per request. Trades arena bytes for CPU; measurable in phase F before adopting.
### 6. NUMA-aware arena placement (`set_mempolicy` / `mbind`)
On multi-socket boxes, bind each shard slice's pages to the owning core's NUMA node — the address is always a local-node load (~80 ns vs ~140 ns remote). No-op on single-socket dev machines; matters at the 16-core scale-out target.
### 7. Multishot recv + huge pages
`IORING_RECV_MULTISHOT` arms one SQE per connection instead of one per request — at 1M connections that is the difference between 1M and ~0 re-arm submissions per tick. `MAP_HUGETLB` (already phase B) cuts TLB pressure: a 16 GB arena is 8.4M × 4 KB entries but only 8K × 2 MB entries.
### Deliberately not suggested
Work stealing (breaks single-writer ACID), shared-heap locking (the doctrine exists to avoid it — and at 1M readers a mutex on the row would serialize everything the seqlock parallelizes), liburing (the prototype's value is the raw syscall sequence), and multi-node distribution (plan 09's single-box stance). See [plan 09 § Non-scope](../../09-concurrency-scaleout.md).
## Cross-references
- [`00-plan.md`](./00-plan.md) — phases A–F that build the architecture described here; improvements 1–7 slot into phases C/F or follow them.
- [`../../docs/plan/09-concurrency-scaleout.md`](../../09-concurrency-scaleout.md) — the thread-per-core doctrine.
- [`../../docs/plan/exploration/linux/07-io_uring.md`](../linux/07-io_uring.md), [`08-mmap.md`](../linux/08-mmap.md) — the two shared-page mechanisms.
- [`../../docs/plan/13-class-model-live-pricing.md`](../../13-class-model-live-pricing.md) — 13e's read-replica alternative, contrasted in improvement 1.
- [`../../docs/writeonce-pl.md`](../../../writeonce-pl.md) — the C/assembly "one address" pedagogy this doc extends to a full runtime.

View file

@ -0,0 +1,90 @@
# 02 — The end goal: the writeonce single binary on this runtime environment
**Context sources:** [`00-plan.md`](./00-plan.md) (the runtime-environment phases, A–B ✅), [`01-architecture.md`](./01-architecture.md) (the one-address trace), [`../../../runtime/wo-language.md`](../../../runtime/wo-language.md) ("one binary per project; no runtime to install on the target host"; `.wo` has "its own lexer, parser, analyzer, and bytecode"), [`../../09-concurrency-scaleout.md`](../../09-concurrency-scaleout.md)–[`12`](../../12-engine-disk-cutover.md) (the Rust product track this proves out), [`../../../runtime/database/02-wo-language.md`](../../../runtime/database/02-wo-language.md) (catalog + transaction semantics the payload carries).
## The end goal, stated once
> A developer writes `.wo` files. `wo build` emits **one static binary**. That binary **is** the runtime environment described in this track — thread-per-core shards, the mlock'd RAM arena, io_uring event loops, WAL dual-write, boot-time recovery — with the application **inside it** as data and bytecode. Deploy = copy one file to a Linux host. Nothing to install; the only "VM" underneath is the Linux kernel.
This resolves the phrase "runs **in** this runtime environment": writeonce is the **Go model, not the JVM model**. The environment is not a process you start and feed programs to — it is a **runtime kernel** statically linked into every application binary. What `wo-rt-c` builds phase by phase is that kernel, proven in C, ported to Rust per plans 09–12.
```mermaid
flowchart TB
subgraph BIN["ONE BINARY (output of `wo build`)"]
subgraph APP["application payload — compiled from .wo"]
CAT["catalog<br/>types/classes → slot geometry"]
RT2["route table<br/>service decls → method+path+op"]
BC["logic bytecode<br/>fn methods, triggers, queries"]
UIA["UI assets<br/>SSR templates + manifest + CSS"]
end
subgraph KERN["runtime kernel — this track"]
BOOT["boot loader<br/>size arena from catalog, WAL replay"]
SCHED["shard scheduler<br/>thread-per-core, SO_REUSEPORT"]
LOOP["event loops<br/>epoll → io_uring"]
ENG["engine primitives<br/>arena slots, bitmaps, indexes"]
WAL["commit pipeline<br/>RAM apply → WAL → group fsync → ack"]
HTTP["HTTP + WS"]
end
end
LINUX["Linux kernel — epoll/io_uring, mmap/mlock, signalfd/eventfd, page cache"]
APP -->|"consumed at boot<br/>and at dispatch"| KERN
KERN --> LINUX
```
## The embedding contract
The seam between the two halves is narrow and data-shaped — the compiler emits **tables**, the kernel consumes them. Five entries:
| Payload artifact | Emitted from | Consumed by | Exists today as |
| --- | --- | --- | --- |
| **Catalog** — per type/class: field layout, slot size, unique keys, shard key | `type`/`class` decls (`crates/rt/src/compile.rs`) | boot loader: arena geometry (`arena_hdr` is its miniature); engine: row codecs | `rt::compile::Catalog`, built per `wo run` |
| **Route table** — `(method, path pattern, operation, type id)` | `service` blocks | HTTP dispatch | `rt::server::router` |
| **Logic bytecode** — method/trigger/query programs | `fn` bodies, `on` blocks, schema-layer DML | a bytecode interpreter running **inside the owning shard's thread** — serial execution = ACID isolation for free | parse-and-discard (13b lands the executor) |
| **UI assets** — SSR templates, manifest, flattened CSS, `wo-runtime.js` | `##ui` triplets (plan 14) | HTTP static + SSR renderer | design (plans 13d/14) |
| **App manifest** — name, version, default port/threads | `wo.toml` | boot banner, env-var defaults | `wo.toml` parsing TBD |
Two consequences worth locking:
1. **The shard key comes from the catalog, not the kernel.** `wo-rt-c` hashes connections (4-tuple) because it has no catalog; the product routes *operations* by the catalog's shard key (plan 09 decision 5: customer id, author id) — a request landing on any thread sends an in-process message to the owning shard. Phase A's `SO_REUSEPORT` spread is the transport layer of that story, not the final routing.
2. **Bytecode runs where the data lives.** A `set_price` method executes on the shard that owns the product row — the interpreter is invoked from the event loop, runs to completion, commits through the WAL pipeline. No cross-thread data access; cross-shard transactions escalate to 2PC (plan 09e).
## Boot sequence of the single binary
What `./myapp` does before serving — each step already prototyped or phased:
| # | Step | Proven by |
| --- | --- | --- |
| 1 | Read embedded catalog → compute arena geometry (shards × slot families) | `arena_hdr` (phase B ✅) |
| 2 | `mmap` + `MAP_POPULATE` + `mlock` the arena (hugepages, 4 K fallback) | phase B ✅ |
| 3 | Replay per-shard WAL/snapshot from disk into the arena, in parallel | phase E |
| 4 | Spawn `WO_THREADS` pinned shard threads, each with ring + `SO_REUSEPORT` listener | phase A ✅ / phase C |
| 5 | Install route table + bytecode programs per shard | 13b / this contract |
| 6 | Listeners open; signalfd→eventfd shutdown armed | phase A ✅ |
Steps 1–2 and 4–6 exist in `wo-rt-c` today with the notes store standing in for the catalog. The end state swaps the hard-coded `slot_note` for catalog-driven slot families — the kernel code does not otherwise change shape.
## What runs where — `.wo` construct → runtime home
| `.wo` construct | At runtime, inside the binary |
| --- | --- |
| `type` / `class` fields | a slot family in each shard's arena slice |
| `class` `fn` methods | bytecode executed serially in the owning shard's thread (`POST /api/<t>/:id/<m>`) |
| `service rest … expose` | route-table entries dispatched by the HTTP layer |
| `on <event>` triggers | bytecode hooked into the commit pipeline, same transaction |
| `LIVE select` / `subscribe` | per-shard subscription registry; deltas fan out one message per shard (09d) |
| `##ui` screens | SSR render + manifest at GET routes; live patches ride the same deltas |
| `policy` rules | predicate rewrites applied before the engine touches slots |
| `main { … }` | bytecode run once at boot step 5½, before listeners open |
## Division of labor (locked by the C-vs-Rust decision)
- **`wo-rt-c` (C)** — proves each kernel syscall sequence first: threads/arena (✅), io_uring, WAL, recovery, bench. It will never parse `.wo`; its notes store is the stand-in payload.
- **`crates/rt` (Rust)** — the product: owns the compiler front-end today (lexer→catalog, Stage 2 shipped) and absorbs each proven kernel sequence per plans 09–12, where ownership makes the shard discipline a compile-time guarantee.
- **Optional phase G** (named in [`00-plan.md`](./00-plan.md)): splice the [`wo-db`](../../../../prototypes/wo-db/) C++ query engine onto `wo-rt-c` as an end-to-end C-family demonstrator of this document — valuable as proof, never the product.
## Cross-references
- [`00-plan.md`](./00-plan.md) — the kernel phases; [`01-architecture.md`](./01-architecture.md) — the one-address trace through the same stack.
- [`../../../runtime/wo-language.md`](../../../runtime/wo-language.md) — the user-facing single-binary promise this document implements.
- [`../../13-class-model-live-pricing.md`](../../13-class-model-live-pricing.md) (13b methods), [`../../14-mvc-ui-implementation.md`](../../14-mvc-ui-implementation.md) (UI assets) — the payload-side tracks.
- [`../../09-concurrency-scaleout.md`](../../09-concurrency-scaleout.md) — shard-key routing and 2PC the contract defers to.

85
justfile Normal file
View file

@ -0,0 +1,85 @@
# writeonce — task runner. `just --list` shows all recipes.
# serve the hello example (docs/examples/hello/main.wo) on :8080
hello:
cargo run --bin wo -- run docs/examples/hello
# C runtime reference (prototypes/wo-rt-c): build, serve, CRUD round-trip, shut down
# Phase A: thread-per-core — each connection hashes to one shard (SO_REUSEPORT),
# so a list may land on a different shard than the create. The counters on /
# show the spread. WO_THREADS=4 keeps the demo output readable.
rt-c-demo port="8085" threads="4":
#!/usr/bin/env bash
set -euo pipefail
make -C prototypes/wo-rt-c
data=$(mktemp -d /tmp/wo-demo-XXXXXX)
WO_PORT={{port}} WO_THREADS={{threads}} WO_DATA=$data ./prototypes/wo-rt-c/wo-rt &
server=$!
trap 'kill $server 2>/dev/null; sleep 0.3; rm -rf $data' EXIT
base=http://127.0.0.1:{{port}}
for _ in $(seq 1 40); do curl -s "$base/healthz" >/dev/null && break; sleep 0.25; done
echo
echo "--- runtime:"; curl -s "$base/"; echo
echo "--- create x4 (each connection may hash to a different shard):"
for i in 1 2 3 4; do curl -s -X POST "$base/api/notes" -d '{"title":"note '$i'"}'; echo; done
echo "--- list (the connection's own shard only — shared-nothing):"
curl -s "$base/api/notes"; echo
echo "--- spread:"; curl -s "$base/"; echo
# phase-F benchmark: reads, durable writes, 10k idle conns (scaled geometry)
rt-c-bench port="8085" threads="8" conns="64":
#!/usr/bin/env bash
set -euo pipefail
make -C prototypes/wo-rt-c clean >/dev/null
make -C prototypes/wo-rt-c CFLAGS="-O2 -Wall -Wextra -std=c11 -DSLOTS_PER_SHARD=262144" wo-rt bench >/dev/null
data=$(mktemp -d /tmp/wo-bench-XXXXXX)
WO_PORT={{port}} WO_THREADS={{threads}} WO_DATA=$data ./prototypes/wo-rt-c/wo-rt >/dev/null 2>&1 &
server=$!
trap 'kill $server 2>/dev/null; sleep 0.3; rm -rf $data; make -C prototypes/wo-rt-c clean >/dev/null; make -C prototypes/wo-rt-c wo-rt bench >/dev/null' EXIT
base=127.0.0.1; for _ in $(seq 1 40); do curl -s "http://$base:{{port}}/healthz" >/dev/null && break; sleep 0.25; done
B=./prototypes/wo-rt-c/bench/bench
echo "wo-rt-c ({{threads}} shards, durable WAL):"
$B $base {{port}} {{conns}} 5 /healthz
$B $base {{port}} {{conns}} 5 /
$B $base {{port}} {{conns}} 3 /api/notes '{"title":"bench"}'
$B $base {{port}} 10000 0 /healthz
# serve the pricing demo (class model — docs/examples/pricing) on :8080
pricing:
cargo run --bin wo -- run docs/examples/pricing
# class model in action: serve pricing, CRUD round-trip on Product, shut down
pricing-demo port="8092":
#!/usr/bin/env bash
set -euo pipefail
cargo build --bin wo
WO_LISTEN=127.0.0.1:{{port}} ./target/debug/wo run docs/examples/pricing &
server=$!
trap 'kill $server 2>/dev/null' EXIT
base=http://127.0.0.1:{{port}}
for _ in $(seq 1 40); do curl -s "$base/healthz" >/dev/null && break; sleep 0.25; done
echo
echo "--- create:"; curl -s -X POST "$base/api/products" -H 'Content-Type: application/json' -d '{"sku":"WO-001","name":"writeonce mug"}'; echo
echo "--- list:"; curl -s "$base/api/products"; echo
echo "--- patch 1:"; curl -s -X PATCH "$base/api/products/1" -d '{"name":"writeonce mug v2"}'; echo
echo "--- live (13c pending, expect 501):"; curl -s -o /dev/null -w '%{http_code}\n' "$base/api/products/live"
echo "--- delete 1 (expect 204):"; curl -s -X DELETE "$base/api/products/1" -o /dev/null -w '%{http_code}\n'
# main.wo in action: serve hello, run the full CRUD round-trip, shut down
hello-demo port="8090":
#!/usr/bin/env bash
set -euo pipefail
cargo build --bin wo
WO_LISTEN=127.0.0.1:{{port}} ./target/debug/wo run docs/examples/hello &
server=$!
trap 'kill $server 2>/dev/null' EXIT
base=http://127.0.0.1:{{port}}
for _ in $(seq 1 40); do curl -s "$base/healthz" >/dev/null && break; sleep 0.25; done
echo
echo "--- create:"; curl -s -X POST "$base/api/notes" -H 'Content-Type: application/json' -d '{"title":"hello","body":"# First note"}'; echo
echo "--- list:"; curl -s "$base/api/notes"; echo
echo "--- get 1:"; curl -s "$base/api/notes/1"; echo
echo "--- patch 1:"; curl -s -X PATCH "$base/api/notes/1" -d '{"pinned":true}'; echo
echo "--- live (Stage 3 stub, expect 501):"; curl -s -o /dev/null -w '%{http_code}\n' "$base/api/notes/live"
echo "--- delete 1 (expect 204):"; curl -s -X DELETE "$base/api/notes/1" -o /dev/null -w '%{http_code}\n'
echo "--- list after delete:"; curl -s "$base/api/notes"; echo

View file

@ -0,0 +1,19 @@
CC ?= cc
CFLAGS ?= -O2 -Wall -Wextra -std=c11
LDFLAGS ?= -pthread
wo-rt: wo-rt.c
$(CC) $(CFLAGS) -o $@ $< $(LDFLAGS)
bench/bench: bench/bench.c
$(CC) $(CFLAGS) -o $@ $< $(LDFLAGS)
bench: bench/bench
run: wo-rt
./wo-rt
clean:
rm -f wo-rt bench/bench
.PHONY: bench run clean

View file

@ -0,0 +1,75 @@
# `wo-rt-c` — the writeonce runtime environment, in C
A single-file C implementation of the **runtime layer** the writeonce language runs on — now at **phase E: a durable RAM database**. Writes follow the dual-write order — RAM apply, framed WAL record (`len|crc32|payload|COMMIT`) to a per-shard `fallocate`'d log, one group-commit `fdatasync` per loop tick, **HTTP ack only after the fsync completion**. Boot performs the **first load, hard drive → RAM**: each shard replays its snapshot + WAL tail into its arena slice in parallel before any accept arms; clean shutdown snapshots each slice and truncates the WAL; a `meta` file pins the shard count so a mismatched `WO_THREADS` refuses to boot. `./wo-rt wal-check <file>` validates a log offline. `WO_THREADS` pinned threads (default = online cores), each owning its own raw io_uring ring (`io_uring_setup` + mmap'd SQ/CQ rings + `io_uring_enter` — **no liburing**), its own `SO_REUSEPORT` listener (multishot accept), its own keep-alive connections, and its own slice of the one mlock'd mmap arena — shared-nothing, no locks. Steady state is one `io_uring_enter` syscall per loop tick. Thread 0 owns the `signalfd`; shutdown broadcasts through per-thread `eventfd`s, both watched via `POLL_ADD` SQEs. **Zero dependencies beyond libc + kernel uapi headers.** The kernel is the runtime.
This is the runtime-layer sibling of [`prototypes/wo-db/`](../wo-db/) (the C++ query-layer prototype): a reference card showing, with no abstraction in the way, exactly which kernel primitives the production Rust runtime (`crates/rt/`) drives through `libc`. Same role, different layer.
```
prototypes/wo-db/ C++ what the LANGUAGE executes (parser, engine, transactions)
prototypes/wo-rt-c/ C what the RUNTIME stands on (epoll, signalfd, sockets)
crates/rt/ Rust the product — both layers, libc only
```
## Build, run, poke
```bash
make # cc -O2 -Wall -Wextra -std=c11 -pthread — no libraries
./wo-rt # 127.0.0.1:8085 (WO_PORT=9000 WO_THREADS=4 ./wo-rt to override)
curl localhost:8085/ # {"runtime":"wo-rt-c","loop":"epoll-et","threads":4,
# "shard":2,"shard_requests":[68,36,44,53]}
curl -X POST localhost:8085/api/notes -d '{"title":"hello"}'
# {"id":2,"title":"hello","shard":1} ← ids interleave per shard
curl localhost:8085/api/notes # the connection's shard only — shared-nothing
# ctrl-C → signalfd on shard 0 → eventfd broadcast → all shards join
```
Each connection hashes to one shard for life (`SO_REUSEPORT` 4-tuple): a list may land on a different shard than the create that preceded it. That is the architecture, not a bug — cross-shard reads are a later phase / design decision (see the [architecture doc's improvements](../../docs/plan/exploration/c-runtime/01-architecture.md)).
Or from the repo root: `just rt-c-demo`.
## Module map
Every block in `wo-rt.c` corresponds one-to-one to a module of the Rust runtime, which in turn mirrors Go's netpoller — the same lineage the docs trace:
| `wo-rt.c` block | Rust (`crates/rt/src/`) | Go (`reference/go/src/runtime/`) | Kernel reference card |
| --- | --- | --- | --- |
| `main` event loop (`epoll_create1` / `epoll_wait`, `EPOLLET`) | `runtime/netpoll_epoll.rs` | `netpoll_epoll.go` | [`linux/01-epoll.md`](../../docs/plan/exploration/linux/01-epoll.md) |
| `sig_setup` (`sigprocmask` + `signalfd`) | `runtime/signalfd.rs` | signal mask handling | [`linux/04-signalfd.md`](../../docs/plan/exploration/linux/04-signalfd.md) |
| `listener_bind` (`SOCK_NONBLOCK`, `accept4`-to-EAGAIN) | `http/listener.rs` | `net.Listen` + accept loop | `socket(7)` |
| `conn_drive` (read-to-EAGAIN, one buffer per fd) | `http/connection.rs` | `conn.Read` loop | the edge-triggered contract |
| `notes[]` store | `engine.rs` (BTreeMaps) | — | [`03-inmemory-engine.md`](../../docs/runtime/database/03-inmemory-engine.md) |
## What it demonstrates
- **One thread owns each shard outright.** Accept, parse, store, respond — no locks, no worker pool, no connection migration. Scaling past one core is more shards ([`09-concurrency-scaleout.md`](../../docs/plan/09-concurrency-scaleout.md)), never shared mutable state. The single cross-thread touch is the relaxed-atomic stats counters on `/` — monotonic, never on the data path.
- **Edge-triggered discipline.** Every registration sets `EPOLLET`; every readiness event is drained to `EAGAIN` (the accept loop and the read loop both). Get this wrong and connections silently hang — the reason the Rust module documents the same contract at the top of `netpoll_epoll.rs`.
- **Signals as fd events.** `SIGINT`/`SIGTERM` are blocked, then read from a `signalfd` on the same epoll — no async-signal-unsafe handler, no self-pipe trick.
- **RAM is the read path.** `GET /api/notes` touches a C array. The production engine is the same idea with MVCC and a WAL behind it.
## Architecture and roadmap
Documentation lives under `docs/` (repo convention) — this README stays here as the directory's orientation page only:
- [`docs/plan/exploration/c-runtime/01-architecture.md`](../../docs/plan/exploration/c-runtime/01-architecture.md) — the runtime defined by tracing **one memory address** through user space, kernel space, and hardware under a million concurrent connections, plus seven improvement proposals (seqlock reads, registered buffers, zero-copy send, SQPOLL, …).
- [`docs/plan/exploration/c-runtime/00-plan.md`](../../docs/plan/exploration/c-runtime/00-plan.md) — the phase sequence: **A → B → C → D → E → F, all ✅ shipped.**
## Measured (phase F, 20-core Linux 6.14, tmpfs data dir, `just rt-c-bench`)
Same C bench client (`bench/bench.c`, keep-alive, only 2xx counted) against both servers:
| Benchmark | **wo-rt-c** (8 shards, durable WAL) | **Go `net/http`** (go1.25.1, 20 cores, no durability) |
| --- | --- | --- |
| `GET /healthz` | **908,916 req/s** · p50 66 µs · p99 154 µs | 495,235 req/s · p50 65 µs · p99 1,310 µs |
| `GET /` (JSON) | **686,738 req/s** · p99 189 µs | 481,666 req/s · p99 1,297 µs |
| `POST` write | **643,250 commits/s** — every one fsync-acked · p99 177 µs | 354,758 req/s — RAM only, no WAL · p99 1,649 µs |
| 10,000 idle conns | 0 errors | 0 errors |
wo-rt-c on 8 cores outpaces Go on 20 with ~8× tighter p99 (Go's GC shows there) — while fsyncing every write Go doesn't. Honest caveats: `net/http` does full general-purpose HTTP; our parser is minimal; .NET was not installed on the box. **ACID under load:** three crash rounds (`kill -9` mid-bench at ~2M commits) all showed WAL records ≥ acked; isolation probe: 300 concurrent commits → 300 distinct ids; torn-tail records drop whole by CRC.
The crash-under-load test **found and fixed two real bugs** the lighter phase-D test missed: an ack-before-fsync race (`conn_continue` armed the send in the same tick the commit was staged) and an fd-reuse ABA hazard in ack parking (fixed with per-connection generation stamps). That is what phase F is for.
- [`docs/plan/exploration/c-runtime/02-single-binary.md`](../../docs/plan/exploration/c-runtime/02-single-binary.md) — the end goal: how the `wo build` **single binary** runs on this runtime environment — the runtime kernel is statically linked into every writeonce app (Go model, nothing to install), with the catalog/routes/bytecode payload consumed at boot.
## Deliberate simplifications
Single-shot RECV re-armed per request (multishot recv + buffer rings are a phase-F improvement), one outstanding SQE per connection, fixed-size buffers, naive `"title"` extraction instead of a JSON parser, no `timerfd`. Requires kernel ≥ 5.19 (multishot accept). This file is for reading; `crates/rt` is for running writeonce.

View file

@ -0,0 +1,201 @@
/*
* bench.c — load client for wo-rt-c (phase F). Zero deps beyond libc.
*
* T threads, each driving ONE keep-alive connection in a tight request/
* response loop (TCP_NODELAY, full-response framing via Content-Length).
* Every request is latency-stamped; the run prints req/s, p50, p99, errors.
*
* usage: ./bench <host> <port> <threads> <seconds> <path> [post-json]
* read: ./bench 127.0.0.1 8085 64 5 /api/notes
* write: ./bench 127.0.0.1 8085 64 5 /api/notes '{"title":"bench"}'
* idle: ./bench 127.0.0.1 8085 10000 0 /healthz # seconds=0: open conns,
* # one request each, hold, exit
*/
#define _GNU_SOURCE
#include <arpa/inet.h>
#include <errno.h>
#include <netinet/in.h>
#include <netinet/tcp.h>
#include <pthread.h>
#include <stdio.h>
#include <stdlib.h>
#include <string.h>
#include <sys/resource.h>
#include <sys/socket.h>
#include <time.h>
#include <unistd.h>
#define MAX_SAMPLES 400000 /* per thread; counting continues past it */
static char g_req[2048];
static int g_reqlen;
static struct sockaddr_in g_addr;
static long g_deadline_us; /* 0 = idle-connection mode */
static int g_hold_secs;
struct worker {
pthread_t tid;
long reqs, errs;
long ok2xx, non2xx; /* honest accounting: only 2xx is success */
long *lat; /* µs samples */
int nlat;
};
static long now_us(void) {
struct timespec ts;
clock_gettime(CLOCK_MONOTONIC, &ts);
return ts.tv_sec * 1000000L + ts.tv_nsec / 1000;
}
static int conn_open(void) {
int fd = socket(AF_INET, SOCK_STREAM, 0);
if (fd < 0) return -1;
int one = 1;
setsockopt(fd, IPPROTO_TCP, TCP_NODELAY, &one, sizeof one);
if (connect(fd, (struct sockaddr *)&g_addr, sizeof g_addr) < 0) { close(fd); return -1; }
return fd;
}
/* One request/response round-trip. Returns HTTP status, or -1 conn-dead. */
static int round_trip(int fd) {
size_t off = 0;
while (off < (size_t)g_reqlen) {
ssize_t n = write(fd, g_req + off, (size_t)g_reqlen - off);
if (n <= 0) { if (n < 0 && errno == EINTR) continue; return -1; }
off += (size_t)n;
}
static _Thread_local char buf[131072];
size_t got = 0, need = 0;
for (;;) {
ssize_t n = read(fd, buf + got, sizeof buf - 1 - got);
if (n <= 0) { if (n < 0 && errno == EINTR) continue; return -1; }
got += (size_t)n;
buf[got] = 0;
if (!need) {
char *he = strstr(buf, "\r\n\r\n");
if (!he) { if (got >= sizeof buf - 1) return -1; continue; }
size_t hdr = (size_t)(he + 4 - buf);
long cl = 0;
char *p = strcasestr(buf, "Content-Length:");
if (p) cl = strtol(p + 15, NULL, 10);
need = hdr + (size_t)cl;
}
if (got >= need) {
int code = 0;
sscanf(buf, "HTTP/1.1 %d", &code);
return code;
}
if (got >= sizeof buf - 1) return -1;
}
}
static void *worker_main(void *arg) {
struct worker *w = arg;
int fd = conn_open();
if (fd < 0) { w->errs++; return NULL; }
if (g_deadline_us == 0) { /* idle-connection mode */
if (round_trip(fd) > 0) w->reqs++; else w->errs++;
sleep((unsigned)g_hold_secs);
close(fd);
return NULL;
}
while (now_us() < g_deadline_us) {
long t0 = now_us();
int code = round_trip(fd);
if (code < 0) { /* reconnect once, then count errs */
close(fd);
fd = conn_open();
if (fd < 0) { w->errs++; break; }
w->errs++;
continue;
}
long dt = now_us() - t0;
if (w->nlat < MAX_SAMPLES) w->lat[w->nlat++] = dt;
w->reqs++;
if (code >= 200 && code < 300) w->ok2xx++; else w->non2xx++;
}
close(fd);
return NULL;
}
static int cmp_long(const void *a, const void *b) {
long x = *(const long *)a, y = *(const long *)b;
return (x > y) - (x < y);
}
int main(int argc, char **argv) {
if (argc < 6) {
fprintf(stderr, "usage: %s <host> <port> <threads> <seconds> <path> [post-json]\n", argv[0]);
return 2;
}
const char *host = argv[1];
int port = atoi(argv[2]);
int threads = atoi(argv[3]);
int secs = atoi(argv[4]);
const char *path = argv[5];
const char *body = argc > 6 ? argv[6] : NULL;
struct rlimit rl;
if (getrlimit(RLIMIT_NOFILE, &rl) == 0 && rl.rlim_cur < rl.rlim_max) {
rl.rlim_cur = rl.rlim_max;
setrlimit(RLIMIT_NOFILE, &rl);
}
memset(&g_addr, 0, sizeof g_addr);
g_addr.sin_family = AF_INET;
g_addr.sin_port = htons((uint16_t)port);
inet_pton(AF_INET, host, &g_addr.sin_addr);
if (body)
g_reqlen = snprintf(g_req, sizeof g_req,
"POST %s HTTP/1.1\r\nHost: %s\r\nContent-Type: application/json\r\n"
"Content-Length: %zu\r\nConnection: keep-alive\r\n\r\n%s",
path, host, strlen(body), body);
else
g_reqlen = snprintf(g_req, sizeof g_req,
"GET %s HTTP/1.1\r\nHost: %s\r\nConnection: keep-alive\r\n\r\n", path, host);
g_hold_secs = 3;
g_deadline_us = secs > 0 ? now_us() + (long)secs * 1000000L : 0;
struct worker *ws = calloc((size_t)threads, sizeof *ws);
for (int i = 0; i < threads; i++) {
ws[i].lat = secs > 0 ? malloc(MAX_SAMPLES * sizeof(long)) : NULL;
pthread_create(&ws[i].tid, NULL, worker_main, &ws[i]);
}
long t0 = now_us();
long total = 0, errs = 0, nlat = 0, ok = 0, bad = 0;
for (int i = 0; i < threads; i++) {
pthread_join(ws[i].tid, NULL);
total += ws[i].reqs;
errs += ws[i].errs;
nlat += ws[i].nlat;
ok += ws[i].ok2xx;
bad += ws[i].non2xx;
}
long wall_us = now_us() - t0;
if (secs == 0) {
printf("idle-conns: opened %ld / %d connections (errs %ld), held %ds, server survived\n",
total, threads, errs, g_hold_secs);
return errs ? 1 : 0;
}
long *all = malloc((size_t)nlat * sizeof(long));
long k = 0;
for (int i = 0; i < threads; i++) {
memcpy(all + k, ws[i].lat, (size_t)ws[i].nlat * sizeof(long));
k += ws[i].nlat;
}
qsort(all, (size_t)nlat, sizeof(long), cmp_long);
double rps = (double)ok / ((double)wall_us / 1e6);
printf("%-22s %d conns %ds %ld ok (2xx) %.0f ok/s p50 %ld µs p99 %ld µs non-2xx %ld errs %ld\n",
body ? "WRITE (POST)" : path, threads, secs, ok, rps,
nlat ? all[nlat / 2] : 0, nlat ? all[(long)((double)nlat * 0.99)] : 0, bad, errs);
return 0;
}

View file

@ -0,0 +1,3 @@
module goref
go 1.25.1

View file

@ -0,0 +1,69 @@
// goref — the Go net/http comparison server for the phase-F benchmark.
// Same endpoints and semantics as wo-rt-c's RAM read path: /healthz, /,
// GET/POST /api/notes against an in-memory store. Durability is NOT
// implemented here (Go side has no WAL), so only READ benchmarks are
// apples-to-apples; the POST comparison measures Go's non-durable path
// against wo-rt-c's fsync-backed path and is labeled as such.
//
// go build -o goref . && ./goref # :8095, GOMAXPROCS = all cores
package main
import (
"encoding/json"
"fmt"
"io"
"net/http"
"sync"
)
type note struct {
ID int `json:"id"`
Title string `json:"title"`
}
var (
mu sync.RWMutex
notes []note
nextID = 1
)
func main() {
http.HandleFunc("/healthz", func(w http.ResponseWriter, r *http.Request) {
io.WriteString(w, "ok")
})
http.HandleFunc("/", func(w http.ResponseWriter, r *http.Request) {
io.WriteString(w, `{"runtime":"go-net-http"}`)
})
http.HandleFunc("/api/notes", func(w http.ResponseWriter, r *http.Request) {
switch r.Method {
case http.MethodGet:
mu.RLock()
b, _ := json.Marshal(notes)
mu.RUnlock()
w.Header().Set("Content-Type", "application/json")
w.Write(b)
case http.MethodPost:
var in struct {
Title string `json:"title"`
}
if json.NewDecoder(r.Body).Decode(&in) != nil || in.Title == "" {
http.Error(w, `{"error":"expected {\"title\":\"...\"}"}`, http.StatusBadRequest)
return
}
mu.Lock()
n := note{ID: nextID, Title: in.Title}
nextID++
if len(notes) < 100000 {
notes = append(notes, n)
}
mu.Unlock()
w.Header().Set("Content-Type", "application/json")
w.WriteHeader(http.StatusCreated)
json.NewEncoder(w).Encode(n)
default:
http.Error(w, "method", http.StatusMethodNotAllowed)
}
})
fmt.Println("[goref] listening on :8095")
http.ListenAndServe("127.0.0.1:8095", nil)
}

979
prototypes/wo-rt-c/wo-rt.c Normal file
View file

@ -0,0 +1,979 @@
/*
* wo-rt.c — the writeonce runtime environment, in C. Phase E: first load.
*
* Phases A–D: N pinned threads with raw io_uring loops, SO_REUSEPORT
* listeners, keep-alive connections, one mlock'd mmap arena, and durable
* commits (RAM apply → framed WAL record → per-tick group fdatasync → ack
* on the fsync CQE). Phase E closes the loop: at boot — BEFORE any accept
* is armed — each shard thread loads its snapshot and replays its WAL into
* its arena slice, in parallel, validating every frame's CRC + COMMIT
* trailer and truncating at the first torn record. Appends resume at the
* validated tail. A clean shutdown writes a per-shard snapshot
* (`shard-<t>.data`) and truncates the WAL; boot prefers snapshot + WAL
* tail. The data directory carries a `meta` file pinning the shard count —
* a restart with a different WO_THREADS refuses to start (resharding is
* plan 09f, not silent data loss). Zero deps beyond libc + kernel uapi.
*
* build: make run: ./wo-rt [WO_PORT=8085 WO_THREADS=4 WO_DATA=./wo-data ./wo-rt]
* poke: curl -X POST localhost:8085/api/notes -d '{"title":"hello"}' # acked after fsync
* wal: ./wo-rt wal-check wo-data/shard-0.wal # offline frame/CRC validation
*
* Phase map: docs/plan/exploration/c-runtime/00-plan.md (A ✅ threads,
* B ✅ arena, C ✅ io_uring, D ✅ WAL, E this file, F bench). One-address
* trace: 01-architecture.md. Single-binary end goal: 02-single-binary.md.
*
* Module map (C ↔ Rust ↔ kernel reference card):
* ring_init/ring_enter ↔ (plan 09 decision 4: per-thread ring) ↔ linux/07-io_uring.md
* arena_init ↔ (plan 10 storage foundations) ↔ linux/08-mmap.md
* wal_flush / OP_FSYNC ↔ (plan 11 WAL + 09c per-shard WAL) ↔ linux/12-pwrite-fsync.md, 09-fallocate.md
* sig/evfd via POLL_ADD↔ runtime/{signalfd,eventfd}.rs ↔ linux/04-signalfd.md, 02-eventfd.md
*
* Requires IORING_FEAT_SINGLE_MMAP (≥5.4) and multishot accept (≥5.19).
* Phase D opens WALs with O_TRUNC (fresh log each boot) — replay-on-boot and
* snapshots are phase E; the crash test inspects the WAL offline via
* `wal-check` BEFORE any restart. Simplifications: single-shot RECV re-armed
* per request, naive JSON extraction, fixed-size WAL payloads.
*/
#define _GNU_SOURCE
#include <arpa/inet.h>
#include <errno.h>
#include <fcntl.h>
#include <linux/io_uring.h>
#include <netinet/in.h>
#include <poll.h>
#include <pthread.h>
#include <sched.h>
#include <signal.h>
#include <stdatomic.h>
#include <stdint.h>
#include <stdio.h>
#include <stdlib.h>
#include <string.h>
#include <sys/eventfd.h>
#include <sys/mman.h>
#include <sys/resource.h>
#include <sys/signalfd.h>
#include <sys/socket.h>
#include <sys/stat.h>
#include <sys/syscall.h>
#include <time.h>
#include <unistd.h>
#ifndef IORING_ACCEPT_MULTISHOT
#define IORING_ACCEPT_MULTISHOT (1U << 0)
#endif
#define MAX_THREADS 64
#ifndef MAX_FDS
#define MAX_FDS 16384 /* conn slots per shard, fd-indexed */
#endif
#define IN_CAP 8192
#define OUT_CAP 65536
#define PAGE 4096
#define HUGE_2M (2u * 1024 * 1024)
#ifndef SLOT_SIZE
#define SLOT_SIZE 256 /* -D overridable for scale runs (phase F) */
#endif
#ifndef SLOTS_PER_SHARD
#define SLOTS_PER_SHARD 256
#endif
#define RING_ENTRIES 1024
#define WAL_PREALLOC (4u * 1024 * 1024) /* fallocate per shard */
#define WAL_COMMIT 0xC0FFEE42u /* frame trailer magic */
#define WAL_BATCH_CAP 65536 /* staged bytes per group commit */
#define WAL_BATCH_CONNS 256 /* acks parked per batch */
/* ---------------------------------------------------------------- crc32 --
* Hand-rolled (poly 0xEDB88320), table built once at boot. Zero deps. */
static uint32_t crc_table[256];
static void crc32_init(void) {
for (uint32_t i = 0; i < 256; i++) {
uint32_t c = i;
for (int k = 0; k < 8; k++) c = (c & 1) ? 0xEDB88320u ^ (c >> 1) : c >> 1;
crc_table[i] = c;
}
}
static uint32_t crc32(const void *buf, size_t len) {
const uint8_t *p = buf;
uint32_t c = 0xFFFFFFFFu;
while (len--) c = crc_table[(c ^ *p++) & 0xFF] ^ (c >> 8);
return c ^ 0xFFFFFFFFu;
}
/* ------------------------------------------------------------ WAL frame --
* [u32 len][u32 crc(payload)][payload][u32 WAL_COMMIT]. A record replays
* whole or not at all: bad len, bad crc, or missing trailer = torn tail.
* The payload carries the (shard,slot) coordinates phase B made stable. */
struct wal_payload {
uint32_t op; /* 1 = insert note */
uint32_t slot;
int32_t id;
char title[128];
};
#define WAL_FRAME_BYTES (4 + 4 + sizeof(struct wal_payload) + 4)
/* ---------------------------------------------------------------- arena --
* Unchanged from phase B: [header page][shard 0: bitmap page + slots]... */
struct arena_hdr {
char magic[8];
uint32_t version;
uint32_t n_shards;
uint32_t slots_per_shard;
uint32_t slot_size;
};
struct slot_note { int32_t id; char title[128]; };
static uint8_t *arena;
static size_t arena_bytes, arena_map_bytes, slice_bytes, bitmap_bytes;
static int arena_huge = 0, arena_locked = 0;
static int arena_init(int n_shards) {
bitmap_bytes = ((size_t)SLOTS_PER_SHARD / 8 + PAGE - 1) & ~((size_t)PAGE - 1);
slice_bytes = bitmap_bytes + (size_t)SLOTS_PER_SHARD * SLOT_SIZE;
arena_bytes = PAGE + (size_t)n_shards * slice_bytes;
arena_map_bytes = (arena_bytes + HUGE_2M - 1) & ~((size_t)HUGE_2M - 1);
arena = mmap(NULL, arena_map_bytes, PROT_READ | PROT_WRITE,
MAP_PRIVATE | MAP_ANONYMOUS | MAP_HUGETLB | MAP_POPULATE, -1, 0);
if (arena != MAP_FAILED) {
arena_huge = 1;
} else {
arena_map_bytes = (arena_bytes + PAGE - 1) & ~((size_t)PAGE - 1);
arena = mmap(NULL, arena_map_bytes, PROT_READ | PROT_WRITE,
MAP_PRIVATE | MAP_ANONYMOUS | MAP_POPULATE, -1, 0);
if (arena == MAP_FAILED) { perror("mmap arena"); return -1; }
}
arena_locked = (mlock(arena, arena_map_bytes) == 0);
if (!arena_locked)
fprintf(stderr, "[wo-rt-c] warn: mlock refused (%s) — arena not pinned\n", strerror(errno));
struct arena_hdr *hdr = (struct arena_hdr *)arena;
memcpy(hdr->magic, "WORTC\0\0", 8);
hdr->version = 3; /* phase C */
hdr->n_shards = (uint32_t)n_shards;
hdr->slots_per_shard = SLOTS_PER_SHARD;
hdr->slot_size = SLOT_SIZE;
return 0;
}
static uint64_t *shard_bitmap(int t) { return (uint64_t *)(arena + PAGE + (size_t)t * slice_bytes); }
static uint8_t *shard_slots (int t) { return arena + PAGE + (size_t)t * slice_bytes + bitmap_bytes; }
static struct slot_note *slot_at(int t, uint32_t i) {
return (struct slot_note *)(shard_slots(t) + (size_t)i * SLOT_SIZE);
}
/* ------------------------------------------------------------- io_uring --
* The raw ring: three pieces of memory shared with the kernel — the SQ/CQ
* ring headers+arrays (one mmap, IORING_FEAT_SINGLE_MMAP) and the SQE array.
* Submission: fill sqes[tail&mask], publish tail with a release store, tell
* the kernel with ONE io_uring_enter that also waits for completions. */
struct ring {
int fd;
unsigned *sq_head, *sq_tail, *sq_mask, *sq_array;
unsigned *cq_head, *cq_tail, *cq_mask;
struct io_uring_sqe *sqes;
struct io_uring_cqe *cqes;
unsigned local_tail; /* SQEs filled, not yet published */
unsigned to_submit;
};
static int ring_init(struct ring *r) {
struct io_uring_params p;
memset(&p, 0, sizeof p);
r->fd = (int)syscall(__NR_io_uring_setup, RING_ENTRIES, &p);
if (r->fd < 0) { perror("io_uring_setup"); return -1; }
if (!(p.features & IORING_FEAT_SINGLE_MMAP)) {
fprintf(stderr, "[wo-rt-c] kernel lacks IORING_FEAT_SINGLE_MMAP (need >= 5.4)\n");
return -1;
}
size_t sq_sz = p.sq_off.array + p.sq_entries * sizeof(unsigned);
size_t cq_sz = p.cq_off.cqes + p.cq_entries * sizeof(struct io_uring_cqe);
size_t sz = sq_sz > cq_sz ? sq_sz : cq_sz;
uint8_t *sqcq = mmap(NULL, sz, PROT_READ | PROT_WRITE, MAP_SHARED | MAP_POPULATE,
r->fd, IORING_OFF_SQ_RING);
if (sqcq == MAP_FAILED) { perror("mmap sq/cq ring"); return -1; }
r->sq_head = (unsigned *)(sqcq + p.sq_off.head);
r->sq_tail = (unsigned *)(sqcq + p.sq_off.tail);
r->sq_mask = (unsigned *)(sqcq + p.sq_off.ring_mask);
r->sq_array = (unsigned *)(sqcq + p.sq_off.array);
r->cq_head = (unsigned *)(sqcq + p.cq_off.head);
r->cq_tail = (unsigned *)(sqcq + p.cq_off.tail);
r->cq_mask = (unsigned *)(sqcq + p.cq_off.ring_mask);
r->cqes = (struct io_uring_cqe *)(sqcq + p.cq_off.cqes);
r->sqes = mmap(NULL, p.sq_entries * sizeof(struct io_uring_sqe),
PROT_READ | PROT_WRITE, MAP_SHARED | MAP_POPULATE,
r->fd, IORING_OFF_SQES);
if (r->sqes == MAP_FAILED) { perror("mmap sqes"); return -1; }
r->local_tail = *r->sq_tail;
r->to_submit = 0;
return 0;
}
static struct io_uring_sqe *sqe_get(struct ring *r) {
unsigned idx = r->local_tail & *r->sq_mask;
struct io_uring_sqe *s = &r->sqes[idx];
memset(s, 0, sizeof *s);
r->sq_array[idx] = idx;
r->local_tail++;
r->to_submit++;
return s;
}
/* user_data = (op << 32) | fd-or-batch-index */
enum { OP_ACCEPT = 1, OP_RECV, OP_SEND, OP_EVFD, OP_SIGFD, OP_WALWR, OP_FSYNC };
static uint64_t ud(int op, int fd) { return ((uint64_t)op << 32) | (uint32_t)fd; }
static int ring_enter(struct ring *r, unsigned wait) {
__atomic_store_n(r->sq_tail, r->local_tail, __ATOMIC_RELEASE);
unsigned n = r->to_submit;
r->to_submit = 0;
for (;;) {
int rc = (int)syscall(__NR_io_uring_enter, r->fd, n, wait,
IORING_ENTER_GETEVENTS, NULL, 0);
if (rc >= 0) return rc;
if (errno == EINTR) { n = 0; continue; } /* already submitted */
perror("io_uring_enter");
return -1;
}
}
/* ----------------------------------------------------------- connection --
* Keep-alive state machine. Exactly one outstanding SQE per connection:
* RECV while a request is being assembled, SEND while a response drains.
* Leftover bytes after a request (pipelining) are carried over and parsed
* before the next RECV is armed. */
struct conn {
char in[IN_CAP]; size_t in_len;
char out[OUT_CAP]; size_t out_len, out_off;
int in_use;
int closing; /* close once the out buffer drains */
int await_durable;/* response parked until this tick's fsync CQE */
uint64_t gen; /* incarnation stamp — kernel fds get reused */
};
/* One group-commit batch: staged WAL bytes + the connections whose acks ride
* its fsync. Double-buffered: while batch[k] is in flight (write→fsync
* linked SQEs), new commits stage into batch[k^1].
* Acks are parked as (fd, gen) pairs: an fd alone is ABA-unsafe — a parked
* connection can die, the kernel reuses its fd for a NEW connection whose
* commit sits in the OTHER batch, and a bare-fd release would ack that new
* connection before ITS record is durable. Found by the phase-F crash test
* (7 acked-but-unwritten records out of ~990k under reconnect churn). */
struct wal_batch {
char buf[WAL_BATCH_CAP];
size_t len;
int conns[WAL_BATCH_CONNS];
uint64_t gens[WAL_BATCH_CONNS];
int n_conns;
};
struct shard {
int id;
int lfd, evfd;
struct ring ring;
pthread_t tid;
int next_id;
_Atomic int used;
int wal_fd;
_Atomic size_t wal_off; /* owner-written; stats-readable cross-shard */
struct wal_batch batch[2];
int active; /* batch being staged */
int in_flight; /* a write→fsync pair is on the ring */
char snap_path[320];
struct conn conns[MAX_FDS];
};
/* Snapshot file: [snap_hdr][bitmap page][slot bytes]. Written on clean
* shutdown, loaded at boot before WAL replay. */
struct snap_hdr {
char magic[8]; /* "WOSNAP\0\0" */
uint32_t version;
int32_t next_id;
int32_t used;
};
static struct shard *shards;
static int n_threads = 1;
static int sigfd = -1;
static _Atomic unsigned long reqs[MAX_THREADS];
static int slot_alloc(struct shard *sh) {
uint64_t *bm = shard_bitmap(sh->id);
for (uint32_t w = 0; w < SLOTS_PER_SHARD / 64; w++) {
if (bm[w] == UINT64_MAX) continue;
uint32_t b = (uint32_t)__builtin_ctzll(~bm[w]);
uint32_t i = w * 64 + b;
if (i >= SLOTS_PER_SHARD) break;
bm[w] |= (1ULL << b);
atomic_fetch_add_explicit(&sh->used, 1, memory_order_relaxed);
return (int)i;
}
return -1;
}
/* --------------------------------------------------------- SQE builders -- */
static void arm_accept(struct shard *sh) { /* multishot: arm once */
struct io_uring_sqe *s = sqe_get(&sh->ring);
s->opcode = IORING_OP_ACCEPT;
s->fd = sh->lfd;
s->ioprio = IORING_ACCEPT_MULTISHOT;
s->user_data = ud(OP_ACCEPT, sh->lfd);
}
static void arm_poll(struct shard *sh, int fd, int op) {
struct io_uring_sqe *s = sqe_get(&sh->ring);
s->opcode = IORING_OP_POLL_ADD;
s->fd = fd;
s->poll_events = POLLIN;
s->user_data = ud(op, fd);
}
static void arm_recv(struct shard *sh, int fd) {
struct conn *c = &sh->conns[fd];
struct io_uring_sqe *s = sqe_get(&sh->ring);
s->opcode = IORING_OP_RECV;
s->fd = fd;
s->addr = (uint64_t)(uintptr_t)(c->in + c->in_len);
s->len = (uint32_t)(IN_CAP - 1 - c->in_len);
s->user_data = ud(OP_RECV, fd);
}
static void arm_send(struct shard *sh, int fd) {
struct conn *c = &sh->conns[fd];
struct io_uring_sqe *s = sqe_get(&sh->ring);
s->opcode = IORING_OP_SEND;
s->fd = fd;
s->addr = (uint64_t)(uintptr_t)(c->out + c->out_off);
s->len = (uint32_t)(c->out_len - c->out_off);
s->msg_flags = MSG_NOSIGNAL;
s->user_data = ud(OP_SEND, fd);
}
/* ------------------------------------------------------------ WAL flush --
* Called once per loop tick. If commits were staged and no batch is in
* flight, submit ONE write SQE for the whole batch at the shard's tail
* offset, hard-linked to ONE fdatasync SQE. Every parked ack in the batch
* is released when the fsync CQE arrives — group commit. */
static void wal_flush(struct shard *sh) {
if (sh->in_flight) return;
struct wal_batch *b = &sh->batch[sh->active];
if (b->len == 0) return;
size_t off = atomic_load_explicit(&sh->wal_off, memory_order_relaxed);
struct io_uring_sqe *w = sqe_get(&sh->ring);
w->opcode = IORING_OP_WRITE;
w->fd = sh->wal_fd;
w->addr = (uint64_t)(uintptr_t)b->buf;
w->len = (uint32_t)b->len;
w->off = off;
w->flags = IOSQE_IO_LINK; /* fsync follows the write */
w->user_data = ud(OP_WALWR, sh->active);
struct io_uring_sqe *f = sqe_get(&sh->ring);
f->opcode = IORING_OP_FSYNC;
f->fd = sh->wal_fd;
f->fsync_flags = IORING_FSYNC_DATASYNC;
f->user_data = ud(OP_FSYNC, sh->active);
sh->in_flight = 1;
sh->active ^= 1; /* new commits stage in the twin */
}
/* Stage one commit's frame + park the connection's ack on the active batch.
* Returns 0 if the batch has no room (caller responds 503, no RAM apply). */
static int wal_append(struct shard *sh, int connfd, uint32_t slot, int32_t id, const char *title) {
struct wal_batch *b = &sh->batch[sh->active];
if (b->len + WAL_FRAME_BYTES > WAL_BATCH_CAP || b->n_conns >= WAL_BATCH_CONNS)
return 0;
struct wal_payload p;
memset(&p, 0, sizeof p);
p.op = 1;
p.slot = slot;
p.id = id;
snprintf(p.title, sizeof p.title, "%s", title);
b->gens[b->n_conns] = sh->conns[connfd].gen;
uint32_t len = (uint32_t)sizeof p;
uint32_t crc = crc32(&p, sizeof p);
uint32_t end = WAL_COMMIT;
char *dst = b->buf + b->len;
memcpy(dst, &len, 4);
memcpy(dst + 4, &crc, 4);
memcpy(dst + 8, &p, sizeof p);
memcpy(dst + 8 + sizeof p, &end, 4);
b->len += WAL_FRAME_BYTES;
b->conns[b->n_conns++] = connfd;
return 1;
}
/* ----------------------------------------------------------------- http -- */
static void respond(struct conn *c, const char *status, const char *ctype, const char *body) {
size_t blen = strlen(body);
int n = snprintf(c->out + c->out_len, OUT_CAP - c->out_len,
"HTTP/1.1 %s\r\nContent-Type: %s\r\nContent-Length: %zu\r\nConnection: %s\r\n\r\n%s",
status, ctype, blen, c->closing ? "close" : "keep-alive", body);
if (n > 0 && (size_t)n < OUT_CAP - c->out_len) c->out_len += (size_t)n;
else c->closing = 1; /* response too big — drop conn */
}
static int json_title(const char *body, char *out, size_t cap) {
const char *p = strstr(body, "\"title\"");
if (!p) return 0;
p = strchr(p + 7, ':'); if (!p) return 0;
p = strchr(p, '"'); if (!p) return 0;
p++;
size_t i = 0;
while (*p && *p != '"' && i + 1 < cap) out[i++] = *p++;
out[i] = 0;
return i > 0;
}
static void route(struct shard *sh, struct conn *c, const char *method, const char *path, const char *body) {
atomic_fetch_add_explicit(&reqs[sh->id], 1, memory_order_relaxed);
if (!strcmp(method, "GET") && !strcmp(path, "/")) {
char out[1024];
size_t off = (size_t)snprintf(out, sizeof out,
"{\"runtime\":\"wo-rt-c\",\"loop\":\"io_uring\",\"threads\":%d,\"shard\":%d,"
"\"arena\":{\"bytes\":%zu,\"mapped\":%zu,\"hugepages\":%s,\"mlocked\":%s,"
"\"slot_size\":%d,\"slots_per_shard\":%d},\"shard_used\":[",
n_threads, sh->id, arena_bytes, arena_map_bytes,
arena_huge ? "true" : "false", arena_locked ? "true" : "false",
SLOT_SIZE, SLOTS_PER_SHARD);
for (int t = 0; t < n_threads; t++)
off += (size_t)snprintf(out + off, sizeof out - off, "%s%d", t ? "," : "",
atomic_load_explicit(&shards[t].used, memory_order_relaxed));
off += (size_t)snprintf(out + off, sizeof out - off, "],\"shard_requests\":[");
for (int t = 0; t < n_threads; t++)
off += (size_t)snprintf(out + off, sizeof out - off, "%s%lu", t ? "," : "",
atomic_load_explicit(&reqs[t], memory_order_relaxed));
off += (size_t)snprintf(out + off, sizeof out - off, "],\"wal_bytes\":[");
for (int t = 0; t < n_threads; t++)
off += (size_t)snprintf(out + off, sizeof out - off, "%s%zu", t ? "," : "",
atomic_load_explicit(&shards[t].wal_off, memory_order_relaxed));
snprintf(out + off, sizeof out - off, "]}");
respond(c, "200 OK", "application/json", out);
} else if (!strcmp(method, "GET") && !strcmp(path, "/healthz")) {
respond(c, "200 OK", "text/plain", "ok");
} else if (!strcmp(method, "GET") && !strcmp(path, "/api/notes")) {
static _Thread_local char out[SLOTS_PER_SHARD * 160 + 64];
uint64_t *bm = shard_bitmap(sh->id);
size_t off = (size_t)snprintf(out, sizeof out, "{\"shard\":%d,\"notes\":[", sh->id);
int first = 1;
for (uint32_t i = 0; i < SLOTS_PER_SHARD; i++) {
if (!(bm[i / 64] & (1ULL << (i % 64)))) continue;
struct slot_note *n = slot_at(sh->id, i);
off += (size_t)snprintf(out + off, sizeof out - off,
"%s{\"id\":%d,\"title\":\"%s\"}", first ? "" : ",", n->id, n->title);
first = 0;
}
snprintf(out + off, sizeof out - off, "]}");
respond(c, "200 OK", "application/json", out);
} else if (!strcmp(method, "POST") && !strcmp(path, "/api/notes")) {
char title[128];
int i;
if (!json_title(body, title, sizeof title)) {
respond(c, "400 Bad Request", "application/json", "{\"error\":\"expected {\\\"title\\\":\\\"...\\\"}\"}");
return;
}
/* Capacity gates BEFORE the RAM apply — an aborted op writes nothing. */
struct wal_batch *b = &sh->batch[sh->active];
if (b->len + WAL_FRAME_BYTES > WAL_BATCH_CAP || b->n_conns >= WAL_BATCH_CONNS) {
respond(c, "503 Service Unavailable", "application/json", "{\"error\":\"commit batch full, retry\"}");
return;
}
if ((i = slot_alloc(sh)) < 0) {
respond(c, "507 Insufficient Storage", "application/json", "{\"error\":\"shard full\"}");
return;
}
struct slot_note *n = slot_at(sh->id, (uint32_t)i); /* 1. the RAM apply */
n->id = sh->next_id;
sh->next_id += n_threads;
snprintf(n->title, sizeof n->title, "%s", title);
char out[224];
snprintf(out, sizeof out, "{\"id\":%d,\"title\":\"%s\",\"shard\":%d,\"slot\":%d}",
n->id, n->title, sh->id, i);
respond(c, "201 Created", "application/json", out); /* built, NOT sent */
int connfd = (int)(c - sh->conns); /* conns is fd-indexed */
wal_append(sh, connfd, (uint32_t)i, n->id, n->title);/* 2. stage WAL frame */
c->await_durable = 1; /* 4. ack rides fsync */
} else {
respond(c, "404 Not Found", "application/json", "{\"error\":\"no such route\"}");
}
}
/* ----------------------------------------------------- state machine ----- */
static void conn_open(struct shard *sh, int fd) {
struct conn *c = &sh->conns[fd];
c->in_len = c->out_len = c->out_off = 0;
c->closing = 0;
c->await_durable = 0;
c->gen++; /* new incarnation — stale parked acks won't match */
c->in_use = 1;
}
static void conn_close(struct shard *sh, int fd) {
if (fd >= 0 && fd < MAX_FDS) sh->conns[fd].in_use = 0;
close(fd);
}
/* Try to consume ONE complete request from the in buffer. Returns 1 if a
* response was produced (out has bytes), 0 if the request is incomplete. */
static int try_process(struct shard *sh, struct conn *c) {
c->in[c->in_len] = 0;
char *hdr_end = strstr(c->in, "\r\n\r\n");
if (!hdr_end) return 0;
char *body = hdr_end + 4;
size_t total = (size_t)(body - c->in);
const char *cl = strcasestr(c->in, "Content-Length:");
if (cl) {
long want = strtol(cl + 15, NULL, 10);
if (want < 0) want = 0;
if (c->in_len < total + (size_t)want) return 0;
total += (size_t)want;
}
/* HTTP/1.1 defaults to keep-alive; honor an explicit close. */
if (strcasestr(c->in, "connection: close") ||
(strstr(c->in, "HTTP/1.0") && !strcasestr(c->in, "connection: keep-alive")))
c->closing = 1;
char method[8] = {0}, path[256] = {0};
if (sscanf(c->in, "%7s %255s", method, path) == 2)
route(sh, c, method, path, body);
else
c->closing = 1;
memmove(c->in, c->in + total, c->in_len - total); /* carry pipelined tail */
c->in_len -= total;
return 1;
}
/* Advance a connection: drain out via SEND, else parse, else arm RECV.
* A parked commit ack arms nothing — the fsync CQE handler resumes it. */
static void conn_continue(struct shard *sh, int fd) {
struct conn *c = &sh->conns[fd];
if (c->await_durable) { return; }
if (c->out_off < c->out_len) { arm_send(sh, fd); return; }
c->out_off = c->out_len = 0;
if (c->closing) { conn_close(sh, fd); return; }
if (try_process(sh, c)) {
/* route() may have JUST parked this response (await set inside
* try_process) — sending now would race the fsync. The fsync CQE
* re-enters here with await cleared and arms the send.
* (Found by the phase-F crash-under-load test: ~6 acked-but-
* unwritten records per ~750k at the kill instant.) */
if (!c->await_durable) arm_send(sh, fd);
return;
}
if (c->in_len >= IN_CAP - 1) { conn_close(sh, fd); return; } /* oversize head */
arm_recv(sh, fd);
}
/* --------------------------------------------------------------- recovery --
* First load: hard drive → RAM, per shard, in parallel, BEFORE accept arms.
* Snapshot (if any) restores the slice wholesale; the WAL tail replays
* commits since that snapshot. Frame validation is wal-check's logic with
* the printf swapped for the arena apply. */
static long now_ms(void) {
struct timespec ts;
clock_gettime(CLOCK_MONOTONIC, &ts);
return ts.tv_sec * 1000 + ts.tv_nsec / 1000000;
}
static int snap_load(struct shard *sh) {
int fd = open(sh->snap_path, O_RDONLY | O_CLOEXEC);
if (fd < 0) return 0;
struct snap_hdr h;
size_t slot_bytes = (size_t)SLOTS_PER_SHARD * SLOT_SIZE;
if (read(fd, &h, sizeof h) != (ssize_t)sizeof h ||
memcmp(h.magic, "WOSNAP\0", 8) != 0 ||
pread(fd, shard_bitmap(sh->id), bitmap_bytes, (off_t)sizeof h) != (ssize_t)bitmap_bytes ||
pread(fd, shard_slots(sh->id), slot_bytes, (off_t)(sizeof h + bitmap_bytes)) != (ssize_t)slot_bytes) {
fprintf(stderr, "[wo-rt-c] shard %d: snapshot unreadable — starting from WAL only\n", sh->id);
memset(shard_bitmap(sh->id), 0, bitmap_bytes + slot_bytes);
close(fd);
return 0;
}
sh->next_id = h.next_id;
atomic_store_explicit(&sh->used, h.used, memory_order_relaxed);
close(fd);
return h.used;
}
static int wal_replay(struct shard *sh) {
size_t off = 0;
int recs = 0;
int32_t maxid = 0;
for (;;) {
uint32_t len, crc, end;
struct wal_payload p;
if (pread(sh->wal_fd, &len, 4, (off_t)off) != 4 || len == 0) break;
if (len != sizeof p ||
pread(sh->wal_fd, &crc, 4, (off_t)(off + 4)) != 4 ||
pread(sh->wal_fd, &p, sizeof p, (off_t)(off + 8)) != (ssize_t)sizeof p ||
pread(sh->wal_fd, &end, 4, (off_t)(off + 8 + sizeof p)) != 4 ||
crc32(&p, sizeof p) != crc || end != WAL_COMMIT) {
fprintf(stderr, "[wo-rt-c] shard %d: torn WAL record at byte %zu — truncating\n",
sh->id, off);
break;
}
if (p.op == 1 && p.slot < SLOTS_PER_SHARD) { /* idempotent apply */
uint64_t *bm = shard_bitmap(sh->id);
if (!(bm[p.slot / 64] & (1ULL << (p.slot % 64)))) {
bm[p.slot / 64] |= (1ULL << (p.slot % 64));
atomic_fetch_add_explicit(&sh->used, 1, memory_order_relaxed);
}
struct slot_note *n = slot_at(sh->id, p.slot);
n->id = p.id;
snprintf(n->title, sizeof n->title, "%s", p.title);
if (p.id > maxid) maxid = p.id;
}
recs++;
off += WAL_FRAME_BYTES;
}
/* Resume appends at the validated tail; drop torn bytes, re-preallocate. */
if (ftruncate(sh->wal_fd, (off_t)off) == 0)
(void)!fallocate(sh->wal_fd, 0, 0, off > WAL_PREALLOC ? off : WAL_PREALLOC);
atomic_store_explicit(&sh->wal_off, off, memory_order_relaxed);
if (maxid > 0 && maxid + n_threads > sh->next_id)
sh->next_id = maxid + n_threads; /* interleaved high-water */
return recs;
}
/* Clean-shutdown snapshot: write slice → fsync → atomic rename → truncate WAL. */
static void snap_write(struct shard *sh) {
char tmp[336];
snprintf(tmp, sizeof tmp, "%s.tmp", sh->snap_path);
int fd = open(tmp, O_WRONLY | O_CREAT | O_TRUNC | O_CLOEXEC, 0644);
if (fd < 0) { perror("snapshot open"); return; }
struct snap_hdr h;
memset(&h, 0, sizeof h);
memcpy(h.magic, "WOSNAP\0", 8);
h.version = 1;
h.next_id = sh->next_id;
h.used = atomic_load_explicit(&sh->used, memory_order_relaxed);
size_t slot_bytes = (size_t)SLOTS_PER_SHARD * SLOT_SIZE;
int ok = write(fd, &h, sizeof h) == (ssize_t)sizeof h
&& write(fd, shard_bitmap(sh->id), bitmap_bytes) == (ssize_t)bitmap_bytes
&& write(fd, shard_slots(sh->id), slot_bytes) == (ssize_t)slot_bytes
&& fsync(fd) == 0;
close(fd);
if (!ok || rename(tmp, sh->snap_path) < 0) { fprintf(stderr, "[wo-rt-c] shard %d: snapshot failed\n", sh->id); unlink(tmp); return; }
if (ftruncate(sh->wal_fd, 0) == 0) { /* WAL now redundant */
(void)!fallocate(sh->wal_fd, 0, 0, WAL_PREALLOC);
fsync(sh->wal_fd);
}
printf("[wo-rt-c] shard %d: snapshot %d rows → %s, wal truncated\n", sh->id, h.used, sh->snap_path);
}
/* ------------------------------------------------------------ shard loop -- */
static void *shard_main(void *arg) {
struct shard *sh = arg;
struct ring *r = &sh->ring;
cpu_set_t set;
CPU_ZERO(&set);
CPU_SET((unsigned)sh->id % (unsigned)sysconf(_SC_NPROCESSORS_ONLN), &set);
pthread_setaffinity_np(pthread_self(), sizeof set, &set);
/* First load: disk → RAM, before any accept is armed. */
long t0 = now_ms();
int srows = snap_load(sh);
int wrecs = wal_replay(sh);
if (srows || wrecs)
printf("[wo-rt-c] shard %d: recovered %d snapshot rows + %d wal records in %ld ms\n",
sh->id, srows, wrecs, now_ms() - t0);
arm_accept(sh);
arm_poll(sh, sh->evfd, OP_EVFD);
if (sh->id == 0) arm_poll(sh, sigfd, OP_SIGFD);
for (;;) {
if (ring_enter(r, 1) < 0) break; /* ONE syscall per tick */
unsigned head = *r->cq_head;
unsigned tail = __atomic_load_n(r->cq_tail, __ATOMIC_ACQUIRE);
for (; head != tail; head++) {
struct io_uring_cqe *cqe = &r->cqes[head & *r->cq_mask];
int op = (int)(cqe->user_data >> 32);
int fd = (int)(uint32_t)cqe->user_data;
int res = cqe->res;
switch (op) {
case OP_EVFD: {
uint64_t v;
(void)!read(sh->evfd, &v, sizeof v);
__atomic_store_n(r->cq_head, head + 1, __ATOMIC_RELEASE);
return NULL;
}
case OP_SIGFD: { /* shard 0 only */
struct signalfd_siginfo si;
if (read(sigfd, &si, sizeof si) == sizeof si)
printf("\n[wo-rt-c] signal %u — broadcasting shutdown to %d shards\n",
si.ssi_signo, n_threads);
uint64_t one = 1;
for (int t = 0; t < n_threads; t++)
(void)!write(shards[t].evfd, &one, sizeof one);
break;
}
case OP_ACCEPT: {
if (res >= 0) {
int cfd = res;
if (cfd >= MAX_FDS) close(cfd);
else { conn_open(sh, cfd); arm_recv(sh, cfd); }
}
if (!(cqe->flags & IORING_CQE_F_MORE)) arm_accept(sh); /* re-arm */
break;
}
case OP_RECV: {
struct conn *c = &sh->conns[fd];
if (!c->in_use) break;
if (res <= 0) { conn_close(sh, fd); break; }
c->in_len += (size_t)res;
conn_continue(sh, fd);
break;
}
case OP_SEND: {
struct conn *c = &sh->conns[fd];
if (!c->in_use) break;
if (res <= 0) { conn_close(sh, fd); break; }
c->out_off += (size_t)res;
conn_continue(sh, fd);
break;
}
case OP_WALWR: { /* fd field carries the batch idx */
struct wal_batch *b = &sh->batch[fd];
if (res != (int)b->len)
fprintf(stderr, "[wo-rt-c] shard %d: WAL write %d != %zu\n", sh->id, res, b->len);
else
atomic_fetch_add_explicit(&sh->wal_off, b->len, memory_order_relaxed);
break;
}
case OP_FSYNC: { /* group commit lands: release acks */
struct wal_batch *b = &sh->batch[fd];
int failed = (res < 0); /* incl. -ECANCELED from a failed link */
if (failed)
fprintf(stderr, "[wo-rt-c] shard %d: fsync failed (%d) — dropping %d acks\n",
sh->id, res, b->n_conns);
for (int k = 0; k < b->n_conns; k++) {
int cfd = b->conns[k];
if (cfd < 0 || cfd >= MAX_FDS) continue;
struct conn *c = &sh->conns[cfd];
if (!c->in_use || !c->await_durable) continue;
if (c->gen != b->gens[k]) continue; /* fd reused — not ours */
c->await_durable = 0;
if (failed) conn_close(sh, cfd); /* never ack non-durable */
else conn_continue(sh, cfd);
}
b->len = 0;
b->n_conns = 0;
sh->in_flight = 0;
break;
}
}
}
__atomic_store_n(r->cq_head, head, __ATOMIC_RELEASE);
wal_flush(sh); /* one write→fsync pair per tick */
}
return NULL;
}
/* ------------------------------------------------------------- listener -- */
static int listener_bind(uint16_t port) {
int fd = socket(AF_INET, SOCK_STREAM | SOCK_CLOEXEC, 0);
if (fd < 0) { perror("socket"); return -1; }
int one = 1;
setsockopt(fd, SOL_SOCKET, SO_REUSEADDR, &one, sizeof one);
setsockopt(fd, SOL_SOCKET, SO_REUSEPORT, &one, sizeof one);
struct sockaddr_in addr = {0};
addr.sin_family = AF_INET;
addr.sin_port = htons(port);
addr.sin_addr.s_addr = htonl(INADDR_LOOPBACK);
if (bind(fd, (struct sockaddr *)&addr, sizeof addr) < 0) { perror("bind"); close(fd); return -1; }
if (listen(fd, SOMAXCONN) < 0) { perror("listen"); close(fd); return -1; }
return fd;
}
static int sig_setup(void) {
sigset_t mask;
sigemptyset(&mask);
sigaddset(&mask, SIGINT);
sigaddset(&mask, SIGTERM);
if (sigprocmask(SIG_BLOCK, &mask, NULL) < 0) { perror("sigprocmask"); return -1; }
int fd = signalfd(-1, &mask, SFD_NONBLOCK | SFD_CLOEXEC);
if (fd < 0) perror("signalfd");
return fd;
}
/* ------------------------------------------------------------ wal-check --
* Offline frame walker: validates every record's len/CRC/COMMIT trailer,
* reports the count and where (if anywhere) the log tears. This is the
* crash test's witness, and the skeleton of phase E's replay loop. */
static int wal_check(const char *path) {
int fd = open(path, O_RDONLY);
if (fd < 0) { fprintf(stderr, "wal-check: %s: %s\n", path, strerror(errno)); return 1; }
size_t off = 0;
int recs = 0;
for (;;) {
uint32_t len, crc, end;
struct wal_payload p;
if (pread(fd, &len, 4, (off_t)off) != 4) break;
if (len == 0) break; /* fallocate'd tail */
if (len != sizeof p) {
printf("%s: TORN at byte %zu (bad len %u) — %d whole records before it\n",
path, off, len, recs);
close(fd);
return 0;
}
if (pread(fd, &crc, 4, (off_t)(off + 4)) != 4 ||
pread(fd, &p, sizeof p, (off_t)(off + 8)) != (ssize_t)sizeof p ||
pread(fd, &end, 4, (off_t)(off + 8 + sizeof p)) != 4 ||
crc32(&p, sizeof p) != crc || end != WAL_COMMIT) {
printf("%s: TORN at byte %zu (bad crc/trailer) — %d whole records before it\n",
path, off, recs);
close(fd);
return 0;
}
printf("%s: rec %d op=%u slot=%u id=%d title=\"%s\"\n", path, recs, p.op, p.slot, p.id, p.title);
recs++;
off += WAL_FRAME_BYTES;
}
printf("%s: %d records, all frames valid, clean tail at byte %zu\n", path, recs, off);
close(fd);
return 0;
}
/* ----------------------------------------------------------------- main -- */
int main(int argc, char **argv) {
crc32_init();
if (argc >= 3 && !strcmp(argv[1], "wal-check")) {
int rc = 0;
for (int a = 2; a < argc; a++) rc |= wal_check(argv[a]);
return rc;
}
uint16_t port = 8085;
const char *env = getenv("WO_PORT");
if (env && atoi(env) > 0) port = (uint16_t)atoi(env);
long cores = sysconf(_SC_NPROCESSORS_ONLN);
n_threads = (int)cores;
env = getenv("WO_THREADS");
if (env && atoi(env) > 0) n_threads = atoi(env);
if (n_threads < 1) n_threads = 1;
if (n_threads > MAX_THREADS) n_threads = MAX_THREADS;
/* Million-connection posture: lift the fd ceiling to the hard max. */
struct rlimit rl;
if (getrlimit(RLIMIT_NOFILE, &rl) == 0 && rl.rlim_cur < rl.rlim_max) {
rl.rlim_cur = rl.rlim_max;
setrlimit(RLIMIT_NOFILE, &rl);
}
sigfd = sig_setup();
if (sigfd < 0) return 1;
if (arena_init(n_threads) < 0) return 1;
const char *data_dir = getenv("WO_DATA");
if (!data_dir || !*data_dir) data_dir = "./wo-data";
if (mkdir(data_dir, 0755) < 0 && errno != EEXIST) { perror("mkdir data dir"); return 1; }
/* The data dir is sharded for exactly n_threads. A different WO_THREADS
* would strand WAL/snapshot files silently — refuse (resharding = 09f). */
char mpath[512];
snprintf(mpath, sizeof mpath, "%s/meta", data_dir);
FILE *mf = fopen(mpath, "r");
if (mf) {
int prev = 0;
if (fscanf(mf, "%d", &prev) == 1 && prev != n_threads) {
fprintf(stderr, "[wo-rt-c] %s was written with WO_THREADS=%d — restart with that, or wipe the dir\n",
data_dir, prev);
fclose(mf);
return 1;
}
fclose(mf);
} else if ((mf = fopen(mpath, "w"))) {
fprintf(mf, "%d\n", n_threads);
fclose(mf);
}
shards = calloc((size_t)n_threads, sizeof(struct shard));
if (!shards) { perror("calloc"); return 1; }
for (int t = 0; t < n_threads; t++) {
struct shard *sh = &shards[t];
sh->id = t;
sh->next_id = t + 1;
sh->lfd = listener_bind(port);
sh->evfd = eventfd(0, EFD_NONBLOCK | EFD_CLOEXEC);
if (sh->lfd < 0 || sh->evfd < 0) return 1;
if (ring_init(&sh->ring) < 0) return 1;
/* Per-shard WAL, preallocated, NOT truncated — boot replays it. */
char path[512];
snprintf(path, sizeof path, "%s/shard-%d.wal", data_dir, t);
sh->wal_fd = open(path, O_RDWR | O_CREAT | O_CLOEXEC, 0644);
if (sh->wal_fd < 0) { perror("open wal"); return 1; }
if (fallocate(sh->wal_fd, 0, 0, WAL_PREALLOC) < 0)
fprintf(stderr, "[wo-rt-c] warn: fallocate %s refused (%s)\n", path, strerror(errno));
snprintf(sh->snap_path, sizeof sh->snap_path, "%s/shard-%d.data", data_dir, t);
}
printf("[wo-rt-c] %d shard%s on http://127.0.0.1:%u — io_uring loops, keep-alive — arena %zu KB (%s pages, %s) — ctrl-C to stop\n",
n_threads, n_threads == 1 ? "" : "s", port,
arena_map_bytes / 1024,
arena_huge ? "2M huge" : "4K",
arena_locked ? "mlocked" : "NOT locked");
printf(" GET / runtime + arena info, per-shard stats\n GET /healthz liveness\n");
printf(" GET /api/notes list (connection's shard)\n POST /api/notes create {\"title\":\"...\"}\n");
fflush(stdout);
for (int t = 0; t < n_threads; t++)
pthread_create(&shards[t].tid, NULL, shard_main, &shards[t]);
for (int t = 0; t < n_threads; t++)
pthread_join(shards[t].tid, NULL);
/* Clean shutdown: persist each slice as a snapshot, truncate the WALs.
* A kill -9 skips this — that's what boot-time WAL replay is for. */
for (int t = 0; t < n_threads; t++)
snap_write(&shards[t]);
printf("[wo-rt-c] all %d shards joined — bye\n", n_threads);
for (int t = 0; t < n_threads; t++) {
close(shards[t].ring.fd); close(shards[t].lfd); close(shards[t].evfd);
close(shards[t].wal_fd);
}
close(sigfd);
munmap(arena, arena_map_bytes);
free(shards);
return 0;
}