writeonce/prototypes/wo-rt-c
shoney.arickathil 929b2802c4 add wo-rt-c: C runtime prototype, phases A-F shipped
prototypes/wo-rt-c — single-file C reference of the writeonce runtime
layer, zero deps beyond libc + kernel uapi: thread-per-core io_uring
event loops (raw syscalls, no liburing), SO_REUSEPORT listeners, one
mlock'd mmap arena sharded by address, WAL dual-write with group commit
(HTTP ack only after the fsync CQE), boot-time snapshot + WAL replay
recovery, and a bench harness with a Go net/http comparison server.

Measured on 20 cores: 908k reads/s p99 154us and 643k fsync-acked
commits/s p99 177us on 8 shards, vs Go net/http 495k/355k (no
durability) on 20 cores. Crash-under-load testing found and fixed an
ack-before-fsync race and an fd-reuse ABA hazard in commit-ack parking.

docs/plan/exploration/c-runtime — the phased plan (00, exit evidence
per phase), the one-address architecture trace (01), and the
single-binary end-goal contract (02). justfile carries the demo and
bench recipes; .gitignore covers binaries and data dirs.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-06-12 23:05:13 +02:00
..
bench add wo-rt-c: C runtime prototype, phases A-F shipped 2026-06-12 23:05:13 +02:00
Makefile add wo-rt-c: C runtime prototype, phases A-F shipped 2026-06-12 23:05:13 +02:00
README.md add wo-rt-c: C runtime prototype, phases A-F shipped 2026-06-12 23:05:13 +02:00
wo-rt.c add wo-rt-c: C runtime prototype, phases A-F shipped 2026-06-12 23:05:13 +02:00

wo-rt-c — the writeonce runtime environment, in C

A single-file C implementation of the runtime layer the writeonce language runs on — now at phase E: a durable RAM database. Writes follow the dual-write order — RAM apply, framed WAL record (len|crc32|payload|COMMIT) to a per-shard fallocate'd log, one group-commit fdatasync per loop tick, HTTP ack only after the fsync completion. Boot performs the first load, hard drive → RAM: each shard replays its snapshot + WAL tail into its arena slice in parallel before any accept arms; clean shutdown snapshots each slice and truncates the WAL; a meta file pins the shard count so a mismatched WO_THREADS refuses to boot. ./wo-rt wal-check <file> validates a log offline. WO_THREADS pinned threads (default = online cores), each owning its own raw io_uring ring (io_uring_setup + mmap'd SQ/CQ rings + io_uring_enter — no liburing), its own SO_REUSEPORT listener (multishot accept), its own keep-alive connections, and its own slice of the one mlock'd mmap arena — shared-nothing, no locks. Steady state is one io_uring_enter syscall per loop tick. Thread 0 owns the signalfd; shutdown broadcasts through per-thread eventfds, both watched via POLL_ADD SQEs. Zero dependencies beyond libc + kernel uapi headers. The kernel is the runtime.

This is the runtime-layer sibling of prototypes/wo-db/ (the C++ query-layer prototype): a reference card showing, with no abstraction in the way, exactly which kernel primitives the production Rust runtime (crates/rt/) drives through libc. Same role, different layer.

prototypes/wo-db/     C++   what the LANGUAGE executes   (parser, engine, transactions)
prototypes/wo-rt-c/   C     what the RUNTIME stands on   (epoll, signalfd, sockets)
crates/rt/            Rust  the product — both layers, libc only

Build, run, poke

make                 # cc -O2 -Wall -Wextra -std=c11 -pthread — no libraries
./wo-rt              # 127.0.0.1:8085   (WO_PORT=9000 WO_THREADS=4 ./wo-rt to override)

curl localhost:8085/   # {"runtime":"wo-rt-c","loop":"epoll-et","threads":4,
                       #  "shard":2,"shard_requests":[68,36,44,53]}
curl -X POST localhost:8085/api/notes -d '{"title":"hello"}'
                       # {"id":2,"title":"hello","shard":1}   ← ids interleave per shard
curl localhost:8085/api/notes        # the connection's shard only — shared-nothing
# ctrl-C → signalfd on shard 0 → eventfd broadcast → all shards join

Each connection hashes to one shard for life (SO_REUSEPORT 4-tuple): a list may land on a different shard than the create that preceded it. That is the architecture, not a bug — cross-shard reads are a later phase / design decision (see the architecture doc's improvements).

Or from the repo root: just rt-c-demo.

Module map

Every block in wo-rt.c corresponds one-to-one to a module of the Rust runtime, which in turn mirrors Go's netpoller — the same lineage the docs trace:

wo-rt.c block Rust (crates/rt/src/) Go (reference/go/src/runtime/) Kernel reference card
main event loop (epoll_create1 / epoll_wait, EPOLLET) runtime/netpoll_epoll.rs netpoll_epoll.go linux/01-epoll.md
sig_setup (sigprocmask + signalfd) runtime/signalfd.rs signal mask handling linux/04-signalfd.md
listener_bind (SOCK_NONBLOCK, accept4-to-EAGAIN) http/listener.rs net.Listen + accept loop socket(7)
conn_drive (read-to-EAGAIN, one buffer per fd) http/connection.rs conn.Read loop the edge-triggered contract
notes[] store engine.rs (BTreeMaps) — 03-inmemory-engine.md

What it demonstrates

  • One thread owns each shard outright. Accept, parse, store, respond — no locks, no worker pool, no connection migration. Scaling past one core is more shards (09-concurrency-scaleout.md), never shared mutable state. The single cross-thread touch is the relaxed-atomic stats counters on / — monotonic, never on the data path.
  • Edge-triggered discipline. Every registration sets EPOLLET; every readiness event is drained to EAGAIN (the accept loop and the read loop both). Get this wrong and connections silently hang — the reason the Rust module documents the same contract at the top of netpoll_epoll.rs.
  • Signals as fd events. SIGINT/SIGTERM are blocked, then read from a signalfd on the same epoll — no async-signal-unsafe handler, no self-pipe trick.
  • RAM is the read path. GET /api/notes touches a C array. The production engine is the same idea with MVCC and a WAL behind it.

Architecture and roadmap

Documentation lives under docs/ (repo convention) — this README stays here as the directory's orientation page only:

Measured (phase F, 20-core Linux 6.14, tmpfs data dir, just rt-c-bench)

Same C bench client (bench/bench.c, keep-alive, only 2xx counted) against both servers:

Benchmark wo-rt-c (8 shards, durable WAL) Go net/http (go1.25.1, 20 cores, no durability)
GET /healthz 908,916 req/s · p50 66 µs · p99 154 µs 495,235 req/s · p50 65 µs · p99 1,310 µs
GET / (JSON) 686,738 req/s · p99 189 µs 481,666 req/s · p99 1,297 µs
POST write 643,250 commits/s — every one fsync-acked · p99 177 µs 354,758 req/s — RAM only, no WAL · p99 1,649 µs
10,000 idle conns 0 errors 0 errors

wo-rt-c on 8 cores outpaces Go on 20 with ~8× tighter p99 (Go's GC shows there) — while fsyncing every write Go doesn't. Honest caveats: net/http does full general-purpose HTTP; our parser is minimal; .NET was not installed on the box. ACID under load: three crash rounds (kill -9 mid-bench at ~2M commits) all showed WAL records ≥ acked; isolation probe: 300 concurrent commits → 300 distinct ids; torn-tail records drop whole by CRC.

The crash-under-load test found and fixed two real bugs the lighter phase-D test missed: an ack-before-fsync race (conn_continue armed the send in the same tick the commit was staged) and an fd-reuse ABA hazard in ack parking (fixed with per-connection generation stamps). That is what phase F is for.

  • docs/plan/exploration/c-runtime/02-single-binary.md — the end goal: how the wo build single binary runs on this runtime environment — the runtime kernel is statically linked into every writeonce app (Go model, nothing to install), with the catalog/routes/bytecode payload consumed at boot.

Deliberate simplifications

Single-shot RECV re-armed per request (multishot recv + buffer rings are a phase-F improvement), one outstanding SQE per connection, fixed-size buffers, naive "title" extraction instead of a JSON parser, no timerfd. Requires kernel ≥ 5.19 (multishot accept). This file is for reading; crates/rt is for running writeonce.