From 5b96a6dee6bef8c8df2d9c6ea0fce69d552256b9 Mon Sep 17 00:00:00 2001 From: "shoney.arickathil" Date: Sat, 15 Aug 2026 22:12:48 +0200 Subject: [PATCH] docs: performance arc as story iterations (9e measure, 9f io_uring) - 9e durability/throughput/scale: the measurement backbone -- run the employee program, restart to prove persistence, benchmark read/write through compiled .wo, ~1M-row mixed load with throughput floor + p99 ceiling + flat RSS; the gate every optimization signs (before/after delta required, no measured delta = not accepted) - 9f io_uring group-commit: replace fsync-per-commit with batched io_uring durability overlapped on shard threads; same ack-after- durable contract, crash battery unchanged, automatic fsync fallback on kernels without it; deliberately LAST (needs 8's threads to overlap and 9e's baseline to beat) - wired the existing levers into the arc: iteration 8 (thread-per-core) = "optimize multithreading", 7b (mark-sweep) = "implement GC" -- each now gated by re-running 9e and recording the delta - explicit sequence recorded in 9e: 9b lands -> 9e baseline -> 7b re-bench -> 8 re-bench -> 9f re-bench - roadmap + board rows for 9e/9f; four forks each for the specs (load generator, absolute vs relative budgets, what "1M" means, durable vs RAM headline; ring model, liburing vs raw, batch boundary, fallback testing) Co-Authored-By: Claude Opus 5 (1M context) --- docs/00-status.md | 4 + .../language-runtime-database/00-story.md | 2 + .../07b-inferred-gc-mark-sweep.md | 5 + .../08-shard-actor-runtime.md | 7 + .../09e-durability-throughput-scale.md | 136 ++++++++++++++++++ .../09f-io-uring-commit.md | 116 +++++++++++++++ 6 files changed, 270 insertions(+) create mode 100644 docs/stories/language-runtime-database/09e-durability-throughput-scale.md create mode 100644 docs/stories/language-runtime-database/09f-io-uring-commit.md diff --git a/docs/00-status.md b/docs/00-status.md index 649bcac..0cfd7ae 100644 --- a/docs/00-status.md +++ b/docs/00-status.md @@ -119,6 +119,8 @@ that sequences its tasks. Read one, approve, then the next starts. | 9b | [`@table`, relations, query](stories/language-runtime-database/09b-table-relations-query.md) | ⬜ spec + plan ready (2026-08-15) | | 9c | [Cross-program tables](stories/language-runtime-database/09c-cross-program-tables.md) | πŸ”„ channel done (branch ipc-attach); manifest+binding pending | | 9d | [Keypair attach auth](stories/language-runtime-database/09d-keypair-attach-auth.md) | πŸ”„ crypto+handshake done (branch keypair-auth); manifest pending | +| 9e | [Durability, throughput, scale](stories/language-runtime-database/09e-durability-throughput-scale.md) | ⬜ needs a spec first | +| 9f | [io_uring group-commit](stories/language-runtime-database/09f-io-uring-commit.md) | ⬜ after 8 + 9e | | 10 | [HTTP service layer](stories/language-runtime-database/10-http-service.md) | ⬜ | Hold | | 11 | [Fibers](stories/language-runtime-database/11-fibers.md) | ⬜ | Hold | | 12 | [Blue-green deploy](stories/language-runtime-database/12-blue-green-deploy.md) | ⬜ | Hold | @@ -327,6 +329,8 @@ Ecommerce sample (verified 2026-06-13): `api.rest` 17/17 expected statuses pass. | 9b | `@table` + relations + language-integrated query β€” comprehension queries, `ref`/`backlink` navigation, GroupBy aggregates; acceptance: new `docs/examples/employee` sample | [spec](superpowers/specs/2026-08-15-table-relations-query-design.md) Β· [plan](plan/compiler/2026-08-15-employee-relations-query.md) | | 9c | Cross-program tables β€” attach to a running program's database (IPC string in wo.toml, manifest-granted rights, owner stays the single writer) | **no spec yet** β€” four open forks recorded in the iteration; brainstorm before planning | | 9d | Keypair attach auth β€” mutual challenge–response, grants name public keys, uid superseded | **no spec yet** β€” four forks recorded; plan folds into 9c's | +| 9e | Durability + throughput + scale β€” restart-persistence, read/write benchmark, ~1M rows; the gate every later optimization re-runs | **no spec yet** β€” four forks recorded; the measurement backbone | +| 9f | io_uring group-commit write path β€” batched durability overlapped on shard threads, fsync fallback | **no spec yet** β€” brainstorm after iterations 8 + 9e | | 10 | HTTP service layer | [plan 6](superpowers/plans/2026-08-01-http-service-layer.md) | | 11 | Fibers | vision Β§3, [blue-green exploration](plan/exploration/blue-green-vm/00-vision.md) | | 12 | Blue-green deploy | [spec](superpowers/specs/2026-08-03-blue-green-vm-design.md) β€” plan authored after iterations 9–10 | diff --git a/docs/stories/language-runtime-database/00-story.md b/docs/stories/language-runtime-database/00-story.md index 4bec861..7d90141 100644 --- a/docs/stories/language-runtime-database/00-story.md +++ b/docs/stories/language-runtime-database/00-story.md @@ -55,6 +55,8 @@ iterations); no commits by agents β€” drafts go to `.dev/commit.md`. | 9b | [`@table`, relations, query](09b-table-relations-query.md) | `@table` becomes real storage; typed `ref`/`backlink`/`multi` relations; compiler-checked LINQ-shaped queries lowered to engine ops | | 9c | [Cross-program tables](09c-cross-program-tables.md) | attach to a running program's database over a local IPC channel: manifest-granted read/write rights, typed statements checked against the owner's shapes, owner stays the single writer | | 9d | [Keypair attach auth](09d-keypair-attach-auth.md) | program identity is a keypair: mutual challenge–response at attach, grants name public keys, replay-proof, rotation is a config change | +| 9e | [Durability, throughput, scale](09e-durability-throughput-scale.md) | restart-persistence proof, read/write benchmark, ~1M-row load; the measurement gate every optimization signs | +| 9f | [io_uring group-commit](09f-io-uring-commit.md) | replace fsync-per-commit with io_uring batched durability, overlapped on the shard threads; fsync fallback kept | | 10 | [HTTP service layer](10-http-service.md) | `service` blocks route to VM methods; REST parity with Stage 2 | | 11 | [Fibers](11-fibers.md) | green threads on the shard scheduler: reduction-budget preemption, park on I/O | | 12 | [Blue-green deploy](12-blue-green-deploy.md) | two VM slots, in-runtime compile, atomic switch, resident rollback | diff --git a/docs/stories/language-runtime-database/07b-inferred-gc-mark-sweep.md b/docs/stories/language-runtime-database/07b-inferred-gc-mark-sweep.md index 9abd261..55f9b61 100644 --- a/docs/stories/language-runtime-database/07b-inferred-gc-mark-sweep.md +++ b/docs/stories/language-runtime-database/07b-inferred-gc-mark-sweep.md @@ -80,6 +80,11 @@ ## Info - Governing spec: [`docs/superpowers/specs/2026-08-11-inferred-gc-mark-sweep-design.md`](../../superpowers/specs/2026-08-11-inferred-gc-mark-sweep-design.md). +- **Gated by the benchmark (2026-08-15):** this is the "implement garbage + collection" lever of the performance arc β€” tri-color mark-sweep replacing + RC changes the write path's tail latency, so landing it means re-running + iteration [9e](09e-durability-throughput-scale.md) and recording the + delta (does tracing help or hurt p99 under write load?). - **Constraint added by the database track (2026-08-15):** a GC-managed value in a `@table` field is a compile error (the engine/heap bulkhead β€” 9b design, section 6). Once GC-ness is inferred rather than annotated, the diff --git a/docs/stories/language-runtime-database/08-shard-actor-runtime.md b/docs/stories/language-runtime-database/08-shard-actor-runtime.md index 38e66f8..c24b93c 100644 --- a/docs/stories/language-runtime-database/08-shard-actor-runtime.md +++ b/docs/stories/language-runtime-database/08-shard-actor-runtime.md @@ -42,6 +42,13 @@ loops, eventfd mail, the machinery this iteration lifts into `wovm`. - The VM's object header has carried a shard id since iteration 2 β€” no relayout. +- **Gated by the benchmark (2026-08-15):** this is the "optimize + multithreading" lever of the performance arc β€” thread-per-core is a + throughput/scale claim, so landing it means re-running iteration + [9e](09e-durability-throughput-scale.md) at the connection/concurrency + scale it unlocks and recording the before/after delta. It is also where + the io_uring write path ([9f](09f-io-uring-commit.md)) gets a thread to + overlap durability against. ## Proposed Solution diff --git a/docs/stories/language-runtime-database/09e-durability-throughput-scale.md b/docs/stories/language-runtime-database/09e-durability-throughput-scale.md new file mode 100644 index 0000000..361813b --- /dev/null +++ b/docs/stories/language-runtime-database/09e-durability-throughput-scale.md @@ -0,0 +1,136 @@ +# Iteration 9e β€” durability proof, throughput, and scale under load + +> Format: fiberloom `product/story-iteration-template`. Part of +> [Story β€” one language, one runtime, one database, one binary](00-story.md). +> +> **Inserted 2026-08-15.** The measurement backbone. Everything after the +> functional engine (9/9b) is an *optimization*, and an optimization without +> a number is a guess β€” this iteration is the number. It comes before the +> optimization iterations (7b GC, 8 shard-actor, 9f io_uring) reopen for +> performance work, because each of those must be gated by re-running THIS +> iteration's benchmark and showing the number moved the right way. +> +> **No spec exists yet.** The forks in *Info* are genuine decisions. + +## Goals + +- **Durability is proven by a restart, not asserted.** The employee program + (iteration 9b) runs, writes rows, is stopped and restarted, and every + acknowledged write is present after replay β€” the WAL's promise turned into + a scripted acceptance on a real program, not just the unit-level crash + battery. +- **Read and write throughput are measured, published, and defended.** A + repeatable benchmark drives the engine through the language (not the C + API): inserts/sec, point-reads/sec, indexed-query/sec, each with p50/p99 + latency, recorded in the tree so a regression is a diff. +- **The scale target is a gate, not a slogan.** "A million users can read and + write" becomes a concrete load: a dataset of ~1M rows across the sample's + tables, a mixed read/write workload at a stated concurrency, sustained for + a stated duration, with throughput and tail latency inside a stated budget + and RSS flat (the log-watcher soak discipline, at database scale). +- **The benchmark is the contract every later optimization signs.** 7b (GC), + 8 (shard-actor threads), and 9f (io_uring) each re-run this and record the + before/after β€” no optimization lands without a measured delta. + +## Acceptance Criteria + +- What to achieve? + - **Given** the employee program seeded with data and then stopped, + - **when** it is restarted and queried, + - **then** every acknowledged row is present with its exact contents, + the ids continue past the persisted maximum, and a query that used an + index before the restart uses it after (the index was rebuilt on + replay). +- What to achieve? + - **Given** the benchmark harness driving inserts, point reads, and + indexed queries through compiled `.wo`, + - **when** it runs to completion, + - **then** it reports ops/sec and p50/p99 for each operation class, writes + the numbers to a tracked results file, and fails if any number crosses + a recorded regression threshold. +- What to achieve? + - **Given** ~1M rows and a mixed read/write workload at the target + concurrency held for the target duration, + - **when** it runs, + - **then** throughput stays above the floor, p99 stays under the ceiling, + RSS is flat between a warmed baseline and the end (no growth beyond + tolerance), zero descriptors leak, and β€” for a write-inclusive run under + a durable configuration β€” a kill mid-load followed by replay loses no + acknowledged write. +- What to achieve? + - **Given** any later optimization iteration (7b, 8, 9f), + - **when** it claims a speedup, + - **then** this benchmark's before/after numbers are in that iteration's + record, and a claim with no measured delta is not accepted. + +## Out Of Scope + +- **The optimizations themselves.** This iteration MEASURES; 7b/8/9f change. + A single-thread RAM-authoritative baseline is a legitimate first number β€” + the point is to have one before anyone tunes. +- **Distributed / multi-machine load.** Same-machine, one process (or one + process per shard once iteration 8 lands). Cross-host is the network layer's + concern, much later. +- **Micro-optimizing the benchmark harness.** It must be honest and + repeatable, not itself fast; if the harness is the bottleneck the spec says + so and fixes that, but a perfect load generator is not the deliverable. +- **A cost-based query planner.** Index selection is 9b's; this iteration + measures what 9b lowers, it does not make the planner smarter. + +## Info + +Forks the spec must settle: + +**1. What generates the load, and in what language?** The doctrine is "the +sample is the test", so the honest generator drives compiled `.wo` β€” a +benchmark mode in the employee program (or a sibling sample) that loops +inserts/reads/queries and times them. The alternative β€” a C harness calling +the engine API directly β€” measures the engine but skips the compiler's +lowering, which is exactly the layer a language-integrated query has to pay +for. Leaning: `.wo` benchmark mode for the headline numbers (the number that +matters is end to end), with the C-API microbench kept only to attribute a +regression to engine vs lowering. + +**2. What are the actual budgets?** Throughput floors and latency ceilings +have to be numbers, and the first run sets them β€” but the spec must decide +whether the gate is absolute (">= N ops/sec on the reference machine") or +relative ("no worse than the last recorded run by more than X%"). Absolute +gates rot across machines; relative gates need a committed baseline file. +Leaning: relative gates against a tracked `bench/baseline.json`, refreshed +deliberately with a commit that says why, plus a loud absolute floor so a +catastrophic regression fails even on a slow machine. + +**3. What does "1M users read and write" concretely mean?** A million +long-lived idle connections is a different test from a million rows under a +churning read/write mix from a bounded connection pool. The sample's shape +(departments, employees) suggests rows, not connections, as the scale axis +for THIS iteration; the connection-scale test belongs with the shard-actor +runtime (iteration 8) and the eventual network layer. Leaning: ~1M rows + +a bounded concurrent read/write workload here; connection scale deferred to +8 with a cross-reference. + +**4. Durable or RAM-only for the throughput headline?** fsync-per-commit +(the current per-statement durability) will dominate write throughput and is +the honest number for a durable workload; RAM-only (no `WO_DATA`) measures +the engine's ceiling. Both matter and mean different things. Leaning: +publish both, labeled β€” durable is the number an operator plans against, and +the gap between them is precisely what iteration 9f (io_uring group-commit) +exists to close. + +## Proposed Solution + +- **Brainstorm the spec**, settling the four forks; then a plan whose first + task is the harness and the baseline file, because nothing downstream means + anything without them. +- **Sequence the whole performance arc around this iteration:** + 1. 9b lands β†’ employee compiles and runs β†’ **9e restart-persistence** and + **9e baseline benchmark** (single-thread, both durable and RAM-only). + 2. **7b** (inferred GC + mark-sweep) β†’ re-run 9e, record the delta (does + tracing change the write path's tail latency?). + 3. **8** (shard-actor, thread-per-core) β†’ re-run 9e at the connection/ + concurrency scale it unlocks, record the delta. + 4. **9f** (io_uring group-commit) β†’ re-run 9e's durable write number, record + the delta against the fsync-per-commit baseline β€” the payoff. +- The benchmark harness and its baseline live under `bench/` (or the existing + `runtime/bench/`), and `just` gets a `db-bench` recipe kept off the fast + path, exactly like `log-watcher::soak`. diff --git a/docs/stories/language-runtime-database/09f-io-uring-commit.md b/docs/stories/language-runtime-database/09f-io-uring-commit.md new file mode 100644 index 0000000..b519768 --- /dev/null +++ b/docs/stories/language-runtime-database/09f-io-uring-commit.md @@ -0,0 +1,116 @@ +# Iteration 9f β€” io_uring group-commit write path + +> Format: fiberloom `product/story-iteration-template`. Part of +> [Story β€” one language, one runtime, one database, one binary](00-story.md). +> +> **Inserted 2026-08-15.** The write-path optimization, and deliberately the +> LAST database performance iteration: it only earns its complexity once +> there is a measured fsync-per-commit baseline to beat (iteration 9e) and a +> multithreaded runtime to overlap against (iteration 8). Doing it earlier +> would optimize a number nobody had measured, against a runtime that +> couldn't use it. +> +> **No spec exists yet.** The forks in *Info* are genuine decisions. + +## Goals + +- **Replace fsync-per-commit with io_uring group-commit** on the WAL write + path: batch a tick's committed records into one submission, let the kernel + overlap the write and the durability barrier, and acknowledge each writer + only after the barrier its record rode has completed β€” the same + ack-after-durable contract, at a fraction of the syscall cost. +- **Overlap durability with work.** With the shard-actor runtime + (iteration 8) the shard thread submits its batch and keeps executing ready + statements while the ring drains, instead of blocking one thread on one + fdatasync β€” the multithreading the throughput number has been waiting for. +- **Keep the durability promise byte-for-byte.** Every guarantee iterations 9 + and 9e proved β€” replay-whole-or-not-at-all, torn-tail drop, no + acknowledged write ever lost β€” holds identically; io_uring changes HOW the + bytes reach the platter, never WHETHER an ack means durable. + +## Acceptance Criteria + +- What to achieve? + - **Given** the io_uring write path under the iteration-9e crash battery + (concurrent writers, kill -9 mid-stream, reboot, replay), + - **when** it runs, + - **then** every acknowledged write is present after replay and no + unacknowledged partial write is ever visible β€” the exact result the + fsync path gives, so durability is provably unchanged. +- What to achieve? + - **Given** the iteration-9e durable write benchmark, + - **when** it is run on the fsync-per-commit path and then the io_uring + group-commit path on the same machine, + - **then** the io_uring path's write throughput is materially higher and + its p99 commit latency lower, with the before/after numbers recorded β€” + the payoff, measured, not asserted. +- What to achieve? + - **Given** a kernel without io_uring (old, or restricted by seccomp), + - **when** the runtime starts, + - **then** it falls back to the pwrite + fdatasync path automatically and + correctly β€” io_uring is an accelerator, never a hard dependency, and a + binary that runs everywhere is the whole project's premise. + +## Out Of Scope + +- **io_uring for the network/accept path.** This iteration is the WAL write + path only; the socket side is the shard-actor runtime's and the network + layer's concern. +- **io_uring for reads.** RAM is authoritative β€” reads never touch a + descriptor (phase-B doctrine), so there is nothing to accelerate on the + read path. This is a write-durability optimization, full stop. +- **Registered buffers / fixed files / SQPOLL tuning** beyond what the + benchmark shows is worth it. Start with the plain submit/complete model; + add ring features only when 9e's number says a specific one pays. +- **Replacing the WAL format or the commit contract.** The bytes on disk and + the meaning of an ack are iteration 9's; this changes the syscall, not the + format. + +## Info + +Forks the spec must settle: + +**1. How much of the ring model, and behind what abstraction?** The write +path today is `pwrite` + `fdatasync` in `database/src/wal.c`; io_uring adds a +submission/completion queue and a durability barrier op +(`IORING_OP_FSYNC`/`IORING_FSYNC_DATASYNC` or `O_DSYNC` writes). The fork: +wrap it behind the existing `wo_wal_commit` boundary (drop-in, the engine +never learns) or expose an async-commit primitive the shard scheduler drives +(faster overlap, but couples the WAL to iteration 8's loop). Leaning: +drop-in behind `wo_wal_commit` first β€” it is the correctness-preserving +step and 9e can measure it standalone β€” then an async variant only if 8's +scheduler shows the blocking boundary is the remaining bottleneck. + +**2. liburing or raw syscalls?** liburing is the ergonomic wrapper but is a +new external dependency, against the libc-only doctrine; the raw +`io_uring_setup`/`io_uring_enter` syscalls are a few hundred lines and keep +the doctrine. Leaning: raw syscalls (the doctrine is load-bearing and this is +a bounded surface), with the mmap'd ring setup written down in the binding +doc the way the WAL format is β€” normative, versioned. + +**3. What is the batch boundary?** Per-statement commit (today) is the +simplest correct thing and the slowest; a group commit needs a boundary β€” a +tick (iteration 8's scheduler quantum), a count, or a short time window. +Leaning: the shard tick once iteration 8 lands (a batch is "everything +committed this tick"), with a single-writer fallback that batches whatever +accumulated between one `wo_wal_commit` call and the ring draining. + +**4. How is the fallback chosen and tested?** A kernel probe at startup +(attempt `io_uring_setup`, fall back on ENOSYS/EPERM) is the mechanism; the +question is how CI proves BOTH paths without two kernels. Leaning: an +environment override (`WO_WAL_MODE=fsync|uring`) so the test matrix runs the +crash battery and the benchmark on both on any capable machine, and the +auto-probe is what production uses. + +## Proposed Solution + +- **Brainstorm the spec** after iterations 8 and 9e exist β€” this iteration is + meaningless without a multithreaded runtime to overlap against and a + measured baseline to beat, and its plan's acceptance is literally "9e's + durable number improved, 9e's crash battery still green, fsync fallback + still correct". +- Expected shape: a `wo_wal` write-mode switch (fsync vs uring), the raw ring + setup + submit/complete in `database/src/wal.c` (or a `wal_uring.c` + beside it), the startup probe + `WO_WAL_MODE` override, the binding doc's + WAL section extended with the ring layout, and iteration 9e re-run on both + paths with the delta committed.