docs: performance arc as story iterations (9e measure, 9f io_uring)

- 9e durability/throughput/scale: the measurement backbone -- run the
  employee program, restart to prove persistence, benchmark read/write
  through compiled .wo, ~1M-row mixed load with throughput floor + p99
  ceiling + flat RSS; the gate every optimization signs (before/after
  delta required, no measured delta = not accepted)
- 9f io_uring group-commit: replace fsync-per-commit with batched
  io_uring durability overlapped on shard threads; same ack-after-
  durable contract, crash battery unchanged, automatic fsync fallback
  on kernels without it; deliberately LAST (needs 8's threads to
  overlap and 9e's baseline to beat)
- wired the existing levers into the arc: iteration 8 (thread-per-core)
  = "optimize multithreading", 7b (mark-sweep) = "implement GC" --
  each now gated by re-running 9e and recording the delta
- explicit sequence recorded in 9e: 9b lands -> 9e baseline -> 7b
  re-bench -> 8 re-bench -> 9f re-bench
- roadmap + board rows for 9e/9f; four forks each for the specs
  (load generator, absolute vs relative budgets, what "1M" means,
  durable vs RAM headline; ring model, liburing vs raw, batch
  boundary, fallback testing)

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
shoney.arickathil 2026-08-15 22:12:48 +02:00
parent 5b96279254
commit 5b96a6dee6
6 changed files with 270 additions and 0 deletions

View file

@ -119,6 +119,8 @@ that sequences its tasks. Read one, approve, then the next starts.
| 9b | [`@table`, relations, query](stories/language-runtime-database/09b-table-relations-query.md) | ⬜ spec + plan ready (2026-08-15) |
| 9c | [Cross-program tables](stories/language-runtime-database/09c-cross-program-tables.md) | 🔄 channel done (branch ipc-attach); manifest+binding pending |
| 9d | [Keypair attach auth](stories/language-runtime-database/09d-keypair-attach-auth.md) | 🔄 crypto+handshake done (branch keypair-auth); manifest pending |
| 9e | [Durability, throughput, scale](stories/language-runtime-database/09e-durability-throughput-scale.md) | ⬜ needs a spec first |
| 9f | [io_uring group-commit](stories/language-runtime-database/09f-io-uring-commit.md) | ⬜ after 8 + 9e |
| 10 | [HTTP service layer](stories/language-runtime-database/10-http-service.md) | ⬜ | Hold |
| 11 | [Fibers](stories/language-runtime-database/11-fibers.md) | ⬜ | Hold |
| 12 | [Blue-green deploy](stories/language-runtime-database/12-blue-green-deploy.md) | ⬜ | Hold |
@ -327,6 +329,8 @@ Ecommerce sample (verified 2026-06-13): `api.rest` 17/17 expected statuses pass.
| 9b | `@table` + relations + language-integrated query — comprehension queries, `ref`/`backlink` navigation, GroupBy aggregates; acceptance: new `docs/examples/employee` sample | [spec](superpowers/specs/2026-08-15-table-relations-query-design.md) · [plan](plan/compiler/2026-08-15-employee-relations-query.md) |
| 9c | Cross-program tables — attach to a running program's database (IPC string in wo.toml, manifest-granted rights, owner stays the single writer) | **no spec yet** — four open forks recorded in the iteration; brainstorm before planning |
| 9d | Keypair attach auth — mutual challenge–response, grants name public keys, uid superseded | **no spec yet** — four forks recorded; plan folds into 9c's |
| 9e | Durability + throughput + scale — restart-persistence, read/write benchmark, ~1M rows; the gate every later optimization re-runs | **no spec yet** — four forks recorded; the measurement backbone |
| 9f | io_uring group-commit write path — batched durability overlapped on shard threads, fsync fallback | **no spec yet** — brainstorm after iterations 8 + 9e |
| 10 | HTTP service layer | [plan 6](superpowers/plans/2026-08-01-http-service-layer.md) |
| 11 | Fibers | vision §3, [blue-green exploration](plan/exploration/blue-green-vm/00-vision.md) |
| 12 | Blue-green deploy | [spec](superpowers/specs/2026-08-03-blue-green-vm-design.md) — plan authored after iterations 9–10 |

View file

@ -55,6 +55,8 @@ iterations); no commits by agents — drafts go to `.dev/commit.md`.
| 9b | [`@table`, relations, query](09b-table-relations-query.md) | `@table` becomes real storage; typed `ref`/`backlink`/`multi` relations; compiler-checked LINQ-shaped queries lowered to engine ops |
| 9c | [Cross-program tables](09c-cross-program-tables.md) | attach to a running program's database over a local IPC channel: manifest-granted read/write rights, typed statements checked against the owner's shapes, owner stays the single writer |
| 9d | [Keypair attach auth](09d-keypair-attach-auth.md) | program identity is a keypair: mutual challenge–response at attach, grants name public keys, replay-proof, rotation is a config change |
| 9e | [Durability, throughput, scale](09e-durability-throughput-scale.md) | restart-persistence proof, read/write benchmark, ~1M-row load; the measurement gate every optimization signs |
| 9f | [io_uring group-commit](09f-io-uring-commit.md) | replace fsync-per-commit with io_uring batched durability, overlapped on the shard threads; fsync fallback kept |
| 10 | [HTTP service layer](10-http-service.md) | `service` blocks route to VM methods; REST parity with Stage 2 |
| 11 | [Fibers](11-fibers.md) | green threads on the shard scheduler: reduction-budget preemption, park on I/O |
| 12 | [Blue-green deploy](12-blue-green-deploy.md) | two VM slots, in-runtime compile, atomic switch, resident rollback |

View file

@ -80,6 +80,11 @@
## Info
- Governing spec: [`docs/superpowers/specs/2026-08-11-inferred-gc-mark-sweep-design.md`](../../superpowers/specs/2026-08-11-inferred-gc-mark-sweep-design.md).
- **Gated by the benchmark (2026-08-15):** this is the "implement garbage
collection" lever of the performance arc — tri-color mark-sweep replacing
RC changes the write path's tail latency, so landing it means re-running
iteration [9e](09e-durability-throughput-scale.md) and recording the
delta (does tracing help or hurt p99 under write load?).
- **Constraint added by the database track (2026-08-15):** a GC-managed value
in a `@table` field is a compile error (the engine/heap bulkhead — 9b
design, section 6). Once GC-ness is inferred rather than annotated, the

View file

@ -42,6 +42,13 @@
loops, eventfd mail, the machinery this iteration lifts into `wovm`.
- The VM's object header has carried a shard id since iteration 2 — no
relayout.
- **Gated by the benchmark (2026-08-15):** this is the "optimize
multithreading" lever of the performance arc — thread-per-core is a
throughput/scale claim, so landing it means re-running iteration
[9e](09e-durability-throughput-scale.md) at the connection/concurrency
scale it unlocks and recording the before/after delta. It is also where
the io_uring write path ([9f](09f-io-uring-commit.md)) gets a thread to
overlap durability against.
## Proposed Solution

View file

@ -0,0 +1,136 @@
# Iteration 9e — durability proof, throughput, and scale under load
> Format: fiberloom `product/story-iteration-template`. Part of
> [Story — one language, one runtime, one database, one binary](00-story.md).
>
> **Inserted 2026-08-15.** The measurement backbone. Everything after the
> functional engine (9/9b) is an *optimization*, and an optimization without
> a number is a guess — this iteration is the number. It comes before the
> optimization iterations (7b GC, 8 shard-actor, 9f io_uring) reopen for
> performance work, because each of those must be gated by re-running THIS
> iteration's benchmark and showing the number moved the right way.
>
> **No spec exists yet.** The forks in *Info* are genuine decisions.
## Goals
- **Durability is proven by a restart, not asserted.** The employee program
(iteration 9b) runs, writes rows, is stopped and restarted, and every
acknowledged write is present after replay — the WAL's promise turned into
a scripted acceptance on a real program, not just the unit-level crash
battery.
- **Read and write throughput are measured, published, and defended.** A
repeatable benchmark drives the engine through the language (not the C
API): inserts/sec, point-reads/sec, indexed-query/sec, each with p50/p99
latency, recorded in the tree so a regression is a diff.
- **The scale target is a gate, not a slogan.** "A million users can read and
write" becomes a concrete load: a dataset of ~1M rows across the sample's
tables, a mixed read/write workload at a stated concurrency, sustained for
a stated duration, with throughput and tail latency inside a stated budget
and RSS flat (the log-watcher soak discipline, at database scale).
- **The benchmark is the contract every later optimization signs.** 7b (GC),
8 (shard-actor threads), and 9f (io_uring) each re-run this and record the
before/after — no optimization lands without a measured delta.
## Acceptance Criteria
- What to achieve?
- **Given** the employee program seeded with data and then stopped,
- **when** it is restarted and queried,
- **then** every acknowledged row is present with its exact contents,
the ids continue past the persisted maximum, and a query that used an
index before the restart uses it after (the index was rebuilt on
replay).
- What to achieve?
- **Given** the benchmark harness driving inserts, point reads, and
indexed queries through compiled `.wo`,
- **when** it runs to completion,
- **then** it reports ops/sec and p50/p99 for each operation class, writes
the numbers to a tracked results file, and fails if any number crosses
a recorded regression threshold.
- What to achieve?
- **Given** ~1M rows and a mixed read/write workload at the target
concurrency held for the target duration,
- **when** it runs,
- **then** throughput stays above the floor, p99 stays under the ceiling,
RSS is flat between a warmed baseline and the end (no growth beyond
tolerance), zero descriptors leak, and — for a write-inclusive run under
a durable configuration — a kill mid-load followed by replay loses no
acknowledged write.
- What to achieve?
- **Given** any later optimization iteration (7b, 8, 9f),
- **when** it claims a speedup,
- **then** this benchmark's before/after numbers are in that iteration's
record, and a claim with no measured delta is not accepted.
## Out Of Scope
- **The optimizations themselves.** This iteration MEASURES; 7b/8/9f change.
A single-thread RAM-authoritative baseline is a legitimate first number —
the point is to have one before anyone tunes.
- **Distributed / multi-machine load.** Same-machine, one process (or one
process per shard once iteration 8 lands). Cross-host is the network layer's
concern, much later.
- **Micro-optimizing the benchmark harness.** It must be honest and
repeatable, not itself fast; if the harness is the bottleneck the spec says
so and fixes that, but a perfect load generator is not the deliverable.
- **A cost-based query planner.** Index selection is 9b's; this iteration
measures what 9b lowers, it does not make the planner smarter.
## Info
Forks the spec must settle:
**1. What generates the load, and in what language?** The doctrine is "the
sample is the test", so the honest generator drives compiled `.wo` — a
benchmark mode in the employee program (or a sibling sample) that loops
inserts/reads/queries and times them. The alternative — a C harness calling
the engine API directly — measures the engine but skips the compiler's
lowering, which is exactly the layer a language-integrated query has to pay
for. Leaning: `.wo` benchmark mode for the headline numbers (the number that
matters is end to end), with the C-API microbench kept only to attribute a
regression to engine vs lowering.
**2. What are the actual budgets?** Throughput floors and latency ceilings
have to be numbers, and the first run sets them — but the spec must decide
whether the gate is absolute (">= N ops/sec on the reference machine") or
relative ("no worse than the last recorded run by more than X%"). Absolute
gates rot across machines; relative gates need a committed baseline file.
Leaning: relative gates against a tracked `bench/baseline.json`, refreshed
deliberately with a commit that says why, plus a loud absolute floor so a
catastrophic regression fails even on a slow machine.
**3. What does "1M users read and write" concretely mean?** A million
long-lived idle connections is a different test from a million rows under a
churning read/write mix from a bounded connection pool. The sample's shape
(departments, employees) suggests rows, not connections, as the scale axis
for THIS iteration; the connection-scale test belongs with the shard-actor
runtime (iteration 8) and the eventual network layer. Leaning: ~1M rows +
a bounded concurrent read/write workload here; connection scale deferred to
8 with a cross-reference.
**4. Durable or RAM-only for the throughput headline?** fsync-per-commit
(the current per-statement durability) will dominate write throughput and is
the honest number for a durable workload; RAM-only (no `WO_DATA`) measures
the engine's ceiling. Both matter and mean different things. Leaning:
publish both, labeled — durable is the number an operator plans against, and
the gap between them is precisely what iteration 9f (io_uring group-commit)
exists to close.
## Proposed Solution
- **Brainstorm the spec**, settling the four forks; then a plan whose first
task is the harness and the baseline file, because nothing downstream means
anything without them.
- **Sequence the whole performance arc around this iteration:**
1. 9b lands → employee compiles and runs → **9e restart-persistence** and
**9e baseline benchmark** (single-thread, both durable and RAM-only).
2. **7b** (inferred GC + mark-sweep) → re-run 9e, record the delta (does
tracing change the write path's tail latency?).
3. **8** (shard-actor, thread-per-core) → re-run 9e at the connection/
concurrency scale it unlocks, record the delta.
4. **9f** (io_uring group-commit) → re-run 9e's durable write number, record
the delta against the fsync-per-commit baseline — the payoff.
- The benchmark harness and its baseline live under `bench/` (or the existing
`runtime/bench/`), and `just` gets a `db-bench` recipe kept off the fast
path, exactly like `log-watcher::soak`.

View file

@ -0,0 +1,116 @@
# Iteration 9f — io_uring group-commit write path
> Format: fiberloom `product/story-iteration-template`. Part of
> [Story — one language, one runtime, one database, one binary](00-story.md).
>
> **Inserted 2026-08-15.** The write-path optimization, and deliberately the
> LAST database performance iteration: it only earns its complexity once
> there is a measured fsync-per-commit baseline to beat (iteration 9e) and a
> multithreaded runtime to overlap against (iteration 8). Doing it earlier
> would optimize a number nobody had measured, against a runtime that
> couldn't use it.
>
> **No spec exists yet.** The forks in *Info* are genuine decisions.
## Goals
- **Replace fsync-per-commit with io_uring group-commit** on the WAL write
path: batch a tick's committed records into one submission, let the kernel
overlap the write and the durability barrier, and acknowledge each writer
only after the barrier its record rode has completed — the same
ack-after-durable contract, at a fraction of the syscall cost.
- **Overlap durability with work.** With the shard-actor runtime
(iteration 8) the shard thread submits its batch and keeps executing ready
statements while the ring drains, instead of blocking one thread on one
fdatasync — the multithreading the throughput number has been waiting for.
- **Keep the durability promise byte-for-byte.** Every guarantee iterations 9
and 9e proved — replay-whole-or-not-at-all, torn-tail drop, no
acknowledged write ever lost — holds identically; io_uring changes HOW the
bytes reach the platter, never WHETHER an ack means durable.
## Acceptance Criteria
- What to achieve?
- **Given** the io_uring write path under the iteration-9e crash battery
(concurrent writers, kill -9 mid-stream, reboot, replay),
- **when** it runs,
- **then** every acknowledged write is present after replay and no
unacknowledged partial write is ever visible — the exact result the
fsync path gives, so durability is provably unchanged.
- What to achieve?
- **Given** the iteration-9e durable write benchmark,
- **when** it is run on the fsync-per-commit path and then the io_uring
group-commit path on the same machine,
- **then** the io_uring path's write throughput is materially higher and
its p99 commit latency lower, with the before/after numbers recorded —
the payoff, measured, not asserted.
- What to achieve?
- **Given** a kernel without io_uring (old, or restricted by seccomp),
- **when** the runtime starts,
- **then** it falls back to the pwrite + fdatasync path automatically and
correctly — io_uring is an accelerator, never a hard dependency, and a
binary that runs everywhere is the whole project's premise.
## Out Of Scope
- **io_uring for the network/accept path.** This iteration is the WAL write
path only; the socket side is the shard-actor runtime's and the network
layer's concern.
- **io_uring for reads.** RAM is authoritative — reads never touch a
descriptor (phase-B doctrine), so there is nothing to accelerate on the
read path. This is a write-durability optimization, full stop.
- **Registered buffers / fixed files / SQPOLL tuning** beyond what the
benchmark shows is worth it. Start with the plain submit/complete model;
add ring features only when 9e's number says a specific one pays.
- **Replacing the WAL format or the commit contract.** The bytes on disk and
the meaning of an ack are iteration 9's; this changes the syscall, not the
format.
## Info
Forks the spec must settle:
**1. How much of the ring model, and behind what abstraction?** The write
path today is `pwrite` + `fdatasync` in `database/src/wal.c`; io_uring adds a
submission/completion queue and a durability barrier op
(`IORING_OP_FSYNC`/`IORING_FSYNC_DATASYNC` or `O_DSYNC` writes). The fork:
wrap it behind the existing `wo_wal_commit` boundary (drop-in, the engine
never learns) or expose an async-commit primitive the shard scheduler drives
(faster overlap, but couples the WAL to iteration 8's loop). Leaning:
drop-in behind `wo_wal_commit` first — it is the correctness-preserving
step and 9e can measure it standalone — then an async variant only if 8's
scheduler shows the blocking boundary is the remaining bottleneck.
**2. liburing or raw syscalls?** liburing is the ergonomic wrapper but is a
new external dependency, against the libc-only doctrine; the raw
`io_uring_setup`/`io_uring_enter` syscalls are a few hundred lines and keep
the doctrine. Leaning: raw syscalls (the doctrine is load-bearing and this is
a bounded surface), with the mmap'd ring setup written down in the binding
doc the way the WAL format is — normative, versioned.
**3. What is the batch boundary?** Per-statement commit (today) is the
simplest correct thing and the slowest; a group commit needs a boundary — a
tick (iteration 8's scheduler quantum), a count, or a short time window.
Leaning: the shard tick once iteration 8 lands (a batch is "everything
committed this tick"), with a single-writer fallback that batches whatever
accumulated between one `wo_wal_commit` call and the ring draining.
**4. How is the fallback chosen and tested?** A kernel probe at startup
(attempt `io_uring_setup`, fall back on ENOSYS/EPERM) is the mechanism; the
question is how CI proves BOTH paths without two kernels. Leaning: an
environment override (`WO_WAL_MODE=fsync|uring`) so the test matrix runs the
crash battery and the benchmark on both on any capable machine, and the
auto-probe is what production uses.
## Proposed Solution
- **Brainstorm the spec** after iterations 8 and 9e exist — this iteration is
meaningless without a multithreaded runtime to overlap against and a
measured baseline to beat, and its plan's acceptance is literally "9e's
durable number improved, 9e's crash battery still green, fsync fallback
still correct".
- Expected shape: a `wo_wal` write-mode switch (fsync vs uring), the raw ring
setup + submit/complete in `database/src/wal.c` (or a `wal_uring.c`
beside it), the startup probe + `WO_WAL_MODE` override, the binding doc's
WAL section extended with the ring layout, and iteration 9e re-run on both
paths with the delta committed.