docs: performance arc as story iterations (9e measure, 9f io_uring)
- 9e durability/throughput/scale: the measurement backbone -- run the employee program, restart to prove persistence, benchmark read/write through compiled .wo, ~1M-row mixed load with throughput floor + p99 ceiling + flat RSS; the gate every optimization signs (before/after delta required, no measured delta = not accepted) - 9f io_uring group-commit: replace fsync-per-commit with batched io_uring durability overlapped on shard threads; same ack-after- durable contract, crash battery unchanged, automatic fsync fallback on kernels without it; deliberately LAST (needs 8's threads to overlap and 9e's baseline to beat) - wired the existing levers into the arc: iteration 8 (thread-per-core) = "optimize multithreading", 7b (mark-sweep) = "implement GC" -- each now gated by re-running 9e and recording the delta - explicit sequence recorded in 9e: 9b lands -> 9e baseline -> 7b re-bench -> 8 re-bench -> 9f re-bench - roadmap + board rows for 9e/9f; four forks each for the specs (load generator, absolute vs relative budgets, what "1M" means, durable vs RAM headline; ring model, liburing vs raw, batch boundary, fallback testing) Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
parent
77b3b06e77
commit
71d3985e81
6 changed files with 270 additions and 0 deletions
|
|
@ -119,6 +119,8 @@ that sequences its tasks. Read one, approve, then the next starts.
|
|||
| 9b | [`@table`, relations, query](stories/language-runtime-database/09b-table-relations-query.md) | ⬜ spec + plan ready (2026-08-15) |
|
||||
| 9c | [Cross-program tables](stories/language-runtime-database/09c-cross-program-tables.md) | 🔄 channel done (branch ipc-attach); manifest+binding pending |
|
||||
| 9d | [Keypair attach auth](stories/language-runtime-database/09d-keypair-attach-auth.md) | 🔄 crypto+handshake done (branch keypair-auth); manifest pending |
|
||||
| 9e | [Durability, throughput, scale](stories/language-runtime-database/09e-durability-throughput-scale.md) | ⬜ needs a spec first |
|
||||
| 9f | [io_uring group-commit](stories/language-runtime-database/09f-io-uring-commit.md) | ⬜ after 8 + 9e |
|
||||
| 10 | [HTTP service layer](stories/language-runtime-database/10-http-service.md) | ⬜ | Hold |
|
||||
| 11 | [Fibers](stories/language-runtime-database/11-fibers.md) | ⬜ | Hold |
|
||||
| 12 | [Blue-green deploy](stories/language-runtime-database/12-blue-green-deploy.md) | ⬜ | Hold |
|
||||
|
|
@ -327,6 +329,8 @@ Ecommerce sample (verified 2026-06-13): `api.rest` 17/17 expected statuses pass.
|
|||
| 9b | `@table` + relations + language-integrated query — comprehension queries, `ref`/`backlink` navigation, GroupBy aggregates; acceptance: new `docs/examples/employee` sample | [spec](superpowers/specs/2026-08-15-table-relations-query-design.md) · [plan](plan/compiler/2026-08-15-employee-relations-query.md) |
|
||||
| 9c | Cross-program tables — attach to a running program's database (IPC string in wo.toml, manifest-granted rights, owner stays the single writer) | **no spec yet** — four open forks recorded in the iteration; brainstorm before planning |
|
||||
| 9d | Keypair attach auth — mutual challenge–response, grants name public keys, uid superseded | **no spec yet** — four forks recorded; plan folds into 9c's |
|
||||
| 9e | Durability + throughput + scale — restart-persistence, read/write benchmark, ~1M rows; the gate every later optimization re-runs | **no spec yet** — four forks recorded; the measurement backbone |
|
||||
| 9f | io_uring group-commit write path — batched durability overlapped on shard threads, fsync fallback | **no spec yet** — brainstorm after iterations 8 + 9e |
|
||||
| 10 | HTTP service layer | [plan 6](superpowers/plans/2026-08-01-http-service-layer.md) |
|
||||
| 11 | Fibers | vision §3, [blue-green exploration](plan/exploration/blue-green-vm/00-vision.md) |
|
||||
| 12 | Blue-green deploy | [spec](superpowers/specs/2026-08-03-blue-green-vm-design.md) — plan authored after iterations 9–10 |
|
||||
|
|
|
|||
|
|
@ -55,6 +55,8 @@ iterations); no commits by agents — drafts go to `.dev/commit.md`.
|
|||
| 9b | [`@table`, relations, query](09b-table-relations-query.md) | `@table` becomes real storage; typed `ref`/`backlink`/`multi` relations; compiler-checked LINQ-shaped queries lowered to engine ops |
|
||||
| 9c | [Cross-program tables](09c-cross-program-tables.md) | attach to a running program's database over a local IPC channel: manifest-granted read/write rights, typed statements checked against the owner's shapes, owner stays the single writer |
|
||||
| 9d | [Keypair attach auth](09d-keypair-attach-auth.md) | program identity is a keypair: mutual challenge–response at attach, grants name public keys, replay-proof, rotation is a config change |
|
||||
| 9e | [Durability, throughput, scale](09e-durability-throughput-scale.md) | restart-persistence proof, read/write benchmark, ~1M-row load; the measurement gate every optimization signs |
|
||||
| 9f | [io_uring group-commit](09f-io-uring-commit.md) | replace fsync-per-commit with io_uring batched durability, overlapped on the shard threads; fsync fallback kept |
|
||||
| 10 | [HTTP service layer](10-http-service.md) | `service` blocks route to VM methods; REST parity with Stage 2 |
|
||||
| 11 | [Fibers](11-fibers.md) | green threads on the shard scheduler: reduction-budget preemption, park on I/O |
|
||||
| 12 | [Blue-green deploy](12-blue-green-deploy.md) | two VM slots, in-runtime compile, atomic switch, resident rollback |
|
||||
|
|
|
|||
|
|
@ -80,6 +80,11 @@
|
|||
## Info
|
||||
|
||||
- Governing spec: [`docs/superpowers/specs/2026-08-11-inferred-gc-mark-sweep-design.md`](../../superpowers/specs/2026-08-11-inferred-gc-mark-sweep-design.md).
|
||||
- **Gated by the benchmark (2026-08-15):** this is the "implement garbage
|
||||
collection" lever of the performance arc — tri-color mark-sweep replacing
|
||||
RC changes the write path's tail latency, so landing it means re-running
|
||||
iteration [9e](09e-durability-throughput-scale.md) and recording the
|
||||
delta (does tracing help or hurt p99 under write load?).
|
||||
- **Constraint added by the database track (2026-08-15):** a GC-managed value
|
||||
in a `@table` field is a compile error (the engine/heap bulkhead — 9b
|
||||
design, section 6). Once GC-ness is inferred rather than annotated, the
|
||||
|
|
|
|||
|
|
@ -42,6 +42,13 @@
|
|||
loops, eventfd mail, the machinery this iteration lifts into `wovm`.
|
||||
- The VM's object header has carried a shard id since iteration 2 — no
|
||||
relayout.
|
||||
- **Gated by the benchmark (2026-08-15):** this is the "optimize
|
||||
multithreading" lever of the performance arc — thread-per-core is a
|
||||
throughput/scale claim, so landing it means re-running iteration
|
||||
[9e](09e-durability-throughput-scale.md) at the connection/concurrency
|
||||
scale it unlocks and recording the before/after delta. It is also where
|
||||
the io_uring write path ([9f](09f-io-uring-commit.md)) gets a thread to
|
||||
overlap durability against.
|
||||
|
||||
## Proposed Solution
|
||||
|
||||
|
|
|
|||
|
|
@ -0,0 +1,136 @@
|
|||
# Iteration 9e — durability proof, throughput, and scale under load
|
||||
|
||||
> Format: `product/story-iteration-template`. Part of
|
||||
> [Story — one language, one runtime, one database, one binary](00-story.md).
|
||||
>
|
||||
> **Inserted 2026-08-15.** The measurement backbone. Everything after the
|
||||
> functional engine (9/9b) is an *optimization*, and an optimization without
|
||||
> a number is a guess — this iteration is the number. It comes before the
|
||||
> optimization iterations (7b GC, 8 shard-actor, 9f io_uring) reopen for
|
||||
> performance work, because each of those must be gated by re-running THIS
|
||||
> iteration's benchmark and showing the number moved the right way.
|
||||
>
|
||||
> **No spec exists yet.** The forks in *Info* are genuine decisions.
|
||||
|
||||
## Goals
|
||||
|
||||
- **Durability is proven by a restart, not asserted.** The employee program
|
||||
(iteration 9b) runs, writes rows, is stopped and restarted, and every
|
||||
acknowledged write is present after replay — the WAL's promise turned into
|
||||
a scripted acceptance on a real program, not just the unit-level crash
|
||||
battery.
|
||||
- **Read and write throughput are measured, published, and defended.** A
|
||||
repeatable benchmark drives the engine through the language (not the C
|
||||
API): inserts/sec, point-reads/sec, indexed-query/sec, each with p50/p99
|
||||
latency, recorded in the tree so a regression is a diff.
|
||||
- **The scale target is a gate, not a slogan.** "A million users can read and
|
||||
write" becomes a concrete load: a dataset of ~1M rows across the sample's
|
||||
tables, a mixed read/write workload at a stated concurrency, sustained for
|
||||
a stated duration, with throughput and tail latency inside a stated budget
|
||||
and RSS flat (the log-watcher soak discipline, at database scale).
|
||||
- **The benchmark is the contract every later optimization signs.** 7b (GC),
|
||||
8 (shard-actor threads), and 9f (io_uring) each re-run this and record the
|
||||
before/after — no optimization lands without a measured delta.
|
||||
|
||||
## Acceptance Criteria
|
||||
|
||||
- What to achieve?
|
||||
- **Given** the employee program seeded with data and then stopped,
|
||||
- **when** it is restarted and queried,
|
||||
- **then** every acknowledged row is present with its exact contents,
|
||||
the ids continue past the persisted maximum, and a query that used an
|
||||
index before the restart uses it after (the index was rebuilt on
|
||||
replay).
|
||||
- What to achieve?
|
||||
- **Given** the benchmark harness driving inserts, point reads, and
|
||||
indexed queries through compiled `.wo`,
|
||||
- **when** it runs to completion,
|
||||
- **then** it reports ops/sec and p50/p99 for each operation class, writes
|
||||
the numbers to a tracked results file, and fails if any number crosses
|
||||
a recorded regression threshold.
|
||||
- What to achieve?
|
||||
- **Given** ~1M rows and a mixed read/write workload at the target
|
||||
concurrency held for the target duration,
|
||||
- **when** it runs,
|
||||
- **then** throughput stays above the floor, p99 stays under the ceiling,
|
||||
RSS is flat between a warmed baseline and the end (no growth beyond
|
||||
tolerance), zero descriptors leak, and — for a write-inclusive run under
|
||||
a durable configuration — a kill mid-load followed by replay loses no
|
||||
acknowledged write.
|
||||
- What to achieve?
|
||||
- **Given** any later optimization iteration (7b, 8, 9f),
|
||||
- **when** it claims a speedup,
|
||||
- **then** this benchmark's before/after numbers are in that iteration's
|
||||
record, and a claim with no measured delta is not accepted.
|
||||
|
||||
## Out Of Scope
|
||||
|
||||
- **The optimizations themselves.** This iteration MEASURES; 7b/8/9f change.
|
||||
A single-thread RAM-authoritative baseline is a legitimate first number —
|
||||
the point is to have one before anyone tunes.
|
||||
- **Distributed / multi-machine load.** Same-machine, one process (or one
|
||||
process per shard once iteration 8 lands). Cross-host is the network layer's
|
||||
concern, much later.
|
||||
- **Micro-optimizing the benchmark harness.** It must be honest and
|
||||
repeatable, not itself fast; if the harness is the bottleneck the spec says
|
||||
so and fixes that, but a perfect load generator is not the deliverable.
|
||||
- **A cost-based query planner.** Index selection is 9b's; this iteration
|
||||
measures what 9b lowers, it does not make the planner smarter.
|
||||
|
||||
## Info
|
||||
|
||||
Forks the spec must settle:
|
||||
|
||||
**1. What generates the load, and in what language?** The doctrine is "the
|
||||
sample is the test", so the honest generator drives compiled `.wo` — a
|
||||
benchmark mode in the employee program (or a sibling sample) that loops
|
||||
inserts/reads/queries and times them. The alternative — a C harness calling
|
||||
the engine API directly — measures the engine but skips the compiler's
|
||||
lowering, which is exactly the layer a language-integrated query has to pay
|
||||
for. Leaning: `.wo` benchmark mode for the headline numbers (the number that
|
||||
matters is end to end), with the C-API microbench kept only to attribute a
|
||||
regression to engine vs lowering.
|
||||
|
||||
**2. What are the actual budgets?** Throughput floors and latency ceilings
|
||||
have to be numbers, and the first run sets them — but the spec must decide
|
||||
whether the gate is absolute (">= N ops/sec on the reference machine") or
|
||||
relative ("no worse than the last recorded run by more than X%"). Absolute
|
||||
gates rot across machines; relative gates need a committed baseline file.
|
||||
Leaning: relative gates against a tracked `bench/baseline.json`, refreshed
|
||||
deliberately with a commit that says why, plus a loud absolute floor so a
|
||||
catastrophic regression fails even on a slow machine.
|
||||
|
||||
**3. What does "1M users read and write" concretely mean?** A million
|
||||
long-lived idle connections is a different test from a million rows under a
|
||||
churning read/write mix from a bounded connection pool. The sample's shape
|
||||
(departments, employees) suggests rows, not connections, as the scale axis
|
||||
for THIS iteration; the connection-scale test belongs with the shard-actor
|
||||
runtime (iteration 8) and the eventual network layer. Leaning: ~1M rows +
|
||||
a bounded concurrent read/write workload here; connection scale deferred to
|
||||
8 with a cross-reference.
|
||||
|
||||
**4. Durable or RAM-only for the throughput headline?** fsync-per-commit
|
||||
(the current per-statement durability) will dominate write throughput and is
|
||||
the honest number for a durable workload; RAM-only (no `WO_DATA`) measures
|
||||
the engine's ceiling. Both matter and mean different things. Leaning:
|
||||
publish both, labeled — durable is the number an operator plans against, and
|
||||
the gap between them is precisely what iteration 9f (io_uring group-commit)
|
||||
exists to close.
|
||||
|
||||
## Proposed Solution
|
||||
|
||||
- **Brainstorm the spec**, settling the four forks; then a plan whose first
|
||||
task is the harness and the baseline file, because nothing downstream means
|
||||
anything without them.
|
||||
- **Sequence the whole performance arc around this iteration:**
|
||||
1. 9b lands → employee compiles and runs → **9e restart-persistence** and
|
||||
**9e baseline benchmark** (single-thread, both durable and RAM-only).
|
||||
2. **7b** (inferred GC + mark-sweep) → re-run 9e, record the delta (does
|
||||
tracing change the write path's tail latency?).
|
||||
3. **8** (shard-actor, thread-per-core) → re-run 9e at the connection/
|
||||
concurrency scale it unlocks, record the delta.
|
||||
4. **9f** (io_uring group-commit) → re-run 9e's durable write number, record
|
||||
the delta against the fsync-per-commit baseline — the payoff.
|
||||
- The benchmark harness and its baseline live under `bench/` (or the existing
|
||||
`runtime/bench/`), and `just` gets a `db-bench` recipe kept off the fast
|
||||
path, exactly like `log-watcher::soak`.
|
||||
116
docs/stories/language-runtime-database/09f-io-uring-commit.md
Normal file
116
docs/stories/language-runtime-database/09f-io-uring-commit.md
Normal file
|
|
@ -0,0 +1,116 @@
|
|||
# Iteration 9f — io_uring group-commit write path
|
||||
|
||||
> Format: `product/story-iteration-template`. Part of
|
||||
> [Story — one language, one runtime, one database, one binary](00-story.md).
|
||||
>
|
||||
> **Inserted 2026-08-15.** The write-path optimization, and deliberately the
|
||||
> LAST database performance iteration: it only earns its complexity once
|
||||
> there is a measured fsync-per-commit baseline to beat (iteration 9e) and a
|
||||
> multithreaded runtime to overlap against (iteration 8). Doing it earlier
|
||||
> would optimize a number nobody had measured, against a runtime that
|
||||
> couldn't use it.
|
||||
>
|
||||
> **No spec exists yet.** The forks in *Info* are genuine decisions.
|
||||
|
||||
## Goals
|
||||
|
||||
- **Replace fsync-per-commit with io_uring group-commit** on the WAL write
|
||||
path: batch a tick's committed records into one submission, let the kernel
|
||||
overlap the write and the durability barrier, and acknowledge each writer
|
||||
only after the barrier its record rode has completed — the same
|
||||
ack-after-durable contract, at a fraction of the syscall cost.
|
||||
- **Overlap durability with work.** With the shard-actor runtime
|
||||
(iteration 8) the shard thread submits its batch and keeps executing ready
|
||||
statements while the ring drains, instead of blocking one thread on one
|
||||
fdatasync — the multithreading the throughput number has been waiting for.
|
||||
- **Keep the durability promise byte-for-byte.** Every guarantee iterations 9
|
||||
and 9e proved — replay-whole-or-not-at-all, torn-tail drop, no
|
||||
acknowledged write ever lost — holds identically; io_uring changes HOW the
|
||||
bytes reach the platter, never WHETHER an ack means durable.
|
||||
|
||||
## Acceptance Criteria
|
||||
|
||||
- What to achieve?
|
||||
- **Given** the io_uring write path under the iteration-9e crash battery
|
||||
(concurrent writers, kill -9 mid-stream, reboot, replay),
|
||||
- **when** it runs,
|
||||
- **then** every acknowledged write is present after replay and no
|
||||
unacknowledged partial write is ever visible — the exact result the
|
||||
fsync path gives, so durability is provably unchanged.
|
||||
- What to achieve?
|
||||
- **Given** the iteration-9e durable write benchmark,
|
||||
- **when** it is run on the fsync-per-commit path and then the io_uring
|
||||
group-commit path on the same machine,
|
||||
- **then** the io_uring path's write throughput is materially higher and
|
||||
its p99 commit latency lower, with the before/after numbers recorded —
|
||||
the payoff, measured, not asserted.
|
||||
- What to achieve?
|
||||
- **Given** a kernel without io_uring (old, or restricted by seccomp),
|
||||
- **when** the runtime starts,
|
||||
- **then** it falls back to the pwrite + fdatasync path automatically and
|
||||
correctly — io_uring is an accelerator, never a hard dependency, and a
|
||||
binary that runs everywhere is the whole project's premise.
|
||||
|
||||
## Out Of Scope
|
||||
|
||||
- **io_uring for the network/accept path.** This iteration is the WAL write
|
||||
path only; the socket side is the shard-actor runtime's and the network
|
||||
layer's concern.
|
||||
- **io_uring for reads.** RAM is authoritative — reads never touch a
|
||||
descriptor (phase-B doctrine), so there is nothing to accelerate on the
|
||||
read path. This is a write-durability optimization, full stop.
|
||||
- **Registered buffers / fixed files / SQPOLL tuning** beyond what the
|
||||
benchmark shows is worth it. Start with the plain submit/complete model;
|
||||
add ring features only when 9e's number says a specific one pays.
|
||||
- **Replacing the WAL format or the commit contract.** The bytes on disk and
|
||||
the meaning of an ack are iteration 9's; this changes the syscall, not the
|
||||
format.
|
||||
|
||||
## Info
|
||||
|
||||
Forks the spec must settle:
|
||||
|
||||
**1. How much of the ring model, and behind what abstraction?** The write
|
||||
path today is `pwrite` + `fdatasync` in `database/src/wal.c`; io_uring adds a
|
||||
submission/completion queue and a durability barrier op
|
||||
(`IORING_OP_FSYNC`/`IORING_FSYNC_DATASYNC` or `O_DSYNC` writes). The fork:
|
||||
wrap it behind the existing `wo_wal_commit` boundary (drop-in, the engine
|
||||
never learns) or expose an async-commit primitive the shard scheduler drives
|
||||
(faster overlap, but couples the WAL to iteration 8's loop). Leaning:
|
||||
drop-in behind `wo_wal_commit` first — it is the correctness-preserving
|
||||
step and 9e can measure it standalone — then an async variant only if 8's
|
||||
scheduler shows the blocking boundary is the remaining bottleneck.
|
||||
|
||||
**2. liburing or raw syscalls?** liburing is the ergonomic wrapper but is a
|
||||
new external dependency, against the libc-only doctrine; the raw
|
||||
`io_uring_setup`/`io_uring_enter` syscalls are a few hundred lines and keep
|
||||
the doctrine. Leaning: raw syscalls (the doctrine is load-bearing and this is
|
||||
a bounded surface), with the mmap'd ring setup written down in the binding
|
||||
doc the way the WAL format is — normative, versioned.
|
||||
|
||||
**3. What is the batch boundary?** Per-statement commit (today) is the
|
||||
simplest correct thing and the slowest; a group commit needs a boundary — a
|
||||
tick (iteration 8's scheduler quantum), a count, or a short time window.
|
||||
Leaning: the shard tick once iteration 8 lands (a batch is "everything
|
||||
committed this tick"), with a single-writer fallback that batches whatever
|
||||
accumulated between one `wo_wal_commit` call and the ring draining.
|
||||
|
||||
**4. How is the fallback chosen and tested?** A kernel probe at startup
|
||||
(attempt `io_uring_setup`, fall back on ENOSYS/EPERM) is the mechanism; the
|
||||
question is how CI proves BOTH paths without two kernels. Leaning: an
|
||||
environment override (`WO_WAL_MODE=fsync|uring`) so the test matrix runs the
|
||||
crash battery and the benchmark on both on any capable machine, and the
|
||||
auto-probe is what production uses.
|
||||
|
||||
## Proposed Solution
|
||||
|
||||
- **Brainstorm the spec** after iterations 8 and 9e exist — this iteration is
|
||||
meaningless without a multithreaded runtime to overlap against and a
|
||||
measured baseline to beat, and its plan's acceptance is literally "9e's
|
||||
durable number improved, 9e's crash battery still green, fsync fallback
|
||||
still correct".
|
||||
- Expected shape: a `wo_wal` write-mode switch (fsync vs uring), the raw ring
|
||||
setup + submit/complete in `database/src/wal.c` (or a `wal_uring.c`
|
||||
beside it), the startup probe + `WO_WAL_MODE` override, the binding doc's
|
||||
WAL section extended with the ring layout, and iteration 9e re-run on both
|
||||
paths with the delta committed.
|
||||
Loading…
Reference in a new issue