writeonce/docs/stories/language-runtime-database/refine/22-durability-throughput-scale.md
shoney.arickathil a798cf6698 docs: pending iterations renumbered by dependency + priority
- developer directive: pending iteration IDs now ARE the priority order;
  LANDED iterations keep historical numbers (code comments and commit
  history cite them — records, not a queue); 8/11 (the half-landed
  arc), 17 (parked, artifacts on a branch), 18 (next, artifacts named)
  also frozen
- mapping (recorded in 00-story): 19<-20 Float+Bytes, 20<-9c attach,
  21<-9d keypair, 22<-9e benchmarks, 23<-9f io_uring WAL, 24<-19 chat,
  25<-10 services, 26<-12 blue-green, 27<-9g query corpus,
  28<-14 skillhost, 29<-13 metaprogramming
- 11 story files renamed; every doc reference re-numbered (word-boundary
  sweep for the lettered 9x ids, phrase-level for numeric ones); the
  iterations table rewritten with Seq == priority and "(was N)" notes;
  story-scoped link check: zero broken
- merge-recovery folded in: the partial master merge had dropped the
  chat story, the fibers exploration note, the arc spec+plan, the
  framework-v2 plan, and the iteration-17 spec+plan — all restored from
  their branches and renumbered consistently

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-20 14:31:09 +02:00

136 lines
7.3 KiB
Markdown

# Iteration 22 — durability proof, throughput, and scale under load
> Format: fiberloom `product/story-iteration-template`. Part of
> [Story — one language, one runtime, one database, one binary](../00-story.md).
>
> **Inserted 2026-08-15.** The measurement backbone. Everything after the
> functional engine (9/9b) is an *optimization*, and an optimization without
> a number is a guess — this iteration is the number. It comes before the
> optimization iterations (7b GC, 8 shard-actor, 23 io_uring) reopen for
> performance work, because each of those must be gated by re-running THIS
> iteration's benchmark and showing the number moved the right way.
>
> **No spec exists yet.** The forks in *Info* are genuine decisions.
## Goals
- **Durability is proven by a restart, not asserted.** The employee program
(iteration 9b) runs, writes rows, is stopped and restarted, and every
acknowledged write is present after replay — the WAL's promise turned into
a scripted acceptance on a real program, not just the unit-level crash
battery.
- **Read and write throughput are measured, published, and defended.** A
repeatable benchmark drives the engine through the language (not the C
API): inserts/sec, point-reads/sec, indexed-query/sec, each with p50/p99
latency, recorded in the tree so a regression is a diff.
- **The scale target is a gate, not a slogan.** "A million users can read and
write" becomes a concrete load: a dataset of ~1M rows across the sample's
tables, a mixed read/write workload at a stated concurrency, sustained for
a stated duration, with throughput and tail latency inside a stated budget
and RSS flat (the log-watcher soak discipline, at database scale).
- **The benchmark is the contract every later optimization signs.** 7b (GC),
8 (shard-actor threads), and 23 (io_uring) each re-run this and record the
before/after — no optimization lands without a measured delta.
## Acceptance Criteria
- What to achieve?
- **Given** the employee program seeded with data and then stopped,
- **when** it is restarted and queried,
- **then** every acknowledged row is present with its exact contents,
the ids continue past the persisted maximum, and a query that used an
index before the restart uses it after (the index was rebuilt on
replay).
- What to achieve?
- **Given** the benchmark harness driving inserts, point reads, and
indexed queries through compiled `.wo`,
- **when** it runs to completion,
- **then** it reports ops/sec and p50/p99 for each operation class, writes
the numbers to a tracked results file, and fails if any number crosses
a recorded regression threshold.
- What to achieve?
- **Given** ~1M rows and a mixed read/write workload at the target
concurrency held for the target duration,
- **when** it runs,
- **then** throughput stays above the floor, p99 stays under the ceiling,
RSS is flat between a warmed baseline and the end (no growth beyond
tolerance), zero descriptors leak, and — for a write-inclusive run under
a durable configuration — a kill mid-load followed by replay loses no
acknowledged write.
- What to achieve?
- **Given** any later optimization iteration (7b, 8, 23),
- **when** it claims a speedup,
- **then** this benchmark's before/after numbers are in that iteration's
record, and a claim with no measured delta is not accepted.
## Out Of Scope
- **The optimizations themselves.** This iteration MEASURES; 7b/8/23 change.
A single-thread RAM-authoritative baseline is a legitimate first number —
the point is to have one before anyone tunes.
- **Distributed / multi-machine load.** Same-machine, one process (or one
process per shard once iteration 8 lands). Cross-host is the network layer's
concern, much later.
- **Micro-optimizing the benchmark harness.** It must be honest and
repeatable, not itself fast; if the harness is the bottleneck the spec says
so and fixes that, but a perfect load generator is not the deliverable.
- **A cost-based query planner.** Index selection is 9b's; this iteration
measures what 9b lowers, it does not make the planner smarter.
## Info
Forks the spec must settle:
**1. What generates the load, and in what language?** The doctrine is "the
sample is the test", so the honest generator drives compiled `.wo` — a
benchmark mode in the employee program (or a sibling sample) that loops
inserts/reads/queries and times them. The alternative — a C harness calling
the engine API directly — measures the engine but skips the compiler's
lowering, which is exactly the layer a language-integrated query has to pay
for. Leaning: `.wo` benchmark mode for the headline numbers (the number that
matters is end to end), with the C-API microbench kept only to attribute a
regression to engine vs lowering.
**2. What are the actual budgets?** Throughput floors and latency ceilings
have to be numbers, and the first run sets them — but the spec must decide
whether the gate is absolute (">= N ops/sec on the reference machine") or
relative ("no worse than the last recorded run by more than X%"). Absolute
gates rot across machines; relative gates need a committed baseline file.
Leaning: relative gates against a tracked `bench/baseline.json`, refreshed
deliberately with a commit that says why, plus a loud absolute floor so a
catastrophic regression fails even on a slow machine.
**3. What does "1M users read and write" concretely mean?** A million
long-lived idle connections is a different test from a million rows under a
churning read/write mix from a bounded connection pool. The sample's shape
(departments, employees) suggests rows, not connections, as the scale axis
for THIS iteration; the connection-scale test belongs with the shard-actor
runtime (iteration 8) and the eventual network layer. Leaning: ~1M rows +
a bounded concurrent read/write workload here; connection scale deferred to
8 with a cross-reference.
**4. Durable or RAM-only for the throughput headline?** fsync-per-commit
(the current per-statement durability) will dominate write throughput and is
the honest number for a durable workload; RAM-only (no `WO_DATA`) measures
the engine's ceiling. Both matter and mean different things. Leaning:
publish both, labeled — durable is the number an operator plans against, and
the gap between them is precisely what iteration 23 (io_uring group-commit)
exists to close.
## Proposed Solution
- **Brainstorm the spec**, settling the four forks; then a plan whose first
task is the harness and the baseline file, because nothing downstream means
anything without them.
- **Sequence the whole performance arc around this iteration:**
1. 9b lands → employee compiles and runs → **22 restart-persistence** and
**22 baseline benchmark** (single-thread, both durable and RAM-only).
2. **7b** (inferred GC + mark-sweep) → re-run 22, record the delta (does
tracing change the write path's tail latency?).
3. **8** (shard-actor, thread-per-core) → re-run 22 at the connection/
concurrency scale it unlocks, record the delta.
4. **23** (io_uring group-commit) → re-run 22's durable write number, record
the delta against the fsync-per-commit baseline — the payoff.
- The benchmark harness and its baseline live under `bench/` (or the existing
`runtime/bench/`), and `just` gets a `db-bench` recipe kept off the fast
path, exactly like `log-watcher::soak`.