- bench/baseline.json: 74 metrics from the first full campaign; tolerances tuned by a two-run repeatability check (mix* 50%, read/query 35%, rest 15% — rationale in _config) - gate bites: --check mode; doctored copy fails, both real runs 74/0 - headline: durable seed 4.5k/s vs ram 297k/s (23's case); reads O(table) at ~1.5k/s; mixread 1280 vs 21 ops/s single-vs-multi (the arc's price); msgrate 13.4M vs 2.45M (mutex-inbox number) - arc delta recorded in story 8; findings in sample README Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
11 KiB
Iteration 22 — db-bench: durability proof, throughput, scale (implementation plan)
Status: ready to execute (2026-08-21). Board: docs/00-status.md.
For agentic workers: REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (
- [ ]) syntax for tracking.Style rule (user convention): concept, reason, and required behavior in words plus verification commands only — no implementation or test code blocks; the executor writes the code.
Goal: the measurement backbone — a .wo benchmark sample, a campaign
driver, a committed baseline contract, and the durability proofs, so
every later optimization signs a measured before/after.
Architecture: four artifacts (spec §1): docs/examples/db-bench
(load generator, pure .wo), scripts/db-bench.sh (campaign driver +
gates), bench/baseline.json (the contract), just db-bench /
just db-bench-quick. One runtime addition: the time.ticks builtin
(CLOCK_MONOTONIC microseconds) — everything else is sample + script.
Tech Stack: .wo (generator), bash + python3 (driver/gate — the
linkcheck.py precedent), C11 libc-only (one builtin case), OCaml
stdlib-only (one stdlib-table row).
Spec: ../specs/2026-08-21-db-bench-design.md
(approved 2026-08-21, normative — the four forks + decisions 5/6 live
there). Story:
22-durability-throughput-scale.md.
Global Constraints
- Branch
db-benchoff the current arc line; commits local only, never push. - Gates that stay green after every task:
just woc-test,just oop-e2e,just deps-accept,just web-app,just log-watcher,just employee,just fibers,just db-actor. - The ONLY runtime change is the
time.ticksbuiltin (spec decision 6); everything else must not touchruntime/srcordatabase/src. bench/results/is gitignored;bench/baseline.jsonis tracked and changes only with a commit that says why.- The headline numbers come from compiled
.woend to end; the C-API microbench is attribution-only and never gated (spec §5).
Task 1 — the time.ticks builtin (the honest clock)
Files:
- Modify:
runtime/src/wob.h(new builtin id 84,WO_B_MAXbump),runtime/src/sysio.c(the case, besideWO_B_TIME_SLEEP),compiler/src/types.ml(the stdlib table row besidem "time" "sleep" 1 46; grep for the sibling tables inowner.ml/emit.mlthat list stdlib names and mirror the row whereversleepappears),docs/plan/oop-vm/08-builtin-surface.md+docs/plan/oop-vm/07-systems-stdlib.md(the contract rows). - Test: a corpus fixture
tests/corpus/run/time-ticksand one line in the existing runtime/compiler suites only if their tables enumerate builtins.
Interfaces:
-
Produces:
time.ticks()— zero args, returns Int microseconds from CLOCK_MONOTONIC (never wall clock: it must be immune to NTP steps; the difference of two calls is a duration). Later tasks time every operation with it. -
Corpus fixture first (
run/time-ticks): RED as WO-E406, then green — asserts monotonicity + nonzero, values stay machine noise. -
Wired: wob.h id 84 + WO_B_MAX bump; the sysio case; ONE compiler table row (types.ml is the single stdlib table — the plan's guess of sibling tables in owner/emit was wrong, no mirror needed); the sys dispatch range needed an explicit
C == WO_B_TIME_TICKSarm (the range gate stops at PROC_RUN). DEVIATION:07-systems-stdlib.mddoes not exist (a planned doc never written) — the row went into08-builtin-surface.mdalone. -
Verified: oop-e2e 104/0 (was 103); full battery green. Commit.
Task 2 — db-bench sample: tables, serial modes, per-op stats
Files:
- Create:
docs/examples/db-bench/wo.toml,docs/examples/db-bench/types.wo(the two related tables — a parent with a@uniqueText column, a child withrefparent + two indexed columns; the employee shape, spec §2),docs/examples/db-bench/main.wo(argv mode dispatch, employee's pattern),docs/examples/db-bench/README.md(what each mode measures and the output-line contract).
Interfaces:
-
Consumes:
time.ticks()from Task 1. -
Produces: modes
seed N,read N,query N,write N,verify; every measured mode prints exactly one line per operation class in the spec's contract:<op> <count> <ops/sec> <p50us> <p99us>.verifyrecounts, checksums contents, runs one indexed probe, exits nonzero on mismatch;seed/writeprint an acknowledged high-water line (acked <n>) the crash battery reads (spec §4). -
Tables +
seed/verifyfirst: seed writes N children (parents amortized 1:100), verify recomputes count + a Int-sum checksum over an indexed column + probes one known unique parent. Prove the pair by hand: seed 10k, verify exits 0; corrupt expectation (verify 10k+1) exits 1. -
Per-op timing: a fixed-size reservoir (spec §2 — honest, not clever) collecting per-op durations from
time.ticks; p50/p99 by sorting the reservoir at report time; ops/sec from total ticks. -
read/query/writemodes over a seeded store, each emitting the contract line; a malformed mode prints usage and exits 2 (employee's shape). -
Verify: run all modes by hand single-shard (
WO_SHARDS=1), lines parse (field count + numeric),justbattery untouched. Commit.
Task 3 — mix + msgrate: the concurrent modes
Files:
- Modify:
docs/examples/db-bench/main.wo(+ atypes.womessage class), README rows.
Interfaces:
-
Consumes: Task 2's tables, stats, output contract.
-
Produces:
mix N C— C spawned actors each running the 90/10 read/write mix, N total ops, one contract line for reads and one for writes plus theackedhigh-water;msgrate N— two actors ping-ponging N messages, printingmsgrate <N> <msgs/sec>. UnderWO_SHARDS=1everything is same-shard (fibers); at default cores placement spreads the actors and every DB statement rides the stage-3 RPC — no bench code may check the shard count (transparency is the point). -
No request/response surface exists (iteration 31): main drives completion the db-actor way — actors bump rows a completion
verifycan count; main sleep-polls the store until the expected count, then settles. Document that as the coordination idiom this side of 31. -
mix: prove single-shard first (deterministic-ish, one thread), then default cores; both emit parseable lines; TSan flavor of the binary runs one short mix clean. -
msgrate: same proof shape; single- and multi-shard runs both print (same-heap vs mutex-inbox comparison, spec §2). -
Verify: hand runs at both shard counts + TSan; battery. Commit.
Task 4 — the campaign driver, gate, and recipes
Files:
- Create:
scripts/db-bench.sh(driver),scripts/db-bench-gate.py(JSON compare — python3, the linkcheck.py precedent),bench/baseline.json(placeholder schema, values filled by Task 5),bench/results/.gitignore. - Modify:
justfile(recipesdb-bench,db-bench-quick),.gitignoreif bench/results needs a root rule.
Interfaces:
-
Consumes: the sample's line contract +
ackedlines. -
Produces: one timestamped JSON per campaign under
bench/results/(structure: flavor → shards → mode → {count, ops_sec, p50us, p99us}); gate exit 0/1 with a per-metric summary tail (value, baseline, delta, tolerance); the baseline schema: per metric{value, tolerance_pct, floor}(spec fork 2). -
Driver: for flavor in ram/durable × shards in 1/default — seed, read, query, write, mix (+ msgrate once per shard count); durable runs under a temp
WO_DATA; RSS + fd sampled during mix, failing on LW_SOAK tolerances (growth > 256 KiB resident or any fd growth). -
Durability teeth in the driver (spec §4): restart proof (seed → clean stop → rerun
verify), crash battery (kill -9 mid-writeK times per shard count, restart,verifyagainst the lastackedhigh-water; K lives in baseline.json). -
Gate: compares every metric to baseline (relative tolerance + absolute floor), prints the standup tail, exit nonzero on breach;
--write-baselinerecords a run as the new contract. -
db-bench-quick: seconds-long counts, loose gate (floors only) — the CI-shaped smoke. -
Verify: quick mode end-to-end green against a freshly written baseline; battery. Commit.
Task 5 — the first campaign: baseline, deltas, gate-bites proof
Files:
-
Modify:
bench/baseline.json(the first honest full run's values + chosen tolerances + floors + K),docs/stories/language-runtime-database/done/08-shard-actor-runtime.md(the arc's recorded delta: single- vs multi-shard columns),docs/examples/db-bench/README.md(the reference-machine numbers). -
Full campaign run twice (20 checks/0 fail each incl. restart proofs and 3x kill -9 batteries per shard count); baseline committed as the first contract. DEVIATION: the two-run repeatability check found ~25% jitter on read/query latencies and scheduling-dependent spread on mix* — per-metric tolerances tuned (mix* 50%, read/query 35%, rest 15%, rationale in the baseline's _config note); both runs pass the tuned contract 74/0.
-
Gate-bites proven:
--checkon a doctored copy (one ops/sec halved) FAILS on exactly that metric; both real runs PASS. Procedure in the README (a--checkgate-only mode was added for this). -
Arc delta + msgrate recorded: story 8 carries the measured table (durable seed 4.5k/s vs ram 297k/s = 23's case; mixread 1280 vs 21 ops/s = the RPC x O(table)-probe price; msgrate 13.4M vs 2.45M = deviation 4's mutex-inbox number). Headline findings in the sample README: point lookups are O(table) — the read-path finding.
-
Verified: two full campaigns green; battery green. Commit.
Task 6 — closeout
Files:
-
Modify: story 22 (landing banner, criteria check, frontmatter
status: done, file moves refine/→done/ with link sweep), the board (standup entry: implemented/findings/learned/unblocked/next/.dev-ref; In-progress row; pending list; stories table),00-story.mdrow,00-dependency-graph.mdnode, the spec's status banner (APPROVED → landed),docs/in-progress/marker deleted. -
Docs synced (the standup answers written from the actual numbers); links verified (the session's link-check loop); frontmatter matches folders.
-
Full battery +
just db-bench-quickonce more after doc edits. Commit.
Self-review notes
- Spec coverage: §1→T4 artifacts + T1 clock; §2 modes→T2/T3; §3 campaign+baseline→T4/T5; §4 durability→T4 driver + T5 run; §5 microbench→deliberately NOT a task until a regression needs blaming (YAGNI — the spec calls it attribution-only; first need creates it, recorded here so the omission is a decision, not a gap); §6 acceptance→T5 (gate-bites, repeatability) + T6.
- The only interface later tasks depend on from T1 is
time.ticks()returning Int µs; T2's line contract is quoted verbatim where T4 parses it. - No code blocks by user convention; every step names its verification command or observable.