writeonce/docs/superpowers/plans/2026-08-21-db-bench.md
shoney.arickathil c0b0dbb846 docs: audit all markdown against the code, fix findings, flatten status folders
- README: shipped concurrency/HTTP/WebSockets sat in the roadmap as "not yet
  available"; "no package manager" contradicted [deps]; the deps example
  would not have compiled (the key IS the module name)
- runtime/README: leads with wovm, wo-rt.c demoted to a historical section;
  dropped 2 nonexistent recipes, crates/rt, @gc refcounting, 13 suites -> 18
- employee + log-watcher READMEs claimed "does not compile"; both are gates
- error catalog: +10 emitted codes incl WO-E250, the only diagnostic the
  shipped query surface raises; recorded why the sweep rotted
- language-surface: group-by parses, then the typechecker refuses it
- 00-code-review + 00-link-audit re-run; history kept, not rewritten
- 48 dead Rust-era exploration links de-linked rather than re-pointed (their
  prose names the retired plan by number); successor map -> discarded.md
- 08-project-structure: compiler/plan/ never existed; corpus has 9 dirs, 5 empty
- releasing.md: dropped a --draft step the workflow never had
- new docs/00-doc-audit.md: findings + disposition, incl one row where the
  audit was wrong and the doc it accused was right
- status folders removed: 34 stories flat, status only in frontmatter; 252
  links recomputed from resolved paths; board/board-views/structure retaught
- story 24 -> in-progress, since frontmatter is now the only truth
- new iteration 38: fs mutation verbs + net.connect, the two capability
  families no iteration owned
- new iteration 39: gofiber/fiber v3.5.0 parity study. The ledger called
  CSRF/sessions unblocked by iteration 34's HMAC, but the runtime has no
  source of randomness at all
- linkcheck skips .dev/.superpowers: 0 broken paths, 0 bad anchors

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-26 19:20:22 +02:00

232 lines
11 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

# Iteration 22 — db-bench: durability proof, throughput, scale (implementation plan)
> **Status: ✅ LANDED 2026-08-21** — all six tasks executed inline,
> deviations disclosed per task. Board:
> [docs/00-status.md](../../stories/00-status.md).
> **For agentic workers:** REQUIRED SUB-SKILL: Use
> superpowers:subagent-driven-development (recommended) or
> superpowers:executing-plans to implement this plan task-by-task. Steps
> use checkbox (`- [ ]`) syntax for tracking.
>
> **Style rule (user convention):** concept, reason, and required
> behavior in words plus verification commands only — no implementation
> or test code blocks; the executor writes the code.
**Goal:** the measurement backbone — a `.wo` benchmark sample, a campaign
driver, a committed baseline contract, and the durability proofs, so
every later optimization signs a measured before/after.
**Architecture:** four artifacts (spec §1): `docs/examples/db-bench`
(load generator, pure `.wo`), `scripts/db-bench.sh` (campaign driver +
gates), `bench/baseline.json` (the contract), `just db-bench` /
`just db-bench-quick`. One runtime addition: the `time.ticks` builtin
(CLOCK_MONOTONIC microseconds) — everything else is sample + script.
**Tech Stack:** `.wo` (generator), bash + python3 (driver/gate — the
linkcheck.py precedent), C11 libc-only (one builtin case), OCaml
stdlib-only (one stdlib-table row).
**Spec:** [`../specs/2026-08-21-db-bench-design.md`](../specs/2026-08-21-db-bench-design.md)
(approved 2026-08-21, normative — the four forks + decisions 5/6 live
there). Story:
[`22-durability-throughput-scale.md`](../../stories/language-runtime-database/22-durability-throughput-scale.md).
## Global Constraints
- Branch `db-bench` off the current arc line; commits local only, never
push.
- Gates that stay green after every task: `just woc-test`,
`just oop-e2e`, `just deps-accept`, `just web-app`,
`just log-watcher`, `just employee`, `just fibers`, `just db-actor`.
- The ONLY runtime change is the `time.ticks` builtin (spec decision 6);
everything else must not touch `runtime/src` or `database/src`.
- `bench/results/` is gitignored; `bench/baseline.json` is tracked and
changes only with a commit that says why.
- The headline numbers come from compiled `.wo` end to end; the C-API
microbench is attribution-only and never gated (spec §5).
---
## Task 1 — the `time.ticks` builtin (the honest clock)
**Files:**
- Modify: `runtime/src/wob.h` (new builtin id 84, `WO_B_MAX` bump),
`runtime/src/sysio.c` (the case, beside `WO_B_TIME_SLEEP`),
`compiler/src/types.ml` (the stdlib table row beside
`m "time" "sleep" 1 46`; grep for the sibling tables in `owner.ml`/
`emit.ml` that list stdlib names and mirror the row wherever `sleep`
appears), `docs/plan/oop-vm/08-builtin-surface.md` +
`docs/plan/oop-vm/07-systems-stdlib.md` (the contract rows).
- Test: a corpus fixture `tests/corpus/run/time-ticks` and one line in
the existing runtime/compiler suites only if their tables enumerate
builtins.
**Interfaces:**
- Produces: `time.ticks()` — zero args, returns Int microseconds from
CLOCK_MONOTONIC (never wall clock: it must be immune to NTP steps;
the difference of two calls is a duration). Later tasks time every
operation with it.
- [x] Corpus fixture first (`run/time-ticks`): RED as WO-E406, then
green — asserts monotonicity + nonzero, values stay machine noise.
- [x] Wired: wob.h id 84 + WO_B_MAX bump; the sysio case; ONE compiler
table row (types.ml is the single stdlib table — the plan's guess of
sibling tables in owner/emit was wrong, no mirror needed); the sys
dispatch range needed an explicit `C == WO_B_TIME_TICKS` arm (the
range gate stops at PROC_RUN). DEVIATION: `07-systems-stdlib.md`
does not exist (a planned doc never written) — the row went into
`08-builtin-surface.md` alone.
- [x] Verified: oop-e2e 104/0 (was 103); full battery green. Commit.
## Task 2 — db-bench sample: tables, serial modes, per-op stats
**Files:**
- Create: `docs/examples/db-bench/wo.toml`,
`docs/examples/db-bench/types.wo` (the two related tables — a parent
with a `@unique` Text column, a child with `ref` parent + two indexed
columns; the employee shape, spec §2),
`docs/examples/db-bench/main.wo` (argv mode dispatch, employee's
pattern), `docs/examples/db-bench/README.md` (what each mode measures
and the output-line contract).
**Interfaces:**
- Consumes: `time.ticks()` from Task 1.
- Produces: modes `seed N`, `read N`, `query N`, `write N`, `verify`;
every measured mode prints exactly one line per operation class in
the spec's contract: `<op> <count> <ops/sec> <p50us> <p99us>`.
`verify` recounts, checksums contents, runs one indexed probe, exits
nonzero on mismatch; `seed`/`write` print an acknowledged high-water
line (`acked <n>`) the crash battery reads (spec §4).
- [ ] Tables + `seed`/`verify` first: seed writes N children (parents
amortized 1:100), verify recomputes count + a Int-sum checksum over an
indexed column + probes one known unique parent. Prove the pair by
hand: seed 10k, verify exits 0; corrupt expectation (verify 10k+1)
exits 1.
- [ ] Per-op timing: a fixed-size reservoir (spec §2 — honest, not
clever) collecting per-op durations from `time.ticks`; p50/p99 by
sorting the reservoir at report time; ops/sec from total ticks.
- [ ] `read`/`query`/`write` modes over a seeded store, each emitting
the contract line; a malformed mode prints usage and exits 2
(employee's shape).
- [ ] Verify: run all modes by hand single-shard (`WO_SHARDS=1`), lines
parse (field count + numeric), `just` battery untouched. Commit.
## Task 3 — mix + msgrate: the concurrent modes
**Files:**
- Modify: `docs/examples/db-bench/main.wo` (+ a `types.wo` message
class), README rows.
**Interfaces:**
- Consumes: Task 2's tables, stats, output contract.
- Produces: `mix N C` — C spawned actors each running the 90/10
read/write mix, N total ops, one contract line for reads and one for
writes plus the `acked` high-water; `msgrate N` — two actors
ping-ponging N messages, printing `msgrate <N> <msgs/sec>`. Under
`WO_SHARDS=1` everything is same-shard (fibers); at default cores
placement spreads the actors and every DB statement rides the stage-3
RPC — no bench code may check the shard count (transparency is the
point).
- No request/response surface exists (iteration 31): main drives
completion the db-actor way — actors bump rows a completion `verify`
can count; main sleep-polls the store until the expected count, then
settles. Document that as the coordination idiom this side of 31.
- [ ] `mix`: prove single-shard first (deterministic-ish, one thread),
then default cores; both emit parseable lines; TSan flavor of the
binary runs one short mix clean.
- [ ] `msgrate`: same proof shape; single- and multi-shard runs both
print (same-heap vs mutex-inbox comparison, spec §2).
- [ ] Verify: hand runs at both shard counts + TSan; battery. Commit.
## Task 4 — the campaign driver, gate, and recipes
**Files:**
- Create: `scripts/db-bench.sh` (driver), `scripts/db-bench-gate.py`
(JSON compare — python3, the linkcheck.py precedent),
`bench/baseline.json` (placeholder schema, values filled by Task 5),
`bench/results/.gitignore`.
- Modify: `justfile` (recipes `db-bench`, `db-bench-quick`),
`.gitignore` if bench/results needs a root rule.
**Interfaces:**
- Consumes: the sample's line contract + `acked` lines.
- Produces: one timestamped JSON per campaign under `bench/results/`
(structure: flavor → shards → mode → {count, ops_sec, p50us, p99us});
gate exit 0/1 with a per-metric summary tail (value, baseline, delta,
tolerance); the baseline schema: per metric `{value, tolerance_pct,
floor}` (spec fork 2).
- [ ] Driver: for flavor in ram/durable × shards in 1/default — seed,
read, query, write, mix (+ msgrate once per shard count); durable
runs under a temp `WO_DATA`; RSS + fd sampled during mix, failing on
LW_SOAK tolerances (growth > 256 KiB resident or any fd growth).
- [ ] Durability teeth in the driver (spec §4): restart proof
(seed → clean stop → rerun `verify`), crash battery (kill -9
mid-`write` K times per shard count, restart, `verify` against the
last `acked` high-water; K lives in baseline.json).
- [ ] Gate: compares every metric to baseline (relative tolerance +
absolute floor), prints the standup tail, exit nonzero on breach;
`--write-baseline` records a run as the new contract.
- [ ] `db-bench-quick`: seconds-long counts, loose gate (floors only) —
the CI-shaped smoke.
- [ ] Verify: quick mode end-to-end green against a freshly written
baseline; battery. Commit.
## Task 5 — the first campaign: baseline, deltas, gate-bites proof
**Files:**
- Modify: `bench/baseline.json` (the first honest full run's values +
chosen tolerances + floors + K),
`docs/stories/language-runtime-database/done/08-shard-actor-runtime.md`
(the arc's recorded delta: single- vs multi-shard columns),
`docs/examples/db-bench/README.md` (the reference-machine numbers).
- [x] Full campaign run twice (20 checks/0 fail each incl. restart
proofs and 3x kill -9 batteries per shard count); baseline committed
as the first contract. DEVIATION: the two-run repeatability check
found ~25% jitter on read/query latencies and scheduling-dependent
spread on mix* — per-metric tolerances tuned (mix* 50%, read/query
35%, rest 15%, rationale in the baseline's _config note); both runs
pass the tuned contract 74/0.
- [x] Gate-bites proven: `--check` on a doctored copy (one ops/sec
halved) FAILS on exactly that metric; both real runs PASS. Procedure
in the README (a `--check` gate-only mode was added for this).
- [x] Arc delta + msgrate recorded: story 8 carries the measured table
(durable seed 4.5k/s vs ram 297k/s = 23's case; mixread 1280 vs 21
ops/s = the RPC x O(table)-probe price; msgrate 13.4M vs 2.45M =
deviation 4's mutex-inbox number). Headline findings in the sample
README: point lookups are O(table) — the read-path finding.
- [x] Verified: two full campaigns green; battery green. Commit.
## Task 6 — closeout
**Files:**
- Modify: story 22 (landing banner, criteria check, frontmatter
`status: done`, file moves refine/→done/ with link sweep), the board
(standup entry: implemented/findings/learned/unblocked/next/.dev-ref;
In-progress row; pending list; stories table), `00-story.md` row,
`00-dependency-graph.md` node, the spec's status banner (APPROVED →
landed), `docs/in-progress/` marker deleted.
- [x] Docs synced (standup from the measured numbers; story/spec/plan
banners; graph node; marker deleted); links verified; frontmatter
matches folders.
- [x] Full battery + `just db-bench-quick` once more after doc edits.
Commit.
## Self-review notes
- Spec coverage: §1→T4 artifacts + T1 clock; §2 modes→T2/T3; §3
campaign+baseline→T4/T5; §4 durability→T4 driver + T5 run; §5
microbench→deliberately NOT a task until a regression needs blaming
(YAGNI — the spec calls it attribution-only; first need creates it,
recorded here so the omission is a decision, not a gap); §6
acceptance→T5 (gate-bites, repeatability) + T6.
- The only interface later tasks depend on from T1 is `time.ticks()`
returning Int µs; T2's line contract is quoted verbatim where T4
parses it.
- No code blocks by user convention; every step names its verification
command or observable.