docs: iteration 22 plan ready; slice active

- plan 2026-08-21-db-bench.md: 6 tasks (time.ticks builtin, sample,
  concurrent modes, driver+gate, first baseline + gate-bites proof,
  closeout); words + verification commands per convention
- spec banner APPROVED; story 22 to in-progress/ (frontmatter synced);
  marker doc created; board standup/rows/pending updated

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
This commit is contained in:
shoney.arickathil 2026-08-21 16:07:51 +02:00
parent cf4fe54c84
commit 551550cea4
8 changed files with 289 additions and 15 deletions

View file

@ -0,0 +1,36 @@
# In progress — iteration 22: db-bench, the measurement backbone
> **Status: 🔄 in progress** (spec + plan approved 2026-08-21) — second
> slice of the concurrency chain **✅ stage 3 → 22 → 31 → 24 → 23 → 32**.
> Board: [../stories/00-status.md](../stories/00-status.md).
>
> One marker doc per active slice; deleted when the slice lands.
## What
Execute [`superpowers/plans/2026-08-21-db-bench.md`](../superpowers/plans/2026-08-21-db-bench.md)
(spec: [`2026-08-21-db-bench-design.md`](../superpowers/specs/2026-08-21-db-bench-design.md),
approved): the `time.ticks` µs builtin, the `docs/examples/db-bench`
sample (seed/read/query/write/mix/msgrate/verify), the campaign driver
with relative gates vs `bench/baseline.json`, the restart proof and
kill -9 battery at both shard counts, and the first committed baseline.
## Why now
Nothing performance-shaped is sourced until this runs — and the arc owes
its before/after. One campaign now covers single- AND multi-shard
honestly (stage 3 landed), prices the RPC, and produces the mutex-inbox
number stage-2's deviation waits on.
## Story
- [22 — durability, throughput, scale](../stories/language-runtime-database/in-progress/22-durability-throughput-scale.md)
## Definition of done
- Plan Tasks 1–6 checked; `just db-bench` exit 0 twice in a row;
gate-bites smoke proven (doctored results FAIL); baseline committed
with rationale; arc delta + msgrate copied into story 8; full battery
green.
- Story 22 → done/; board standup written from the actual numbers; this
file deleted. Next slice: iteration 31 (actor lifecycle).

View file

@ -35,10 +35,13 @@ machine-readable truth behind this board; live Obsidian Dataview views:
## ▶ NEXT PLAN
**The concurrency + fiber chain — ✅ stage 3 → 22 → 31 → 24 → 23 → 32**
(directive 2026-08-21). Next slice: **iteration 22, the measurement
backbone** — its spec brainstorm is the next act (four forks recorded in
[the story](language-runtime-database/refine/22-durability-throughput-scale.md));
no marker doc until it starts.
(directive 2026-08-21). Active slice: **iteration 22, the measurement
backbone** — spec + plan approved 2026-08-21
([spec](../superpowers/specs/2026-08-21-db-bench-design.md) ·
[plan](../superpowers/plans/2026-08-21-db-bench.md) ·
[marker](../in-progress/2026-08-21-db-bench.md)); the four forks settled
as their leanings, plus `time.ticks` (µs clock) as the one runtime
addition and a new `db-bench` sample as the vehicle.
**Implemented last time (2026-08-21):** the arc's **stage 3 — the
transparent DB actor landed, the arc is COMPLETE** (stories
@ -230,7 +233,7 @@ that sequences its tasks. Read one, approve, then the next starts.
| 9b | [`@table`, relations, query](language-runtime-database/done/09b-table-relations-query.md) | 🔄 query surface + relations + FK done (branch query-surface); group-by parked |
| 19 | [Float + Bytes](language-runtime-database/done/19-missing-scalar-types.md) | ✅ **landed 2026-08-20** — `.wob` v5: Float constant tag, field kinds 6/7, opcodes 34-41 (IEEE-quiet f64), builtins 70-83. Full stack: literals, arithmetic, `@table` column, WAL bit-exact replay, json fractions in / shortest-round-trip out, `?Float` reserved-NaN nil, total-order index (NaN last, `-0.0` == `+0.0`), Bytes + base64. No implicit Int/Float mixing (WO-E201); `float`/`trunc` are the only bridges. Proof: web-app price is a real Float (`{"price":9.99}`), `just web-app` 23/0; corpus 103/0 |
| 11 | [Fibers](language-runtime-database/done/11-fibers.md) | ✅ **landed 2026-08-21** with the arc (`just fibers` 10/0); fs-park re-scoped out of v1, disclosed in the story |
| 22 | [Durability, throughput, scale](language-runtime-database/refine/22-durability-throughput-scale.md) | ⬜ needs a spec first — second in chain, after arc stage 3 |
| 22 | [Durability, throughput, scale](language-runtime-database/in-progress/22-durability-throughput-scale.md) | 🔄 **spec + plan approved 2026-08-21, executing** — db-bench sample + campaign gates; forks settled |
| 31 | [Actor lifecycle](language-runtime-database/refine/31-actor-lifecycle.md) | ⬜ needs a spec first — third in chain (story written 2026-08-21) |
| 24 | [chat: WebSocket workload](language-runtime-database/refine/24-chat-websocket-workload.md) | ⬜ fourth in chain — the arc's acceptance; after 31 |
| 23 | [io_uring group-commit](language-runtime-database/refine/23-io-uring-commit.md) | ⬜ fifth in chain, after stage 3 + 22 |
@ -253,7 +256,7 @@ that sequences its tasks. Read one, approve, then the next starts.
| Track | Item | Where |
| -------- | --------------------------------------------------------------------------- | ---------------------------------------------------------- |
| Runtime | nothing active — arc stage 3 landed 2026-08-21; next per the chain: iteration 22's spec brainstorm (four forks recorded in its story) | [order](#implementation-order-re-sequenced-2026-08-21--concurrency-chain) |
| Runtime | **iteration 22: db-bench** — spec + plan approved 2026-08-21; executing | [marker](../in-progress/2026-08-21-db-bench.md) · [plan](../superpowers/plans/2026-08-21-db-bench.md) |
The active slice's marker doc lives in [`in-progress/`](../in-progress/) —
one file, deleted when the slice lands. Everything else pending is the
@ -441,7 +444,8 @@ precedence notes for resumption.
database is broken today. Plan of record:
[`2026-08-20-shard-fiber-arc.md`](../superpowers/plans/2026-08-20-shard-fiber-arc.md)
(stages 1+2 landed 2026-08-20, branch `concurrency-arc`).
2. **22** — the measurement backbone: restart-persistence proof + baseline
2. 🔄 **22** — IN PROGRESS (spec + plan approved 2026-08-21) — the
measurement backbone: restart-persistence proof + baseline
benchmark (durable + RAM-only), single- AND multi-shard in one
campaign, plus the stage-2 mutex-inbox number (rings only if the mutex
costs). It has never run — no `bench/baseline.json`, no `just db-bench`;

View file

@ -100,7 +100,7 @@ still pending IS the runtime-concurrency chain; order:
| 11 | 15 | [deps: `wo.toml [deps]`](done/15-deps-package-manager.md) | exact-rev git deps + `wo.lock` + `.wo-deps`; flat-only, offline once locked |
| 12 | 16 | [web framework](done/16-web-framework.md) | the `.wo` framework v1 (router, middleware, auth, all three body hooks) consumed via `[deps]` |
| 13 | 8+11 | [Shard-actor runtime](done/08-shard-actor-runtime.md) · [Fibers](done/11-fibers.md) | ✅ **THE ARC LANDED 2026-08-21** — stages 1+2 (fibers/budget/actors/io_uring plane; pinned shards, envelope sends, home-routed frees, WO-E222) + stage 3's transparent DB actor: worker statements marshal to shard 0, ack-after-owner-fsync, materialized replies (`just db-actor` 8/0, ASan/TSan, WAL replay pair). fs-park re-scoped out (disclosed in story 11). |
| 14 | 22 | [Durability, throughput, scale](refine/22-durability-throughput-scale.md) | restart-persistence proof, benchmarks, ~1M rows — the baseline the arc and 23 sign against. It has never run, so every performance claim on this project is currently unsourced; runs after stage 3 so one campaign covers single- and multi-shard, plus the stage-2 mutex-inbox number. *(was 9e)* |
| 14 | 22 | [Durability, throughput, scale](in-progress/22-durability-throughput-scale.md) | 🔄 **spec + plan approved 2026-08-21, executing** — db-bench sample (`time.ticks` µs clock, seed/read/query/write/mix/msgrate/verify), campaign gates vs `bench/baseline.json`, restart + kill -9 proofs at both shard counts; the arc's delta and the mutex-inbox number come out of the first run. *(was 9e)* |
| 15 | 30 | Observability, CI, fuzz *(no story file yet)* | **NEW** — runtime counters + a profiler hook, 22's harness wired to run per change instead of by hand, and a fuzz target on the parser and `.wob` loader. The whole proof-maturity gap had no iteration to point at. |
| 16 | 19 | [Float + Bytes](done/19-missing-scalar-types.md) | **LANDED 2026-08-20** — `.wob` v5; the full stack: IEEE-quiet f64 through literals/VM/@table/WAL/json + Bytes as the binary carrier, no implicit mixing, total-order indexes. Unblocks 24 (WS frames) and the crypto fork (digests). *(was 20)* |
| 17 | 31 | [Actor lifecycle](refine/31-actor-lifecycle.md) | request/response (today `send` is one-way and callers `sleep` to await), bounded mailboxes with backpressure (today the FIFO just grows), actor death/supervision, and timers beyond `time.sleep`. 24 cannot be written honestly without these. *(story written 2026-08-21)* |

View file

@ -97,7 +97,7 @@ status: done
- **Gated by the benchmark (2026-08-15):** this is the "implement garbage
collection" lever of the performance arc — tri-color mark-sweep replacing
RC changes the write path's tail latency, so landing it means re-running
iteration [22](../refine/22-durability-throughput-scale.md) and recording the
iteration [22](../in-progress/22-durability-throughput-scale.md) and recording the
delta (does tracing help or hurt p99 under write load?).
- **Constraint added by the database track (2026-08-15):** a GC-managed value
in a `@table` field is a compile error (the engine/heap bulkhead — 9b

View file

@ -159,7 +159,7 @@ the slice's marker doc when it landed):
- The VM's object header has carried a shard id since iteration 2 — no
relayout.
- **Gated by the benchmark:** landing the arc means re-running
[22](../refine/22-durability-throughput-scale.md) at the concurrency
[22](../in-progress/22-durability-throughput-scale.md) at the concurrency
scale it unlocks and recording the before/after delta; it is also
where [23](../refine/23-io-uring-commit.md) gets a thread to overlap
durability against.

View file

@ -1,6 +1,6 @@
---
iteration: "22"
status: refine
status: in-progress
chain: 2
---
@ -16,7 +16,16 @@ chain: 2
> performance work, because each of those must be gated by re-running THIS
> iteration's benchmark and showing the number moved the right way.
>
> **No spec exists yet.** The forks in *Info* are genuine decisions.
> ~~**No spec exists yet.** The forks in *Info* are genuine decisions.~~
>
> **SPEC APPROVED 2026-08-21** — the four forks below are SETTLED as
> their recorded leanings (developer confirmation), plus two new
> decisions: the vehicle is a NEW sample `docs/examples/db-bench`
> (employee stays a teaching sample) and `time.ticks` (CLOCK_MONOTONIC
> µs) is the iteration's one runtime addition. Spec:
> [`2026-08-21-db-bench-design.md`](../../../superpowers/specs/2026-08-21-db-bench-design.md)
> · plan: [`2026-08-21-db-bench.md`](../../../superpowers/plans/2026-08-21-db-bench.md)
> — **in progress** (second slice of the chain).
>
> **RE-SEQUENCED 2026-08-21** (developer decision): runs AFTER the arc's
> stage 3 — the transparent DB actor is a correctness hole (a multi-shard

View file

@ -0,0 +1,224 @@
# Iteration 22 — db-bench: durability proof, throughput, scale (implementation plan)
> **Status: ready to execute (2026-08-21).** Board:
> [docs/00-status.md](../../stories/00-status.md).
> **For agentic workers:** REQUIRED SUB-SKILL: Use
> superpowers:subagent-driven-development (recommended) or
> superpowers:executing-plans to implement this plan task-by-task. Steps
> use checkbox (`- [ ]`) syntax for tracking.
>
> **Style rule (user convention):** concept, reason, and required
> behavior in words plus verification commands only — no implementation
> or test code blocks; the executor writes the code.
**Goal:** the measurement backbone — a `.wo` benchmark sample, a campaign
driver, a committed baseline contract, and the durability proofs, so
every later optimization signs a measured before/after.
**Architecture:** four artifacts (spec §1): `docs/examples/db-bench`
(load generator, pure `.wo`), `scripts/db-bench.sh` (campaign driver +
gates), `bench/baseline.json` (the contract), `just db-bench` /
`just db-bench-quick`. One runtime addition: the `time.ticks` builtin
(CLOCK_MONOTONIC microseconds) — everything else is sample + script.
**Tech Stack:** `.wo` (generator), bash + python3 (driver/gate — the
linkcheck.py precedent), C11 libc-only (one builtin case), OCaml
stdlib-only (one stdlib-table row).
**Spec:** [`../specs/2026-08-21-db-bench-design.md`](../specs/2026-08-21-db-bench-design.md)
(approved 2026-08-21, normative — the four forks + decisions 5/6 live
there). Story:
[`22-durability-throughput-scale.md`](../../stories/language-runtime-database/in-progress/22-durability-throughput-scale.md).
## Global Constraints
- Branch `db-bench` off the current arc line; commits local only, never
push.
- Gates that stay green after every task: `just woc-test`,
`just oop-e2e`, `just deps-accept`, `just web-app`,
`just log-watcher`, `just employee`, `just fibers`, `just db-actor`.
- The ONLY runtime change is the `time.ticks` builtin (spec decision 6);
everything else must not touch `runtime/src` or `database/src`.
- `bench/results/` is gitignored; `bench/baseline.json` is tracked and
changes only with a commit that says why.
- The headline numbers come from compiled `.wo` end to end; the C-API
microbench is attribution-only and never gated (spec §5).
---
## Task 1 — the `time.ticks` builtin (the honest clock)
**Files:**
- Modify: `runtime/src/wob.h` (new builtin id 84, `WO_B_MAX` bump),
`runtime/src/sysio.c` (the case, beside `WO_B_TIME_SLEEP`),
`compiler/src/types.ml` (the stdlib table row beside
`m "time" "sleep" 1 46`; grep for the sibling tables in `owner.ml`/
`emit.ml` that list stdlib names and mirror the row wherever `sleep`
appears), `docs/plan/oop-vm/08-builtin-surface.md` +
`docs/plan/oop-vm/07-systems-stdlib.md` (the contract rows).
- Test: a corpus fixture `tests/corpus/run/time-ticks` and one line in
the existing runtime/compiler suites only if their tables enumerate
builtins.
**Interfaces:**
- Produces: `time.ticks()` — zero args, returns Int microseconds from
CLOCK_MONOTONIC (never wall clock: it must be immune to NTP steps;
the difference of two calls is a duration). Later tasks time every
operation with it.
- [ ] Corpus fixture first: two `time.ticks()` calls around a spin loop;
assert the difference is non-negative and the second call is >= the
first (exact values are machine noise — the fixture asserts ordering
and that the builtin exists). Verify it FAILS today (unknown builtin).
- [ ] Wire the id (84), the sysio case (clock_gettime MONOTONIC,
seconds*1e6 + nsec/1e3, as Int), the compiler table rows, the two
contract-doc rows (name, arity 0, return Int µs, monotonic-not-wall
wording).
- [ ] Verify: fixture green under `just oop-e2e`; full battery green.
Commit.
## Task 2 — db-bench sample: tables, serial modes, per-op stats
**Files:**
- Create: `docs/examples/db-bench/wo.toml`,
`docs/examples/db-bench/types.wo` (the two related tables — a parent
with a `@unique` Text column, a child with `ref` parent + two indexed
columns; the employee shape, spec §2),
`docs/examples/db-bench/main.wo` (argv mode dispatch, employee's
pattern), `docs/examples/db-bench/README.md` (what each mode measures
and the output-line contract).
**Interfaces:**
- Consumes: `time.ticks()` from Task 1.
- Produces: modes `seed N`, `read N`, `query N`, `write N`, `verify`;
every measured mode prints exactly one line per operation class in
the spec's contract: `<op> <count> <ops/sec> <p50us> <p99us>`.
`verify` recounts, checksums contents, runs one indexed probe, exits
nonzero on mismatch; `seed`/`write` print an acknowledged high-water
line (`acked <n>`) the crash battery reads (spec §4).
- [ ] Tables + `seed`/`verify` first: seed writes N children (parents
amortized 1:100), verify recomputes count + a Int-sum checksum over an
indexed column + probes one known unique parent. Prove the pair by
hand: seed 10k, verify exits 0; corrupt expectation (verify 10k+1)
exits 1.
- [ ] Per-op timing: a fixed-size reservoir (spec §2 — honest, not
clever) collecting per-op durations from `time.ticks`; p50/p99 by
sorting the reservoir at report time; ops/sec from total ticks.
- [ ] `read`/`query`/`write` modes over a seeded store, each emitting
the contract line; a malformed mode prints usage and exits 2
(employee's shape).
- [ ] Verify: run all modes by hand single-shard (`WO_SHARDS=1`), lines
parse (field count + numeric), `just` battery untouched. Commit.
## Task 3 — mix + msgrate: the concurrent modes
**Files:**
- Modify: `docs/examples/db-bench/main.wo` (+ a `types.wo` message
class), README rows.
**Interfaces:**
- Consumes: Task 2's tables, stats, output contract.
- Produces: `mix N C` — C spawned actors each running the 90/10
read/write mix, N total ops, one contract line for reads and one for
writes plus the `acked` high-water; `msgrate N` — two actors
ping-ponging N messages, printing `msgrate <N> <msgs/sec>`. Under
`WO_SHARDS=1` everything is same-shard (fibers); at default cores
placement spreads the actors and every DB statement rides the stage-3
RPC — no bench code may check the shard count (transparency is the
point).
- No request/response surface exists (iteration 31): main drives
completion the db-actor way — actors bump rows a completion `verify`
can count; main sleep-polls the store until the expected count, then
settles. Document that as the coordination idiom this side of 31.
- [ ] `mix`: prove single-shard first (deterministic-ish, one thread),
then default cores; both emit parseable lines; TSan flavor of the
binary runs one short mix clean.
- [ ] `msgrate`: same proof shape; single- and multi-shard runs both
print (same-heap vs mutex-inbox comparison, spec §2).
- [ ] Verify: hand runs at both shard counts + TSan; battery. Commit.
## Task 4 — the campaign driver, gate, and recipes
**Files:**
- Create: `scripts/db-bench.sh` (driver), `scripts/db-bench-gate.py`
(JSON compare — python3, the linkcheck.py precedent),
`bench/baseline.json` (placeholder schema, values filled by Task 5),
`bench/results/.gitignore`.
- Modify: `justfile` (recipes `db-bench`, `db-bench-quick`),
`.gitignore` if bench/results needs a root rule.
**Interfaces:**
- Consumes: the sample's line contract + `acked` lines.
- Produces: one timestamped JSON per campaign under `bench/results/`
(structure: flavor → shards → mode → {count, ops_sec, p50us, p99us});
gate exit 0/1 with a per-metric summary tail (value, baseline, delta,
tolerance); the baseline schema: per metric `{value, tolerance_pct,
floor}` (spec fork 2).
- [ ] Driver: for flavor in ram/durable × shards in 1/default — seed,
read, query, write, mix (+ msgrate once per shard count); durable
runs under a temp `WO_DATA`; RSS + fd sampled during mix, failing on
LW_SOAK tolerances (growth > 256 KiB resident or any fd growth).
- [ ] Durability teeth in the driver (spec §4): restart proof
(seed → clean stop → rerun `verify`), crash battery (kill -9
mid-`write` K times per shard count, restart, `verify` against the
last `acked` high-water; K lives in baseline.json).
- [ ] Gate: compares every metric to baseline (relative tolerance +
absolute floor), prints the standup tail, exit nonzero on breach;
`--write-baseline` records a run as the new contract.
- [ ] `db-bench-quick`: seconds-long counts, loose gate (floors only) —
the CI-shaped smoke.
- [ ] Verify: quick mode end-to-end green against a freshly written
baseline; battery. Commit.
## Task 5 — the first campaign: baseline, deltas, gate-bites proof
**Files:**
- Modify: `bench/baseline.json` (the first honest full run's values +
chosen tolerances + floors + K),
`docs/stories/language-runtime-database/done/08-shard-actor-runtime.md`
(the arc's recorded delta: single- vs multi-shard columns),
`docs/examples/db-bench/README.md` (the reference-machine numbers).
- [ ] Full campaign on this machine; inspect the tail; commit the
baseline with a message that says it IS the first contract.
- [ ] Gate-bites smoke (spec acceptance): doctor a copy of the results
(halve one ops/sec) and run the gate against it — must FAIL; then the
real results — must PASS. Record the procedure in the README.
- [ ] Copy the arc delta + msgrate numbers into story 8's record and
the mutex-inbox note (arc plan deviation 4 references it).
- [ ] Verify: `just db-bench` exit 0 twice in a row (repeatability);
battery. Commit.
## Task 6 — closeout
**Files:**
- Modify: story 22 (landing banner, criteria check, frontmatter
`status: done`, file moves refine/→done/ with link sweep), the board
(standup entry: implemented/findings/learned/unblocked/next/.dev-ref;
In-progress row; pending list; stories table), `00-story.md` row,
`00-dependency-graph.md` node, the spec's status banner (APPROVED →
landed), `docs/in-progress/` marker deleted.
- [ ] Docs synced (the standup answers written from the actual numbers);
links verified (the session's link-check loop); frontmatter matches
folders.
- [ ] Full battery + `just db-bench-quick` once more after doc edits.
Commit.
## Self-review notes
- Spec coverage: §1→T4 artifacts + T1 clock; §2 modes→T2/T3; §3
campaign+baseline→T4/T5; §4 durability→T4 driver + T5 run; §5
microbench→deliberately NOT a task until a regression needs blaming
(YAGNI — the spec calls it attribution-only; first need creates it,
recorded here so the omission is a decision, not a gap); §6
acceptance→T5 (gate-bites, repeatability) + T6.
- The only interface later tasks depend on from T1 is `time.ticks()`
returning Int µs; T2's line contract is quoted verbatim where T4
parses it.
- No code blocks by user convention; every step names its verification
command or observable.

View file

@ -1,12 +1,13 @@
# Iteration 22 — durability proof, throughput, and scale under load (design)
**Date:** 2026-08-21
**Status:** proposed (story iteration 22) — awaiting developer review; the
plan follows approval. Board: [docs/00-status.md](../../stories/00-status.md)
**Status:** ✅ APPROVED 2026-08-21 (developer review) — plan ready:
[`2026-08-21-db-bench.md`](../plans/2026-08-21-db-bench.md).
Board: [docs/00-status.md](../../stories/00-status.md)
**Scope:** the measurement backbone — a benchmark workload in `.wo`, a
campaign driver script, a tracked baseline contract, and the durability
proofs (restart persistence + crash battery), single- AND multi-shard.
**Relates to:** [story 22](../../stories/language-runtime-database/refine/22-durability-throughput-scale.md)
**Relates to:** [story 22](../../stories/language-runtime-database/in-progress/22-durability-throughput-scale.md)
(the four forks settled below), the landed arc
([story 8's guarantee contract](../../stories/language-runtime-database/done/08-shard-actor-runtime.md)
— the stage-3 delta this iteration records), iteration 23 (the durable