databasev2 4 part A, task 4. Scope extended with developer approval: the plan authorised touching the sample only for observability, but no existing leg has enough concurrent durable writes to exercise group commit at all, so the payoff was unevaluable either way. The finding that forced it: - `mix` writes on one op in ten with C=4 (all_mode calls mix_mode(n/10, 4); Mixer writes on i % 10 == 9), so the quick run performs 20 writes total. Measured mean batch 1.01 over 3112 barriers, peak 3 - that is a property of the WORKLOAD, not the mechanism: peak 3 of a possible 4 shows batches form whenever writes actually coincide - `wmix N C` added: every op a durable write, C at once. Updates rather than inserts, so it is comparable to mixwrite and the row count stays flat. Histogram kind 2 — a replayed store still holds the seeding run's kind-0/1 Hist rows and merging those would report someone else's latencies - WO_WAL_STATS=1 prints one line at exit: batches, records, peak_batch, peak_staged. Opt-in, because it would otherwise pollute every durable program's output. Counters live in wo_wal; no builtin, the numbers are diagnostic and not part of the language Measured, and it scales with concurrency exactly as designed: - C = 4 / 16 / 64 -> mean batch 1.13 / 1.76 / 5.35, peak 3 / 10 / 39 - the gate's own legs: durable.s1 5412 records over 5412 barriers (mean 1.0, peak 1 — the inline path, one barrier per statement BY DESIGN), durable.sN 7757 over 2296 (mean 3.38, peak 28) at 2x the throughput - peak staged 1372 B settles the no-cap decision with a number: the batch is tiny, so the upstream mailbox bound is sufficient - mean_batch/peak_batch are higher-is-better (the default detector would have called bigger batches worse) - only the batch SHAPE metrics are waived to 100%; wmix throughput and latency keep real tolerances (15% s1, 50% sN) — a blanket waiver would have left the entire new leg ungated - the live assertion `mean > 1.0` on the sN leg is what catches inertness - gate bites: sN wmix ops_sec -60% -> FAIL on exactly that metric, 1 of 86 Verified: db-bench-quick 89 checks 0 failures; baseline refreshed (86 metrics). Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|---|---|---|
| .. | ||
| main.wo | ||
| README.md | ||
| types.wo | ||
| wo.toml | ||
db-bench — iteration 22's load generator
The measurement backbone (spec:
docs/superpowers/specs/2026-08-21-db-bench-design.md). Pure .wo;
every measured mode prints one machine-parsable line per operation
class:
<op> <count> <ops/sec> <p50us> <p99us>
Timing is per-operation via time.ticks (CLOCK_MONOTONIC µs).
Percentiles come from a 1µs-bucket histogram clamped at 20000µs — exact
to the microsecond below the clamp; a p99 AT 20000 means "clamp or
worse". (A histogram, not the spec's reservoir: the language has no
container element-write or sort, and the histogram's tail fidelity is
strictly better. Recorded as a plan deviation.)
Modes
| mode | what it prices |
|---|---|
all N |
the throughput campaign in ONE process: seed N, read N/2, query N/10, write N/2. Without WO_DATA the store is RAM and dies with the process, so the measured modes must share the seeding run. |
seed N |
timed inserts: one parent per 100 children (FK probe each insert, unique-index maintenance per parent), k non-unique (10 rows/key), deterministic v. Writes Meta expectation rows. |
read N |
indexed take-1 point lookups, LCG-spread keys. |
query N |
full equality probes on the k index (≈10 rows each), materialized and counted. |
write N |
alternating inserts (disjoint k range 2e6+) and updates through query results. Corrupts the checksum by design — durability legs run on a fresh store. |
wal N |
the crash battery's vehicle: insert-only (k range 1e6+), acked <i> printed AFTER each insert returns — the return IS the ack (RAM applied, WAL record staged, ONE commit done). |
verify |
store vs its own Meta rows: count, checksum, one unique probe. Exit 3 on mismatch. |
verify-acked M |
after kill -9 mid-wal: rows 1..M exist with the right v; rows beyond M allowed (acked after the last print flushed). Exit 3 on mismatch. |
Coordination idiom (this side of iteration 31)
There is no request/response surface yet: concurrent modes drive completion the db-actor way — actors write rows, main polls the store until the expected count, then settles. Retired when 31 lands.
Standing finding (2026-08-21, first run)
A hand-built multi Bucket of insert results SEGVs on drop: the
compiler classifies the elements OWNED while table refs are scalar ids.
Query-built multis are runtime-typed and safe. Worked around here
(single ref local, bucket-major seeding); the compiler fix is its own
slice.
Reference-machine numbers (first campaign, 2026-08-21)
bench/baseline.json is the contract; headline readings:
- ram seed 245–290k inserts/s; durable seed ≈4.5k/s (fsync-per-commit ≈220µs each — the gap iteration 23 exists to close).
- reads/queries ≈1.1–1.3M ops/s at p50 1µs since the read-path index
slice (2026-08-22, engine
wo_idx_probe+ emitter index selection) — up from ≈1.5k/s at p50 600µs when point lookups walked every slab (~×850). mixread 89k ops/s single-shard, ~1.9k multi-shard (was 1,280 / 21): the RPC round-trip is now the visible cost, as designed. - msgrate ≈13M msgs/s same-heap vs ≈2.4M cross-shard (the mutex-inbox number, stage-2 deviation 4).
- Tolerance policy lives in the DRIVER (
tolerance_for), not hand-edits — a baseline refresh regenerates it: mix*/sN/read/query 50% (scheduling + µs-scale jitter), rest 15%; latency floorsmax(4×value, 100µs)— the tripwire means "µs became ms".
The gate must bite (proven 2026-08-21)
scripts/db-bench.py --check <results.json> evaluates a recorded run:
the real results pass 74/0; a doctored copy FAILS on exactly the
doctored metrics — use a 15%-class metric (seed) halved plus a
50%-class metric (read) quartered, so both tolerance classes prove they
bite. Re-run the smoke after any gate or policy change.