The branch was 17 ahead / 25 behind with 11 conflicting files, and drifting further: db.c had been rewritten twice on master since (group commit, then compaction). Resolved rather than rebased so both histories stay legible. Conflicts, and how each was settled: - db.c: BOTH semantics kept. Master's fatal path and compaction check now sit behind the branch's `table_is_durable` predicate, in all three inline arms — a volatile table reaches neither the barrier nor the compaction check - db-bench sample: every mode from both sides (growth, growth-verify, randread, replayseed, wmix) and ONE `boot` mode, which both sides had added independently - db-bench.py: all six legs kept. Both sides had also grown the same WAL-size helper under different names; collapsed into one - perf-targets: the branch's §5 (RAM ceiling) then master's §6/§7 — master's numbering had already assumed a §5 it did not have - story frontmatter: master's `status` (the landing truth) plus the branch's `readiness` axis. 03 would have read `done` + `refine`, which is a contradiction — it was brainstormed and landed on master, so `ready` - board: both standup blocks newest-first; master's chain rows (a superset); the branch's databasev2 1-2 rows with master's 3-4. Fixed a stray `|` in master's row 3 - baseline: master's, then REGENERATED from a full campaign — 143 metrics, 132 checks, 0 failures with both sides' legs present TWO HALF-EXPOSED FEATURES FIXED, because the merge rule is that master gets no feature that is honoured in name only: - `resident: keys` PARSED, set a .wob flag, and did nothing: rows stayed fully resident. A developer could declare a 120 GB table keys-resident, watch it compile, and be OOM-killed. The loader now REFUSES it with a message naming what to write instead, until tasks 5c/5d land. The compiler still parses it and its AST golden still passes, so the grammar work stays tested - `durable: false` was honoured ONLY on the inline path. wo_db_exec_req had no guard at all, so a volatile table written from an actor on a worker shard would still be logged — precisely porch's session-table case, and precisely what iteration 2 exists to provide. All three request-path arms now carry the same predicate. Found by reading the merged code, not by a test: the obvious probe runs main() on the primary and therefore only exercises the inline path Verified on the merged tree: wovm-test 0, woc-test 0, oop-e2e 122/0, residency-accept 8/0, db-bench 132/0, linkcheck clean. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|---|---|---|
| .. | ||
| main.wo | ||
| README.md | ||
| types.wo | ||
| wo.toml | ||
db-bench — iteration 22's load generator
The measurement backbone (spec:
docs/superpowers/specs/2026-08-21-db-bench-design.md). Pure .wo;
every measured mode prints one machine-parsable line per operation
class:
<op> <count> <ops/sec> <p50us> <p99us>
Timing is per-operation via time.ticks (CLOCK_MONOTONIC µs).
Percentiles come from a 1µs-bucket histogram clamped at 20000µs — exact
to the microsecond below the clamp; a p99 AT 20000 means "clamp or
worse". (A histogram, not the spec's reservoir: the language has no
container element-write or sort, and the histogram's tail fidelity is
strictly better. Recorded as a plan deviation.)
Modes
| mode | what it prices |
|---|---|
all N |
the throughput campaign in ONE process: seed N, read N/2, query N/10, write N/2. Without WO_DATA the store is RAM and dies with the process, so the measured modes must share the seeding run. |
seed N |
timed inserts: one parent per 100 children (FK probe each insert, unique-index maintenance per parent), k non-unique (10 rows/key), deterministic v. Writes Meta expectation rows. |
read N |
indexed take-1 point lookups, LCG-spread keys. |
query N |
full equality probes on the k index (≈10 rows each), materialized and counted. |
write N |
alternating inserts (disjoint k range 2e6+) and updates through query results. Corrupts the checksum by design — durability legs run on a fresh store. |
wal N |
the crash battery's vehicle: insert-only (k range 1e6+), acked <i> printed AFTER each insert returns — the return IS the ack (RAM applied, WAL record staged, ONE commit done). |
wmix N C |
databasev2 4: every op a durable write (update through a query result), C at once. Exists because mix writes on one op in ten with C=4 — 20 writes in a quick run, measured mean batch 1.01 — so no existing leg could show whether group commit engages. Histogram kind 2, because a replayed store still holds the seeding run's kind-0/1 Hist rows. Seed first. |
boot |
databasev2 3: does NOTHING. With WO_DATA set the runtime replays the whole log before main runs, so a mode with no work of its own is the only honest way to price boot |
verify |
store vs its own Meta rows: count, checksum, one unique probe. Exit 3 on mismatch. |
verify-acked M |
after kill -9 mid-wal: rows 1..M exist with the right v; rows beyond M allowed (acked after the last print flushed). Exit 3 on mismatch. |
Env knobs
| var | effect |
|---|---|
WO_DATA=<dir> |
durability on: replay <dir>/shard-0.wal at boot, log every write. Without it the store is RAM-only |
WO_SHARDS=<n> |
shard count. 1 means every statement runs inline on shard 0 and group commit cannot engage — batches form only where writes queue from other shards |
WO_CHECKPOINT_BYTES / WO_CHECKPOINT_RATIO |
databasev2 3: the checkpoint trigger — the log must exceed the floor AND exceed the ratio times the last compaction's own size. A tiny floor forces compaction in a few writes, which is how the gate tests the policy at all; an enormous one disables it, which is how the checkpoint leg measures the same workload with and without |
WO_WAL_STATS=1 |
databasev2 4: print one line at exit — walstats batches=… records=… peak_batch=… peak_staged=… compactions=… compact_us_max=… compact_us_total=… compacted_bytes=…. Opt-in so it does not pollute every durable program's output. Mean batch is records/batches; mean 1.0 means group commit is not engaging, which is expected for a serial writer or WO_SHARDS=1 and a bug anywhere else |
Do not put WO_DATA on /tmp. It is tmpfs on the reference machine,
where fdatasync is free: the same wmix run measured 195 000 ops/s at p50
1 µs there against 2200 ops/s at p50 7200 µs on ext4. There is no
durability barrier to price on a memory filesystem. The driver keeps its stores
under bench/ for exactly this reason.
Coordination idiom (this side of iteration 31)
There is no request/response surface yet: concurrent modes drive completion the db-actor way — actors write rows, main polls the store until the expected count, then settles. Retired when 31 lands.
Standing finding (2026-08-21, first run)
A hand-built multi Bucket of insert results SEGVs on drop: the
compiler classifies the elements OWNED while table refs are scalar ids.
Query-built multis are runtime-typed and safe. Worked around here
(single ref local, bucket-major seeding); the compiler fix is its own
slice.
Reference-machine numbers (first campaign, 2026-08-21)
bench/baseline.json is the contract; headline readings:
- ram seed 245–290k inserts/s; durable seed ≈4.5k/s (fsync-per-commit ≈220µs each — the gap iteration 23 exists to close).
- reads/queries ≈1.1–1.3M ops/s at p50 1µs since the read-path index
slice (2026-08-22, engine
wo_idx_probe+ emitter index selection) — up from ≈1.5k/s at p50 600µs when point lookups walked every slab (~×850). mixread 89k ops/s single-shard, ~1.9k multi-shard (was 1,280 / 21): the RPC round-trip is now the visible cost, as designed. - msgrate ≈13M msgs/s same-heap vs ≈2.4M cross-shard (the mutex-inbox number, stage-2 deviation 4).
- Tolerance policy lives in the DRIVER (
tolerance_for), not hand-edits — a baseline refresh regenerates it: mix*/sN/read/query 50% (scheduling + µs-scale jitter), rest 15%; latency floorsmax(4×value, 100µs)— the tripwire means "µs became ms".
The gate must bite (proven 2026-08-21)
scripts/db-bench.py --check <results.json> evaluates a recorded run:
the real results pass 74/0; a doctored copy FAILS on exactly the
doctored metrics — use a 15%-class metric (seed) halved plus a
50%-class metric (read) quartered, so both tolerance classes prove they
bite. Re-run the smoke after any gate or policy change.