From 6183a67dfc3c934e19eed5aacc04e367b45726b7 Mon Sep 17 00:00:00 2001 From: "shoney.arickathil" Date: Fri, 28 Aug 2026 09:54:52 +0200 Subject: [PATCH] =?UTF-8?q?perf(db):=20group=20commit=20measured=20?= =?UTF-8?q?=E2=80=94=20~2.9x=20durable=20write=20throughput,=20T5?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit databasev2 4 part A, task 5. Controlled before/after — same machine, same workload (wmix 4000 32), same build except db.c and vm.c, two runs each interleaved: - per-statement barrier: 2213 / 2177 ops/sec, p50 7183 / 7251us - group commit: 6216 / 6525 ops/sec, p50 3458 / 3444us - ~2.9x throughput, ~2.1x lower p50 The full campaign confirms it a second way: s1 takes the inline path and commits per statement BY DESIGN, so within one build the shard configs are batching-off vs batching-on — 1467 -> 5117 ops/sec, mean batch 1.0 -> 5.43, peak 1 -> 57. 3.5x, agreeing with the 2.9x above. Recorded honestly: - the BEFORE p99 is at the histogram ceiling (hist_add clamps at 20000us and both runs pinned there), so the true figure is >=20ms and unknown. The improvement is AT LEAST 2.3x; the old p99 was off the instrument - durable.sN.mixwrite went 480 -> 492 ops/sec, i.e. UNCHANGED. That was the spec's original payoff metric and correcting it was part of the brainstorm: mix performs 20 writes at C=4, mean batch 1.01. A workload that never has two writes in flight cannot be helped by batching them - seed is likewise unchanged: a serial writer has nothing to batch with - so the payoff is real but CONDITIONAL — it appears where concurrent durable writes fan into the owner shard, and nowhere else Two traps recorded in perf-targets §6: - do not benchmark durability on /tmp: it is tmpfs here, where fdatasync is free. The same run reported 195000 ops/sec at p50 1us there against 2200 at p50 7200us on ext4 — no barrier to amortise, so the measurement measures nothing. db-bench keeps its stores under bench/ for this reason - the record count is not the update count: 7755 records for 4000 updates, because hist_dump and the done-marker are themselves durable inserts - FIXED a regression I introduced in T4: master's committed baseline is FULL mode (N=20000, crash_reps=3, msg_n=200000) and I had overwritten it with quick-mode values. Regenerated from a full campaign; the full run now passes 106 checks 0 failures against it - gate still bites: sN wmix ops_sec -70% -> FAIL on exactly that metric Co-Authored-By: Claude Opus 5 (1M context) --- bench/baseline.json | 252 +++++++++++++++++++------------------- docs/plan/perf-targets.md | 71 +++++++++++ 2 files changed, 197 insertions(+), 126 deletions(-) diff --git a/bench/baseline.json b/bench/baseline.json index 4828462..b158229 100644 --- a/bench/baseline.json +++ b/bench/baseline.json @@ -1,52 +1,52 @@ { "_config": { - "N": 2000, - "crash_reps": 1, - "msg_n": 20000, + "N": 20000, + "crash_reps": 3, + "msg_n": 200000, "note": "refresh only with a commit that says why; tolerances come from tolerance_for() in the driver", - "wal_n": 800 + "wal_n": 4000 }, "durable.s1.mixread.ops_sec": { "dir": "higher", - "floor": 2240, + "floor": 2070, "tolerance_pct": 50, - "value": 8961 + "value": 8281 }, "durable.s1.mixread.p50us": { "dir": "lower", "floor": 100, "tolerance_pct": 50, - "value": 1 + "value": 2 }, "durable.s1.mixread.p99us": { "dir": "lower", "floor": 100, "tolerance_pct": 50, - "value": 4 + "value": 21 }, "durable.s1.mixwrite.ops_sec": { "dir": "higher", - "floor": 248, + "floor": 230, "tolerance_pct": 50, - "value": 995 + "value": 920 }, "durable.s1.mixwrite.p50us": { "dir": "lower", - "floor": 824, + "floor": 1784, "tolerance_pct": 50, - "value": 206 + "value": 446 }, "durable.s1.mixwrite.p99us": { "dir": "lower", - "floor": 904, + "floor": 2252, "tolerance_pct": 50, - "value": 226 + "value": 563 }, "durable.s1.query.ops_sec": { "dir": "higher", - "floor": 233644, + "floor": 210260, "tolerance_pct": 50, - "value": 934579 + "value": 841042 }, "durable.s1.query.p50us": { "dir": "lower", @@ -58,13 +58,13 @@ "dir": "lower", "floor": 100, "tolerance_pct": 50, - "value": 3 + "value": 4 }, "durable.s1.read.ops_sec": { "dir": "higher", - "floor": 190114, + "floor": 284155, "tolerance_pct": 50, - "value": 760456 + "value": 1136621 }, "durable.s1.read.p50us": { "dir": "lower", @@ -80,21 +80,21 @@ }, "durable.s1.seed.ops_sec": { "dir": "higher", - "floor": 1080, + "floor": 1076, "tolerance_pct": 15, - "value": 4323 + "value": 4306 }, "durable.s1.seed.p50us": { "dir": "lower", - "floor": 848, + "floor": 868, "tolerance_pct": 15, - "value": 212 + "value": 217 }, "durable.s1.seed.p99us": { "dir": "lower", - "floor": 2108, + "floor": 2080, "tolerance_pct": 15, - "value": 527 + "value": 520 }, "durable.s1.wmix.mean_batch": { "dir": "higher", @@ -104,21 +104,21 @@ }, "durable.s1.wmix.ops_sec": { "dir": "higher", - "floor": 810, + "floor": 366, "tolerance_pct": 15, - "value": 3241 + "value": 1467 }, "durable.s1.wmix.p50us": { "dir": "lower", - "floor": 852, + "floor": 1820, "tolerance_pct": 15, - "value": 213 + "value": 455 }, "durable.s1.wmix.p99us": { "dir": "lower", - "floor": 2032, + "floor": 2884, "tolerance_pct": 15, - "value": 508 + "value": 721 }, "durable.s1.wmix.peak_batch": { "dir": "higher", @@ -134,63 +134,63 @@ }, "durable.s1.write.ops_sec": { "dir": "higher", - "floor": 1103, + "floor": 554, "tolerance_pct": 15, - "value": 4415 + "value": 2216 }, "durable.s1.write.p50us": { "dir": "lower", - "floor": 856, + "floor": 1828, "tolerance_pct": 15, - "value": 214 + "value": 457 }, "durable.s1.write.p99us": { "dir": "lower", - "floor": 1932, + "floor": 2880, "tolerance_pct": 15, - "value": 483 + "value": 720 }, "durable.sN.mixread.ops_sec": { "dir": "higher", - "floor": 1111, + "floor": 1107, "tolerance_pct": 50, - "value": 4445 + "value": 4429 }, "durable.sN.mixread.p50us": { "dir": "lower", - "floor": 348, + "floor": 308, "tolerance_pct": 50, - "value": 87 + "value": 77 }, "durable.sN.mixread.p99us": { "dir": "lower", - "floor": 17540, + "floor": 4172, "tolerance_pct": 50, - "value": 4385 + "value": 1043 }, "durable.sN.mixwrite.ops_sec": { "dir": "higher", "floor": 123, "tolerance_pct": 50, - "value": 493 + "value": 492 }, "durable.sN.mixwrite.p50us": { "dir": "lower", - "floor": 1460, + "floor": 2156, "tolerance_pct": 50, - "value": 365 + "value": 539 }, "durable.sN.mixwrite.p99us": { "dir": "lower", - "floor": 3888, + "floor": 6952, "tolerance_pct": 50, - "value": 972 + "value": 1738 }, "durable.sN.query.ops_sec": { "dir": "higher", - "floor": 211864, + "floor": 237529, "tolerance_pct": 50, - "value": 847457 + "value": 950118 }, "durable.sN.query.p50us": { "dir": "lower", @@ -202,13 +202,13 @@ "dir": "lower", "floor": 100, "tolerance_pct": 50, - "value": 2 + "value": 3 }, "durable.sN.read.ops_sec": { "dir": "higher", - "floor": 204415, + "floor": 276701, "tolerance_pct": 50, - "value": 817661 + "value": 1106806 }, "durable.sN.read.p50us": { "dir": "lower", @@ -224,81 +224,81 @@ }, "durable.sN.seed.ops_sec": { "dir": "higher", - "floor": 1111, + "floor": 1085, "tolerance_pct": 50, - "value": 4446 + "value": 4343 }, "durable.sN.seed.p50us": { "dir": "lower", - "floor": 848, + "floor": 868, "tolerance_pct": 50, - "value": 212 + "value": 217 }, "durable.sN.seed.p99us": { "dir": "lower", - "floor": 1996, + "floor": 2040, "tolerance_pct": 50, - "value": 499 + "value": 510 }, "durable.sN.wmix.mean_batch": { "dir": "higher", - "floor": 0.0, + "floor": 1.0, "tolerance_pct": 100, - "value": 3.38 + "value": 5.43 }, "durable.sN.wmix.ops_sec": { "dir": "higher", - "floor": 1575, + "floor": 1279, "tolerance_pct": 50, - "value": 6302 + "value": 5117 }, "durable.sN.wmix.p50us": { "dir": "lower", - "floor": 13776, + "floor": 32832, "tolerance_pct": 50, - "value": 3444 + "value": 8208 }, "durable.sN.wmix.p99us": { "dir": "lower", - "floor": 51212, + "floor": 48676, "tolerance_pct": 50, - "value": 12803 + "value": 12169 }, "durable.sN.wmix.peak_batch": { "dir": "higher", - "floor": 7, + "floor": 14, "tolerance_pct": 100, - "value": 28 + "value": 57 }, "durable.sN.wmix.peak_staged": { "dir": "lower", - "floor": 5488, + "floor": 11172, "tolerance_pct": 100, - "value": 1372 + "value": 2793 }, "durable.sN.write.ops_sec": { "dir": "higher", - "floor": 1117, + "floor": 552, "tolerance_pct": 50, - "value": 4470 + "value": 2210 }, "durable.sN.write.p50us": { "dir": "lower", - "floor": 852, + "floor": 1828, "tolerance_pct": 50, - "value": 213 + "value": 457 }, "durable.sN.write.p99us": { "dir": "lower", - "floor": 2020, + "floor": 2948, "tolerance_pct": 50, - "value": 505 + "value": 737 }, "ram.s1.mixread.ops_sec": { "dir": "higher", - "floor": 2235, + "floor": 22322, "tolerance_pct": 50, - "value": 8940 + "value": 89290 }, "ram.s1.mixread.p50us": { "dir": "lower", @@ -314,9 +314,9 @@ }, "ram.s1.mixwrite.ops_sec": { "dir": "higher", - "floor": 248, + "floor": 2480, "tolerance_pct": 50, - "value": 993 + "value": 9921 }, "ram.s1.mixwrite.p50us": { "dir": "lower", @@ -328,19 +328,19 @@ "dir": "lower", "floor": 100, "tolerance_pct": 50, - "value": 1 + "value": 2 }, "ram.s1.msgrate.msgs_sec": { "dir": "higher", - "floor": 384911, + "floor": 1568873, "tolerance_pct": 15, - "value": 3079291 + "value": 12550988 }, "ram.s1.query.ops_sec": { "dir": "higher", - "floor": 274725, + "floor": 256016, "tolerance_pct": 50, - "value": 1098901 + "value": 1024065 }, "ram.s1.query.p50us": { "dir": "lower", @@ -352,13 +352,13 @@ "dir": "lower", "floor": 100, "tolerance_pct": 50, - "value": 1 + "value": 2 }, "ram.s1.read.ops_sec": { "dir": "higher", - "floor": 312500, + "floor": 233448, "tolerance_pct": 50, - "value": 1250000 + "value": 933794 }, "ram.s1.read.p50us": { "dir": "lower", @@ -370,91 +370,91 @@ "dir": "lower", "floor": 100, "tolerance_pct": 50, - "value": 1 + "value": 2 }, "ram.s1.seed.ops_sec": { "dir": "higher", - "floor": 335570, + "floor": 53529, "tolerance_pct": 15, - "value": 1342281 + "value": 214119 }, "ram.s1.seed.p50us": { "dir": "lower", "floor": 100, "tolerance_pct": 15, - "value": 1 + "value": 4 }, "ram.s1.seed.p99us": { "dir": "lower", "floor": 100, "tolerance_pct": 15, - "value": 2 + "value": 12 }, "ram.s1.write.ops_sec": { "dir": "higher", - "floor": 237416, + "floor": 48021, "tolerance_pct": 15, - "value": 949667 + "value": 192086 }, "ram.s1.write.p50us": { "dir": "lower", "floor": 100, "tolerance_pct": 15, - "value": 1 + "value": 6 }, "ram.s1.write.p99us": { "dir": "lower", "floor": 100, "tolerance_pct": 15, - "value": 2 + "value": 15 }, "ram.sN.mixread.ops_sec": { "dir": "higher", - "floor": 2243, + "floor": 7481, "tolerance_pct": 50, - "value": 8975 + "value": 29924 }, "ram.sN.mixread.p50us": { "dir": "lower", - "floor": 264, + "floor": 320, "tolerance_pct": 50, - "value": 66 + "value": 80 }, "ram.sN.mixread.p99us": { "dir": "lower", - "floor": 2080, + "floor": 624, "tolerance_pct": 50, - "value": 520 + "value": 156 }, "ram.sN.mixwrite.ops_sec": { "dir": "higher", - "floor": 249, + "floor": 831, "tolerance_pct": 50, - "value": 997 + "value": 3324 }, "ram.sN.mixwrite.p50us": { "dir": "lower", - "floor": 288, + "floor": 348, "tolerance_pct": 50, - "value": 72 + "value": 87 }, "ram.sN.mixwrite.p99us": { "dir": "lower", - "floor": 344, + "floor": 608, "tolerance_pct": 50, - "value": 86 + "value": 152 }, "ram.sN.msgrate.msgs_sec": { "dir": "higher", - "floor": 176056, + "floor": 248897, "tolerance_pct": 50, - "value": 1408450 + "value": 1991179 }, "ram.sN.query.ops_sec": { "dir": "higher", - "floor": 308641, + "floor": 227790, "tolerance_pct": 50, - "value": 1234567 + "value": 911161 }, "ram.sN.query.p50us": { "dir": "lower", @@ -466,13 +466,13 @@ "dir": "lower", "floor": 100, "tolerance_pct": 50, - "value": 1 + "value": 2 }, "ram.sN.read.ops_sec": { "dir": "higher", - "floor": 310945, + "floor": 216919, "tolerance_pct": 50, - "value": 1243781 + "value": 867678 }, "ram.sN.read.p50us": { "dir": "lower", @@ -484,42 +484,42 @@ "dir": "lower", "floor": 100, "tolerance_pct": 50, - "value": 1 + "value": 2 }, "ram.sN.seed.ops_sec": { "dir": "higher", - "floor": 363372, + "floor": 60469, "tolerance_pct": 50, - "value": 1453488 + "value": 241878 }, "ram.sN.seed.p50us": { "dir": "lower", "floor": 100, "tolerance_pct": 50, - "value": 1 + "value": 4 }, "ram.sN.seed.p99us": { "dir": "lower", "floor": 100, "tolerance_pct": 50, - "value": 2 + "value": 10 }, "ram.sN.write.ops_sec": { "dir": "higher", - "floor": 202922, + "floor": 41677, "tolerance_pct": 50, - "value": 811688 + "value": 166708 }, "ram.sN.write.p50us": { "dir": "lower", "floor": 100, "tolerance_pct": 50, - "value": 1 + "value": 8 }, "ram.sN.write.p99us": { "dir": "lower", "floor": 100, "tolerance_pct": 50, - "value": 4 + "value": 15 } } \ No newline at end of file diff --git a/docs/plan/perf-targets.md b/docs/plan/perf-targets.md index 9d12aea..386c6e8 100644 --- a/docs/plan/perf-targets.md +++ b/docs/plan/perf-targets.md @@ -67,3 +67,74 @@ not the limiting factor for any current workload). the ~55× gap is one fdatasync per statement (~220µs each). **Owner: iteration 23** (io_uring group-commit) — its acceptance is literally this number moving while the crash battery stays green. + +## 6. WAL group commit: one barrier per drain (databasev2 4 part A) + +**Measured 2026-08-28.** Before this, the engine committed per *statement*: +`db.c` called `wo_wal_commit` immediately after every append, so each row +change bought its own `pwrite` + `fdatasync`. Now shard 0 stages every queued +write request, issues one barrier, and only then releases the held replies. + +### The controlled before/after + +Same machine, same workload (`wmix 4000 32` — every op a durable update, 32 +concurrent), same build except `db.c` and `vm.c`, two runs each, interleaved: + +| | ops/sec | p50 | p99 | +| --- | --- | --- | --- | +| per-statement barrier | 2213 · 2177 | 7183 · 7251 µs | **20000 · 20000 µs** | +| group commit | **6216 · 6525** | **3458 · 3444 µs** | 11139 · 5971 µs | + +**≈2.9× throughput, ≈2.1× lower p50.** + +**The p99 "before" figure is at the histogram ceiling, not a measurement.** +`hist_add` clamps at 20000 µs, and both before-runs pinned there — so the true +before p99 is ≥20 ms and unknown. The improvement is *at least* 2.3×; the +honest statement is that the old p99 was off the end of the instrument. + +### Confirmation from the committed baseline + +The full campaign gives the same answer a second way. `s1` takes the inline +path, which commits per statement **by design**, so within one build the two +shard configurations are batching-off against batching-on: + +| Leg | ops/sec | p50 | p99 | mean batch | peak batch | +| --- | --- | --- | --- | --- | --- | +| `durable.s1.wmix` (inline, unbatched) | 1467 | 455 µs | 721 µs | **1.0** | 1 | +| `durable.sN.wmix` (batched) | **5117** | 8208 µs | 12169 µs | **5.43** | 57 | + +3.5× throughput, agreeing with the 2.9× above. Note `sN` latency is *higher* +while throughput is 3.5× better: 64 writers queueing behind one owner shard +trade per-op latency for barrier amortisation, which is what group commit is. + +Batching scales with write concurrency exactly as designed — mean batch at +C = 4 / 16 / 64 was **1.13 / 1.76 / 5.35**, peak **3 / 10 / 39**. + +### What did NOT improve, and why that was predicted + +`durable.sN.mixwrite` went **480 → 492 ops/s** — unchanged. That is the metric +the spec *originally* named as the payoff, and correcting it was part of the +brainstorm: `mix` writes on one op in ten with C=4, so a quick run performs +**20 writes** and mean batch measured **1.01** over 3112 barriers. A workload +that never has two writes in flight cannot be helped by batching them. +`durable.*.seed` is likewise unchanged: a serial single writer has nothing to +batch with under any scheme. + +**So the payoff is real but conditional: it appears exactly where concurrent +durable writes fan into the owner shard, and nowhere else.** + +### Two traps worth recording + +**Do not benchmark durability on `/tmp`.** It is `tmpfs` here, where +`fdatasync` is free — the same `wmix` run reported **195 000 ops/s at p50 1 µs** +there against **2200 ops/s at p50 7200 µs** on ext4. There is no barrier to +amortise on a memory filesystem, so a group-commit measurement taken there +measures nothing. `db-bench` gets this right by keeping its stores under +`bench/`. + +**The record count is not the update count.** `wmix` staged 7755 records for +4000 updates because the histogram dump and the done-marker are themselves +durable inserts. They arrive as an end-of-run burst, which is batch-friendly, +so `mean_batch` is not purely update-driven. Peak staged bytes stayed small +(2793 B at C=64), which is what settled the decision to ship **no batch cap**: +the request queue's existing upstream bound is sufficient.