perf(db): group commit measured — ~2.9x durable write throughput, T5

databasev2 4 part A, task 5.

Controlled before/after — same machine, same workload (wmix 4000 32),
same build except db.c and vm.c, two runs each interleaved:

- per-statement barrier: 2213 / 2177 ops/sec, p50 7183 / 7251us
- group commit:          6216 / 6525 ops/sec, p50 3458 / 3444us
- ~2.9x throughput, ~2.1x lower p50

The full campaign confirms it a second way: s1 takes the inline path and
commits per statement BY DESIGN, so within one build the shard configs are
batching-off vs batching-on — 1467 -> 5117 ops/sec, mean batch 1.0 -> 5.43,
peak 1 -> 57. 3.5x, agreeing with the 2.9x above.

Recorded honestly:

- the BEFORE p99 is at the histogram ceiling (hist_add clamps at 20000us
  and both runs pinned there), so the true figure is >=20ms and unknown.
  The improvement is AT LEAST 2.3x; the old p99 was off the instrument
- durable.sN.mixwrite went 480 -> 492 ops/sec, i.e. UNCHANGED. That was
  the spec's original payoff metric and correcting it was part of the
  brainstorm: mix performs 20 writes at C=4, mean batch 1.01. A workload
  that never has two writes in flight cannot be helped by batching them
- seed is likewise unchanged: a serial writer has nothing to batch with
- so the payoff is real but CONDITIONAL — it appears where concurrent
  durable writes fan into the owner shard, and nowhere else

Two traps recorded in perf-targets §6:

- do not benchmark durability on /tmp: it is tmpfs here, where fdatasync
  is free. The same run reported 195000 ops/sec at p50 1us there against
  2200 at p50 7200us on ext4 — no barrier to amortise, so the measurement
  measures nothing. db-bench keeps its stores under bench/ for this reason
- the record count is not the update count: 7755 records for 4000 updates,
  because hist_dump and the done-marker are themselves durable inserts

- FIXED a regression I introduced in T4: master's committed baseline is
  FULL mode (N=20000, crash_reps=3, msg_n=200000) and I had overwritten it
  with quick-mode values. Regenerated from a full campaign; the full run
  now passes 106 checks 0 failures against it
- gate still bites: sN wmix ops_sec -70% -> FAIL on exactly that metric

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
shoney.arickathil 2026-08-28 09:54:52 +02:00
parent 5f9598af6a
commit 6183a67dfc
2 changed files with 197 additions and 126 deletions

View file

@ -1,52 +1,52 @@
{
"_config": {
"N": 2000,
"crash_reps": 1,
"msg_n": 20000,
"N": 20000,
"crash_reps": 3,
"msg_n": 200000,
"note": "refresh only with a commit that says why; tolerances come from tolerance_for() in the driver",
"wal_n": 800
"wal_n": 4000
},
"durable.s1.mixread.ops_sec": {
"dir": "higher",
"floor": 2240,
"floor": 2070,
"tolerance_pct": 50,
"value": 8961
"value": 8281
},
"durable.s1.mixread.p50us": {
"dir": "lower",
"floor": 100,
"tolerance_pct": 50,
"value": 1
"value": 2
},
"durable.s1.mixread.p99us": {
"dir": "lower",
"floor": 100,
"tolerance_pct": 50,
"value": 4
"value": 21
},
"durable.s1.mixwrite.ops_sec": {
"dir": "higher",
"floor": 248,
"floor": 230,
"tolerance_pct": 50,
"value": 995
"value": 920
},
"durable.s1.mixwrite.p50us": {
"dir": "lower",
"floor": 824,
"floor": 1784,
"tolerance_pct": 50,
"value": 206
"value": 446
},
"durable.s1.mixwrite.p99us": {
"dir": "lower",
"floor": 904,
"floor": 2252,
"tolerance_pct": 50,
"value": 226
"value": 563
},
"durable.s1.query.ops_sec": {
"dir": "higher",
"floor": 233644,
"floor": 210260,
"tolerance_pct": 50,
"value": 934579
"value": 841042
},
"durable.s1.query.p50us": {
"dir": "lower",
@ -58,13 +58,13 @@
"dir": "lower",
"floor": 100,
"tolerance_pct": 50,
"value": 3
"value": 4
},
"durable.s1.read.ops_sec": {
"dir": "higher",
"floor": 190114,
"floor": 284155,
"tolerance_pct": 50,
"value": 760456
"value": 1136621
},
"durable.s1.read.p50us": {
"dir": "lower",
@ -80,21 +80,21 @@
},
"durable.s1.seed.ops_sec": {
"dir": "higher",
"floor": 1080,
"floor": 1076,
"tolerance_pct": 15,
"value": 4323
"value": 4306
},
"durable.s1.seed.p50us": {
"dir": "lower",
"floor": 848,
"floor": 868,
"tolerance_pct": 15,
"value": 212
"value": 217
},
"durable.s1.seed.p99us": {
"dir": "lower",
"floor": 2108,
"floor": 2080,
"tolerance_pct": 15,
"value": 527
"value": 520
},
"durable.s1.wmix.mean_batch": {
"dir": "higher",
@ -104,21 +104,21 @@
},
"durable.s1.wmix.ops_sec": {
"dir": "higher",
"floor": 810,
"floor": 366,
"tolerance_pct": 15,
"value": 3241
"value": 1467
},
"durable.s1.wmix.p50us": {
"dir": "lower",
"floor": 852,
"floor": 1820,
"tolerance_pct": 15,
"value": 213
"value": 455
},
"durable.s1.wmix.p99us": {
"dir": "lower",
"floor": 2032,
"floor": 2884,
"tolerance_pct": 15,
"value": 508
"value": 721
},
"durable.s1.wmix.peak_batch": {
"dir": "higher",
@ -134,63 +134,63 @@
},
"durable.s1.write.ops_sec": {
"dir": "higher",
"floor": 1103,
"floor": 554,
"tolerance_pct": 15,
"value": 4415
"value": 2216
},
"durable.s1.write.p50us": {
"dir": "lower",
"floor": 856,
"floor": 1828,
"tolerance_pct": 15,
"value": 214
"value": 457
},
"durable.s1.write.p99us": {
"dir": "lower",
"floor": 1932,
"floor": 2880,
"tolerance_pct": 15,
"value": 483
"value": 720
},
"durable.sN.mixread.ops_sec": {
"dir": "higher",
"floor": 1111,
"floor": 1107,
"tolerance_pct": 50,
"value": 4445
"value": 4429
},
"durable.sN.mixread.p50us": {
"dir": "lower",
"floor": 348,
"floor": 308,
"tolerance_pct": 50,
"value": 87
"value": 77
},
"durable.sN.mixread.p99us": {
"dir": "lower",
"floor": 17540,
"floor": 4172,
"tolerance_pct": 50,
"value": 4385
"value": 1043
},
"durable.sN.mixwrite.ops_sec": {
"dir": "higher",
"floor": 123,
"tolerance_pct": 50,
"value": 493
"value": 492
},
"durable.sN.mixwrite.p50us": {
"dir": "lower",
"floor": 1460,
"floor": 2156,
"tolerance_pct": 50,
"value": 365
"value": 539
},
"durable.sN.mixwrite.p99us": {
"dir": "lower",
"floor": 3888,
"floor": 6952,
"tolerance_pct": 50,
"value": 972
"value": 1738
},
"durable.sN.query.ops_sec": {
"dir": "higher",
"floor": 211864,
"floor": 237529,
"tolerance_pct": 50,
"value": 847457
"value": 950118
},
"durable.sN.query.p50us": {
"dir": "lower",
@ -202,13 +202,13 @@
"dir": "lower",
"floor": 100,
"tolerance_pct": 50,
"value": 2
"value": 3
},
"durable.sN.read.ops_sec": {
"dir": "higher",
"floor": 204415,
"floor": 276701,
"tolerance_pct": 50,
"value": 817661
"value": 1106806
},
"durable.sN.read.p50us": {
"dir": "lower",
@ -224,81 +224,81 @@
},
"durable.sN.seed.ops_sec": {
"dir": "higher",
"floor": 1111,
"floor": 1085,
"tolerance_pct": 50,
"value": 4446
"value": 4343
},
"durable.sN.seed.p50us": {
"dir": "lower",
"floor": 848,
"floor": 868,
"tolerance_pct": 50,
"value": 212
"value": 217
},
"durable.sN.seed.p99us": {
"dir": "lower",
"floor": 1996,
"floor": 2040,
"tolerance_pct": 50,
"value": 499
"value": 510
},
"durable.sN.wmix.mean_batch": {
"dir": "higher",
"floor": 0.0,
"floor": 1.0,
"tolerance_pct": 100,
"value": 3.38
"value": 5.43
},
"durable.sN.wmix.ops_sec": {
"dir": "higher",
"floor": 1575,
"floor": 1279,
"tolerance_pct": 50,
"value": 6302
"value": 5117
},
"durable.sN.wmix.p50us": {
"dir": "lower",
"floor": 13776,
"floor": 32832,
"tolerance_pct": 50,
"value": 3444
"value": 8208
},
"durable.sN.wmix.p99us": {
"dir": "lower",
"floor": 51212,
"floor": 48676,
"tolerance_pct": 50,
"value": 12803
"value": 12169
},
"durable.sN.wmix.peak_batch": {
"dir": "higher",
"floor": 7,
"floor": 14,
"tolerance_pct": 100,
"value": 28
"value": 57
},
"durable.sN.wmix.peak_staged": {
"dir": "lower",
"floor": 5488,
"floor": 11172,
"tolerance_pct": 100,
"value": 1372
"value": 2793
},
"durable.sN.write.ops_sec": {
"dir": "higher",
"floor": 1117,
"floor": 552,
"tolerance_pct": 50,
"value": 4470
"value": 2210
},
"durable.sN.write.p50us": {
"dir": "lower",
"floor": 852,
"floor": 1828,
"tolerance_pct": 50,
"value": 213
"value": 457
},
"durable.sN.write.p99us": {
"dir": "lower",
"floor": 2020,
"floor": 2948,
"tolerance_pct": 50,
"value": 505
"value": 737
},
"ram.s1.mixread.ops_sec": {
"dir": "higher",
"floor": 2235,
"floor": 22322,
"tolerance_pct": 50,
"value": 8940
"value": 89290
},
"ram.s1.mixread.p50us": {
"dir": "lower",
@ -314,9 +314,9 @@
},
"ram.s1.mixwrite.ops_sec": {
"dir": "higher",
"floor": 248,
"floor": 2480,
"tolerance_pct": 50,
"value": 993
"value": 9921
},
"ram.s1.mixwrite.p50us": {
"dir": "lower",
@ -328,19 +328,19 @@
"dir": "lower",
"floor": 100,
"tolerance_pct": 50,
"value": 1
"value": 2
},
"ram.s1.msgrate.msgs_sec": {
"dir": "higher",
"floor": 384911,
"floor": 1568873,
"tolerance_pct": 15,
"value": 3079291
"value": 12550988
},
"ram.s1.query.ops_sec": {
"dir": "higher",
"floor": 274725,
"floor": 256016,
"tolerance_pct": 50,
"value": 1098901
"value": 1024065
},
"ram.s1.query.p50us": {
"dir": "lower",
@ -352,13 +352,13 @@
"dir": "lower",
"floor": 100,
"tolerance_pct": 50,
"value": 1
"value": 2
},
"ram.s1.read.ops_sec": {
"dir": "higher",
"floor": 312500,
"floor": 233448,
"tolerance_pct": 50,
"value": 1250000
"value": 933794
},
"ram.s1.read.p50us": {
"dir": "lower",
@ -370,91 +370,91 @@
"dir": "lower",
"floor": 100,
"tolerance_pct": 50,
"value": 1
"value": 2
},
"ram.s1.seed.ops_sec": {
"dir": "higher",
"floor": 335570,
"floor": 53529,
"tolerance_pct": 15,
"value": 1342281
"value": 214119
},
"ram.s1.seed.p50us": {
"dir": "lower",
"floor": 100,
"tolerance_pct": 15,
"value": 1
"value": 4
},
"ram.s1.seed.p99us": {
"dir": "lower",
"floor": 100,
"tolerance_pct": 15,
"value": 2
"value": 12
},
"ram.s1.write.ops_sec": {
"dir": "higher",
"floor": 237416,
"floor": 48021,
"tolerance_pct": 15,
"value": 949667
"value": 192086
},
"ram.s1.write.p50us": {
"dir": "lower",
"floor": 100,
"tolerance_pct": 15,
"value": 1
"value": 6
},
"ram.s1.write.p99us": {
"dir": "lower",
"floor": 100,
"tolerance_pct": 15,
"value": 2
"value": 15
},
"ram.sN.mixread.ops_sec": {
"dir": "higher",
"floor": 2243,
"floor": 7481,
"tolerance_pct": 50,
"value": 8975
"value": 29924
},
"ram.sN.mixread.p50us": {
"dir": "lower",
"floor": 264,
"floor": 320,
"tolerance_pct": 50,
"value": 66
"value": 80
},
"ram.sN.mixread.p99us": {
"dir": "lower",
"floor": 2080,
"floor": 624,
"tolerance_pct": 50,
"value": 520
"value": 156
},
"ram.sN.mixwrite.ops_sec": {
"dir": "higher",
"floor": 249,
"floor": 831,
"tolerance_pct": 50,
"value": 997
"value": 3324
},
"ram.sN.mixwrite.p50us": {
"dir": "lower",
"floor": 288,
"floor": 348,
"tolerance_pct": 50,
"value": 72
"value": 87
},
"ram.sN.mixwrite.p99us": {
"dir": "lower",
"floor": 344,
"floor": 608,
"tolerance_pct": 50,
"value": 86
"value": 152
},
"ram.sN.msgrate.msgs_sec": {
"dir": "higher",
"floor": 176056,
"floor": 248897,
"tolerance_pct": 50,
"value": 1408450
"value": 1991179
},
"ram.sN.query.ops_sec": {
"dir": "higher",
"floor": 308641,
"floor": 227790,
"tolerance_pct": 50,
"value": 1234567
"value": 911161
},
"ram.sN.query.p50us": {
"dir": "lower",
@ -466,13 +466,13 @@
"dir": "lower",
"floor": 100,
"tolerance_pct": 50,
"value": 1
"value": 2
},
"ram.sN.read.ops_sec": {
"dir": "higher",
"floor": 310945,
"floor": 216919,
"tolerance_pct": 50,
"value": 1243781
"value": 867678
},
"ram.sN.read.p50us": {
"dir": "lower",
@ -484,42 +484,42 @@
"dir": "lower",
"floor": 100,
"tolerance_pct": 50,
"value": 1
"value": 2
},
"ram.sN.seed.ops_sec": {
"dir": "higher",
"floor": 363372,
"floor": 60469,
"tolerance_pct": 50,
"value": 1453488
"value": 241878
},
"ram.sN.seed.p50us": {
"dir": "lower",
"floor": 100,
"tolerance_pct": 50,
"value": 1
"value": 4
},
"ram.sN.seed.p99us": {
"dir": "lower",
"floor": 100,
"tolerance_pct": 50,
"value": 2
"value": 10
},
"ram.sN.write.ops_sec": {
"dir": "higher",
"floor": 202922,
"floor": 41677,
"tolerance_pct": 50,
"value": 811688
"value": 166708
},
"ram.sN.write.p50us": {
"dir": "lower",
"floor": 100,
"tolerance_pct": 50,
"value": 1
"value": 8
},
"ram.sN.write.p99us": {
"dir": "lower",
"floor": 100,
"tolerance_pct": 50,
"value": 4
"value": 15
}
}

View file

@ -67,3 +67,74 @@ not the limiting factor for any current workload).
the ~55× gap is one fdatasync per statement (~220µs each).
**Owner: iteration 23** (io_uring group-commit) — its acceptance is
literally this number moving while the crash battery stays green.
## 6. WAL group commit: one barrier per drain (databasev2 4 part A)
**Measured 2026-08-28.** Before this, the engine committed per *statement*:
`db.c` called `wo_wal_commit` immediately after every append, so each row
change bought its own `pwrite` + `fdatasync`. Now shard 0 stages every queued
write request, issues one barrier, and only then releases the held replies.
### The controlled before/after
Same machine, same workload (`wmix 4000 32` — every op a durable update, 32
concurrent), same build except `db.c` and `vm.c`, two runs each, interleaved:
| | ops/sec | p50 | p99 |
| --- | --- | --- | --- |
| per-statement barrier | 2213 · 2177 | 7183 · 7251 µs | **20000 · 20000 µs** |
| group commit | **6216 · 6525** | **3458 · 3444 µs** | 11139 · 5971 µs |
**≈2.9× throughput, ≈2.1× lower p50.**
**The p99 "before" figure is at the histogram ceiling, not a measurement.**
`hist_add` clamps at 20000 µs, and both before-runs pinned there — so the true
before p99 is ≥20 ms and unknown. The improvement is *at least* 2.3×; the
honest statement is that the old p99 was off the end of the instrument.
### Confirmation from the committed baseline
The full campaign gives the same answer a second way. `s1` takes the inline
path, which commits per statement **by design**, so within one build the two
shard configurations are batching-off against batching-on:
| Leg | ops/sec | p50 | p99 | mean batch | peak batch |
| --- | --- | --- | --- | --- | --- |
| `durable.s1.wmix` (inline, unbatched) | 1467 | 455 µs | 721 µs | **1.0** | 1 |
| `durable.sN.wmix` (batched) | **5117** | 8208 µs | 12169 µs | **5.43** | 57 |
3.5× throughput, agreeing with the 2.9× above. Note `sN` latency is *higher*
while throughput is 3.5× better: 64 writers queueing behind one owner shard
trade per-op latency for barrier amortisation, which is what group commit is.
Batching scales with write concurrency exactly as designed — mean batch at
C = 4 / 16 / 64 was **1.13 / 1.76 / 5.35**, peak **3 / 10 / 39**.
### What did NOT improve, and why that was predicted
`durable.sN.mixwrite` went **480 → 492 ops/s** — unchanged. That is the metric
the spec *originally* named as the payoff, and correcting it was part of the
brainstorm: `mix` writes on one op in ten with C=4, so a quick run performs
**20 writes** and mean batch measured **1.01** over 3112 barriers. A workload
that never has two writes in flight cannot be helped by batching them.
`durable.*.seed` is likewise unchanged: a serial single writer has nothing to
batch with under any scheme.
**So the payoff is real but conditional: it appears exactly where concurrent
durable writes fan into the owner shard, and nowhere else.**
### Two traps worth recording
**Do not benchmark durability on `/tmp`.** It is `tmpfs` here, where
`fdatasync` is free — the same `wmix` run reported **195 000 ops/s at p50 1 µs**
there against **2200 ops/s at p50 7200 µs** on ext4. There is no barrier to
amortise on a memory filesystem, so a group-commit measurement taken there
measures nothing. `db-bench` gets this right by keeping its stores under
`bench/`.
**The record count is not the update count.** `wmix` staged 7755 records for
4000 updates because the histogram dump and the done-marker are themselves
durable inserts. They arrive as an end-of-run burst, which is batch-friendly,
so `mean_batch` is not purely update-driven. Peak staged bytes stayed small
(2793 B at C=64), which is what settled the decision to ship **no batch cap**:
the request queue's existing upstream bound is sufficient.