writeonce/docs/plan/perf-targets.md
shoney.arickathil 6183a67dfc perf(db): group commit measured — ~2.9x durable write throughput, T5
databasev2 4 part A, task 5.

Controlled before/after — same machine, same workload (wmix 4000 32),
same build except db.c and vm.c, two runs each interleaved:

- per-statement barrier: 2213 / 2177 ops/sec, p50 7183 / 7251us
- group commit:          6216 / 6525 ops/sec, p50 3458 / 3444us
- ~2.9x throughput, ~2.1x lower p50

The full campaign confirms it a second way: s1 takes the inline path and
commits per statement BY DESIGN, so within one build the shard configs are
batching-off vs batching-on — 1467 -> 5117 ops/sec, mean batch 1.0 -> 5.43,
peak 1 -> 57. 3.5x, agreeing with the 2.9x above.

Recorded honestly:

- the BEFORE p99 is at the histogram ceiling (hist_add clamps at 20000us
  and both runs pinned there), so the true figure is >=20ms and unknown.
  The improvement is AT LEAST 2.3x; the old p99 was off the instrument
- durable.sN.mixwrite went 480 -> 492 ops/sec, i.e. UNCHANGED. That was
  the spec's original payoff metric and correcting it was part of the
  brainstorm: mix performs 20 writes at C=4, mean batch 1.01. A workload
  that never has two writes in flight cannot be helped by batching them
- seed is likewise unchanged: a serial writer has nothing to batch with
- so the payoff is real but CONDITIONAL — it appears where concurrent
  durable writes fan into the owner shard, and nowhere else

Two traps recorded in perf-targets §6:

- do not benchmark durability on /tmp: it is tmpfs here, where fdatasync
  is free. The same run reported 195000 ops/sec at p50 1us there against
  2200 at p50 7200us on ext4 — no barrier to amortise, so the measurement
  measures nothing. db-bench keeps its stores under bench/ for this reason
- the record count is not the update count: 7755 records for 4000 updates,
  because hist_dump and the done-marker are themselves durable inserts

- FIXED a regression I introduced in T4: master's committed baseline is
  FULL mode (N=20000, crash_reps=3, msg_n=200000) and I had overwritten it
  with quick-mode values. Regenerated from a full campaign; the full run
  now passes 106 checks 0 failures against it
- gate still bites: sN wmix ops_sec -70% -> FAIL on exactly that metric

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-28 09:54:52 +02:00

140 lines
7 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

# Performance targets — measured, named, waiting
A register like [`discarded.md`](discarded.md)/[`learnings.md`](learnings.md):
optimization candidates that exist because a NUMBER says so, not a
hunch. Every row cites its measurement (the db-bench campaign,
`bench/baseline.json`, or the [go-sqlite comparison](../../bench/compare/go-sqlite/README.md))
and names an owner iteration when one exists. A target leaves this file
by landing (delta recorded in the baseline) or by being rejected into
`discarded.md` with its reason.
## 1. The write path — update-through-query re-probes, insert re-encodes
**Measured 2026-08-22** (go-sqlite comparison, N=20k, same machine,
ext4): ram mixed writes 195,465 ops/s vs SQLite's 380,069 (×1.9
behind); durable mixed writes 2,324 vs 3,257 (×1.4 behind) — while
writeonce WINS durable seed ×1.4 and reads ×2.6–6.4. The write gap is
specifically the UPDATE half of the mix.
Suspected costs, in probable order (attribute before optimizing — the
C-API microbench the 22 spec reserves exists for exactly this):
1. **Update-through-query runs a whole query statement per update**:
probe (now O(1)) + materialize an id `multi` (arena alloc) +
`DB_GET_FIELD`/`DB_UPDATE_FIELD` builtin round-trips per touched
field. SQLite's equivalent is one page write inside one statement.
2. **`wo_row_update_field` walks every index three times** (shadow
unique check, old-entry removal, new-entry add — three
`touches`-loops over `t->indexes` per update; see
`database/src/table.c`).
3. **Insert encodes per field with a malloc per text/owned value**
(`db_val_encode`) — visible as ram seed ×1.2 behind SQLite (245k vs
297k) even though the durable flavor wins.
4. **A WAL update record re-encodes the whole row**
(`wo_wal_append_update` writes the row image, not a delta).
**Owner:** none yet. Sequence note: iteration 23 (io_uring
group-commit) rewrites the durable write path's syscall story anyway —
re-measure after 23 lands, then decide whether the RAM-side costs
(1–3) earn their own slice. Acceptance shape: ram write ops/s closes
on SQLite's number with reads unharmed; baseline refreshed with the
delta recorded.
## 2. Cross-shard DB RPC halves concurrent read throughput
**Measured 2026-08-22** (db-bench campaign): ram mixread 89,538 ops/s
single-shard vs 44,918 at default cores; durable 9,211 vs 4,324. The
DB actor serializes every statement on shard 0 and each op pays an
envelope + park/unpark round-trip.
**Owner: by design, priced deliberately** (story 8's settled decision
1 — rejected alternatives: engine lock, partitioned tables "wait for a
measured need"). THIS is the measured need's first data point; the
recorded escalation path is partitioned/replicated read state, only if
a real workload (iteration 24's chat) hurts. Not actionable before 24.
## 3. The mutex inbox costs ~6× on cross-shard message rate
**Measured 2026-08-22**: 16.7M msgs/s same-heap vs 2.85M cross-shard
(`msgrate`). **Owner: iteration 31** (mailbox/backpressure decisions
consume this number) and stage-2 deviation 4 (lock-free rings arrive
only if the mutex is the measured bottleneck — at 2.85M msgs/s it is
not the limiting factor for any current workload).
## 4. Durable write throughput is fsync-bound at ~4.5k/s
**Measured 2026-08-21**: durable seed 4,460 inserts/s vs ram 245k —
the ~55× gap is one fdatasync per statement (~220µs each).
**Owner: iteration 23** (io_uring group-commit) — its acceptance is
literally this number moving while the crash battery stays green.
## 6. WAL group commit: one barrier per drain (databasev2 4 part A)
**Measured 2026-08-28.** Before this, the engine committed per *statement*:
`db.c` called `wo_wal_commit` immediately after every append, so each row
change bought its own `pwrite` + `fdatasync`. Now shard 0 stages every queued
write request, issues one barrier, and only then releases the held replies.
### The controlled before/after
Same machine, same workload (`wmix 4000 32` — every op a durable update, 32
concurrent), same build except `db.c` and `vm.c`, two runs each, interleaved:
| | ops/sec | p50 | p99 |
| --- | --- | --- | --- |
| per-statement barrier | 2213 · 2177 | 7183 · 7251 µs | **20000 · 20000 µs** |
| group commit | **6216 · 6525** | **3458 · 3444 µs** | 11139 · 5971 µs |
**≈2.9× throughput, ≈2.1× lower p50.**
**The p99 "before" figure is at the histogram ceiling, not a measurement.**
`hist_add` clamps at 20000 µs, and both before-runs pinned there — so the true
before p99 is ≥20 ms and unknown. The improvement is *at least* 2.3×; the
honest statement is that the old p99 was off the end of the instrument.
### Confirmation from the committed baseline
The full campaign gives the same answer a second way. `s1` takes the inline
path, which commits per statement **by design**, so within one build the two
shard configurations are batching-off against batching-on:
| Leg | ops/sec | p50 | p99 | mean batch | peak batch |
| --- | --- | --- | --- | --- | --- |
| `durable.s1.wmix` (inline, unbatched) | 1467 | 455 µs | 721 µs | **1.0** | 1 |
| `durable.sN.wmix` (batched) | **5117** | 8208 µs | 12169 µs | **5.43** | 57 |
3.5× throughput, agreeing with the 2.9× above. Note `sN` latency is *higher*
while throughput is 3.5× better: 64 writers queueing behind one owner shard
trade per-op latency for barrier amortisation, which is what group commit is.
Batching scales with write concurrency exactly as designed — mean batch at
C = 4 / 16 / 64 was **1.13 / 1.76 / 5.35**, peak **3 / 10 / 39**.
### What did NOT improve, and why that was predicted
`durable.sN.mixwrite` went **480 → 492 ops/s** — unchanged. That is the metric
the spec *originally* named as the payoff, and correcting it was part of the
brainstorm: `mix` writes on one op in ten with C=4, so a quick run performs
**20 writes** and mean batch measured **1.01** over 3112 barriers. A workload
that never has two writes in flight cannot be helped by batching them.
`durable.*.seed` is likewise unchanged: a serial single writer has nothing to
batch with under any scheme.
**So the payoff is real but conditional: it appears exactly where concurrent
durable writes fan into the owner shard, and nowhere else.**
### Two traps worth recording
**Do not benchmark durability on `/tmp`.** It is `tmpfs` here, where
`fdatasync` is free — the same `wmix` run reported **195 000 ops/s at p50 1 µs**
there against **2200 ops/s at p50 7200 µs** on ext4. There is no barrier to
amortise on a memory filesystem, so a group-commit measurement taken there
measures nothing. `db-bench` gets this right by keeping its stores under
`bench/`.
**The record count is not the update count.** `wmix` staged 7755 records for
4000 updates because the histogram dump and the done-marker are themselves
durable inserts. They arrive as an end-of-run burst, which is batch-friendly,
so `mean_batch` is not purely update-driven. Peak staged bytes stayed small
(2793 B at C=64), which is what settled the decision to ship **no batch cap**:
the request queue's existing upstream bound is sufficient.