TESTS DELIBERATELY HELD at the developer's instruction — logic only. The existing suite passes (36 suites, 0 fail) but exercises NEITHER new behaviour: nothing builds a 16-deep chain, and no checkpoint test uses a log near 64 MiB. Green here means "did not break what existed". - tier 1: wo_wal_fold_row_at gains hops_out. The walk already visits every hop, so the depth is free — this is the design's pd_prune_xid, a cheap "is work worth doing" hint taken from work already happening - the update path branches on it: past WO_DELTA_MAX_HOPS (16) it writes a full-row image instead of a delta, terminating the chain. `r` already holds the complete post-update row because index maintenance required folding it, so flattening costs bytes, not an extra read - wo_wal_append_row_image encodes from a caller-held row, as WO_WAL_INSERT: a chain's base must replay into a database where nothing precedes it, so replay/compaction/fold need no change - tier 2: should_compact gains a TRIGGERING absolute term and a ceiling. Our `floor` SUPPRESSES on a small log — the opposite of postgres's vac_base_thresh, which triggers on a small absolute problem the proportion hides. We had the proportion and the suppressor and neither real guard - verified by construction, not test: both update entry points converge on row_apply_field_keys; db.c captures next_offset BEFORE calling in, so the re-point is transparent to which record type was written Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> (cherry picked from commit 1b808abd5942de81c3a6416714d1302384103040)
148 lines
7.7 KiB
Markdown
148 lines
7.7 KiB
Markdown
---
|
|
track: databasev2
|
|
iteration: "11"
|
|
status: in-progress
|
|
readiness: ready
|
|
---
|
|
|
|
# databasev2 11 — bounding a keys-resident row's delta chain
|
|
|
|
> Part of [Story — databasev2: the database beyond RAM](00-story.md).
|
|
> Spec: [`2026-08-30-bounded-delta-chains-design.md`](../../superpowers/specs/2026-08-30-bounded-delta-chains-design.md).
|
|
>
|
|
> **Why this exists.** [Iteration 2](02-table-storage-modes.md) shipped delta
|
|
> updates with a deliberate decision not to cap chain length, on the reasoning
|
|
> that compaction bounds it. A whole-branch review showed that reasoning does
|
|
> not hold for the one workload the feature is motivated by. This iteration
|
|
> closes it. `readiness: ready` — every fork is settled in the spec.
|
|
|
|
## The finding this iteration answers
|
|
|
|
A read of a keys-resident row costs `1 + chain length` preads, and replay costs
|
|
**O(N²)** per chain. The shipped mitigation is compaction, which flattens chains
|
|
to zero. But `wo_wal_should_compact` triggers on `used > last * ratio` — a byte
|
|
ratio over the whole log — and cannot see that one row has a very long chain.
|
|
|
|
One popular SKU whose stock moves on every order, in a catalogue that is
|
|
otherwise quiet, grows an unbounded chain without ever moving that ratio. The
|
|
guard that bounds replay in general is structurally blind to the single case
|
|
that makes replay quadratic.
|
|
|
|
## The design, in one sentence each
|
|
|
|
**Tier 1 — flatten on update.** The update path already folds the row, because
|
|
it needs the old values for index maintenance, and the fold already walks the
|
|
chain hop by hop; so it reports the depth for free, and when that depth reaches
|
|
**K** the update appends a full-row record instead of a delta. Read cost becomes
|
|
at most `K + 1` reads and replay `O(K²)` per row, independent of when a
|
|
checkpoint fires.
|
|
|
|
**Tier 2 — give the compaction policy an absolute term and a ceiling.** Our
|
|
current policy has a proportional term and a *suppressor* misleadingly named a
|
|
floor; it lacks the triggering floor and the ceiling that keep a size-based
|
|
policy honest.
|
|
|
|
## Where the design came from
|
|
|
|
Read from PostgreSQL's source at `.dev/reference/postgresql`, not recalled:
|
|
|
|
- **`heap_page_prune_opt`** collapses HOT chains opportunistically, on a page the
|
|
process already holds, gated by an O(1) on-page hint and then by page fullness
|
|
against `Max(fillfactor, BLCKSZ/10)`. Tier 1 is this shape: do the work while
|
|
you already hold the thing, using a signal you already computed.
|
|
- **autovacuum** thresholds on
|
|
`vac_base_thresh + vac_scale_factor * reltuples`, clamped by a maximum —
|
|
defaults 50, 0.2, 100 000 000. A count with a floor, a proportion and a
|
|
ceiling, per table. Tier 2 borrows the floor and the ceiling.
|
|
- **Postgres never thresholds on new-bytes-versus-old-bytes**, despite knowing
|
|
exactly what a chain costs. Its space test is "will the next version fit" — an
|
|
operational constraint, not an economic comparison. That ruled out the
|
|
byte-ratio shape here too.
|
|
|
|
## Progress
|
|
|
|
| Part | State |
|
|
| --- | --- |
|
|
| Tier 1 — the fold reports hop count | ✅ `wo_wal_fold_row_at` takes `hops_out`; the walk already visited each hop, so it costs nothing |
|
|
| Tier 1 — the update branches on depth | ✅ `row_apply_field_keys` writes a full-row image past `WO_DELTA_MAX_HOPS` (16) instead of a delta |
|
|
| Tier 1 — the chain-terminating write | ✅ `wo_wal_append_row_image`, encoded as `WO_WAL_INSERT` so replay, compaction and the fold need no change |
|
|
| Tier 2 — absolute garbage term | ✅ `WO_CKPT_ABS_BYTES` (64 MiB) triggers regardless of proportion |
|
|
| Tier 2 — proportional ceiling | ✅ `WO_CKPT_MAX_GARBAGE` (256 MiB) caps the ratio term |
|
|
| **Tests** | ⏸ **DELIBERATELY HELD** — see below |
|
|
|
|
**Verified by construction, not by test.** Both update entry points converge on
|
|
`row_apply_field_keys` (`table.c:1039` and `:1319`), so one branch covers both.
|
|
The re-point is transparent to flattening because `db.c` captures
|
|
`wo_wal_next_offset(w)` *before* calling into `table.c` — it targets wherever
|
|
the next record lands, delta or full row alike. And a fold that reaches a
|
|
flattened record terminates there, so the next update sees depth 0.
|
|
|
|
**What holding the tests costs, stated plainly.** The existing suite passes
|
|
(36 suites, 0 failures) but that proves only that threading `hops_out` through
|
|
the fold, `keys_fold_into` and their callers broke nothing — which is the change
|
|
most likely to break something silently, so it is worth having. It does **not**
|
|
exercise either new behaviour:
|
|
|
|
- No existing test builds a chain 16 deep, so the flatten branch is almost
|
|
certainly never executed by the suite.
|
|
- Existing checkpoint tests use logs far below 64 MiB, so the two new
|
|
compaction terms never fire either.
|
|
|
|
A green run here means "did not break what existed", not "works".
|
|
|
|
## Acceptance Criteria
|
|
|
|
Outstanding — none verified, because the tests are held. The logic for every
|
|
one of them is implemented; nothing is proven.
|
|
|
|
- **Given** a row updated K times, **when** updated once more, **then** the
|
|
record its offset names is a full row and its chain length is zero.
|
|
- **Given** a row updated far more than K times, **when** it is read, **then** it
|
|
performs at most K + 1 record reads, asserted by counting rather than timing.
|
|
- **Given** the same row, **when** the process restarts, **then** replay is
|
|
correct and its cost does not grow with the total updates ever applied.
|
|
- **Given** a flattening update, **when** replayed, **then** the row matches the
|
|
same row in a `resident: all` table under the same update sequence — the
|
|
resident table is the oracle.
|
|
- **Given** a flattening update to an indexed column, **when** queried through
|
|
that index, **then** the row is found by its new value and not its old, before
|
|
and after a restart.
|
|
- **Given** reclaimable bytes past the absolute threshold but inside the ratio,
|
|
**when** the policy is evaluated, **then** compaction fires. *(Tier 2.)*
|
|
- **Given** a `resident: all` table, **when** any of this runs, **then** nothing
|
|
about its behaviour or log records changes.
|
|
|
|
## Out Of Scope
|
|
|
|
- **Varying K by row width.** Hop count is what bounds read and replay cost;
|
|
width would optimise only write amplification. Revisit with a measurement, not
|
|
before.
|
|
- **A time-based compaction trigger.** Records are durable at commit, so an idle
|
|
log does not grow.
|
|
- **Whether `resident: keys` earns its place at all.** That is iteration 2's
|
|
task 7, and it should arguably run *before* this work — see below.
|
|
|
|
## Info — the forks, settled
|
|
|
|
1. **Where to fix it: the update path, not the checkpoint.** Making compaction
|
|
depth-aware would mean one hot row triggering a stop-the-world rewrite of the
|
|
entire log — a 2 651 µs pause that scales with total live rows, not with the
|
|
row that misbehaved. Postgres reaches for the local, opportunistic fix first
|
|
for the same reason.
|
|
2. **The metric is hop count, not bytes.** Each hop is one `pread` whose cost
|
|
barely varies with the bytes it carries, so hops are what our read cost is
|
|
made of. Bytes govern write amplification, which is the secondary concern.
|
|
3. **K is a fixed constant and does not scale with table size.** Postgres scales
|
|
by `reltuples` because it thresholds a table-level aggregate with
|
|
proportional harm. Ours is per-row with additive cost — reading one product
|
|
costs the same whether the catalogue holds a hundred rows or ten million, and
|
|
total replay is the sum across rows. Scaling K up with size would make the
|
|
largest databases boot worst.
|
|
|
|
## Sequencing note
|
|
|
|
This iteration is **ready but arguably should not be next**. Iteration 2's
|
|
task 7 has still never measured whether `resident: keys` beats the kernel's own
|
|
paging, and everything built on it — including this — assumes it does. If that
|
|
measurement comes back poorly, this work is optimising something that should be
|
|
deleted. Recommended order: measure first, then this.
|