docs(db2-chains): spec + story for bounding a row's delta chain

- fixes a limitation iteration 2 shipped: compaction was supposed to
  bound chain length, but wo_wal_should_compact triggers on a whole-log
  byte ratio and cannot see one hot row's chain
- tier 1, flatten on update: the update path ALREADY folds the row for
  index maintenance and the fold already walks hop by hop, so it reports
  depth for free. Past a fixed K it writes a full row instead of a
  delta. Read <= K+1 reads, replay O(K^2) per row. No format change, no
  per-row RAM, no new trigger
- tier 2: our compaction policy has a proportional term and a
  SUPPRESSOR misleadingly called a floor; postgres's floor TRIGGERS on
  small absolute garbage. Add that term and a ceiling
- design read from .dev/reference/postgresql, not recalled:
  heap_page_prune_opt gates on an O(1) on-page hint then page fullness
  against Max(fillfactor, BLCKSZ/10); autovacuum uses base + scale *
  reltuples clamped by a max (50, 0.2, 1e8). Neither thresholds on
  new-bytes-versus-old-bytes
- K deliberately does NOT scale with table size: postgres scales a
  table-level aggregate with proportional harm, ours is per-row with
  additive cost, so scaling up would make big databases boot worst
- the story says plainly it should NOT be next: task 7 has still never
  measured whether resident: keys beats the kernel's own paging

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
(cherry picked from commit f667cad2cfbe187b5973440ab1af015b2df288f8)
This commit is contained in:
shoney.arickathil 2026-08-30 15:11:42 +02:00
parent b058feb517
commit cfd660a5e6
4 changed files with 294 additions and 0 deletions

View file

@ -160,6 +160,7 @@ before its mechanism existed; the history is in
| 8 | [Query grammar from corpora](08-query-grammar-corpus.md) *(was language 27)* | whole-query `count`, `exists` | independent |
| 9 | [Cross-program tables](09-cross-program-tables.md) *(was language 20)* | attach to a running program's database over local IPC | independent |
| 10 | [Keypair attach auth](10-keypair-attach-auth.md) *(was language 21)* | program identity as a keypair; mutual challenge–response | 9 |
| 11 | [Bounded delta chains](11-bounded-delta-chains.md) | cap a keys-resident row's delta chain in the update path, and give the compaction policy an absolute term + ceiling | 2 (fixes a limitation it shipped) |
```
An arrow points AT the iteration that NEEDS the other.

View file

@ -201,6 +201,13 @@ Met:
not to cap chain length rests on compaction bounding it instead; for
this shape it does not.
**Answered by [iteration 11](11-bounded-delta-chains.md)** (spec written
2026-08-30): the update path already folds the row and the fold already
walks hop by hop, so it reports the depth for free — past a fixed K the
update writes a full row instead of a delta, and the chain resets. Read
cost becomes at most K+1 reads and replay O(K²) per row, independent of
when a checkpoint fires. Limitations 2 and 3 above both fall to it.
Outstanding:
- **Given** `durable: true` and no `WO_DATA`, **when** the program starts,

View file

@ -0,0 +1,116 @@
---
track: databasev2
iteration: "11"
status: pending
readiness: ready
---
# databasev2 11 — bounding a keys-resident row's delta chain
> Part of [Story — databasev2: the database beyond RAM](00-story.md).
> Spec: [`2026-08-30-bounded-delta-chains-design.md`](../../superpowers/specs/2026-08-30-bounded-delta-chains-design.md).
>
> **Why this exists.** [Iteration 2](02-table-storage-modes.md) shipped delta
> updates with a deliberate decision not to cap chain length, on the reasoning
> that compaction bounds it. A whole-branch review showed that reasoning does
> not hold for the one workload the feature is motivated by. This iteration
> closes it. `readiness: ready` — every fork is settled in the spec.
## The finding this iteration answers
A read of a keys-resident row costs `1 + chain length` preads, and replay costs
**O(N²)** per chain. The shipped mitigation is compaction, which flattens chains
to zero. But `wo_wal_should_compact` triggers on `used > last * ratio` — a byte
ratio over the whole log — and cannot see that one row has a very long chain.
One popular SKU whose stock moves on every order, in a catalogue that is
otherwise quiet, grows an unbounded chain without ever moving that ratio. The
guard that bounds replay in general is structurally blind to the single case
that makes replay quadratic.
## The design, in one sentence each
**Tier 1 — flatten on update.** The update path already folds the row, because
it needs the old values for index maintenance, and the fold already walks the
chain hop by hop; so it reports the depth for free, and when that depth reaches
**K** the update appends a full-row record instead of a delta. Read cost becomes
at most `K + 1` reads and replay `O(K²)` per row, independent of when a
checkpoint fires.
**Tier 2 — give the compaction policy an absolute term and a ceiling.** Our
current policy has a proportional term and a *suppressor* misleadingly named a
floor; it lacks the triggering floor and the ceiling that keep a size-based
policy honest.
## Where the design came from
Read from PostgreSQL's source at `.dev/reference/postgresql`, not recalled:
- **`heap_page_prune_opt`** collapses HOT chains opportunistically, on a page the
process already holds, gated by an O(1) on-page hint and then by page fullness
against `Max(fillfactor, BLCKSZ/10)`. Tier 1 is this shape: do the work while
you already hold the thing, using a signal you already computed.
- **autovacuum** thresholds on
`vac_base_thresh + vac_scale_factor * reltuples`, clamped by a maximum —
defaults 50, 0.2, 100 000 000. A count with a floor, a proportion and a
ceiling, per table. Tier 2 borrows the floor and the ceiling.
- **Postgres never thresholds on new-bytes-versus-old-bytes**, despite knowing
exactly what a chain costs. Its space test is "will the next version fit" — an
operational constraint, not an economic comparison. That ruled out the
byte-ratio shape here too.
## Acceptance Criteria
Outstanding — none met; this iteration has not started.
- **Given** a row updated K times, **when** updated once more, **then** the
record its offset names is a full row and its chain length is zero.
- **Given** a row updated far more than K times, **when** it is read, **then** it
performs at most K + 1 record reads, asserted by counting rather than timing.
- **Given** the same row, **when** the process restarts, **then** replay is
correct and its cost does not grow with the total updates ever applied.
- **Given** a flattening update, **when** replayed, **then** the row matches the
same row in a `resident: all` table under the same update sequence — the
resident table is the oracle.
- **Given** a flattening update to an indexed column, **when** queried through
that index, **then** the row is found by its new value and not its old, before
and after a restart.
- **Given** reclaimable bytes past the absolute threshold but inside the ratio,
**when** the policy is evaluated, **then** compaction fires. *(Tier 2.)*
- **Given** a `resident: all` table, **when** any of this runs, **then** nothing
about its behaviour or log records changes.
## Out Of Scope
- **Varying K by row width.** Hop count is what bounds read and replay cost;
width would optimise only write amplification. Revisit with a measurement, not
before.
- **A time-based compaction trigger.** Records are durable at commit, so an idle
log does not grow.
- **Whether `resident: keys` earns its place at all.** That is iteration 2's
task 7, and it should arguably run *before* this work — see below.
## Info — the forks, settled
1. **Where to fix it: the update path, not the checkpoint.** Making compaction
depth-aware would mean one hot row triggering a stop-the-world rewrite of the
entire log — a 2 651 µs pause that scales with total live rows, not with the
row that misbehaved. Postgres reaches for the local, opportunistic fix first
for the same reason.
2. **The metric is hop count, not bytes.** Each hop is one `pread` whose cost
barely varies with the bytes it carries, so hops are what our read cost is
made of. Bytes govern write amplification, which is the secondary concern.
3. **K is a fixed constant and does not scale with table size.** Postgres scales
by `reltuples` because it thresholds a table-level aggregate with
proportional harm. Ours is per-row with additive cost — reading one product
costs the same whether the catalogue holds a hundred rows or ten million, and
total replay is the sum across rows. Scaling K up with size would make the
largest databases boot worst.
## Sequencing note
This iteration is **ready but arguably should not be next**. Iteration 2's
task 7 has still never measured whether `resident: keys` beats the kernel's own
paging, and everything built on it — including this — assumes it does. If that
measurement comes back poorly, this work is optimising something that should be
deleted. Recommended order: measure first, then this.

View file

@ -0,0 +1,170 @@
# Bounding a keys-resident row's delta chain
Design settled 2026-08-30. Fixes the limitation shipped with
[databasev2 2](../../stories/databasev2/02-table-storage-modes.md)'s delta
updates and recorded in
[`2026-08-30-keys-resident-delta-updates-design.md`](2026-08-30-keys-resident-delta-updates-design.md).
## The problem, and why the shipped mitigation does not fire
An update to a `resident: keys` row appends a delta — one field's new value plus
a back-pointer. Reading the row folds the chain backward, so a read costs
`1 + chain length` preads, and replay costs **O(N²)** per chain because it folds
once per delta and each fold walks back to the base.
The shipped design chose not to cap chain length, on the reasoning that
compaction flattens every chain and that deltas grow the log, pulling the next
checkpoint forward. **That reasoning is wrong for the case that matters.**
`wo_wal_should_compact` decides on `used > last * ratio` — a byte ratio over the
whole log. It cannot see that one row has a five-thousand-delta chain. A single
hot row taking many small updates barely moves that ratio in a large database,
so the checkpoint never fires, that row's chain grows without bound, and its
replay cost grows as the square.
The motivating workload is precisely this shape: one popular SKU whose stock
moves on every order while the rest of the catalogue sits still.
## What PostgreSQL does, read from source
Verified against `.dev/reference/postgresql`, not recalled. Postgres solves the
same class of problem — chains of row versions that must be collapsed — with
**two tiers**, and neither is a size ratio.
**Tier 1, `heap_page_prune_opt` in `pruneheap.c`** — opportunistic and local.
Three gates, cheapest first: an O(1) `pd_prune_xid` hint stored on the page; a
visibility test; then `PageIsFull(page) || PageGetHeapFreeSpace(page) < minfree`
where `minfree = Max(fillfactor target, BLCKSZ / 10)`. The work happens on a
page the process **already holds** because it is reading or updating it anyway.
The source is explicit that the check is deliberately approximate — it reads
free space without taking a lock, because "avoiding taking a lock seems more
important than sometimes getting a wrong answer in what is after all just a
heuristic estimate."
**Tier 2, autovacuum** — background and per-table:
`vacthresh = vac_base_thresh + vac_scale_factor * reltuples`, clamped by
`autovacuum_vacuum_max_threshold`. Shipped defaults are 50, 0.2 and 100 000 000.
A **count** with a floor, a proportional term and a ceiling — computed per
table, never per database.
Three lessons, and one correction to our own vocabulary:
- **Do the work while you already hold the thing.** That is the whole of tier 1.
- **The floor exists to catch what the proportion hides.** `base_thresh = 50`
fires on a small table where 20% would not. `Max(…, BLCKSZ/10)` does the same
for space.
- **The ceiling exists so scale does not defer forever.**
- **Our "floor" is the opposite of theirs, despite the name.**
`wo_wal_should_compact` reads `if (used < floor) return 0` — ours *suppresses*
compaction on a small log. Postgres's floor *triggers* cleanup on a small
absolute problem. We have the proportional term and the suppressor; we have
neither the triggering floor nor the ceiling.
Postgres also, notably, does **not** threshold on "new bytes versus old bytes",
despite knowing exactly what every chain costs. Its space check is "will the
next version physically fit", a hard operational constraint, not an economic
comparison. That rules out the byte-ratio shape for us as well — and our cost is
worse suited to it still, since each hop is one `pread` whose cost barely varies
with the bytes it carries.
## Tier 1 — flatten on update
**The update path already folds the row.** It must: it needs the old values to
maintain indexes. And the fold already walks the chain hop by hop. So it can
report how many hops it took, and the update path learns the chain's depth for
free — no new record field, no extra read, no per-row RAM. That reported hop
count is our `pd_prune_xid`: the cheap signal that says whether work is worth
doing, obtained from something we were doing anyway.
The rule is one branch. When the fold reports a depth at or beyond **K**, the
update appends a **full-row record** instead of a delta, and the chain resets to
zero. Otherwise it appends a delta as today.
Consequences:
- A read costs at most **K + 1** preads, always, independent of when a
checkpoint fires.
- Replay costs **O(K²) per row**, bounded rather than unbounded.
- Write cost rises by one row-sized record per K updates — amortised, under
`1/K` extra bytes against today.
- Compaction, replay and the fold are untouched. A full-row record is a shape
all three already handle, because it is what an insert writes.
**K is a fixed constant, not a per-table knob.** Postgres ships `fillfactor` and
`base_thresh` as documented constants that are rarely tuned, and that is the
right precedent: K is a *bound*, not a dial. Anything from 8 to 64 caps the
pathology, and being wrong by a factor of two costs one extra row-write per K
updates.
**K does NOT scale with table size, and that is deliberate.** Postgres scales its
threshold by `reltuples` because it thresholds a table-level aggregate whose harm
is proportional. Ours is a per-row property with additive cost: reading one
product costs `1 + depth` preads whether the catalogue holds a hundred rows or
ten million, and total replay is the sum over every row's chain. Scaling K up
with table size would make the largest databases boot worst — exactly backwards.
**Row width is the one thing that might justify varying K**, since flattening
writes a whole row while a delta writes one field, so the write-amplification
break-even genuinely depends on row size. Deliberately **not** done now: hop
count is what bounds read and replay cost, which are the costs actually hurting,
and width would optimise only the write side. Revisit if measurement shows write
amplification matters.
## Tier 2 — give the checkpoint the trigger shape it is missing
Smaller, and separable from tier 1. Tier 1 bounds one row; tier 2 corrects the
whole-log policy's shape so it stops being blind to absolute garbage.
`wo_wal_should_compact` gains, alongside its existing ratio:
- **An absolute garbage term** — compact when reclaimable bytes exceed an
absolute threshold regardless of ratio. This is postgres's `base_thresh`, and
it is what our current "floor" is not.
- **A ceiling** — cap the proportional term so a very large live set does not
defer compaction indefinitely. This is `autovacuum_vacuum_max_threshold`.
The existing floor keeps its current meaning — do not bother with a tiny log —
but the doc comment must stop calling it a floor in postgres's sense, because it
does the opposite thing.
## Acceptance criteria
- **Given** a keys-resident row updated K times, **when** it is updated once
more, **then** the record its offset names is a full row, not a delta, and its
chain length is zero.
- **Given** a row updated many times more than K, **when** it is read, **then**
the read performs at most K + 1 record reads — asserted by counting, not by
timing.
- **Given** a row updated many times more than K, **when** the process restarts,
**then** replay reconstructs it correctly and its cost does not grow with the
total number of updates ever applied to it.
- **Given** a flattening update, **when** it is replayed, **then** the row is
identical to the same row in a `resident: all` table subjected to the same
update sequence. The resident table is the oracle.
- **Given** a flattening update that changes an indexed column, **when** the row
is queried through that index, **then** it is found by the new value and not
the old — before and after a restart.
- **Given** reclaimable bytes past the absolute threshold but within the ratio,
**when** the policy is evaluated, **then** compaction fires. *(Tier 2.)*
- **Given** a `resident: all` table, **when** anything here runs, **then**
nothing about its behaviour or its log records changes.
## Out of scope
- Varying K by row width. Reasoned above; revisit only with a measurement.
- A time-based compaction trigger. Records are durable at commit, so an idle log
does not grow — the existing design's reasoning still holds.
- The mid-drain stale-read limitation, which a separate fix already closed.
- Whether `resident: keys` is worth having at all. That is
[databasev2 2](../../stories/databasev2/02-table-storage-modes.md)'s task 7,
and this design does not answer it.
## Risks
- **K is a constant chosen without measurement.** The bound is right in shape;
its value is a judgement. The mitigation is that being wrong is cheap and
symmetric — too small costs write amplification, too large costs read latency,
and neither is a correctness failure.
- **Flattening makes one update in K expensive.** A burst of updates to one row
pays a row-sized write on every Kth. Acceptable, and the alternative is an
unbounded read path, but it should be visible in the measurement rather than
discovered in production.