docs(db2-chains): spec + story for bounding a row's delta chain
- fixes a limitation iteration 2 shipped: compaction was supposed to bound chain length, but wo_wal_should_compact triggers on a whole-log byte ratio and cannot see one hot row's chain - tier 1, flatten on update: the update path ALREADY folds the row for index maintenance and the fold already walks hop by hop, so it reports depth for free. Past a fixed K it writes a full row instead of a delta. Read <= K+1 reads, replay O(K^2) per row. No format change, no per-row RAM, no new trigger - tier 2: our compaction policy has a proportional term and a SUPPRESSOR misleadingly called a floor; postgres's floor TRIGGERS on small absolute garbage. Add that term and a ceiling - design read from .dev/reference/postgresql, not recalled: heap_page_prune_opt gates on an O(1) on-page hint then page fullness against Max(fillfactor, BLCKSZ/10); autovacuum uses base + scale * reltuples clamped by a max (50, 0.2, 1e8). Neither thresholds on new-bytes-versus-old-bytes - K deliberately does NOT scale with table size: postgres scales a table-level aggregate with proportional harm, ours is per-row with additive cost, so scaling up would make big databases boot worst - the story says plainly it should NOT be next: task 7 has still never measured whether resident: keys beats the kernel's own paging Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> (cherry picked from commit f667cad2cfbe187b5973440ab1af015b2df288f8)
This commit is contained in:
parent
b058feb517
commit
cfd660a5e6
4 changed files with 294 additions and 0 deletions
|
|
@ -160,6 +160,7 @@ before its mechanism existed; the history is in
|
||||||
| 8 | [Query grammar from corpora](08-query-grammar-corpus.md) *(was language 27)* | whole-query `count`, `exists` | independent |
|
| 8 | [Query grammar from corpora](08-query-grammar-corpus.md) *(was language 27)* | whole-query `count`, `exists` | independent |
|
||||||
| 9 | [Cross-program tables](09-cross-program-tables.md) *(was language 20)* | attach to a running program's database over local IPC | independent |
|
| 9 | [Cross-program tables](09-cross-program-tables.md) *(was language 20)* | attach to a running program's database over local IPC | independent |
|
||||||
| 10 | [Keypair attach auth](10-keypair-attach-auth.md) *(was language 21)* | program identity as a keypair; mutual challenge–response | 9 |
|
| 10 | [Keypair attach auth](10-keypair-attach-auth.md) *(was language 21)* | program identity as a keypair; mutual challenge–response | 9 |
|
||||||
|
| 11 | [Bounded delta chains](11-bounded-delta-chains.md) | cap a keys-resident row's delta chain in the update path, and give the compaction policy an absolute term + ceiling | 2 (fixes a limitation it shipped) |
|
||||||
|
|
||||||
```
|
```
|
||||||
An arrow points AT the iteration that NEEDS the other.
|
An arrow points AT the iteration that NEEDS the other.
|
||||||
|
|
|
||||||
|
|
@ -201,6 +201,13 @@ Met:
|
||||||
not to cap chain length rests on compaction bounding it instead; for
|
not to cap chain length rests on compaction bounding it instead; for
|
||||||
this shape it does not.
|
this shape it does not.
|
||||||
|
|
||||||
|
**Answered by [iteration 11](11-bounded-delta-chains.md)** (spec written
|
||||||
|
2026-08-30): the update path already folds the row and the fold already
|
||||||
|
walks hop by hop, so it reports the depth for free — past a fixed K the
|
||||||
|
update writes a full row instead of a delta, and the chain resets. Read
|
||||||
|
cost becomes at most K+1 reads and replay O(K²) per row, independent of
|
||||||
|
when a checkpoint fires. Limitations 2 and 3 above both fall to it.
|
||||||
|
|
||||||
Outstanding:
|
Outstanding:
|
||||||
|
|
||||||
- **Given** `durable: true` and no `WO_DATA`, **when** the program starts,
|
- **Given** `durable: true` and no `WO_DATA`, **when** the program starts,
|
||||||
|
|
|
||||||
116
docs/stories/databasev2/11-bounded-delta-chains.md
Normal file
116
docs/stories/databasev2/11-bounded-delta-chains.md
Normal file
|
|
@ -0,0 +1,116 @@
|
||||||
|
---
|
||||||
|
track: databasev2
|
||||||
|
iteration: "11"
|
||||||
|
status: pending
|
||||||
|
readiness: ready
|
||||||
|
---
|
||||||
|
|
||||||
|
# databasev2 11 — bounding a keys-resident row's delta chain
|
||||||
|
|
||||||
|
> Part of [Story — databasev2: the database beyond RAM](00-story.md).
|
||||||
|
> Spec: [`2026-08-30-bounded-delta-chains-design.md`](../../superpowers/specs/2026-08-30-bounded-delta-chains-design.md).
|
||||||
|
>
|
||||||
|
> **Why this exists.** [Iteration 2](02-table-storage-modes.md) shipped delta
|
||||||
|
> updates with a deliberate decision not to cap chain length, on the reasoning
|
||||||
|
> that compaction bounds it. A whole-branch review showed that reasoning does
|
||||||
|
> not hold for the one workload the feature is motivated by. This iteration
|
||||||
|
> closes it. `readiness: ready` — every fork is settled in the spec.
|
||||||
|
|
||||||
|
## The finding this iteration answers
|
||||||
|
|
||||||
|
A read of a keys-resident row costs `1 + chain length` preads, and replay costs
|
||||||
|
**O(N²)** per chain. The shipped mitigation is compaction, which flattens chains
|
||||||
|
to zero. But `wo_wal_should_compact` triggers on `used > last * ratio` — a byte
|
||||||
|
ratio over the whole log — and cannot see that one row has a very long chain.
|
||||||
|
|
||||||
|
One popular SKU whose stock moves on every order, in a catalogue that is
|
||||||
|
otherwise quiet, grows an unbounded chain without ever moving that ratio. The
|
||||||
|
guard that bounds replay in general is structurally blind to the single case
|
||||||
|
that makes replay quadratic.
|
||||||
|
|
||||||
|
## The design, in one sentence each
|
||||||
|
|
||||||
|
**Tier 1 — flatten on update.** The update path already folds the row, because
|
||||||
|
it needs the old values for index maintenance, and the fold already walks the
|
||||||
|
chain hop by hop; so it reports the depth for free, and when that depth reaches
|
||||||
|
**K** the update appends a full-row record instead of a delta. Read cost becomes
|
||||||
|
at most `K + 1` reads and replay `O(K²)` per row, independent of when a
|
||||||
|
checkpoint fires.
|
||||||
|
|
||||||
|
**Tier 2 — give the compaction policy an absolute term and a ceiling.** Our
|
||||||
|
current policy has a proportional term and a *suppressor* misleadingly named a
|
||||||
|
floor; it lacks the triggering floor and the ceiling that keep a size-based
|
||||||
|
policy honest.
|
||||||
|
|
||||||
|
## Where the design came from
|
||||||
|
|
||||||
|
Read from PostgreSQL's source at `.dev/reference/postgresql`, not recalled:
|
||||||
|
|
||||||
|
- **`heap_page_prune_opt`** collapses HOT chains opportunistically, on a page the
|
||||||
|
process already holds, gated by an O(1) on-page hint and then by page fullness
|
||||||
|
against `Max(fillfactor, BLCKSZ/10)`. Tier 1 is this shape: do the work while
|
||||||
|
you already hold the thing, using a signal you already computed.
|
||||||
|
- **autovacuum** thresholds on
|
||||||
|
`vac_base_thresh + vac_scale_factor * reltuples`, clamped by a maximum —
|
||||||
|
defaults 50, 0.2, 100 000 000. A count with a floor, a proportion and a
|
||||||
|
ceiling, per table. Tier 2 borrows the floor and the ceiling.
|
||||||
|
- **Postgres never thresholds on new-bytes-versus-old-bytes**, despite knowing
|
||||||
|
exactly what a chain costs. Its space test is "will the next version fit" — an
|
||||||
|
operational constraint, not an economic comparison. That ruled out the
|
||||||
|
byte-ratio shape here too.
|
||||||
|
|
||||||
|
## Acceptance Criteria
|
||||||
|
|
||||||
|
Outstanding — none met; this iteration has not started.
|
||||||
|
|
||||||
|
- **Given** a row updated K times, **when** updated once more, **then** the
|
||||||
|
record its offset names is a full row and its chain length is zero.
|
||||||
|
- **Given** a row updated far more than K times, **when** it is read, **then** it
|
||||||
|
performs at most K + 1 record reads, asserted by counting rather than timing.
|
||||||
|
- **Given** the same row, **when** the process restarts, **then** replay is
|
||||||
|
correct and its cost does not grow with the total updates ever applied.
|
||||||
|
- **Given** a flattening update, **when** replayed, **then** the row matches the
|
||||||
|
same row in a `resident: all` table under the same update sequence — the
|
||||||
|
resident table is the oracle.
|
||||||
|
- **Given** a flattening update to an indexed column, **when** queried through
|
||||||
|
that index, **then** the row is found by its new value and not its old, before
|
||||||
|
and after a restart.
|
||||||
|
- **Given** reclaimable bytes past the absolute threshold but inside the ratio,
|
||||||
|
**when** the policy is evaluated, **then** compaction fires. *(Tier 2.)*
|
||||||
|
- **Given** a `resident: all` table, **when** any of this runs, **then** nothing
|
||||||
|
about its behaviour or log records changes.
|
||||||
|
|
||||||
|
## Out Of Scope
|
||||||
|
|
||||||
|
- **Varying K by row width.** Hop count is what bounds read and replay cost;
|
||||||
|
width would optimise only write amplification. Revisit with a measurement, not
|
||||||
|
before.
|
||||||
|
- **A time-based compaction trigger.** Records are durable at commit, so an idle
|
||||||
|
log does not grow.
|
||||||
|
- **Whether `resident: keys` earns its place at all.** That is iteration 2's
|
||||||
|
task 7, and it should arguably run *before* this work — see below.
|
||||||
|
|
||||||
|
## Info — the forks, settled
|
||||||
|
|
||||||
|
1. **Where to fix it: the update path, not the checkpoint.** Making compaction
|
||||||
|
depth-aware would mean one hot row triggering a stop-the-world rewrite of the
|
||||||
|
entire log — a 2 651 µs pause that scales with total live rows, not with the
|
||||||
|
row that misbehaved. Postgres reaches for the local, opportunistic fix first
|
||||||
|
for the same reason.
|
||||||
|
2. **The metric is hop count, not bytes.** Each hop is one `pread` whose cost
|
||||||
|
barely varies with the bytes it carries, so hops are what our read cost is
|
||||||
|
made of. Bytes govern write amplification, which is the secondary concern.
|
||||||
|
3. **K is a fixed constant and does not scale with table size.** Postgres scales
|
||||||
|
by `reltuples` because it thresholds a table-level aggregate with
|
||||||
|
proportional harm. Ours is per-row with additive cost — reading one product
|
||||||
|
costs the same whether the catalogue holds a hundred rows or ten million, and
|
||||||
|
total replay is the sum across rows. Scaling K up with size would make the
|
||||||
|
largest databases boot worst.
|
||||||
|
|
||||||
|
## Sequencing note
|
||||||
|
|
||||||
|
This iteration is **ready but arguably should not be next**. Iteration 2's
|
||||||
|
task 7 has still never measured whether `resident: keys` beats the kernel's own
|
||||||
|
paging, and everything built on it — including this — assumes it does. If that
|
||||||
|
measurement comes back poorly, this work is optimising something that should be
|
||||||
|
deleted. Recommended order: measure first, then this.
|
||||||
170
docs/superpowers/specs/2026-08-30-bounded-delta-chains-design.md
Normal file
170
docs/superpowers/specs/2026-08-30-bounded-delta-chains-design.md
Normal file
|
|
@ -0,0 +1,170 @@
|
||||||
|
# Bounding a keys-resident row's delta chain
|
||||||
|
|
||||||
|
Design settled 2026-08-30. Fixes the limitation shipped with
|
||||||
|
[databasev2 2](../../stories/databasev2/02-table-storage-modes.md)'s delta
|
||||||
|
updates and recorded in
|
||||||
|
[`2026-08-30-keys-resident-delta-updates-design.md`](2026-08-30-keys-resident-delta-updates-design.md).
|
||||||
|
|
||||||
|
## The problem, and why the shipped mitigation does not fire
|
||||||
|
|
||||||
|
An update to a `resident: keys` row appends a delta — one field's new value plus
|
||||||
|
a back-pointer. Reading the row folds the chain backward, so a read costs
|
||||||
|
`1 + chain length` preads, and replay costs **O(N²)** per chain because it folds
|
||||||
|
once per delta and each fold walks back to the base.
|
||||||
|
|
||||||
|
The shipped design chose not to cap chain length, on the reasoning that
|
||||||
|
compaction flattens every chain and that deltas grow the log, pulling the next
|
||||||
|
checkpoint forward. **That reasoning is wrong for the case that matters.**
|
||||||
|
`wo_wal_should_compact` decides on `used > last * ratio` — a byte ratio over the
|
||||||
|
whole log. It cannot see that one row has a five-thousand-delta chain. A single
|
||||||
|
hot row taking many small updates barely moves that ratio in a large database,
|
||||||
|
so the checkpoint never fires, that row's chain grows without bound, and its
|
||||||
|
replay cost grows as the square.
|
||||||
|
|
||||||
|
The motivating workload is precisely this shape: one popular SKU whose stock
|
||||||
|
moves on every order while the rest of the catalogue sits still.
|
||||||
|
|
||||||
|
## What PostgreSQL does, read from source
|
||||||
|
|
||||||
|
Verified against `.dev/reference/postgresql`, not recalled. Postgres solves the
|
||||||
|
same class of problem — chains of row versions that must be collapsed — with
|
||||||
|
**two tiers**, and neither is a size ratio.
|
||||||
|
|
||||||
|
**Tier 1, `heap_page_prune_opt` in `pruneheap.c`** — opportunistic and local.
|
||||||
|
Three gates, cheapest first: an O(1) `pd_prune_xid` hint stored on the page; a
|
||||||
|
visibility test; then `PageIsFull(page) || PageGetHeapFreeSpace(page) < minfree`
|
||||||
|
where `minfree = Max(fillfactor target, BLCKSZ / 10)`. The work happens on a
|
||||||
|
page the process **already holds** because it is reading or updating it anyway.
|
||||||
|
The source is explicit that the check is deliberately approximate — it reads
|
||||||
|
free space without taking a lock, because "avoiding taking a lock seems more
|
||||||
|
important than sometimes getting a wrong answer in what is after all just a
|
||||||
|
heuristic estimate."
|
||||||
|
|
||||||
|
**Tier 2, autovacuum** — background and per-table:
|
||||||
|
`vacthresh = vac_base_thresh + vac_scale_factor * reltuples`, clamped by
|
||||||
|
`autovacuum_vacuum_max_threshold`. Shipped defaults are 50, 0.2 and 100 000 000.
|
||||||
|
A **count** with a floor, a proportional term and a ceiling — computed per
|
||||||
|
table, never per database.
|
||||||
|
|
||||||
|
Three lessons, and one correction to our own vocabulary:
|
||||||
|
|
||||||
|
- **Do the work while you already hold the thing.** That is the whole of tier 1.
|
||||||
|
- **The floor exists to catch what the proportion hides.** `base_thresh = 50`
|
||||||
|
fires on a small table where 20% would not. `Max(…, BLCKSZ/10)` does the same
|
||||||
|
for space.
|
||||||
|
- **The ceiling exists so scale does not defer forever.**
|
||||||
|
- **Our "floor" is the opposite of theirs, despite the name.**
|
||||||
|
`wo_wal_should_compact` reads `if (used < floor) return 0` — ours *suppresses*
|
||||||
|
compaction on a small log. Postgres's floor *triggers* cleanup on a small
|
||||||
|
absolute problem. We have the proportional term and the suppressor; we have
|
||||||
|
neither the triggering floor nor the ceiling.
|
||||||
|
|
||||||
|
Postgres also, notably, does **not** threshold on "new bytes versus old bytes",
|
||||||
|
despite knowing exactly what every chain costs. Its space check is "will the
|
||||||
|
next version physically fit", a hard operational constraint, not an economic
|
||||||
|
comparison. That rules out the byte-ratio shape for us as well — and our cost is
|
||||||
|
worse suited to it still, since each hop is one `pread` whose cost barely varies
|
||||||
|
with the bytes it carries.
|
||||||
|
|
||||||
|
## Tier 1 — flatten on update
|
||||||
|
|
||||||
|
**The update path already folds the row.** It must: it needs the old values to
|
||||||
|
maintain indexes. And the fold already walks the chain hop by hop. So it can
|
||||||
|
report how many hops it took, and the update path learns the chain's depth for
|
||||||
|
free — no new record field, no extra read, no per-row RAM. That reported hop
|
||||||
|
count is our `pd_prune_xid`: the cheap signal that says whether work is worth
|
||||||
|
doing, obtained from something we were doing anyway.
|
||||||
|
|
||||||
|
The rule is one branch. When the fold reports a depth at or beyond **K**, the
|
||||||
|
update appends a **full-row record** instead of a delta, and the chain resets to
|
||||||
|
zero. Otherwise it appends a delta as today.
|
||||||
|
|
||||||
|
Consequences:
|
||||||
|
|
||||||
|
- A read costs at most **K + 1** preads, always, independent of when a
|
||||||
|
checkpoint fires.
|
||||||
|
- Replay costs **O(K²) per row**, bounded rather than unbounded.
|
||||||
|
- Write cost rises by one row-sized record per K updates — amortised, under
|
||||||
|
`1/K` extra bytes against today.
|
||||||
|
- Compaction, replay and the fold are untouched. A full-row record is a shape
|
||||||
|
all three already handle, because it is what an insert writes.
|
||||||
|
|
||||||
|
**K is a fixed constant, not a per-table knob.** Postgres ships `fillfactor` and
|
||||||
|
`base_thresh` as documented constants that are rarely tuned, and that is the
|
||||||
|
right precedent: K is a *bound*, not a dial. Anything from 8 to 64 caps the
|
||||||
|
pathology, and being wrong by a factor of two costs one extra row-write per K
|
||||||
|
updates.
|
||||||
|
|
||||||
|
**K does NOT scale with table size, and that is deliberate.** Postgres scales its
|
||||||
|
threshold by `reltuples` because it thresholds a table-level aggregate whose harm
|
||||||
|
is proportional. Ours is a per-row property with additive cost: reading one
|
||||||
|
product costs `1 + depth` preads whether the catalogue holds a hundred rows or
|
||||||
|
ten million, and total replay is the sum over every row's chain. Scaling K up
|
||||||
|
with table size would make the largest databases boot worst — exactly backwards.
|
||||||
|
|
||||||
|
**Row width is the one thing that might justify varying K**, since flattening
|
||||||
|
writes a whole row while a delta writes one field, so the write-amplification
|
||||||
|
break-even genuinely depends on row size. Deliberately **not** done now: hop
|
||||||
|
count is what bounds read and replay cost, which are the costs actually hurting,
|
||||||
|
and width would optimise only the write side. Revisit if measurement shows write
|
||||||
|
amplification matters.
|
||||||
|
|
||||||
|
## Tier 2 — give the checkpoint the trigger shape it is missing
|
||||||
|
|
||||||
|
Smaller, and separable from tier 1. Tier 1 bounds one row; tier 2 corrects the
|
||||||
|
whole-log policy's shape so it stops being blind to absolute garbage.
|
||||||
|
|
||||||
|
`wo_wal_should_compact` gains, alongside its existing ratio:
|
||||||
|
|
||||||
|
- **An absolute garbage term** — compact when reclaimable bytes exceed an
|
||||||
|
absolute threshold regardless of ratio. This is postgres's `base_thresh`, and
|
||||||
|
it is what our current "floor" is not.
|
||||||
|
- **A ceiling** — cap the proportional term so a very large live set does not
|
||||||
|
defer compaction indefinitely. This is `autovacuum_vacuum_max_threshold`.
|
||||||
|
|
||||||
|
The existing floor keeps its current meaning — do not bother with a tiny log —
|
||||||
|
but the doc comment must stop calling it a floor in postgres's sense, because it
|
||||||
|
does the opposite thing.
|
||||||
|
|
||||||
|
## Acceptance criteria
|
||||||
|
|
||||||
|
- **Given** a keys-resident row updated K times, **when** it is updated once
|
||||||
|
more, **then** the record its offset names is a full row, not a delta, and its
|
||||||
|
chain length is zero.
|
||||||
|
- **Given** a row updated many times more than K, **when** it is read, **then**
|
||||||
|
the read performs at most K + 1 record reads — asserted by counting, not by
|
||||||
|
timing.
|
||||||
|
- **Given** a row updated many times more than K, **when** the process restarts,
|
||||||
|
**then** replay reconstructs it correctly and its cost does not grow with the
|
||||||
|
total number of updates ever applied to it.
|
||||||
|
- **Given** a flattening update, **when** it is replayed, **then** the row is
|
||||||
|
identical to the same row in a `resident: all` table subjected to the same
|
||||||
|
update sequence. The resident table is the oracle.
|
||||||
|
- **Given** a flattening update that changes an indexed column, **when** the row
|
||||||
|
is queried through that index, **then** it is found by the new value and not
|
||||||
|
the old — before and after a restart.
|
||||||
|
- **Given** reclaimable bytes past the absolute threshold but within the ratio,
|
||||||
|
**when** the policy is evaluated, **then** compaction fires. *(Tier 2.)*
|
||||||
|
- **Given** a `resident: all` table, **when** anything here runs, **then**
|
||||||
|
nothing about its behaviour or its log records changes.
|
||||||
|
|
||||||
|
## Out of scope
|
||||||
|
|
||||||
|
- Varying K by row width. Reasoned above; revisit only with a measurement.
|
||||||
|
- A time-based compaction trigger. Records are durable at commit, so an idle log
|
||||||
|
does not grow — the existing design's reasoning still holds.
|
||||||
|
- The mid-drain stale-read limitation, which a separate fix already closed.
|
||||||
|
- Whether `resident: keys` is worth having at all. That is
|
||||||
|
[databasev2 2](../../stories/databasev2/02-table-storage-modes.md)'s task 7,
|
||||||
|
and this design does not answer it.
|
||||||
|
|
||||||
|
## Risks
|
||||||
|
|
||||||
|
- **K is a constant chosen without measurement.** The bound is right in shape;
|
||||||
|
its value is a judgement. The mitigation is that being wrong is cheap and
|
||||||
|
symmetric — too small costs write amplification, too large costs read latency,
|
||||||
|
and neither is a correctness failure.
|
||||||
|
- **Flattening makes one update in K expensive.** A burst of updates to one row
|
||||||
|
pays a row-sized write on every Kth. Acceptable, and the alternative is an
|
||||||
|
unbounded read path, but it should be visible in the measurement rather than
|
||||||
|
discovered in production.
|
||||||
Loading…
Reference in a new issue