docs(db2-chains): spec + story for bounding a row's delta chain
- fixes a limitation iteration 2 shipped: compaction was supposed to bound chain length, but wo_wal_should_compact triggers on a whole-log byte ratio and cannot see one hot row's chain - tier 1, flatten on update: the update path ALREADY folds the row for index maintenance and the fold already walks hop by hop, so it reports depth for free. Past a fixed K it writes a full row instead of a delta. Read <= K+1 reads, replay O(K^2) per row. No format change, no per-row RAM, no new trigger - tier 2: our compaction policy has a proportional term and a SUPPRESSOR misleadingly called a floor; postgres's floor TRIGGERS on small absolute garbage. Add that term and a ceiling - design read from .dev/reference/postgresql, not recalled: heap_page_prune_opt gates on an O(1) on-page hint then page fullness against Max(fillfactor, BLCKSZ/10); autovacuum uses base + scale * reltuples clamped by a max (50, 0.2, 1e8). Neither thresholds on new-bytes-versus-old-bytes - K deliberately does NOT scale with table size: postgres scales a table-level aggregate with proportional harm, ours is per-row with additive cost, so scaling up would make big databases boot worst - the story says plainly it should NOT be next: task 7 has still never measured whether resident: keys beats the kernel's own paging Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> (cherry picked from commit f667cad2cfbe187b5973440ab1af015b2df288f8)
This commit is contained in:
parent
b058feb517
commit
cfd660a5e6
4 changed files with 294 additions and 0 deletions
|
|
@ -160,6 +160,7 @@ before its mechanism existed; the history is in
|
|||
| 8 | [Query grammar from corpora](08-query-grammar-corpus.md) *(was language 27)* | whole-query `count`, `exists` | independent |
|
||||
| 9 | [Cross-program tables](09-cross-program-tables.md) *(was language 20)* | attach to a running program's database over local IPC | independent |
|
||||
| 10 | [Keypair attach auth](10-keypair-attach-auth.md) *(was language 21)* | program identity as a keypair; mutual challenge–response | 9 |
|
||||
| 11 | [Bounded delta chains](11-bounded-delta-chains.md) | cap a keys-resident row's delta chain in the update path, and give the compaction policy an absolute term + ceiling | 2 (fixes a limitation it shipped) |
|
||||
|
||||
```
|
||||
An arrow points AT the iteration that NEEDS the other.
|
||||
|
|
|
|||
|
|
@ -201,6 +201,13 @@ Met:
|
|||
not to cap chain length rests on compaction bounding it instead; for
|
||||
this shape it does not.
|
||||
|
||||
**Answered by [iteration 11](11-bounded-delta-chains.md)** (spec written
|
||||
2026-08-30): the update path already folds the row and the fold already
|
||||
walks hop by hop, so it reports the depth for free — past a fixed K the
|
||||
update writes a full row instead of a delta, and the chain resets. Read
|
||||
cost becomes at most K+1 reads and replay O(K²) per row, independent of
|
||||
when a checkpoint fires. Limitations 2 and 3 above both fall to it.
|
||||
|
||||
Outstanding:
|
||||
|
||||
- **Given** `durable: true` and no `WO_DATA`, **when** the program starts,
|
||||
|
|
|
|||
116
docs/stories/databasev2/11-bounded-delta-chains.md
Normal file
116
docs/stories/databasev2/11-bounded-delta-chains.md
Normal file
|
|
@ -0,0 +1,116 @@
|
|||
---
|
||||
track: databasev2
|
||||
iteration: "11"
|
||||
status: pending
|
||||
readiness: ready
|
||||
---
|
||||
|
||||
# databasev2 11 — bounding a keys-resident row's delta chain
|
||||
|
||||
> Part of [Story — databasev2: the database beyond RAM](00-story.md).
|
||||
> Spec: [`2026-08-30-bounded-delta-chains-design.md`](../../superpowers/specs/2026-08-30-bounded-delta-chains-design.md).
|
||||
>
|
||||
> **Why this exists.** [Iteration 2](02-table-storage-modes.md) shipped delta
|
||||
> updates with a deliberate decision not to cap chain length, on the reasoning
|
||||
> that compaction bounds it. A whole-branch review showed that reasoning does
|
||||
> not hold for the one workload the feature is motivated by. This iteration
|
||||
> closes it. `readiness: ready` — every fork is settled in the spec.
|
||||
|
||||
## The finding this iteration answers
|
||||
|
||||
A read of a keys-resident row costs `1 + chain length` preads, and replay costs
|
||||
**O(N²)** per chain. The shipped mitigation is compaction, which flattens chains
|
||||
to zero. But `wo_wal_should_compact` triggers on `used > last * ratio` — a byte
|
||||
ratio over the whole log — and cannot see that one row has a very long chain.
|
||||
|
||||
One popular SKU whose stock moves on every order, in a catalogue that is
|
||||
otherwise quiet, grows an unbounded chain without ever moving that ratio. The
|
||||
guard that bounds replay in general is structurally blind to the single case
|
||||
that makes replay quadratic.
|
||||
|
||||
## The design, in one sentence each
|
||||
|
||||
**Tier 1 — flatten on update.** The update path already folds the row, because
|
||||
it needs the old values for index maintenance, and the fold already walks the
|
||||
chain hop by hop; so it reports the depth for free, and when that depth reaches
|
||||
**K** the update appends a full-row record instead of a delta. Read cost becomes
|
||||
at most `K + 1` reads and replay `O(K²)` per row, independent of when a
|
||||
checkpoint fires.
|
||||
|
||||
**Tier 2 — give the compaction policy an absolute term and a ceiling.** Our
|
||||
current policy has a proportional term and a *suppressor* misleadingly named a
|
||||
floor; it lacks the triggering floor and the ceiling that keep a size-based
|
||||
policy honest.
|
||||
|
||||
## Where the design came from
|
||||
|
||||
Read from PostgreSQL's source at `.dev/reference/postgresql`, not recalled:
|
||||
|
||||
- **`heap_page_prune_opt`** collapses HOT chains opportunistically, on a page the
|
||||
process already holds, gated by an O(1) on-page hint and then by page fullness
|
||||
against `Max(fillfactor, BLCKSZ/10)`. Tier 1 is this shape: do the work while
|
||||
you already hold the thing, using a signal you already computed.
|
||||
- **autovacuum** thresholds on
|
||||
`vac_base_thresh + vac_scale_factor * reltuples`, clamped by a maximum —
|
||||
defaults 50, 0.2, 100 000 000. A count with a floor, a proportion and a
|
||||
ceiling, per table. Tier 2 borrows the floor and the ceiling.
|
||||
- **Postgres never thresholds on new-bytes-versus-old-bytes**, despite knowing
|
||||
exactly what a chain costs. Its space test is "will the next version fit" — an
|
||||
operational constraint, not an economic comparison. That ruled out the
|
||||
byte-ratio shape here too.
|
||||
|
||||
## Acceptance Criteria
|
||||
|
||||
Outstanding — none met; this iteration has not started.
|
||||
|
||||
- **Given** a row updated K times, **when** updated once more, **then** the
|
||||
record its offset names is a full row and its chain length is zero.
|
||||
- **Given** a row updated far more than K times, **when** it is read, **then** it
|
||||
performs at most K + 1 record reads, asserted by counting rather than timing.
|
||||
- **Given** the same row, **when** the process restarts, **then** replay is
|
||||
correct and its cost does not grow with the total updates ever applied.
|
||||
- **Given** a flattening update, **when** replayed, **then** the row matches the
|
||||
same row in a `resident: all` table under the same update sequence — the
|
||||
resident table is the oracle.
|
||||
- **Given** a flattening update to an indexed column, **when** queried through
|
||||
that index, **then** the row is found by its new value and not its old, before
|
||||
and after a restart.
|
||||
- **Given** reclaimable bytes past the absolute threshold but inside the ratio,
|
||||
**when** the policy is evaluated, **then** compaction fires. *(Tier 2.)*
|
||||
- **Given** a `resident: all` table, **when** any of this runs, **then** nothing
|
||||
about its behaviour or log records changes.
|
||||
|
||||
## Out Of Scope
|
||||
|
||||
- **Varying K by row width.** Hop count is what bounds read and replay cost;
|
||||
width would optimise only write amplification. Revisit with a measurement, not
|
||||
before.
|
||||
- **A time-based compaction trigger.** Records are durable at commit, so an idle
|
||||
log does not grow.
|
||||
- **Whether `resident: keys` earns its place at all.** That is iteration 2's
|
||||
task 7, and it should arguably run *before* this work — see below.
|
||||
|
||||
## Info — the forks, settled
|
||||
|
||||
1. **Where to fix it: the update path, not the checkpoint.** Making compaction
|
||||
depth-aware would mean one hot row triggering a stop-the-world rewrite of the
|
||||
entire log — a 2 651 µs pause that scales with total live rows, not with the
|
||||
row that misbehaved. Postgres reaches for the local, opportunistic fix first
|
||||
for the same reason.
|
||||
2. **The metric is hop count, not bytes.** Each hop is one `pread` whose cost
|
||||
barely varies with the bytes it carries, so hops are what our read cost is
|
||||
made of. Bytes govern write amplification, which is the secondary concern.
|
||||
3. **K is a fixed constant and does not scale with table size.** Postgres scales
|
||||
by `reltuples` because it thresholds a table-level aggregate with
|
||||
proportional harm. Ours is per-row with additive cost — reading one product
|
||||
costs the same whether the catalogue holds a hundred rows or ten million, and
|
||||
total replay is the sum across rows. Scaling K up with size would make the
|
||||
largest databases boot worst.
|
||||
|
||||
## Sequencing note
|
||||
|
||||
This iteration is **ready but arguably should not be next**. Iteration 2's
|
||||
task 7 has still never measured whether `resident: keys` beats the kernel's own
|
||||
paging, and everything built on it — including this — assumes it does. If that
|
||||
measurement comes back poorly, this work is optimising something that should be
|
||||
deleted. Recommended order: measure first, then this.
|
||||
170
docs/superpowers/specs/2026-08-30-bounded-delta-chains-design.md
Normal file
170
docs/superpowers/specs/2026-08-30-bounded-delta-chains-design.md
Normal file
|
|
@ -0,0 +1,170 @@
|
|||
# Bounding a keys-resident row's delta chain
|
||||
|
||||
Design settled 2026-08-30. Fixes the limitation shipped with
|
||||
[databasev2 2](../../stories/databasev2/02-table-storage-modes.md)'s delta
|
||||
updates and recorded in
|
||||
[`2026-08-30-keys-resident-delta-updates-design.md`](2026-08-30-keys-resident-delta-updates-design.md).
|
||||
|
||||
## The problem, and why the shipped mitigation does not fire
|
||||
|
||||
An update to a `resident: keys` row appends a delta — one field's new value plus
|
||||
a back-pointer. Reading the row folds the chain backward, so a read costs
|
||||
`1 + chain length` preads, and replay costs **O(N²)** per chain because it folds
|
||||
once per delta and each fold walks back to the base.
|
||||
|
||||
The shipped design chose not to cap chain length, on the reasoning that
|
||||
compaction flattens every chain and that deltas grow the log, pulling the next
|
||||
checkpoint forward. **That reasoning is wrong for the case that matters.**
|
||||
`wo_wal_should_compact` decides on `used > last * ratio` — a byte ratio over the
|
||||
whole log. It cannot see that one row has a five-thousand-delta chain. A single
|
||||
hot row taking many small updates barely moves that ratio in a large database,
|
||||
so the checkpoint never fires, that row's chain grows without bound, and its
|
||||
replay cost grows as the square.
|
||||
|
||||
The motivating workload is precisely this shape: one popular SKU whose stock
|
||||
moves on every order while the rest of the catalogue sits still.
|
||||
|
||||
## What PostgreSQL does, read from source
|
||||
|
||||
Verified against `.dev/reference/postgresql`, not recalled. Postgres solves the
|
||||
same class of problem — chains of row versions that must be collapsed — with
|
||||
**two tiers**, and neither is a size ratio.
|
||||
|
||||
**Tier 1, `heap_page_prune_opt` in `pruneheap.c`** — opportunistic and local.
|
||||
Three gates, cheapest first: an O(1) `pd_prune_xid` hint stored on the page; a
|
||||
visibility test; then `PageIsFull(page) || PageGetHeapFreeSpace(page) < minfree`
|
||||
where `minfree = Max(fillfactor target, BLCKSZ / 10)`. The work happens on a
|
||||
page the process **already holds** because it is reading or updating it anyway.
|
||||
The source is explicit that the check is deliberately approximate — it reads
|
||||
free space without taking a lock, because "avoiding taking a lock seems more
|
||||
important than sometimes getting a wrong answer in what is after all just a
|
||||
heuristic estimate."
|
||||
|
||||
**Tier 2, autovacuum** — background and per-table:
|
||||
`vacthresh = vac_base_thresh + vac_scale_factor * reltuples`, clamped by
|
||||
`autovacuum_vacuum_max_threshold`. Shipped defaults are 50, 0.2 and 100 000 000.
|
||||
A **count** with a floor, a proportional term and a ceiling — computed per
|
||||
table, never per database.
|
||||
|
||||
Three lessons, and one correction to our own vocabulary:
|
||||
|
||||
- **Do the work while you already hold the thing.** That is the whole of tier 1.
|
||||
- **The floor exists to catch what the proportion hides.** `base_thresh = 50`
|
||||
fires on a small table where 20% would not. `Max(…, BLCKSZ/10)` does the same
|
||||
for space.
|
||||
- **The ceiling exists so scale does not defer forever.**
|
||||
- **Our "floor" is the opposite of theirs, despite the name.**
|
||||
`wo_wal_should_compact` reads `if (used < floor) return 0` — ours *suppresses*
|
||||
compaction on a small log. Postgres's floor *triggers* cleanup on a small
|
||||
absolute problem. We have the proportional term and the suppressor; we have
|
||||
neither the triggering floor nor the ceiling.
|
||||
|
||||
Postgres also, notably, does **not** threshold on "new bytes versus old bytes",
|
||||
despite knowing exactly what every chain costs. Its space check is "will the
|
||||
next version physically fit", a hard operational constraint, not an economic
|
||||
comparison. That rules out the byte-ratio shape for us as well — and our cost is
|
||||
worse suited to it still, since each hop is one `pread` whose cost barely varies
|
||||
with the bytes it carries.
|
||||
|
||||
## Tier 1 — flatten on update
|
||||
|
||||
**The update path already folds the row.** It must: it needs the old values to
|
||||
maintain indexes. And the fold already walks the chain hop by hop. So it can
|
||||
report how many hops it took, and the update path learns the chain's depth for
|
||||
free — no new record field, no extra read, no per-row RAM. That reported hop
|
||||
count is our `pd_prune_xid`: the cheap signal that says whether work is worth
|
||||
doing, obtained from something we were doing anyway.
|
||||
|
||||
The rule is one branch. When the fold reports a depth at or beyond **K**, the
|
||||
update appends a **full-row record** instead of a delta, and the chain resets to
|
||||
zero. Otherwise it appends a delta as today.
|
||||
|
||||
Consequences:
|
||||
|
||||
- A read costs at most **K + 1** preads, always, independent of when a
|
||||
checkpoint fires.
|
||||
- Replay costs **O(K²) per row**, bounded rather than unbounded.
|
||||
- Write cost rises by one row-sized record per K updates — amortised, under
|
||||
`1/K` extra bytes against today.
|
||||
- Compaction, replay and the fold are untouched. A full-row record is a shape
|
||||
all three already handle, because it is what an insert writes.
|
||||
|
||||
**K is a fixed constant, not a per-table knob.** Postgres ships `fillfactor` and
|
||||
`base_thresh` as documented constants that are rarely tuned, and that is the
|
||||
right precedent: K is a *bound*, not a dial. Anything from 8 to 64 caps the
|
||||
pathology, and being wrong by a factor of two costs one extra row-write per K
|
||||
updates.
|
||||
|
||||
**K does NOT scale with table size, and that is deliberate.** Postgres scales its
|
||||
threshold by `reltuples` because it thresholds a table-level aggregate whose harm
|
||||
is proportional. Ours is a per-row property with additive cost: reading one
|
||||
product costs `1 + depth` preads whether the catalogue holds a hundred rows or
|
||||
ten million, and total replay is the sum over every row's chain. Scaling K up
|
||||
with table size would make the largest databases boot worst — exactly backwards.
|
||||
|
||||
**Row width is the one thing that might justify varying K**, since flattening
|
||||
writes a whole row while a delta writes one field, so the write-amplification
|
||||
break-even genuinely depends on row size. Deliberately **not** done now: hop
|
||||
count is what bounds read and replay cost, which are the costs actually hurting,
|
||||
and width would optimise only the write side. Revisit if measurement shows write
|
||||
amplification matters.
|
||||
|
||||
## Tier 2 — give the checkpoint the trigger shape it is missing
|
||||
|
||||
Smaller, and separable from tier 1. Tier 1 bounds one row; tier 2 corrects the
|
||||
whole-log policy's shape so it stops being blind to absolute garbage.
|
||||
|
||||
`wo_wal_should_compact` gains, alongside its existing ratio:
|
||||
|
||||
- **An absolute garbage term** — compact when reclaimable bytes exceed an
|
||||
absolute threshold regardless of ratio. This is postgres's `base_thresh`, and
|
||||
it is what our current "floor" is not.
|
||||
- **A ceiling** — cap the proportional term so a very large live set does not
|
||||
defer compaction indefinitely. This is `autovacuum_vacuum_max_threshold`.
|
||||
|
||||
The existing floor keeps its current meaning — do not bother with a tiny log —
|
||||
but the doc comment must stop calling it a floor in postgres's sense, because it
|
||||
does the opposite thing.
|
||||
|
||||
## Acceptance criteria
|
||||
|
||||
- **Given** a keys-resident row updated K times, **when** it is updated once
|
||||
more, **then** the record its offset names is a full row, not a delta, and its
|
||||
chain length is zero.
|
||||
- **Given** a row updated many times more than K, **when** it is read, **then**
|
||||
the read performs at most K + 1 record reads — asserted by counting, not by
|
||||
timing.
|
||||
- **Given** a row updated many times more than K, **when** the process restarts,
|
||||
**then** replay reconstructs it correctly and its cost does not grow with the
|
||||
total number of updates ever applied to it.
|
||||
- **Given** a flattening update, **when** it is replayed, **then** the row is
|
||||
identical to the same row in a `resident: all` table subjected to the same
|
||||
update sequence. The resident table is the oracle.
|
||||
- **Given** a flattening update that changes an indexed column, **when** the row
|
||||
is queried through that index, **then** it is found by the new value and not
|
||||
the old — before and after a restart.
|
||||
- **Given** reclaimable bytes past the absolute threshold but within the ratio,
|
||||
**when** the policy is evaluated, **then** compaction fires. *(Tier 2.)*
|
||||
- **Given** a `resident: all` table, **when** anything here runs, **then**
|
||||
nothing about its behaviour or its log records changes.
|
||||
|
||||
## Out of scope
|
||||
|
||||
- Varying K by row width. Reasoned above; revisit only with a measurement.
|
||||
- A time-based compaction trigger. Records are durable at commit, so an idle log
|
||||
does not grow — the existing design's reasoning still holds.
|
||||
- The mid-drain stale-read limitation, which a separate fix already closed.
|
||||
- Whether `resident: keys` is worth having at all. That is
|
||||
[databasev2 2](../../stories/databasev2/02-table-storage-modes.md)'s task 7,
|
||||
and this design does not answer it.
|
||||
|
||||
## Risks
|
||||
|
||||
- **K is a constant chosen without measurement.** The bound is right in shape;
|
||||
its value is a judgement. The mitigation is that being wrong is cheap and
|
||||
symmetric — too small costs write amplification, too large costs read latency,
|
||||
and neither is a correctness failure.
|
||||
- **Flattening makes one update in K expensive.** A burst of updates to one row
|
||||
pays a row-sized write on every Kth. Acceptable, and the alternative is an
|
||||
unbounded read path, but it should be visible in the measurement rather than
|
||||
discovered in production.
|
||||
Loading…
Reference in a new issue