diff --git a/docs/stories/databasev2/00-story.md b/docs/stories/databasev2/00-story.md index a5a59ab..65eca09 100644 --- a/docs/stories/databasev2/00-story.md +++ b/docs/stories/databasev2/00-story.md @@ -160,6 +160,7 @@ before its mechanism existed; the history is in | 8 | [Query grammar from corpora](08-query-grammar-corpus.md) *(was language 27)* | whole-query `count`, `exists` | independent | | 9 | [Cross-program tables](09-cross-program-tables.md) *(was language 20)* | attach to a running program's database over local IPC | independent | | 10 | [Keypair attach auth](10-keypair-attach-auth.md) *(was language 21)* | program identity as a keypair; mutual challenge–response | 9 | +| 11 | [Bounded delta chains](11-bounded-delta-chains.md) | cap a keys-resident row's delta chain in the update path, and give the compaction policy an absolute term + ceiling | 2 (fixes a limitation it shipped) | ``` An arrow points AT the iteration that NEEDS the other. diff --git a/docs/stories/databasev2/02-table-storage-modes.md b/docs/stories/databasev2/02-table-storage-modes.md index 2a6565a..3d4681b 100644 --- a/docs/stories/databasev2/02-table-storage-modes.md +++ b/docs/stories/databasev2/02-table-storage-modes.md @@ -201,6 +201,13 @@ Met: not to cap chain length rests on compaction bounding it instead; for this shape it does not. + **Answered by [iteration 11](11-bounded-delta-chains.md)** (spec written + 2026-08-30): the update path already folds the row and the fold already + walks hop by hop, so it reports the depth for free — past a fixed K the + update writes a full row instead of a delta, and the chain resets. Read + cost becomes at most K+1 reads and replay O(K²) per row, independent of + when a checkpoint fires. Limitations 2 and 3 above both fall to it. + Outstanding: - **Given** `durable: true` and no `WO_DATA`, **when** the program starts, diff --git a/docs/stories/databasev2/11-bounded-delta-chains.md b/docs/stories/databasev2/11-bounded-delta-chains.md new file mode 100644 index 0000000..3dc1694 --- /dev/null +++ b/docs/stories/databasev2/11-bounded-delta-chains.md @@ -0,0 +1,116 @@ +--- +track: databasev2 +iteration: "11" +status: pending +readiness: ready +--- + +# databasev2 11 — bounding a keys-resident row's delta chain + +> Part of [Story — databasev2: the database beyond RAM](00-story.md). +> Spec: [`2026-08-30-bounded-delta-chains-design.md`](../../superpowers/specs/2026-08-30-bounded-delta-chains-design.md). +> +> **Why this exists.** [Iteration 2](02-table-storage-modes.md) shipped delta +> updates with a deliberate decision not to cap chain length, on the reasoning +> that compaction bounds it. A whole-branch review showed that reasoning does +> not hold for the one workload the feature is motivated by. This iteration +> closes it. `readiness: ready` — every fork is settled in the spec. + +## The finding this iteration answers + +A read of a keys-resident row costs `1 + chain length` preads, and replay costs +**O(N²)** per chain. The shipped mitigation is compaction, which flattens chains +to zero. But `wo_wal_should_compact` triggers on `used > last * ratio` — a byte +ratio over the whole log — and cannot see that one row has a very long chain. + +One popular SKU whose stock moves on every order, in a catalogue that is +otherwise quiet, grows an unbounded chain without ever moving that ratio. The +guard that bounds replay in general is structurally blind to the single case +that makes replay quadratic. + +## The design, in one sentence each + +**Tier 1 — flatten on update.** The update path already folds the row, because +it needs the old values for index maintenance, and the fold already walks the +chain hop by hop; so it reports the depth for free, and when that depth reaches +**K** the update appends a full-row record instead of a delta. Read cost becomes +at most `K + 1` reads and replay `O(K²)` per row, independent of when a +checkpoint fires. + +**Tier 2 — give the compaction policy an absolute term and a ceiling.** Our +current policy has a proportional term and a *suppressor* misleadingly named a +floor; it lacks the triggering floor and the ceiling that keep a size-based +policy honest. + +## Where the design came from + +Read from PostgreSQL's source at `.dev/reference/postgresql`, not recalled: + +- **`heap_page_prune_opt`** collapses HOT chains opportunistically, on a page the + process already holds, gated by an O(1) on-page hint and then by page fullness + against `Max(fillfactor, BLCKSZ/10)`. Tier 1 is this shape: do the work while + you already hold the thing, using a signal you already computed. +- **autovacuum** thresholds on + `vac_base_thresh + vac_scale_factor * reltuples`, clamped by a maximum — + defaults 50, 0.2, 100 000 000. A count with a floor, a proportion and a + ceiling, per table. Tier 2 borrows the floor and the ceiling. +- **Postgres never thresholds on new-bytes-versus-old-bytes**, despite knowing + exactly what a chain costs. Its space test is "will the next version fit" — an + operational constraint, not an economic comparison. That ruled out the + byte-ratio shape here too. + +## Acceptance Criteria + +Outstanding — none met; this iteration has not started. + +- **Given** a row updated K times, **when** updated once more, **then** the + record its offset names is a full row and its chain length is zero. +- **Given** a row updated far more than K times, **when** it is read, **then** it + performs at most K + 1 record reads, asserted by counting rather than timing. +- **Given** the same row, **when** the process restarts, **then** replay is + correct and its cost does not grow with the total updates ever applied. +- **Given** a flattening update, **when** replayed, **then** the row matches the + same row in a `resident: all` table under the same update sequence — the + resident table is the oracle. +- **Given** a flattening update to an indexed column, **when** queried through + that index, **then** the row is found by its new value and not its old, before + and after a restart. +- **Given** reclaimable bytes past the absolute threshold but inside the ratio, + **when** the policy is evaluated, **then** compaction fires. *(Tier 2.)* +- **Given** a `resident: all` table, **when** any of this runs, **then** nothing + about its behaviour or log records changes. + +## Out Of Scope + +- **Varying K by row width.** Hop count is what bounds read and replay cost; + width would optimise only write amplification. Revisit with a measurement, not + before. +- **A time-based compaction trigger.** Records are durable at commit, so an idle + log does not grow. +- **Whether `resident: keys` earns its place at all.** That is iteration 2's + task 7, and it should arguably run *before* this work — see below. + +## Info — the forks, settled + +1. **Where to fix it: the update path, not the checkpoint.** Making compaction + depth-aware would mean one hot row triggering a stop-the-world rewrite of the + entire log — a 2 651 µs pause that scales with total live rows, not with the + row that misbehaved. Postgres reaches for the local, opportunistic fix first + for the same reason. +2. **The metric is hop count, not bytes.** Each hop is one `pread` whose cost + barely varies with the bytes it carries, so hops are what our read cost is + made of. Bytes govern write amplification, which is the secondary concern. +3. **K is a fixed constant and does not scale with table size.** Postgres scales + by `reltuples` because it thresholds a table-level aggregate with + proportional harm. Ours is per-row with additive cost — reading one product + costs the same whether the catalogue holds a hundred rows or ten million, and + total replay is the sum across rows. Scaling K up with size would make the + largest databases boot worst. + +## Sequencing note + +This iteration is **ready but arguably should not be next**. Iteration 2's +task 7 has still never measured whether `resident: keys` beats the kernel's own +paging, and everything built on it — including this — assumes it does. If that +measurement comes back poorly, this work is optimising something that should be +deleted. Recommended order: measure first, then this. diff --git a/docs/superpowers/specs/2026-08-30-bounded-delta-chains-design.md b/docs/superpowers/specs/2026-08-30-bounded-delta-chains-design.md new file mode 100644 index 0000000..082aac1 --- /dev/null +++ b/docs/superpowers/specs/2026-08-30-bounded-delta-chains-design.md @@ -0,0 +1,170 @@ +# Bounding a keys-resident row's delta chain + +Design settled 2026-08-30. Fixes the limitation shipped with +[databasev2 2](../../stories/databasev2/02-table-storage-modes.md)'s delta +updates and recorded in +[`2026-08-30-keys-resident-delta-updates-design.md`](2026-08-30-keys-resident-delta-updates-design.md). + +## The problem, and why the shipped mitigation does not fire + +An update to a `resident: keys` row appends a delta — one field's new value plus +a back-pointer. Reading the row folds the chain backward, so a read costs +`1 + chain length` preads, and replay costs **O(N²)** per chain because it folds +once per delta and each fold walks back to the base. + +The shipped design chose not to cap chain length, on the reasoning that +compaction flattens every chain and that deltas grow the log, pulling the next +checkpoint forward. **That reasoning is wrong for the case that matters.** +`wo_wal_should_compact` decides on `used > last * ratio` — a byte ratio over the +whole log. It cannot see that one row has a five-thousand-delta chain. A single +hot row taking many small updates barely moves that ratio in a large database, +so the checkpoint never fires, that row's chain grows without bound, and its +replay cost grows as the square. + +The motivating workload is precisely this shape: one popular SKU whose stock +moves on every order while the rest of the catalogue sits still. + +## What PostgreSQL does, read from source + +Verified against `.dev/reference/postgresql`, not recalled. Postgres solves the +same class of problem — chains of row versions that must be collapsed — with +**two tiers**, and neither is a size ratio. + +**Tier 1, `heap_page_prune_opt` in `pruneheap.c`** — opportunistic and local. +Three gates, cheapest first: an O(1) `pd_prune_xid` hint stored on the page; a +visibility test; then `PageIsFull(page) || PageGetHeapFreeSpace(page) < minfree` +where `minfree = Max(fillfactor target, BLCKSZ / 10)`. The work happens on a +page the process **already holds** because it is reading or updating it anyway. +The source is explicit that the check is deliberately approximate — it reads +free space without taking a lock, because "avoiding taking a lock seems more +important than sometimes getting a wrong answer in what is after all just a +heuristic estimate." + +**Tier 2, autovacuum** — background and per-table: +`vacthresh = vac_base_thresh + vac_scale_factor * reltuples`, clamped by +`autovacuum_vacuum_max_threshold`. Shipped defaults are 50, 0.2 and 100 000 000. +A **count** with a floor, a proportional term and a ceiling — computed per +table, never per database. + +Three lessons, and one correction to our own vocabulary: + +- **Do the work while you already hold the thing.** That is the whole of tier 1. +- **The floor exists to catch what the proportion hides.** `base_thresh = 50` + fires on a small table where 20% would not. `Max(…, BLCKSZ/10)` does the same + for space. +- **The ceiling exists so scale does not defer forever.** +- **Our "floor" is the opposite of theirs, despite the name.** + `wo_wal_should_compact` reads `if (used < floor) return 0` — ours *suppresses* + compaction on a small log. Postgres's floor *triggers* cleanup on a small + absolute problem. We have the proportional term and the suppressor; we have + neither the triggering floor nor the ceiling. + +Postgres also, notably, does **not** threshold on "new bytes versus old bytes", +despite knowing exactly what every chain costs. Its space check is "will the +next version physically fit", a hard operational constraint, not an economic +comparison. That rules out the byte-ratio shape for us as well — and our cost is +worse suited to it still, since each hop is one `pread` whose cost barely varies +with the bytes it carries. + +## Tier 1 — flatten on update + +**The update path already folds the row.** It must: it needs the old values to +maintain indexes. And the fold already walks the chain hop by hop. So it can +report how many hops it took, and the update path learns the chain's depth for +free — no new record field, no extra read, no per-row RAM. That reported hop +count is our `pd_prune_xid`: the cheap signal that says whether work is worth +doing, obtained from something we were doing anyway. + +The rule is one branch. When the fold reports a depth at or beyond **K**, the +update appends a **full-row record** instead of a delta, and the chain resets to +zero. Otherwise it appends a delta as today. + +Consequences: + +- A read costs at most **K + 1** preads, always, independent of when a + checkpoint fires. +- Replay costs **O(K²) per row**, bounded rather than unbounded. +- Write cost rises by one row-sized record per K updates — amortised, under + `1/K` extra bytes against today. +- Compaction, replay and the fold are untouched. A full-row record is a shape + all three already handle, because it is what an insert writes. + +**K is a fixed constant, not a per-table knob.** Postgres ships `fillfactor` and +`base_thresh` as documented constants that are rarely tuned, and that is the +right precedent: K is a *bound*, not a dial. Anything from 8 to 64 caps the +pathology, and being wrong by a factor of two costs one extra row-write per K +updates. + +**K does NOT scale with table size, and that is deliberate.** Postgres scales its +threshold by `reltuples` because it thresholds a table-level aggregate whose harm +is proportional. Ours is a per-row property with additive cost: reading one +product costs `1 + depth` preads whether the catalogue holds a hundred rows or +ten million, and total replay is the sum over every row's chain. Scaling K up +with table size would make the largest databases boot worst — exactly backwards. + +**Row width is the one thing that might justify varying K**, since flattening +writes a whole row while a delta writes one field, so the write-amplification +break-even genuinely depends on row size. Deliberately **not** done now: hop +count is what bounds read and replay cost, which are the costs actually hurting, +and width would optimise only the write side. Revisit if measurement shows write +amplification matters. + +## Tier 2 — give the checkpoint the trigger shape it is missing + +Smaller, and separable from tier 1. Tier 1 bounds one row; tier 2 corrects the +whole-log policy's shape so it stops being blind to absolute garbage. + +`wo_wal_should_compact` gains, alongside its existing ratio: + +- **An absolute garbage term** — compact when reclaimable bytes exceed an + absolute threshold regardless of ratio. This is postgres's `base_thresh`, and + it is what our current "floor" is not. +- **A ceiling** — cap the proportional term so a very large live set does not + defer compaction indefinitely. This is `autovacuum_vacuum_max_threshold`. + +The existing floor keeps its current meaning — do not bother with a tiny log — +but the doc comment must stop calling it a floor in postgres's sense, because it +does the opposite thing. + +## Acceptance criteria + +- **Given** a keys-resident row updated K times, **when** it is updated once + more, **then** the record its offset names is a full row, not a delta, and its + chain length is zero. +- **Given** a row updated many times more than K, **when** it is read, **then** + the read performs at most K + 1 record reads — asserted by counting, not by + timing. +- **Given** a row updated many times more than K, **when** the process restarts, + **then** replay reconstructs it correctly and its cost does not grow with the + total number of updates ever applied to it. +- **Given** a flattening update, **when** it is replayed, **then** the row is + identical to the same row in a `resident: all` table subjected to the same + update sequence. The resident table is the oracle. +- **Given** a flattening update that changes an indexed column, **when** the row + is queried through that index, **then** it is found by the new value and not + the old — before and after a restart. +- **Given** reclaimable bytes past the absolute threshold but within the ratio, + **when** the policy is evaluated, **then** compaction fires. *(Tier 2.)* +- **Given** a `resident: all` table, **when** anything here runs, **then** + nothing about its behaviour or its log records changes. + +## Out of scope + +- Varying K by row width. Reasoned above; revisit only with a measurement. +- A time-based compaction trigger. Records are durable at commit, so an idle log + does not grow — the existing design's reasoning still holds. +- The mid-drain stale-read limitation, which a separate fix already closed. +- Whether `resident: keys` is worth having at all. That is + [databasev2 2](../../stories/databasev2/02-table-storage-modes.md)'s task 7, + and this design does not answer it. + +## Risks + +- **K is a constant chosen without measurement.** The bound is right in shape; + its value is a judgement. The mitigation is that being wrong is cheap and + symmetric — too small costs write amplification, too large costs read latency, + and neither is a correctness failure. +- **Flattening makes one update in K expensive.** A burst of updates to one row + pays a row-sized write on every Kth. Acceptable, and the alternative is an + unbounded read path, but it should be visible in the measurement rather than + discovered in production.