- docs/stories/databasev2/, numbered from 1. Six PENDING database iterations
moved from the language track and renumbered, keeping the old id in
`was_language_iteration:` so a search for "iteration 32" still finds it:
32 -> 3 WAL checkpoint, 23 -> 4 io_uring commit, 33 -> 7 single-file store,
27 -> 8 query grammar, 20 -> 9 cross-program, 21 -> 10 keypair auth.
Done work (9, 9b, 22) stays as v1 history; language 18 left whole
- the problem, read off the engine not guessed: rows are malloc'd slabs with
addresses stable forever, NO eviction/spill/paging anywhere in database/src,
the WAL never checkpoints so boot replays all history, and durability is one
process-global WO_DATA so no table can say it matters more than another.
An allocation failure IS a clean catchable WO_T_OOM — but swap thrash
arrives first and carries no error signal at all, which is the real hazard
- four new iterations:
1 measure the ceiling FIRST (curve not cliff; the three exits; kill -9 at
exhaustion) — every later default should follow from a number
2 `@table(mode: ram | durable | cold)` — the grammar ask. Small surface
(Ast.table_cfg gains a key, the parser already rejects unknown args), big
semantics: `durable` defaults so nothing changes silently, and the
compiler refuses a durable row holding a `ref` into a ram table
5 bounded tables + refuse/evict/back-pressure, shedding BEFORE the OS acts
6 cold tiering — mostly forks, incl. whether the language surfaces the
fault cost and whether @unique on cold is refused outright. A paged
B-tree stays rejected: if tiering needs one, reject tiering
- 39 links repointed, link TEXT renumbered to track-local ids; arc gains one
pointer row replacing the six moved; board + board-views cover three tracks
- linkcheck 0 broken / 0 anchors; no code blocks in any story
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
10 KiB
10 KiB
| track | iteration | status |
|---|---|---|
| databasev2 | 6 | refine |
databasev2 6 — cold tiering: rows that leave RAM and come back
Part of Story — databasev2: the database beyond RAM. Needs 2 for the
coldmode declaration, 3 so the log this builds on does not grow forever, and 5 for the policy machinery.The iteration that actually raises the ceiling, and the one most likely to go wrong. Everything before it makes the limit visible, declared and managed. This one removes it — for tables that opt in — and in doing so touches the project's most load-bearing principle. It should be approached with more suspicion than enthusiasm.
Goals
- A
coldtable may hold more rows than fit in memory. Recently-used rows are resident; the rest live on disk and are faulted back on access. This is the whole feature and every other goal is a constraint on it. - Do it without becoming a paged storage engine.
discarded.mdrecords that the disk story is the WAL and that a paged B-tree engine was rejected.coldmust not be a licence to rebuild SQLite insidedatabase/src. The design that respects the doctrine reuses the log that already exists plus an index into it — a log-structured read path, not a page cache. - Keep the resident path exactly as fast as it is. A
durableorramtable must not pay one instruction for a feature it does not use. Iteration 22's 1.3M ops/s read baseline is the regression gate, and a measurable read regression on non-coldtables is grounds to reject the design, not to tune it. - Be honest in the query surface about what a fault costs. A scan over a
coldtable can touch disk per row. The engine currently answers every read from memory at p50 1µs; acoldscan is a different animal and the language should not pretend otherwise — see fork 3, which is the most important question in this iteration. - Never lose an acked write. Every guarantee iterations 9 and 22 established holds byte-for-byte: ack-after-fsync, whole-or-nothing replay, torn tails dropped by CRC. A tiering layer that weakens any of those is a regression disguised as a feature.
Phases
Phase A — settle the design before writing any of it
- This iteration needs a spec more than any other in the track. The candidate shapes are genuinely different: (a) the WAL becomes the primary store with an in-memory id→offset index and a resident row cache; (b) a separate append-only row file per cold table, checkpointed by 3's machinery; (c) eviction to disk with a free-space map, which is the paged engine wearing a hat.
- Whichever wins must state its read amplification, its recovery story, and what happens when the index itself does not fit — an id→offset map for a billion rows is not free either, and a design that only moves the ceiling is worth knowing about before it is built.
- Verify: the spec names the shape, the amplification, and the failure modes. No code in this phase.
Phase B — the resident/cold boundary
- Which rows are resident: reuse 5's policy and accounting rather than inventing a second notion of "least valuable".
- Eviction becomes write-then-drop instead of drop, and it must be atomic with respect to a concurrent reader — a row that is being written out must not be briefly unreachable.
- The fault path: a lookup that misses resident memory reads from disk, materialises the row, and admits it under the resident policy.
- Verify: a table larger than the resident bound serves correct rows for every
id; a row evicted and faulted back is byte-identical, including every heap
value (
db_text,db_rec,db_multi,db_mapeach round-trip).
Phase C — indexes and constraints across the boundary
- The hard part, and the reason this is late in the track. A secondary index
over a
coldtable either stays fully resident (bounding the table by index size rather than row size — which may be the honest answer) or is itself tiered. A@uniqueconstraint must hold across rows nobody has in memory: the shadow check cannot scan a slab that is not there. - Foreign-key restrict must also hold — a delete has to know whether any cold row references it.
- Verify:
@uniquerefuses a duplicate whose only conflicting row is cold; FK restrict refuses a delete whose only referrer is cold. These two criteria are the correctness core of the iteration.
Phase D — recovery
- Crash mid-eviction, crash mid-fault, crash mid-checkpoint-of-a-cold-table. Each must recover to a consistent state with no acked write lost and no row visible twice.
- Interaction with 3's snapshot: a cold table's on-disk rows are part of the durable state a checkpoint must account for, not something it can truncate past.
- Verify:
kill -9at each of the three points, replayed, with every acked write present and the resident/cold split re-derived correctly.
Phase E — measure it, then decide whether to keep it
- Read/write throughput and p99 for a
coldtable at several resident ratios, and a regression check that non-coldtables did not move. - Publish the amplification honestly in
perf-targets.md: how much slower a cold fault is than a resident read, as a number. - Verify:
just db-benchgreen with new cold-path rows; the resident baseline unchanged;just employee,just db-actor,oop-acceptgreen; ASan and TSan clean on the fault path.
Acceptance Criteria
- Given a
coldtable with more rows than the resident bound, when any row is looked up by id, then it is returned correctly whether resident or faulted, byte-identical including every heap-valued column. - Given a
coldtable under a read workload, when the resident set is smaller than the working set, then the process holds steady state without approaching the RAM ceiling iteration 1 measured. - Given a
@uniquecolumn on acoldtable, when a duplicate is inserted whose conflicting row is not resident, then the insert is refused — the constraint holds across the boundary or it does not hold. - Given a
refinto acoldtable, when the referenced row's owner is deleted and the only referrer is cold, then FK restrict refuses the delete. - Given
kill -9during an eviction, a fault, and a checkpoint, when the program restarts, then every acked write is present, no row appears twice, and the resident/cold split is re-derived correctly. - Given a
durableorramtable, when the read benchmark runs after this iteration, then its throughput and p99 are inside the existing baseline tolerance — no cost for a feature not used. - Given a cold fault, when its latency is measured, then the
amplification versus a resident read is recorded in
perf-targets.mdas a number a developer can plan around.
Out Of Scope
- A paged B-tree storage engine. Explicitly rejected in
discarded.mdand not reopened by this iteration. If the spec phase concludes that tiering requires one, the correct outcome is to reject tiering and say so — not to quietly build the thing the project decided against. - Making
coldthe default, or applying it to a table that did not ask. Opt-in per table, forever. - Tiering to anything but the local filesystem. Object storage needs outbound sockets (language iteration 38) and would change the latency story by orders of magnitude.
- Compression of cold rows. Composes with porch 7's codec if that lands first; not a dependency either way and not this slice.
- Cross-shard cold tables. The owner shard owns the store and the WAL; a cold table is more of the same. Per-shard storage is a separate architectural question noted in 2's forks.
- Tiering the query planner's behaviour. If a scan over a cold table is expensive, the answer for now is that it is expensive and documented — not a cost-based planner.
Info
Forks the spec must settle — this iteration is mostly forks, which is why phase A produces no code:
- Which shape? WAL-as-primary-store with an id→offset index and a row cache reuses machinery that exists and keeps the doctrine ("the disk story is the WAL") literally true. A separate per-table row file is cleaner to reason about and duplicates the log. Eviction with a free-space map is the rejected paged design. Leaning (a), with the caveat in fork 2.
- What if the index does not fit either? An id→offset entry per row is far
smaller than a row, so this moves the ceiling by a large constant — but it
does not remove it. Say so plainly in the spec:
coldbuys an order of magnitude, not infinity. A design sold as unlimited will be deployed as if it were. - Does the language surface the cost? Three positions. Silent — a cold table reads like any other and the developer discovers the latency in production. Annotated — the mode is at the declaration, so an attentive reader knows, which is the status quo of this design. Or explicit at the use site, where a query over a cold table must acknowledge it somehow. The third is most in keeping with a language whose whole thesis is that the compiler tells you the truth — and it is also the most intrusive. This is the fork with the largest effect on what writeonce is, and it deserves the brainstorm more than any implementation detail here.
- Is
@uniqueon a cold table simply refused? Keeping a unique index fully resident is a bound on the table by index size, which is honest and simple. Refusing@uniqueoncoldoutright is even simpler and might be right for a first version — a constraint that silently only checks resident rows would be a correctness hole, and that is the one outcome that must not ship.