- 05: hand-rolled JSON - 06: bespoke error type - 07: inotify content watcher - 08: sendfile static assets - 09: concurrency scaleout - 10: storage foundations - 11: WAL and recovery - 12: engine disk cutover - 14: MVC UI implementation - 15: MCP streamable HTTP - 16: Postgres mirror
11 KiB
12 — Engine cutover: rows live on disk
Kanban: ⬜ not started — the C prototype's phase B (mmap arena) is the proving ground. Board: 00-kanban.md
Context sources: ./10-storage-foundations.md, ./11-wal-and-recovery.md, ../runtime/database/03-inmemory-engine.md, ../runtime/database/07-wo-seg-migration.md, ./exploration/postgresql/buffer-and-checkpoint.md, ./exploration/postgresql/page-format.md, ./exploration/linux/12-pwrite-fsync.md.
Goal
Engine's row payload is no longer in RAM. The in-memory map is HashMap<type, BTreeMap<i64, SegmentOffset>>. Reads pread against the segment file and verify the CRC. RAM footprint is bounded by id count + per-id overhead, independent of row payload size — the runtime can now serve a dataset 10× larger than RAM.
A periodic checkpoint flushes dirty segments and advances the control-file LSN, bounding recovery time on restart.
Design decisions (locked)
- In-memory index is
BTreeMap<i64, SegmentOffset>. Keeps the existingEngine::listinsertion-order iteration. Roughly 24 bytes per entry (i64 key + u64 value + tree node overhead) — a million rows fits in 24 MiB regardless of row size. - Reads via
pread+ decode + CRC verify. No user-space buffer pool — the OS page cache is the cache (per./exploration/postgresql/buffer-and-checkpoint.md). Hot rows hit cached pages and the syscall returns memcpy-fast. - Bounded LRU on top of pread. Optional small
HashMap<(ty, offset), Row>capped atWO_CACHE_ROWS=10000(configurable). Avoids re-decoding on hot reads. Eviction on insert when full. Phase 12 ships without it if the bench numbers are fine; included here as a follow-on hatch. - Tombstoned offsets stay in the BTreeMap until compaction. A delete writes a tombstone to the segment + marks the BTreeMap entry as
SegmentOffset::Tombstone. List skips them. Counts as wasted space until a future compaction phase rewrites the segment. - Checkpoint =
fsyncevery active segment fd + advance control file. Runs everyCHECKPOINT_INTERVAL_SECS=60(configurable) and at clean shutdown. - MVCC stays out of scope. Subscriber pre-commit views are tick-boundary semantics, not version chains (per
../runtime/database/03-inmemory-engine.md). When a future phase adds LIVE subscriber predicate matching, version chains may join the engine — until then, the single-thread invariant gives us the same visibility guarantees for free. - Secondary indexes deferred. Phase 12 ships only the primary
idBTree.unique+indexschema attributes get their own follow-on phase.
Scope
Files rewritten inside crates/rt/src/
| File | Change |
|---|---|
engine.rs |
BTreeMap<i64, Row> → BTreeMap<i64, SegmentOffset>. Engine::get becomes seg_store.read(ty, offset)?. Engine::list walks the BTreeMap and preads each record (sequential — page cache makes it fast for the sample workload). Engine::create / update / delete keep the phase-10 segment append + phase-11 WAL append, but no longer keep the Row in memory. |
bin/wo.rs |
After WAL recovery, populate the BTreeMap with (id → offset) pairs by walking the recovered records. Also: spawn a TimerFd::periodic(CHECKPOINT_INTERVAL_SECS) registered on the event loop; the checkpoint step runs when the timer fires. |
New file inside crates/db/src/
| File | Responsibility | Approx LOC |
|---|---|---|
checkpoint.rs |
Checkpoint::run(seg_store, wal, control) — fsync every segment fd, write last_durable_lsn = wal.tail_lsn to the control file (rename-on-write), prune retired WAL segments older than the new LSN. |
~150 |
What SegmentOffset looks like
#[derive(Debug, Clone, Copy)]
enum SegmentOffset {
Live(u64), // byte offset in the segment file
Tombstone(u64), // ditto, but the row is logically deleted
}
A BTreeMap<i64, SegmentOffset> consumes ~24 B per entry (key + 16-byte enum). 10M rows → ~240 MiB index. Order-of-magnitude bigger than O(rowcount × pointer) because the enum carries a discriminant; collapse to u64 with a high-bit tombstone flag if memory pressure justifies it later.
Recovery (phase 11) becomes
fn recover(data_dir: &Path) -> Result<Engine> {
let ctl = ControlFile::read_or_initialize(data_dir)?;
let seg_store = SegStore::open(data_dir)?;
let wal = Wal::open(data_dir.join("wal"), ctl.last_durable_lsn)?;
let mut engine = Engine::new(catalog);
engine.attach(seg_store, wal);
// Walk the segments first to populate the offset index from durable rows.
for ty in engine.catalog().order.iter() {
for (id, offset, flags) in seg_store.iter(ty)? {
engine.index_mut(ty).insert(id, match flags {
Flags::ACTIVE => SegmentOffset::Live(offset),
Flags::TOMBSTONE => SegmentOffset::Tombstone(offset),
});
}
}
// Then replay any WAL records past the last checkpoint to catch up.
for rec in Replay::from(data_dir.join("wal"), ctl.last_durable_lsn)? {
engine.apply_replay(rec?)?;
}
Ok(engine)
}
The WAL replay still runs but covers a much smaller range — only what's been written since the last checkpoint. Recovery time is bounded by WAL volume between checkpoints, not by the entire history.
Checkpoint as a loop step
let cp_timer = TimerFd::periodic(Duration::from_secs(60))?;
eloop.register(cp_timer.as_raw_fd(), Interest::READABLE, Token(cp_timer.as_raw_fd() as u64))?;
// In serve_loop:
fd if fd == cp_timer.as_raw_fd() => {
let _ = cp_timer.read(); // drain timerfd's expirations
let mut eng = engine.lock().unwrap();
Checkpoint::run(&eng.seg_store, &eng.wal, &mut eng.control)?;
println!("[wo] checkpoint at LSN {}", eng.control.last_durable_lsn);
}
The phase-02 TimerFd::periodic already exists; this is the first runtime caller for it.
Bench
A small criterion-style microbench in crates/rt/benches/engine_disk.rs:
| Test | Target |
|---|---|
| Insert 100k rows (50-byte payload) | < 5 s wall, < 50 MiB RSS at end |
| Random read 100k rows under steady-state load | < 5 µs p50, < 100 µs p99 (page cache hot) |
| Cold-cache read 100k rows | < 200 µs p50 (one disk seek per read) |
| Recovery time after kill -9 mid-bench | < WAL_volume / disk_throughput, dominated by fdatasync round-trips |
criterion is normally an external crate; we're not adding deps. The bench is a #[test] with a --release runner — coarse but enough to catch regressions.
Cargo.toml delta
None — db and wal are already in from phases 10 and 11.
Exit criteria
cargo buildat root, four direct deps unchanged (anyhow,serde,serde_json,libc).- All existing unit tests still pass after the engine rewrite. The two heaviest are
engine::tests::crud_roundtrip_auto_id(port to verify offset semantics) andserver::tests::*(HTTP-level CRUD — should be unaffected). - End-to-end api.rest battery passes byte-identically — same status codes, same JSON bodies, same key ordering.
- Integration test
crates/rt/tests/disk_engine.rs:- Seed 10k rows of a 1 KiB payload type. Memory after seed (
/proc/self/statusVmRSS) is bounded byid_count × 24 B + listener_overhead, NOT by10000 × 1024. Specifically: less than 40 MiB. - Restart with kill -9 mid-write; recovery completes in < 1 s for a 16-MiB-WAL-segment workload.
- GET random ids — every read returns the right row, CRC verified.
- Seed 10k rows of a 1 KiB payload type. Memory after seed (
- Checkpoint smoke — start the binary, write 5 rows, wait
CHECKPOINT_INTERVAL_SECS+1seconds, verifydata/controlis updated (mtime moved,last_durable_lsnadvanced).strace -e fsync,renameduring the wait shows the checkpoint sequence. - Cold start with no
data/— a freshwo runon an empty data dir just works (creates the dir, no replay needed). Same fordata/+ empty WAL.
Non-scope
- No secondary indexes.
uniqueandindexschema attributes still trigger no extra storage. Future phase. - No compaction. Tombstoned offsets and old segment bytes accumulate. Trigger compaction is a separate phase keyed on a
dead-bytes / live-bytesratio. - No MVCC. Subscriber pre-commit views are tick-boundary semantics (
docs/runtime/database/03-inmemory-engine.md). Version chains land alongside the cross-shard subscription work in09c-per-shard-wal/09d-cross-shard-subscriptions. - No
O_DIRECT. Page cache is the cache. Per./exploration/linux/12-pwrite-fsync.md. - No
io_uringreads.preadsyscalls are short and the loop has no other work waiting; an async batched read API isn't worth its own complexity at this size. - No streaming list.
Engine::listreturns all rows for a type in one call. Pagination + cursor support is a future phase keyed on a real workload that hits the wall.
Verification
cargo build
cargo test --lib # rt + db + wal unit tests
cargo test --test disk_engine # the new integration test
cargo test --release --test disk_engine -- --nocapture # bench numbers visible
# manual end-to-end
cargo run --release --bin wo -- run docs/examples/blog &
PID=$!
sleep 1
# Seed 10000 rows
for i in $(seq 1 10000); do
curl -sf -X POST http://127.0.0.1:8080/api/articles \
-H 'Content-Type: application/json' \
-d "{\"slug\":\"s$i\",\"title\":\"T$i\",\"author\":1,\"published\":true,\"meta\":{\"excerpt\":\"\",\"body_md\":\"\"}}" \
> /dev/null
done
# Memory check
ps -o rss= -p $PID # expect under ~50 MiB even with 10k rows × 1 KiB each
# Wait for checkpoint
sleep 65
ls -la docs/examples/blog/data/control # mtime should be recent
kill -INT $PID
# Cold restart
cargo run --release --bin wo -- run docs/examples/blog &
PID=$!
sleep 1
curl -s 'http://127.0.0.1:8080/api/articles' | python3 -c 'import json,sys;print(len(json.load(sys.stdin)))'
# expect: 10000
kill -INT $PID
cd .dev/reference/crates && cargo build && cargo test # v1 untouched
After this phase
The single-thread runtime is durable, RAM-bounded, and recovery-fast. Phases 13+ pivot to layering features on top: secondary indexes, compaction, query-layer integration, then the 09a-09f scaleout sequence which lifts the same primitives into per-shard form. The empty crates/{value, engine, txn} skeletons get populated as their phases activate; wal/ and db/ are now real code, used by rt/.
The crates/rt/Cargo.toml direct dep list at the end of phase 12 is anyhow + serde + serde_json + libc + db + wal. Phase 05 collapses serde + serde_json into the hand-rolled JSON module; phase 06 collapses anyhow into a bespoke error type. The libc + path-deps end state from ./done/01-scafolding-crates.md is reachable in two more phases past 12.