- 05: hand-rolled JSON - 06: bespoke error type - 07: inotify content watcher - 08: sendfile static assets - 09: concurrency scaleout - 10: storage foundations - 11: WAL and recovery - 12: engine disk cutover - 14: MVC UI implementation - 15: MCP streamable HTTP - 16: Postgres mirror
11 KiB
11 — WAL + crash recovery
Kanban: ⬜ not started (scope reduced) — replay + ack-after-fsync + group commit landed via 09c and its follow-ups; remaining here: snapshots (
.data), compaction, WAL rotation. Board: 00-kanban.md
Context sources: ./10-storage-foundations.md, ../runtime/database/02-wo-language.md#concurrency-model, ../runtime/database/03-inmemory-engine.md, ./exploration/postgresql/wal.md, ./exploration/postgresql/buffer-and-checkpoint.md, ./exploration/linux/12-pwrite-fsync.md, ../02-recovery.md.
Goal
kill -9 mid-write loses nothing acknowledged. On restart, recovery replays the WAL into the in-memory HashMap and the engine serves traffic exactly as if nothing had happened. The engine is still HashMap-backed in this phase — phase 12 changes that. Phase 11 wires durability without changing the engine's read shape.
This is the phase where the locked architecture statement from 02-wo-language.md § Concurrency Model becomes code:
"On
COMMIT, the WAL record must befsync'd before the client gets acknowledgment. Group commit drains many pending commits into one fsync SQE per tick."
Design decisions (locked)
- WAL is separate from the segment files. Segments hold post-recovery row data; the WAL is the durability log we replay from. Per
./exploration/postgresql/wal.md. Different fsync cadence (every commit for WAL, every checkpoint for segments). - LSN = monotonic byte offset across all WAL segments. Postgres convention. 64-bit. Simple
<comparisons. Matches the placeholder slot phase 10 reserved in the frame header. - Group commit via the loop tick. No separate writer thread. Every loop tick: drain all pending commits, one
fdatasynccovers all of them, then ack each request. Same effect as Postgres' group-commit fence; the fence is the tick boundary. fdatasync, notfsync, for the WAL. WAL files areposix_fallocate'd up front to a fixed segment size — writes never extend them, so the inode metadata doesn't change andfdatasyncis sufficient. Per./exploration/linux/12-pwrite-fsync.md.- Control file via rename-on-write.
data/control.tmp→fsync→rename→ parent-dirfsync. Atomic across crashes. - Replay is idempotent. Recovery replays records starting at
last_durable_lsn; partial replay (crash mid-recovery) re-replays from the same anchor with the same effect. - Module at
crates/wal/.crates/wal/'s placeholder doc-comment names this work. Populated here.
Scope
New files inside crates/wal/src/
| File | Responsibility | Approx LOC |
|---|---|---|
lib.rs |
Re-exports Wal, Lsn, Replay, ControlFile. |
~30 |
lsn.rs |
pub struct Lsn(pub u64) — newtype with Display, ordering, segment-id + offset accessors. |
~50 |
wal.rs |
Wal { dir, active_fd, active_seg_id, tail_lsn, pending: Vec<PendingCommit> }. append(rec) -> Lsn, commit() -> io::Result<Lsn> (issues fdatasync), enqueue_ack(fd) / drain_acks() -> Vec<RawFd>. |
~250 |
segment.rs |
WAL-file rollover: open new <seg-id>.wal, posix_fallocate to 16 MiB, switch active fd, retire the previous segment. |
~120 |
control.rs |
ControlFile { magic, version, last_durable_lsn, crc } — read on startup, write on checkpoint. Rename-on-write. |
~120 |
replay.rs |
Replay::from(dir, last_lsn) -> Iterator<Item = Result<Record>> — walks WAL forward, yields decoded records to the caller. |
~150 |
Total: ~720 LOC. No v1 precedent — wo-wal doesn't exist (despite being named in the database series). New territory.
File layout
docs/examples/blog/
└── data/
├── control ← 32-byte fixed-size; updated atomically
├── Article.seg ← phase 10's segment files
├── ...
└── wal/
├── 0000000000000001.wal ← active WAL segment (16 MiB fallocated)
└── 0000000000000002.wal ← created at rollover
Control file format (32 bytes)
[u8 magic[4] = b"WOCT"]
[u8 version = 1]
[u8 _pad[3] = 0]
[u64 last_durable_lsn LE]
[u64 created_unix_seconds LE]
[u32 crc32c]
Atomic update sequence (per ./exploration/postgresql/buffer-and-checkpoint.md):
fs::write("control.tmp", &bytes)?;
let f = File::open("control.tmp")?;
unsafe { libc::fsync(f.as_raw_fd()); }
fs::rename("control.tmp", "control")?;
let dfd = unsafe { libc::open(data_dir.as_ptr(), libc::O_RDONLY) };
unsafe { libc::fsync(dfd); libc::close(dfd); }
Engine integration
Engine::create / update / delete — each calls into the WAL after the in-memory mutation succeeds and the segment append (phase 10) succeeds:
let lsn = self.wal.append(WalRecord::Mutation { ty, op, row })?;
self.wal.enqueue_ack(/* request fd */ fd);
// loop tick later: drain_acks() runs after commit() fsyncs.
The HTTP handler in crates/rt/src/server.rs becomes:
fn create_h(engine: &Shared, ty: &str, req: &Request, _params: &RouteParams) -> Response {
let body = parse_json_body(req);
let mut eng = engine.lock().unwrap();
match eng.create(ty, body) {
Ok(row) => {
// Engine's create now returns *after* the WAL append, but BEFORE
// the fsync. The response is held until the next tick's group commit.
Response::deferred(Status::CREATED, json!(row))
}
Err(e) => Response::status(Status::BAD_REQUEST).text(e.to_string()),
}
}
Response::deferred is a new variant — the response object is stashed on the connection, but the wire bytes aren't sent until wal.drain_acks() returns this fd. Phase 11 introduces this concept; phase 12 keeps it.
Recovery on startup
fn recover(data_dir: &Path) -> Result<Engine> {
let ctl = ControlFile::read_or_initialize(data_dir)?;
let mut engine = Engine::new(catalog);
let mut max_seen = ctl.last_durable_lsn;
for rec in Replay::from(data_dir.join("wal"), ctl.last_durable_lsn)? {
let rec = rec?;
match rec.payload {
WalRecord::Mutation { ty, op, row } => engine.apply_replay(ty, op, row)?,
}
max_seen = rec.lsn;
}
// Don't advance the control file yet — checkpoint (phase 12+) does that.
println!("[wo] recovered {} records, tail LSN {}", count, max_seen);
Ok(engine)
}
Engine::apply_replay is Engine::create / update / delete minus the wal.append callback (already-replayed records re-applied don't get re-WAL'd).
Group commit — the loop integration
crates/rt/src/bin/wo.rs's serve_loop gains:
'outer: loop {
let events = eloop.wait_once(Some(Duration::from_secs(60)))?;
for ev in events { /* dispatch as before */ }
// Group commit fence — runs once per tick after request dispatch.
if engine.lock().unwrap().wal.pending_commits() > 0 {
let _ = engine.lock().unwrap().wal.commit(); // one fdatasync
for fd in engine.lock().unwrap().wal.drain_acks() {
// Mark the connection writable; its queued response now flushes.
eloop.modify(fd, Interest::READ_WRITE, Token(fd as u64))?;
}
}
}
Cargo.toml delta
crates/rt/Cargo.toml adds wal as a path dep alongside db. Workspace adds crates/wal to the members list.
Exit criteria
cargo buildat root —rt,db,walall compile.- Unit tests in
crates/wal/src/:wal_append_assigns_monotonic_lsn— successive appends produce strictly increasing LSNs.wal_rollover_at_segment_cap— appending pastWAL_SEG_SIZEopens segment 2 without losing tail.replay_yields_records_in_order— write 100, replay returns 100 in LSN order.control_file_rename_on_write— kill -9 between tmp-write and rename leaves old control intact.crc_mismatch_aborts_replay— corrupting one byte in WAL aborts replay withCrcMismatch.
- Integration test
crates/rt/tests/wal_recovery.rs:- Starts
wo run docs/examples/blogwithWO_LISTEN=127.0.0.1:0. - POSTs 100 articles via the http stack.
- Sends
kill -9to the binary. - Restarts; GETs all 100 back.
- Starts
- The api.rest 20-assertion battery still passes byte-identically under
WO_FSYNC=on(default). WO_FSYNC=offenv var — when set, skips thefdatasyncfor tests that don't care about durability. Drops cold-restart-recovery latency to zero.strace -e fdatasync,fsync,renameduring a 5-commit run shows ~5fdatasynccalls (one per tick), nofsync(no rollover, no checkpoint yet), norename(no control update yet — that's phase 12+).
Non-scope
- No checkpoint loop yet. Phase 11 reads the control file at startup and writes it at clean shutdown only; periodic checkpoint lands with phase 12 or shortly after. Until then, recovery walks the entire WAL on every restart — fine for a sample workload, expensive for production.
- No segment compaction. Old WAL segments stay on disk forever in this phase. A future phase truncates after a checkpoint advances the control file past them.
- No
io_uring. Synchronouspwrite+fdatasync.io_uringbecomes interesting when the loop drains many fds per tick; phase 11's commit cadence doesn't need it. Layered on later. - No partial-record handling on torn writes. A WAL segment is
posix_fallocate'd up front, so partial-write torn-record on the leading edge of the file is the only scenario; the CRC trailer detects it and replay stops cleanly. - No multi-process recovery. Single-binary invariant.
- No engine cutover to disk reads.
Engine::list / getstill walk the in-memoryHashMap— phase 12.
Verification
cargo build # rt + db + wal
cargo test --lib # all unit tests green
cargo test --test wal_recovery # the kill-9 integration test
# durability smoke (manual)
WO_LISTEN=127.0.0.1:8765 cargo run --bin wo -- run docs/examples/blog &
PID=$!
sleep 1
for i in 1 2 3 4 5; do
curl -sf -X POST http://127.0.0.1:8765/api/articles \
-H 'Content-Type: application/json' \
-d "{\"slug\":\"a$i\",\"title\":\"A$i\",\"author\":1,\"published\":true,\"meta\":{\"excerpt\":\"\",\"body_md\":\"\"}}" \
> /dev/null
done
kill -9 $PID
WO_LISTEN=127.0.0.1:8765 cargo run --bin wo -- run docs/examples/blog &
PID=$!
sleep 1
curl -s http://127.0.0.1:8765/api/articles | python3 -c "import json,sys;print(len(json.load(sys.stdin)))"
# expect: 5
kill -INT $PID
# strace check
strace -e fdatasync,fsync,rename -f -p $(pgrep -f 'target/debug/wo run') 2>&1 | head -20
# v1 untouched
cd .dev/reference/crates && cargo build && cargo test
After this phase
Durability is real but the engine is still HashMap<type, BTreeMap<i64, Row>> — every row, in full, lives in RAM. Phase 12 swaps the in-memory Row for an offset into the segment file and adds checkpoints, completing the transition to a durable, RAM-bounded engine. After phase 12 the runtime can serve a 10× larger dataset than fits in RAM without a redesign.