explore postgres code base

This commit is contained in:
shoney.arickathil 2026-05-05 00:26:49 +02:00
parent ce02f742fc
commit c8f76484f4
14 changed files with 1064 additions and 12 deletions

2
.gitignore vendored
View file

@ -17,8 +17,10 @@
# Each contributor sets their own via:
# ln -s <path-to-linux-src> reference/linux
# ln -s <path-to-go-src> reference/go
# ln -s <path-to-postgresql-src> reference/postgresql
/reference/linux
/reference/go
/reference/postgresql
# Editor / OS noise — left broad on purpose so a contributor doesn't
# accidentally commit their IDE scratch or macOS metadata.

View file

@ -156,3 +156,17 @@ The transition from current to target doesn't have to be all-or-nothing:
3. **Phase 3** — Collapse repositories. Move frontend into the unified codebase. Ship as a single binary that serves both API and static assets.
Each phase produces a working system. The current architecture can run in parallel until the new one is ready.
## Implementation phases
The "embedded storage engine" of Phase 1 above lands in three numbered plan docs under [`docs/plan/`](./plan/):
| Phase | Doc | What it ships |
| --- | --- | --- |
| 10 | [`plan/10-storage-foundations.md`](./plan/10-storage-foundations.md) | On-disk row codec (length-prefix + flags + LSN + CRC32C); per-type segment files (`data/<TypeName>.seg`); `posix_fallocate` preallocation; `pwrite`-only append path. Reads still in-memory. |
| 11 | [`plan/11-wal-and-recovery.md`](./plan/11-wal-and-recovery.md) | WAL log with `fdatasync` at commit; group commit per loop tick; control file with `last_durable_lsn` (rename-on-write); replay loop on startup. `kill -9` mid-write loses nothing acknowledged. |
| 12 | [`plan/12-engine-disk-cutover.md`](./plan/12-engine-disk-cutover.md) | `Engine`'s row payload moves to disk; in-memory map becomes `BTreeMap<i64, SegmentOffset>`. Periodic checkpoint flushes segments + advances the control file. RAM bounded by id-count, not row size. |
Postgres' storage subsystem is the design reference — see [`docs/plan/exploration/postgresql/`](./plan/exploration/postgresql/) for which Postgres modules informed which decision and what writeonce skips (multi-process IPC, latches, separate writer processes).
The durability syscalls themselves live in [`docs/plan/exploration/linux/12-pwrite-fsync.md`](./plan/exploration/linux/12-pwrite-fsync.md).

View file

@ -8,7 +8,7 @@ Phases 02–08 produce a single-threaded event-loop runtime with zero external R
This phase is the refinement of the Phase-2 concurrency doctrine "shard to scale past one core" into a concrete architecture. The stance stays the same — **no Go-style goroutines, no work-stealing across threads, no shared mutable heap** — but now we have multiple event loops, each owning its core, its share of connections, and its slice of engine state.
This doc is a **master plan**. It outlines sub-phases 10–15 at a high level; each sub-phase lands as its own numbered plan doc when implementation starts. No code changes in this pass.
This doc is a **master plan**. It outlines sub-phases A–F at a high level; each sub-phase lands as its own plan doc (`09a-…`, `09b-…`, …) when implementation starts. No code changes in this pass. The numerical phase slots 10/11/12 are taken by the storage roadmap ([`./10-storage-foundations.md`](./10-storage-foundations.md), [`./11-wal-and-recovery.md`](./11-wal-and-recovery.md), [`./12-engine-disk-cutover.md`](./12-engine-disk-cutover.md)) — single-thread durability has to land before the per-shard WAL of `09c`.
## Goal
@ -63,31 +63,31 @@ Reference cards already exist for most; this phase adds the ones that are cross-
Each one lands as its own numbered plan doc when ready for implementation. Smoke test (`cargo run --bin wo -- run docs/examples/ecommerce` serves correctly) stays green after every sub-phase.
### `10-thread-per-core.md` — N event loops, `SO_REUSEPORT`
### `09a-thread-per-core.md` — N event loops, `SO_REUSEPORT`
Introduce a thread-pool manager at `crates/rt/src/runtime/scheduler.rs` (Go parallel: `proc.go`). Spawn `WO_THREADS` OS threads at boot; each pins itself and runs an `EventLoop`. Replace the single `Listener` with per-thread listeners bound `SO_REUSEPORT` to the same port. State is still global at first (shared `Arc<Mutex<Engine>>`) — one thing at a time. Exit criterion: `wo run` boots N threads visible in `ps -T`, accepts load balanced across them per `ss -tnp`, no regression in the 20-assertion blog smoke.
### `11-sharded-engine.md` — per-thread engine state
### `09b-sharded-engine.md` — per-thread engine state
Partition the in-memory engine catalog + row BTreeMaps by shard id (= thread id). Shard key is `customer.id` for ecommerce / `author.id` for blog / per-type default for anything else. Add a shard router in front of every REST/WS handler: resolve the shard from the request's identifying field, send an in-process message to that thread's mailbox, await response. Shared `Arc<Mutex<Engine>>` goes away; each thread owns its slice. Cross-shard reads (admin `list orders`) fan out to every thread and merge results.
### `12-per-shard-wal.md` — one WAL file per shard
### `09c-per-shard-wal.md` — one WAL file per shard
Each thread has its own `foo.wal` + `foo.data` + per-thread `io_uring` ring (phase 3's durability work, repeated per shard). Recovery is parallel across threads. No shared WAL writer thread. Group commit is per-thread.
Each thread has its own `foo.wal` + `foo.data` + per-thread `io_uring` ring ([phase 11's durability work](./11-wal-and-recovery.md), repeated per shard). Recovery is parallel across threads. No shared WAL writer thread. Group commit is per-thread.
### `13-cross-shard-subscriptions.md` — LIVE fanout
### `09d-cross-shard-subscriptions.md` — LIVE fanout
A commit on shard K that creates/updates rows of type T needs to wake subscribers on every shard watching T. Via broadcast: K writes the delta to a per-subscriber-thread mailbox — one message per destination thread, not per subscriber. The destination thread then does the fine-grained predicate match against its local subscription table. Avoids N² traffic when N connections watch the same stream.
### `14-cross-shard-txn.md` — 2PC for transactions that span shards
### `09e-cross-shard-txn.md` — 2PC for transactions that span shards
`fn checkout(customer, product, qty)` might touch shards A (customer), B (product), and C (order) if they hash differently. The transaction coordinator (already designed in [`../runtime/database/02-wo-language.md`](../runtime/database/02-wo-language.md) § Cross-Paradigm Transaction Coordinator) generalises to cross-shard: `begin(snapshot_ts)` broadcasts to all participating shards, `prepare()` collects votes, `commit(wal_lsn)` atomically flips markers, `abort()` if any participant refuses. The per-shard WAL entries carry the 2PC state machine.
### `15-observability-and-rebalance.md` — ops
### `09f-observability-and-rebalance.md` — ops
Per-shard metrics (connections, ops/s, p99, WAL lag), Prometheus scrape endpoint on one well-known thread. A `WO_RESHARD` admin command migrates a contiguous customer-id range from shard K to shard K′ via state snapshot → replay → cutover. For a fixed-core deployment this is rare; matters when `WO_THREADS` changes between runs.
## Verification targets (after 15 lands)
## Verification targets (after `09f` lands)
Ecommerce sample on an 8-core box with `WO_THREADS=8`:
@ -104,7 +104,7 @@ Ecommerce sample on an 8-core box with `WO_THREADS=8`:
## Non-scope
- **No Go-style goroutines, even after this phase.** Adding M:N scheduling is not on the roadmap. When one core runs out, add more cores (more threads) — horizontally, thread-per-core.
- **No distributed (multi-node) sharding.** This phase is single-box only. Redis-Cluster-style network sharding is a separate future phase; the in-process shard bus (phase 10's mailboxes) is not the same thing as a cluster membership protocol.
- **No distributed (multi-node) sharding.** This phase is single-box only. Redis-Cluster-style network sharding is a separate future phase; the in-process shard bus (`09a`'s mailboxes) is not the same thing as a cluster membership protocol.
- **No dynamic thread count at runtime.** `WO_THREADS` is set at boot and pinned. Adding/removing a thread means a rolling restart. Acceptable for a database; fundamental to the zero-contention model.
- **No work-stealing.** A slow handler on thread A does not get rebalanced to thread B. Back-pressure is the thread-local queue filling up. If one thread hot-spots because of a bad shard key, the fix is to reshard — not to steal.
- **No `std::thread::available_parallelism` on exotic hosts.** `WO_THREADS` override covers kubernetes CFS-bound pods, NUMA partitioning, and single-core debug runs.

View file

@ -0,0 +1,146 @@
# 10 — Storage Foundations: on-disk row codec + segment append path
**Context sources:** [`./done/04-cutover-remove-tokio-axum.md`](./done/04-cutover-remove-tokio-axum.md), [`../runtime/database/03-inmemory-engine.md`](../runtime/database/03-inmemory-engine.md), [`../runtime/database/07-wo-seg-migration.md`](../runtime/database/07-wo-seg-migration.md), [`./exploration/postgresql/smgr-and-md.md`](./exploration/postgresql/smgr-and-md.md), [`./exploration/postgresql/page-format.md`](./exploration/postgresql/page-format.md), [`./exploration/linux/12-pwrite-fsync.md`](./exploration/linux/12-pwrite-fsync.md), [`./exploration/linux/09-fallocate.md`](./exploration/linux/09-fallocate.md), [`reference/crates/wo-seg/src/`](../../reference/crates/wo-seg/src/).
## Goal
Every engine mutation appends a typed record to a per-type segment file on disk. **Reads still hit the in-memory `HashMap` — no behaviour change visible to clients yet.** Killing the process after a write leaves a real `data/<TypeName>.seg` on disk; restart re-creates an empty `HashMap` and ignores the segment (recovery is phase 11). This phase only proves the **on-disk row format**.
Lays the codec + filesystem layout that phase 11 (WAL + recovery) and phase 12 (disk-backed engine) build on top of.
## Design decisions (locked)
1. **One segment file per type.** `data/<TypeName>.seg`. No per-record file proliferation, no per-database tablespaces, no relfilenode indirection (per [`./exploration/postgresql/smgr-and-md.md`](./exploration/postgresql/smgr-and-md.md) — Postgres' multi-file model exists for multi-tenant ops; writeonce binds to one data dir per `wo run`).
2. **Append-only with tombstone byte.** Updates and deletes append a new record (with the old one's id + a `TOMBSTONE` flag); compaction is a follow-on phase. Same model as v1 wo-seg.
3. **Length-prefix framing with CRC32C trailer.** `[u32 length LE][u8 flags][u8 record_kind][u64 LSN][payload bytes][u32 CRC32C]`. The CRC trailer is the **one design point where writeonce diverges from v1 wo-seg**: wo-seg skipped checksums; we don't.
4. **Payload codec is `serde_json` for now.** Phase 05 (hand-rolled JSON) swaps it; the codec slot is a single `RowCodec` trait so the swap is mechanical.
5. **`posix_fallocate` to 1 MiB at file creation.** Doubles when full. Avoids `ENOSPC` mid-write and minimizes filesystem-level fragmentation. Per [`./exploration/linux/09-fallocate.md`](./exploration/linux/09-fallocate.md).
6. **`pwrite` for the append, no fsync yet.** This phase does not commit a durability barrier — the bytes land in the OS page cache and that's it. Phase 11 adds the fsync. Lets us validate the format without conflating it with fsync semantics.
7. **Module at `crates/db/`, not extracted from `rt`.** `crates/db/` has been a placeholder since the scaffolding phase — this phase populates it. Other crates (`engine`, `value`, `wal`, `txn`) stay placeholders until their phases activate.
## Scope
### New files inside `crates/db/src/`
| File | Responsibility | Approx LOC |
| --- | --- | --- |
| `lib.rs` | Re-exports `SegStore`, `Frame`, `Flags`, `RecordKind`, `RowCodec`, `LSN`. Replaces today's empty `lib.rs` doc-comment. | ~30 |
| `frame.rs` | `Frame` struct + `encode(payload, flags, kind, lsn) -> Vec<u8>` + `decode(bytes) -> Result<Frame>` with CRC verification. | ~150 |
| `crc.rs` | CRC32C via the SSE 4.2 `crc32c.h` algorithm. Software fallback for older CPUs. ~80 lines hand-rolled vs. pulling a crate. | ~80 |
| `codec.rs` | `trait RowCodec { fn encode(&self, row: &Row, buf: &mut Vec<u8>); fn decode(&self, bytes: &[u8]) -> Result<Row>; }` + `JsonCodec` impl backed by today's `serde_json`. | ~50 |
| `seg.rs` | `SegStore { dir: PathBuf, fds: HashMap<String, RawFd>, tails: HashMap<String, u64> }`. `open(dir)`, `append(ty, &Row) -> Result<u64-offset>`, `read(ty, offset) -> Result<Row>` (used by phase 11 recovery, not by the engine yet). | ~250 |
Total: ~560 LOC. The framing math + fallocate + pwrite plumbing is ported from [`reference/crates/wo-seg/src/{writer.rs,reader.rs,header.rs}`](../../reference/crates/wo-seg/src/) with the CRC trailer added.
### File layout written under `<wo_run_dir>/`
```
docs/examples/blog/
├── app.wo
├── ui/...
└── data/ ← created by phase 10
├── Article.seg
├── Author.seg
├── Comment.seg
└── Tag.seg
```
`data/` is gitignored (already covered by `/data` and `/docs/examples/*/data` in `.gitignore`). Empty when no rows exist; created lazily on first write.
### Record framing (illustrated)
```text
┌─ length excludes itself; covers flags..CRC.
▼
[u32 length LE][u8 flags][u8 kind][u64 LSN][payload bytes ...][u32 CRC32C]
│ │
│ └─ 0x00 = ROW, 0x01 = TOMBSTONE, others reserved
└─ 0x00 = ACTIVE, 0x01 = DELETED (per-record live bit)
```
`flags` is a per-record live bit — flip it to `DELETED` to soft-delete in place without rewriting the payload. `kind` is the discriminator for upcoming record kinds (phase 11 introduces `WAL_BEGIN`, `WAL_COMMIT`); for phase 10 every record is `ROW`. `LSN` is `0` until phase 11 starts assigning real LSNs — it's a placeholder slot now so phase 11 doesn't reshape the format.
### Engine integration
The `Engine::create / update / delete` methods in `crates/rt/src/engine.rs` get a `seg_store: Arc<Mutex<SegStore>>` field plumbed through `Engine::new`. After every successful in-memory mutation:
```rust
self.seg_store.lock().unwrap()
.append(ty, &row)
.map_err(|e| anyhow!("seg append: {e}"))?;
```
Failure aborts the whole mutation — the in-memory write is rolled back. This phase does NOT introduce a "best-effort persistence" mode.
`Engine::list / get` remain unchanged; reads stay in-memory.
### `Cargo.toml` delta
```diff
[dependencies]
anyhow = "1"
serde = { version = "1", features = ["derive"] }
serde_json = "1"
libc = "0.2"
+
+[dependencies.db]
+path = "../db"
```
`crates/db/Cargo.toml` itself stays at `libc + serde_json` (the latter via `RowCodec`'s `JsonCodec`). When phase 05 lands, the `serde_json` import collapses into the runtime's hand-rolled `Value`.
The root workspace member list also activates: `crates/db` joins `crates/rt` as a non-empty member.
## Exit criteria
1. **`cargo build`** at root — both `crates/rt` and `crates/db` compile. Five direct deps (`anyhow`, `serde`, `serde_json`, `libc`, `db`).
2. **`cargo test --lib`** — all existing 37 `rt` tests still green; new `db` tests cover:
- `frame_roundtrip` — encode then decode produces the same `Frame`.
- `crc_detects_corruption` — flipping one byte in the payload makes `decode` return `CrcMismatch`.
- `seg_append_writes_to_disk` — `append` then re-`open` reads the same row back.
- `seg_grows_when_full` — appending past the initial 1 MiB triggers a fallocate-grow without losing existing records.
3. **End-to-end** — `cargo run --bin wo -- run docs/examples/blog`, `curl -X POST /api/articles` with a body, then `xxd docs/examples/blog/data/Article.seg | head -3` — output shows the magic length prefix and the JSON payload.
4. **`reference/rest/blog.rest`** — 20-assertion battery still passes byte-identically.
5. **Restart leaves the segment on disk but ignores it.** `wo run`, write 5 rows, ctrl-C, `wo run` again, `GET /api/articles` returns `[]`. The segment file still exists. Phase 11 will start replaying it.
## Non-scope
- **No fsync.** Pure write path; durability barrier is phase 11.
- **No WAL.** Mutations go straight to the segment. Phase 11 introduces a separate WAL log; segments become the post-checkpoint home for replayed records.
- **No reads from disk.** `Engine::get` stays in-memory. Phase 12 cuts over.
- **No secondary indexes.** Phase 12 introduces a primary `id` BTree on disk; secondary indexes (`unique`, `index` schema attributes) are a later phase.
- **No compaction.** Tombstoned records pile up. Compaction lands when a benchmark says it has to.
- **No cross-type transactions / RETURNING aliases.** The locked schema design (`02-wo-language.md`) names cross-paradigm transactions; the runtime gets there in a later phase.
- **No `crates/db` API stability.** Internal-only until `crates/db/Cargo.toml` declares `[lib]`-level external surfaces.
## Verification
```bash
cargo build # rt + db both compile
cargo test --lib # rt + db unit tests
cargo test -p db # db-only
# manual end-to-end
cargo run --bin wo -- run docs/examples/blog &
PID=$!
sleep 1
curl -s -X POST http://127.0.0.1:8080/api/articles \
-H 'Content-Type: application/json' \
-d '{"slug":"a","title":"A","author":1,"published":true,"meta":{"excerpt":"e","body_md":"b"}}'
ls -la docs/examples/blog/data/
xxd docs/examples/blog/data/Article.seg | head -5
kill -INT $PID
# restart sanity — phase 10 is "format-only", no replay
cargo run --bin wo -- run docs/examples/blog &
PID=$!
sleep 1
curl -s http://127.0.0.1:8080/api/articles # expect []
kill -INT $PID
cd reference/crates && cargo build && cargo test # v1 untouched
```
## After this phase
The on-disk format exists but is dead weight — written, never read. Phase 11 brings it to life: introduces a separate WAL log, fsync at commit, group commit per loop tick, and a recovery loop that replays the WAL into the in-memory `HashMap` on startup. Phase 12 then cuts the engine over to read from segments instead of from RAM, completing the transition from in-memory to durable storage.

View file

@ -0,0 +1,209 @@
# 11 — WAL + crash recovery
**Context sources:** [`./10-storage-foundations.md`](./10-storage-foundations.md), [`../runtime/database/02-wo-language.md#concurrency-model`](../runtime/database/02-wo-language.md#concurrency-model), [`../runtime/database/03-inmemory-engine.md`](../runtime/database/03-inmemory-engine.md), [`./exploration/postgresql/wal.md`](./exploration/postgresql/wal.md), [`./exploration/postgresql/buffer-and-checkpoint.md`](./exploration/postgresql/buffer-and-checkpoint.md), [`./exploration/linux/12-pwrite-fsync.md`](./exploration/linux/12-pwrite-fsync.md), [`../02-recovery.md`](../02-recovery.md).
## Goal
`kill -9` mid-write loses nothing acknowledged. On restart, recovery replays the WAL into the in-memory `HashMap` and the engine serves traffic exactly as if nothing had happened. **The engine is still `HashMap`-backed in this phase** — phase 12 changes that. Phase 11 wires durability without changing the engine's read shape.
This is the phase where the locked architecture statement from [`02-wo-language.md` § Concurrency Model](../runtime/database/02-wo-language.md#concurrency-model) becomes code:
> "On `COMMIT`, the WAL record must be `fsync`'d before the client gets acknowledgment. Group commit drains many pending commits into one fsync SQE per tick."
## Design decisions (locked)
1. **WAL is separate from the segment files.** Segments hold post-recovery row data; the WAL is the durability log we replay from. Per [`./exploration/postgresql/wal.md`](./exploration/postgresql/wal.md). Different fsync cadence (every commit for WAL, every checkpoint for segments).
2. **LSN = monotonic byte offset across all WAL segments.** Postgres convention. 64-bit. Simple `<` comparisons. Matches the placeholder slot phase 10 reserved in the frame header.
3. **Group commit via the loop tick.** No separate writer thread. Every loop tick: drain all pending commits, **one** `fdatasync` covers all of them, then ack each request. Same effect as Postgres' group-commit fence; the fence is the tick boundary.
4. **`fdatasync`, not `fsync`, for the WAL.** WAL files are `posix_fallocate`'d up front to a fixed segment size — writes never extend them, so the inode metadata doesn't change and `fdatasync` is sufficient. Per [`./exploration/linux/12-pwrite-fsync.md`](./exploration/linux/12-pwrite-fsync.md).
5. **Control file via rename-on-write.** `data/control.tmp` → `fsync` → `rename` → parent-dir `fsync`. Atomic across crashes.
6. **Replay is idempotent.** Recovery replays records starting at `last_durable_lsn`; partial replay (crash mid-recovery) re-replays from the same anchor with the same effect.
7. **Module at `crates/wal/`.** `crates/wal/`'s placeholder doc-comment names this work. Populated here.
## Scope
### New files inside `crates/wal/src/`
| File | Responsibility | Approx LOC |
| --- | --- | --- |
| `lib.rs` | Re-exports `Wal`, `Lsn`, `Replay`, `ControlFile`. | ~30 |
| `lsn.rs` | `pub struct Lsn(pub u64)` — newtype with `Display`, ordering, segment-id + offset accessors. | ~50 |
| `wal.rs` | `Wal { dir, active_fd, active_seg_id, tail_lsn, pending: Vec<PendingCommit> }`. `append(rec) -> Lsn`, `commit() -> io::Result<Lsn>` (issues `fdatasync`), `enqueue_ack(fd) / drain_acks() -> Vec<RawFd>`. | ~250 |
| `segment.rs` | WAL-file rollover: open new `<seg-id>.wal`, `posix_fallocate` to 16 MiB, switch active fd, retire the previous segment. | ~120 |
| `control.rs` | `ControlFile { magic, version, last_durable_lsn, crc }` — read on startup, write on checkpoint. Rename-on-write. | ~120 |
| `replay.rs` | `Replay::from(dir, last_lsn) -> Iterator<Item = Result<Record>>` — walks WAL forward, yields decoded records to the caller. | ~150 |
Total: ~720 LOC. No v1 precedent — wo-wal doesn't exist (despite being named in the database series). New territory.
### File layout
```
docs/examples/blog/
└── data/
├── control ← 32-byte fixed-size; updated atomically
├── Article.seg ← phase 10's segment files
├── ...
└── wal/
├── 0000000000000001.wal ← active WAL segment (16 MiB fallocated)
└── 0000000000000002.wal ← created at rollover
```
### Control file format (32 bytes)
```text
[u8 magic[4] = b"WOCT"]
[u8 version = 1]
[u8 _pad[3] = 0]
[u64 last_durable_lsn LE]
[u64 created_unix_seconds LE]
[u32 crc32c]
```
Atomic update sequence (per [`./exploration/postgresql/buffer-and-checkpoint.md`](./exploration/postgresql/buffer-and-checkpoint.md)):
```rust
fs::write("control.tmp", &bytes)?;
let f = File::open("control.tmp")?;
unsafe { libc::fsync(f.as_raw_fd()); }
fs::rename("control.tmp", "control")?;
let dfd = unsafe { libc::open(data_dir.as_ptr(), libc::O_RDONLY) };
unsafe { libc::fsync(dfd); libc::close(dfd); }
```
### Engine integration
`Engine::create / update / delete` — each calls into the WAL after the in-memory mutation succeeds and the segment append (phase 10) succeeds:
```rust
let lsn = self.wal.append(WalRecord::Mutation { ty, op, row })?;
self.wal.enqueue_ack(/* request fd */ fd);
// loop tick later: drain_acks() runs after commit() fsyncs.
```
The HTTP handler in `crates/rt/src/server.rs` becomes:
```rust
fn create_h(engine: &Shared, ty: &str, req: &Request, _params: &RouteParams) -> Response {
let body = parse_json_body(req);
let mut eng = engine.lock().unwrap();
match eng.create(ty, body) {
Ok(row) => {
// Engine's create now returns *after* the WAL append, but BEFORE
// the fsync. The response is held until the next tick's group commit.
Response::deferred(Status::CREATED, json!(row))
}
Err(e) => Response::status(Status::BAD_REQUEST).text(e.to_string()),
}
}
```
`Response::deferred` is a new variant — the response object is stashed on the connection, but the wire bytes aren't sent until `wal.drain_acks()` returns this fd. Phase 11 introduces this concept; phase 12 keeps it.
### Recovery on startup
```rust
fn recover(data_dir: &Path) -> Result<Engine> {
let ctl = ControlFile::read_or_initialize(data_dir)?;
let mut engine = Engine::new(catalog);
let mut max_seen = ctl.last_durable_lsn;
for rec in Replay::from(data_dir.join("wal"), ctl.last_durable_lsn)? {
let rec = rec?;
match rec.payload {
WalRecord::Mutation { ty, op, row } => engine.apply_replay(ty, op, row)?,
}
max_seen = rec.lsn;
}
// Don't advance the control file yet — checkpoint (phase 12+) does that.
println!("[wo] recovered {} records, tail LSN {}", count, max_seen);
Ok(engine)
}
```
`Engine::apply_replay` is `Engine::create / update / delete` minus the `wal.append` callback (already-replayed records re-applied don't get re-WAL'd).
### Group commit — the loop integration
`crates/rt/src/bin/wo.rs`'s `serve_loop` gains:
```rust
'outer: loop {
let events = eloop.wait_once(Some(Duration::from_secs(60)))?;
for ev in events { /* dispatch as before */ }
// Group commit fence — runs once per tick after request dispatch.
if engine.lock().unwrap().wal.pending_commits() > 0 {
let _ = engine.lock().unwrap().wal.commit(); // one fdatasync
for fd in engine.lock().unwrap().wal.drain_acks() {
// Mark the connection writable; its queued response now flushes.
eloop.modify(fd, Interest::READ_WRITE, Token(fd as u64))?;
}
}
}
```
### `Cargo.toml` delta
`crates/rt/Cargo.toml` adds `wal` as a path dep alongside `db`. Workspace adds `crates/wal` to the members list.
## Exit criteria
1. **`cargo build`** at root — `rt`, `db`, `wal` all compile.
2. **Unit tests in `crates/wal/src/`:**
- `wal_append_assigns_monotonic_lsn` — successive appends produce strictly increasing LSNs.
- `wal_rollover_at_segment_cap` — appending past `WAL_SEG_SIZE` opens segment 2 without losing tail.
- `replay_yields_records_in_order` — write 100, replay returns 100 in LSN order.
- `control_file_rename_on_write` — kill -9 between tmp-write and rename leaves old control intact.
- `crc_mismatch_aborts_replay` — corrupting one byte in WAL aborts replay with `CrcMismatch`.
3. **Integration test `crates/rt/tests/wal_recovery.rs`:**
- Starts `wo run docs/examples/blog` with `WO_LISTEN=127.0.0.1:0`.
- POSTs 100 articles via the http stack.
- Sends `kill -9` to the binary.
- Restarts; GETs all 100 back.
4. **The api.rest 20-assertion battery still passes byte-identically** under `WO_FSYNC=on` (default).
5. **`WO_FSYNC=off` env var** — when set, skips the `fdatasync` for tests that don't care about durability. Drops cold-restart-recovery latency to zero.
6. **`strace -e fdatasync,fsync,rename`** during a 5-commit run shows ~5 `fdatasync` calls (one per tick), no `fsync` (no rollover, no checkpoint yet), no `rename` (no control update yet — that's phase 12+).
## Non-scope
- **No checkpoint loop yet.** Phase 11 reads the control file at startup and writes it at clean shutdown only; periodic checkpoint lands with phase 12 or shortly after. Until then, recovery walks the entire WAL on every restart — fine for a sample workload, expensive for production.
- **No segment compaction.** Old WAL segments stay on disk forever in this phase. A future phase truncates after a checkpoint advances the control file past them.
- **No `io_uring`.** Synchronous `pwrite` + `fdatasync`. `io_uring` becomes interesting when the loop drains many fds per tick; phase 11's commit cadence doesn't need it. Layered on later.
- **No partial-record handling on torn writes.** A WAL segment is `posix_fallocate`'d up front, so partial-write torn-record on the leading edge of the file is the only scenario; the CRC trailer detects it and replay stops cleanly.
- **No multi-process recovery.** Single-binary invariant.
- **No engine cutover to disk reads.** `Engine::list / get` still walk the in-memory `HashMap` — phase 12.
## Verification
```bash
cargo build # rt + db + wal
cargo test --lib # all unit tests green
cargo test --test wal_recovery # the kill-9 integration test
# durability smoke (manual)
WO_LISTEN=127.0.0.1:8765 cargo run --bin wo -- run docs/examples/blog &
PID=$!
sleep 1
for i in 1 2 3 4 5; do
curl -sf -X POST http://127.0.0.1:8765/api/articles \
-H 'Content-Type: application/json' \
-d "{\"slug\":\"a$i\",\"title\":\"A$i\",\"author\":1,\"published\":true,\"meta\":{\"excerpt\":\"\",\"body_md\":\"\"}}" \
> /dev/null
done
kill -9 $PID
WO_LISTEN=127.0.0.1:8765 cargo run --bin wo -- run docs/examples/blog &
PID=$!
sleep 1
curl -s http://127.0.0.1:8765/api/articles | python3 -c "import json,sys;print(len(json.load(sys.stdin)))"
# expect: 5
kill -INT $PID
# strace check
strace -e fdatasync,fsync,rename -f -p $(pgrep -f 'target/debug/wo run') 2>&1 | head -20
# v1 untouched
cd reference/crates && cargo build && cargo test
```
## After this phase
Durability is real but the engine is still `HashMap<type, BTreeMap<i64, Row>>` — every row, in full, lives in RAM. Phase 12 swaps the in-memory `Row` for an offset into the segment file and adds checkpoints, completing the transition to a durable, RAM-bounded engine. After phase 12 the runtime can serve a 10× larger dataset than fits in RAM without a redesign.

View file

@ -0,0 +1,173 @@
# 12 — Engine cutover: rows live on disk
**Context sources:** [`./10-storage-foundations.md`](./10-storage-foundations.md), [`./11-wal-and-recovery.md`](./11-wal-and-recovery.md), [`../runtime/database/03-inmemory-engine.md`](../runtime/database/03-inmemory-engine.md), [`../runtime/database/07-wo-seg-migration.md`](../runtime/database/07-wo-seg-migration.md), [`./exploration/postgresql/buffer-and-checkpoint.md`](./exploration/postgresql/buffer-and-checkpoint.md), [`./exploration/postgresql/page-format.md`](./exploration/postgresql/page-format.md), [`./exploration/linux/12-pwrite-fsync.md`](./exploration/linux/12-pwrite-fsync.md).
## Goal
`Engine`'s row payload is no longer in RAM. The in-memory map is `HashMap<type, BTreeMap<i64, SegmentOffset>>`. Reads `pread` against the segment file and verify the CRC. RAM footprint is bounded by **id count + per-id overhead**, independent of row payload size — the runtime can now serve a dataset 10× larger than RAM.
A periodic checkpoint flushes dirty segments and advances the control-file LSN, bounding recovery time on restart.
## Design decisions (locked)
1. **In-memory index is `BTreeMap<i64, SegmentOffset>`.** Keeps the existing `Engine::list` insertion-order iteration. Roughly 24 bytes per entry (i64 key + u64 value + tree node overhead) — a million rows fits in 24 MiB regardless of row size.
2. **Reads via `pread` + decode + CRC verify.** No user-space buffer pool — the OS page cache is the cache (per [`./exploration/postgresql/buffer-and-checkpoint.md`](./exploration/postgresql/buffer-and-checkpoint.md)). Hot rows hit cached pages and the syscall returns memcpy-fast.
3. **Bounded LRU on top of pread.** Optional small `HashMap<(ty, offset), Row>` capped at `WO_CACHE_ROWS=10000` (configurable). Avoids re-decoding on hot reads. Eviction on insert when full. **Phase 12 ships without it** if the bench numbers are fine; included here as a follow-on hatch.
4. **Tombstoned offsets stay in the BTreeMap until compaction.** A delete writes a tombstone to the segment + marks the BTreeMap entry as `SegmentOffset::Tombstone`. List skips them. Counts as wasted space until a future compaction phase rewrites the segment.
5. **Checkpoint = `fsync` every active segment fd + advance control file.** Runs every `CHECKPOINT_INTERVAL_SECS=60` (configurable) and at clean shutdown.
6. **MVCC stays out of scope.** Subscriber pre-commit views are tick-boundary semantics, not version chains (per [`../runtime/database/03-inmemory-engine.md`](../runtime/database/03-inmemory-engine.md)). When a future phase adds LIVE subscriber predicate matching, version chains may join the engine — until then, the single-thread invariant gives us the same visibility guarantees for free.
7. **Secondary indexes deferred.** Phase 12 ships only the primary `id` BTree. `unique` + `index` schema attributes get their own follow-on phase.
## Scope
### Files rewritten inside `crates/rt/src/`
| File | Change |
| --- | --- |
| `engine.rs` | `BTreeMap<i64, Row>` → `BTreeMap<i64, SegmentOffset>`. `Engine::get` becomes `seg_store.read(ty, offset)?`. `Engine::list` walks the BTreeMap and `pread`s each record (sequential — page cache makes it fast for the sample workload). `Engine::create / update / delete` keep the phase-10 segment append + phase-11 WAL append, but no longer keep the `Row` in memory. |
| `bin/wo.rs` | After WAL recovery, populate the BTreeMap with `(id → offset)` pairs by walking the recovered records. Also: spawn a `TimerFd::periodic(CHECKPOINT_INTERVAL_SECS)` registered on the event loop; the checkpoint step runs when the timer fires. |
### New file inside `crates/db/src/`
| File | Responsibility | Approx LOC |
| --- | --- | --- |
| `checkpoint.rs` | `Checkpoint::run(seg_store, wal, control)` — fsync every segment fd, write `last_durable_lsn = wal.tail_lsn` to the control file (rename-on-write), prune retired WAL segments older than the new LSN. | ~150 |
### What `SegmentOffset` looks like
```rust
#[derive(Debug, Clone, Copy)]
enum SegmentOffset {
Live(u64), // byte offset in the segment file
Tombstone(u64), // ditto, but the row is logically deleted
}
```
A `BTreeMap<i64, SegmentOffset>` consumes ~24 B per entry (key + 16-byte enum). 10M rows → ~240 MiB index. Order-of-magnitude bigger than `O(rowcount × pointer)` because the enum carries a discriminant; collapse to `u64` with a high-bit tombstone flag if memory pressure justifies it later.
### Recovery (phase 11) becomes
```rust
fn recover(data_dir: &Path) -> Result<Engine> {
let ctl = ControlFile::read_or_initialize(data_dir)?;
let seg_store = SegStore::open(data_dir)?;
let wal = Wal::open(data_dir.join("wal"), ctl.last_durable_lsn)?;
let mut engine = Engine::new(catalog);
engine.attach(seg_store, wal);
// Walk the segments first to populate the offset index from durable rows.
for ty in engine.catalog().order.iter() {
for (id, offset, flags) in seg_store.iter(ty)? {
engine.index_mut(ty).insert(id, match flags {
Flags::ACTIVE => SegmentOffset::Live(offset),
Flags::TOMBSTONE => SegmentOffset::Tombstone(offset),
});
}
}
// Then replay any WAL records past the last checkpoint to catch up.
for rec in Replay::from(data_dir.join("wal"), ctl.last_durable_lsn)? {
engine.apply_replay(rec?)?;
}
Ok(engine)
}
```
The WAL replay still runs but covers a much smaller range — only what's been written since the last checkpoint. Recovery time is bounded by WAL volume between checkpoints, not by the entire history.
### Checkpoint as a loop step
```rust
let cp_timer = TimerFd::periodic(Duration::from_secs(60))?;
eloop.register(cp_timer.as_raw_fd(), Interest::READABLE, Token(cp_timer.as_raw_fd() as u64))?;
// In serve_loop:
fd if fd == cp_timer.as_raw_fd() => {
let _ = cp_timer.read(); // drain timerfd's expirations
let mut eng = engine.lock().unwrap();
Checkpoint::run(&eng.seg_store, &eng.wal, &mut eng.control)?;
println!("[wo] checkpoint at LSN {}", eng.control.last_durable_lsn);
}
```
The phase-02 `TimerFd::periodic` already exists; this is the first runtime caller for it.
### Bench
A small criterion-style microbench in `crates/rt/benches/engine_disk.rs`:
| Test | Target |
| --- | --- |
| Insert 100k rows (50-byte payload) | < 5 s wall, < 50 MiB RSS at end |
| Random read 100k rows under steady-state load | < 5 µs p50, < 100 µs p99 (page cache hot) |
| Cold-cache read 100k rows | < 200 µs p50 (one disk seek per read) |
| Recovery time after kill -9 mid-bench | < WAL_volume / disk_throughput, dominated by `fdatasync` round-trips |
`criterion` is normally an external crate; we're not adding deps. The bench is a `#[test]` with a `--release` runner — coarse but enough to catch regressions.
### `Cargo.toml` delta
None — `db` and `wal` are already in from phases 10 and 11.
## Exit criteria
1. **`cargo build`** at root, four direct deps unchanged (`anyhow`, `serde`, `serde_json`, `libc`).
2. **All existing unit tests still pass** after the engine rewrite. The two heaviest are `engine::tests::crud_roundtrip_auto_id` (port to verify offset semantics) and `server::tests::*` (HTTP-level CRUD — should be unaffected).
3. **End-to-end api.rest battery passes byte-identically** — same status codes, same JSON bodies, same key ordering.
4. **Integration test `crates/rt/tests/disk_engine.rs`:**
- Seed 10k rows of a 1 KiB payload type. Memory after seed (`/proc/self/status` `VmRSS`) is bounded by `id_count × 24 B + listener_overhead`, NOT by `10000 × 1024`. Specifically: less than 40 MiB.
- Restart with kill -9 mid-write; recovery completes in < 1 s for a 16-MiB-WAL-segment workload.
- GET random ids — every read returns the right row, CRC verified.
5. **Checkpoint smoke** — start the binary, write 5 rows, wait `CHECKPOINT_INTERVAL_SECS+1` seconds, verify `data/control` is updated (mtime moved, `last_durable_lsn` advanced). `strace -e fsync,rename` during the wait shows the checkpoint sequence.
6. **Cold start with no `data/`** — a fresh `wo run` on an empty data dir just works (creates the dir, no replay needed). Same for `data/` + empty WAL.
## Non-scope
- **No secondary indexes.** `unique` and `index` schema attributes still trigger no extra storage. Future phase.
- **No compaction.** Tombstoned offsets and old segment bytes accumulate. Trigger compaction is a separate phase keyed on a `dead-bytes / live-bytes` ratio.
- **No MVCC.** Subscriber pre-commit views are tick-boundary semantics (`docs/runtime/database/03-inmemory-engine.md`). Version chains land alongside the cross-shard subscription work in `09c-per-shard-wal` / `09d-cross-shard-subscriptions`.
- **No `O_DIRECT`.** Page cache is the cache. Per [`./exploration/linux/12-pwrite-fsync.md`](./exploration/linux/12-pwrite-fsync.md).
- **No `io_uring` reads.** `pread` syscalls are short and the loop has no other work waiting; an async batched read API isn't worth its own complexity at this size.
- **No streaming list.** `Engine::list` returns all rows for a type in one call. Pagination + cursor support is a future phase keyed on a real workload that hits the wall.
## Verification
```bash
cargo build
cargo test --lib # rt + db + wal unit tests
cargo test --test disk_engine # the new integration test
cargo test --release --test disk_engine -- --nocapture # bench numbers visible
# manual end-to-end
cargo run --release --bin wo -- run docs/examples/blog &
PID=$!
sleep 1
# Seed 10000 rows
for i in $(seq 1 10000); do
curl -sf -X POST http://127.0.0.1:8080/api/articles \
-H 'Content-Type: application/json' \
-d "{\"slug\":\"s$i\",\"title\":\"T$i\",\"author\":1,\"published\":true,\"meta\":{\"excerpt\":\"\",\"body_md\":\"\"}}" \
> /dev/null
done
# Memory check
ps -o rss= -p $PID # expect under ~50 MiB even with 10k rows × 1 KiB each
# Wait for checkpoint
sleep 65
ls -la docs/examples/blog/data/control # mtime should be recent
kill -INT $PID
# Cold restart
cargo run --release --bin wo -- run docs/examples/blog &
PID=$!
sleep 1
curl -s 'http://127.0.0.1:8080/api/articles' | python3 -c 'import json,sys;print(len(json.load(sys.stdin)))'
# expect: 10000
kill -INT $PID
cd reference/crates && cargo build && cargo test # v1 untouched
```
## After this phase
The single-thread runtime is durable, RAM-bounded, and recovery-fast. Phases 13+ pivot to layering features on top: secondary indexes, compaction, query-layer integration, then the `09a-09f` scaleout sequence which lifts the same primitives into per-shard form. The empty `crates/{value, engine, txn}` skeletons get populated as their phases activate; `wal/` and `db/` are now real code, used by `rt/`.
The `crates/rt/Cargo.toml` direct dep list at the end of phase 12 is `anyhow + serde + serde_json + libc + db + wal`. Phase 05 collapses `serde + serde_json` into the hand-rolled JSON module; phase 06 collapses `anyhow` into a bespoke error type. The `libc` + path-deps end state from [`./done/01-scafolding-crates.md`](./done/01-scafolding-crates.md) is reachable in two more phases past 12.

View file

@ -19,6 +19,7 @@ Each primitive has its own numbered file with the kernel source path (into [`ref
| [09](./09-fallocate.md) | `fallocate` + `pread` + `pwritev2` — positional I/O & pre-allocation | phase 3 (WAL + SSTables) |
| [10](./10-pidfd.md) | `pidfd` — process as fd, race-free supervision | future supervisor |
| [11](./11-memfd_create.md) | `memfd_create` — anonymous shared memory | phase 3 (index build) |
| [12](./12-pwrite-fsync.md) | `pwrite` + `fsync`/`fdatasync`/`sync_file_range`/`posix_fadvise` — durability barriers | phases 10, 11, 12 (storage foundations, WAL, engine cutover) |
The list below is the original overview kept for context and for a handful of adjacent primitives (`fanotify`, `splice`/`tee`) that don't yet have their own reference card.

View file

@ -0,0 +1,152 @@
# 12 — `pwrite` + `fsync` durability syscalls
The previous cards cover positional I/O ([`09-fallocate.md`](./09-fallocate.md)) and the page cache backstory ([`08-mmap.md`](./08-mmap.md)) but skip the actual durability primitives. This card fills the gap. Every persistent-storage phase (10, 11, 12) leans on these.
## The four syscalls
| # | Syscall | What it guarantees | When to use |
| --- | --- | --- | --- |
| 1 | `pwrite(2)` / `pwritev2(2)` | Bytes are in the OS page cache at the given offset. NOT on disk. | Every write — the durability barrier is `fsync`, not `write`. |
| 2 | `fdatasync(2)` | Bytes for this fd are on durable storage; metadata except size NOT guaranteed. | WAL commits — we don't read inode timestamps for replay, so metadata is irrelevant. |
| 3 | `fsync(2)` | Bytes + all inode metadata for this fd are on durable storage. | After file rename / extension where the metadata change matters (control file, segment rollover, parent directory). |
| 4 | `sync_file_range(2)` | Linux-specific async hint — start I/O on a byte range without committing the metadata. | Optional optimization — kick off WAL writeback early, then `fdatasync` at commit. |
## Postgres uses all four
| Postgres call | Wraps | Where |
| --- | --- | --- |
| `pg_pwrite()` | `pwrite64` | [`storage/file/fd.c`](../../../../reference/postgresql/src/backend/storage/file/fd.c) — every block-aligned write. |
| `pg_fsync()` | `fsync` (or platform variant) | [`storage/file/fd.c`](../../../../reference/postgresql/src/backend/storage/file/fd.c) — wraps `wal_sync_method` GUC dispatch. |
| `pg_fdatasync()` | `fdatasync` | Same. Selected when `wal_sync_method = fdatasync`. |
| Async writeback | `sync_file_range` | [`access/transam/xlog.c`](../../../../reference/postgresql/src/backend/access/transam/xlog.c) — `issue_xlog_fsync` calls `sync_file_range(SYNC_FILE_RANGE_WRITE)` to start I/O on the WAL ahead of the durability barrier. |
The Postgres GUC matrix (`wal_sync_method`) lets the operator pick between `fsync`, `fdatasync`, `open_sync`, `open_datasync`, `fsync_writethrough`. **Writeonce picks one** — `fdatasync` for the WAL, `fsync` for control files and segment rollovers — and ships it.
## Kernel source
| Path | What |
| --- | --- |
| [`reference/linux/fs/read_write.c`](../../../reference/linux/fs/read_write.c) | `SYSCALL_DEFINE4(pread64, ...)`, `SYSCALL_DEFINE4(pwrite64, ...)`, `SYSCALL_DEFINE6(pwritev2, ...)`. |
| [`reference/linux/fs/sync.c`](../../../reference/linux/fs/sync.c) | `SYSCALL_DEFINE1(fsync, ...)`, `SYSCALL_DEFINE1(fdatasync, ...)`, `SYSCALL_DEFINE4(sync_file_range, ...)`. |
| [`reference/linux/include/uapi/asm-generic/fcntl.h`](../../../reference/linux/include/uapi/asm-generic/fcntl.h) | `O_SYNC`, `O_DSYNC`, `O_DIRECT`. |
| [`reference/linux/Documentation/filesystems/ext4/journal.rst`](../../../reference/linux/Documentation/filesystems/ext4/journal.rst) | What ext4's journal commits when `fsync` runs. Worth understanding what the kernel actually does on the durability path. |
## Man pages
`man 2 pwrite`, `man 2 pwritev2`, `man 2 fsync`, `man 2 fdatasync`, `man 2 sync_file_range`, `man 2 posix_fadvise`.
## Rust FFI via `libc`
```rust
use libc::{
pwrite, pwritev2,
fsync, fdatasync,
sync_file_range,
posix_fadvise,
iovec, off_t,
SYNC_FILE_RANGE_WRITE, SYNC_FILE_RANGE_WAIT_AFTER, SYNC_FILE_RANGE_WAIT_BEFORE,
POSIX_FADV_DONTNEED, POSIX_FADV_RANDOM, POSIX_FADV_SEQUENTIAL,
RWF_DSYNC, RWF_SYNC,
};
```
## Direct-syscall example — WAL group commit
```rust
// 1. Append a record to the WAL via pwrite — bytes land in the page cache.
let n = unsafe {
libc::pwrite(wal_fd, rec.as_ptr() as *const _, rec.len(), wal_tail as i64)
};
if n < 0 { return Err(io::Error::last_os_error()); }
// 2. Group-commit fence: drain all pending commits this tick, then ONE barrier.
// fdatasync is enough — the WAL file is fixed-size (fallocated), no metadata change
// to commit. Saves a journal-update on ext4.
if unsafe { libc::fdatasync(wal_fd) } < 0 {
return Err(io::Error::last_os_error());
}
// 3. Now ack each pending commit.
for fd in commits.drain(..) { send_response(fd, /* 200 OK */); }
```
## Direct-syscall example — control file rename-on-write
```rust
// 1. Write the new control file content to a sibling temp file.
let tmp = data_dir.join("control.tmp");
let fd = unsafe { libc::open(tmp.as_ptr(), libc::O_WRONLY | libc::O_CREAT | libc::O_TRUNC, 0o644) };
let _ = unsafe { libc::pwrite(fd, ctl.as_ptr() as *const _, ctl.len(), 0) };
// 2. fsync (NOT fdatasync) — the file's size + inode metadata must be durable
// before the rename, otherwise the rename could publish a zero-byte control
// file after a crash.
if unsafe { libc::fsync(fd) } < 0 { return Err(...); }
unsafe { libc::close(fd); }
// 3. rename(2) is atomic on POSIX-compliant filesystems.
unsafe { libc::rename(tmp.as_ptr(), final_path.as_ptr()); }
// 4. fsync the parent directory so the rename's directory-entry update is durable.
let dfd = unsafe { libc::open(data_dir.as_ptr(), libc::O_RDONLY) };
unsafe { libc::fsync(dfd); libc::close(dfd); }
```
## `fsync` vs `fdatasync` — when each matters
`fsync` flushes:
- All dirty page-cache pages backing the file.
- The inode metadata: size, atime/mtime/ctime, link count.
`fdatasync` flushes:
- All dirty pages.
- The inode metadata **only if it would change a future read** — i.e. file size if the write extended the file.
Why this matters for writeonce:
- **WAL writes**: the file is `posix_fallocate`'d up front to its full segment size, so writes never extend it. `fdatasync` is enough. Saves the metadata journal update, which on ext4 is roughly half the cost of an `fsync`.
- **Control file**: `fsync` because we follow with a `rename(2)` whose effect we want to be durably visible — and renames involve directory inode metadata.
- **Segment rollover** (creating a new `.seg` file or extending past the fallocate'd size): `fsync` once at rollover, `fdatasync` for steady-state writes.
## `sync_file_range` — when the optimization is worth it
`sync_file_range(fd, offset, length, SYNC_FILE_RANGE_WRITE)` starts I/O on a byte range without waiting. Use case: between WAL appends and the commit fence, kick the kernel to start writing the just-appended bytes. By the time `fdatasync` runs, much of the work is already in flight, so the latency of the barrier shrinks.
Postgres uses this in `issue_xlog_fsync`. Writeonce skips it in phase 11 — the loop tick boundary is short enough (microseconds in the steady state) that the bytes are still in the page cache and the kernel writeback hasn't kicked in. Revisit if benchmarks show fsync latency dominating.
## `posix_fadvise` — drop pages from the cache after writes
`posix_fadvise(fd, offset, len, POSIX_FADV_DONTNEED)` tells the kernel "I won't read these bytes again soon" — the kernel can reclaim the page cache for them. Useful after **bulk inserts** (`COPY` in Postgres, future `wo bulk-load` in writeonce) when you don't want the page cache to evict useful working-set pages to make room for cold inserts.
Not needed in phase 10/11. Possibly useful later for archival writes.
## Why writeonce does NOT use `O_DIRECT`
`O_DIRECT` bypasses the page cache: writes go straight to the block device, reads come straight from it. Saves a memcpy at the cost of:
- **No read-ahead.** Every read is a disk seek.
- **No write coalescing.** Two adjacent `pwrite`s of the same page hit the device twice.
- **Strict alignment.** Buffer pointers, offsets, and lengths must be multiples of the device's logical block size.
- **Recovery cost.** With `O_DIRECT`, the page cache is dark. Recovery has to physically read every WAL byte. With buffered I/O, the recently-written WAL is already in the page cache — replay is a memcpy.
Writeonce trades the memcpy for the recovery speed and the simpler programming model. If a future phase identifies a workload where the memcpy is the bottleneck (it almost never is), `O_DIRECT` can be added per-fd.
## Gotchas
- **`fsync` returning success doesn't always mean the bytes are durable.** Some consumer SSDs lie about cache-flush completion. Postgres has documented this extensively (`wal_sync_method = open_datasync` was added partly as a workaround). For writeonce on a real server-class SSD, `fdatasync` + `O_DSYNC` give correct semantics; on flaky hardware, no software can compensate.
- **`fdatasync` is enough only because we fallocated.** If a future phase appends to a non-fallocated file, the size change requires `fsync` to commit the new size to the inode — otherwise a crash recovery sees the old size and our written bytes appear missing.
- **`pwrite`'s atomicity is limited.** On regular files, writes ≤ `PIPE_BUF` (4 KiB) are atomic with respect to other writes. Larger writes can be torn — another reader interleaved with the writer might see half the new bytes. The single-thread loop has no other writers, so this is moot, but it's worth knowing if a future phase introduces a second writer.
- **Don't `close` then `fsync`**. The fsync is a no-op against a closed fd. If you must close, fsync first.
- **`rename` requires a parent-directory fsync to be durable across a crash.** A common bug — Postgres' `fsync_parent_path` exists exactly for this.
## Used by
- [`docs/plan/10-storage-foundations.md`](../../10-storage-foundations.md) — segment append uses `pwrite` + later `fdatasync`.
- [`docs/plan/11-wal-and-recovery.md`](../../11-wal-and-recovery.md) — group-commit uses `pwrite` + `fdatasync`; control file uses `fsync` + `rename` + parent-dir `fsync`.
- [`docs/plan/12-engine-disk-cutover.md`](../../12-engine-disk-cutover.md) — checkpoint uses `fsync` per active segment fd.
Pair with [`postgresql/wal.md`](../postgresql/wal.md), [`postgresql/buffer-and-checkpoint.md`](../postgresql/buffer-and-checkpoint.md) for the design context.
## v1 port source
**Partial.** `reference/crates/wo-seg/src/writer.rs:92` calls `file.sync_all()` (Rust stdlib's `fsync` wrapper). Phase 11 replaces with explicit `libc::fdatasync` for the WAL path; segment files keep `fsync` semantics for rollover events.

View file

@ -0,0 +1,48 @@
# PostgreSQL — storage subsystem reference
These cards exist to make the Postgres backend a useful **library of patterns** for writeonce's persistent-storage phases (10–12) without inviting a multi-process port. Each card pulls one subsystem out of [`reference/postgresql/src/backend/`](../../../../reference/postgresql/src/backend/) — paths into the Postgres tree, the underlying *idea*, and the writeonce translation.
The symlink is user-specific:
```bash
ln -s /home/shoney/projects/postgresql reference/postgresql
```
Gitignored — see [`.gitignore`](../../../../.gitignore). Pair it with [`reference/linux`](../../../../reference/linux) and [`reference/go`](../../../../reference/go) if not already linked.
## Per-subsystem cards
| # | Postgres area | What writeonce takes | What writeonce skips |
| --- | --- | --- | --- |
| [wal](./wal.md) | `access/transam/xlog*.c` | append-only sequential log, LSN-as-byte-offset, segment rollover, group commit, `pwrite` + `fsync` at commit | replication/archiver, multi-process WAL writer, GUC matrix |
| [smgr-and-md](./smgr-and-md.md) | `storage/smgr/{md,smgr,bulk_write}.c` | one file per relation, segments capped at `RELSEG_SIZE`, immediate vs deferred fsync | multi-fork abstraction (main/fsm/vm), shared-memory descriptor cache |
| [buffer-and-checkpoint](./buffer-and-checkpoint.md) | `storage/buffer/{bufmgr,freelist}.c` + `postmaster/{checkpointer,bgwriter}.c` | page cache + dirty bit + LRU; checkpoint flushes then advances control-file LSN | shared-buffer pinning/unpinning, separate writer processes, latches |
| [page-format](./page-format.md) | `storage/page/{bufpage,checksum}.c` | page header (LSN, checksum, free-space markers); CRC32C trailers | MVCC visibility (xmin/xmax/ctid), access-method-specific opaque space |
## The lift-vs-skip filter
Postgres is multi-process by birth: a postmaster forks one backend per connection plus dedicated checkpointer / bgwriter / walwriter / archiver / autovacuum processes. Most of `src/backend/storage/ipc/`, `storage/lmgr/`, the latch system, and the `proc.c` family exist to coordinate **between those processes** — shared-memory regions, semaphores, condition variables, lock manager partitions, signal forwarding. Writeonce is single-process and single-threaded, so all of that mechanism is dead weight here. The *concepts* underneath (fairness, deadlock detection, request batching) generalize anyway, but writeonce satisfies them with single-thread invariants instead of IPC primitives.
What carries over cleanly:
1. **Sequential WAL with fsync at commit** — applicable to any durable store regardless of process model.
2. **Page cache abstraction** — even single-threaded engines need a dirty/clean bit and an LRU eviction story; the kernel page cache covers most of it via `mmap` / buffered I/O, but the dirty-tracking + flush-batching policy is something we own.
3. **Control file with last-safe-LSN** — small, fixed-size, atomically updated via rename-on-write. Survives multi-process and single-process alike.
4. **Recovery = replay WAL from last checkpoint** — the algorithm is identical; what writeonce skips is the postmaster signaling that says "ok, recovery is done, accept connections."
5. **CRC32C on every record + page** — the cost is a few cycles per write, the pay-off is silent-corruption detection. Worth it.
What stays out:
- **Shared-memory / dynamic-shmem coordination** (`storage/ipc/dsm*.c`, `storage/lmgr/`). Single-thread loop has no co-tenants.
- **Multi-version concurrency control** (`access/transam/clog.c`, xmin/xmax tuple headers). The locked architecture (`docs/runtime/database/02-wo-language.md` § Concurrency Model) commits to MVCC for snapshot isolation, but the version chains are not what makes single-thread writes durable. Layered in later, when LIVE subscribers want pre-commit views.
- **Separate writer processes** (`postmaster/{walwriter,bgwriter,checkpointer,archiver,autovacuum}.c`). Each becomes a per-tick chunk of work in the same loop, gated by deadlines.
## Phase mapping
The implementation phases that lean on this material:
- [`docs/plan/10-storage-foundations.md`](../../10-storage-foundations.md) — page format and segment files. Lifts ideas from `smgr/md.c` and `page/bufpage.h`.
- [`docs/plan/11-wal-and-recovery.md`](../../11-wal-and-recovery.md) — WAL framing, group commit, recovery loop. Lifts ideas from `access/transam/xlog.c` and the xlog-recovery family.
- [`docs/plan/12-engine-disk-cutover.md`](../../12-engine-disk-cutover.md) — buffer cache + dirty tracking. Lifts ideas from `storage/buffer/bufmgr.c` and `postmaster/checkpointer.c`.
Pair each card with [`docs/plan/exploration/linux/12-pwrite-fsync.md`](../linux/12-pwrite-fsync.md) for the actual syscalls — these cards are about *design patterns*, that one is about *kernel calls*.

View file

@ -0,0 +1,71 @@
# Buffer cache & checkpoint
`storage/buffer/bufmgr.c` is Postgres' page cache: pages live in shared memory, are pinned/unpinned by readers, and are written back to disk lazily. `postmaster/checkpointer.c` is the dedicated process that periodically flushes all dirty pages, then advances the **redo pointer** in the control file — a guarantee that "everything before LSN X is on disk; recovery can start from X."
Writeonce takes the same two ideas, single-thread:
- A **clean/dirty bit per page** kept in-process. Page reads and writes hit the OS page cache directly via `pread` / `pwrite` (writeonce does not maintain its own user-space buffer pool; the kernel page cache is good enough for the workload size we target).
- A **checkpoint routine** that runs as a periodic step in the loop. It walks dirty pages, issues `fsync` on the affected files, and rewrites the control file's `last_durable_lsn`.
No separate process. No shared-buffer pinning. No dynamic-shared-memory coordination.
## Postgres source
| File | Responsibility |
| --- | --- |
| [`storage/buffer/bufmgr.c`](../../../../reference/postgresql/src/backend/storage/buffer/bufmgr.c) | Page cache front-door: `ReadBuffer`, `BufferGetPage`, `MarkBufferDirty`, `FlushBuffer`. Tracks dirty bit per buffer; pinning prevents eviction. |
| [`storage/buffer/freelist.c`](../../../../reference/postgresql/src/backend/storage/buffer/freelist.c) | Clock-sweep eviction policy. Buffers with `usage_count = 0` and `pin_count = 0` are eviction candidates; usage decremented on every sweep pass, incremented on access. |
| [`storage/buffer/buf_table.c`](../../../../reference/postgresql/src/backend/storage/buffer/buf_table.c) | Hash table from `(file, block)` → buffer slot. The lookup that `ReadBuffer` does. |
| [`postmaster/checkpointer.c`](../../../../reference/postgresql/src/backend/postmaster/checkpointer.c) | The checkpointer process. Triggered by time (`checkpoint_timeout`), WAL volume (`max_wal_size`), or signal. Runs `BufferSync()` to flush dirty buffers, then `CreateCheckPoint()` to update the control file. |
| [`postmaster/bgwriter.c`](../../../../reference/postgresql/src/backend/postmaster/bgwriter.c) | Continuously trickles dirty pages to disk between checkpoints. Smooths the I/O burst the checkpointer would cause. |
| [`storage/buffer/README`](../../../../reference/postgresql/src/backend/storage/buffer/README) | Overview of the pinning, locking, and replacement policy. Worth reading. |
## The page-cache idea worth porting
A buffer in Postgres is a `(file_id, block_number, page_data, dirty_bit, pin_count, usage_count, content_lock, io_lock)`. Strip out everything that exists for multi-process coordination (`pin_count`, locks) and you get the per-block state you need in any persistent store: **the bytes, where they came from on disk, and whether they're dirty since last fsync.**
Writeonce's phase 12 `Engine` keeps an `HashMap<(TypeName, SegmentOffset), CachedRow>` where `CachedRow = { bytes: Vec<u8>, dirty: bool }`. Rows are read on-demand (cache miss → `pread` + decode + CRC verify), written through to the segment but not flushed to disk until the next commit's WAL fsync covers them. Dirty rows accumulate; a periodic checkpoint flushes the segment fds and advances the control-file LSN.
The kernel page cache does most of the work. `pread` against an fd that already has its page cached is a memcpy. `pwrite` populates the page cache without going to disk until pressure or `fsync`. This is why writeonce explicitly does NOT use `O_DIRECT` (see [`linux/12-pwrite-fsync.md`](../linux/12-pwrite-fsync.md)) — the page cache is the one cache we want.
## Checkpoint — the writeonce shape
Postgres' checkpoint runs in a separate process and signals the postmaster when done. Writeonce's runs as a periodic loop step:
```text
loop tick (every CHECKPOINT_INTERVAL, e.g. 60s):
fsync(every active segment fd) // metadata + data barrier
fsync(wal_dir_fd) // ensure recent WAL writes are visible
write control.tmp { last_durable_lsn = current_wal_tail }
fsync(control.tmp)
rename(control.tmp, control)
fsync(data_dir_fd) // make the rename durable
```
The `rename(2)` is atomic on POSIX-compliant filesystems — at any crash point, either `control.tmp` is missing (the rename hasn't happened) or `control` reflects the new content. `fsync(parent_dir)` is needed because rename's atomicity is in-kernel; the directory entry isn't durable until its parent inode is synced. (Postgres does the same in `BasicOpenFile` + `fsync_parent_path`.)
Recovery on startup reads `control`, finds the `last_durable_lsn`, and replays WAL forward from there. Records before that LSN are *known* to be in the segment files; records after are replayed.
## What writeonce skips
- **Pinning + content locks.** Single-thread loop has one reader and one writer of the cache: itself. No need for `LockBuffer(BUFFER_LOCK_SHARE)` etc.
- **`bgwriter` continuous trickle.** Postgres has a *separate* process slowly cleaning the buffer pool to avoid I/O spikes at checkpoint. Writeonce's checkpoints are infrequent enough (60s default) that a spike is fine; if it becomes a problem, the same loop can do "soft flush K pages per tick" without spawning anything.
- **Hash partitioning of the buffer table.** Postgres partitions `buf_table` to reduce lock contention — single-thread doesn't have lock contention.
- **`shared_buffers` GUC.** Postgres lets the operator size the buffer pool. Writeonce trusts the OS page cache and bounds its in-process cache by an LRU with a simple count limit (`WO_CACHE_ROWS=10000` default, configurable).
## Where to look in the Postgres source
For the page cache:
- `BufferAlloc()` in `bufmgr.c` — read the function header. Strip the locks and you've got the cache-miss path.
- `BufferSync()` in `bufmgr.c` — read the prologue. The dirty-buffer-walk + per-relation fsync coalescing is the checkpoint algorithm.
For the checkpoint:
- `CreateCheckPoint()` in `xlog.c` — the control-file update sequence. Read the comments around `WriteControlFile` and the surrounding `pg_fsync`s. That's the rename-on-write pattern in practice.
- The `checkpointer.c` main loop is short and worth scanning for the time-vs-WAL-volume trigger logic.
## Used by
- [`docs/plan/12-engine-disk-cutover.md`](../../12-engine-disk-cutover.md) — disk-backed engine reads and dirty-row tracking.
- [`docs/plan/11-wal-and-recovery.md`](../../11-wal-and-recovery.md) — control file write sequence (phase 11 ships the control file; checkpoint as a periodic step lands with phase 12 or shortly after).
Pair with [`linux/12-pwrite-fsync.md`](../linux/12-pwrite-fsync.md) for the fsync semantics and [`linux/08-mmap.md`](../linux/08-mmap.md) for the OS page-cache backstory.

View file

@ -0,0 +1,82 @@
# Page format & checksums
`storage/page/bufpage.h` defines Postgres' on-disk page layout: a fixed `BLCKSZ` (8 KiB by default), with a 24-byte header at the front and tuples filling from the back, line pointers pointing into them. `storage/page/checksum.c` adds an optional CRC32C-derived checksum embedded in the header — turned on at `initdb --data-checksums` time, off by default for historical performance reasons.
Writeonce's phase 10 starts simpler — variable-length records, no pages. Phase 12+ may revisit a page-style layout if range scans become hot enough that record-level reads aren't enough. Either way, the **header + checksum pattern** carries over and the `bufpage.h` header is worth understanding.
## Postgres source
| File | Responsibility |
| --- | --- |
| [`storage/page/bufpage.c`](../../../../reference/postgresql/src/backend/storage/page/bufpage.c) | Page initialization (`PageInit`), line-pointer manipulation, free-space accounting. |
| [`storage/page/checksum.c`](../../../../reference/postgresql/src/backend/storage/page/checksum.c) | The page checksum algorithm — CRC32C-style with a Postgres-specific finalization. Optional, enabled at cluster init. |
| [`storage/page/itemptr.c`](../../../../reference/postgresql/src/backend/storage/page/itemptr.c) | Item pointer (`ItemPointerData`) — `(block_number, offset_within_page)` 6-byte tuple address. The on-disk equivalent of writeonce's `(TypeName, SegmentOffset)`. |
| [`include/storage/bufpage.h`](../../../../reference/postgresql/src/include/storage/bufpage.h) | The header-file definition. Read this first — it's the spec. |
| [`storage/page/README`](../../../../reference/postgresql/src/backend/storage/page/README) | One-page overview of the slotted-page model and how checksums interact with WAL. |
## The Postgres page header (24 bytes)
```text
struct PageHeaderData {
PageXLogRecPtr pd_lsn; // 8 bytes — the LSN that last modified this page
uint16 pd_checksum; // 2 bytes — CRC32C over the page (set if data_checksums)
uint16 pd_flags; // 2 bytes — has-free-space, etc.
LocationIndex pd_lower; // 2 bytes — offset to start of free space (line ptr end)
LocationIndex pd_upper; // 2 bytes — offset to end of free space (tuple start)
LocationIndex pd_special; // 2 bytes — offset to access-method specific area
uint16 pd_pagesize_version; // 2 bytes — page size + layout version
TransactionId pd_prune_xid; // 4 bytes — oldest XID to prune (vacuum hint)
};
```
The body of the page after this header holds **line pointers** (`ItemIdData`, 4 bytes each) growing forward and **tuples** growing backward. The gap between `pd_lower` and `pd_upper` is the free space.
## What's worth porting (eventually)
1. **`pd_lsn` field at the head of every page.** When a page is read back, the LSN tells you "this page reflects WAL records up to LSN N." Recovery can skip records ≤ N for this page (they're already applied). Postgres uses this to avoid double-applying WAL during recovery; writeonce's phase 12+ would too if it goes page-based.
2. **CRC32C trailer/embedded checksum.** Postgres puts it in the header and zeroes the field while computing. Writeonce's phase 10 record framing puts a CRC32C **trailer** (last 4 bytes of the record) — same algorithm, different position. The trailer position is simpler when records are variable-length: the length field in the header tells you exactly where the CRC ends.
3. **Slotted-page line pointers (later).** When phase 12+ wants page-locality for range scans, the slotted-page model — line pointers near the page header, tuples backward from the end — gives O(1) tuple access by index without resizing copies. Postgres' implementation is well-trodden ground.
## What writeonce does instead (phase 10)
Variable-length records, length-prefixed. The framing is in [`docs/plan/10-storage-foundations.md`](../../10-storage-foundations.md):
```text
[u32 length LE][u8 flags][u8 record_kind][u64 LSN][payload bytes][u32 CRC32C]
```
Compared to a Postgres page:
| Concern | Postgres page | Writeonce record |
| --- | --- | --- |
| Granularity | 8 KiB fixed | variable, typical row size |
| Address | `(file, BLCKSZ × block)` | `(type, byte_offset)` |
| LSN | `pd_lsn` in header | embedded after flags |
| Checksum | `pd_checksum` in header | CRC32C trailer |
| Free-space tracking | `pd_lower` / `pd_upper` | none — append-only, segment growth via fallocate |
| Line pointers | yes — relocatable tuples | no — record offset is permanent until tombstoned |
Phase 10 trades range-scan locality for simplicity and append-only commit semantics. The trade is reversible: a future phase can introduce a page layer above the segment without breaking the WAL/recovery contract.
## Why writeonce starts without slotted pages
Postgres' page format earns its complexity:
- **MVCC tuple visibility** needs in-place updates of `xmin/xmax/ctid` on individual tuples — line pointers let one page hold versions across many transactions without rewriting tuples on every `UPDATE`.
- **Free-space recovery** within a page (after a tuple is dead and pruned) is essential when 99% of pages are partly empty.
- **Range queries on a B-tree leaf** want all the keys in one page, sorted, so a 4-KiB read returns dozens of matches.
Phase 10 hits none of these:
- No MVCC yet (`docs/plan/12-engine-disk-cutover.md` defers).
- Append-only segments — a tombstoned record is wasted bytes until compaction, which is fine for the workload.
- Reads go through the in-memory `BTreeMap<i64, SegmentOffset>` index — the segment file isn't scanned linearly; we know exactly where each row lives.
Slotted pages are the answer when those assumptions break. Until then, the framing above is enough.
## Used by
- [`docs/plan/10-storage-foundations.md`](../../10-storage-foundations.md) — record framing borrows the **header + checksum** pattern from `bufpage.h`.
- [`docs/plan/12-engine-disk-cutover.md`](../../12-engine-disk-cutover.md) — when reading rows back from disk, CRC verification is the silent-corruption safety net the page header gives Postgres.
Pair with [`wal.md`](./wal.md) for the LSN convention and [`buffer-and-checkpoint.md`](./buffer-and-checkpoint.md) for the dirty-page semantics that pages need.

View file

@ -0,0 +1,80 @@
# Storage manager — relation files on disk
`storage/smgr/{smgr.c,md.c}` is the layer between "this is relation X" and "this is a set of OS files." It maps each relation to a sequence of **segment files** on disk, capped at `RELSEG_SIZE` (default 1 GiB) — a relation that grows past 1 GiB rolls over into `<oid>.1`, `<oid>.2`, …. Reads and writes go through `smgrread` / `smgrwrite` / `smgrextend`, which delegate to the magnetic-disk implementation in `md.c`. There's a per-relation cache (`SMgrRelation`) so the OS file descriptors persist across many calls.
The writeonce equivalent is **per-type segment files** (`data/<TypeName>.seg`). Same shape, simpler: only one storage manager (no kerb-style abstractions for shared/local/whatever), no segment number — single file per type until a type's data grows past a configured cap.
## Postgres source
| File | Responsibility |
| --- | --- |
| [`storage/smgr/smgr.c`](../../../../reference/postgresql/src/backend/storage/smgr/smgr.c) | Front-door API. `smgropen`, `smgrread`, `smgrwrite`, `smgrextend`, `smgrdounlink`. Holds the `SMgrRelation` cache. |
| [`storage/smgr/md.c`](../../../../reference/postgresql/src/backend/storage/smgr/md.c) | The actual implementation against the kernel. Manages `MdfdVec` (open file descriptor handles per segment number), opens missing segments lazily. |
| [`storage/smgr/bulk_write.c`](../../../../reference/postgresql/src/backend/storage/smgr/bulk_write.c) | Optimized path for bulk-loading: writes directly to `smgrwrite` without going through shared buffers. Useful for `COPY` / `CREATE INDEX` + the recovery path's wal-replay-rebuilds-pages flow. |
| [`storage/smgr/README`](../../../../reference/postgresql/src/backend/storage/smgr/README) | Brief but worth reading — explains the relfilenode → file naming convention and how `RELSEG_SIZE` interacts with 32-bit-fs-size historical limits. |
## What `md.c` actually does
For a request to read block B of relation R, `md.c`:
1. Looks up R's `MdfdVec`. Each entry holds `(segment_number, fd)`.
2. Computes `segment_number = B / RELSEG_SIZE_BLOCKS`, `block_within_segment = B % RELSEG_SIZE_BLOCKS`.
3. If the fd for that segment isn't cached, `open(2)` the file (`<oid>.<seg>`).
4. `pread(fd, buf, BLCKSZ, block_within_segment * BLCKSZ)` — single positional read, no shared cursor.
For an `smgrextend(R, B)` (grow):
1. If we're crossing a segment boundary, `open(O_CREAT)` the next segment file.
2. `pwrite(fd, zero_buf, BLCKSZ, block_within_segment * BLCKSZ)` — extend the file by writing a zero page at the new offset. (Postgres preallocates pages explicitly because some filesystems sparsely allocate when extending without writing — `md.c` wants the page allocated *now*.)
3. Optionally `posix_fallocate` the segment up front instead.
`md.c` is also where `mdsyncfiletag` lives — Postgres' deferred-fsync mechanism. The checkpointer hands `md.c` "the list of files that need to be `fsync`'d before this checkpoint completes," `md.c` deduplicates and issues the syscalls.
## What writeonce keeps
1. **One file per relation/type.** `data/Article.seg`, `data/Comment.seg`. The Postgres-style `relfilenode` numbering is overkill — type names are stable and globally unique within a `wo run` directory.
2. **Segment rollover when files get big.** Postgres caps at 1 GiB. Writeonce inherits a sensible cap (start at 1 GiB; no observable difference until then). Beyond the cap: `<TypeName>.1.seg`, `.2.seg`, …
3. **Lazy fd open + fd caching.** First write to a type opens the file; the fd lives in an `HashMap<TypeName, RawFd>` for the lifetime of the engine. Closed on drop. No per-tick syscall churn.
4. **`pwrite` for positional writes; `pread` for positional reads.** No `lseek` cursor, so concurrent reads are safe. Pairs naturally with [`io_uring`](../linux/07-io_uring.md) when we want batched I/O.
5. **`posix_fallocate` to reserve space at segment creation.** Avoids `ENOSPC` mid-write and minimizes filesystem-level fragmentation. See [`linux/09-fallocate.md`](../linux/09-fallocate.md).
## What writeonce drops
- **Multi-fork relations.** Postgres has `main`, `fsm` (free-space map), `vm` (visibility map) forks per relation, each its own file. Writeonce starts with one fork per type — secondary indexes get their own files (`<TypeName>.<col>.idx`) when phase 12 needs them, not as a forks-of-the-same-relation abstraction.
- **`RelFileNode` indirection.** Postgres runs `(database_oid, tablespace_oid, relation_oid)` through a layer that resolves to a path. Writeonce names files by type directly. Layered indirection is something we add when (if) we ever support multi-database.
- **Tablespaces.** The whole concept of "relation X lives in tablespace Y which is a directory at path P" is a multi-tenant ops feature. Writeonce binds to one data dir per `wo run` invocation.
- **`storage/large_object/`.** TOAST + LOB. JSON values bigger than ~8 KiB pages get sliced into pieces stored in a `pg_toast_<oid>` table. Writeonce records aren't page-bound until phase 12+ adds page-level layout, and even then large records can stay in the segment as one variable-length entry.
## Hot reads in `md.c`
The single function worth porting in spirit is `mdread` / `mdwrite`. The Postgres versions are ~80 LOC each; the writeonce versions collapse to ~30 once you strip the `BLCKSZ`/segment-arithmetic + the multi-fork abstraction. The recipe:
```rust
fn read(&self, ty: &str, offset: u64, len: usize) -> io::Result<Vec<u8>> {
let fd = self.fd(ty)?; // opens lazily, caches
let mut buf = vec![0u8; len];
let n = unsafe {
libc::pread(fd, buf.as_mut_ptr() as *mut _, len, offset as i64)
};
if n < 0 { return Err(io::Error::last_os_error()); }
buf.truncate(n as usize);
Ok(buf)
}
fn append(&self, ty: &str, bytes: &[u8]) -> io::Result<u64> {
let fd = self.fd(ty)?;
let offset = self.tail(ty)?; // tracked in-memory by SegStore
let n = unsafe {
libc::pwrite(fd, bytes.as_ptr() as *const _, bytes.len(), offset as i64)
};
if n < 0 { return Err(io::Error::last_os_error()); }
self.advance_tail(ty, n as u64);
Ok(offset)
}
```
That's the core of phase 10's `SegStore`.
## Used by
[`docs/plan/10-storage-foundations.md`](../../10-storage-foundations.md) — segment file layout, append path. Pair with [`linux/09-fallocate.md`](../linux/09-fallocate.md) for preallocation, [`linux/12-pwrite-fsync.md`](../linux/12-pwrite-fsync.md) for the syscall details.

View file

@ -0,0 +1,66 @@
# WAL — append-only durability log
`access/transam/xlog*.c` is Postgres' write-ahead log: every mutation writes a record to a sequential log on disk *before* the in-memory page changes are considered durable. On `COMMIT`, the WAL is `fsync`'d up to the commit's LSN and only then is the client acknowledged. The data files themselves can be flushed lazily — recovery rebuilds them from the WAL.
Writeonce mirrors the algorithm. The single-thread loop replaces multi-process coordination with per-tick group-commit drainage; everything else carries over.
## Postgres source
| File | Responsibility |
| --- | --- |
| [`access/transam/xlog.c`](../../../../reference/postgresql/src/backend/access/transam/xlog.c) | Top-level WAL machinery: insertion locks, segment rollover, flush coordination, control-file rendezvous. |
| [`access/transam/xloginsert.c`](../../../../reference/postgresql/src/backend/access/transam/xloginsert.c) | Build a WAL record (header + payload + backup-block deltas) and place it into the in-memory WAL buffer. |
| [`access/transam/xlogreader.c`](../../../../reference/postgresql/src/backend/access/transam/xlogreader.c) | Decode WAL records during recovery — pure parser, no I/O. Useful as the read-side spec. |
| [`access/transam/xlogrecovery.c`](../../../../reference/postgresql/src/backend/access/transam/xlogrecovery.c) | The replay loop. Walks the WAL from the last-checkpoint LSN, replays each record into shared buffers, advances the redo pointer. |
| [`postmaster/walwriter.c`](../../../../reference/postgresql/src/backend/postmaster/walwriter.c) | Background process that flushes the WAL buffer to disk asynchronously. Writeonce does this **inline in the loop tick**. |
## The five Postgres WAL ideas writeonce keeps
1. **LSN = monotonic byte offset across all segments.** A 64-bit cursor that combines `(segment_id, offset_within_segment)` into one value. Cheap arithmetic — comparing two LSNs is a single `<` on `u64`. Gives every record a permanent address.
2. **Segment rollover.** WAL is broken into fixed-size files (`pg_wal/<24-hex-name>`, default 16 MiB) so old segments can be archived/recycled without touching the active write head. Writeonce inherits the size choice — 16 MiB is a Postgres-tested sweet spot between rollover frequency and archived-file granularity.
3. **Group commit.** Many backends call `XLogFlush(commit_lsn)` concurrently; only the first one issues the `fsync`, the rest wait on the result. The fsync covers everyone's records up to the highest LSN flushed. **In writeonce, this is automatic** — the loop drains all pending commits between two `wait_once`s, then issues one `fsync` covering all of them. No coordination primitive needed.
4. **`pwrite` + `fsync` (or `fdatasync`) at commit.** Postgres uses `pg_pwrite()` (a wrapper over `pwrite64`) for buffer-aligned positional writes and `pg_fsync()` for the durability barrier. Writeonce uses the same syscall pair. See [`linux/12-pwrite-fsync.md`](../linux/12-pwrite-fsync.md).
5. **Control file as the recovery anchor.** A small fixed-size file (`global/pg_control`) stores the **last redo-safe LSN** — the point where recovery starts. Updated atomically (Postgres uses a careful write-fsync-rename sequence). On startup, recovery reads it, then scans WAL forward from that LSN.
## What writeonce drops
- **WAL buffer + walwriter process.** Postgres has an in-memory ring of WAL pages (`XLogCtl`) and a separate process that flushes it asynchronously between commits. Writeonce keeps a per-tick pending-commit list and flushes inline; no buffer ring, no extra process.
- **Replication slots / archive command.** WAL streaming and `archive_command` belong to the multi-machine story. Writeonce is single-binary; replication is a phase past 12.
- **Backup blocks (`XLOG_FPI`, full-page images).** Postgres logs whole pages on first modification after a checkpoint to defend against torn writes (page = 8 KiB, kernel write atomicity = 4 KiB on most fs). Writeonce's record-level CRC32C and per-record framing replace this — no torn-write hazard at the page granularity because writeonce doesn't have pages until a later phase.
- **WAL levels (`wal_level = minimal | replica | logical`).** Postgres tunes how much detail to log based on whether replicas exist. Writeonce always logs the same shape.
## Record framing — Postgres vs. writeonce
Postgres records: `XLogRecord` header (24 bytes: total length, xid, prev LSN, info, rmgr id, CRC32C) + per-rmgr block headers + payload. Variable length. Compact.
Writeonce records (per `docs/plan/10-storage-foundations.md`): `[u32 length LE][u8 flags][u8 record_kind][u64 LSN][payload bytes][u32 CRC32C]`. Simpler — no resource manager indirection, no backup blocks, no XID. The single-thread loop owns all the schema metadata so the rmgr layer collapses into a one-byte `record_kind`.
## Group commit — the writeonce shape
```text
loop tick:
events = epoll_wait_once();
for each readable conn fd:
read request, run handler, mutate engine
if handler committed: append WAL record, push fd to commits[]
if !commits.empty():
fsync(wal_fd)
for fd in commits: send 200 OK, mark conn writable
drain writable fds
```
Same effect as Postgres' group-commit fence (one `fsync` flushes many commits) without the IPC. The fence is the loop tick boundary itself.
## Pointers when implementing phase 11
- [`xloginsert.c:XLogInsert()`](../../../../reference/postgresql/src/backend/access/transam/xloginsert.c) — entry point for "insert this record into the WAL." Read the prologue + the LSN-assignment loop, ignore the buffer-juggling.
- [`xlog.c:XLogFlush()`](../../../../reference/postgresql/src/backend/access/transam/xlog.c) — "make this LSN durable on disk." Read the early-out for "already flushed" and the group-commit waiter logic.
- [`xlogrecovery.c:PerformWalRecovery()`](../../../../reference/postgresql/src/backend/access/transam/xlogrecovery.c) — the replay loop. Read the redo-pointer advance logic; ignore the multi-process startup signaling.
## Used by
[`docs/plan/11-wal-and-recovery.md`](../../11-wal-and-recovery.md) — WAL framing, group commit, control file, replay loop. Pair with [`linux/12-pwrite-fsync.md`](../linux/12-pwrite-fsync.md) for syscall details.

8
prompt.md Normal file
View file

@ -0,0 +1,8 @@
# db storage-runtime
- add symlink of ~/projects/postgresql to /reference
- explore src/backend/ — that's where the server lives. Key subdirectories:
eg storage/ — buffer manager, locks, shared memory (great for learning linux kernel usage)
- update docs/plan with numerically numbered markdown files. Creating a plan to have persistant storage.