databasev2 3, chain 6. Brainstormed 2026-08-28 after databasev2 4 part A
landed.
Design: compact the log by rewriting it as one record per live row into a
temp file, fsync, rename over the live WAL, fsync the parent dir, reopen.
Recovery is COMPLETELY UNCHANGED — boot still opens one file and replays
it — and the crash criterion ("the same store as if the checkpoint had
never started") is satisfied by rename, not by code we must get right.
Read .dev/reference/postgresql for this. The finding is that PG's design
is UNAVAILABLE to us, which is what makes the simpler option legitimate:
- PG never compacts its WAL; segments before the redo point are recycled
by rename or unlinked. Its records are page deltas, so a compacted redo
log is not a store — hence heap files, a control file, a redo pointer,
a second recovery source and a separate process
- ours are FULL ROW IMAGES (apply_record implements UPDATE as
remove-then-recreate), so a compacted log IS a complete store. That one
difference deletes all of the above from the design
- what IS worth porting is the ordering discipline: publish the new
"recovery starts here" atomically and LAST, so a crash falls back. PG
needs a start-of-checkpoint redo pointer plus an end-of-checkpoint
control file update; we get the same property from one rename, because
we can swap the whole data set atomically and PG cannot
Forks settled:
- no snapshot format — the compacted log is the snapshot, existing grammar,
so no new encoder or decoder and the dump reuses wo_wal_append_insert
- one source, not two
- volume-only trigger, as a ratio against the LAST compaction's measured
output (the denominator is known exactly; estimating the live set would
mean estimating Text) with an absolute floor. NO TIMER — PG's exists to
bound loss from unflushed buffers and we have none; an idle log does not
grow. Copying the mechanism without the reason was the trap
- stop-the-world, with the pause measured against a stated budget rather
than assumed acceptable; alternatives are bought against a number
- compaction may run ONLY where nothing is staged (right after a barrier),
or a staged record lands in a file about to be replaced. Normative
Recorded before it can be found late: compaction invalidates every WAL
offset iteration 2's `resident: keys` stores, so the compactor rebuilds the
offset map as it writes. Nothing breaks today because that storage half is
unimplemented — it would break later, looking like corruption.
Also corrected exploration/postgresql/buffer-and-checkpoint.md, which was
wrong on two counts: PG does NOT update its control file by rename (in-place
full-block write + CRC32C), and its checkpoint sketch assumes writeonce has
segment files, which it does not and deliberately will not.
Grounding measured on master: seed 20000 leaves a 986614-byte log; 20000
updates take it to 2590262 bytes with the SAME live rows, and boot+verify on
that store is 155ms.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
96 lines
8.2 KiB
Markdown
96 lines
8.2 KiB
Markdown
# Buffer cache & checkpoint
|
|
|
|
`storage/buffer/bufmgr.c` is Postgres' page cache: pages live in shared memory, are pinned/unpinned by readers, and are written back to disk lazily. `postmaster/checkpointer.c` is the dedicated process that periodically flushes all dirty pages, then advances the **redo pointer** in the control file — a guarantee that "everything before LSN X is on disk; recovery can start from X."
|
|
|
|
Writeonce takes the same two ideas, single-thread:
|
|
|
|
- A **clean/dirty bit per page** kept in-process. Page reads and writes hit the OS page cache directly via `pread` / `pwrite` (writeonce does not maintain its own user-space buffer pool; the kernel page cache is good enough for the workload size we target).
|
|
- A **checkpoint routine** that runs as a periodic step in the loop. It walks dirty pages, issues `fsync` on the affected files, and rewrites the control file's `last_durable_lsn`.
|
|
|
|
No separate process. No shared-buffer pinning. No dynamic-shared-memory coordination.
|
|
|
|
## Postgres source
|
|
|
|
| File | Responsibility |
|
|
| --- | --- |
|
|
| [`storage/buffer/bufmgr.c`](../../../../.dev/reference/postgresql/src/backend/storage/buffer/bufmgr.c) | Page cache front-door: `ReadBuffer`, `BufferGetPage`, `MarkBufferDirty`, `FlushBuffer`. Tracks dirty bit per buffer; pinning prevents eviction. |
|
|
| [`storage/buffer/freelist.c`](../../../../.dev/reference/postgresql/src/backend/storage/buffer/freelist.c) | Clock-sweep eviction policy. Buffers with `usage_count = 0` and `pin_count = 0` are eviction candidates; usage decremented on every sweep pass, incremented on access. |
|
|
| [`storage/buffer/buf_table.c`](../../../../.dev/reference/postgresql/src/backend/storage/buffer/buf_table.c) | Hash table from `(file, block)` → buffer slot. The lookup that `ReadBuffer` does. |
|
|
| [`postmaster/checkpointer.c`](../../../../.dev/reference/postgresql/src/backend/postmaster/checkpointer.c) | The checkpointer process. Triggered by time (`checkpoint_timeout`), WAL volume (`max_wal_size`), or signal. Runs `BufferSync()` to flush dirty buffers, then `CreateCheckPoint()` to update the control file. |
|
|
| [`postmaster/bgwriter.c`](../../../../.dev/reference/postgresql/src/backend/postmaster/bgwriter.c) | Continuously trickles dirty pages to disk between checkpoints. Smooths the I/O burst the checkpointer would cause. |
|
|
| [`storage/buffer/README`](../../../../.dev/reference/postgresql/src/backend/storage/buffer/README) | Overview of the pinning, locking, and replacement policy. Worth reading. |
|
|
|
|
## The page-cache idea worth porting
|
|
|
|
A buffer in Postgres is a `(file_id, block_number, page_data, dirty_bit, pin_count, usage_count, content_lock, io_lock)`. Strip out everything that exists for multi-process coordination (`pin_count`, locks) and you get the per-block state you need in any persistent store: **the bytes, where they came from on disk, and whether they're dirty since last fsync.**
|
|
|
|
Writeonce's phase 12 `Engine` keeps an `HashMap<(TypeName, SegmentOffset), CachedRow>` where `CachedRow = { bytes: Vec<u8>, dirty: bool }`. Rows are read on-demand (cache miss → `pread` + decode + CRC verify), written through to the segment but not flushed to disk until the next commit's WAL fsync covers them. Dirty rows accumulate; a periodic checkpoint flushes the segment fds and advances the control-file LSN.
|
|
|
|
The kernel page cache does most of the work. `pread` against an fd that already has its page cached is a memcpy. `pwrite` populates the page cache without going to disk until pressure or `fsync`. This is why writeonce explicitly does NOT use `O_DIRECT` (see [`linux/12-pwrite-fsync.md`](../linux/12-pwrite-fsync.md)) — the page cache is the one cache we want.
|
|
|
|
## Checkpoint — the writeonce shape
|
|
> **⚠ TWO CORRECTIONS, 2026-08-28** (found while brainstorming
|
|
> [databasev2 3](../../../stories/databasev2/03-wal-checkpoint.md); spec:
|
|
> [`2026-08-28-wal-checkpoint-design.md`](../../../superpowers/specs/2026-08-28-wal-checkpoint-design.md)).
|
|
>
|
|
> 1. **Postgres does NOT update its control file by rename.** The claim below
|
|
> that "Postgres does the same in `BasicOpenFile` + `fsync_parent_path`" is
|
|
> wrong: `update_controlfile` (`src/common/controldata_utils.c`) opens the
|
|
> existing file `O_WRONLY`, writes a zero-padded **full block in place**, and
|
|
> relies on **CRC32C** over the struct to detect a torn write. The
|
|
> `fsync(parent_dir)` reasoning below is still correct *for renames* — it is
|
|
> just not what Postgres does here.
|
|
> 2. **The checkpoint sketch below assumes writeonce has segment files.** It
|
|
> says records before the LSN are "*known* to be in the segment files". There
|
|
> are none: the WAL is writeonce's only durable form, replayed into RAM, and
|
|
> [databasev2 2](../../../stories/databasev2/02-table-storage-modes.md)
|
|
> deliberately rejected adding a paged store. This document predates the
|
|
> databasev2 direction, so read the loop below as a design for an
|
|
> architecture that was not chosen.
|
|
>
|
|
> What survived the comparison is the **ordering discipline**, not the
|
|
> architecture: publish the new "recovery starts here" atomically and last, so a
|
|
> crash falls back. writeonce gets that from one `rename` of the whole log —
|
|
> possible only because its records are full row images, where Postgres' are
|
|
> page deltas.
|
|
|
|
|
|
Postgres' checkpoint runs in a separate process and signals the postmaster when done. Writeonce's runs as a periodic loop step:
|
|
|
|
```text
|
|
loop tick (every CHECKPOINT_INTERVAL, e.g. 60s):
|
|
fsync(every active segment fd) // metadata + data barrier
|
|
fsync(wal_dir_fd) // ensure recent WAL writes are visible
|
|
write control.tmp { last_durable_lsn = current_wal_tail }
|
|
fsync(control.tmp)
|
|
rename(control.tmp, control)
|
|
fsync(data_dir_fd) // make the rename durable
|
|
```
|
|
|
|
The `rename(2)` is atomic on POSIX-compliant filesystems — at any crash point, either `control.tmp` is missing (the rename hasn't happened) or `control` reflects the new content. `fsync(parent_dir)` is needed because rename's atomicity is in-kernel; the directory entry isn't durable until its parent inode is synced. (Postgres does the same in `BasicOpenFile` + `fsync_parent_path`.)
|
|
|
|
Recovery on startup reads `control`, finds the `last_durable_lsn`, and replays WAL forward from there. Records before that LSN are *known* to be in the segment files; records after are replayed.
|
|
|
|
## What writeonce skips
|
|
|
|
- **Pinning + content locks.** Single-thread loop has one reader and one writer of the cache: itself. No need for `LockBuffer(BUFFER_LOCK_SHARE)` etc.
|
|
- **`bgwriter` continuous trickle.** Postgres has a *separate* process slowly cleaning the buffer pool to avoid I/O spikes at checkpoint. Writeonce's checkpoints are infrequent enough (60s default) that a spike is fine; if it becomes a problem, the same loop can do "soft flush K pages per tick" without spawning anything.
|
|
- **Hash partitioning of the buffer table.** Postgres partitions `buf_table` to reduce lock contention — single-thread doesn't have lock contention.
|
|
- **`shared_buffers` GUC.** Postgres lets the operator size the buffer pool. Writeonce trusts the OS page cache and bounds its in-process cache by an LRU with a simple count limit (`WO_CACHE_ROWS=10000` default, configurable).
|
|
|
|
## Where to look in the Postgres source
|
|
|
|
For the page cache:
|
|
- `BufferAlloc()` in `bufmgr.c` — read the function header. Strip the locks and you've got the cache-miss path.
|
|
- `BufferSync()` in `bufmgr.c` — read the prologue. The dirty-buffer-walk + per-relation fsync coalescing is the checkpoint algorithm.
|
|
|
|
For the checkpoint:
|
|
- `CreateCheckPoint()` in `xlog.c` — the control-file update sequence. Read the comments around `WriteControlFile` and the surrounding `pg_fsync`s. That's the rename-on-write pattern in practice.
|
|
- The `checkpointer.c` main loop is short and worth scanning for the time-vs-WAL-volume trigger logic.
|
|
|
|
## Used by
|
|
|
|
- `docs/plan/12-engine-disk-cutover.md` (Rust-era, removed 2026-08-18) — disk-backed engine reads and dirty-row tracking.
|
|
- `docs/plan/11-wal-and-recovery.md` (Rust-era, removed 2026-08-18) — control file write sequence (phase 11 ships the control file; checkpoint as a periodic step lands with phase 12 or shortly after).
|
|
|
|
Pair with [`linux/12-pwrite-fsync.md`](../linux/12-pwrite-fsync.md) for the fsync semantics and [`linux/08-mmap.md`](../linux/08-mmap.md) for the OS page-cache backstory.
|