writeonce/docs/plan/exploration/postgresql/00-postgresql.md
shoney.arickathil 746dc2b42b docs(databasev2): third track — the database beyond RAM, with per-table storage modes
- docs/stories/databasev2/, numbered from 1. Six PENDING database iterations
  moved from the language track and renumbered, keeping the old id in
  `was_language_iteration:` so a search for "iteration 32" still finds it:
  32 -> 3 WAL checkpoint, 23 -> 4 io_uring commit, 33 -> 7 single-file store,
  27 -> 8 query grammar, 20 -> 9 cross-program, 21 -> 10 keypair auth.
  Done work (9, 9b, 22) stays as v1 history; language 18 left whole
- the problem, read off the engine not guessed: rows are malloc'd slabs with
  addresses stable forever, NO eviction/spill/paging anywhere in database/src,
  the WAL never checkpoints so boot replays all history, and durability is one
  process-global WO_DATA so no table can say it matters more than another.
  An allocation failure IS a clean catchable WO_T_OOM — but swap thrash
  arrives first and carries no error signal at all, which is the real hazard
- four new iterations:
  1 measure the ceiling FIRST (curve not cliff; the three exits; kill -9 at
    exhaustion) — every later default should follow from a number
  2 `@table(mode: ram | durable | cold)` — the grammar ask. Small surface
    (Ast.table_cfg gains a key, the parser already rejects unknown args), big
    semantics: `durable` defaults so nothing changes silently, and the
    compiler refuses a durable row holding a `ref` into a ram table
  5 bounded tables + refuse/evict/back-pressure, shedding BEFORE the OS acts
  6 cold tiering — mostly forks, incl. whether the language surfaces the
    fault cost and whether @unique on cold is refused outright. A paged
    B-tree stays rejected: if tiering needs one, reject tiering
- 39 links repointed, link TEXT renumbered to track-local ids; arc gains one
  pointer row replacing the six moved; board + board-views cover three tracks
- linkcheck 0 broken / 0 anchors; no code blocks in any story

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-26 20:52:48 +02:00

5.7 KiB
Raw Blame History

PostgreSQL — storage subsystem reference

These cards exist to make the Postgres backend a useful library of patterns for writeonce's storage, constraint, and index work without inviting a multi-process port. (The storage cards originally fed the Rust-era plans 10–12, removed with that track 2026-08-18; the patterns fed the shipped C engine and remain the reference.) Each card pulls one subsystem out of reference/postgresql/src/backend/ — paths into the Postgres tree, the underlying idea, and the writeonce translation.

The symlink is user-specific:

ln -s /home/shoney/projects/postgresql reference/postgresql

Gitignored — see .gitignore. Pair it with reference/linux and reference/go if not already linked.

Per-subsystem cards

# Postgres area What writeonce takes What writeonce skips
wal access/transam/xlog*.c append-only sequential log, LSN-as-byte-offset, segment rollover, group commit, pwrite + fsync at commit replication/archiver, multi-process WAL writer, GUC matrix
smgr-and-md storage/smgr/{md,smgr,bulk_write}.c one file per relation, segments capped at RELSEG_SIZE, immediate vs deferred fsync multi-fork abstraction (main/fsm/vm), shared-memory descriptor cache
buffer-and-checkpoint storage/buffer/{bufmgr,freelist}.c + postmaster/{checkpointer,bgwriter}.c page cache + dirty bit + LRU; checkpoint flushes then advances control-file LSN shared-buffer pinning/unpinning, separate writer processes, latches
page-format storage/page/{bufpage,checksum}.c page header (LSN, checksum, free-space markers); CRC32C trailers MVCC visibility (xmin/xmax/ctid), access-method-specific opaque space
constraints-and-grammar parser/gram.y, catalog/pg_constraint.h, utils/adt/ri_triggers.c PK = blessed unique index; FK forward-only catalog + inline-check semantics; ON DELETE action set; backlink-implies-index (our improvement) trigger machinery, deferrable constraints, MATCH PARTIAL, composite keys
indexing-and-point-lookup access/{nbtree,hash}/README, optimizer/path/costsize.c, storage/itemptr.h hash-bucket point lookup (expected O(1)), index-entry-as-row-address (TID ↔ our slot), selectivity beats seqscan by arithmetic btree/gin/gist/spgist/brin AMs, cost-based planner, index paging

The lift-vs-skip filter

Postgres is multi-process by birth: a postmaster forks one backend per connection plus dedicated checkpointer / bgwriter / walwriter / archiver / autovacuum processes. Most of src/backend/storage/ipc/, storage/lmgr/, the latch system, and the proc.c family exist to coordinate between those processes — shared-memory regions, semaphores, condition variables, lock manager partitions, signal forwarding. Writeonce is single-process and single-threaded, so all of that mechanism is dead weight here. The concepts underneath (fairness, deadlock detection, request batching) generalize anyway, but writeonce satisfies them with single-thread invariants instead of IPC primitives.

What carries over cleanly:

  1. Sequential WAL with fsync at commit — applicable to any durable store regardless of process model.
  2. Page cache abstraction — even single-threaded engines need a dirty/clean bit and an LRU eviction story; the kernel page cache covers most of it via mmap / buffered I/O, but the dirty-tracking + flush-batching policy is something we own.
  3. Control file with last-safe-LSN — small, fixed-size, atomically updated via rename-on-write. Survives multi-process and single-process alike.
  4. Recovery = replay WAL from last checkpoint — the algorithm is identical; what writeonce skips is the postmaster signaling that says "ok, recovery is done, accept connections."
  5. CRC32C on every record + page — the cost is a few cycles per write, the pay-off is silent-corruption detection. Worth it.

What stays out:

  • Shared-memory / dynamic-shmem coordination (storage/ipc/dsm*.c, storage/lmgr/). Single-thread loop has no co-tenants.
  • Multi-version concurrency control (access/transam/clog.c, xmin/xmax tuple headers). The locked architecture (docs/runtime/database/02-wo-language.md § Concurrency Model) commits to MVCC for snapshot isolation, but the version chains are not what makes single-thread writes durable. Layered in later, when LIVE subscribers want pre-commit views.
  • Separate writer processes (postmaster/{walwriter,bgwriter,checkpointer,archiver,autovacuum}.c). Each becomes a per-tick chunk of work in the same loop, gated by deadlines.

Phase mapping

The implementation phases that lean on this material:

  • The Rust-era consumers (plans 10/11/12: storage foundations, WAL and recovery, disk cutover) were removed with that track 2026-08-18; their ideas shipped in database/src/ (typed WAL + replay) and the rest wait on databasev2 3 (checkpoint) — the wal/buffer cards are its entry material.
  • Current consumers: constraints-and-grammar (the @table PK/FK grammar direction) and indexing-and-point-lookup (the O(1) read-path slice iteration 22's numbers demand).

Pair each card with docs/plan/exploration/linux/12-pwrite-fsync.md for the actual syscalls — these cards are about design patterns, that one is about kernel calls.