writeonce/database/src/wal.h
shoney.arickathil 18a56f24ec feat(db): wo_wal_read_row_at — materialise a row from a log offset
Task 5b of docs/superpowers/plans/2026-08-26-table-residency.md, whose Task 5
is now split 5a-5d (plan updated in this commit).

- the offset twin of wo_row_read (table.c:721): same out-gate contract —
  every value handed back is a FRESH VM allocation — but resolved from a file
  position instead of the id hash
- fits entirely in wal.c because everything it needs was already public:
  scan_record and dec_val are local, and wo_val_decode_vm / wo_db_val_free
  are exported at table.h:183-187. Two decode stages, since the record and
  the VM speak different dialects: dec_val -> engine slots -> VM copies, with
  the engine slots freed as scratch on every path
- ZERO storage change. Nothing calls it yet; that is the point of separating
  it from 5c, so the read path can be proven before the slabs are touched
- refuses rather than guessing, each case distinguishable: no intact record
  at the offset, a malformed header, a decode failure, trailing bytes, and a
  REMOVE tombstone. That last one matters most — handing a tombstone back as
  a row would read a deleted row as live

Tested by deep field comparison, not by "it parsed": 24 rows with a nil Text
every third row, each read back BY OFFSET and compared field by field,
including the string bytes. Plus all three refusal paths — tombstone, a
mid-record offset (the silent-wrong-row failure this guards), and past the
intact prefix.

The free-on-every-path claim is VERIFIED, not assumed: removing the free
produced 3 LeakSanitizer reports; restoring it returns to 0. Worth doing
because "ASan is clean" only means something if the harness would have
complained.

PLAN SPLIT: Task 5's storage half was written as if it were plumbing. Measured
instead: wo_row_ptr returns a db_row* into a slab with 11 call sites, table.c
has 37 slab references, db.c:105-181 scans slabs directly, enc_val serialises
FROM the slab, and no operation exists that drops a payload while keeping
index entries. Note this is the OPPOSITE half from the earlier retraction —
the record FORMAT needed nothing, the record STORAGE genuinely is deep. 5c
(id->offset map + drop-payload-keep-index) and 5d (rewiring the call sites,
scans, @unique/FK across the boundary) get their own write-ups.

Gates: test_wal 3654/0 (was 3428), all 18 runtime suites 0 fail under
ASan+UBSan, oop-e2e 119/0, residency 8/0, employee 8/0, db-actor 8/0.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-27 10:27:40 +02:00

137 lines
6.7 KiB
C

/* wal.h — typed-row write-ahead log + boot replay (iteration 9, Task 2).
*
* The c-runtime plan's shipped pattern (phases D/E), generalized to typed
* rows. The commit order is doctrine, verbatim:
*
* RAM apply → wal_append (staged) → wal_commit (write + fdatasync)
* → only then is the write ACKNOWLEDGED
*
* Record framing — replay-whole-or-not-at-all:
*
* record := len u32 | crc u32 | payload | mark u32
* len = payload byte count (never 0; 0 = preallocated tail, stop)
* crc = CRC32 of payload
* mark = 0x574F4C31 "WOL1" — written LAST, so a record without its
* mark is torn by definition
* payload := kind u8 | class_id u32 | row_id u64 | body
* kind : 1 insert (body = the row's fields, engine encoding below)
* 2 remove (no body)
* 3 update (reserved for Task 5)
*
* Field encoding in a body walks the class table's kinds:
* SCALAR 8 bytes
* TEXT u32 len | bytes (0xFFFFFFFF = nil)
* OWNED u8 0 = nil, or u8 1 | u32 class_id | fields recursively
* MULTI u8 0 = nil, or u8 1 | u8 elem_kind | u32 len | elements
* MAP u8 0 = nil, or u8 1 | u8 kk | u8 vk | u32 len | k v pairs
*
* Replay decodes payloads STRAIGHT into engine-owned values — the VM heap
* is never involved (boot must not depend on a VM existing yet), and rows
* re-enter through the same choke-point row API, so Task 4's indexes are
* rebuilt for free. A torn tail (short record, bad CRC, missing mark) drops
* everything from the tear onward — never a partial record, never a record
* after a tear. Little-endian on-disk, matching the .wob loader's platform
* note.
*
* wo_wal_check is the offline oracle the crash battery verifies with: it
* walks a WAL file with no engine at all and reports how many records are
* intact and where the intact prefix ends. */
#ifndef WO_WAL_H
#define WO_WAL_H
#include "table.h"
#define WO_WAL_MARK 0x574F4C31u /* "WOL1" LE */
enum { WO_WAL_INSERT = 1, WO_WAL_REMOVE = 2, WO_WAL_UPDATE = 3 };
typedef struct wo_wal {
int fd;
uint64_t off; /* next write offset (the intact tail) */
/* staged batch: appended by wal_append_*, flushed by wal_commit */
uint8_t *buf;
size_t len, cap;
} wo_wal;
/* databasev2 2: the file offset the NEXT staged record will occupy.
*
* Exact, and knowable at append time — no deferral to flush is needed, which
* is what the design spec feared. `off` is the durable tail and `len` the
* bytes staged but not yet written, and wo_wal_commit pwrites the whole batch
* AT `off` before advancing it, so a record staged now lands at off+len.
*
* Correct across the two awkward cases:
* - a failed commit leaves `off` unadvanced and `len` intact, so the batch
* is rewritten from the same place and previously-reported offsets stay
* valid;
* - a torn tail is handled by wo_wal_open, which positions `off` at the end
* of the INTACT prefix, so offsets are always relative to validated data.
*
* Call it BEFORE the append whose offset you want, and only trust the value
* after the matching wo_wal_commit returns 0 — a record whose commit failed
* was never durable and its offset must not be recorded anywhere. */
static inline uint64_t wo_wal_next_offset(const wo_wal *w) { return w->off + w->len; }
/* Open (create if missing) and preallocate [prealloc] bytes (best-effort;
* a filesystem without fallocate still works). Positions the write offset
* at the end of the INTACT record prefix — an existing file is scanned the
* same way replay scans it, so a torn tail is overwritten, not appended
* after. 0 ok, -1 errno-style failure. */
int wo_wal_open(wo_wal *w, const char *path, uint64_t prealloc);
void wo_wal_close(wo_wal *w);
/* Stage a record for the row that MUST already be applied to RAM (the
* commit-order doctrine). Insert/update read the row via wo_row_ptr.
* 0 ok, -1 OOM / no such row. */
int wo_wal_append_insert(wo_wal *w, wo_db *db, uint32_t class_id, uint64_t id);
int wo_wal_append_remove(wo_wal *w, uint32_t class_id, uint64_t id);
/* UPDATE re-logs the whole row (KISS: replay replaces — remove + re-create
* with the same id; the prefix/suffix delta trick from the survey is a
* later optimization, recorded). Call AFTER the RAM update. */
int wo_wal_append_update(wo_wal *w, wo_db *db, uint32_t class_id, uint64_t id);
/* Write the staged batch and fdatasync — the ack line. Empty batch = ok,
* no syscall. 0 ok, -1 write/sync failure (the batch stays staged). */
int wo_wal_commit(wo_wal *w);
/* Boot replay: apply every intact record to [db] in order. Ids re-enter
* exactly as logged; each table's next_id advances past the replayed ids
* that belong to this shard. Returns the number of records applied, or -1
* on open failure / a record naming an unknown class (corruption beyond
* what a torn tail explains). A torn tail is NOT an error: replay applies
* the intact prefix and reports it. */
int64_t wo_wal_replay(const char *path, wo_db *db);
/* databasev2 2: as wo_wal_replay, but distinguishes the two failure kinds.
* Returns the applied count on success; -1 on corruption beyond a torn tail;
* -2 when the log holds records for a class the loaded image declares
* `durable: false`, writing that class id through [volatile_cid] if non-NULL.
* The plain wo_wal_replay above is this with NULL, kept so the existing
* callers and the 156 WAL unit checks are untouched. */
int64_t wo_wal_replay_ex(const char *path, wo_db *db, uint32_t *volatile_cid);
/* databasev2 2: read one row straight from a log offset — the offset twin of
* wo_row_read. [out_vals] must have room for the class's field_cnt values and
* receives FRESH VM allocations (the out-gate: always copies). [class_out] and
* [id_out] are optional. Offsets come from wo_wal_next_offset, recorded at
* append time.
*
* 0 ok
* -1 no intact record at that offset, a malformed header, a record that
* does not decode, trailing bytes, or a REMOVE tombstone (which carries
* no fields — refused rather than decoded, since returning a deleted row
* as live is the worst outcome available here)
* -2 out of memory (*msg set)
*
* Nothing in the engine calls this yet: it is the read half of `resident:
* keys`, landed ahead of the storage change so it can be tested alone. */
int wo_wal_read_row_at(wo_wal *w, wo_db *db, wo_rt *rt, uint64_t off,
uint32_t *class_out, uint64_t *id_out, uint64_t *out_vals,
const char **msg);
/* Offline verification (no engine): scan [path], count intact records.
* *intact_bytes (optional) = where the intact prefix ends. -1 = open
* failure. */
int64_t wo_wal_check(const char *path, uint64_t *intact_bytes);
#endif /* WO_WAL_H */