fix(db2-keys): stage the schema head before the first offset capture

- `seed` of docs/examples/residency (`resident: keys`) on a FRESH WO_DATA
  segfaulted rc 139 in both the dir and the file form; pre-existing.
- Root cause: the databasev2 12 head record was staged lazily INSIDE the
  first `wo_wal_append_*` (wal.c `stage()`), after db.c:78/293 had read
  `koff = wo_wal_next_offset(w)`; the first keys-resident row was re-pointed
  at the schema record and its first read folded "record header is
  malformed"; `wo_idx_probe` borrows with `msg == NULL` -> zero-page write.
- Fix: one helper `stage_schema_head` shared by `stage()`,
  `wo_wal_ensure_schema` and `wo_wal_next_offset` (no longer a pure inline):
  the head is staged before any caller observes `off + len`. Still lazy,
  never for a log that stays empty; head-stage OOM is `wo_wal_stage_fatal`.
  db.c untouched; compaction/migrate stage the head explicitly, unaffected.
- Failing test first: `test_keys_resident_fresh_log_first_row` (test_wal.c),
  the db.c:78 sequence call for call, then read-by-id, `wo_idx_probe`,
  replay. Pre-fix: `koff != 0` FAIL, `wo_row_read` -1 "record header is
  malformed", ASan SEGV `wo_wal_fold_row_at wal.c:1871` via `table.c:373`.
- Gates: test_wal 6650/0 (was 6629); `make -C runtime test` 21 suites
  8452/0 (was 8431); wovm-asan clean; residency `seed` + restart `order`
  under wovm_asan rc 0 on a fresh dir AND a fresh app.db; a control build
  with wal.c/wal.h reverted reproduces the SEGV.
- CODE-LOGIC §Schema migrations: "Head before any offset capture" bullet.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
(cherry picked from commit 631007839451fb970e9dec83338d19c09b3043ab)
This commit is contained in:
shoney.arickathil 2026-09-10 14:50:23 +02:00
parent 58dcb1f969
commit 0b9f09fc12
4 changed files with 153 additions and 28 deletions

View file

@ -355,10 +355,25 @@ A `@table` class is the schema; the log is the database; boot compares them.
- **The log describes itself.** `WO_WAL_SCHEMA` (kind 5) is the head record - **The log describes itself.** `WO_WAL_SCHEMA` (kind 5) is the head record
of every fresh and every compacted log: per class its NAME, storage flags, of every fresh and every compacted log: per class its NAME, storage flags,
and per field name + kind + the two encoding-relevant metadata words. and per field name + kind + the two encoding-relevant metadata words.
Written lazily by `stage()` ahead of the FIRST real record — never for a Written lazily ahead of the FIRST real record — never for a log that
log that stays empty, because `durable: false` programs have a documented stays empty, because `durable: false` programs have a documented
zero-bytes contract. `apply_record` skips it before reading cid/id (its zero-bytes contract. `apply_record` skips it before reading cid/id (its
class count would be misread as a cid); replay does not count it. class count would be misread as a cid); replay does not count it.
- **Head before any offset capture (defect fix 2026-09-10).** One helper,
`stage_schema_head`, stages the pending head; `stage()` calls it on the
first append and `wo_wal_next_offset()` calls it BEFORE answering, so the
offset a caller records for a keys-resident row (`db.c`'s `koff`/`roff`,
taken before the append) can never name the head. It used to: boot sets
the schema (`main.c`, `wo_wal_set_schema`) and never forces the head, so
the first `resident: keys` row of a fresh log was re-pointed at the schema
record — its first read folded "record header is malformed", and through
`wo_idx_probe` (a borrow with `msg == NULL`) that was a zero-page write:
the residency example's `seed` died rc 139 in both `WO_DATA` forms.
`wo_wal_next_offset` is therefore no longer pure; a head-stage OOM there is
`wo_wal_stage_fatal`. Compaction and migration stage the head explicitly
on a schema-less replacement log and were never exposed. Pinned by
`test_keys_resident_fresh_log_first_row` (test_wal.c): the db.c:78
sequence call for call, then read-by-id, `wo_idx_probe`, and replay.
- **The diff is name-keyed** (`wo_schema_diff`). Classes match by name, - **The diff is name-keyed** (`wo_schema_diff`). Classes match by name,
fields by name + kind, owned references (`fclass`) by the NAME the number fields by name + kind, owned references (`fclass`) by the NAME the number
resolves to — so pure declaration reordering costs only a cid remap, which resolves to — so pure declaration reordering costs only a cid remap, which

View file

@ -473,25 +473,45 @@ void wo_wal_close(wo_wal *w) {
w->fd = -1; w->fd = -1;
} }
/* frame one payload into the staged batch */ static int stage(wo_wal *w, const wbuf *payload);
static int stage(wo_wal *w, const wbuf *payload) {
if (payload->oom) return -1; /* databasev2 12: the schema head is written LAZILY, ahead of the first real
/* databasev2 12: the schema head is written LAZILY, ahead of the first * record — never for a log that stays empty. A program whose only tables are
* real record — never for a log that stays empty. A program whose only * `durable: false` opens a WAL and appends nothing, and the documented
* tables are `durable: false` opens a WAL and appends nothing, and the * contract is that such a run writes ZERO bytes; an eager head record broke
* documented contract is that such a run writes ZERO bytes; an eager * that by 75 bytes and the residency gate caught it.
* head record broke that by 75 bytes and the residency gate caught it. *
* The flag, not the emptiness check, breaks the recursion. */ * Two callers, and the ORDER between them is the invariant: stage() (the
if (w->schema && !w->schema_written) { * first append) and wo_wal_next_offset (a caller capturing that first
* record's offset BEFORE appending it — db.c's koff/roff pattern). The head
* must be staged before either observes off + len; staged only inside the
* append, the captured offset named the head instead of the row, and the
* first keys-resident row of a fresh log read back as "record header is
* malformed" (2026-09-10). The flag, not the emptiness check, breaks the
* recursion through stage(). */
static int stage_schema_head(wo_wal *w) {
if (!w->schema || w->schema_written) return 0;
w->schema_written = 1; w->schema_written = 1;
if (w->off == 0 && w->len == 0) { if (w->off != 0 || w->len != 0) return 0; /* records exist: legacy until compaction */
wbuf sp = {0}; wbuf sp = {0};
wput(&sp, w->schema, w->schema_len); wput(&sp, w->schema, w->schema_len);
int rc = stage(w, &sp); int rc = stage(w, &sp);
free(sp.b); free(sp.b);
if (rc != 0) return -1; return rc;
} }
uint64_t wo_wal_next_offset(wo_wal *w) {
/* a head-stage failure here is the death the following append would
* have taken for the same OOM — never answer with an offset the head
* would then displace */
if (stage_schema_head(w) != 0) wo_wal_stage_fatal(w);
return w->off + w->len;
} }
/* frame one payload into the staged batch */
static int stage(wo_wal *w, const wbuf *payload) {
if (payload->oom) return -1;
if (stage_schema_head(w) != 0) return -1;
wbuf rec = {0}; wbuf rec = {0};
wput_u32(&rec, (uint32_t)payload->len); wput_u32(&rec, (uint32_t)payload->len);
wput_u32(&rec, crc32(payload->b, payload->len)); wput_u32(&rec, crc32(payload->b, payload->len));
@ -821,14 +841,9 @@ int wo_wal_ensure_schema(wo_wal *w) {
if (!w->schema) return 0; /* never set: legacy behaviour */ if (!w->schema) return 0; /* never set: legacy behaviour */
if (w->off != 0 || w->len != 0) return 0; /* records exist or staged */ if (w->off != 0 || w->len != 0) return 0; /* records exist or staged */
if (w->schema_written) return 0; if (w->schema_written) return 0;
/* go through stage() so the lazy-head flag and this path can never /* the same helper stage() and wo_wal_next_offset use, so no path can
* double-write; stage() itself emits the head when it sees the flag */ * double-write the head */
w->schema_written = 1; if (stage_schema_head(w) != 0) return -1;
wbuf p = {0};
wput(&p, w->schema, w->schema_len);
int rc = stage(w, &p);
free(p.b);
if (rc != 0) return -1;
return wo_wal_commit(w); return wo_wal_commit(w);
} }

View file

@ -236,8 +236,15 @@ int wo_wal_read_schema(const char *path, uint8_t **payload_out, uint32_t *len_ou
* *
* Call it BEFORE the append whose offset you want, and only trust the value * Call it BEFORE the append whose offset you want, and only trust the value
* after the matching wo_wal_commit returns 0 — a record whose commit failed * after the matching wo_wal_commit returns 0 — a record whose commit failed
* was never durable and its offset must not be recorded anywhere. */ * was never durable and its offset must not be recorded anywhere.
static inline uint64_t wo_wal_next_offset(const wo_wal *w) { return w->off + w->len; } *
* Not pure: on a fresh log with a schema set, the first call stages the
* lazy schema head (databasev2 12) so the answer names the CALLER's record.
* Staged inside the append instead, the head displaced the first
* keys-resident row: db.c's koff pointed at the schema record and the row
* read back as "record header is malformed" (2026-09-10). A head-stage OOM
* is wo_wal_stage_fatal — the death the append would have taken. */
uint64_t wo_wal_next_offset(wo_wal *w);
/* Open (create if missing) and preallocate [prealloc] bytes (best-effort; /* Open (create if missing) and preallocate [prealloc] bytes (best-effort;
* a filesystem without fallocate still works). Positions the write offset * a filesystem without fallocate still works). Positions the write offset

View file

@ -993,6 +993,93 @@ static const wo_classdesc KEYS_IDX_CLASSES[] = {
.idx_cnt = 1, .idx_meta = keys_idx_meta}, .idx_cnt = 1, .idx_meta = keys_idx_meta},
}; };
/* 2026-09-10 defect: the FIRST keys-resident row of a FRESH log. Boot adopts
* the compiled schema (main.c: wo_wal_set_schema, never wo_wal_ensure_schema)
* and the head record is staged lazily ahead of the first real record.
* db.c's insert arm captures the row's offset with wo_wal_next_offset BEFORE
* the append, so the head must already be staged by then — otherwise `koff`
* names the schema record and the row's first read folds "record header is
* malformed"; through wo_idx_probe (borrow with msg == NULL) that was a
* zero-page write — the residency example's `seed` died rc 139. Call for
* call the db.c:78 sequence, then the two reads `seed` performs. */
static void test_keys_resident_fresh_log_first_row(void) {
char path[128];
snprintf(path, sizeof path, "%s/keysfresh.wal", g_dir);
wo_rt rt;
T_EQ(wo_rt_init(&rt, 1 << 20, KEYS_IDX_CLASSES, 1), 0);
wo_db db;
T_EQ(wo_db_init(&db, KEYS_IDX_CLASSES, 1, 0, 1), 0);
wo_wal w;
T_EQ(wo_wal_open(&w, path, 1 << 16), 0);
db.rt = &rt; rt.wal = &w; rt.db = &db;
const char *msg = "";
static wo_schema_field f[] = {
{(const uint8_t *)"n", 1, WO_K_SCALAR, WO_SCHEMA_NONE, WO_SCHEMA_NONE},
{(const uint8_t *)"label", 5, WO_K_TEXT, WO_SCHEMA_NONE, WO_SCHEMA_NONE},
};
static wo_schema_class cls[] = {{(const uint8_t *)"row", 3, WO_CLASSF_RESIDENT_KEYS, 2, f}};
wo_schema sc = {1, cls, NULL};
T_EQ(wo_wal_set_schema(&w, &sc), 0);
T_EQ(wo_wal_read_schema(path, NULL, NULL), 1); /* still zero bytes: lazy */
/* db.c's insert arm, call for call */
wo_str *s = wo_str_new(&rt, "sku", 3);
uint64_t vals[2] = {10, (uint64_t)(uintptr_t)s};
uint64_t id = wo_row_insert(&db, 0, vals, &msg, NULL);
T_CHECK(id != 0);
uint64_t koff = wo_wal_next_offset(&w);
T_EQ(wo_wal_append_insert(&w, &db, 0, id), 0);
T_EQ(wo_wal_pend_drop(&w, 0, id, koff), 0);
T_EQ(wo_wal_commit(&w), 0);
wo_db_flush_drops(&db, &w);
wo_str_free(&rt, s);
/* the log describes itself AND koff names the row, not the head */
T_EQ(wo_wal_read_schema(path, NULL, NULL), 0);
T_CHECK(koff != 0);
uint32_t at_cid = 99;
uint64_t at_id = 0, at[2];
T_EQ(wo_wal_read_row_at(&w, &db, &rt, koff, &at_cid, &at_id, at, &msg), 0);
T_CHECK(at_cid == 0 && at_id == id);
if (at_cid == 0 && at_id == id) wo_str_free(&rt, (wo_str *)(uintptr_t)at[1]);
/* seed's first read: through the id map */
uint64_t out[2];
msg = "";
int rrc = wo_row_read(&db, &rt, 0, id, out, &msg);
T_EQ(rrc, 0);
if (rrc != 0) fprintf(stderr, " wo_row_read: %s\n", msg);
else {
T_EQ(out[0], 10);
wo_str_free(&rt, (wo_str *)(uintptr_t)out[1]);
}
/* seed's second read: the index probe borrows with msg == NULL */
uint64_t *ids;
uint32_t cnt;
T_EQ(wo_idx_probe(&db, 0, 0, 10, NULL, 0, &ids, &cnt), 1);
T_CHECK(cnt == 1 && ids[0] == id);
free(ids);
/* and a restart sees one row behind the head, readable by offset again */
wo_wal_close(&w);
wo_db_destroy(&db);
wo_db db2;
T_EQ(wo_db_init(&db2, KEYS_IDX_CLASSES, 1, 0, 1), 0);
db2.rt = &rt; rt.wal = NULL; rt.db = &db2;
T_EQ(wo_wal_replay(path, &db2), 1);
wo_wal w2;
T_EQ(wo_wal_open(&w2, path, 1 << 16), 0);
rt.wal = &w2;
db_row *r = wo_row_borrow(&db2, 0, id, &msg);
T_CHECK(r != NULL && r->slots[0] == 10);
wo_row_release(&db2, 0, r);
wo_wal_close(&w2);
wo_db_destroy(&db2);
wo_rt_destroy(&rt);
}
/* Task 3, the test that matters: updating an INDEXED column on a /* Task 3, the test that matters: updating an INDEXED column on a
* keys-resident row must move the row in the index too, not just in the * keys-resident row must move the row in the index too, not just in the
* log — queried through wo_idx_probe, the row is found by its NEW value and * log — queried through wo_idx_probe, the row is found by its NEW value and
@ -3581,6 +3668,7 @@ int main(void) {
test_fold_refuses_forward_pointing_delta(); test_fold_refuses_forward_pointing_delta();
test_keys_resident_update_field(); test_keys_resident_update_field();
test_keys_resident_update_indexed(); test_keys_resident_update_indexed();
test_keys_resident_fresh_log_first_row();
test_keys_resident_indexed_across_flatten(); test_keys_resident_indexed_across_flatten();
test_migrate_reorder_owned(); test_migrate_reorder_owned();
test_migrate_delta_splice(); test_migrate_delta_splice();