feat(db2-delta): lift the resident:keys refusal, prove it end to end

- loader.c: delete the INCOMPLETE-update BAIL; durable:false +
  resident:keys stays refused (nowhere to read from)
- table.c: root-cause fix for the Text-index gap — a keys-resident
  borrow now holds ENGINE values, matching wo_row_ptr's contract
  (table.h's "no VM pointer" doctrine), not a VM-decoded row. Fixes
  idx_hash/idx_cols_equal/wo_idx_probe AND db.c's GET_FIELD/PROBE
  arms with one change; reproduced pre-fix as an ASan
  heap-buffer-overflow
- docs/examples/residency: Product is genuinely resident:keys;
  residency-accept.sh's refusal leg replaced by proving the program
  runs and stock survives a restart (11/0)
- test_wal.c: oracle test drives resident:all and resident:keys
  through the same update sequence and asserts identical rows;
  Text-indexed-update test catches the representation bug; five
  pre-existing tests corrected to the fixed contract (4746/0)
- story, README, status board, CODE-LOGIC.md updated; three known
  limitations documented: mid-drain stale reads, O(N^2) replay in
  chain length, compaction blind to per-row chain length

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
(cherry picked from commit b87c68f950f01aa5e572fbb86a0f374adc83d813)
This commit is contained in:
shoney.arickathil 2026-08-30 11:09:45 +02:00
parent f138ac0abe
commit e08c26309a
9 changed files with 517 additions and 191 deletions

View file

@ -251,3 +251,76 @@ compaction count, the stop-the-world pause (max and total) and the last
compaction's size. `WO_CHECKPOINT_BYTES` and `WO_CHECKPOINT_RATIO` move the compaction's size. `WO_CHECKPOINT_BYTES` and `WO_CHECKPOINT_RATIO` move the
policy; setting a tiny floor forces compaction in a few writes, which is how the policy; setting a tiny floor forces compaction in a few writes, which is how the
gate tests it at all. gate tests it at all.
## Keys-resident updates: read-modify-append, stage-here/commit-in-caller (databasev2 2/3, 2026-08-30)
**The shape.** A keys-resident row has no slab slot to mutate — its payload
lives in the log — so `row_apply_field_keys` (table.c) does read-modify-
**append** instead of a slot swap: borrow (folds the row's current value),
append a WAL delta record (id, field, new value) chained off the row's
current offset via a back-pointer, RAM-apply the index swap. `wo_wal_fold_row_at`
is THE fold — written once, called by every reader (`wo_row_borrow`), by
replay, and by compaction — so a read, a boot, and a checkpoint can never
disagree about a chain's current value.
**Stage-here, commit-in-caller — mirrors insert exactly.** `row_apply_field_keys`
stages the delta but does **not** commit and does **not** move the id map:
table.c applies RAM and appends; `db.c` owns the barrier and the post-barrier
map move, the same split insert already used (`wo_wal_pend_drop` /
`wo_db_flush_drops` for insert; `wo_wal_pend_repoint` / `wo_db_flush_drops`
for update). The caller captures the delta's own offset via
`wo_wal_next_offset()` **before** calling in — insert's own `koff` pattern —
since nothing between that capture and `wo_wal_append_delta` stages any other
bytes on the WAL. `back_off` — the back-pointer a new delta chains from —
checks a PENDING re-point (`wo_wal_repoint_offset1`) before falling back to
the durable `wo_row_offset1`: two updates to the same row staged behind one
drain's barrier must chain to each other, not both to the row's pre-drain
offset, or the first update would be orphaned from the chain.
**The unique shadow-check runs against a THROWAWAY buffer, never `t->scratch`.**
The row under update already occupies the table's one scratch buffer
(`wo_row_borrow` refuses a nested borrow on the same table), so a candidate
probe needs a buffer of its own — `keys_fold_into`, the fold-into-a-caller-
supplied-buffer half of `wo_row_borrow`, bypasses the scratch gate for exactly
this. A candidate updated earlier in the SAME uncommitted drain has its
re-point only pending, so the candidate probe also consults
`wo_wal_repoint_offset1` — and `wo_wal_fold_row_at` itself reads the WAL's
staging buffer (not yet durable) for an offset that falls inside it, so a
same-drain candidate's NEW value is what a real `@unique` clash sees.
**A keys-resident borrow holds ENGINE values, exactly `wo_row_ptr`'s contract
— restored 2026-08-30.** `table.h`'s opening doctrine: "the engine and the VM
heap are two memory worlds crossed only by copy... a row stores NO VM
pointer." `keys_fold_into` used to decode the fold's engine output to a VM
value before handing the row back, which every OTHER reader of a borrowed row
(`db.c`'s GET_FIELD/PROBE, `wo_row_read`, and `idx_hash`/`idx_cols_equal`/
`wo_idx_probe`) was NOT written to expect — they all decode engine→VM
themselves, on the assumption a borrow is engine-encoded like a slab row.
Invisible for SCALAR/FLOAT (decode is identity either way), and un-exercised
for TEXT/BYTES because the loader refused `resident: keys` outright until
this task lifted it — nothing had ever read a keys-resident Text field
through `db.c` at all. Fixed by making `keys_fold_into` stop decoding: the
fold's engine output lands straight in the borrowed row's slots,
`wo_row_release` frees them with `db_val_free` (not `wo_drop_kind`) exactly
like `table_destroy` frees a slab row's fields, and `row_apply_field_keys`
uses its already-engine-encoded `nv` directly instead of decoding a throwaway
VM copy. No index function needed to change, and neither did `db.c`.
Reproduced as a genuine ASan heap-buffer-overflow (a `wo_str*` read through
the `db_text*` layout) before the fix, pinned by
`test_keys_resident_update_indexed_text` (`runtime/test/test_wal.c`) after it.
**Three limitations, shipped and documented rather than fixed:**
1. *Mid-drain stale reads.* A request reading a row inside the same uncommitted
drain as an earlier request's in-flight update to it may see the last
durable value. Read-your-writes holds within a request, not across requests
sharing a drain; closing it needs the fold to consult the staging buffer
generally, not only for the same-drain unique shadow-check above.
2. *Replay is O(N²) in a row's delta-chain length* — `apply_delta` folds the
pre-delta row, and `wo_row_remove` (called internally) folds the SAME
offset again, so each replayed delta re-walks its whole chain.
3. *Compaction triggers on byte ratio only* — `wo_wal_should_compact` has no
per-row delta-count signal, so one hot row (a single popular SKU) can grow
a long personal chain without moving the aggregate ratio enough to fire a
checkpoint. The no-chain-cap design decision rests on compaction bounding
length; for this shape it does not.

View file

@ -751,16 +751,28 @@ db_row *wo_row_ptr(wo_db *db, uint32_t class_id, uint64_t id) {
return slot_row(t, (uint32_t)(s1 - 1)); return slot_row(t, (uint32_t)(s1 - 1));
} }
/* keys-resident fold+decode: reads the row at [off] (the row's current /* keys-resident fold: reads the row at [off] (the row's current record) and
* record) and decodes it into VM values inside [buf] (t->row_size bytes, * folds it into [buf] (t->row_size bytes, caller-owned) — the piece
* caller-owned) — the piece wo_row_borrow and a unique shadow-check's * wo_row_borrow and a unique shadow-check's candidate probe both need,
* candidate probe both need, factored out because they cannot share a * factored out because they cannot share a buffer: wo_row_borrow writes into
* buffer: wo_row_borrow writes into t->scratch and holds it busy for the * t->scratch and holds it busy for the whole life of the borrow, so a
* whole life of the borrow, so a shadow-check that needs to look at OTHER * shadow-check that needs to look at OTHER rows of the SAME table while the
* rows of the SAME table while the row under test is still borrowed must * row under test is still borrowed must use a buffer of its own, never
* use a buffer of its own, never t->scratch. [id] is checked against what * t->scratch. [id] is checked against what the fold actually names, same as
* the fold actually names, same as wo_row_borrow always did. NULL on any * wo_row_borrow always did. NULL on any failure, *msg set.
* failure, *msg set. */ *
* table.h's opening doctrine: "the engine and the VM heap are two memory
* worlds crossed only by copy... a row stores NO VM pointer." wo_wal_fold_row_at
* hands back ENGINE-owned values (dec_val's representation, exactly what a
* slab row's own slots hold, per wal.h) — those land straight in r->slots,
* with no VM decode stage, so a keys-resident borrow matches wo_row_ptr's
* contract exactly instead of a second, divergent one. Every existing
* out-gate (wo_row_read, db.c's GET_FIELD/PROBE, idx_hash/idx_cols_equal/
* wo_idx_probe) already decodes engine->VM itself on the assumption that a
* borrowed row is engine-encoded; a decode done AGAIN here used to hand them
* a VM wo_str* reinterpreted as an engine db_text* — same bug either
* direction, invisible for scalars (decode is identity there) and silent
* wrong-bytes for Text/Bytes, which is exactly what stayed unexercised. */
static db_row *keys_fold_into(wo_db *db, uint32_t class_id, uint64_t id, static db_row *keys_fold_into(wo_db *db, uint32_t class_id, uint64_t id,
uint64_t off, uint8_t *buf, const char **msg) { uint64_t off, uint8_t *buf, const char **msg) {
const wo_classdesc *c = &db->classes[class_id]; const wo_classdesc *c = &db->classes[class_id];
@ -768,36 +780,19 @@ static db_row *keys_fold_into(wo_db *db, uint32_t class_id, uint64_t id,
uint32_t got_cid = 0; uint32_t got_cid = 0;
uint64_t got_id = 0; uint64_t got_id = 0;
/* keys-resident delta updates, Task 2: the fold, not a single-record /* keys-resident delta updates, Task 2: the fold, not a single-record
* read — a row's current offset may point at a delta, not a base row. */ * read — a row's current offset may point at a delta, not a base row.
uint64_t *eng = c->field_cnt ? calloc(c->field_cnt, sizeof *eng) : NULL; * Folds straight into r->slots: field_cnt uint64_t slots is exactly
if (c->field_cnt && !eng) { * what out_vals expects, and what a db_row already provides. */
if (msg) *msg = "out of memory"; if (wo_wal_fold_row_at((wo_wal *)db->rt->wal, db, off, &got_cid, &got_id, r->slots, msg) != 0)
return NULL; return NULL;
}
if (wo_wal_fold_row_at((wo_wal *)db->rt->wal, db, off, &got_cid, &got_id, eng, msg) != 0) {
free(eng);
return NULL;
}
if (got_cid != class_id || got_id != id) { if (got_cid != class_id || got_id != id) {
/* the offset pointed at someone else's record — a compaction that /* the offset pointed at someone else's record — a compaction that
* moved records without rebuilding this map would land here, which is * moved records without rebuilding this map would land here, which is
* exactly the obligation recorded at wo_wal_compact */ * exactly the obligation recorded at wo_wal_compact */
for (uint32_t i = 0; i < c->field_cnt; i++) wo_db_val_free(db, c->kinds[i], eng[i]); for (uint32_t i = 0; i < c->field_cnt; i++) wo_db_val_free(db, c->kinds[i], r->slots[i]);
free(eng);
if (msg) *msg = "log offset does not hold the expected row"; if (msg) *msg = "log offset does not hold the expected row";
return NULL; return NULL;
} }
/* two decode stages, same reason as wo_wal_read_row_at: the fold hands
back ENGINE-owned values, and the VM never sees those, so each one is
copied into a fresh VM value here before the engine originals free. */
int ok = 1;
for (uint32_t i = 0; i < c->field_cnt; i++) {
r->slots[i] = wo_val_decode_vm(db, db->rt, c->kinds[i], eng[i], &ok, msg);
if (!ok) break;
}
for (uint32_t i = 0; i < c->field_cnt; i++) wo_db_val_free(db, c->kinds[i], eng[i]);
free(eng);
if (!ok) return NULL;
r->id = id; r->id = id;
r->class_id = class_id; r->class_id = class_id;
r->flags = 0; r->flags = 0;
@ -850,13 +845,10 @@ void wo_row_release(wo_db *db, uint32_t class_id, db_row *r) {
db_table *t = &db->tables[class_id]; db_table *t = &db->tables[class_id];
if (!t->scratch_busy || (uint8_t *)r != t->scratch) return; /* slab-backed */ if (!t->scratch_busy || (uint8_t *)r != t->scratch) return; /* slab-backed */
const wo_classdesc *c = &db->classes[class_id]; const wo_classdesc *c = &db->classes[class_id];
/* These are VM values, not engine values. wo_wal_read_row_at is the /* These are ENGINE values, exactly what a slab row holds (keys_fold_into's
* out-gate — it always COPIES, producing fresh runtime allocations — so * contract) — freed the same way table_destroy frees a slab row's fields,
* they must be dropped through the runtime. Freeing them with the engine's * not through the runtime. */
* allocator (as this did while the materialising path was still a stub) for (uint32_t i = 0; i < c->field_cnt; i++) db_val_free(c->kinds[i], r->slots[i]);
* is a bad-free the moment a keys-resident row is actually read back. */
if (db->rt)
for (uint32_t i = 0; i < c->field_cnt; i++) wo_drop_kind(db->rt, c->kinds[i], r->slots[i]);
t->scratch_busy = 0; t->scratch_busy = 0;
} }
@ -1137,10 +1129,10 @@ static int row_apply_field_slot(wo_db *db, db_table *t, const wo_classdesc *c,
* back-pointer. [nv] is already engine-encoded (same convention as * back-pointer. [nv] is already engine-encoded (same convention as
* row_apply_field_slot); consumed on every path. * row_apply_field_slot); consumed on every path.
* *
* The borrow's materialised row holds VM values (wo_row_release drops every * The borrow's materialised row holds ENGINE values now (keys_fold_into's
* slot through the runtime), so [nv] is decoded to a VM value up front and * contract matches wo_row_ptr's), so [nv] lands in r->slots[field] directly —
* that is what ever lands in r->slots[field] — the engine encoding is used * no VM decode stage, same representation the WAL record and the index
* only for the WAL record and freed once staged. * functions already expect.
* *
* Task 4 (keys-resident delta updates) ruling: this function stages the * Task 4 (keys-resident delta updates) ruling: this function stages the
* delta but does NOT commit and does NOT move the id map — mirroring * delta but does NOT commit and does NOT move the id map — mirroring
@ -1182,15 +1174,6 @@ static int row_apply_field_keys(wo_db *db, uint32_t class_id, uint64_t id,
uint64_t pending1 = wo_wal_repoint_offset1(w, class_id, id); uint64_t pending1 = wo_wal_repoint_offset1(w, class_id, id);
uint64_t back_off = (pending1 ? pending1 : wo_row_offset1(db, class_id, id)) - 1; uint64_t back_off = (pending1 ? pending1 : wo_row_offset1(db, class_id, id)) - 1;
int ok = 1;
uint64_t nv_vm = wo_val_decode_vm(db, db->rt, c->kinds[field], nv, &ok, msg);
if (!ok) {
wo_row_release(db, class_id, r);
db_val_free(c->kinds[field], nv);
if (err_kind) *err_kind = DB_ERR_OOM;
return -1;
}
/* unique shadow-check: run with the NEW value before anything durable or /* unique shadow-check: run with the NEW value before anything durable or
indexed moves, exactly row_apply_field_slot's promise. indexed moves, exactly row_apply_field_slot's promise.
CRITICAL: candidates are probed into a THROWAWAY buffer, never CRITICAL: candidates are probed into a THROWAWAY buffer, never
@ -1201,8 +1184,8 @@ static int row_apply_field_keys(wo_db *db, uint32_t class_id, uint64_t id,
would accept duplicates). Candidates are always in this same, would accept duplicates). Candidates are always in this same,
keys-resident table, so keys_fold_into (bypassing wo_row_borrow and keys-resident table, so keys_fold_into (bypassing wo_row_borrow and
its scratch_busy gate) is safe to call directly. */ its scratch_busy gate) is safe to call directly. */
uint64_t old_vm = r->slots[field]; uint64_t old_eng = r->slots[field];
r->slots[field] = nv_vm; r->slots[field] = nv;
uint8_t *cand_buf = NULL; uint8_t *cand_buf = NULL;
for (uint32_t x = 0; x < t->index_cnt; x++) { for (uint32_t x = 0; x < t->index_cnt; x++) {
db_index *ix = &t->indexes[x]; db_index *ix = &t->indexes[x];
@ -1216,9 +1199,8 @@ static int row_apply_field_keys(wo_db *db, uint32_t class_id, uint64_t id,
if (!cand_buf) { if (!cand_buf) {
cand_buf = malloc(t->row_size); cand_buf = malloc(t->row_size);
if (!cand_buf) { if (!cand_buf) {
r->slots[field] = old_vm; r->slots[field] = old_eng;
wo_row_release(db, class_id, r); wo_row_release(db, class_id, r);
wo_drop_kind(db->rt, c->kinds[field], nv_vm);
db_val_free(c->kinds[field], nv); db_val_free(c->kinds[field], nv);
if (err_kind) *err_kind = DB_ERR_OOM; if (err_kind) *err_kind = DB_ERR_OOM;
*msg = "out of memory"; *msg = "out of memory";
@ -1244,16 +1226,16 @@ static int row_apply_field_keys(wo_db *db, uint32_t class_id, uint64_t id,
db_row *other = db_row *other =
keys_fold_into(db, class_id, b->ids[i], cand_off1 - 1, cand_buf, &obmsg); keys_fold_into(db, class_id, b->ids[i], cand_off1 - 1, cand_buf, &obmsg);
int clash = other && idx_cols_equal(c, ix, r, other); int clash = other && idx_cols_equal(c, ix, r, other);
/* keys_fold_into decoded fresh VM values for EVERY field, same /* keys_fold_into decoded fresh ENGINE values for EVERY field,
as a real borrow — nobody else owns them, so drop them here */ same as a real borrow — nobody else owns them, so drop them
here the same way wo_row_release would */
if (other) if (other)
for (uint32_t k = 0; k < c->field_cnt; k++) for (uint32_t k = 0; k < c->field_cnt; k++)
wo_drop_kind(db->rt, c->kinds[k], other->slots[k]); db_val_free(c->kinds[k], other->slots[k]);
if (clash) { if (clash) {
r->slots[field] = old_vm; /* untouched, promised */ r->slots[field] = old_eng; /* untouched, promised */
free(cand_buf); free(cand_buf);
wo_row_release(db, class_id, r); wo_row_release(db, class_id, r);
wo_drop_kind(db->rt, c->kinds[field], nv_vm);
db_val_free(c->kinds[field], nv); db_val_free(c->kinds[field], nv);
if (err_kind) *err_kind = DB_ERR_UNIQUE; if (err_kind) *err_kind = DB_ERR_UNIQUE;
*msg = "unique index violation"; *msg = "unique index violation";
@ -1262,23 +1244,21 @@ static int row_apply_field_keys(wo_db *db, uint32_t class_id, uint64_t id,
} }
} }
free(cand_buf); free(cand_buf);
r->slots[field] = old_vm; /* restored: still the OLD row until applied */ r->slots[field] = old_eng; /* restored: still the OLD row until applied */
/* RAM apply (Task 4 ruling): the borrowed row is the OLD row — out of /* RAM apply (Task 4 ruling): the borrowed row is the OLD row — out of
every index under the OLD value, then in again under the NEW one. every index under the OLD value, then in again under the NEW one.
Unconditional from here: a failure below is fatal, not a trap. */ Unconditional from here: a failure below is fatal, not a trap. */
idx_remove_row(db, t, r); idx_remove_row(db, t, r);
r->slots[field] = nv_vm; r->slots[field] = nv;
(void)idx_add_row(db, t, r); /* cannot violate uniqueness: the shadow (void)idx_add_row(db, t, r); /* cannot violate uniqueness: the shadow
check above already cleared it */ check above already cleared it */
if (wo_wal_append_delta(w, db, class_id, id, field, back_off, nv) != 0) if (wo_wal_append_delta(w, db, class_id, id, field, back_off, nv) != 0)
wo_wal_stage_fatal(w); /* RAM already moved; see the insert arm */ wo_wal_stage_fatal(w); /* RAM already moved; see the insert arm */
db_val_free(c->kinds[field], nv); /* staged now; the engine copy served
the log record */
wo_drop_kind(db->rt, c->kinds[field], old_vm); db_val_free(c->kinds[field], old_eng); /* old value done: r now holds nv */
wo_row_release(db, class_id, r); wo_row_release(db, class_id, r); /* frees r's slots, including nv, as engine values */
if (err_kind) *err_kind = DB_ERR_NONE; if (err_kind) *err_kind = DB_ERR_NONE;
return 0; return 0;
} }

View file

@ -53,47 +53,57 @@ field change does.
| `durable: true` (default) | WAL-logged, replayed at boot | ✅ works | | `durable: true` (default) | WAL-logged, replayed at boot | ✅ works |
| `durable: false` | never written to the log; costs no disk and no fsync; empty after a restart | ✅ works | | `durable: false` | never written to the log; costs no disk and no fsync; empty after a restart | ✅ works |
| `resident: all` (default) | every row's payload lives in RAM | ✅ works | | `resident: all` (default) | every row's payload lives in RAM | ✅ works |
| `resident: keys` | the id map stays resident, the payload lives in the WAL and is read back by offset | ⛔ **refused at load** | | `resident: keys` | the id map stays resident, the payload lives in the WAL and is read back by offset | ✅ works, including update |
## Why `resident: keys` is refused — and why this example is the argument ## `resident: keys`, and what it costs
It is the mode the track exists for: a catalogue is the table that outgrows RAM `Product` above is declared `resident: keys` — the mode the track exists for. A
first, so `Product` is exactly what you would want to declare keys-resident. catalogue is the table that outgrows RAM first: only the `sku -> row` id map
Storage, the read paths, scans, `@unique`, deletes and checkpoint survival all stays in memory, and each row's payload is read back from the log. Storage, the
work. **Updating such a row does not**, and `place_order` is precisely why that read paths, scans (including through the `sku` index, a Text column), `@unique`,
matters — the row has no slab slot to mutate, so the write would land in a deletes, checkpoint survival, and update all work.
materialised scratch buffer and be discarded *silently*.
The shape of the fix follows from the same example. When an order is placed **A read costs one `pread` plus every delta since the row's last checkpoint.**
only `stock` changes; `sku`, `name` and `price` do not. Appending the whole row Updating a keys-resident row has no slab slot to mutate, so it is
per sale would rewrite every field to move one integer, on the hottest write read-modify-**append**: `place_order` moving `stock` appends a small delta
path a shop has — which is the argument for appending a **delta** (id, field, record (id, field, new value) chained off the row's previous record, rather
new value) and folding it on read, with the existing checkpoint doing the fold than rewriting `sku`, `name` and `price` to change one integer — the argument
that keeps delta chains short. That design is being settled now; the loader for a delta at all, on the hottest write path a shop has. Reading the row back
refusal stands until it lands. folds that chain: the base row plus every delta not yet superseded or
checkpointed away. A row updated once costs a `pread` and one small decode on
top of the base read; a row updated many times between checkpoints costs one
decode per delta still in the chain.
So the loader refuses the annotation rather than honouring it in name only. Three limitations ship with this, on purpose documented rather than fixed:
Uncomment the `AuditEntry` block in `main.wo` and you get:
``` 1. **Mid-drain stale reads.** A request reading a row inside the same
wovm: class 0 declares `resident: keys`, which is INCOMPLETE: rows are stored uncommitted drain, while an earlier request in that drain has an in-flight
and read keys-only, but UPDATING one is not implemented (it needs update to it, may see the last durable value, not that request's write.
read-modify-append). Remove it until databasev2 2 lands updates; Read-your-writes holds within a request, not across requests sharing a
`resident: all` is what runs drain. Closing it needs the fold to consult the WAL's staging buffer
``` generally, which is materially bigger than this feature.
2. **Replay is O(N²) in a row's delta-chain length.** Each replayed delta
re-folds the whole chain back to its base record, so boot cost for one long
chain is quadratic in that chain's length.
3. **Compaction cannot see chain length.** The checkpoint that flattens delta
chains triggers on the log's overall byte ratio, not on any one row's delta
count — so a single hot row taking many small updates (a popular SKU,
exactly this example's workload) can grow a long personal chain without
moving the aggregate ratio enough to fire a checkpoint. This mode's design
deliberately does not cap chain length, trusting compaction to bound it
instead; for a hot-row workload, it may not.
Note **where** that comes from: `woc` compiles it happily and emits a `.wob`. The refusal that used to stand here was earned, not reflexive: an audit before
The annotation is a load-time property, so the compiler is green and `wovm` lifting it found that `delete` on a keys-resident table was reading a WAL byte
exits 2. offset as a slab index and freeing whatever it landed on — memory corruption,
not a missing feature — fixed and pinned by a test that SEGVs against the old
Refusing at load rather than at the first update is deliberate. A developer who code. The same audit, repeated before lifting the update refusal, found a
declared a 120 GB table keys-resident, saw it compile, and shipped would find second bug of the same shape: three index functions (and `db.c`'s field-read
the gap in production. That judgement earned its keep in a way nobody had and probe paths) were reading a keys-resident row's Text column through the
written down: an audit before relaxing the refusal found that `delete` on such wrong struct layout, reproduced as a genuine ASan heap-buffer-overflow. Fixed
a table was reading a WAL byte offset as a slab index and freeing whatever it at the root — a keys-resident row now holds the same engine-encoded values a
landed on — memory corruption, not a missing feature. It is fixed and pinned by `resident: all` row always has — and pinned by a test that reproduces the
a test that SEGVs against the old code, but the refusal is what stood in front overflow against the pre-fix code.
of it.
## What this example does NOT show ## What this example does NOT show

View file

@ -18,9 +18,22 @@
-- The second run places an order. Stock comes back decremented after a -- The second run places an order. Stock comes back decremented after a
-- restart; the cart does not come back at all. -- restart; the cart does not come back at all.
-- Precious: WAL-logged, replayed at boot. `durable: true` is the default and -- Precious AND read-selectively resident: WAL-logged, replayed at boot, but
-- is written out here only because this example is about the annotation. -- only the id -> log-offset map lives in RAM — each row's payload is read
@table(name: "products", index: [sku], durable: true, resident: all) -- back from the log. A real catalogue is the table that outgrows RAM first,
-- so this is the mode you would actually reach for one. Storage, reads,
-- scans through the `sku` index below (a Text column — exercised here, not
-- just the scalar columns other tests stuck to), `@unique`, deletes and
-- checkpoint survival all work, and so does UPDATE: `place_order` below moves
-- `stock` through a WAL delta record — read-modify-APPEND, not a slab
-- mutation, since a keys-resident row has no slab slot to mutate. See the
-- README for what a delta update costs a reader.
--
-- Only `stock` changes when an order is placed; `sku`, `name` and `price` do
-- not. Appending the whole row on every sale would rewrite every field to
-- change one integer, on the hottest write path a shop has — which is why
-- the delta is one field, not a full-row rewrite.
@table(name: "products", index: [sku], durable: true, resident: keys)
class Product { class Product {
sku: Text sku: Text
name: Text name: Text
@ -36,42 +49,6 @@ class Cart {
sku: Text sku: Text
} }
-- WHAT DOES NOT COMPILE YET, and why it is written here rather than omitted.
--
-- A real catalogue is the table that outgrows RAM first, so `Product` above is
-- exactly what you would want to declare keys-resident:
--
-- @table(name: "products", index: [sku], durable: true, resident: keys)
-- class Product {
-- sku: Text
-- name: Text
-- price: Int
-- stock: Int
-- }
--
-- `resident: keys` keeps the id map in RAM and leaves each row's payload in the
-- WAL, read back by offset. Storage, reads, scans, `@unique`, deletes and
-- checkpoint survival all work today. **Updating such a row does not**, and
-- `place_order` below is precisely why that matters: the row has no slab slot
-- to mutate, so the write would land in a materialised scratch buffer and be
-- discarded SILENTLY.
--
-- The loader therefore refuses the annotation rather than honouring it in name
-- only:
--
-- wovm: class 0 declares `resident: keys`, which is INCOMPLETE: rows are
-- stored and read keys-only, but UPDATING one is not implemented (it needs
-- read-modify-append). Remove it until databasev2 2 lands updates;
-- `resident: all` is what runs
--
-- Note WHERE it is refused: the compiler accepts it and emits a .wob — the
-- annotation is a load-time property, so `woc` is green and `wovm` exits 2.
--
-- This example is also the argument for HOW updates should land. Only `stock`
-- changes when an order is placed; `sku`, `name` and `price` do not. Appending
-- the whole row on every sale would rewrite every field to change one integer,
-- on the hottest write path a shop has.
fn count_products() -> Int { fn count_products() -> Int {
let n = 0; let n = 0;
for _p in from p in Product select p { n = n + 1; } for _p in from p in Product select p { n = n + 1; }

View file

@ -67,6 +67,68 @@ behind this board; live Obsidian Dataview views:
## ▶ NEXT PLAN ## ▶ NEXT PLAN
### Landed 2026-08-30 — keys-resident delta updates DONE, loader refusal lifted
**Implemented last time (2026-08-30):** the six-task
[keys-resident delta updates](../superpowers/plans/2026-08-30-keys-resident-delta-updates.md)
plan's final task — lifting the `runtime/src/loader.c` refusal of
`resident: keys` and proving update end to end. The refusal (databasev2 2's
Outstanding criterion) is now Met: a keys-resident row updates through a WAL
delta record, read-modify-**append**, folded back to a value by
`wo_wal_fold_row_at` on every read, replay and compaction. Proven four ways —
the fold itself (earlier tasks), group-commit staging with the id-map re-point
deferred to the post-barrier flush, replay/compaction folding delta chains the
same way reads do, and this task's oracle test
(`test_oracle_all_vs_keys_same_update_sequence`, `runtime/test/test_wal.c`)
driving the SAME update sequence against a `resident: all` table and a
`resident: keys` table and asserting byte-identical rows at every step.
`docs/examples/residency`'s `Product` table is genuinely `resident: keys` now;
`scripts/residency-accept.sh`'s gate leg inverted from "the annotation is
refused" to "the program runs and `place_order`'s stock decrement survives a
restart" (11 checks, 0 failures).
**A second bug surfaced auditing the request path before lifting the
refusal** — the same audit class that caught `delete`'s memory corruption
in the prior session. `idx_hash`, `idx_cols_equal` and `wo_idx_probe`
(`database/src/table.c`) read a TEXT column's slot as an engine `db_text*`,
but a keys-resident borrow was handing back VM-decoded `wo_str*` — a
different struct layout, reproduced as a genuine ASan heap-buffer-overflow,
not merely wrong values. The same bug was independently present in `db.c`'s
`GET_FIELD` and `PROBE` arms (inline and request-path), unaudited until now
because nothing could reach a keys-resident row through them while the
annotation was refused. Fixed at the root: a keys-resident borrow now hands
back engine values, exactly `wo_row_ptr`'s contract for `resident: all`
(`table.h`'s own "a row stores NO VM pointer" doctrine) — no index function
needed to change, and `db.c` needed none either. Pinned by
`test_keys_resident_update_indexed_text`, which reproduces the overflow
against the pre-fix code; all five pre-existing tests that read a
keys-resident Text field directly were auditing the OLD (wrong) contract and
are corrected alongside it. `test_wal` 4746/0 throughout.
**What did NOT land, by design — three limitations documented, not fixed:**
(1) mid-drain stale reads — a request reading a row in the same uncommitted
drain as an earlier request's in-flight update to it may see the last durable
value, not that write; (2) replay is O(N²) in a row's delta-chain length,
since each replayed delta re-folds the whole chain; (3) compaction triggers on
byte ratio only, with no per-row delta-count signal, so one hot row (a single
popular SKU — this feature's own motivating workload) can grow a long chain
without moving the aggregate ratio enough to checkpoint. Item 3 is the
sharper finding: the design's decision not to cap chain length rests on
compaction bounding it, and for a hot-row workload it does not. Recorded in
[the story](databasev2/02-table-storage-modes.md) and the example's README.
**.dev / reference projects used:** none — internal-only, `table.c`/`wal.c`/
`db.c` read directly to audit the request path and trace the representation
mismatch.
**Dependencies unblocked:** none newly technical — databasev2 2's own tasks 6
(the two runtime refusals: no-`WO_DATA`, the byte budget) and 7 (measure, gate,
close out) were already the next items and do not depend on this.
**Next steps:** databasev2 2 tasks 6/7, as before. `database/src/CODE-LOGIC.md`
is current with the stage-here/commit-in-caller update contract and the
engine-representation fix.
### Landed 2026-08-29 — databasev2 2 tasks 5c/5d, and a branch consolidation ### Landed 2026-08-29 — databasev2 2 tasks 5c/5d, and a branch consolidation
**Implemented last time (2026-08-29):** `resident: keys` storage and every read **Implemented last time (2026-08-29):** `resident: keys` storage and every read

View file

@ -152,13 +152,55 @@ Outstanding:
Found by asking whether the read-modify-append plan was ready, not by a gate — Found by asking whether the read-modify-append plan was ready, not by a gate —
it is unreachable today only because the loader refuses the annotation. it is unreachable today only because the loader refuses the annotation.
- **Given** an `update` to a row on a `resident: keys` table, **when** it runs, - **Given** an `update` to a row on a `resident: keys` table, **when** it runs,
**then** it is applied. ❌ **refused explicitly** by **then** it is applied. ✅ **lifted 2026-08-30.** Read-modify-**append**: a
`wo_row_update_field{,_slot}`. A keys row lives in the log with no slab slot WAL delta record chains off the row's previous offset, and `wo_wal_fold_row_at`
to mutate; writing into the borrow's scratch would discard the write — the ONE fold every reader, replay and compaction call — walks the chain
*silently*, which is the one failure this iteration must not ship. Doing it back to a value. Verified four ways: the fold itself, on a chain built by
properly is read-modify-**append** — a new record, then re-point the offset — hand (databasev2 2 tasks); the request path stages the delta under group
and that is its own piece of work. **The loader's refusal of `resident: keys` commit and defers the id-map re-point to the post-barrier flush, so a hot
stays until it lands**, so no program can reach the half-feature. row costs one fsync per DRAIN, not per update; replay and compaction fold
delta chains the same way an ordinary read does; and the oracle test
(`test_oracle_all_vs_keys_same_update_sequence`, `test_wal.c`) drives the
SAME sequence of updates against a `resident: all` table and a
`resident: keys` table and asserts the rows read byte-identical at every
step — the strongest available check that the fold agrees with ordinary
storage, since the resident table IS the oracle. `docs/examples/residency`'s
`Product` table is genuinely `resident: keys` now; `place_order`'s stock
decrement survives a restart, gated end-to-end by
`scripts/residency-accept.sh`.
**A second gap surfaced auditing the request path before lifting the
refusal — the same audit class that caught the `delete` memory corruption
below.** `idx_hash`, `idx_cols_equal` and `wo_idx_probe` (`table.c`) read a
TEXT column's slot as an engine `db_text*`, but the keys-resident fold was
handing back VM-decoded `wo_str*` — a different struct layout. Reproduced
as a genuine ASan heap-buffer-overflow, not merely wrong values, and present
too in `db.c`'s `GET_FIELD` and `PROBE` arms (inline and request-path
alike) — nobody had audited those against a keys-resident row because
nothing could reach one while the annotation was refused. Fixed at the
root rather than patched at each reader: a keys-resident borrow now hands
back engine values, exactly `wo_row_ptr`'s contract for `resident: all`
(`table.h`'s own "a row stores NO VM pointer" doctrine) — no index function
needed to change. Pinned by `test_keys_resident_update_indexed_text`, which
reproduces the heap-buffer-overflow against the pre-fix code.
**Three limitations shipped, not fixed — documented, not papered over:**
1. *Mid-drain stale reads.* A request reading a row inside the same
uncommitted drain, while an earlier request in that drain has an
in-flight update to it, may see the last durable value — read-your-writes
holds within a request, not across requests in one drain. Closing it
needs the fold to consult the WAL staging buffer generally, which is
materially bigger.
2. *Replay is O(N²) in a row's delta-chain length.* Each replayed delta
re-folds the whole chain back to its base record, so boot cost for one
long chain is quadratic.
3. *Compaction cannot see chain length.* `wo_wal_should_compact` triggers on
a byte ratio only, with no per-row delta-count trigger, so one hot row
taking many small updates — a single popular SKU, this feature's own
motivating workload — can grow a long chain without moving the aggregate
ratio enough to fire a checkpoint. The delta-updates design's decision
not to cap chain length rests on compaction bounding it instead; for
this shape it does not.
- **Given** `durable: true` and no `WO_DATA`, **when** the program starts, - **Given** `durable: true` and no `WO_DATA`, **when** the program starts,
**then** it refuses. *(task 6 — today this combination silently discards **then** it refuses. *(task 6 — today this combination silently discards
every write)* every write)*

View file

@ -183,23 +183,14 @@ int wo_load_buf(wo_module *m, const uint8_t *buf, size_t len, char *err,
* nowhere to live. woc refuses this at compile time; the loader * nowhere to live. woc refuses this at compile time; the loader
* refuses it again because what the loader accepts, the interpreter * refuses it again because what the loader accepts, the interpreter
* trusts — this combination must never reach the engine. */ * trusts — this combination must never reach the engine. */
/* databasev2 2: `resident: keys` PARSES and sets this bit, but the /* databasev2 2 (5c/5d) + databasev2 3 (keys-resident delta updates):
* storage half (tasks 5c/5d) is not implemented — rows are still fully * `resident: keys` landed in full — storage, boot, the read paths,
* resident. Accepting it would be an annotation the compiler honours * checkpoint survival, and now UPDATE (read-modify-APPEND: a delta
* in name only. * record chains off the row's current offset; wo_row_borrow folds the
* * chain back to a value on every read). The annotation used to be
* Narrowed 2026-08-29 (databasev2 2, 5c/5d): storage, boot, the read * refused here because a table you could insert into and read but not
* paths and checkpoint survival all landed. What has NOT landed is * update was a worse promise than one that never compiled — that gap
* UPDATE — a keys-resident row lives in the log with no slab slot to * is closed, so this class is accepted like any other. */
* mutate, so an update needs read-modify-append. Until that exists the
* annotation is still refused, because a table you can insert into and
* read but not update is a worse promise than one that never compiled. */
if (flags & WO_CLASSF_RESIDENT_KEYS)
BAIL("class %u declares `resident: keys`, which is INCOMPLETE: rows "
"are stored and read keys-only, but UPDATING one is not "
"implemented (it needs read-modify-append). Remove it until "
"databasev2 2 lands updates; `resident: all` is what runs",
(unsigned)i);
if ((flags & WO_CLASSF_VOLATILE) && (flags & WO_CLASSF_RESIDENT_KEYS)) if ((flags & WO_CLASSF_VOLATILE) && (flags & WO_CLASSF_RESIDENT_KEYS))
BAIL("class %u: durable:false with resident:keys — rows would have " BAIL("class %u: durable:false with resident:keys — rows would have "
"nowhere to be read from", (unsigned)i); "nowhere to be read from", (unsigned)i);

View file

@ -559,8 +559,10 @@ static void test_keys_resident_round_trip(void) {
T_CHECK(r != NULL); T_CHECK(r != NULL);
T_CHECK(r->id == id); T_CHECK(r->id == id);
T_CHECK(r->slots[0] == 4242); T_CHECK(r->slots[0] == 4242);
wo_str *back = (wo_str *)(uintptr_t)r->slots[1]; /* engine-encoded, matching wo_row_ptr's contract (table.h's "a row
T_CHECK(back != NULL && back->len == 5 && memcmp(back->data, "hello", 5) == 0); stores NO VM pointer" doctrine) — db_text, not wo_str */
db_text *back = (db_text *)(uintptr_t)r->slots[1];
T_CHECK(back != NULL && back->len == 5 && memcmp(back->bytes, "hello", 5) == 0);
wo_row_release(&db, 0, r); wo_row_release(&db, 0, r);
/* the scratch is reusable: a second borrow must succeed, which it cannot /* the scratch is reusable: a second borrow must succeed, which it cannot
@ -897,8 +899,8 @@ static void test_keys_resident_update_field(void) {
db_row *r = wo_row_borrow(&db, 0, id, &msg); db_row *r = wo_row_borrow(&db, 0, id, &msg);
T_CHECK(r != NULL); T_CHECK(r != NULL);
T_CHECK(r->slots[0] == 999); T_CHECK(r->slots[0] == 999);
wo_str *back = (wo_str *)(uintptr_t)r->slots[1]; db_text *back = (db_text *)(uintptr_t)r->slots[1];
T_CHECK(back != NULL && back->len == 5 && memcmp(back->data, "hello", 5) == 0); T_CHECK(back != NULL && back->len == 5 && memcmp(back->bytes, "hello", 5) == 0);
wo_row_release(&db, 0, r); wo_row_release(&db, 0, r);
/* the scratch must be free again — a release that skipped clearing /* the scratch must be free again — a release that skipped clearing
@ -998,6 +1000,104 @@ static void test_keys_resident_update_indexed(void) {
wo_rt_destroy(&rt); wo_rt_destroy(&rt);
} }
/* class 0: Row { n: scalar, sku: Text @unique } — the Text-representation
* gap flagged across Tasks 3, 4 and 5 and fixed in Task 6: every existing
* keys-resident index test above indexes the SCALAR column, never
* exercising idx_hash/idx_cols_equal/wo_idx_probe's WO_K_TEXT arm (nor
* db.c's GET_FIELD/PROBE arms) against a keys-resident row. The root cause
* was keys_fold_into handing back VM wo_str* where a borrowed row's slots
* are supposed to hold engine db_text* — table.h's own "a row stores NO VM
* pointer" doctrine, true for `resident: all` and silently false for
* `resident: keys` until this task. This is the test that proves the fix:
* index the TEXT column, update it, and probe by both the OLD and NEW
* value — the same shape as test_keys_resident_update_indexed, on the
* column that used to misread. */
static const uint8_t keys_text_idx_kinds[] = {WO_K_SCALAR, WO_K_TEXT};
static const uint32_t keys_text_idx_meta[] = {1 /*unique*/, 1, 1 /*col: sku (field 1)*/};
static const wo_classdesc KEYS_TEXT_IDX_CLASSES[] = {
{.name = 0, .flags = WO_CLASSF_RESIDENT_KEYS, .field_cnt = 2, .kinds = keys_text_idx_kinds,
.idx_cnt = 1, .idx_meta = keys_text_idx_meta},
};
static void test_keys_resident_update_indexed_text(void) {
char path[128];
snprintf(path, sizeof path, "%s/keystextidx.wal", g_dir);
wo_rt rt;
T_EQ(wo_rt_init(&rt, 1 << 20, KEYS_TEXT_IDX_CLASSES, 1), 0);
wo_db db;
T_EQ(wo_db_init(&db, KEYS_TEXT_IDX_CLASSES, 1, 0, 1), 0);
wo_wal w;
T_EQ(wo_wal_open(&w, path, 1 << 16), 0);
db.rt = &rt; rt.wal = &w; rt.db = &db;
const char *msg = "";
wo_str *sa = wo_str_new(&rt, "SKU-AAA", 7);
wo_str *sb = wo_str_new(&rt, "SKU-BBB", 7);
uint64_t va[2] = {1, (uint64_t)(uintptr_t)sa};
uint64_t vb[2] = {2, (uint64_t)(uintptr_t)sb};
uint64_t a = wo_row_insert(&db, 0, va, &msg, NULL);
uint64_t b = wo_row_insert(&db, 0, vb, &msg, NULL);
T_CHECK(a != 0 && b != 0);
uint64_t off_a = wo_wal_next_offset(&w);
T_EQ(wo_wal_append_insert(&w, &db, 0, a), 0);
T_EQ(wo_wal_commit(&w), 0);
T_EQ(wo_row_drop_payload(&db, 0, a, off_a), 0);
uint64_t off_b = wo_wal_next_offset(&w);
T_EQ(wo_wal_append_insert(&w, &db, 0, b), 0);
T_EQ(wo_wal_commit(&w), 0);
T_EQ(wo_row_drop_payload(&db, 0, b, off_b), 0);
/* before the update: probing "SKU-AAA" finds a — through wo_idx_probe's
verify step, which borrows the row and reads its Text slot, exactly
the path the representation bug corrupted */
uint64_t *ids;
uint32_t cnt;
T_EQ(wo_idx_probe(&db, 0, 0, 0, "SKU-AAA", 7, &ids, &cnt), 1);
T_CHECK(cnt == 1 && ids[0] == a);
free(ids);
/* update a's Text column: "SKU-AAA" -> "SKU-CCC" */
wo_str *sc = wo_str_new(&rt, "SKU-CCC", 7);
int ek = 0;
uint64_t roff = wo_wal_next_offset(&w);
T_EQ(wo_row_update_field(&db, 0, a, 1, (uint64_t)(uintptr_t)sc, &msg, &ek), 0);
T_EQ(ek, DB_ERR_NONE);
T_EQ(wo_wal_commit(&w), 0);
T_EQ(wo_row_set_offset(&db, 0, a, roff), 0);
/* found by the NEW value */
T_EQ(wo_idx_probe(&db, 0, 0, 0, "SKU-CCC", 7, &ids, &cnt), 1);
T_CHECK(cnt == 1 && ids[0] == a);
free(ids);
/* gone from the OLD one */
T_EQ(wo_idx_probe(&db, 0, 0, 0, "SKU-AAA", 7, &ids, &cnt), 1);
T_CHECK(cnt == 0 && ids == NULL);
/* b, untouched, still finds by its own value */
T_EQ(wo_idx_probe(&db, 0, 0, 0, "SKU-BBB", 7, &ids, &cnt), 1);
T_CHECK(cnt == 1 && ids[0] == b);
free(ids);
/* a genuine duplicate is still refused: updating b's sku to a's NEW
value must trip @unique — proving idx_cols_equal reads the correct
engine bytes on BOTH sides, not a coincidental symmetric misread */
wo_str *sdupe = wo_str_new(&rt, "SKU-CCC", 7);
T_EQ(wo_row_update_field(&db, 0, b, 1, (uint64_t)(uintptr_t)sdupe, &msg, &ek), -1);
T_EQ(ek, DB_ERR_UNIQUE);
/* the row itself reads back correctly through wo_row_borrow */
db_row *r = wo_row_borrow(&db, 0, a, &msg);
T_CHECK(r != NULL);
db_text *back = (db_text *)(uintptr_t)r->slots[1];
T_CHECK(back != NULL && back->len == 7 && memcmp(back->bytes, "SKU-CCC", 7) == 0);
wo_row_release(&db, 0, r);
wo_wal_close(&w);
wo_db_destroy(&db);
wo_rt_destroy(&rt);
}
/* class 0: Row { n: scalar @unique, label: Text } — same shape as /* class 0: Row { n: scalar @unique, label: Text } — same shape as
* KEYS_IDX_CLASSES, but the index is genuinely unique this time. */ * KEYS_IDX_CLASSES, but the index is genuinely unique this time. */
static const uint8_t keys_uniq_kinds[] = {WO_K_SCALAR, WO_K_TEXT}; static const uint8_t keys_uniq_kinds[] = {WO_K_SCALAR, WO_K_TEXT};
@ -1128,8 +1228,8 @@ static void test_keys_resident_two_updates_one_drain(void) {
/* both updates visible, in order */ /* both updates visible, in order */
db_row *r = wo_row_borrow(&db, 0, id, &msg); db_row *r = wo_row_borrow(&db, 0, id, &msg);
T_CHECK(r != NULL && r->slots[0] == 333); T_CHECK(r != NULL && r->slots[0] == 333);
wo_str *back = (wo_str *)(uintptr_t)r->slots[1]; db_text *back = (db_text *)(uintptr_t)r->slots[1];
T_CHECK(back != NULL && back->len == 3 && memcmp(back->data, "sku", 3) == 0); T_CHECK(back != NULL && back->len == 3 && memcmp(back->bytes, "sku", 3) == 0);
wo_row_release(&db, 0, r); wo_row_release(&db, 0, r);
/* the chain itself: delta 2's back-pointer names delta 1's OWN offset, /* the chain itself: delta 2's back-pointer names delta 1's OWN offset,
@ -1371,9 +1471,9 @@ static void test_keys_resident_survives_compaction(void) {
db_row *r = wo_row_borrow(&db, 0, ids[i], &msg); db_row *r = wo_row_borrow(&db, 0, ids[i], &msg);
T_CHECK(r != NULL); T_CHECK(r != NULL);
T_CHECK(r->slots[0] == (uint64_t)(i * 101 + 7)); T_CHECK(r->slots[0] == (uint64_t)(i * 101 + 7));
wo_str *back = (wo_str *)(uintptr_t)r->slots[1]; db_text *back = (db_text *)(uintptr_t)r->slots[1];
T_CHECK(back != NULL && back->len == (size_t)(1 + i)); T_CHECK(back != NULL && back->len == (size_t)(1 + i));
T_CHECK(memcmp(back->data, texts[i], (size_t)(1 + i)) == 0); T_CHECK(memcmp(back->bytes, texts[i], (size_t)(1 + i)) == 0);
wo_row_release(&db, 0, r); wo_row_release(&db, 0, r);
} }
wo_wal_close(&w); wo_wal_close(&w);
@ -1436,8 +1536,8 @@ static void test_keys_resident_replay(void) {
db_row *r = wo_row_borrow(&db2, 0, ids[i], &msg); db_row *r = wo_row_borrow(&db2, 0, ids[i], &msg);
T_CHECK(r != NULL); T_CHECK(r != NULL);
T_CHECK(r->slots[0] == (uint64_t)(i * 11 + 1)); T_CHECK(r->slots[0] == (uint64_t)(i * 11 + 1));
wo_str *back = (wo_str *)(uintptr_t)r->slots[1]; db_text *back = (db_text *)(uintptr_t)r->slots[1];
T_CHECK(back != NULL && back->len == 3 && memcmp(back->data, "abc", 3) == 0); T_CHECK(back != NULL && back->len == 3 && memcmp(back->bytes, "abc", 3) == 0);
wo_row_release(&db2, 0, r); wo_row_release(&db2, 0, r);
} }
wo_wal_close(&w2); wo_wal_close(&w2);
@ -2105,6 +2205,104 @@ static void test_read_row_at(void) {
wo_rt_destroy(&rt); wo_rt_destroy(&rt);
} }
/* Task 6, Step 5: the oracle test. `resident: all` never goes near a delta —
* every update is a direct slab mutation — so running the SAME sequence of
* updates against a `resident: all` table and a `resident: keys` table and
* asserting the rows read identically at every step is the strongest
* available proof that the fold agrees with ordinary storage: the resident
* table is the oracle, exactly what an independent implementation would be,
* without needing to write one. CLASSES (flags=0) and KEYS_CLASSES
* (WO_CLASSF_RESIDENT_KEYS) share the same shape — {n: scalar, label: Text}
* — already declared above for other tests. */
static void assert_rows_equal(wo_db *db_all, uint64_t id_all, wo_db *db_keys,
uint64_t id_keys, wo_rt *rt, const char *step) {
uint64_t out_all[2], out_keys[2];
const char *msg = "";
T_EQ(wo_row_read(db_all, rt, 0, id_all, out_all, &msg), 0);
T_EQ(wo_row_read(db_keys, rt, 0, id_keys, out_keys, &msg), 0);
T_CHECK(out_all[0] == out_keys[0]);
wo_str *sa = (wo_str *)(uintptr_t)out_all[1];
wo_str *sk = (wo_str *)(uintptr_t)out_keys[1];
int text_eq = (!sa && !sk) ||
(sa && sk && sa->len == sk->len && memcmp(sa->data, sk->data, sa->len) == 0);
if (!text_eq) fprintf(stderr, "oracle mismatch at %s\n", step);
T_CHECK(text_eq);
if (sa) wo_str_free(rt, sa);
if (sk) wo_str_free(rt, sk);
}
static void test_oracle_all_vs_keys_same_update_sequence(void) {
char path[128];
snprintf(path, sizeof path, "%s/oracle.wal", g_dir);
wo_rt rt;
T_EQ(wo_rt_init(&rt, 1 << 20, KEYS_CLASSES, 1), 0);
/* the oracle needs no WAL at all: row_apply_field_slot never touches one */
wo_db db_all;
T_EQ(wo_db_init(&db_all, CLASSES, 1, 0, 1), 0);
wo_db db_keys;
T_EQ(wo_db_init(&db_keys, KEYS_CLASSES, 1, 0, 1), 0);
wo_wal w;
T_EQ(wo_wal_open(&w, path, 1 << 16), 0);
db_keys.rt = &rt; rt.wal = &w; rt.db = &db_keys;
const char *msg = "";
wo_str *s0a = wo_str_new(&rt, "start", 5);
wo_str *s0k = wo_str_new(&rt, "start", 5);
uint64_t va[2] = {10, (uint64_t)(uintptr_t)s0a};
uint64_t vk[2] = {10, (uint64_t)(uintptr_t)s0k};
uint64_t id_all = wo_row_insert(&db_all, 0, va, &msg, NULL);
uint64_t id_keys = wo_row_insert(&db_keys, 0, vk, &msg, NULL);
T_CHECK(id_all != 0 && id_keys != 0);
/* drop the keys row's payload to the log now, exactly as a post-barrier
flush would — every update below folds it back out of the WAL */
uint64_t koff = wo_wal_next_offset(&w);
T_EQ(wo_wal_append_insert(&w, &db_keys, 0, id_keys), 0);
T_EQ(wo_wal_commit(&w), 0);
T_EQ(wo_row_drop_payload(&db_keys, 0, id_keys, koff), 0);
assert_rows_equal(&db_all, id_all, &db_keys, id_keys, &rt, "insert");
/* the sequence: scalar and Text fields both move, more than once each,
so the fold is exercised on a multi-hop chain the same shape a real
catalogue would build one small update at a time */
struct { int field; uint64_t scalar; const char *text; } steps[] = {
{0, 20, NULL}, {1, 0, "mid1"}, {0, 30, NULL},
{1, 0, "mid2"}, {0, 40, NULL}, {1, 0, "end"},
};
for (size_t i = 0; i < sizeof steps / sizeof steps[0]; i++) {
int ek = 0;
uint64_t val_all, val_keys;
if (steps[i].field == 0) {
val_all = steps[i].scalar;
val_keys = steps[i].scalar;
} else {
uint32_t tl = (uint32_t)strlen(steps[i].text);
val_all = (uint64_t)(uintptr_t)wo_str_new(&rt, steps[i].text, tl);
val_keys = (uint64_t)(uintptr_t)wo_str_new(&rt, steps[i].text, tl);
}
T_EQ(wo_row_update_field(&db_all, 0, id_all, (uint32_t)steps[i].field, val_all,
&msg, &ek),
0);
T_EQ(ek, DB_ERR_NONE);
uint64_t roff = wo_wal_next_offset(&w);
T_EQ(wo_row_update_field(&db_keys, 0, id_keys, (uint32_t)steps[i].field, val_keys,
&msg, &ek),
0);
T_EQ(ek, DB_ERR_NONE);
T_EQ(wo_wal_commit(&w), 0);
T_EQ(wo_row_set_offset(&db_keys, 0, id_keys, roff), 0);
char label[32];
snprintf(label, sizeof label, "step %zu", i);
assert_rows_equal(&db_all, id_all, &db_keys, id_keys, &rt, label);
}
wo_wal_close(&w);
wo_db_destroy(&db_all);
wo_db_destroy(&db_keys);
wo_rt_destroy(&rt);
}
int main(void) { int main(void) {
snprintf(g_dir, sizeof g_dir, "/tmp/wo-wal-test-XXXXXX"); snprintf(g_dir, sizeof g_dir, "/tmp/wo-wal-test-XXXXXX");
if (!mkdtemp(g_dir)) return 1; if (!mkdtemp(g_dir)) return 1;
@ -2119,6 +2317,7 @@ int main(void) {
test_fold_refuses_forward_pointing_delta(); test_fold_refuses_forward_pointing_delta();
test_keys_resident_update_field(); test_keys_resident_update_field();
test_keys_resident_update_indexed(); test_keys_resident_update_indexed();
test_keys_resident_update_indexed_text();
test_keys_resident_update_unique_violation_refused(); test_keys_resident_update_unique_violation_refused();
test_keys_resident_two_updates_one_drain(); test_keys_resident_two_updates_one_drain();
test_keys_resident_unique_clash_pending_repoint(); test_keys_resident_unique_clash_pending_repoint();
@ -2139,6 +2338,7 @@ int main(void) {
test_read_row_at(); test_read_row_at();
test_crash_battery(); test_crash_battery();
test_compact_crash_battery(); test_compact_crash_battery();
test_oracle_all_vs_keys_same_update_sequence();
/* leave the dir for a failed run's forensics only */ /* leave the dir for a failed run's forensics only */
if (!t_fail) { if (!t_fail) {
char cmd[128]; char cmd[128];

View file

@ -133,10 +133,13 @@ grep -q 'WO-E224' <<<"$dref_out" \
&& ok "durable ref into a volatile table is WO-E224" \ && ok "durable ref into a volatile table is WO-E224" \
|| bad "dangling ref refused" "got: $dref_out" || bad "dangling ref refused" "got: $dref_out"
# ---- 5. the doc example actually runs, and its README's claim holds -------- # ---- 5. the doc example actually runs, and Product really is resident: keys
# The example is the readable half of this gate. An example no gate runs rots, # The example is the readable half of this gate. An example no gate runs
# and its README quotes the loader's refusal verbatim — so that message is # rots. `Product` is declared `resident: keys` in main.wo — checks 2 and 3
# checked here too, not merely trusted. # below are this leg's whole point: the program runs, and `place_order`'s
# stock decrement on a keys-resident row survives a restart, replayed out of
# the log rather than out of a slab. That is the property a load-time refusal
# used to stand in for; now the example proves it directly instead.
LOG=/tmp/residency.log LOG=/tmp/residency.log
: > "$LOG" : > "$LOG"
echo "residency-accept: example output -> $LOG (tail -F it)" echo "residency-accept: example output -> $LOG (tail -F it)"
@ -167,18 +170,6 @@ else
bad "the doc example compiles" "see $LOG" bad "the doc example compiles" "see $LOG"
fi fi
# the README quotes this message; drift between them is a doc bug
printf '@table(name: "audit", durable: true, resident: keys)\nclass A {\n at: Int\n}\nfn main() -> Int { return 0; }\n' > "$WORK/keys.wo"
if "$WOC" --emit "$WORK/keys.wo" -o "$WORK/keys.wob" >>"$LOG" 2>&1; then
keys_out="$(WO_DATA="$WORK/exdata" "$WOVM" "$WORK/keys.wob" 2>&1)"
printf '%s\n' "$keys_out" >> "$LOG"
grep -q 'resident: keys.*INCOMPLETE' <<<"$keys_out" \
&& ok "resident: keys is refused at LOAD with the documented message" \
|| bad "keys refusal message" "got: $keys_out"
else
bad "resident: keys compiles (the refusal is load-time, not compile-time)" "woc rejected it"
fi
echo echo
echo "residency-accept: $((pass + fail)) checks, $fail failures" echo "residency-accept: $((pass + fail)) checks, $fail failures"
[[ $fail -eq 0 ]] || exit 1 [[ $fail -eq 0 ]] || exit 1