docs: retract the row-encoding rewrite — the flat record format already exists

- found during the pre-execution review of the plan, before any code
- the claim was wrong in both spec and plan: table.c's db_val_encode builds
  the IN-MEMORY slot; the FILE record is a separate encoding in wal.c and has
  been flat since iteration 9. enc_val inlines every kind recursively with no
  pointer anywhere; dec_val reads it back; a record is
  `WO_WAL_INSERT | class_id | id | <value per field>` in the
  len|crc|payload|mark frame; scan_record already preads and CRC-verifies a
  record at an arbitrary offset
- so the row encoding needs NO change, and Task 5 (a "self-contained,
  offset-based" rewrite billed as the iteration's substantive engineering) is
  DELETED, not reduced. 8 tasks -> 7, and the highest-risk task is gone
- the real difficulty is where the spec never looked: wo_wal_append_insert
  stages into a 1 MiB buffer, so a record's final offset is unknown until
  flush. Threading an accurate offset back through a buffered writer —
  correct across partial flush, failed commit and torn tail — is now Task 5's
  first two steps, with a unit test that straddles a buffer boundary and a
  case asserting no offset is published for a record that never reached disk
- dependent claims corrected: the mmap alternative's premise, the read-path
  bullet (now names scan_record/dec_val), and the self-review coverage table,
  which records the retraction rather than quietly dropping the row
- root cause worth noting: reading one layer and inferring another. Second
  time this iteration — the first was assuming WO_HEAP_MB bounded table
  storage when it bounds the VM arena
- no code written yet; linkcheck 0 broken / 0 anchors

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
shoney.arickathil 2026-08-26 23:00:18 +02:00
parent 3586650baa
commit 1ee7cce597
2 changed files with 88 additions and 74 deletions

View file

@ -47,7 +47,7 @@ engine, libc only), the existing `.wob` image format, `tests/corpus` +
(`docs/plan/oop-vm/00-wob-format.md`) in the same commit. (`docs/plan/oop-vm/00-wob-format.md`) in the same commit.
- Gates that must be green at the end of every task: `just woc-test`, - Gates that must be green at the end of every task: `just woc-test`,
`just wovm-test`, `just oop-e2e`, `just employee`, `just db-actor`. `just wovm-test`, `just oop-e2e`, `just employee`, `just db-actor`.
Tasks 6 onward add `just db-bench`. Tasks 5 onward add `just db-bench`.
--- ---
@ -173,7 +173,7 @@ create fixtures under `tests/corpus/run/`.
**Interfaces:** **Interfaces:**
- Consumes: Task 3's descriptor fields. - Consumes: Task 3's descriptor fields.
- Produces: the observable behaviour Task 8's gate asserts — no WAL growth for - Produces: the observable behaviour Task 7's gate asserts — no WAL growth for
a volatile table, and an empty table after restart. a volatile table, and an empty table after restart.
- [ ] Reach the per-table durability property at the three append sites in - [ ] Reach the per-table durability property at the three append sites in
@ -202,57 +202,47 @@ create fixtures under `tests/corpus/run/`.
--- ---
## Task 5 — self-contained row records ## Task 5 — the `resident: keys` read path
**Files:** modify `database/src/table.c` (`db_val_encode` at :56-141, > **Retraction, 2026-08-26.** This plan originally had a Task 5 that rewrote
`db_val_decode` at :149-), `database/src/table.h` (the `db_text`/`db_rec`/ > `db_val_encode`/`db_val_decode` into a "self-contained, offset-based" record
`db_multi`/`db_map` shapes), `database/src/wal.c` (record payload framing); > format, described as the iteration's one real rewrite. **That task was
create a unit suite under `runtime/test/`. > fictional and has been deleted.** `table.c`'s `db_val_encode` builds the
> *in-memory slot*; the *file* record is a separate encoding in `wal.c`, and it
> has been flat since iteration 9: `enc_val` inlines every kind recursively with
> no pointer anywhere, `dec_val` reads it back into fresh engine values, a record
> is `WO_WAL_INSERT | class_id | id | <value per field>` inside the
> `len|crc|payload|mark` frame, and `scan_record(fd, off, …)` already `pread`s
> and CRC-verifies a record at an arbitrary offset. Nothing about the row
> encoding needs to change.
>
> **The real difficulty is offset capture, and it lives in this task.**
> `wo_wal_append_insert` calls `stage()` into a buffer (opened at 1 MiB in
> `main.c:210`), so a record's final file offset is unknown at append time and
> known only when that buffer flushes. Threading an accurate offset back to the
> caller through a buffered writer — correct across a partial flush, a failed
> commit, and a torn tail — is where to expect the bugs.
**Interfaces:**
- Produces: a record encoding readable standalone from a file, and a decode
that materialises VM values from it. Task 6 depends on both.
- [ ] Establish the problem precisely before changing anything: `table.c:70-75`
allocates a `db_text` and returns its **address** as the slot word, and the
owned/multi/map cases do the same. Pointers minted by a dead process are
meaningless in a file. Write this down in `database/src/CODE-LOGIC.md` as the
reason the encoding changes.
- [ ] Define the self-contained record layout: every heap value inlined into
the record body, internal references expressed as offsets from the record's
own start, so the whole record is position-independent and copyable.
- [ ] Rewrite `db_val_encode` to emit into a caller-provided buffer in that
layout rather than returning pointers, keeping the per-kind structure
(scalar/float pass through; text and bytes inline; owned recurses; multi and
map recurse element-wise) and keeping the existing refusal of the GCREF kind
— the GC bulkhead the engine enforces even though the compiler should have
made it impossible.
- [ ] Rewrite `db_val_decode` to read that layout and allocate fresh VM values,
preserving the existing rule that decode always copies and no interior
pointer ever escapes.
- [ ] Keep the in-memory path working: a `resident: all` table still holds rows
as it does today. The encoding change is about what reaches the *file*; the
resident representation is not this task's subject and must not regress.
- [ ] Add a unit suite that round-trips every field kind — including a text
containing the record's own delimiter bytes, an empty multi, a nested owned
record, and a map with text keys — asserting byte-identical recovery.
- [ ] Add an ASan leg for the new decode path. This is where the bugs are.
- [ ] Verify: `make -C runtime test`, `make -C runtime test-iso`,
`just oop-e2e`, `just employee`, `just db-actor` green; ASan clean.
- [ ] Commit. Draft: `refactor(db): self-contained, offset-based row records`.
---
## Task 6 — the `resident: keys` read path
**Files:** modify `database/src/table.c` / `table.h` (the per-table id map, the **Files:** modify `database/src/table.c` / `table.h` (the per-table id map, the
slab path), `database/src/db.c` (insert/read/update/delete), `database/src/wal.c` slab path), `database/src/db.c` (insert/read/update/delete), `database/src/wal.c`
(boot map rebuild); create fixtures under `tests/corpus/run/`. (boot map rebuild); create fixtures under `tests/corpus/run/`.
**Interfaces:** **Interfaces:**
- Consumes: Task 3's descriptor fields, Task 5's encode/decode. - Consumes: Task 3's descriptor fields, and `wal.c`'s existing `enc_val`/`dec_val`/`scan_record` — see the note below.
- Produces: a table whose rows are not resident, serving reads by offset. - Produces: a table whose rows are not resident, serving reads by offset.
- [ ] **Offset capture first, before any map exists.** Make the staging path in
`wal.c` able to report the file offset a record will occupy. Decide between
computing it as the buffer's base file offset plus the record's position
within the buffer, or deferring the report until flush; whichever is chosen,
the offset must be wrong in *no* case — a wrong offset reads a neighbouring
record and passes its CRC.
- [ ] Prove offset capture in isolation before it has a consumer: a unit test
that appends records straddling a buffer boundary, flushes, then reads each
back by its reported offset via `scan_record` and asserts the recovered id
matches the one appended. Include a failed-commit case, where no offset must
be published for a record that never reached disk.
- [ ] For a `resident: keys` table, replace the slab retention with an - [ ] For a `resident: keys` table, replace the slab retention with an
id→offset map. Keep every index resident: the id map, each secondary index, id→offset map. Keep every index resident: the id map, each secondary index,
and each `@unique` shadow. That residency is what makes the arithmetic work and each `@unique` shadow. That residency is what makes the arithmetic work
@ -261,7 +251,7 @@ slab path), `database/src/db.c` (insert/read/update/delete), `database/src/wal.c
- [ ] Insert: append the record as Task 4 leaves it, then record id→offset - [ ] Insert: append the record as Task 4 leaves it, then record id→offset
instead of retaining a slab row. instead of retaining a slab row.
- [ ] Read by id: map lookup, `pread` at the offset, verify the CRC the frame - [ ] Read by id: map lookup, `pread` at the offset, verify the CRC the frame
already carries, decode via Task 5. Use `pread` and **not** `O_DIRECT` — the already carries, decode via `dec_val`. Use `pread` and **not** `O_DIRECT` — the
kernel page cache is deliberately the hot copy. kernel page cache is deliberately the hot copy.
- [ ] Update: append a new record, repoint the offset. The superseded record - [ ] Update: append a new record, repoint the offset. The superseded record
becomes garbage; do not attempt reclamation here — that is databasev2 3. becomes garbage; do not attempt reclamation here — that is databasev2 3.
@ -287,13 +277,13 @@ slab path), `database/src/db.c` (insert/read/update/delete), `database/src/wal.c
--- ---
## Task 7 — the two runtime refusals ## Task 6 — the two runtime refusals
**Files:** modify `runtime/src/main.c` (the `WO_DATA` block at :199-215), **Files:** modify `runtime/src/main.c` (the `WO_DATA` block at :199-215),
plus wherever per-table accounting lands from Task 6. plus wherever per-table accounting lands from Task 6.
**Interfaces:** **Interfaces:**
- Consumes: Task 3's descriptor fields, Task 6's accounting. - Consumes: Task 3's descriptor fields, Task 5's accounting.
- [ ] **Refuse `durable: true` with no `WO_DATA`.** Today `main.c:199` opens a - [ ] **Refuse `durable: true` with no `WO_DATA`.** Today `main.c:199` opens a
WAL only when the variable is set, and `db.c` skips appends when it is not — WAL only when the variable is set, and `db.c` skips appends when it is not —
@ -326,7 +316,7 @@ plus wherever per-table accounting lands from Task 6.
--- ---
## Task 8 — measure, gate, document, close out ## Task 7 — measure, gate, document, close out
**Files:** modify `scripts/db-bench.py` and `docs/examples/db-bench/`, **Files:** modify `scripts/db-bench.py` and `docs/examples/db-bench/`,
`bench/baseline.json`, `docs/plan/perf-targets.md`, `bench/baseline.json`, `docs/plan/perf-targets.md`,
@ -378,19 +368,19 @@ was not covered.
| Format / descriptor / version bump | 3 | | Format / descriptor / version bump | 3 |
| `durable: false` skips the WAL | 4 | | `durable: false` skips the WAL | 4 |
| Replay skips or refuses on mismatch | 4 | | Replay skips or refuses on mismatch | 4 |
| Self-contained offset-based records | 5 | | Self-contained offset-based records | **none — already exists in `wal.c` since iteration 9. Claim retracted 2026-08-26; see the spec section of the same name.** |
| Read / update / delete / scan / boot for non-resident | 6 | | Read / update / delete / scan / boot for non-resident, incl. offset capture | 5 |
| `@unique` and FK-restrict across the boundary | 6 | | `@unique` and FK-restrict across the boundary | 5 |
| Startup refusal: durable with no `WO_DATA` | 7 | | Startup refusal: durable with no `WO_DATA` | 6 |
| Byte budget, default fraction, breach diagnostic | 7 | | Byte budget, default fraction, breach diagnostic | 6 |
| Proof plan: baseline, amplification, crash battery | 8 | | Proof plan: baseline, amplification, crash battery | 7 |
| Docs: catalog, language surface, binding contract, CODE-LOGIC | 1, 3, 8 | | Docs: catalog, language surface, binding contract, CODE-LOGIC | 1, 3, 7 |
**Gaps found and closed during review:** the spec's escape hatch for an **Gaps found and closed during review:** the spec's escape hatch for an
intentionally ephemeral run was implied but never stated — added as an explicit intentionally ephemeral run was implied but never stated — added as an explicit
step in Task 7, because a refusal with no way forward is worse than the silent step in Task 6, because a refusal with no way forward is worse than the silent
loss it replaces. The spec's note that databasev2 3's snapshot should persist loss it replaces. The spec's note that databasev2 3's snapshot should persist
the offset map is now a recorded step in Task 6 rather than prose only. the offset map is now a recorded step in Task 5 rather than prose only.
**Deliberately not in this plan:** checkpoint and compaction (databasev2 3), **Deliberately not in this plan:** checkpoint and compaction (databasev2 3),
eviction and a resident row cache (databasev2 5), io_uring on the read path eviction and a resident row cache (databasev2 5), io_uring on the read path

View file

@ -69,7 +69,7 @@ grows two fields beside `table_name` and `indexes`.
| `durable` | `true`, `false` | `true` | `false` skips the WAL append entirely: no record, no fsync, ack from RAM, table empty after restart. | | `durable` | `true`, `false` | `true` | `false` skips the WAL append entirely: no record, no fsync, ack from RAM, table empty after restart. |
| `resident` | `all`, `keys` | `all` | `keys` keeps the id map, every secondary index and every unique shadow in RAM; rows are read from the log by offset. | | `resident` | `all`, `keys` | `all` | `keys` keeps the id map, every secondary index and every unique shadow in RAM; rows are read from the log by offset. |
`true`/`false` are already keyword tokens; `all`/`index` are parsed as the same `true`/`false` are already keyword tokens; `all`/`keys` are parsed as the same
bare identifiers the `index:` argument's column list already accepts. No lexer bare identifiers the `index:` argument's column list already accepts. No lexer
change. change.
@ -98,10 +98,11 @@ indexes are very much resident.
append the record as today, then record id→offset in the resident map instead append the record as today, then record id→offset in the resident map instead
of retaining the row in a slab. The WAL append is already the durable write; of retaining the row in a slab. The WAL append is already the durable write;
this stops discarding its payload. this stops discarding its payload.
- **Read by id** — resident map lookup, then `pread` at the offset, verify CRC, - **Read by id** — resident map lookup, then `scan_record` at the offset (which
decode into fresh VM values. The decode path already exists already `pread`s and verifies the CRC), skip the record header, and `dec_val`
(`wo_val_decode_vm` always copies; rows never hand out interior pointers), so each field. Both functions already exist in `wal.c` and are already exercised
the change is where the bytes come from. by replay; the change is that they are called on demand rather than only at
boot.
- **Update** — append a new record, repoint the offset. The superseded record - **Update** — append a new record, repoint the offset. The superseded record
becomes garbage, reclaimed by the checkpoint. becomes garbage, reclaimed by the checkpoint.
- **Delete** — append a tombstone, drop the id from the map and every index. - **Delete** — append a tombstone, drop the id from the map and every index.
@ -112,20 +113,43 @@ indexes are very much resident.
O(entire history), which is the honest cost of shipping this before O(entire history), which is the honest cost of shipping this before
[databasev2 3](../../stories/databasev2/03-wal-checkpoint.md). [databasev2 3](../../stories/databasev2/03-wal-checkpoint.md).
### Row encoding: the one real rewrite ### Row encoding: nothing to build — corrected 2026-08-26
A row slot today holds raw pointers. `table.c`'s `WO_K_TEXT` case allocates a **An earlier draft of this section was wrong and claimed the opposite.** It said
`db_text` and returns its address as the slot word; owned, multi and map do the the on-disk record was pointer-bearing and that re-encoding it was "the one real
same. Pointers minted by a dead process are meaningless in a file, so the rewrite" and the substantive engineering of this iteration. That came from
on-disk record must be **self-contained and offset-based**: every heap value reading `table.c`'s `db_val_encode`, which builds the **in-memory slot**, and
inlined into the record with internal references expressed as offsets from the inferring the file format from it. The file format is a *separate* encoding in
record's own start. `wal.c`, and it has been flat since iteration 9.
This is confined to `db_val_encode`/`db_val_decode` and the record framing. It What is already there, verified:
is the substantive engineering in this iteration and the place to expect the
bugs. `wal.c` already frames records as `len|crc|payload|mark`, so the framing - `wal.c`'s `enc_val` inlines every kind recursively with no pointer anywhere —
exists; what changes is that the payload must be readable standalone rather than text and bytes as length-then-bytes, owned as class id then fields, multi as
only replayable. element kind, length, items, map as key kind, value kind, length, pairs.
GCREF is never stored and never logged.
- `dec_val` reads that back and allocates fresh engine-owned values.
- A record is `WO_WAL_INSERT | class_id | id | <value per field>`, wrapped in
the `len|crc|payload|mark` frame.
- `scan_record(fd, off, …)` already `pread`s the record at an arbitrary offset
and verifies its CRC.
So the record is already position-independent, already carries the class id and
row id, and is already randomly addressable. The in-memory slot representation
needs **no change at all**, because it was never what reached the file.
**Where the real work is instead: capturing the offset.**
`wo_wal_append_insert` calls `stage()` into a buffer (opened at 1 MiB in
`main.c`), so a record's final file offset is not known at append time — only
when that buffer flushes. Threading an accurate offset back to the caller
through a buffered writer, and keeping it correct across a partial flush and a
torn tail, is the delicate piece of this iteration. It is a much better-defined
problem than the rewrite this section used to describe, and it is bounded to
`wal.c`'s staging path plus the map that consumes it.
Consequence for the plan: this iteration is cheaper and lower-risk than first
estimated. The task that was to perform the rewrite is deleted rather than
reduced.
### Constraints across the residency boundary ### Constraints across the residency boundary
@ -215,7 +239,7 @@ commit as the code. An older image is refused on version rather than misread.
| Alternative | Why not | | Alternative | Why not |
| --- | --- | | --- | --- |
| **`mmap` the row file** | Requires offset-based rows *and* gives up precise ack-after-fsync for the kernel's flush schedule. The repo's own mmap study only ever proposed it read-only for segment lookups. If rows become self-contained anyway, mmap is a possible later optimisation of the read path — recorded, not adopted. | | **`mmap` the row file** | Gives up precise ack-after-fsync for the kernel's flush schedule, which is the one guarantee this design will not trade. The repo's own mmap study only ever proposed it read-only for segment lookups. Since records are *already* flat and position-independent, mmap remains available later as a pure read-path optimisation over the same file — recorded, not adopted. |
| **Buffer pool with dirty-page tracking** | This is the Rust-era phase-12 design in `exploration/postgresql/buffer-and-checkpoint.md` (`CachedRow { bytes, dirty }`, `WO_CACHE_ROWS` LRU) that died with that track. It duplicates the kernel page cache, and the page cache is explicitly the cache this project wants. | | **Buffer pool with dirty-page tracking** | This is the Rust-era phase-12 design in `exploration/postgresql/buffer-and-checkpoint.md` (`CachedRow { bytes, dirty }`, `WO_CACHE_ROWS` LRU) that died with that track. It duplicates the kernel page cache, and the page cache is explicitly the cache this project wants. |
| **Paged B-tree engine** | Rejected 2026-08-18 and still rejected. Reading rows from the log we already write is not this. | | **Paged B-tree engine** | Rejected 2026-08-18 and still rejected. Reading rows from the log we already write is not this. |
| **A three-valued `mode:` enum** | Needs a name per combination. The brainstorm demonstrated that the third name is unwriteable before its mechanism is decided. | | **A three-valued `mode:` enum** | Needs a name per combination. The brainstorm demonstrated that the third name is unwriteable before its mechanism is decided. |