databasev2 4 part A, task 6. Mostly documentation, plus one real fix the
full battery caught.
THE FIX. The drain held EVERY DB reply until the barrier — including
reads, which stage nothing and have no stake in durability. That parked
readers behind an fsync for no reason: durable.sN.mixread.p99 rose from
~1043us to 4057us. Only a statement that actually staged a record now has
its reply held. Caught by the gate, not by review.
THE TRADE, recorded rather than smoothed over. What remains is inherent: a
barrier blocks the owner shard LONGER (more records per fsync) though LESS
OFTEN, so anything queued behind one waits. Three full runs of the same
build gave durable.sN.mixread.p99 of 1043 / 2318 / 4147us and wmix.p99 of
8758 / 20000us — a 2-4x spread with the box near idle. So part A buys ~3x
write throughput at the cost of a longer, noisier tail on the owner shard,
and that is the strongest argument for part B (submit and keep serving).
- durable.sN.*.p99us tolerance widened to 100% WITH the reason in the
code: a 2-4x-variable tail gated at 50% gates the disk, not the engine.
The floor is the real guard and is not slack — mixread's (4172us) came
within 25us of tripping on the worst run. Baseline refreshed; a fresh
full run then passed 106 checks 0 failures
EXIT STATUS MOVED 3 -> 74 (sysexits EX_IOERR). 3 and 4 are already used by
SAMPLES for their own meanings — db-bench's own `verify` exits 3 on a
checksum mismatch, and it is the gate that exercises durability, so a
durability abort exiting 3 would have been indistinguishable from the
mismatch it should help diagnose. The low range belongs to programs.
Docs:
- story: progress, the payoff measured two ways, the cost side, criteria
split met/outstanding, and a "part B — its premise changed" section:
it was justified by "close the 66x gap", but that gap is two problems
and only the concurrent one was a batching problem
- board: standup entry in the six-question shape; both databasev2 4 rows
rewritten. They had said "close the 66x gap" — recorded as MIS-STATED
rather than quietly renumbered
- 00-wob-format.md and 04-db-binding.md: the normative failure contract
("a failed WAL commit traps WO_T_IO after un-applying the row") was
false; corrected, along with the tick-scoped group commit that never
happened
- database/src/CODE-LOGIC.md: where the barrier runs and why there, why
replies are held, why the inline path is asymmetric, the one failure
rule, and how to measure it
- db-bench README: the wmix mode, the env knobs, and the tmpfs warning
Battery: wovm-test 36 suites 0 fail, woc-test, oop-e2e 119/0,
db-bench 106/0, linkcheck clean.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
155 lines
8 KiB
Markdown
155 lines
8 KiB
Markdown
# DB binding — row format, id discipline, WAL layout, query subset
|
||
|
||
> Normative companion to the engine plan
|
||
> ([`2026-08-01-db-engine-binding.md`](../../superpowers/plans/2026-08-01-db-engine-binding.md)),
|
||
> the way `00-wob-format.md` is normative for the image. Grows with the
|
||
> plan's tasks; this revision covers **Task 1 (row storage)**. Memory-safety
|
||
> doctrine lives in the 9b design's section 6 (the copy bulkhead) — this doc
|
||
> is the *format*.
|
||
|
||
## Two memory worlds, one crossing rule
|
||
|
||
Rows store **no VM pointer**, ever. Values cross from VM heap to row storage
|
||
by copy on insert, and back by copy on read (`wo_row_read` allocates fresh VM
|
||
values from the shard's runtime). The engine's own allocations are plain
|
||
malloc — never the VM arena, so table growth cannot eat the program's heap
|
||
cap, and a heap-exhausted program can still read its data.
|
||
|
||
## Row format
|
||
|
||
```
|
||
row := header slots
|
||
header := id u64 | class_id u32 | flags u32 (16 bytes)
|
||
slots := field_cnt × u64, declaration order (the VM object shape)
|
||
```
|
||
|
||
One 8-byte slot per field, kind-driven — the same kind bytes the `.wob`
|
||
class table carries, walked the same way the VM walks them:
|
||
|
||
| kind | slot holds | engine-owned shape |
|
||
| --- | --- | --- |
|
||
| `SCALAR` | the 8 bytes themselves | — (`WO_NIL_SCALAR` spells a `?scalar` nil) |
|
||
| `TEXT` | pointer, 0 = nil | `db_text { len u32; bytes[] }` |
|
||
| `OWNED` | pointer, 0 = nil | `db_rec { class_id u32; slots[] }` — flattened by value, recursively through these same rules |
|
||
| `MULTI` | pointer, 0 = nil | `db_multi { elem_kind u8; len u32; items[] }`, elements encoded element-wise |
|
||
| `MAP` | pointer, 0 = nil | `db_map { key_kind, val_kind u8; len u32; kv pairs }` |
|
||
| `GCREF` | **never stored** | compile error upstream (the GC bulkhead); the engine refuses it defensively as an encode error |
|
||
|
||
`ref T` is a `SCALAR` at this layer — the target row's id. The engine learns
|
||
what it references only when the FK checks land (9b plan, Task 3).
|
||
|
||
## Storage
|
||
|
||
Per shard, per class, created lazily on first insert:
|
||
|
||
- **Slabs** of 256 rows (`DB_SLAB_ROWS`), malloc'd, **never moved or freed
|
||
while the table lives** — a row's address is stable for its lifetime,
|
||
which is the property 9b's loop-scoped row views stand on.
|
||
- An **occupancy bitmap** (one bit per slot, slab-major) and a LIFO
|
||
**free-slot list**: removal recycles the slot; a recycled slot is always
|
||
used before a new slab grows. Ids are never reused; slots are.
|
||
- The **primary index**: an open-addressing hash, id → slot, splitmix64
|
||
finalizer, power-of-two capacity, 0.7 load, tombstoned deletes (ids are
|
||
never 0 and never reused, so the all-ones sentinel cannot collide).
|
||
|
||
## Id discipline
|
||
|
||
Per table, per shard: shard S of N allocates `S+1, S+1+N, S+1+2N, …` — the
|
||
c-runtime plan's shipped interleave. Creation is coordination-free; a row's
|
||
owner shard is `(id-1) % N`. Milestone 1 runs at N=1 and everything
|
||
degenerates to `1, 2, 3, …`. Id 0 does not exist (it is the hash's "empty"
|
||
and the `?ref`'s nil).
|
||
|
||
## Choke points
|
||
|
||
`wo_row_insert` and `wo_row_remove` are the only functions that mutate a
|
||
table. Task 4's secondary indexes hook exactly these two sites (marked
|
||
`INDEX HOOK` in `database/src/table.c`); the WAL (Task 2) stages its record
|
||
beside the same calls. Anything else touching a slab is a defect by
|
||
definition — the doctrine the Rust engine learned and this engine enforces.
|
||
|
||
## WAL (Task 2) — `database/src/wal.{c,h}`
|
||
|
||
Record framing, replay-whole-or-not-at-all:
|
||
|
||
```
|
||
record := len u32 | crc u32 | payload | mark u32
|
||
len = payload bytes (never 0: a zero length is the preallocated tail)
|
||
crc = CRC32 (poly 0xEDB88320) of the payload
|
||
mark = 0x574F4C31 "WOL1", the last bytes of the record — a record
|
||
without its mark is torn by definition
|
||
payload := kind u8 | class_id u32 | row_id u64 | body
|
||
kind = 1 insert (body = fields), 2 remove (no body), 3 update (Task 5)
|
||
```
|
||
|
||
Body fields walk the class table's kinds: `SCALAR` 8 bytes; `TEXT` u32 len +
|
||
bytes (`0xFFFFFFFF` = nil); `OWNED` presence u8 then class id + fields
|
||
recursively; `MULTI` presence + elem kind + len + elements; `MAP` presence +
|
||
both kinds + len + pairs. Little-endian, same platform note as the loader.
|
||
|
||
**Commit order (doctrine, verbatim from the shipped phase-D pattern):** RAM
|
||
apply → stage record → `wo_wal_commit` (one pwrite of the batch + one
|
||
fdatasync) → only then acknowledge. Group commit = everything staged since
|
||
the last commit rides one sync.
|
||
|
||
**Replay** decodes payloads straight into engine-owned values — no VM heap
|
||
involved, boot cannot depend on a VM existing — and rows re-enter through
|
||
the choke-point row API, so Task 4's indexes rebuild for free. A torn tail
|
||
(short record, bad CRC, missing mark, zero length) ends the intact prefix:
|
||
everything from the tear on is dropped whole, and `wo_wal_open` positions
|
||
its write offset AT the tear so the next commit overwrites it. A record that
|
||
CRC-passes but does not decode is corruption, not a tear — replay fails
|
||
loudly. A missing file is a fresh boot, not an error. After replay each
|
||
table's `next_id` sits past every replayed id this shard owns.
|
||
|
||
**Oracle:** `wo_wal_check(path)` walks a file with no engine and reports the
|
||
intact record count and prefix end — the crash battery's verifier
|
||
(`runtime/test/test_wal.c`: five rounds of insert/commit/ack-over-pipe with
|
||
SIGKILL mid-stream; every acked row present and exact after replay).
|
||
|
||
## Insert (Task 3) — builtin 61, `database/src/db.c`
|
||
|
||
`insert Class { field: expr, … }` is a typed expression (statement position
|
||
included): fields validate like a constructor literal (defaults and `?`
|
||
fields omittable — an omitted `?scalar` gets `WO_NIL_SCALAR`, other omitted
|
||
optionals the zero word, declared defaults their value), and the result is
|
||
the new row's id. Lowering emits builtin **61**: R[B] = class-id constant,
|
||
R[B+1..] = one slot per declared field in declaration order (the literal's
|
||
order is irrelevant — slots are the class table's).
|
||
|
||
Execution: `wo_row_insert` (RAM, engine copies every value), then — when
|
||
durability is on — stage, then a barrier before the acknowledgment. **Updated
|
||
2026-08-28 (databasev2 4 part A): group commit landed, and the barrier's
|
||
location now depends on which path the statement takes.**
|
||
|
||
A statement arriving from a worker shard marshals to shard 0 and parks; shard 0
|
||
stages every such request, issues **one** barrier when its queue empties, and
|
||
only then releases the held replies — so each writer is acknowledged after the
|
||
barrier that carried *its* record. A statement already running on shard 0 takes
|
||
the inline path and still commits before the builtin returns, because it has no
|
||
reply to hold: it returns into its own fiber, and batching it would require
|
||
parking that fiber on the barrier (deferred to part B). The boundary is the
|
||
queue draining, **not** the tick this document previously anticipated — a tick
|
||
would add latency to a lone writer, taxing an idle system to serve a busy one.
|
||
|
||
Measured: ~2.9× durable write throughput and ~2.1× lower p50 on a
|
||
write-concurrent workload; unchanged for a serial writer, which has nothing to
|
||
batch with.
|
||
|
||
A failed commit **no longer traps — it ends the process** (exit 74, with a
|
||
diagnostic naming the operation, log path, `errno` and batch size). So does a
|
||
failed staging. `WO_T_IO` is unreachable from a DB write. One rule: once a
|
||
statement has mutated RAM, the outcomes are durable or death. Engine failures
|
||
still trap `WO_T_DB`. Durability is opt-in: `WO_DATA=<dir>` makes the CLI replay
|
||
`<dir>/shard-0.wal` before the entry runs and commit every insert; without
|
||
it the engine is RAM-only (every corpus fixture runs that way).
|
||
|
||
Ownership: the engine copies at the row API, so an insert **borrows** its
|
||
field values — no transfer, no E304; freshly built values are dropped at the
|
||
site (emit.ml mirrors the push/set reap). The insert node is trap-capable
|
||
(unique violations arrive with Task 4) and carries a live-mask drop entry.
|
||
|
||
## Still to come in this document
|
||
|
||
- **Task 4**: secondary-index format, `@unique` trap code.
|
||
- **Task 5**: the select subset, update record semantics, and its builtins.
|