- wal.h/wal.c: wo_wal_fold_row_at — THE fold. Walks BACKWARD from an
offset through WO_WAL_DELTA records, remembering the first value
seen per field index (newest wins, since newest is seen first),
stops at the first INSERT/UPDATE, decodes it, overlays resolved
fields. Returns ENGINE-owned values so reads, replay, and
compaction (Tasks 3/5) can all build on the same output.
- Cycle guard: caps the walk at what the log up to the starting
offset could possibly hold (13 = scan_record's own record-size
floor), so a corrupt or malicious back-pointer fails loudly
instead of spinning.
- table.c: wo_row_borrow's keys arm now calls the fold instead of
wo_wal_read_row_at directly, then VM-decodes the result — same
two-stage pattern wo_wal_read_row_at used internally. Per-table
scratch, scratch_busy nested-borrow refusal, and the cid/id
identity check all preserved unchanged.
- resident: all path (wo_row_ptr) untouched.
- test_wal.c: two new tests — deltas on two different fields (changed
fields take the new value, the untouched field keeps its original)
and two deltas on the SAME field (the newer wins, pinning direction
— a reversed fold would pass with the older value instead).
Verified failing pre-implementation (wo_row_borrow returned NULL
since a delta record isn't INSERT/UPDATE) and passing after.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
(cherry picked from commit a60231cde1d49d74743cedbd134d2da11158b70b)
- review finding: field_idx=0 and back_off=0 (fresh WAL, offset 0) meant
a u32/u64 swap of these two values wrote identical zero bytes either
way — undetectable by the prior assertions
- test_delta_record: new dedicated 3-scalar-field class (not shared
KEYS_CLASSES) so field_idx can be a nonzero, fixed-8-byte value without
a Text field's variable-length encoding complicating the fixed body
size assertion
- stage+commit a filler row first so the target row's insert record (the
delta's back-pointer) lands at a nonzero offset, not the WAL's initial 0
- delta now targets field_idx=2 with back_off=base_off, both nonzero and
distinct from each other and from class_id=0
- class_id stays 0: this fixture registers exactly one class, so there is
no other value to give it without an unused second class purely to
shift an index
- verified live: temporarily swapped the field_idx/back_off wput calls in
wal.c, confirmed test_wal now fails (fidx==49 want 2, back==2 want 49),
then reverted — wal.c diff is a no-op, only the test changed
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
(cherry picked from commit 20ba0965e15e200641b8b63893a2f238a0279fba)
- enum: add WO_WAL_DELTA = 4, existing 1/2/3 untouched (on-disk logs)
- wal.h: document kind 4's payload shape in the format docblock
- wal.h: declare wo_wal_append_delta(w, db, class_id, id, field_idx,
back_off, value) — back-pointer taken as a parameter, not looked up,
keeping the encoder ignorant of table/map state
- wal.c: implement it, modeled on wo_wal_append_insert's shape —
wput_u8/u32/u64 the header fields, enc_val the one field, stage()
- test_wal.c: new test_delta_record — stages a delta after an insert,
commits, then preads the raw record and asserts kind/class/id/
field_idx/back-pointer/value all round-trip; registered in main()
- nothing reads deltas back yet — decode/apply is a later task
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
(cherry picked from commit 9c6f832c0534a59e644c53b7cd2850581da12159)
- wo_row_remove's keys arm borrows the row from the log to find its
index entries, and a borrow reads through db->rt->wal. At boot that
pointer is not wired yet: main.c replays first (main.c:226) and
assigns rt.wal afterwards (main.c:268)
- so the borrow found no log, the remove failed, and replay reported a
valid tombstone as CORRUPTION. An UPDATE record would have failed the
same way, since replay applies it as remove-then-recreate
- replay now lends the runtime a read-only view over the fd it already
has open, for the replay's duration only, and restores what was there
- broken by the delete fix in 76b8fd9 — deletes worked in-process but
their tombstones broke the next boot. Unreachable in production only
because the loader still refuses the annotation
- pinned by test_keys_resident_delete_then_replay, verified failing
against the unfixed code (2 failures) and clean with it
- found by asking whether the read-modify-append plan was ready, not by
a gate
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
(cherry picked from commit dc25462461b9f79d70c803f7174adc90fa16c90e)
- wo_row_remove read the id map's value as a slot, but on a keys table
that value is a LOG OFFSET (hput(t, id, wal_off + 1)). slot_row does
no bounds check, so a delete indexed t->slabs[] with a byte offset and
then called db_val_free on whatever it landed on — arbitrary frees,
not a wrong answer
- keys tables now take their own arm: no slab slot, no bitmap bit, no
free-list entry to return. The index hook needs the row's values, so
the row is borrowed from the log for exactly that long
- wo_row_ptr carried the same trap and is public. It cannot refuse keys
tables outright (insert legitimately calls it while the map still
holds a slot), so it now detects the offset case — index past the
slabs, or bitmap bit clear — and returns NULL. Callers all handle NULL
- test_keys_resident_delete pins it; it SEGVs against the old code,
verified by reverting the fix rather than assumed
- found while auditing every hget() reader before narrowing the loader
refusal to allow benchmarking. The refusal was justified in the docs
by "updates are unimplemented" while actually standing in front of
this too: a guard whose stated reason is narrower than its real one
gets removed by someone who believes the stated reason
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
(cherry picked from commit 76b8fd944af9ed062467bdf9ab93c2e96dd198cf)
- loader's resident:keys refusal said "rows are still fully resident"
and "until tasks 5c/5d land". Both false since f606fc9. Corrected to
name the real blocker: UPDATE needs read-modify-append
- databasev2 00-story: the sequence graph drew 2->3->4, which reads as
3 needing 2 and 4 needing 3. Both backwards, and it still drew the
2->5->6 path the 2026-08-27 amendment retired. Redrawn stating only
real dependencies, with 4 and 3 shown as composing rather than
ordered, and the execution order that actually happened
- databasev2 03: the hazard and its Outstanding entry both claimed
nothing fails "because iteration 2's storage half is unimplemented".
Marked discharged, and recorded that the hazard named only half the
danger — the bitmap walk would have dropped keys rows outright
- databasev2 06: pending -> hold (largely superseded, revisit only on
a measurement); dated its 5c/5d references
- porch 01: rewritten to the settled shape. readiness ready, status
in-progress, phases B and C marked superseded with why
- porch 01 claimed time.after "is still a reserved builtin id". False —
builtin 90, implemented. That claim is what made the iteration look
cheaper than it is
- porch README gains honest ledger rows for both features (partial,
being rebuilt), not shipped
- skill-catalog README pointed at a story path that moved tracks;
linkcheck now 0 broken
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
(cherry picked from commit b3d8c403e1d19ac27ec966de85cb293e0765795c)
- wo_row_read and the @unique shadow probe go through borrow/release;
release runs before every exit, including wo_row_read's early return
- updates on a keys table refused explicitly in wo_row_update_field and
the slot variant: no slab slot to mutate, and writing the borrow's
scratch would discard the write silently. Needs read-modify-append
- compaction walked the bitmap, which a keys row has no bit in — every
such row would have been dropped from the new log. Now walks
wo_row_next_id and re-points each row to where it lands
- moves records byte-for-byte (copy_record) rather than decoding: a
borrowed row holds VM values, enc_val expects engine values, and ASan
caught that mismatch as a 4294967292-byte memcpy
- wo_row_set_offset updates a value in place and never rehashes, so a
wo_row_next_id cursor stays valid while compaction re-points
- a compaction that fails after moving rows is fatal: the map would name
an unlinked temp file, and the intact log replays correctly
- test_keys_resident_survives_compaction pins both failure modes; rows
rewrite in hash order so offsets really move
- loader still refuses resident: keys — updates are not implemented
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
(cherry picked from commit f606fc9b76d983cac2b348f03f7d4f01433cd905)
databasev2 2, task 5c. The write and boot halves. Still not exposed: the
loader refuses `resident: keys` until 5d rewires the readers.
GROUP COMMIT FORCED THE DESIGN. A keys-resident payload can only be dropped
once its record is durable, but databasev2 4 deferred the barrier to the drain
— so at append time the bytes are still in the staging buffer and the recorded
offset would pread ZEROS. Dropping at append would have produced rows that
read as garbage, intermittently, only under multi-shard load.
So the drop is recorded, not performed:
- wo_wal gains a pending-drop list, the same shape as the drain's held replies
and for the same reason
- both write paths take the offset BEFORE the append (wo_wal_next_offset) and
record it; the inline path flushes right after its own commit, the request
path's flush runs in the drain immediately after the barrier
- if the process dies before the barrier the list dies with it, which is
correct: nothing was dropped and nothing was lost
- an out-of-memory pend is ignored on purpose — the row simply stays resident,
which is safe
Boot: replay now leaves a keys-resident table pointing at the LOG. Each record
is applied normally, so indexes and uniqueness are built exactly as for any
other table, and the payload is then dropped with THAT record's offset. For an
update the later record wins, because each apply overwrites the map in order —
the rule replay already follows.
Tests: the round trip (insert, commit, drop, read back with Text intact) and
now BOOT — a fresh wo_db replays the store and every row materialises from the
log, count intact, nothing in a slab.
Verified: just wovm-test — 36 suites 0 fail, test_wal 4301 pass, cli_smoke OK.
REMAINING (5d), and precise: every reader still goes through wo_row_ptr, which
for a keys table would index a freed slot. The scans in db.c walk the BITMAP,
and a keys table's bitmap is empty by construction — so a query over one would
today return no rows at all. That, FK restrict, and the @unique shadow are 5d,
and the loader refusal stays until they land.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
(cherry picked from commit 08abd09bf88918f2582e74713dc7903beb8aaeb8)
databasev2 2, task 5c step 2. The storage half the accessor was left waiting
for. Not yet wired into insert (that and 5d remain), and the loader still
refuses `resident: keys`, so nothing is exposed to a program yet.
- wo_db gains an `rt` back-pointer, set in main.c beside VM.rt.db. wo_rt
already carries `db` and `wal` as opaque handles, so this closes the loop
and a borrow can reach the log WITHOUT threading a wal pointer through
eleven call sites — which is the whole reason 5c is one accessor
- wo_row_drop_payload: the operation the plan recorded as MISSING. Frees the
slot and its engine-owned values, then re-points the id map at the record's
log offset (off + 1, reusing the same 0-is-empty trick as slot + 1). It
deliberately does NOT touch the secondary indexes (they store row ids, so
they stay correct), does NOT decrement count (the row is still live, only
its backing moved), and does NOT remove the id (that is how it is found)
- wo_row_borrow materialises for a keys table: reads the offset from the id
map, calls 5b's wo_wal_read_row_at into the per-table scratch, and checks
the record actually holds the expected class and id — a compaction that
moved records without rebuilding the map lands exactly there, which is the
obligation recorded at wo_wal_compact
- fully-resident tables keep today's path and pay one predicate
A REAL BUG, exposed the first time the path was used: wo_row_release freed the
materialised values with the ENGINE's allocator. They are VM values —
wo_wal_read_row_at is the out-gate and always copies — so ASan reported a
bad-free immediately. It now drops them through the runtime. That stub was
written in 5c step 1 for a path that did not exist yet.
Recorded while implementing: wo_wal_next_offset's contract says to trust an
offset "only after the matching commit returns 0". Group commit (databasev2 4)
defers that barrier to the drain, so db.c can no longer check inline — but part
A also made a failed commit FATAL, so no execution can record an offset whose
record never became durable. Same guarantee, different mechanism.
Test: a heap-valued row is inserted, committed, has its payload dropped, and is
read back out of the log with its Text intact; count is unchanged (still live);
and a second borrow succeeds, which fails if release did not clear the scratch.
Verified: just wovm-test — 36 suites 0 fail, test_wal 4273 pass, cli_smoke OK.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
(cherry picked from commit 125bd09218d616b2a16b140de770d3f38b45f0ac)
The branch was 17 ahead / 25 behind with 11 conflicting files, and drifting
further: db.c had been rewritten twice on master since (group commit, then
compaction). Resolved rather than rebased so both histories stay legible.
Conflicts, and how each was settled:
- db.c: BOTH semantics kept. Master's fatal path and compaction check now sit
behind the branch's `table_is_durable` predicate, in all three inline arms —
a volatile table reaches neither the barrier nor the compaction check
- db-bench sample: every mode from both sides (growth, growth-verify, randread,
replayseed, wmix) and ONE `boot` mode, which both sides had added
independently
- db-bench.py: all six legs kept. Both sides had also grown the same
WAL-size helper under different names; collapsed into one
- perf-targets: the branch's §5 (RAM ceiling) then master's §6/§7 — master's
numbering had already assumed a §5 it did not have
- story frontmatter: master's `status` (the landing truth) plus the branch's
`readiness` axis. 03 would have read `done` + `refine`, which is a
contradiction — it was brainstormed and landed on master, so `ready`
- board: both standup blocks newest-first; master's chain rows (a superset);
the branch's databasev2 1-2 rows with master's 3-4. Fixed a stray `|` in
master's row 3
- baseline: master's, then REGENERATED from a full campaign — 143 metrics,
132 checks, 0 failures with both sides' legs present
TWO HALF-EXPOSED FEATURES FIXED, because the merge rule is that master gets
no feature that is honoured in name only:
- `resident: keys` PARSED, set a .wob flag, and did nothing: rows stayed fully
resident. A developer could declare a 120 GB table keys-resident, watch it
compile, and be OOM-killed. The loader now REFUSES it with a message naming
what to write instead, until tasks 5c/5d land. The compiler still parses it
and its AST golden still passes, so the grammar work stays tested
- `durable: false` was honoured ONLY on the inline path. wo_db_exec_req had no
guard at all, so a volatile table written from an actor on a worker shard
would still be logged — precisely porch's session-table case, and precisely
what iteration 2 exists to provide. All three request-path arms now carry the
same predicate. Found by reading the merged code, not by a test: the obvious
probe runs main() on the primary and therefore only exercises the inline path
Verified on the merged tree: wovm-test 0, woc-test 0, oop-e2e 122/0,
residency-accept 8/0, db-bench 132/0, linkcheck clean.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
databasev2 3, task 5.
Full campaign, same workload twice, differing only in whether
checkpointing may fire:
- WAL used 1962358 -> 907094 bytes (2.16x reclaimed)
- boot 114 -> 64 ms (1.78x), median of 3
- stop-the-world pause max 2651us against a STATED 50ms budget
The budget is asserted, not assumed: 50ms is a stall a serving process
can absorb without a client seeing a timeout, and the leg fails if it is
exceeded. The pause is O(live rows) — at ~181 MB/s a 1GB live set implies
~5.5s, which is the number an incremental design must be bought against.
The spec deliberately did not buy it in advance.
FOUND BY MEASURING: the dump was 8x slower than it needed to be. It
flushed through wo_wal_commit, which fdatasyncs, so it paid one barrier
per 256 records. Intermediate durability there is worthless — the temp is
not authoritative until the rename and is fsynced once immediately before
it. With a single final barrier:
- ~107KB live: 23948us -> 2903us
- ~500KB live: 36361us -> 7526us
- ~1.98MB live: 107649us -> 13212us
- marginal ~22 MB/s -> ~181 MB/s, sync-bound to bandwidth-bound
Correctness re-proven after that change: wovm-test 36 suites 0 fail,
test_wal 760 pass including the 40-round kill-during-compaction battery.
Two measurement defects of my own, fixed rather than reported:
- boot measured through the driver's run() helper reported 251ms both
with and without checkpointing — run() samples RSS on a 250ms poll, so
every timing floors at the quantum. Measured directly instead, median
of 3
- ckpt.reclaim_x was recorded as lower-is-better by the default detector,
which would have PASSED "reclaimed nothing" and FAILED an improvement:
the feature's central claim, gated backwards. Now higher-is-better,
gated at 15% while the wall-clock metrics stay wide — waiving them all
would have left the leg ungated, part A's task 4 mistake
- sample gains a `boot` mode that does nothing, so boot time is boot time
- walstats now reports compactions, pause max/total and compacted bytes
- baseline refreshed from the FULL campaign (N=20000, crash_reps=3), and
a fresh full run passes 116 checks 0 failures
- gate bites: reclaim_x doctored to 1.0 -> FAIL on exactly that metric
One flake seen and checked, not papered over: durable.sN.query.ops_sec
failed once at 53% below baseline. It is a read-only metric that touches
no WAL code, and a re-run passed 116/0 with the box at load 1.85.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
databasev2 3, task 4. The plan called this the riskiest task because a race
can pass by luck, so it is argued with mutants rather than green runs.
The battery: a forked child inserts, acks, deletes the oldest so HISTORY
grows while the live set stays ~17, and compacts every 24 iterations. The
parent SIGKILLs at varied instants so kills land before, inside and after
rewrites, then replays and checks the ACKED LIVE SET. The existing
battery's "records >= acks" oracle cannot be reused: collapsing history is
exactly what compaction is for.
A REAL DEFECT IN MY FIRST VERSION, found by the failures and fixed in the
TEST, not by weakening it:
- the child acked deletes AFTER committing them, so a kill in between left
the row legitimately gone on disk while the last ack still said
"inserted" — the parent then demanded a row the engine was right to
remove. Symptom was an acked insert missing near the end of the stream,
~1 run in 3
- deletes now announce INTENT BEFORE committing, so such a row's fate is
simply UNKNOWN to the parent, which is the honest thing to assert. Every
acked insert never marked for deletion must still be present with its
acked value
- the stale-temp assertion was also wrong: it checked for absence after
wo_wal_replay, which never opens the WAL. The guarantee is "removed AT
OPEN", so the test now opens and then asserts. A temp surviving a kill
is expected debris, not a defect
Proven to have teeth, which matters because assertions were softened:
- against the design's rejected alternative (in-place rewrite instead of
the atomic rename) it fails EVERY run, reporting log_records=0 — the
kill landed mid-copy and destroyed the log. That is the corruption
rename exists to prevent
- on correct code: 10 consecutive runs x 40 rounds clean, plus the suite
Also carried the log's PREALLOCATION to the replacement. The WAL is
preallocated so appends never extend the file, which is what lets
fdatasync alone be the ack barrier; a replacement opened with prealloc 0
silently changes that property and the zero-padded tail the scan relies
on. Stated honestly: this is hygiene making the replacement equivalent to
what open() would have produced — I could NOT prove it was the cause of
the observed loss, and the ack race above explains it.
Verified: just wovm-test — 36 suites 0 fail, test_wal 760 pass, cli_smoke OK.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
databasev2 3, task 3.
- wo_wal_should_compact is a PURE decision (used bytes, last compaction's
measured output, floor, ratio) so it is testable without a store —
which is the only way a policy like this gets tested at all. Denominator
is the last compaction's real output, not an estimate of the live set:
estimating would mean estimating Text
- 8 boundary assertions incl. "exactly 3x is not MORE than 3x" and a zero
ratio disabling the policy rather than dividing by nothing
- MUTATION-TESTED instead of observing RED: implementation and test were
written together, so removing the floor check was verified to fail
exactly the two floor assertions. Equivalent evidence, stated plainly
- WO_CHECKPOINT_BYTES / WO_CHECKPOINT_RATIO at boot beside WO_MAILBOX.
The knobs are what make the policy testable — a gate sets a tiny floor
and forces compaction in a few writes instead of megabytes
- NO timer, per the spec: Postgres' CheckPointTimeout bounds loss from
unflushed buffers; our records are durable at commit and an idle log
does not grow
- the ordering rule is now asserted, not trusted: a test stages a record,
requests compaction, and requires REFUSAL with the log untouched and
the staged record still committable afterwards
FOUND AND FIXED a gap in my own wiring. The plan said to call the check
"after the drain's barrier", and I did — but a statement running ON the
owner shard never enters that drain, so WO_SHARDS=1 never compacted and
its log grew forever: measured 536086 bytes where the multi-shard run
held 446024. Now checked after the inline path's commit too (db.c
maybe_compact), where the buffer is equally empty. WO_SHARDS=1 went
536086 -> 260657 bytes. For a checkpoint this mattered more than part A's
equivalent gap: an unbounded log is an operational failure, not just lost
throughput.
Also corrected a measurement of my own: multi-shard logs looked unbounded
(448KB -> 1013KB -> 1647KB across 8k/24k/48k updates). They are not.
Instrumentation showed compaction ran 25 times with zero failures, each
writing MORE than the last, because the live set genuinely grows — wmix's
hist_dump and done-markers are themselves durable inserts. Final log
1631040 against a last compaction of 866432 is a ratio of 1.88, just under
the 2x threshold: the policy holding exactly.
Replies are released BEFORE compaction runs, deliberately: their records
are already durable, and holding them across a stop-the-world rewrite
would add its full duration to their latency for nothing.
Verified: wovm-test 36 suites 0 fail, test_wal 360 pass; db-bench-quick
crash.s1/crash.sN and both restart legs green, and part A still batches
(sN mean 4.16, peak 30).
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
databasev2 3, task 2.
- wo_wal_open removes `<log>.compact` before reading anything. The only
way one exists is a crash before the rename, which means its records
were never authoritative
- deleted rather than ignored, deliberately: a file full of well-formed
records sitting beside the log is exactly what a future reader mistakes
for data
Test uses PLAUSIBLE content, not garbage — a byte copy of a real log —
because garbage would be rejected by the CRC anyway and would prove
nothing. It asserts the temp is present before the open, gone after, and
that the live log still replays to exactly what it said.
RED was an assertion failure (`access(tmp, F_OK) != 0` unmet), not a
compile error, so the test was proven to exercise the behaviour before the
behaviour existed.
Verified: just wovm-test — 36 suites 0 fail, test_wal 340 pass, cli_smoke OK.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
databasev2 3, task 1.
- walks each class's live rows via the bitmap-over-slabs pattern db.c
already uses in three places, appending one INSERT per live row through
the EXISTING append path. No second encoder, no new format, and ids are
preserved exactly because wo_wal_append_insert takes the id and reads
the row from the store
- FLUSHES EVERY 256 RECORDS rather than staging the whole store: stage()
grows the staging buffer by doubling and never shrinks it, so a
one-buffer dump would hold the entire store in RAM on top of the store
— the unbounded growth databasev2 1 identified as how this engine dies
- the switch, in order: fsync the temp file, rename over the live path,
fsync the PARENT DIRECTORY (rename's atomicity is in-kernel; the
directory entry is not durable until the parent is synced — Postgres
does the same for the same reason), then reopen the descriptor, because
the old one refers to an unlinked inode
- REFUSES when anything is staged: those records would land in a file
about to be replaced. The caller-side guard is task 3; this is the
backstop
- a failure leaves the ORIGINAL log intact and usable and returns -1. A
failed checkpoint is a missed optimisation, not a durability event, so
it deliberately does NOT take databasev2 4's fatal path
- records the bytes written, so task 3's trigger can compare against a
measured denominator instead of estimating the live set (which would
mean estimating Text)
Recovery is untouched — the result is an ordinary log in the ordinary
grammar, replayed from byte 0. Crash safety comes from rename, not from
code of ours.
Test asserts BOTH halves, on purpose:
- the log shrinks: 43 records (3 inserts + 40 updates of the SAME row, so
history grows while the live set does not) -> 3 records, fewer bytes
- AND a fresh replay reproduces the store: every id present, and row 0
carries the 40th update's value rather than its original. "It got
shorter" is also true of a truncating bug, so the replay comparison is
what actually proves it
- and the WAL stays usable after the swap: a further append lands after
the compacted records, giving 4 on the next check
Verified: just wovm-test — 36 suites 0 fail, test_wal 315 pass (was 165),
cli_smoke OK.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
databasev2 4 part A, task 6. Mostly documentation, plus one real fix the
full battery caught.
THE FIX. The drain held EVERY DB reply until the barrier — including
reads, which stage nothing and have no stake in durability. That parked
readers behind an fsync for no reason: durable.sN.mixread.p99 rose from
~1043us to 4057us. Only a statement that actually staged a record now has
its reply held. Caught by the gate, not by review.
THE TRADE, recorded rather than smoothed over. What remains is inherent: a
barrier blocks the owner shard LONGER (more records per fsync) though LESS
OFTEN, so anything queued behind one waits. Three full runs of the same
build gave durable.sN.mixread.p99 of 1043 / 2318 / 4147us and wmix.p99 of
8758 / 20000us — a 2-4x spread with the box near idle. So part A buys ~3x
write throughput at the cost of a longer, noisier tail on the owner shard,
and that is the strongest argument for part B (submit and keep serving).
- durable.sN.*.p99us tolerance widened to 100% WITH the reason in the
code: a 2-4x-variable tail gated at 50% gates the disk, not the engine.
The floor is the real guard and is not slack — mixread's (4172us) came
within 25us of tripping on the worst run. Baseline refreshed; a fresh
full run then passed 106 checks 0 failures
EXIT STATUS MOVED 3 -> 74 (sysexits EX_IOERR). 3 and 4 are already used by
SAMPLES for their own meanings — db-bench's own `verify` exits 3 on a
checksum mismatch, and it is the gate that exercises durability, so a
durability abort exiting 3 would have been indistinguishable from the
mismatch it should help diagnose. The low range belongs to programs.
Docs:
- story: progress, the payoff measured two ways, the cost side, criteria
split met/outstanding, and a "part B — its premise changed" section:
it was justified by "close the 66x gap", but that gap is two problems
and only the concurrent one was a batching problem
- board: standup entry in the six-question shape; both databasev2 4 rows
rewritten. They had said "close the 66x gap" — recorded as MIS-STATED
rather than quietly renumbered
- 00-wob-format.md and 04-db-binding.md: the normative failure contract
("a failed WAL commit traps WO_T_IO after un-applying the row") was
false; corrected, along with the tick-scoped group commit that never
happened
- database/src/CODE-LOGIC.md: where the barrier runs and why there, why
replies are held, why the inline path is asymmetric, the one failure
rule, and how to measure it
- db-bench README: the wmix mode, the env knobs, and the tmpfs warning
Battery: wovm-test 36 suites 0 fail, woc-test, oop-e2e 119/0,
db-bench 106/0, linkcheck clean.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
databasev2 4 part A, task 4. Scope extended with developer approval: the
plan authorised touching the sample only for observability, but no
existing leg has enough concurrent durable writes to exercise group
commit at all, so the payoff was unevaluable either way.
The finding that forced it:
- `mix` writes on one op in ten with C=4 (all_mode calls mix_mode(n/10,
4); Mixer writes on i % 10 == 9), so the quick run performs 20 writes
total. Measured mean batch 1.01 over 3112 barriers, peak 3
- that is a property of the WORKLOAD, not the mechanism: peak 3 of a
possible 4 shows batches form whenever writes actually coincide
- `wmix N C` added: every op a durable write, C at once. Updates rather
than inserts, so it is comparable to mixwrite and the row count stays
flat. Histogram kind 2 — a replayed store still holds the seeding run's
kind-0/1 Hist rows and merging those would report someone else's
latencies
- WO_WAL_STATS=1 prints one line at exit: batches, records, peak_batch,
peak_staged. Opt-in, because it would otherwise pollute every durable
program's output. Counters live in wo_wal; no builtin, the numbers are
diagnostic and not part of the language
Measured, and it scales with concurrency exactly as designed:
- C = 4 / 16 / 64 -> mean batch 1.13 / 1.76 / 5.35, peak 3 / 10 / 39
- the gate's own legs: durable.s1 5412 records over 5412 barriers (mean
1.0, peak 1 — the inline path, one barrier per statement BY DESIGN),
durable.sN 7757 over 2296 (mean 3.38, peak 28) at 2x the throughput
- peak staged 1372 B settles the no-cap decision with a number: the batch
is tiny, so the upstream mailbox bound is sufficient
- mean_batch/peak_batch are higher-is-better (the default detector would
have called bigger batches worse)
- only the batch SHAPE metrics are waived to 100%; wmix throughput and
latency keep real tolerances (15% s1, 50% sN) — a blanket waiver would
have left the entire new leg ungated
- the live assertion `mean > 1.0` on the sN leg is what catches inertness
- gate bites: sN wmix ops_sec -60% -> FAIL on exactly that metric, 1 of 86
Verified: db-bench-quick 89 checks 0 failures; baseline refreshed (86
metrics).
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
databasev2 4 part A, task 2. The core change, and mostly deletion.
- the REQUEST path (wo_db_exec_req) no longer commits after each append.
Applying to RAM and staging stay exactly where they were
- wo_vm_adopt holds each DB reply envelope in a local FIFO instead of
pushing it as the statement finishes. Pushing there would unpark the
requester before its record is durable — the ack contract this
iteration exists to make literally true rather than true by accident
of every batch having one member
- at the end of the drain: ONE wo_wal_commit_fatal for everything staged,
then every held reply. Locals rather than per-shard state: nothing
needs to outlive the batch it describes
- "did this statement stage anything" is asked of the buffer, not guessed
from the opcode, and that count is what the failure diagnostic reports
- the drain commits unconditionally when anything is staged, because the
inline path relies on finding the buffer empty (task 3 documents that)
- staging failure on the request path is now FATAL via wo_wal_stage_fatal:
the row is already in RAM and of the three verbs only insert could undo
itself, so continuing means RAM ahead of disk. One rule
- wal_die is now shared by both fatal points
Verified — the ack contract is the thing that could break, so it is what
was tested:
- just wovm-test: 36 suites (18 x both dispatch flavors) 0 fail, cli_smoke OK
- just db-bench-quick: 85 checks, 0 failures. The legs that matter:
crash.sN.0 — 612 acked rows all present after kill -9, which is the
BATCHING path (multi-shard requests, held replies, one barrier);
crash.s1.0 — 800 acked rows; restart.s1 and restart.sN replay byte-true
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
databasev2 4 part A, task 1.
- wo_wal gains `path`: the abort diagnostic is worthless without naming
the file it could not write. strdup'd in open, freed in close; NULL is
tolerated so the message degrades rather than crashes
- wo_wal_commit now reports WHICH half failed — WO_WAL_ERR_WRITE for
pwrite, WO_WAL_ERR_SYNC for fdatasync. A short write and a device
refusing the flush are different operational problems and the operator
needs the right one named
- wo_wal_commit_fatal(w, nrec): commits, or prints one diagnostic naming
the operation, path, errno and record count, then exits
WO_EXIT_DURABILITY (3 — 1 is a trap, 2 is a refusal, so this takes a
third of its own)
- retrying is not offered, deliberately: on Linux a failed fsync may have
already discarded the dirty pages, so a second call can report success
having written nothing. Replay is the recovery that works
- test_wal: a failed commit is DETECTED, reports the write error
specifically, keeps the batch staged (a failed commit consumes
nothing), and the WAL knows its own path. 165 pass (was 156)
DISCLOSED GAP: the exit path itself is not exercised. Forcing a real
fdatasync failure needs a full or read-only filesystem, which the gate
cannot arrange without mount privileges. No fault-injection switch was
added — shipping a binary that can be told to kill itself is the worse
trade, and the spec rejected it.
Verified: just wovm-test — 36 suites (18 x both dispatch flavors) 0 fail,
cli_smoke OK.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Iteration 24 closes, absorbing 31 and 34. No code in this commit.
- stories 24, 31, 34 -> `status: done`, each with a landing banner. 24's
records the gate numbers and BOTH disclosed deviations: monitor takes
three arguments (the caller may be `main`, which has no mailbox) and a
v1 `call` reply is a typed scalar (which is what let the agreement be
checked at compile time, WO-E226). 31's notes it landed INSIDE 24 and
that a fifth mechanism it never anticipated came out of proving the
gate — the drain guarantee (40). 34's names the gap it did NOT close:
still no RNG, so CSRF/sessions stay blocked
- board: in-progress row cleared, marker doc deleted (convention), the
standup entry in the six-question shape, chain note — next link is
databasev2 4 (io_uring group-commit, chain 5)
- graph: PUBSUB2 (pub/sub + WebSockets, "rejected until here") -> done
- porch ledger: a WebSocket/pub-sub row added; the cancellation row now
says what it actually waits on rather than repeating "the arc"; the
README's "no WebSockets/SSE" limitation was stale — WebSockets are
supported, SSE and chunked encoding are not
- CODE-LOGIC: runtime/src gains the actor-lifecycle section (call, death,
the cap counter's sender/home-thread split, the monitor walk, the timer
list), the drain guarantee, and the digest section; docs/examples/chat
gains its own — actor topology, WHY two actors per connection, fd
ownership, and the shutdown choreography
Battery after the doc edits: wovm-test 36 suites 0 fail, woc-test exit 0,
oop-e2e 119/0, chat 11/0, web-app 46/0, linkcheck clean.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
A message sent before the stop flag is observed must be delivered and run
before the engine stops. One rule; a spin count could never express it.
- root cause in `shard_main` (runtime/src/vm.c): NEXT_RUNNABLE() already
stated the contract — "a WORKER on stop keeps DRAINING ... so queued
shutdown messages (close frames!) still run" — but the IDLE branch
contradicted it, calling fib_reap_all and breaking on WO_IO_STOP,
abandoning its inbox for wo_engine_stop() to free wholesale
- an actor between messages is exactly that idle case, which is why a WARM
soak server hid it: warm shards held live fibers and took the right path
- fix: while the primary's drain window is open, an idle worker adopts its
inbox and runs what arrives; sched_yield on an empty poll so a drain
cannot burn a core per shard and starve the actors it exists to let run
- unreachable at WO_SHARDS=1: wo_engine_stop returns early at nshards <= 1
Measured:
- fresh-server SIGTERM drain: 5 of 16 failing before, 20 of 20 clean after
- `just chat` at the FULL 1000-client soak: 11 checks, 0 failures, both
WO_IO backends, ASan clean with zero leaks
- the fd leg settled at scale too: 1000 connections left the count at 44,
unchanged after 20 more — lazy per-shard init, not a leak
- runtime battery 36 suites (18 x both dispatch flavors) 0 fail;
compiler 556 checks 0 fail
- story: docs/stories/language-runtime-database/40-shutdown-drain-guarantee.md
(chain 3 with 31, status done), board row, slice marker updated
- outstanding and named: a pin below the gate needs new multithreaded test
infrastructure — nothing in runtime/test/ drives wo_engine_start/stop and
no corpus fixture can trigger a stop
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Task 5b of docs/superpowers/plans/2026-08-26-table-residency.md, whose Task 5
is now split 5a-5d (plan updated in this commit).
- the offset twin of wo_row_read (table.c:721): same out-gate contract —
every value handed back is a FRESH VM allocation — but resolved from a file
position instead of the id hash
- fits entirely in wal.c because everything it needs was already public:
scan_record and dec_val are local, and wo_val_decode_vm / wo_db_val_free
are exported at table.h:183-187. Two decode stages, since the record and
the VM speak different dialects: dec_val -> engine slots -> VM copies, with
the engine slots freed as scratch on every path
- ZERO storage change. Nothing calls it yet; that is the point of separating
it from 5c, so the read path can be proven before the slabs are touched
- refuses rather than guessing, each case distinguishable: no intact record
at the offset, a malformed header, a decode failure, trailing bytes, and a
REMOVE tombstone. That last one matters most — handing a tombstone back as
a row would read a deleted row as live
Tested by deep field comparison, not by "it parsed": 24 rows with a nil Text
every third row, each read back BY OFFSET and compared field by field,
including the string bytes. Plus all three refusal paths — tombstone, a
mid-record offset (the silent-wrong-row failure this guards), and past the
intact prefix.
The free-on-every-path claim is VERIFIED, not assumed: removing the free
produced 3 LeakSanitizer reports; restoring it returns to 0. Worth doing
because "ASan is clean" only means something if the harness would have
complained.
PLAN SPLIT: Task 5's storage half was written as if it were plumbing. Measured
instead: wo_row_ptr returns a db_row* into a slab with 11 call sites, table.c
has 37 slab references, db.c:105-181 scans slabs directly, enc_val serialises
FROM the slab, and no operation exists that drops a payload while keeping
index entries. Note this is the OPPOSITE half from the earlier retraction —
the record FORMAT needed nothing, the record STORAGE genuinely is deep. 5c
(id->offset map + drop-payload-keep-index) and 5d (rewiring the call sites,
scans, @unique/FK across the boundary) get their own write-ups.
Gates: test_wal 3654/0 (was 3428), all 18 runtime suites 0 fail under
ASan+UBSan, oop-e2e 119/0, residency 8/0, employee 8/0, db-actor 8/0.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Task 5a of docs/superpowers/plans/2026-08-26-table-residency.md. The read
path itself is NOT in this commit; see the note below.
- the offset problem is far smaller than the spec feared. `w->off` is the
durable tail and `w->len` the staged bytes, and wo_wal_commit pwrites the
whole batch AT off before advancing it — so a record staged now lands at
exactly off+len, knowable at append time with no deferral to flush
- shipped as an inline accessor rather than new out-params on the three
append functions, so the 156 existing WAL checks keep their signatures
- correct across both awkward cases, and both are now unit-pinned:
a failed commit leaves off unadvanced so the record still lands where it
was promised, and wo_wal_open positions off at the end of the INTACT
prefix so offsets are always relative to validated data
- test_offset_capture asserts the recovered ID per record, not merely that a
record parses — a wrong offset reads a NEIGHBOURING record, which passes
its own CRC and returns the wrong row silently. 400 records across
repeated buffer growth (stage() doubles from 4096) and uneven commit
batches, so offsets are exercised mid-buffer and right after a flush
The failed-commit test caught MY OWN misunderstanding: I asserted
next_offset was unchanged after a failed commit. It is not, and should not
be — the record is still staged, so next_offset correctly points PAST it.
The invariant that matters is that the durable tail did not move, which is
what it now asserts.
Gates: test_wal 3428/0 (was 3426), all 18 runtime suites 0 fail under
ASan+UBSan, cli_smoke OK, oop-e2e 119/0, residency 8/0.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Task 4 of docs/superpowers/plans/2026-08-26-table-residency.md — the first
behavioural change in the iteration.
- db.c: one predicate, `table_is_durable`, gating the three EXISTING mutation
sites. Kept as a function rather than an inlined condition so
database/src/CODE-LOGIC.md's "nothing else may mutate storage" claim keeps
holding — the choke points stayed three
- the ack contract is untouched for durable tables: RAM applied, record
staged, one commit before the ack, and a failed commit still removes the row
- replay: a log holding records for a class the image now declares volatile is
a real migration case, not corruption. apply_record returns -2 (distinct
from -1), wo_wal_replay_ex reports the class id, and main.c names it and
exits 2. `wo_wal_replay` stays as the NULL wrapper, so all 156 WAL unit
checks are untouched
- measured, not asserted: 50 inserts wrote 1500 WAL bytes into a durable
table and ZERO into a volatile one. The file's SIZE proves nothing (it is
fallocate'd to 1 MiB up front), so the gate measures the non-zero prefix
BUG I INTRODUCED AND CAUGHT: the mismatch message first printed the class name
with %s, but wo_str.data is `char data[]` with NO NUL terminator (obj.h) — a
buffer over-read. Now %.*s with the explicit length, and re-verified under
ASan.
New gate `just residency` (8 checks), because everything above was otherwise
a one-off manual measurement: restart behaviour, the zero-byte write path, the
mismatch refusal (exit 2, names the class, NOT reported as corruption), and
both compile-time refusals. Its own first run failed two checks for a bug in
the script rather than the feature — `woc | grep` under `set -o pipefail`
returns woc's exit 1 even when grep matches, since woc exits 1 whenever it
reports diagnostics. Captures first now, with the reason noted inline.
Also new: corpus run/table-volatile-inprocess pins that a volatile table is a
FULL table in-process — same @unique enforcement, same index probe, same query
surface. Only survival differs, and that is unobservable from inside one
process.
Gates: woc-test 557/0, 18 runtime suites 0 fail, cli_smoke OK, oop-e2e 119/0
(was 118), residency 8/0, employee 8/0, db-actor 8/0, site 21/0, ASan clean on
the new replay path.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Task 3 of docs/superpowers/plans/2026-08-26-table-residency.md.
- NO LAYOUT CHANGE. The plan said to add descriptor fields; the descriptor
already had a `flags` u32 with only bit0 used, so both properties ride
spare bits (WO_CLASSF_VOLATILE 0x02, WO_CLASSF_RESIDENT_KEYS 0x04). A v7
class record is byte-identical in shape to a v6 one, which is a much
smaller and safer change than the plan assumed
- both spelled as the NON-default, so a zero flags word means exactly what
every pre-v7 image meant: durable, every row resident. A non-@table class
has both clear by construction
- the loader refuses the meaningless pair (bit1+bit2) independently of woc,
on the standing principle that what the loader accepts the interpreter
trusts. Verified by FORGING the flags word in an otherwise valid image,
since woc will not emit one: flags=6 gives "durable:false with
resident:keys", flags=8 still gives "unknown flags"
- WOB_VERSION 6 -> 7. Kept because an OLDER runtime reading a v7 image would
otherwise treat a volatile table as durable and quietly disagree with its
own source. loader.c's check is exact-match, so a v6 image is refused
rather than read with the bits clear — verified by patching a v7 header
back down to 6
GAP FOUND AND CLOSED: woc ACCEPTED `durable: false, resident: keys`. Task 1's
steps covered duplicates and bad values but never the combination, and the
plan had only put that refusal in the loader. The spec wants both, so the
compiler now refuses it too (WO-E102, checked after the argument list is
complete since it is a property of the pair). A compile error is the one a
developer can act on.
VERSION DRIFT: the constant lives in FOUR places, not one. wob.h,
emit.ml:157, disasm.ml:186, and compiler/test/runner.ml:2405 — the last is a
deliberately independent reimplementation of the loader battery, and it
caught the drift as 14 failures rather than silently passing. Its flags mask
and the combination refusal are now in sync too, which is the point of it
being independent rather than shared.
Gates: woc-test 557/0, 18 runtime suites 0 fail, 18 ISO-flavour suites 0
fail, cli_smoke OK, oop-e2e 118/0, employee 8/0, db-actor 8/0, site 21/0.
Zero goldens moved (git diff over golden/ empty).
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
- README: shipped concurrency/HTTP/WebSockets sat in the roadmap as "not yet
available"; "no package manager" contradicted [deps]; the deps example
would not have compiled (the key IS the module name)
- runtime/README: leads with wovm, wo-rt.c demoted to a historical section;
dropped 2 nonexistent recipes, crates/rt, @gc refcounting, 13 suites -> 18
- employee + log-watcher READMEs claimed "does not compile"; both are gates
- error catalog: +10 emitted codes incl WO-E250, the only diagnostic the
shipped query surface raises; recorded why the sweep rotted
- language-surface: group-by parses, then the typechecker refuses it
- 00-code-review + 00-link-audit re-run; history kept, not rewritten
- 48 dead Rust-era exploration links de-linked rather than re-pointed (their
prose names the retired plan by number); successor map -> discarded.md
- 08-project-structure: compiler/plan/ never existed; corpus has 9 dirs, 5 empty
- releasing.md: dropped a --draft step the workflow never had
- new docs/00-doc-audit.md: findings + disposition, incl one row where the
audit was wrong and the doc it accused was right
- status folders removed: 34 stories flat, status only in frontmatter; 252
links recomputed from resolved paths; board/board-views/structure retaught
- story 24 -> in-progress, since frontmatter is now the only truth
- new iteration 38: fs mutation verbs + net.connect, the two capability
families no iteration owned
- new iteration 39: gofiber/fiber v3.5.0 parity study. The ledger called
CSRF/sessions unblocked by iteration 34's HMAC, but the runtime has no
source of randomness at all
- linkcheck skips .dev/.superpowers: 0 broken paths, 0 bad anchors
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
- docs/examples/chat: registry (call consumer) / room / reader+writer
actor pair per connection over ws_accept + wsframe; presence,
broadcast, cross-room isolation, mailbox-full = drop-from-room;
reader tail sends hardened (a full writer no longer orphans the fd)
- RUNTIME SEMANTICS CHANGE (the drain): SIGTERM no longer kills parked
fibers from outside — the plane WAKES them and each wait RESOLVES
(deadline'd waits answer their timeout result, sleeps return early,
plain waits answer WO_SYS_STOPPED and unwind THAT fiber alone; main's
STOPPED still ends the program). Workers keep adopting their inboxes
after stop until eng_shutdown. This is what lets a program drain:
chat's close frames now reach clients (byte-verified 0x88), then
main returns and the reap runs
- also: SIGPIPE ignored process-wide (EPIPE trap instead of death);
two-phase engine teardown (real drops while arenas+routing live,
settle passes for routed frees) — fixes the registry-map leak and
the drain UAF ASan found
- gate scripts/chat-accept.sh + just chat: handshake independently
verified, functional matrix on BOTH backends, 1k-hot-room soak
(1000/1000 in ~35ms), drain close-frames, SIGTERM exit 0, ASan leg
clean. OPEN: soak-fds check (18 fds settle slower than the window)
+ full battery after the semantics change — NOT yet run
- committed for manual testing at the user's request
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
- monitor(watched, observer, msg): registration lives on the watched
actor's home thread (kind-7 envelope cross-shard); actor_die walks
the list; already-dead fires NOW; the notice msg moves; a full
observer's notice drops with a stderr line (no fiber to trap)
- time.after(ms, addr, msg): per-shard timer list riding the deadline
machinery (uring tick min + epoll timeout both include timers;
fired from the same sweep); ms <= 0 delivers now; NO cancel — the
generation-counter idiom is pinned by run/timer-generation
- runtime_notify: one runtime-sourced delivery path (notices, timers) —
reserve-or-drop, cross-shard via kind-0 envelopes
- compiler: monitor typed as a bespoke free fn (notice typed against
the OBSERVER's mailbox — the three-argument deviation, disclosed);
time.after as a stdlib row whose msg arg is EXEMPT from the module-
call fresh-arg drop (it moves — the double-own bug the timer fixture
caught); owner move slots for both
- corpus: run/monitor-death (trap-death + already-dead notices),
run/timer-delivery (armed + immediate), run/timer-generation
- teardown drops undelivered notices and unfired timers; battery 13/13
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
- parse.wo: ANY Transfer-Encoding header is 400-and-close (RFC 9112
§6.1) — silently treating chunked as body-less was the smuggling
door the dup-CL fix left open
- net.listen/listen_unix backlog 64 -> 1024: the soak's connect bursts
overflowed the kernel accept queue and BLACK-HOLED clients (three-way
handshake done, server never sees the conn — 35-70 stuck per run,
fully reproduced then gone at 1024; kernel clamps via somaxconn)
- web-app gate grows to 46 checks: TE-reject; the 1k soak — 500 real
conns all served + 500 idle conns all evicted, server fds home
(45 -> 45), RSS 24MB, healthy after
- battery green (site restart + fibers-TSan legs flaked under parallel
battery load, both clean serially — the standing flake pair)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
- runtime ids 91-95: net.read_dl/accept_dl/write_dl (per-call deadline,
nil/false = the EXPECTED timeout; ms<=0 = old behavior bit for bit),
net.listen_unix (unlink-before-bind, O_NONBLOCK on the listener —
probe-found: accept4's flag covers accepted sockets only), net.peer
- plane: one-op-per-park stays law — deadlines ride one per-shard
TIMEOUT tick (sentinel user_data) + post-CQE expiry sweep +
POLL_REMOVE tombstone; epoll's deadline scan grew the fd-park case;
fibers POOL instead of freeing mid-run (stale-CQE UAF); plain parks
zero park_deadline (no stale sleep deadlines)
- probe: all five seams verified on BOTH WO_IO backends (timeout
timing exact, peer round-trip, unix rebind)
- framework: parse_request grows first_ms/read_ms; serve_conn — the
keep-alive loop with deadlines where parked idle conns are LEGAL
(close-when-idle RETIRED); App.handle_conn exposes it; plain serve()
unchanged for simple apps
- web-app: app-owned accept_dl loop + ConnWorker actor per connection
(each builds its own App; cross-shard placement rides the DB actor);
WA_IDLE_MS knob; gate grows to 41 checks — two slow requests served
in PARALLEL, stalled client evicted at the idle deadline, slow-loris
torn at the read deadline (400)
- docs: story 35 -> done with banner; SQE/CQE design spec LANDED (was
the review doc); ledger rows (timeouts/unix/keep-alive/peer), graph
(NETSEAM cleared, KEEPAL done), builtin-surface rows, runtime
CODE-LOGIC section, board entry
- battery 13/13 fresh-built
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
- runtime: mailbox slots grow caller metadata (wo_msg), call parks on
WO_PARK_INBOX (the DB-RPC protocol) and the resume consumes a SCALAR
reply; FIBER_DONE ships the receive's return value home (same-shard
unpark or kind-6 envelope); kind-5 carries cross-shard calls
- actor death is real now: a receive trapping uncaught marks the actor
dead, error-unparks the in-flight caller AND every queued caller,
drops queued payloads + state, releases cap slots; send-to-dead
drops silently, call-to-dead traps — a call never hangs. Fixes the
pre-existing leak/dangle in TRAPF's fiber-death path (cur_msg leaked,
a->active dangled, the mailbox rotted)
- compiler: reply typing through actor-M erasure — every receive(M)
program-wide must agree on one return type and it must be a copyable
scalar (v1); WO-E226 names disagreeing classes / void receives /
non-scalar replies; call's message moves exactly like send's (owner)
- corpus: run/call-echo (park + ordered replies), run/call-dead-trap
(mid-call + to-dead, both catchable), compile-fail/call-void-receive,
compile-fail/call-reply-disagree; cross-shard call proof rides the
chat gate next
- battery 12/12 fresh-built (ASan+TSan lanes in fibers/db-actor green)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
- cap 1024 (WO_MAILBOX override at boot): sender-side atomic
reserve/release on every path — same-shard, cross-shard envelope,
OOM rollbacks; full mailbox traps the SENDER catchably; delivery
pop releases; overshoot bounded by in-flight sends (disclosed)
- test_mailbox 12/0: exact cap single-threaded, two racing senders win
exactly cap slots, drain/refill clean
- corpus run/mailbox-full-trap: parked sleeper, send loop catches
"actor mailbox full" after >= 1024 sends
- pre-existing compiler bug found + fixed: a try ARM yielding a Text
PLACE (bare e.msg, try box.field) aliased a register the arm's scope
end freed — ASan use-after-free, SEGV on the next unwind's
double-walk; emit_try now applies copy_place_text to both arm
results; pinned by corpus run/catch-msg-place
- db-bench driver: msgrate keeps iteration 22's unbounded-flood
contract via WO_MAILBOX=MSG_N (the cap is 24's policy, not 22's)
- battery 12/12 fresh-built
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
- brings time.ticks builtin (id 84), bench sample + campaign driver
(scripts/db-bench.py), just db-bench/db-bench-quick recipes, first
baseline recorded, story 22 to done/, postgres study cards
- conflicts resolved: board in-progress table (iteration 36 +
framework rows kept, db-bench row now "22 landed"; dangling order
anchor repointed); story 36 moved back to in-progress/ (dir-rename
inference dragged it to done/ — 36 still awaits the manual pass)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
- engine: wo_idx_probe answers single-column equality from the index
hash buckets (idx_hash_key1 reproduces idx_hash bit for bit; verify
compares exactly as the slab walk did, so results identical);
composite indexes keep the walk; both executors wired (local + DB
actor RPC)
- compiler: probe_key_of_where lowers "var.col == key" on an indexed
column to DB_PROBE; all where guards still run (guard stays the
final arbiter); keys = ident/int-literal only; Float/Bytes excluded
(engine raw-eq narrower than VM float-eq)
- measured: reads 1.3k -> 1.3M ops/s, p50 600us -> 1us (~x850);
query x830; mixread 1.3k -> 89k s1, 21 -> ~1.9k sN
- gate policy moved into the driver (tolerance_for: refresh-proof);
latency floors max(4x,100us); quick mode skips poll-bound mix
floors; both tolerance classes proven to bite
- proof: test_table wo_idx_probe suite (RED first), corpus
query-index-probe 105/0, full battery green, TSan clean, two
campaigns pass the refreshed baseline
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
- worker DB builtins marshal to shard 0: requester-side slot encode
(VM heaps never read cross-shard), owner executes serialized in
adopt, reply unparks via new WO_PARK_INBOX park + envelope 3/4
- engine gains thread-agnostic slot entry points (insert_slots,
update_field_slot, val_encode/clone, wo_db_exec_req); traps and
messages byte-identical to the local path
- main.c: engine + replay boot BEFORE shards spawn; workers assert
rt.db/rt.wal NULL; busy shard adopts inbox once per slice
- latent stage-1 bug fixed: shared io_uring params static raced by
lazy worker init lost park wakes (~1/20 hangs); params per-vm,
short submit now fails loud
- new sample docs/examples/db-actor + just db-actor gate 8/0 (multi
x3, uring/epoll forced, single byte-exact, WAL replay pair);
ASan+TSan 6/6; full battery green
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
- Float full stack: literals (fraction/exponent; `0..10` still a range), f64
opcodes 34-41, @table column, WAL bit-exact replay, json fractions in and
shortest-round-trip out. IEEE-quiet — FDIV never traps where DIV does.
- Bytes: a wo_str with its own class id, so alloc/free/copy are shared but no
Text builtin accepts one; len/at/slice/eq/concat, base64 both ways, json
boundary as base64; TEXT_COPY preserves the kind.
- No implicit Int/Float mixing (WO-E201 in the typechecker, not the emitter,
which picks the opcode from one side and would misread the other).
- One IEEE deviation: float_cmp total order (NaN last, -0.0 == +0.0) for
indexes and order-by, keys canonicalized to match. `?Float` nil is a
reserved quiet NaN — the zero word is +0.0, WO_NIL_SCALAR's bits are -2.0.
- Renderer prefers fixed over exponential in 1e-6..1e21: pure shortest makes
a price of 900.0 read `9e+02`. One renderer for interp/json/float_to_text.
- Fixed en route: lexer double-counted the leading digit; is_scalar_shaped
took Float/Bytes as Int-shaped; Bytes ownership needed a shared heap-scalar
predicate or temps never dropped; order-by bit-compared negatives backwards.
- Iteration 17: `kind = "library"` (absent = program; bad value = WO-E109),
entry-less check mode retiring the `--emit` workaround, Go's `internal/` as
WO-E108 at the consumer's `use`. Driver-only; VM/.wob/GC untouched.
- Framework reorg: internal/{parse,serve}.wo; http/form.wo split out to keep
media_type/form_values public (parse.wo had grown public surface).
- Docs: link audit (97 -> 88 broken, conflict markers resolved, 2 duplicate
stories removed), 00-code-review verified 26/27, iterations re-sequenced.
- Also carries the pre-staged pub(read)/using/#if work from the index.
- Gates: corpus 103/0, test_wal 156/0, web-app 26/0, oop-accept ALL MET.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
- spawn placement: round-robin across shards (same-shard when the
engine is absent/single); the actor's mailbox and delivery belong to
its HOME thread — a spawn to another shard travels as an ADOPT
envelope, a send as a SEND envelope (mutex-guarded inbox + eventfd
wake; the spec's lock-free rings stay a disclosed deviation until
9e measures the mutex)
- workers: first envelope triggers lazy full-vm init UNDER the inbox
mutex (TSan caught the memset racing a concurrent push, twice — the
second was inbox_push reading wake_efd outside the lock; both fixed,
gate x8 + battery clean); serve loop = adopt -> run to drained ->
wait on the plane (the wake eventfd is watched by io_uring POLL_ADD
oneshot / epoll level-triggered on BOTH backends)
- ownership across heaps: every allocation stamps rt->shard_id into
the header (the field reserved since iteration 2); a drop on the
wrong shard routes home as a FREE envelope — the owner's arena stays
single-threaded by construction; at teardown routed frees become
no-ops (arenas die wholesale) which is what un-danced the freed-mutex
ASan SEGV the first ordering had
- WO-E222: an actor's state or message type that is (or transitively
contains) an inferred-traced class refuses at the spawn/send — with
round-robin every actor is potentially remote; corpus-pinned
(compile-fail/traced-send, inference-aware: Box contains ?Node)
- determinism narrowed per spec: oop-e2e pins WO_SHARDS=1 (exact
outputs); the fibers gate grows multi-shard SET assertions + a TSan
run (wovm-tsan target; setarch -R fallback for kernel 6.5+ ASLR)
- NEXT_RUNNABLE honors engine shutdown for parked workers (deadlock
hole closed); io_wait's adopt-wake (rc 1) no longer reads as fatal
- battery: oop-e2e 93/0, fibers 10/0 x8 (+WO_IO=epoll), log-watcher
7/0, employee 8/0, web-app 21/0, deps 8/0, runtime tests 16/16
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
- wo_engine + wo_engine_start/stop: one pinned pthread per extra core
(pthread_setaffinity_np); shard 0 is the primary (entry + database);
workers idle on a wake eventfd until fibers arrive (T6) or shutdown
- default = all cores (the arc's brave landing), WO_SHARDS=1..64
overrides; N=1 spawns no threads — byte-identical to stage 1
- worker vms are LAZY: identity + wake fd only until their first fiber
arrives — 20 idle shards must not cost 1.25 GiB of eager arenas
(they did: the web-app gate flaked on exactly that before the fix;
3 consecutive green runs after)
- engine stops (join + destroy) before the primary's teardown
- battery at the 20-core default: oop-e2e 92/0, fibers 8/0,
log-watcher 7/0 (+WO_SHARDS=1 identical), web-app 21/0 x3,
employee 8/0, deps 8/0, runtime tests 16/16 files green
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
- park.c/park.h: one event loop per shard. io_uring PRIMARY (raw
io_uring_setup/io_uring_enter, uapi structs mirrored, 5.4-floor ops:
POLL_ADD for fd readiness, TIMEOUT for sleeps, user_data = the fiber);
epoll+deadline-scan FALLBACK behind the startup probe; WO_IO=
uring|epoll forces either so CI proves both on one kernel
- park protocol: a blocking builtin fills cur->park_* and returns
WO_SYS_PARKED; resume either RE-EXECUTES it (fd readiness: accept/
read/write retry) or continues PAST it (sleep: result preset,
park_done=1 — re-executing would restart the full duration)
- sysio: listener + accepted fds nonblocking (accept4 SOCK_NONBLOCK);
accept/read park on EAGAIN; write parks on EAGAIN with its partial
progress carried across the retry in park_wr_at; sleep parks on a
deadline — with ONE fiber the plane's wait IS the blocking call,
program mode is the degenerate case, not a special one
- scheduler: NEXT_RUNNABLE waits on the plane when the queue empties;
a stop interrupting the wait reaps EVERY fiber (queued and parked)
and returns the clean-stop status; parked fibers are GC roots and
fib_reap_all drains them
- proof: full battery green on the uring path (oop-e2e 92/0,
log-watcher 7/0 incl. the mcp accept/read/write loop, employee 8/0,
web-app 21/0, deps 8/0), WO_IO=epoll battery green (log-watcher 7/0,
web-app 21/0), WO_IO=uring forced green, LW_SOAK=8 10/0 (fd + RSS
flatness holds over parked I/O)
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
- language: `spawn Cls { fields }` expression (ctor semantics — fields
MOVE; result is the address); `actor M` parametric field type
(contextual like multi/map — actor stays a legal identifier); `send`
is a builtin free-fn name, not a keyword (shadowing rule applies)
- typing: M inferred from Cls's receive(msg: M); WO-E221 when receive
is missing, mis-armed, or M is not a class/record/union; send checks
addr is `actor M` and the message IS an M (silent when underivable);
ctor half of spawn delegates to the Ctor arm (completeness, ?T, E207)
- ownership: send's message TRANSFERS (sender's later use = WO-E301,
corpus-pinned); spawn's fields move via the ctor machinery; an
address is Copy
- emit: spawn lowers to ctor + LOADK receive's method index + BUILTIN
68; send is BUILTIN 69 with the message excluded from fresh-arg drops
(the runtime owns it now)
- runtime: wo_actor (moved-in instance, receive idx, growable FIFO
mailbox, one delivery fiber at a time); delivery reuses the fiber
context across messages and re-queues per message (fairness — an
actor never monopolizes); the runtime drops each message after its
receive returns; actor state/queued/in-flight messages are GC roots;
teardown drops everything (main-return reap included); loader knows
the two arities
- corpus: run/actor-echo (typed spawn/send, one-at-a-time delivery
interleaved with main by budget — output exact, ASan-clean),
compile-fail/spawn-no-receive (WO-E221), send-after-move (WO-E301)
- battery green: oop-e2e 92/0, woc-test, wovm-test, log-watcher 7/0,
employee 8/0, web-app 21/0, deps-accept 8/0
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
- fiber states (RUNNABLE/PARKED/DONE), intrusive FIFO run queue,
wo_vm_spawn_fiber (calloc'd context, frame 0 set up like wo_vm_call)
- reduction budget: WO_REDUCTIONS (default 4000), checked at loop
BACK-EDGES AFTER the jump lands so the saved pc is the loop head —
a pre-instruction save at budget 1 re-executes the jump into the
same decrement and livelocks (found by reasoning, pinned by the
budget-1 test; deviation from the spec's three-site wording,
recorded in the yield macro's comment)
- FIBER_DONE: main returning ends the program and reaps every
remaining fiber through vm_unwind (drop maps run); a spawned fiber
ending frees silently; its return value is discarded by contract
- TRAPF: an uncaught trap in a spawned fiber kills that fiber ALONE
(stderr report, program lives); in main it stays the program's death
- WO_SYS_STOPPED reaps all fibers wherever it lands (main unlinked
from the queue and unwound if a spawned fiber caught the stop)
- vm_gc_roots walks the live fiber plus every queued one
- test_fiber (45 checks, ASan): EXACT round-robin interleave at budget
1 across three fibers pushing tags into one shared multi;
main-return reaps a spinning fiber holding an owned Big (ASan proves
the free); a DIV0 fiber dies alone, main answers 0
- full battery green: wovm-test, oop-e2e 89/0, woc-test, log-watcher
7/0, employee 8/0, web-app 21/0, deps-accept 8/0 (scheduler dormant
= one branch per back-edge)
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>