docs: stage 3 refined with guarantee contract; story 32 born

- five-property map in marker doc: atomicity/durability/recovery
  proven or held (18); concurrency control = stage 3's property;
  space reclamation RAM done (slot reuse), disk = new story 32
- story 08: three stage-3 criteria (one commit per write RPC +
  ack-after-owner-fsync, workers WAL-free + replay-before-serve,
  no torn reads under TSan corpus); arc plan stage 3 carries them
- refine/32-wal-checkpoint.md: snapshot + truncate, bounded replay,
  crash-during-checkpoint safe; four forks; after 23
- chain now stage 3 -> 22 -> 31 -> 24 -> 23 -> 32 in all 10 docs;
  boards + seq bumps synced

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
This commit is contained in:
shoney.arickathil 2026-08-21 11:26:45 +02:00
parent 8232a3fcd1
commit 788e026739
11 changed files with 171 additions and 20 deletions

View file

@ -1,7 +1,7 @@
# In progress — the 8+11 arc, stage 3: the transparent DB actor
> **Status: 🔄 in progress** (started 2026-08-21) — first slice of the
> concurrency chain **stage 3 → 22 → 31 → 24 → 23**. Board:
> concurrency chain **stage 3 → 22 → 31 → 24 → 23 → 32**. Board:
> [../00-status.md](../stories/00-status.md).
>
> This folder holds ONE marker doc: the slice being executed right now,
@ -33,8 +33,24 @@ single- AND multi-shard honestly.
still blocks instead of parking, land it here or re-scope 11's
criterion explicitly at close)
## Guarantee contract (refined 2026-08-21)
The five store properties, mapped honestly — what this slice owes vs
what is already proven or deferred:
| property | state |
| --- | --- |
| Atomicity | per-statement ✅ (WAL record replays whole-or-not-at-all); multi-statement = `transaction { }`, iteration 18, ⏸ held. **Stage-3 obligation:** a worker write RPC is exactly ONE owner-shard commit — a crash between send and commit leaves no ack and no partial state. |
| Durability | ✅ fsync-per-commit, ack-after-durable. **Stage-3 obligation:** the ack crosses shards only AFTER the owner's fsync completes. Power-loss rides fdatasync semantics; 22's kill battery is the scripted proof. |
| Crash recovery | ✅ boot replay, torn-tail drop, index rebuild; 22 scripts the restart proof. **Stage-3 obligation:** workers never open the WAL or data dir; replay completes on the primary before any worker serves. |
| Concurrency control | **THIS SLICE.** The DB actor serializes every statement; replies are materialized copies — no torn read can exist by construction. Each statement sees the serialized moment its envelope executes; cross-statement snapshots arrive with 18. |
| Space reclamation | RAM ✅ — deleted rows free their slot ([`04-db-binding.md`](../plan/oop-vm/04-db-binding.md): "Ids are never reused; slots are"). Disk ✖ — the WAL grows unbounded, no checkpoint exists → [story 32](../stories/language-runtime-database/refine/32-wal-checkpoint.md), end of chain. |
## Definition of done
- The guarantee obligations above hold: one commit per write RPC with
ack-after-owner-fsync, workers WAL-free with replay-before-serve, and
no torn reads under concurrent multi-shard load.
- Plan Tasks 7–8 checked off; full battery green (`just woc-test`,
`just oop-e2e`, `just deps-accept`, `just web-app`, `just log-watcher`,
`just employee`, `just fibers`) — multi-shard DB answers byte-identical

View file

@ -30,7 +30,7 @@ Statuses: ✅ **done** · 🔄 **in progress** · ⬜ **pending** · ⏸ **hold*
## ▶ NEXT PLAN
**The concurrency + fiber chain — stage 3 → 22 → 31 → 24 → 23**
**The concurrency + fiber chain — stage 3 → 22 → 31 → 24 → 23 → 32**
(directive 2026-08-21). Active slice: the 8+11 arc's **stage 3, the
transparent DB actor** — marker doc
[`in-progress/2026-08-21-arc-stage-3.md`](../in-progress/2026-08-21-arc-stage-3.md),
@ -66,7 +66,9 @@ stage 3 unblocks 22's multi-shard campaign.
**Next steps:** stage 3 → 22 (baselines + mutex-inbox number) → 31
(lifecycle) → 24 (chat, the arc's acceptance) → 23 (io_uring
group-commit). Held tail resumes on its own precedence notes.
group-commit) → 32 (WAL checkpoint — disk reclamation, added
2026-08-21 by the stage-3 guarantee refinement). Held tail resumes on
its own precedence notes.
**`.dev/reference` used:** `linux` (the "single event loop" card behind
the io_uring-first directive, arc T4). Record the ones each iteration
@ -227,7 +229,8 @@ that sequences its tasks. Read one, approve, then the next starts.
| 22 | [Durability, throughput, scale](language-runtime-database/refine/22-durability-throughput-scale.md) | ⬜ needs a spec first — second in chain, after arc stage 3 |
| 31 | [Actor lifecycle](language-runtime-database/refine/31-actor-lifecycle.md) | ⬜ needs a spec first — third in chain (story written 2026-08-21) |
| 24 | [chat: WebSocket workload](language-runtime-database/refine/24-chat-websocket-workload.md) | ⬜ fourth in chain — the arc's acceptance; after 31 |
| 23 | [io_uring group-commit](language-runtime-database/refine/23-io-uring-commit.md) | ⬜ last in chain, after stage 3 + 22 |
| 23 | [io_uring group-commit](language-runtime-database/refine/23-io-uring-commit.md) | ⬜ fifth in chain, after stage 3 + 22 |
| 32 | [WAL checkpoint](language-runtime-database/refine/32-wal-checkpoint.md) | ⬜ last in chain, after 23 — disk reclamation + bounded replay (story written 2026-08-21) |
| 20 | [Cross-program tables](language-runtime-database/hold/20-cross-program-tables.md) | ⏸ hold (2026-08-21); channel done (branch ipc-attach keeps its manifest) |
| 21 | [Keypair attach auth](language-runtime-database/hold/21-keypair-attach-auth.md) | ⏸ hold (2026-08-21); crypto+handshake done (branch keypair-auth keeps its manifest) |
| 25 | [HTTP service layer](../superpowers/plans/2026-08-01-http-service-layer.md) | ⏸ hold (2026-08-21) — story file removed; the plan doc remains |
@ -447,6 +450,11 @@ precedence notes for resumption.
2026-08-20 — Bytes carries the frames).
5. **23** — io_uring group-commit; the WAL's WRITE+FSYNC chains ride the
arc's per-shard ring (T4); after 22's baseline — the payoff, measured.
6. **32** — WAL checkpoint
([story](language-runtime-database/refine/32-wal-checkpoint.md),
written 2026-08-21): the WAL is append-only forever — snapshot +
truncate reclaims disk and bounds replay; after 23 (composes with
group-commit), policy set by 22's aged-store numbers.
**30** — observability, CI, fuzz: named 2026-08-20, still row-only (no
story file); slots in when scheduled — nothing in the chain depends on it.

View file

@ -73,7 +73,7 @@ adding surface, and stop stacking features on unmeasured ground.
RE-SEQUENCED 2026-08-21 (third pass — the concurrency chain). Everything
still pending IS the runtime-concurrency chain; order:
**stage 3 → 22 → 31 → 24 → 23**. Changes from the second pass:
**stage 3 → 22 → 31 → 24 → 23 → 32**. Changes from the second pass:
- **The arc's stage 3 moves ahead of 22** — correctness before
measurement: a multi-shard program touching the database traps
@ -106,14 +106,15 @@ still pending IS the runtime-concurrency chain; order:
| 17 | 31 | [Actor lifecycle](refine/31-actor-lifecycle.md) | request/response (today `send` is one-way and callers `sleep` to await), bounded mailboxes with backpressure (today the FIFO just grows), actor death/supervision, and timers beyond `time.sleep`. 24 cannot be written honestly without these. *(story written 2026-08-21)* |
| 18 | 24 | [chat: WebSocket workload](refine/24-chat-websocket-workload.md) | the arc's acceptance: WS upgrade + frames (SHA-1 via crypto fork, Bytes via 19), rooms/broadcast, 1k clients, drain-clean. *(was 19)* |
| 19 | 23 | [io_uring group-commit](refine/23-io-uring-commit.md) | WAL WRITE+FSYNC chains on the arc's per-shard rings; fsync fallback kept (after 22 + the arc). *(was 9f)* |
| 20 | 25 | [HTTP service layer](../../superpowers/plans/2026-08-01-http-service-layer.md) | `service` blocks lower onto the framework (after 9b + 20 by their own precedence notes). **HELD 2026-08-21** — story file removed; the plan doc remains. *(was 10)* |
| 21 | 18 | [framework v2: memory-rich features](hold/18-memory-db-features.md) | spec+plan approved: TTL cache, @table flags, durable job queue, `transaction { }` over the WAL's staged batch. **Demoted from seq 14**: more surface on a framework with one consumer, and the cache still stores `Text` because there are no generics |
| 22 | 27 | [Query grammar corpus](hold/27-query-grammar-corpus.md) | grow the query grammar from real corpora; likely collapses to "confirm `len(query)` + add `exists`"; precedes 28. *(was 9g)* |
| 23 | 26 | [Blue-green deploy](hold/26-blue-green-deploy.md) | two VM slots, in-runtime compile, atomic switch, resident rollback (plan authored after 9 + 25). *(was 12)* |
| 24 | 20 | [Cross-program tables](hold/20-cross-program-tables.md) | attach to a running program's database over local IPC; owner stays the single writer (channel half-built). **Demoted from seq 16**: new distribution surface while there is no TLS, no crypto, and the multi-shard DB still traps. *(was 9c)* |
| 25 | 21 | [Keypair attach auth](hold/21-keypair-attach-auth.md) | program identity is a keypair; mutual challenge–response at attach (crypto half-built; plan folds into 20's). **Demoted with 20** — and it needs crypto primitives that do not exist. *(was 9d)* |
| 26 | 28 | [skillhost host workload](hold/28-skillhost-host-workload.md) | host-shaped driving workload naming runtime gaps — demoted with the framework goal. *(was 14)* |
| 27 | 29 | [Compile-time metaprogramming](hold/29-compile-time-metaprogramming.md) | `@derive(...)` from class-table metadata; held with the parked drain by the 2026-08-08 scope directive. *(was 13)* |
| 20 | 32 | [WAL checkpoint](refine/32-wal-checkpoint.md) | **NEW 2026-08-21** (stage-3 guarantee refinement found the hole) — the WAL is append-only forever: snapshot + truncate reclaims disk and bounds replay time; every durability guarantee byte-identical; crash mid-checkpoint recovers from the previous snapshot + full tail. After 23 (composes with group-commit); RAM slot-reuse already contracted in `04-db-binding.md`. |
| 21 | 25 | [HTTP service layer](../../superpowers/plans/2026-08-01-http-service-layer.md) | `service` blocks lower onto the framework (after 9b + 20 by their own precedence notes). **HELD 2026-08-21** — story file removed; the plan doc remains. *(was 10)* |
| 22 | 18 | [framework v2: memory-rich features](hold/18-memory-db-features.md) | spec+plan approved: TTL cache, @table flags, durable job queue, `transaction { }` over the WAL's staged batch. **Demoted from seq 14**: more surface on a framework with one consumer, and the cache still stores `Text` because there are no generics |
| 23 | 27 | [Query grammar corpus](hold/27-query-grammar-corpus.md) | grow the query grammar from real corpora; likely collapses to "confirm `len(query)` + add `exists`"; precedes 28. *(was 9g)* |
| 24 | 26 | [Blue-green deploy](hold/26-blue-green-deploy.md) | two VM slots, in-runtime compile, atomic switch, resident rollback (plan authored after 9 + 25). *(was 12)* |
| 25 | 20 | [Cross-program tables](hold/20-cross-program-tables.md) | attach to a running program's database over local IPC; owner stays the single writer (channel half-built). **Demoted from seq 16**: new distribution surface while there is no TLS, no crypto, and the multi-shard DB still traps. *(was 9c)* |
| 26 | 21 | [Keypair attach auth](hold/21-keypair-attach-auth.md) | program identity is a keypair; mutual challenge–response at attach (crypto half-built; plan folds into 20's). **Demoted with 20** — and it needs crypto primitives that do not exist. *(was 9d)* |
| 27 | 28 | [skillhost host workload](hold/28-skillhost-host-workload.md) | host-shaped driving workload naming runtime gaps — demoted with the framework goal. *(was 14)* |
| 28 | 29 | [Compile-time metaprogramming](hold/29-compile-time-metaprogramming.md) | `@derive(...)` from class-table metadata; held with the parked drain by the 2026-08-08 scope directive. *(was 13)* |
| ✅ | 17 | [library projects + `internal/`](done/17-library-projects-internal.md) | **LANDED 2026-08-20** — `kind = "library"` + entry-less check mode (retires the `--emit` workaround) and Go's `internal/` rule as WO-E108 at the consumer's `use`; driver-only, VM/GC untouched. `just web-app` 26/0 |

View file

@ -22,7 +22,7 @@
> scope: **stage 3, the transparent DB actor** — a correctness fix, not
> an optimization: worker VMs are zero-initialized, so a DB statement
> off the primary shard traps `WO_T_DB`. Concurrency-chain order:
> **stage 3 → 22 → 31 → 24 → 23**.
> **stage 3 → 22 → 31 → 24 → 23 → 32**.
## Settled decisions (2026-08-20)
@ -47,7 +47,7 @@
4. **Order — SUPERSEDED 2026-08-21.** Originally 22 → the 8+11 arc → 23;
in fact the arc's stages 1+2 landed before 22 ever ran (accepted
deviation — the before/after delta is owed and lands as 22's
multi-shard pass). Current order: **stage 3 → 22 → 31 → 24 → 23**.
multi-shard pass). Current order: **stage 3 → 22 → 31 → 24 → 23 → 32**.
23 still waits for the arc's tick boundary and 22's baseline.
## Goals
@ -90,6 +90,27 @@
- **when** the employee/web-app matrices run multi-shard,
- **then** every answer is byte-identical to the single-shard run and
the WAL's ack-after-durable contract is unchanged.
- What to achieve? (stage-3 refinement, 2026-08-21)
- **Given** a worker shard issuing a write through the DB actor,
- **when** the process is killed between the worker's send and the
owner's commit,
- **then** the write was never acknowledged AND replay shows no
partial state — a write RPC is exactly one owner-shard commit,
and the ack crosses shards only after the owner's fsync.
- What to achieve? (stage-3 refinement, 2026-08-21)
- **Given** a multi-shard boot,
- **when** workers start serving,
- **then** WAL replay has already completed on the primary, and no
worker ever opens the WAL or the data directory (asserted in
debug builds).
- What to achieve? (stage-3 refinement, 2026-08-21)
- **Given** concurrent workers hammering reads and writes at one
table,
- **when** the deterministic multi-shard corpus runs under TSan,
- **then** no torn read exists — every statement sees the serialized
moment its envelope executes on the owner shard (replies are
materialized copies). Full guarantee map:
[the marker doc](../../../in-progress/2026-08-21-arc-stage-3.md).
## Out Of Scope

View file

@ -34,7 +34,7 @@
>
> **RE-SEQUENCED 2026-08-21**: what remains for the chain is the arc's
> stage 3 (transparent DB actor — [iteration 8](08-shard-actor-runtime.md)),
> then measurement. Order: **stage 3 → 22 → 31 → 24 → 23**.
> then measurement. Order: **stage 3 → 22 → 31 → 24 → 23 → 32**.
## Goals

View file

@ -21,7 +21,7 @@
> baseline. New measurement target since stage 2: the mutex-guarded
> inbox + eventfd (the plan's deviation — lock-free rings arrive only if
> this number says the mutex costs). Chain order:
> **stage 3 → 22 → 31 → 24 → 23**.
> **stage 3 → 22 → 31 → 24 → 23 → 32**.
## Goals

View file

@ -23,7 +23,8 @@
> paths on one kernel — AMENDED: the override is the arc-wide
> `WO_IO=uring|epoll` (the arc's T4 owns the probe and the per-shard
> ring; `WO_WAL_MODE` is subsumed). Position — RE-SEQUENCED 2026-08-21:
> LAST in the concurrency chain, **stage 3 → 22 → 31 → 24 → 23**
> FIFTH in the concurrency chain (32, WAL checkpoint, follows it —
> added 2026-08-21), **stage 3 → 22 → 31 → 24 → 23 → 32**
> (supersedes the 2026-08-20 old-id ordering "9e → 8+11 → 9f"); the
> per-shard ring already exists (arc T4 landed 2026-08-20,
> `WO_IO=uring|epoll`) — this iteration adds the WAL's WRITE+FSYNC

View file

@ -8,7 +8,7 @@
> its spec AFTER the arc's — it lands at the arc's end and proves it.
>
> **RE-SEQUENCED 2026-08-21**: fourth in the chain,
> **stage 3 → 22 → 31 → 24 → 23** — chat cannot be written honestly
> **stage 3 → 22 → 31 → 24 → 23 → 32** — chat cannot be written honestly
> before [iteration 31](31-actor-lifecycle.md) (request/response,
> bounded mailboxes, actor death, timers). Iteration 19 LANDED
> 2026-08-20, so Bytes is available for frame parse/serialize.

View file

@ -6,7 +6,7 @@
> **Inserted 2026-08-21** (concurrency-chain re-sequence; the iteration
> was named as "new 31" in the 2026-08-20 code-review re-sequence — this
> is its story file). Third in the chain,
> **stage 3 → 22 → 31 → 24 → 23**: chat
> **stage 3 → 22 → 31 → 24 → 23 → 32**: chat
> ([iteration 24](24-chat-websocket-workload.md)) cannot be written
> honestly without these four mechanisms.

View file

@ -0,0 +1,92 @@
# Iteration 32 — WAL checkpoint: disk space reclamation and bounded replay
> Format: fiberloom `product/story-iteration-template`. Part of
> [Story — one language, one runtime, one database, one binary](../00-story.md).
>
> **Inserted 2026-08-21** (stage-3 guarantee refinement found the hole):
> the WAL is append-only FOREVER — no checkpoint, no truncation exists
> in the engine or anywhere on the roadmap. Disk grows without bound and
> replay time grows with history, so restart cost rises with every write
> the program ever made. RAM reclamation already exists (deleted rows
> free their slot — [`04-db-binding.md`](../../../plan/oop-vm/04-db-binding.md):
> "Ids are never reused; slots are"); this iteration is the DISK half.
> LAST in the concurrency chain:
> **stage 3 → 22 → 31 → 24 → 23 → 32** — it wants 22's measured
> replay/restart numbers to justify its policy and must compose with
> 23's group-commit write path.
## Goals
- **Disk space is reclaimed.** A checkpoint writes the live store as a
snapshot and truncates the WAL behind it; deleted rows and
overwritten versions stop occupying disk forever.
- **Replay is bounded.** Startup replays snapshot + WAL tail, not the
program's whole write history — restart time becomes a function of
store size, not store age.
- **Every existing guarantee holds byte-for-byte.** Ack-after-durable,
replay-whole-or-not-at-all, torn-tail drop, ids never reused — a
checkpoint changes where bytes live, never what an ack means. A crash
DURING checkpoint recovers from the previous snapshot + full tail:
the old WAL is not truncated until the new snapshot is durable.
## Acceptance Criteria (draft — the spec refines)
- **Given** a store with N rows after many writes and deletes, **when**
a checkpoint completes, **then** disk usage reflects the live rows
(plus the WAL tail), and a restart replays snapshot + tail to the
byte-identical store.
- **Given** kill -9 at ANY instant during a checkpoint, **when** the
process restarts, **then** recovery produces the same consistent
store as if the checkpoint had never started — no acknowledged write
lost, no partial snapshot ever read.
- **Given** the iteration-22 restart benchmark re-run after checkpoint
lands, **when** replay time is measured on an aged store, **then**
the bounded-replay improvement is recorded as a before/after delta.
- **Given** writes arriving while a checkpoint runs (the DB actor
serializes statements; the checkpoint must not stall them beyond the
stated budget), **when** the mixed load completes, **then** every ack
held its durability contract and the tail contains exactly the
post-snapshot writes.
## Out Of Scope
- MVCC / multi-version reads — the store is update-in-place RAM; "old
versions" exist only as WAL history, which is exactly what truncation
reclaims.
- Incremental/streaming backup, point-in-time recovery — a snapshot is
a recovery artifact here, not a backup product.
- Cross-shard checkpoint coordination — the WAL is owner-shard-only
(stage 3's rule); one shard, one checkpoint.
- Compression, dedup, tiering — measure first (22), add only what a
number justifies.
## Info
Forks the spec must settle:
1. **Snapshot format** — a row-image dump of the live store (simple,
O(live rows)) vs a rewritten-compacted WAL (reuses replay machinery,
O(live rows) too but stays in one format). Leaning: row-image dump
in the WAL's existing record grammar, so replay needs no second
decoder.
2. **Trigger policy** — size threshold (WAL bytes vs snapshot bytes
ratio), boot-time compaction, explicit call, or some mix. Leaning:
ratio threshold checked at commit, plus manual trigger for tests;
decided against 22's numbers.
3. **Write availability during checkpoint** — stop-the-world dump
(simplest; the DB actor just runs one long "statement") vs
fork-and-dump vs incremental copy. Leaning: measure the
stop-the-world pause on the 1M-row store first (22); complexity only
if the pause breaks a stated budget.
4. **Composition with 23** — the snapshot's durability barrier rides
the same per-shard ring (WRITE+FSYNC chain, then the truncate);
ordering vs in-flight group commits must be stated normatively in
[`04-db-binding.md`](../../../plan/oop-vm/04-db-binding.md)'s WAL
section.
## Proposed Solution
Brainstorm → spec → plan after 23 lands (the write path it composes
with) using 22's aged-store replay numbers as the policy input; extend
`04-db-binding.md`'s WAL section with the snapshot format the way the
record grammar is documented today.

View file

@ -185,6 +185,18 @@ transitively-traced check), corpus + TSan.
## Stage 3 — the DB actor + closing the arc
> **Guarantee obligations (2026-08-21 refinement, developer-approved)**
> — Tasks 7–8 build against these, in addition to their own checkboxes:
> (1) a worker write RPC is exactly ONE owner-shard commit; the ack
> crosses shards only AFTER the owner's fsync — a kill between send and
> commit leaves no ack and no partial state; (2) workers never open the
> WAL or data directory (debug-build assert); replay completes on the
> primary before any worker serves; (3) statements are serialized by the
> DB actor — replies are materialized copies, no torn reads under the
> concurrent multi-shard corpus (TSan). The five-property map lives in
> [the marker doc](../../in-progress/2026-08-21-arc-stage-3.md); disk
> space reclamation is story 32, not this stage.
### Task 7 — transparent DB RPC
**Files:** `runtime/src/builtin.c` (db cases marshal when not on shard