writeonce/docs/in-progress/2026-08-21-arc-stage-3.md
shoney.arickathil e7ca2fdd7c docs: stage 3 refined with guarantee contract; story 32 born
- five-property map in marker doc: atomicity/durability/recovery
  proven or held (18); concurrency control = stage 3's property;
  space reclamation RAM done (slot reuse), disk = new story 32
- story 08: three stage-3 criteria (one commit per write RPC +
  ack-after-owner-fsync, workers WAL-free + replay-before-serve,
  no torn reads under TSan corpus); arc plan stage 3 carries them
- refine/32-wal-checkpoint.md: snapshot + truncate, bounded replay,
  crash-during-checkpoint safe; four forks; after 23
- chain now stage 3 -> 22 -> 31 -> 24 -> 23 -> 32 in all 10 docs;
  boards + seq bumps synced

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-21 11:26:45 +02:00

3.6 KiB
Raw Blame History

In progress — the 8+11 arc, stage 3: the transparent DB actor

Status: 🔄 in progress (started 2026-08-21) — first slice of the concurrency chain stage 3 → 22 → 31 → 24 → 23 → 32. Board: ../00-status.md.

This folder holds ONE marker doc: the slice being executed right now, so the active work is findable without reading the board. When the slice lands, the board's done section takes the record and this file is deleted — the plan and stories stay the authority.

What

Stage 3 (Tasks 7–8) of the plan of record, superpowers/plans/2026-08-20-shard-fiber-arc.md: engine calls off the owner shard become message sends with parked replies. A correctness fix, not an optimization — rt.db is set only on the primary (runtime/src/main.c), worker VMs are zero-initialized, so any DB statement off the primary traps WO_T_DB today.

Why now

Developer decision 2026-08-21: correctness before measurement. Fixing the hole first lets iteration 22's single benchmark campaign cover single- AND multi-shard honestly.

Stories

  • 8 — shard-actor runtime (stage 3 is its remaining scope)
  • 11 — fibers (substance landed stages 1+2; closes with this slice — note: fs still blocks instead of parking, land it here or re-scope 11's criterion explicitly at close)

Guarantee contract (refined 2026-08-21)

The five store properties, mapped honestly — what this slice owes vs what is already proven or deferred:

property state
Atomicity per-statement ✅ (WAL record replays whole-or-not-at-all); multi-statement = transaction { }, iteration 18, ⏸ held. Stage-3 obligation: a worker write RPC is exactly ONE owner-shard commit — a crash between send and commit leaves no ack and no partial state.
Durability ✅ fsync-per-commit, ack-after-durable. Stage-3 obligation: the ack crosses shards only AFTER the owner's fsync completes. Power-loss rides fdatasync semantics; 22's kill battery is the scripted proof.
Crash recovery ✅ boot replay, torn-tail drop, index rebuild; 22 scripts the restart proof. Stage-3 obligation: workers never open the WAL or data dir; replay completes on the primary before any worker serves.
Concurrency control THIS SLICE. The DB actor serializes every statement; replies are materialized copies — no torn read can exist by construction. Each statement sees the serialized moment its envelope executes; cross-statement snapshots arrive with 18.
Space reclamation RAM ✅ — deleted rows free their slot (04-db-binding.md: "Ids are never reused; slots are"). Disk ✖ — the WAL grows unbounded, no checkpoint exists → story 32, end of chain.

Definition of done

  • The guarantee obligations above hold: one commit per write RPC with ack-after-owner-fsync, workers WAL-free with replay-before-serve, and no torn reads under concurrent multi-shard load.
  • Plan Tasks 7–8 checked off; full battery green (just woc-test, just oop-e2e, just deps-accept, just web-app, just log-watcher, just employee, just fibers) — multi-shard DB answers byte-identical to single-shard, WAL ack-after-durable unchanged.
  • Stories 08 + 11 move to stories/language-runtime-database/done/ together; board updated in the same change.
  • Next slice: iteration 22 (measurement backbone) replaces this file.