docs(db2-keys): reconcile databasev2 and porch markdown with the code

- loader's resident:keys refusal said "rows are still fully resident"
  and "until tasks 5c/5d land". Both false since f606fc9. Corrected to
  name the real blocker: UPDATE needs read-modify-append
- databasev2 00-story: the sequence graph drew 2->3->4, which reads as
  3 needing 2 and 4 needing 3. Both backwards, and it still drew the
  2->5->6 path the 2026-08-27 amendment retired. Redrawn stating only
  real dependencies, with 4 and 3 shown as composing rather than
  ordered, and the execution order that actually happened
- databasev2 03: the hazard and its Outstanding entry both claimed
  nothing fails "because iteration 2's storage half is unimplemented".
  Marked discharged, and recorded that the hazard named only half the
  danger — the bitmap walk would have dropped keys rows outright
- databasev2 06: pending -> hold (largely superseded, revisit only on
  a measurement); dated its 5c/5d references
- porch 01: rewritten to the settled shape. readiness ready, status
  in-progress, phases B and C marked superseded with why
- porch 01 claimed time.after "is still a reserved builtin id". False —
  builtin 90, implemented. That claim is what made the iteration look
  cheaper than it is
- porch README gains honest ledger rows for both features (partial,
  being rebuilt), not shipped
- skill-catalog README pointed at a story path that moved tracks;
  linkcheck now 0 broken

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
(cherry picked from commit b3d8c403e1d19ac27ec966de85cb293e0765795c)
This commit is contained in:
shoney.arickathil 2026-08-29 22:09:27 +02:00
parent 533fc5294f
commit 8311330531
8 changed files with 244 additions and 33 deletions

View file

@ -152,6 +152,8 @@ first (pure `.wo` cannot express it yet).
| Item | State |
| --- | --- |
| Rate limiting (fixed window, durable) | 🔶 tables + a first limiter exist (`middleware/store.wo`, `limiter.wo`); the counting path is being rebuilt to serialize through an actor pool, because read-modify-write from a handler fiber loses increments — [porch 1](../../stories/porch/01-store-backed-middleware.md) |
| Idempotent replay of unsafe requests | 🔶 tables + a first middleware exist (`middleware/idempotent.wo`); being rebuilt to block on the in-flight owner rather than store-after-completion, which cannot detect a collision at all — [porch 1](../../stories/porch/01-store-backed-middleware.md) |
| Transaction-per-request middleware (commit on 2xx, roll back otherwise) | ⏸ **v2** — needs iteration 18's `transaction { }` |
| Cancellation → rollback | ⏸ arc landed; still needs v2's `transaction { }` (iteration 18) |
| Migration generation + review workflow | ⬜ recorded future story (script-based destructive migrations) |

View file

@ -67,6 +67,55 @@ behind this board; live Obsidian Dataview views:
## ▶ NEXT PLAN
### Landed 2026-08-29 — databasev2 2 tasks 5c/5d, and a branch consolidation
**Implemented last time (2026-08-29):** `resident: keys` storage and every read
path. Rows are stored keys-only — the payload dropped after the WAL barrier,
the id map holding a log offset instead of a slab slot — and read back through
a borrow/release accessor. The scans go through a row iterator that walks the
id map for a keys table and the bitmap for a resident one, deliberately, since
hash order would reorder every unordered query.
**The obligation databasev2 3 left at the compactor is discharged**, and it
caught more than it predicted. The recorded hazard was that compaction moves
records and invalidates stored offsets. True — but the compactor walked the
slab *bitmap*, which a keys-resident row has no bit in, so every such row would
have been omitted from the new log outright. Silent data loss, not a bad
pointer, and rebuilding offsets would never have caught it. Records are now
moved byte-for-byte and re-pointed as they land.
**What is NOT done, and why the annotation is still refused:** updating a
keys-resident row. It lives in the log with no slab slot to mutate, so writing
into the borrow's scratch would discard the write *silently* — the one failure
this iteration must not ship. It needs read-modify-append. `wo_row_update_field`
and its slot variant refuse explicitly, and the loader still rejects the
annotation, with its message corrected to say so.
**Also this session:** every branch consolidated to `dev` and `master` only.
Three could not be replayed and are preserved as annotated tags rather than
merged or discarded — `archive/cleanup-pre-existing-changes` carries the Rust
runtime master deleted, and `archive/ipc-attach` + `archive/keypair-auth`
refactor the same row-API functions `db2-keys` rewrote. That last one is not a
merge conflict but an integration task: iteration 9c transfers ownership of
`vals` on failure, while the keys-resident arm returns early without freeing, so
a merge that compiles and passes could still leak or double-free.
**porch 1 brainstormed and specced.** Phases B and C are superseded before
review — both store the response after the handler returns, which cannot detect
an in-flight collision at all. Design settled: serialize through a sharded actor
pool, persist in a `@table`. Found and fixed en route: `docs/examples/porch/`
did not typecheck at all, for want of a `use json`.
**.dev / reference projects used:** the PostgreSQL checkpoint study
(`docs/plan/exploration/postgresql/buffer-and-checkpoint.md`) for the compaction
shape; the Fiber parity study for porch 1's scope.
**Dependencies unblocked:** nothing was blocked. databasev2 2's remaining tasks
6 and 7 depend only on iteration 2.
**Next steps:** databasev2 2 task 6 (the two runtime refusals) and task 7
(measure, gate, close out), then the porch 1 implementation plan.
### Landed 2026-08-29 — databasev2 3, WAL checkpoint (the chain's last link)
**Implemented last time (2026-08-29):** compaction. The log used to grow forever

View file

@ -162,10 +162,17 @@ before its mechanism existed; the history is in
| 10 | [Keypair attach auth](10-keypair-attach-auth.md) *(was language 21)* | program identity as a keypair; mutual challenge–response | 9 |
```
1 ──▶ 2 ──▶ 5 ──▶ 6
│ ▲
3 ──▶ 4 ─────┘
7, 8 independent
An arrow points AT the iteration that NEEDS the other.
1 ──▶ 2 ◀── 3 2 needs 1 (budget from a measurement) and 3 (the
│ offset map survives compaction). Since 5d, 3 also
▼ calls 2's row API — the coupling runs both ways.
5 5 needs 2. Nothing needs 5.
4 composes with 3 on the WAL commit path; NEITHER
needs the other. Executed 4 then 3 (chain 5, then 6).
6 superseded by 2 — not sequenced.
7, 8 independent.
9 ──▶ 10
```
@ -179,6 +186,21 @@ after 3, 4 and 5. `resident: keys` took that role into iteration 2, so 6 is
largely superseded and 5 is no longer a prerequisite for anything on the
critical path.
**Corrected 2026-08-29 — the graph above used to say the opposite of this
prose.** It drew `2 ──▶ 3 ──▶ 4`, which reads as 3 needing 2 and 4 needing 3.
Both are backwards. 2 needs 3 (the offset map), and the execution order that
actually happened is **4 before 3** — 4's part A landed 2026-08-28, 3 landed
2026-08-29, which is also what the `chain` field says (4 is chain 5, 3 is
chain 6). The retired `2 ──▶ 5 ──▶ 6` path was still drawn as well. Arrows now
point at the dependency, not at the reader's guess.
**The coupling between 2 and 3 runs both ways as of 5d.** 3's compactor calls
2's row API — the iterator, the offset accessor and its setter — because
compaction moves every record and must re-point the map it invalidates. That
was the hazard 3 recorded; it is now discharged, and it means compaction is not
a pure file operation. Full review in
[`00-databasev2-chain-review.md`](../../00-databasev2-chain-review.md).
## What this track does NOT own
| Not databasev2's | Owner |

View file

@ -58,12 +58,13 @@ chain: 6
> none; stop-the-world, with the pause measured against a stated budget rather
> than assumed acceptable.
>
> **The coupling that would otherwise be found late:** compaction moves every
> record, so it **invalidates every WAL offset**
> **The coupling that would otherwise be found late — and was found in time:**
> compaction moves every record, so it **invalidates every WAL offset**
> [iteration 2](02-table-storage-modes.md)'s `resident: keys` stores. The
> compactor rebuilds the offset map as it writes. Recorded now because iteration
> 2's storage half is unimplemented, so nothing breaks today — it would break
> later, looking like corruption rather than a design gap.
> compactor rebuilds the offset map as it writes. Recorded here while iteration
> 2's storage half was still unimplemented; **it landed 2026-08-29 and the
> obligation was discharged** (`f606fc9`), including a worse failure this note
> did not predict — see the hazard section at the end.
>
> **Measured on master 2026-08-28, grounding the whole iteration:** `seed 20000`
> leaves a 986 614-byte log; 20 000 updates take it to **2 590 262 bytes with the
@ -126,12 +127,11 @@ Met:
Outstanding:
- **The `resident: keys` offset map.** Compaction moves every record, so it
invalidates every WAL offset [iteration 2](02-table-storage-modes.md) stores.
The compactor must rebuild that map as it writes. **Nothing fails today**
because iteration 2's storage half is unimplemented — which is exactly why the
obligation is written at the compactor in `wal.c`, where the next implementer
hits it, rather than only in a spec they may not read.
- ~~**The `resident: keys` offset map.**~~ **Discharged 2026-08-29 by
iteration 2's task 5d** (`f606fc9`). The obligation written at the compactor
in `wal.c` did its job: the implementer hit it there. See the hazard section
below for what it caught — and for the second, worse failure it did not
predict.
- **The pause is O(live rows).** At ~181 MB/s a 1 GB live set implies ~5.5 s,
past any interactive budget. Incremental or forked copying was deliberately
not bought in advance; this is the number to buy it against.
@ -242,3 +242,30 @@ the snapshot persists the map and compaction is forbidden while any
`resident: keys` table is live. **The first is almost certainly right**,
but it means compaction cannot be written as a pure file operation that
ignores in-memory table state.
### Settled 2026-08-29 — and the hazard was only half the danger
The first shape was implemented, in iteration 2's task 5d (`f606fc9`).
Compaction re-points each row as it writes it, using a value-only map update
that cannot rehash, so a walk in progress stays valid and no per-row buffer of
new offsets is needed. Compaction is therefore **not** a pure file operation,
exactly as predicted above.
**What this section did not predict is the failure that would actually have
struck first.** It described stored offsets going stale — a pointer into a
rewritten file. But the compactor walked the slab **bitmap**, and a
keys-resident row holds no bitmap bit: its slot returns to the free list when
the payload is dropped. Every such row would therefore have been omitted from
the new log altogether. That is silent data loss, not a bad pointer, and
rebuilding offsets would never have caught it — the rows would simply have been
gone.
Both failure modes are now pinned by `test_keys_resident_survives_compaction`,
which rewrites rows in hash order so the offsets genuinely move; a map left
un-repointed lands on another row and fails the identity check rather than
passing by luck.
A failure *after* any row has been re-pointed is fatal by design: the map would
name offsets inside a temp file that the failure path unlinks, and the intact
original log replays correctly, so stopping is strictly better than serving
wrong rows.

View file

@ -1,7 +1,7 @@
---
track: databasev2
iteration: "6"
status: pending
status: hold
readiness: refine
---
@ -16,7 +16,8 @@ readiness: refine
> This iteration was written to implement a `cold` mode. That mode no longer
> exists: the brainstorm replaced it with `resident: all | keys`, and
> **`resident: keys` is the ceiling-raising mechanism** — indexes resident, rows
> read from the log by offset. It is iteration 2's tasks 5c/5d, not this file's.
> read from the log by offset. It is iteration 2's tasks 5c/5d — which landed
> 2026-08-29 — not this file's.
>
> Its premise was also specifically *rejected*, not merely relocated. This
> iteration assumed a **user-space resident working set** with faulting and an
@ -26,7 +27,8 @@ readiness: refine
> position `exploration/postgresql/buffer-and-checkpoint.md` already argued and
> the reason the engine avoids `O_DIRECT`.
>
> **What may still be left:** if measurement after 5c/5d shows the page cache
> **What may still be left:** if measurement after 5c/5d (landed 2026-08-29,
> still unmeasured — that is iteration 2's task 7) shows the page cache
> insufficient for some workload, a user-space working set becomes arguable
> again — but only with that number in hand, which is the opposite of how this
> file was written.

View file

@ -49,7 +49,7 @@ risky work starts.
| # | Iteration | Delivers | Needs |
| --- | --- | --- | --- |
| 1 | [Store-backed middleware](01-store-backed-middleware.md) | rate limiting + idempotency over a `@table` store | nothing new — starts today |
| 1 | [Store-backed middleware](01-store-backed-middleware.md) | rate limiting + idempotency over a `@table` store, serialized through a sharded actor pool | 🔄 in progress; needs no new primitive (`call`/`send`/`monitor`/`time.after` all landed) |
| 2 | [Randomness and cookies](02-randomness-and-cookies.md) | a `random_bytes` runtime builtin, repeated response headers, `Cookie:` parsing, signed cookies | a language-track builtin (phase A) |
| 3 | [Sessions](03-sessions.md) | server-side sessions, idle + absolute timeout, revocation | 2 |
| 4 | [CSRF](04-csrf.md) | token mint/verify, trusted origins, single-use tokens | 2, 3 |

View file

@ -1,14 +1,22 @@
---
track: porch
iteration: "1"
status: pending
readiness: refine
status: in-progress
readiness: ready
---
# porch 1 — store-backed middleware: rate limiting and idempotency
> Part of [Story — `porch`, the writeonce web framework](00-story.md).
> Source: [the Fiber parity study](../../plan/exploration/fiber/00-fiber-parity.md) §2.
> Spec: [`2026-08-29-porch-store-backed-middleware-design.md`](../../superpowers/specs/2026-08-29-porch-store-backed-middleware-design.md).
>
> **Rewritten 2026-08-29** after the brainstorm settled every fork. The
> iteration's premise changed: it was scoped as the cheapest slice because it
> needed "only a `@table` and `time.ticks`", and it now serializes through an
> actor pool. That is not new runtime surface — `spawn`, `send`, `call`,
> `monitor` and `time.after` all landed with the actor-lifecycle work — but it
> is more than the original framing, and the reason is in History below.
>
> **First deliberately because it is the cheapest.** Both features need only a
> `@table` and `time.ticks`, both of which already exist — no new builtin, no
@ -31,6 +39,35 @@ readiness: refine
- **Establish the store convention for iterations 2–4.** Sessions and CSRF will
want the same shape. Decide it once, here, on the cheap slice.
## The design, as settled
One rule, inherited by porch 2 (sessions) and 3 (CSRF):
> **Serialize through an actor. Persist in a `@table`. Never read-modify-write
> from a handler fiber.**
| Concern | Owner |
| --- | --- |
| per-key ordering, in-flight ownership, waiter lists | a sharded pool of actors, selected by hash of the key |
| counters, stored responses, request digests | `@table` rows — WAL-durable, replayed at boot |
| deciding whether a request passes | the actor, never the handler fiber |
The read-modify-write is the defect both features shared: a handler reads a
count, adds one and writes it back, so two interleaved fibers lose an
increment. An actor processes one message at a time, so routing both features
through the pool buys per-key serialization with no locks and no polling. A
pool rather than one actor because ordering is needed *per key*, never
globally.
## Progress
| Phase | State |
| --- | --- |
| A — store convention | ✅ `519d411` — two purpose-shaped tables. **Outstanding**: the request-digest column the refusal criterion needs |
| B — rate limiter | ⚠️ **superseded** — `5b1e82a` built the store-after shape; see History |
| C — idempotency | ⚠️ **superseded** — `aee7926` built the same shape; `5c3544d` fixed its missing `use json`, which had made the whole porch library uncompilable |
| D — gate and ledger | ⬜ not started; the restart leg is the one that must not be skipped |
## Phases
### Phase A — the store convention
@ -101,8 +138,20 @@ readiness: refine
arrives, **then** it is refused rather than answered with the other request's
response.
- **Given** two identical keyed requests in flight at once, **when** both are
dispatched, **then** exactly one executes and the other gets the decided
answer (replay or 409), never a partial write.
dispatched, **then** exactly one executes and the other receives that one's
stored response — never a refusal, never a partial write. *(Restated
2026-08-29: this used to permit 409. Blocking supersedes it — the duplicate
parks on `call` until the owner reports.)*
- **Given** N concurrent requests for one limiter key, **when** they are
counted, **then** the total is exactly N and no increment is lost. *(Added:
unreachable before the pool, and the defect that most undermines a limiter.)*
- **Given** a saturated actor pool, **when** a request arrives, **then** it is
refused with 503 rather than served uncounted. *(Added: fail-closed, because
saturating the pool must not become the limiter's bypass.)*
None of these are met yet — Phase D owns the gate, and B and C are being
rebuilt. The concurrency legs need genuine parallelism: a test that cannot
fail before the fix is not a test.
## Out Of Scope
@ -117,15 +166,23 @@ readiness: refine
- **Distributed limiting across processes.** One program owns its database;
cross-program state is language
[databasev2 9](../databasev2/09-cross-program-tables.md).
- **A background expiry sweeper.** No timer exists (`time.after` is still a
reserved builtin id in `wob.h`). Lazy pruning on access, deliberately.
- **A background expiry sweeper.** Lazy pruning on access, deliberately —
porch has no scheduler. **Corrected 2026-08-29**: this used to say "no timer
exists (`time.after` is still a reserved builtin id)". That is false —
`time.after` is builtin 90 and implemented (`runtime/src/builtin.c`, via
`wo_vm_timer_after`), along with `spawn` (68), `send` (69), `call` (88) and
`monitor` (89). The exclusion stands on its own merits: `time.after` is a
one-shot timer aimed at an actor, not a recurring sweep. But the *reason*
given was wrong, and it is the claim that made this iteration look cheaper
than it is.
- **The TTL cache middleware** — language
[iteration 18](../language-runtime-database/18-memory-db-features.md) owns it,
spec already approved. Do not build a second cache here.
## Info
Forks the spec must settle:
**All three settled 2026-08-29** — kept with their outcomes rather than
deleted, so the reasoning survives:
1. **One store or two?** A single generic key/value/expiry table serving both
features, or a purpose-shaped table each. Leaning two: the columns genuinely
@ -142,4 +199,50 @@ Forks the spec must settle:
cookies yet (iteration 2), so this is cheap to decide now and expensive to
retrofit later.
Nothing here needs a new runtime primitive, which is the point of going first.
Nothing here needs a new runtime primitive — still true, and now verified
rather than assumed: `call` is "a send that WAITS", whose park/reply protocol
lives in the VM, and that is exactly the blocking primitive the design needs.
**Outcomes:**
1. **Two stores**, as leaned. Confirmed by construction in `519d411`; a
generic table would have forced a counter and a response body through the
same `Text` column.
2. **The peer address, unless the app declares otherwise.** A `trust_proxy`
flag defaulting to off: off keys on `net.peer(req.conn)`, which cannot be
forged; on keys on the left-most `X-Forwarded-For` entry. Only the deployer
knows the topology, so the declaration belongs in their code. The built
version branched on `req.ctx["verified_proxy"]`, which nothing anywhere
sets — a dead branch. Verifying the proxy is genuinely story 35's, and
`client_ip` says so in its own comment.
3. **`content-type` only.** Settled as built, and settled correctly: replaying
a stored `Set-Cookie` or a stale `Date` is wrong, and cookies arrive in
iteration 2, so this is cheap now and expensive later.
## History
**2026-08-29 — Phases B and C superseded before review.** Both were built
against the original framing and both store the response *after* the handler
returns. Three of the seven original criteria cannot hold in that shape, which
is why this is a rebuild rather than a patch:
- **In-flight collision was undetectable.** `before` finds nothing and passes;
the row appears in `after`, once the handler has already run. Concurrent
duplicates both miss and both execute — there is no reservation to collide
on.
- **The 10-second in-flight heuristic was inverted.** The timestamp is stamped
when the response is *stored*, not when the request *starts*, so it fired on
legitimate fast replays — the common case — answering 409 where the stored
response was owed, and could never fire on a genuinely concurrent request.
- **"Reused key, different body is refused" was unreachable.** The body digest
was folded into the lookup key, so a different body was a different key and
simply missed. Safe — the wrong response is never served — but nothing looks
the bare key up, so nothing can refuse. The digest becomes a column.
Two defects independent of the redesign, both since fixed or scheduled: the
limiter emitted a monotonic tick into `X-RateLimit-Reset`, where a client
expects a Unix timestamp; and it used delete-then-insert where assigning to a
row field writes through (`compiler/src/emit.ml`), which doubled WAL traffic
and left a window where a failed insert after a successful delete silently lost
the counter — handing out a free window, the exact inverse of the durability
this iteration exists to demonstrate.

View file

@ -186,13 +186,19 @@ int wo_load_buf(wo_module *m, const uint8_t *buf, size_t len, char *err,
/* databasev2 2: `resident: keys` PARSES and sets this bit, but the
* storage half (tasks 5c/5d) is not implemented — rows are still fully
* resident. Accepting it would be an annotation the compiler honours
* in name only: a developer could declare a 120 GB table keys-resident,
* see it compile, and be OOM-killed. Refuse until the storage lands. */
* in name only.
*
* Narrowed 2026-08-29 (databasev2 2, 5c/5d): storage, boot, the read
* paths and checkpoint survival all landed. What has NOT landed is
* UPDATE — a keys-resident row lives in the log with no slab slot to
* mutate, so an update needs read-modify-append. Until that exists the
* annotation is still refused, because a table you can insert into and
* read but not update is a worse promise than one that never compiled. */
if (flags & WO_CLASSF_RESIDENT_KEYS)
BAIL("class %u declares `resident: keys`, which is NOT IMPLEMENTED "
"yet — rows are still fully resident, so the annotation would "
"be honoured in name only. Remove it until databasev2 2 tasks "
"5c/5d land; `resident: all` is what actually runs",
BAIL("class %u declares `resident: keys`, which is INCOMPLETE: rows "
"are stored and read keys-only, but UPDATING one is not "
"implemented (it needs read-modify-append). Remove it until "
"databasev2 2 lands updates; `resident: all` is what runs",
(unsigned)i);
if ((flags & WO_CLASSF_VOLATILE) && (flags & WO_CLASSF_RESIDENT_KEYS))
BAIL("class %u: durable:false with resident:keys — rows would have "