docs: borrow/GC analysis for the database engine, into the iterations

- 9b spec gains section 6 "Ownership, borrows, and GC across the
  engine boundary": two one-way copy gates (no VM pointer enters a
  row, everything a select returns is copied out), so the collector
  never traces engine memory and the engine never touches refcounts
- row views are borrows WITHOUT a runtime net: rows share the VM's
  field encoding but not its header, so no borrow word backs them --
  the compile-time escape rule is load-bearing alone
- cursor stability settled: scans materialize their id list before
  the body, row updates through the view stay legal (raise mode
  updates an indexed column mid-scan and is the proving fixture),
  insert/delete on a table with an open cursor is a new WO-E5xx
- GC-pause interaction recorded: collector runs between statements,
  a long scan delays slices -- accepted, documented
- iteration-7b ordering constraint: GC inference must classify before
  table-field validation, diagnostic names the inference reason --
  noted in 7b story, iteration-9 plan constraints, 9b plan tasks
- stories 09/09b Info sections point at the analysis; 9b plan Tasks
  3/5 carry the enforceable checkboxes (ASan boundary assertion)

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
shoney.arickathil 2026-08-15 09:57:53 +02:00
parent 1dc2912047
commit 4ddb83c2e1
6 changed files with 122 additions and 5 deletions

View file

@ -119,6 +119,19 @@ discarded — a check is an index probe, never query text).
- [ ] Cursor builtins + borrowed-row-view lifetime rules written into the
binding doc (a row view never escapes the loop that opened the cursor —
the ownership pass enforces it, mirror of the container-read borrow).
**Rows have no borrow word** (they share the VM's field encoding, not
its header), so unlike every VM-heap borrow there is no runtime trap
behind this rule — the compile-time check is load-bearing alone, which
is why it gets its own diagnostic and goldens rather than riding on
E30x.
- [ ] Cursor stability per spec section 6: scans **materialize their id list**
before the body runs and point-read per iteration; updates through the
row view stay legal (exclusive row borrow, index maintenance at the row
API); `insert`/`delete` targeting a table with an open cursor is a new
WO-E5xx (the ownership pass carries the open-cursor table set through
the loop body); read-only nested queries over the same table stay legal.
The `raise` mode — updating an indexed column mid-scan — is the fixture
that proves the materialized-id semantics.
- [ ] FK trap + restrict trap wired through the row API; a debug-build probe
counter exposed for Task 5's index-selection proof.
- [ ] The `delete` statement (point delete of a row value) lands here too:
@ -164,14 +177,18 @@ no new ownership classes, and the ownership pass's existing drop machinery
covers the query's temporaries because the lowering IS ordinary loops.
- [ ] Desugar + lowering for every clause; disassembly of the sample's report
mode shows loops and builtins, no plan tree, no text.
mode shows loops and builtins, no plan tree, no text. The select
boundary is the ownership bulkhead (spec section 6): everything a query
returns is copied or freshly built, so no value anywhere points into a
row slab after the query ends — asserted under ASan by mutating rows
after a query and re-reading the query's results.
- [ ] Index selection proven: the acceptance asserts probe-counter deltas for
the indexed `staff <department>` path versus a full-scan query.
- [ ] `oop-e2e`, `woc-test`, ASan-corpus gates green; commit locally.
### Task 6: The employee sample is the acceptance
**Concept & reason:** spec section 6 verbatim — `docs/examples/employee` with
**Concept & reason:** spec section 7 verbatim — `docs/examples/employee` with
`Department`/`Employee` (`@table`, `@unique` name, composite `[dept, salary]`
index, `ref`/`backlink` pair), wo.toml manifest so `woc .` builds it, a `just`
module beside it (the log-watcher convention), and `scripts/employee-accept.sh`
@ -198,7 +215,7 @@ and `report` proving replay on this workload.
## Out of scope — deferred by name
- Everything spec section 7 lists: set operators, outer joins, subqueries,
- Everything spec section 8 lists: set operators, outer joins, subqueries,
composite group keys, groups as values, FK cascade/set-nil, deferred
checks, SQL text in any role, cross-shard queries, `LIVE`, migrations,
cost-based planning.

View file

@ -80,6 +80,13 @@
## Info
- Governing spec: [`docs/superpowers/specs/2026-08-11-inferred-gc-mark-sweep-design.md`](../../superpowers/specs/2026-08-11-inferred-gc-mark-sweep-design.md).
- **Constraint added by the database track (2026-08-15):** a GC-managed value
in a `@table` field is a compile error (the engine/heap bulkhead — 9b
design, section 6). Once GC-ness is inferred rather than annotated, the
inference pass must classify every class **before** table-field validation,
and the diagnostic must name the inference reason ("class X is
garbage-collected via Y and cannot be stored in a table field") — otherwise
the error becomes unactionable exactly when it stops being self-evident.
- **Why the annotation was insufficient, not merely inconvenient:** the OOP
spec's own example, `@gc class PriceCache { entries: map<SKU, Money> }`, is
acyclic. It needs GC because it is shared, and second-class borrows cannot

View file

@ -47,6 +47,14 @@
lesson, the C engine enforces it.
- The wo-db overlap manifest keeps the C++ prototype and this engine
answer-compatible where features overlap.
- **Ownership/GC analysis (2026-08-15):** the engine and the VM heap are two
memory worlds crossed only by copy — rows store no VM pointers, GC-managed
values in stored fields are a compile error, and everything a query returns
is copied out — so the collector never traces rows and the engine never
counts references. The full analysis (row views as borrows without a
runtime net, cursor stability, GC-pause interaction) lives in the 9b
design's section 6:
[`2026-08-15-table-relations-query-design.md`](../../superpowers/specs/2026-08-15-table-relations-query-design.md).
## Proposed Solution

View file

@ -116,6 +116,16 @@ delegates, and `IQueryable`'s runtime expression trees, the last of which
depends on reflection that principle 13 forbids outright. Take the vocabulary
and the semantics; leave the plumbing.
**4. (Settled with the spec, 2026-08-15) How do queries interact with the
borrow checker and the GC?** Row views are borrows of engine memory with **no
runtime borrow word behind them** — the compile-time escape rule is
load-bearing alone. Scans materialize their id list up front, so updating a
row (even an indexed column) inside the loop is sound, while `insert`/`delete`
on a table with an open cursor is a compile error. The GC never meets the
engine at all: both directions across the boundary are copies, and GC-managed
values cannot be stored — spec section 6 is the full analysis, including the
iteration-7b ordering constraint (inference before table-field validation).
Also relevant: `@table(name:, index:)` already parses today with known-key
validation (`WO-E102`), the Rust runtime already ships secondary indexes and
`find_by` behind that annotation, and `ref T` already classifies as a scalar

View file

@ -20,6 +20,7 @@
- **Id discipline:** ids interleave per shard (`t+1, t+1+N, …`); a row's owner shard is `(id-1) % N`; point ops on a foreign row hop once via the plan-4 mailbox — creates are always local.
- **Index doctrine:** secondary indexes update only inside the engine's insert/remove path; direct storage mutation is a defect by definition.
- **RAM is authoritative:** reads never touch a file descriptor (phase-B doctrine); disk exists for durability and boot.
- **The GC bulkhead (analysis 2026-08-15, spec'd in the 9b design section 6):** values cross between VM heap and row storage only by copy, and a GC-managed value in a stored field is a compile error — so the collector never traces engine memory and the engine never touches reference counts. When iteration 7b makes GC-ness inferred, inference must classify classes **before** table-field validation so this error keeps firing, with the message naming the inference reason.
- **Format changes go through the format doc:** DB operations extend the builtin table (ids appended to `docs/plan/oop-vm/00-wob-format.md`); no opcode-space or version change.
---

View file

@ -259,7 +259,81 @@ cursor/group/probe set), no new opcodes, no version bump beyond iteration 9's.
---
## 6. Acceptance workload — `docs/examples/employee`
## 6. Ownership, borrows, and GC across the engine boundary
The engine and the VM heap are two memory worlds, and the whole safety story
is that values only ever CROSS between them by copy. Analysis recorded here
because both iterations' correctness hangs on it (2026-08-15).
### The bulkhead: two one-way gates, both already doctrine
- **Into storage:** a row stores no VM pointer — scalars copy, Texts copy,
owned objects flatten by value, containers copy element-wise, `ref` is an
id, and a GC-managed value in a `@table` field is a **compile error**
(iteration 9's field-encoding rules). So no row ever points at a GC object.
- **Out of storage:** everything a `select` emits is copied or freshly built
at the boundary — Texts via the established copy rule, projections as new
records. So no GC root, no local, and no container ever points into a row
slab once the query ends.
Consequence: **the collector never traces engine memory and the engine never
touches reference counts.** Iteration 7b (inferred GC, mark-sweep) does not
change this — it changes only *when* the "into" gate's error fires: GC-ness
becomes inferred, so inference must classify every class **before** table-field
validation runs, and the diagnostic reads "class X is garbage-collected
(inferred via Y) and cannot be stored in a table field." A class stored in a
table is thereby constrained to ownership-expressible shapes; that is a
feature, not a limitation — tables are the language's answer to shared
long-lived data, which is most of what `@gc` exists for.
### Row views are borrows without a runtime net
A cursor yields a **row view**: a borrow of engine-owned memory, valid until
the cursor advances or closes. Two things make this different from every
borrow the language has today:
- VM-heap borrows have a runtime defense (the object header's borrow word,
`WO_T_BORROW` traps). Rows share the VM's field *encoding* but not its
header — there is no borrow word in a row slab, so **the compile-time rule
is load-bearing alone**. The ownership pass enforces: a row view never
escapes the query loop that produced it (the container-read-borrow mirror),
and anything that leaves does so as a copy through `select`.
- The program can mutate the table it is iterating — single-writer per shard
removes concurrent writers, not the program's own hand.
### Cursor stability: materialize ids, allow row updates, forbid structural
The `raise` mode is the honest case: it updates `salary` — an **indexed**
column — while iterating an index scan. Naive cursor-over-index breaks here
(entries move mid-scan). The semantics, chosen for KISS and enforceability:
- **A scan materializes its matching id list before the body runs**, then
point-reads each row per iteration. O(matches) ids of memory, recorded as
the cost; index-order iteration falls out for free.
- **Updates through the row view are allowed** — the view is an exclusive
borrow of that row for the iteration (the `mut` analog); index maintenance
for the changed column happens at the row API as always, and cannot disturb
the already-collected id list.
- **`insert` into or `delete` from a table with an open cursor is a compile
error** (new WO-E5xx): a materialized id list cannot defend a point-read
against a row deleted mid-loop, and silently skipping a vanished id is the
kind of quiet wrongness this language exists to refuse. The ownership pass
carries an open-cursor table set through the loop body, statically — insert
and delete name their target class at compile time. Read-only nested
queries over the same table remain legal (shared borrows).
### Query temporaries and GC pressure
Group hash tables and join build sides are engine-side C allocations scoped
to the statement — freed when the query ends, invisible to both the drop
tables and the collector. Query results are ordinary owned VM values, freed
by the existing drop machinery. Nothing on the query path allocates a GC
object or an RC operation. One accepted interaction: the budgeted collector
runs between statements, so a long full-table scan delays GC slices for its
duration — acceptable at this scale, recorded so nobody rediscovers it as a
latency mystery.
## 7. Acceptance workload — `docs/examples/employee`
A new sample, structured like `log-watcher` (wo.toml manifest, program mode,
`just` module, acceptance script), small enough to read in one sitting and
@ -293,7 +367,7 @@ The sample is 9b's acceptance the way log-watcher was iterations 1–7's: no new
corpus fixtures beyond the db corpus iteration 9 already plans; the sample is
the test.
## 7. Out of scope (inherited and new)
## 8. Out of scope (inherited and new)
- Everything 9b's story already excludes: cross-shard queries and distributed
joins, `LIVE` subscriptions, migrations, cost-based planning.