diff --git a/.gitignore b/.gitignore
index ee9ec15..e5975fb 100644
--- a/.gitignore
+++ b/.gitignore
@@ -3,6 +3,7 @@
# `woc .` manifest builds (wo.toml [build] target)
/docs/examples/log-watcher/target
+/docs/examples/employee/target
# Rust runtime (crates/rt/): compiled binary + build artifacts
/crates/rt/target
diff --git a/README.md b/README.md
index 113ddcc..8954eb0 100644
--- a/README.md
+++ b/README.md
@@ -1,77 +1,323 @@
# writeonce
-A declarative full-stack programming language. You write `.wo` files; the runtime compiles them into a binary that owns the database, serves REST, and pushes live subscriptions — no external database, no external web server, no frontend framework.
+**A small compiled language with a database built in.** You write `.wo`
+files; one command turns them into a single native binary that carries its
+own storage engine — a typed, WAL-durable, crash-recoverable database — with
+no server to install, no ORM, and no query strings. Tables are just classes,
+queries are written in the language and checked by the compiler, and the whole
+program ships as one file that depends only on the system C library.
-Think **Go + Postgres + `net/http` + Phoenix LiveView, folded into one language and one binary.**
+> **Status: early, honest.** Everything documented on this page compiles and
+> runs today and is exercised by the acceptance tests in this repository.
+> Features that are planned but **not yet available** are listed separately
+> under [Roadmap](#roadmap) — they are not described as if they work. Nothing
+> here is API-stable yet.
-# persistant database
+---
-- reads and writes database to RAM, persist data to postgres SQL.
-- The entire database lives in RAM; every committed write is mirrored to PostgreSQL **as a backup** — asynchronously, behind the runtime's own WAL, never in the read or ack path. Set `WO_PG=postgres://user@host:5432/db` and every type's rows appear as a Postgres table (named by its `@table(name: ...)` annotation) that you can query with plain `psql`. Plan and phases: [`docs/plan/16-postgres-mirror.md`](docs/plan/16-postgres-mirror.md); try it: `just pricing-pg-demo`.
+## Why writeonce
-## Quickstart
+- **The database is part of the language.** A `class` marked `@table` *is* a
+ table. Its rows persist through a write-ahead log, survive a restart, and are
+ reached by navigating typed relations — not by assembling SQL text.
+- **Queries are compiled, not interpreted.** `from e in Employee where
+ e.salary > 90000 select e` lowers to bytecode loops over the engine. A
+ mistyped field name is a **compile error**, not a runtime surprise. There is
+ no SQL string anywhere in the shipped binary.
+- **One binary, no runtime dependencies.** `woc .` produces a self-contained
+ executable (~100 KB for the sample programs) that links only libc. Copy it to
+ a server and run it.
+- **Small on purpose.** No FFI, no package manager, no framework. The standard
+ library is a handful of OS modules. The language is designed to be read.
+
+writeonce is **not** a web framework and does not (yet) serve HTTP, WebSockets,
+or a UI. It is a systems language whose distinguishing feature is the embedded
+database. If you have seen an older "writeonce" that served REST from `cargo
+run`, that was a separate, earlier runtime; this page documents the current
+`woc`/`wovm` toolchain.
+
+---
+
+## System requirements
+
+**To run a compiled writeonce program:**
+
+- Linux on x86-64. The produced binary is a native executable that links only
+ the system C library (`libc`); nothing else is required at runtime.
+
+**To build programs from source (the toolchain), you need:**
+
+| Tool | Version tested | Purpose |
+| --- | --- | --- |
+| OCaml | 4.14+ | builds `woc`, the compiler front end |
+| dune | 3.14+ | OCaml build driver |
+| A C11 compiler | gcc 13 / clang | builds `wovm`, the runtime VM |
+| just | 1.x | task runner for the build/test recipes |
+| make | any | drives the runtime build |
+
+Other POSIX platforms (macOS, BSD) are untested. The toolchain itself has no
+network or package-download step — it builds entirely from the checked-in
+source.
+
+---
+
+## Getting the toolchain
+
+Two artifacts make up the toolchain:
+
+- **`woc`** — the compiler (OCaml). Reads `.wo` source, type-checks it, runs
+ the ownership pass, and emits a `.wob` image or a standalone binary.
+- **`wovm`** — the runtime (C11). Loads a `.wob` image and executes it. When
+ `woc` builds a standalone binary, it embeds the image into a copy of `wovm`.
+
+Build both from the repository root:
```bash
-git clone https://github.com/shoneyJ/writeonce
-cd writeonce
-cargo run --bin wo -- run docs/examples/blog # serve the sample blog on :8080
-curl http://127.0.0.1:8080/api/articles # it's a real REST API now
+just woc-build # builds compiler/_build/default/bin/woc
+just wovm-build # builds runtime/wovm
+
+# gate them (optional but recommended)
+just woc-test # compiler unit + golden suites
+just wovm-test # runtime unit suites, both dispatch flavors, ASan-clean
```
-See [`.dev/reference/rest/blog.rest`](.dev/reference/rest/blog.rest) for a preconfigured HTTP-request file that drives the whole sample — open it in VS Code (with the REST Client extension) or JetBrains and click "Send Request" on each block.
+---
-## What this repository contains
+## Your first program
-| Path | What it is |
-| ------------------------------------------------------------------------------------------ | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
-| [`crates/rt/`](crates/rt/) | The new `.wo` language runtime — lexer, type-DSL parser, in-memory engine, axum REST server. Produces the `wo` binary. |
-| [`crates/{ql,value,engine,txn,db,wal,sub,http,gen,policy,logic,service,ui,app}/`](crates/) | 14 empty placeholder crates scaffolded for Phases 2–6. Real code extracts from `rt/` as each phase activates. |
-| [`docs/runtime/wo-language.md`](docs/runtime/wo-language.md) | **Start here.** The language overview: toolchain, hello-world, stdlib, client model. |
-| [`docs/runtime/database.md`](docs/runtime/database.md) | The 7-phase engineering series that drives the runtime's design. |
-| [`docs/examples/blog/`](docs/examples/blog/) | Sample `.wo` project: blog with articles, authors, tags, comments. ~200 lines. |
-| [`docs/examples/ecommerce/`](docs/examples/ecommerce/) | Sample `.wo` project: storefront + live order-ops table + cross-paradigm checkout. ~300 lines. |
-| [`prototypes/wo-db/`](prototypes/wo-db/) | C++ prototype of the query-layer engine (SQL + Cypher + document paths, `RETURNING` aliases, `LIVE` stub). ~2k lines, smoke tests pass. Reference implementation the Rust port follows. |
-| [`.dev/reference/rest/`](.dev/reference/rest/) | `.rest` files (VS Code REST Client / JetBrains HTTP format) for manually testing the running prototype. |
-| [`.dev/reference/crates/`](.dev/reference/crates/) | The v1 writeonce blog — 13 Rust crates implementing the original `.seg` + sidecar-index storage engine and `.htmlx` templating. Preserved as a nested workspace; see [`.dev/reference/README.md`](.dev/reference/README.md). |
+A writeonce project is a directory with a `wo.toml` manifest and one or more
+`.wo` files. Every program has an entry point:
-## Current stage
+```
+-- hello/main.wo
+fn main(args: multi Text) -> Int {
+ print("hello, writeonce");
+ return 0;
+}
+```
-The runtime is under active development. Each stage lands as an independently shippable cut:
+```toml
+# hello/wo.toml
+name = "hello"
+version = "0.1.0"
-| Stage | What works | Status |
-| ------ | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------------------------------------------------------- |
-| **1** | `wo run
` discovers every `.wo` file under a directory | ✅ shipped |
-| **2** | Type-DSL parser, in-memory engine, REST CRUD (`list` / `get` / `create` / `update` / `delete`) generated from `service rest` blocks, JSON bodies with auto-id, default-value seeding, partial-update PATCH | ✅ shipped — `cargo run -- run docs/examples/blog` |
-| **3** | LIVE subscriptions over WebSocket, delta frames on commit, `me` / session layer | pending |
-| **4+** | Transactional fns (`fn checkout in txn snapshot`), row-level policies, type-attached triggers, `##ui` SSR, WAL durability, codegen | see [docs/runtime/database.md](docs/runtime/database.md) |
+[runtime]
+wo = ">= 0.1"
+```
-`cargo test --lib` at the root runs 14 unit tests covering the lexer, parser, compiler, and engine. Stage-3 endpoints respond `501 Not Implemented` until they land.
-
-## Build & test
+Compile the directory into a single binary and run it:
```bash
-cargo build # builds all 15 crates (only `rt` has real code)
-cargo test --lib # 14 unit tests
-
-cargo run --bin wo -- run docs/examples/blog # serve the blog sample
-cargo run --bin wo -- run docs/examples/ecommerce # serve the ecommerce sample
-
-# Override the listen address
-WO_LISTEN=127.0.0.1:9000 cargo run --bin wo -- run docs/examples/blog
+woc hello/ # produces hello/target/hello
+./hello/target/hello
+# hello, writeonce
```
-## The v1 codebase (reference)
+`main` returns an `Int` — that value is the process **exit code**. `args` is
+the command-line arguments (the program name is not included).
-The original writeonce blog engine — 13 crates, flat-file `.seg` storage, sidecar indexes, `.htmlx` templates, hand-rolled `epoll` event loop — moved to [`.dev/reference/crates/`](.dev/reference/crates/) when the new runtime was scaffolded. It's a nested Cargo workspace:
+### The two build paths
```bash
-cd .dev/reference/crates
-cargo build # all 13 v1 crates still compile
-cargo test # 12 unit tests, 1 ignored integration test
+# 1. standalone binary (what you ship): woc reads wo.toml, emits target/
+woc myproject/
+
+# 2. image + VM (handy while developing): emit a .wob, run it with wovm
+woc --emit myproject/ -o app.wob
+wovm app.wob arg1 arg2
```
-V1 crates keep the `wo-` prefix (`wo-seg`, `wo-store`, …). The new runtime crates dropped it (`ql`, `value`, `engine`, …). [`docs/runtime/database/07-wo-seg-migration.md`](docs/runtime/database/07-wo-seg-migration.md) is the phased coexistence plan for replacing v1 with the new runtime — abstract behind a trait, dual-write, cut over, decommission.
+Both paths run the same program. The standalone binary is the release artifact;
+the image path lets you inspect or move the image around.
-## License & status
+---
-Work in progress. Nothing here is stable. Read the language overview in [`docs/runtime/wo-language.md`](docs/runtime/wo-language.md) if you want to know the shape; read the phase docs if you want to see the engineering plan; look in [`docs/examples/`](docs/examples/) if you want to see what the end product feels like.
+## Language at a glance
+
+writeonce is statically typed with a compile-time ownership model — every value
+has a known owner, memory is freed deterministically, and values that form
+cycles are collected by an inferred garbage collector (you never annotate GC-
+ness; the compiler infers it). The surface will look familiar:
+
+- **Types:** `Int`, `Text`, `Bool`, and user `class` types. `?T` marks an
+ optional (nullable) value; `nil` is the empty case.
+- **Containers:** `multi T` (a growable list) and `map`. Literals:
+ `[]`, `[a, b]`, `{}`.
+- **Classes & records:** classes with fields and methods, `static const` /
+ `static fn` members, module-scoped across files.
+- **Control flow:** `if`/`else`, `for x in xs`, `for k, v in m`, `switch`
+ expressions, and `try { … } catch (e) { … }` (also an expression form).
+- **Strings:** interpolation with `${expr}` inside a `"…"` literal.
+- **Functions:** free functions and methods; arguments and returns are typed.
+
+```
+fn classify(n: Int) -> Text {
+ if n < 0 { return "negative"; }
+ return switch n {
+ case 0: "zero";
+ default: "positive";
+ };
+}
+```
+
+### Standard library
+
+A compact set of OS modules, reached by their reserved names — no imports:
+
+| Module | What it does |
+| --- | --- |
+| `fs` | `exists`, `list`, `stat`, `read_all`, `read_at`, `append` |
+| `time` | `sleep`, `now`, `local`, `iso` |
+| `env` | `get`, `stopping` (a cooperative shutdown flag) |
+| `net` | TCP `listen` / `accept` / `read` / `write` / `close` (host + port) |
+| `proc` | `run` a child process, capture stdout/stderr/exit |
+| `json` | `encode` / `decode` (`json.decode(t) as T` yields `?T`) |
+
+These are deliberately minimal — the surface a real program needs, and no more.
+
+---
+
+## The database
+
+This is the point of the language. Declaring storage is declaring a class:
+
+```
+@table(name: "departments", index: [name])
+class Department {
+ name: Text @unique
+ staff: backlink Employee.dept -- reverse relation, not a stored column
+}
+
+@table(name: "employees", index: [dept], index: [dept, salary])
+class Employee {
+ name: Text
+ salary: Int
+ hired: Int
+ dept: ref Department -- foreign key: stored as the row id
+}
+```
+
+- **`@table`** makes a class persistent — named storage plus declared secondary
+ indexes. Every instance you `insert` is written to a write-ahead log **before**
+ it is acknowledged, so an acked write survives a crash; on the next start the
+ log is replayed.
+- **`ref T`** is a typed foreign key (a forward relation). **`backlink T.f`** is
+ its inverse — a virtual field, no stored column, resolved by an index scan.
+- **`@unique`** enforces uniqueness at insert/update; a violation is a
+ **catchable** trap.
+- **Foreign keys restrict deletes**: deleting a row that another row still
+ references traps rather than orphaning it.
+
+### Writing and reading data
+
+Mutation is direct; queries are a comprehension the compiler lowers to engine
+operations:
+
+```
+-- insert (WAL-durable); @unique makes a re-insert trap, and try/catch it:
+let eng = try insert Department { name: "Engineering" } catch (e) nil;
+insert Employee { name: "Asha", salary: 9200000, hired: 1704067200000, dept: eng };
+
+-- query: filter, order, limit, project — checked at compile time
+for e in from s in Employee where s.salary > 8000000 order by s.salary desc select s {
+ print("${e.name} ${e.salary} (${e.dept.name})"); -- ref navigation
+}
+
+-- navigate a backlink (the department's staff), update through the result
+for e in from s in dept.staff select s {
+ e.salary = e.salary + e.salary * 5 / 100; -- update-through-row
+}
+
+-- delete (restricted if still referenced)
+let ok = try delete row catch (e) nil;
+```
+
+The query surface available today is **`from v in
where …
+[order by k [desc]] [take n] select v | v.field`**, plus `insert`, delete, and
+update-through-a-row. It is proven end to end by the `employee` sample, whose
+data survives a process restart via log replay.
+
+---
+
+## Project layout & the manifest
+
+```
+myproject/
+├── wo.toml # manifest: name, version, [runtime], [build]
+├── main.wo # entry point (fn main)
+├── types.wo # your @table classes, other types
+└── target/ # build output (the standalone binary lands here)
+```
+
+```toml
+name = "myproject"
+version = "0.1.0"
+
+[runtime]
+wo = ">= 0.1"
+
+[build]
+runtime = "../../../runtime/wovm" # path to the wovm the binary is built from
+```
+
+`woc myproject/` compiles every `.wo` file under the directory as one program.
+
+Programs that create tables read their data directory from the `WO_DATA`
+environment variable at run time:
+
+```bash
+WO_DATA=./data ./target/myproject seed
+WO_DATA=./data ./target/myproject report # a fresh process still sees the data
+```
+
+---
+
+## Worked examples
+
+Two complete sample programs live in the repository and double as the language's
+acceptance tests:
+
+- **`docs/examples/employee/`** — departments and employees related by
+ `ref`/`backlink`, `@unique`, foreign-key restrict on delete, per-department
+ reports, and persistence across a restart. Run it:
+
+ ```bash
+ just employee # compile + run every mode against a durable database
+ ```
+
+- **`docs/examples/log-watcher/`** — a long-running daemon that watches log
+ files for silent death, using the `fs`/`time`/`net`/`proc` stdlib. Run it:
+
+ ```bash
+ just log-watcher
+ ```
+
+Read either program's `main.wo` for idiomatic, working writeonce.
+
+---
+
+## Roadmap
+
+Planned, **not yet available** — listed so the shipped surface above stays
+honest. These exist as design iterations and/or work-in-progress branches, not
+as features you can use today:
+
+- **Query aggregates** — `group … by … into g` with `count`/`avg`/`min`/`max`
+ and projection records. (Today the same result is written by hand from the
+ shipped primitives.)
+- **HTTP service layer** — `service` blocks that route requests to methods.
+- **Concurrency** — a shard-actor runtime and green-threaded fibers.
+- **Cross-program database access** — one program attaching to another's
+ database over a local channel, with keypair authentication and per-client
+ rights.
+- **Blue-green deployment** — in-process recompile and atomic version switch.
+- **Compile-time metaprogramming** — `@derive(Json/Csv/Eq/…)` generated from a
+ class's own metadata, no reflection.
+
+Known current limits worth naming: `net` is TCP host+port only; `proc.run` has
+no timeout or signal control; there is no stdin/stdout byte I/O and no FFI.
+
+---
+
+*writeonce is a work in progress. Interfaces will change. If you build
+something with it, pin to a commit.*
diff --git a/compiler/bin/main.ml b/compiler/bin/main.ml
index a8f1276..a3240bb 100644
--- a/compiler/bin/main.ml
+++ b/compiler/bin/main.ml
@@ -600,10 +600,23 @@ let manifest_parse (path : string) : (string * string) list =
if line = "" || (String.length line >= 1 && line.[0] = '#') then ()
else if line.[0] = '[' then begin
if line.[String.length line - 1] <> ']' then fail !lineno "malformed section header";
- section := String.sub line 1 (String.length line - 2);
- if !section <> "runtime" && !section <> "build" then
- fail !lineno (Printf.sprintf "unknown section [%s] (runtime and build exist)" !section)
+ (* accept `[[table.array]]` headers too (iteration 9c's
+ [[share.clients]]) by trimming the doubled brackets *)
+ let inner = String.sub line 1 (String.length line - 2) in
+ let inner =
+ if String.length inner >= 2 && inner.[0] = '[' && inner.[String.length inner - 1] = ']'
+ then String.sub inner 1 (String.length inner - 2)
+ else inner
+ in
+ section := inner;
+ if !section <> "runtime" && !section <> "build" && !section <> "share"
+ && !section <> "share.clients"
+ then
+ fail !lineno
+ (Printf.sprintf "unknown section [%s] (runtime and build exist)" !section)
end
+ else if !section = "share" || !section = "share.clients" then
+ () (* iteration 9c manifest keys — parsed by the attach feature, ignored here *)
else
match String.index_opt line '=' with
| None -> fail !lineno "expected `key = \"value\"`"
diff --git a/compiler/src/ast.ml b/compiler/src/ast.ml
index d02879b..840f2f0 100644
--- a/compiler/src/ast.ml
+++ b/compiler/src/ast.ml
@@ -69,6 +69,9 @@ type field_ty =
| Ref of string
| Multi of string
| Map of string * string (* key type, value type: map *)
+ | Backlink of string * string (* backlink C.f: the computed inverse of a
+ `ref` — NOT a stored column; reading it
+ scans C's index on f. Types as multi C. *)
| Nullable of field_ty (* ?T wrapper *)
(* Parameter passing convention (spec section 3, rule 2): default is an
@@ -211,6 +214,15 @@ and expr_kind =
| Binary of binop * expr * expr
| Ctor of string * (string * expr) list
| DbStub of Token.t list
+ (* `insert Class { field: expr, ... }` — the FIRST DB statement to leave
+ the stub behind (iteration 9, Task 3). Typed like a constructor
+ literal, returns the new row's id (Int), legal in statement and
+ expression position both. `select` stays a DbStub until Task 5. *)
+ | Insert of string * (string * expr) list
+ (* `delete ` (iteration 9b): removes the row a table-class value
+ names; an expression yielding the deleted id (restrict/trap surfaces
+ through the engine like any DB fault, catchable). *)
+ | Delete of expr
(* haxe-parity Task 2: one `${expr}` interpolation site, produced only
by the string-interpolation desugar (parser.ml) — never written
directly by a parse rule the way every other expr_kind is. Its
@@ -274,6 +286,31 @@ and expr_kind =
ename : string;
handler : stmt list;
}
+ (* iteration 9b: a language-integrated query. `from in
+ where * [group by into ] [order by [desc]] [take ]
+ select ` — lowered to a bytecode loop over engine cursor builtins,
+ never SQL text. A table-class value is its row id at runtime, so field
+ access on a range variable reads through the engine. Slice scope today:
+ from/where/order/take/select and group-by aggregation; join is later. *)
+ | Query of query
+
+and query_source =
+ | QTable of string (* a table class by name: `from e in Employee` *)
+ | QNav of expr (* a backlink/multi navigation: `from s in d.staff` *)
+
+and query = {
+ q_var : string;
+ q_src : query_source;
+ q_wheres : expr list;
+ (* group by ... into : present iff this is an aggregating
+ query. q_group_key is the whole grouped element (`e`), q_group_by the
+ key, q_gvar the group binding whose `.f` columns feed aggregates. *)
+ q_group : (string * expr) option; (* (gvar, key_expr) *)
+ q_order : (expr * bool) option; (* (key, desc?) *)
+ q_take : expr option;
+ q_select : expr;
+ q_pos : pos;
+}
(* ---- statements (Task 5) ---------------------------------------------
diff --git a/compiler/src/disasm.ml b/compiler/src/disasm.ml
index a30ca0c..ecd0f86 100644
--- a/compiler/src/disasm.ml
+++ b/compiler/src/disasm.ml
@@ -164,7 +164,7 @@ let dump (img : string) : string =
let line fmt = Buffer.add_string out (fmt ^ "\n") in
if u32 img 0 <> magic then raise (Bad "bad magic");
let ver = u32 img 4 in
- if ver <> 2 then raise (Bad (Printf.sprintf "unsupported version %d" ver));
+ if ver <> 3 then raise (Bad (Printf.sprintf "unsupported version %d" ver));
let coff = u32 img 8 and ccnt = u32 img 12 in
let koff = u32 img 16 and kcnt = u32 img 20 in
let ioff = u32 img 24 and icnt = u32 img 28 in
@@ -213,6 +213,14 @@ let dump (img : string) : string =
renders as keys, so a wrong one is worth seeing. *)
let names = List.init fcnt (fun j -> u32 img (!o + (j * 4))) in
o := !o + (fcnt * 12);
+ (* v3 index tail: walk past (the disassembly prints class shape, not
+ indexes — dump goldens stay byte-stable across the version bump) *)
+ let icnt = u32 img !o in
+ o := !o + 4;
+ for _ = 1 to icnt do
+ let ccnt = u32 img (!o + 4) in
+ o := !o + 8 + (ccnt * 4)
+ done;
let fields =
List.map2
(fun nmk k -> if nmk = 0xFFFFFFFF then k else Printf.sprintf "%s:%s" (kname nmk) k)
diff --git a/compiler/src/dump.ml b/compiler/src/dump.ml
index b2530f2..3cfe237 100644
--- a/compiler/src/dump.ml
+++ b/compiler/src/dump.ml
@@ -150,6 +150,7 @@ let rec field_ty_str : Ast.field_ty -> string = function
| Ast.Ref s -> Printf.sprintf "ref %s" s
| Ast.Multi s -> Printf.sprintf "multi %s" s
| Ast.Map (k, v) -> Printf.sprintf "map<%s, %s>" k v
+ | Ast.Backlink (c, f) -> Printf.sprintf "backlink %s.%s" c f
| Ast.Nullable t -> "?" ^ field_ty_str t
let param_str (p : Ast.param) : string = Printf.sprintf "%s%s: %s" (conv_str p.conv) p.name (field_ty_str p.ty)
@@ -222,6 +223,19 @@ let rec expr_str (e : Ast.expr) : string =
Printf.sprintf "%s { %s }" name
(String.concat ", "
(List.map (fun (fname, fval) -> Printf.sprintf "%s: %s" fname (expr_str fval)) fields))
+ | Ast.Insert (name, fields) ->
+ Printf.sprintf "INSERT %s { %s }" name
+ (String.concat ", "
+ (List.map (fun (fname, fval) -> Printf.sprintf "%s: %s" fname (expr_str fval)) fields))
+ | Ast.Query q ->
+ let src = match q.Ast.q_src with Ast.QTable cn -> cn | Ast.QNav e -> expr_str e in
+ Printf.sprintf "QUERY from %s in %s%s%s select %s" q.Ast.q_var src
+ (String.concat "" (List.map (fun w -> " where " ^ expr_str w) q.Ast.q_wheres))
+ (match q.Ast.q_group with
+ | Some (g, k) -> Printf.sprintf " group by %s into %s" (expr_str k) g
+ | None -> "")
+ (expr_str q.Ast.q_select)
+ | Ast.Delete t -> Printf.sprintf "DELETE %s" (expr_str t)
| Ast.DbStub toks -> Printf.sprintf "DB_STUB(%s)" (dbstub_tokens_str toks)
| Ast.Interp inner -> Printf.sprintf "INTERP(%s)" (expr_str inner)
| Ast.ListLit items -> Printf.sprintf "[%s]" (String.concat ", " (List.map expr_str items))
diff --git a/compiler/src/emit.ml b/compiler/src/emit.ml
index 70255cc..5dcf173 100644
--- a/compiler/src/emit.ml
+++ b/compiler/src/emit.ml
@@ -152,7 +152,7 @@ let stdlib_not_linked_code = Diag.emitter_prefix ^ "06"
============================================================ *)
let wob_magic = 0x31424F57 (* "WOB1" read as an LE u32 *)
-let wob_version = 2 (* v2: per-field class-table metadata *)
+let wob_version = 3 (* v3: v2 + per-class secondary-index metadata *)
let wob_hdr_size = 44
let wob_none = 0xFFFFFFFF
let k_int = 0
@@ -257,6 +257,7 @@ let b_map_val_at = 38
let b_multi_set = 39
let b_map_get_opt = 59
let b_text_copy = 60
+let b_db_insert = 61
(* json (runtime/src/json.c): encode takes the value's static kind as its
second argument, decode the class id to build as its second. *)
@@ -348,6 +349,16 @@ type clsrec = {
cr_gc : bool;
cr_fields : (string * Ast.field_ty) array;
cr_methods : string list; (* method names, declaration order *)
+ (* iteration 9 Task 4: (unique, column indices) per secondary index —
+ `@table(index: [a, b])` entries (non-unique, composite) plus one
+ unique single-column entry per `@unique` field. Serialized as the v3
+ class-record tail; the engine builds its runtime indexes from this. *)
+ cr_indexes : (bool * int array) list;
+ cr_is_table : bool; (* has @table — its instances are row ids (iteration 9b) *)
+ (* backlink fields (iteration 9b): name -> (source class, source field).
+ Virtual — not in cr_fields, no stored column; `d.staff` reads them by
+ probing the source class's index on the source field. *)
+ cr_backlinks : (string * (string * string)) list;
}
type ifacerec = {
@@ -772,6 +783,44 @@ let field_kind (p : pctx) (ft : Ast.field_ty) : int =
let class_of_name (p : pctx) (n : string) : int option = SM.find_opt n p.p_class_id
+(* iteration 9b: a @table class's instances are row ids, so field access on
+ one reads through the engine (DB_GET_FIELD) rather than GETF. *)
+let is_table_class (p : pctx) (cid : int) : bool =
+ cid >= 0 && cid < Array.length p.p_classes && p.p_classes.(cid).cr_is_table
+
+let b_str_lt = 67
+let b_db_update_field = 62
+let b_db_delete = 63
+let b_db_scan = 64
+let b_db_get_field = 65
+let b_db_probe = 66
+
+(* iteration 9b: `d.staff` where staff is `backlink Employee.dept` reads by
+ probing Employee's index on its `dept` column. Resolve to (source cid,
+ index number) — None if the source field is not a declared index (a
+ backlink without a backing index has no efficient read and is rejected). *)
+let backlink_target (p : pctx) (base_cid : int) (fname : string) : (int * int) option =
+ match List.assoc_opt fname p.p_classes.(base_cid).cr_backlinks with
+ | None -> None
+ | Some (src_class, src_field) -> (
+ match class_of_name p src_class with
+ | None -> None
+ | Some scid ->
+ let sc = p.p_classes.(scid) in
+ (* stored column index of the source field *)
+ let col = ref (-1) in
+ Array.iteri (fun i (n, _) -> if n = src_field then col := i) sc.cr_fields;
+ if !col < 0 then None
+ else
+ (* the index whose single column is that field *)
+ let rec find n = function
+ | [] -> None
+ | (_, cols) :: tl ->
+ if Array.length cols = 1 && cols.(0) = !col then Some (scid, n)
+ else find (n + 1) tl
+ in
+ find 0 sc.cr_indexes)
+
let field_of (p : pctx) (cid : int) (fname : string) : (int * Ast.field_ty) option =
let fs = p.p_classes.(cid).cr_fields in
let rec go i = if i >= Array.length fs then None else
@@ -941,6 +990,22 @@ let variant_tag_value (p : pctx) (u : Types.union_info) (vi : Types.variant_info
| None -> 0 (* unreachable: pass 1 registers every payload-union variant *)
else vi.Types.vi_tag
+(* iteration 9b: a query's element type, as the name a `Multi` carries.
+ `select x` yields the source class (a row id typed as the class);
+ `select x.field` yields that field's type; anything else falls back to
+ Int (the slice's shapes are these two). *)
+let query_elem_scalar (p : pctx) (q : Ast.query) ~(src : string) : string =
+ match q.Ast.q_select.Ast.kind with
+ | Ast.Ident v when v = q.Ast.q_var -> src (* select the whole row: element = source class *)
+ | Ast.Field ({ Ast.kind = Ast.Ident v; _ }, fname) when v = q.Ast.q_var -> (
+ match class_of_name p src with
+ | Some cid -> (
+ match field_of p cid fname with
+ | Some (_, ty) -> ( match unwrap ty with Scalar n -> n | _ -> "Int")
+ | None -> "Int")
+ | None -> "Int")
+ | _ -> "Int"
+
let rec ty_of_expr (p : pctx) (f : fstate) (e : Ast.expr) : Ast.field_ty option =
match e.kind with
| IntLit _ -> Some (Scalar "Int")
@@ -970,10 +1035,17 @@ let rec ty_of_expr (p : pctx) (f : fstate) (e : Ast.expr) : Ast.field_ty option
| Field (base, fname) -> (
match ty_of_expr p f base with
| Some bt -> (
- match unwrap bt with
+ (* a `ref C` navigates into C: the target is a table row id *)
+ match (match unwrap bt with Ref c -> Scalar c | other -> other) with
| Scalar cn -> (
match class_of_name p cn with
- | Some cid -> ( match field_of p cid fname with Some (_, t) -> Some t | None -> None)
+ | Some cid -> (
+ match field_of p cid fname with
+ | Some (_, t) -> Some t
+ | None -> (
+ match List.assoc_opt fname p.p_classes.(cid).cr_backlinks with
+ | Some (sc, _) -> Some (Multi sc)
+ | None -> None))
| None -> None)
| _ -> None)
| None -> None)
@@ -1054,6 +1126,16 @@ let rec ty_of_expr (p : pctx) (f : fstate) (e : Ast.expr) : Ast.field_ty option
| Eq | Ne | Lt | Le | Gt | Ge | And | Or -> Some (Scalar "Bool")
| Add | Sub | Mul | Div | Mod -> ( match ty_of_expr p f l with Some t -> Some t | None -> Some (Scalar "Int")))
| Ctor (cn, _) -> Some (Scalar cn)
+ | Insert _ -> Some (Scalar "Int")
+ | Delete _ -> Some (Scalar "Int")
+ | Query q ->
+ let src =
+ match q.Ast.q_src with
+ | Ast.QTable cn -> cn
+ | Ast.QNav nav -> (
+ match ty_of_expr p f nav with Some t -> (match unwrap t with Multi c -> c | Scalar c -> c | _ -> "") | None -> "")
+ in
+ Some (Multi (query_elem_scalar p q ~src))
| Interp _ -> Some (Scalar "Text")
| DbStub _ -> None
| Switch (subject, arms) -> (
@@ -1344,7 +1426,8 @@ let field_class_meta (p : pctx) (ty : Ast.field_ty) : int =
match name_of (Ast.Scalar e) with
| Some n -> ( match class_of_name p n with Some cid -> cid | None -> wob_none)
| None -> wob_none)
- | Ast.Ref _ | Ast.Nullable _ -> wob_none
+ | Ast.Ref n -> ( match class_of_name p n with Some cid -> cid | None -> wob_none)
+ | Ast.Backlink _ | Ast.Nullable _ -> wob_none
let field_elem_meta (p : pctx) (ty : Ast.field_ty) : int =
match unwrap ty with
@@ -1583,9 +1666,22 @@ let rec emit_expr (p : pctx) (f : fstate) (v : views) ~(dst : int) ?expected (e
| Field (base, fname) -> (
match ty_of_expr p f base with
| Some bt -> (
- match unwrap bt with
+ match (match unwrap bt with Ref c -> Scalar c | other -> other) with
| Scalar cn -> (
match class_of_name p cn with
+ | Some cid when is_table_class p cid && backlink_target p cid fname <> None -> (
+ (* `d.staff`: probe the source class's index for rows referencing
+ this row's id. Window: [class, index, key(=base id)]. *)
+ match backlink_target p cid fname with
+ | Some (scid, ino) ->
+ let b = emit_operand p f v base in
+ let w = alloc_temps p f e.pos 3 in
+ put f (ins_abx op_loadk w (check_bx p f e.pos "constant" (const_int p scid)));
+ put f (ins_abx op_loadk (w + 1) (check_bx p f e.pos "constant" (const_int p ino)));
+ put f (ins_abc op_move (w + 2) b 0);
+ sync_mask p f v e.id;
+ put f (ins_abc op_builtin dst w b_db_probe)
+ | None -> ())
| Some cid -> (
match field_of p cid fname with
| Some (idx, _) ->
@@ -1603,7 +1699,17 @@ let rec emit_expr (p : pctx) (f : fstate) (v : views) ~(dst : int) ?expected (e
f.f_stmt_drops <- g :: f.f_stmt_drops;
f.f_esc_drops <- g :: f.f_esc_drops
end;
- put f (ins_abc op_getf dst b (check_field_idx p f e.pos idx))
+ if is_table_class p cid then begin
+ (* a table-class value is its row id; read the column from the
+ engine. Window: [class-id, id, field-idx]. *)
+ let w = alloc_temps p f e.pos 3 in
+ put f (ins_abx op_loadk w (check_bx p f e.pos "constant" (const_int p cid)));
+ put f (ins_abc op_move (w + 1) b 0);
+ put f (ins_abx op_loadk (w + 2) (check_bx p f e.pos "constant" (const_int p idx)));
+ sync_mask p f v e.id;
+ put f (ins_abc op_builtin dst w b_db_get_field)
+ end
+ else put f (ins_abc op_getf dst b (check_field_idx p f e.pos idx))
| None ->
err p ~code:cannot_lower_code ~file:f.f_file ~pos:e.pos
~message:(Printf.sprintf "`%s` has no field `%s`" cn fname);
@@ -1654,6 +1760,43 @@ let rec emit_expr (p : pctx) (f : fstate) (v : views) ~(dst : int) ?expected (e
put f (ins_abc op_neg dst b 0)
| Binary (op, l, r) -> emit_binary p f v ~dst op l r
| Ctor (cn, fields) -> emit_ctor p f v ~dst e cn fields
+ | Insert (cn, fields) -> emit_insert p f v ~dst e cn fields
+ | Delete target -> (
+ match ty_of_expr p f target with
+ | Some bt -> (
+ match (match unwrap bt with Ref c -> Scalar c | o -> o) with
+ | Scalar cn -> (
+ match class_of_name p cn with
+ | Some cid when is_table_class p cid ->
+ (* reserve dst past the window: in tail position dst == the first
+ window reg, and moving the id into dst would clobber the class
+ id — the disassembly-caught bug *)
+ let outer = f.f_temp in
+ if f.f_temp <= dst then f.f_temp <- dst + 1;
+ let w = alloc_temps p f e.pos 2 in
+ put f (ins_abx op_loadk w (check_bx p f e.pos "constant" (const_int p cid)));
+ let save = f.f_temp in
+ emit_expr p f v ~dst:(w + 1) target;
+ f.f_temp <- save;
+ (* keep the id so `delete x` can be used as an expression *)
+ put f (ins_abc op_move dst (w + 1) 0);
+ sync_mask p f v e.id;
+ f.f_cur_line <- e.pos.line;
+ put f (ins_abc op_builtin w w b_db_delete);
+ f.f_temp <- outer
+ | _ ->
+ err p ~code:cannot_lower_code ~file:f.f_file ~pos:e.pos
+ ~message:"`delete` target is not a table row";
+ put f (ins_abx op_loadk dst (const_int p 0)))
+ | _ ->
+ err p ~code:cannot_lower_code ~file:f.f_file ~pos:e.pos
+ ~message:"`delete` target is not a table row";
+ put f (ins_abx op_loadk dst (const_int p 0)))
+ | None ->
+ err p ~code:cannot_lower_code ~file:f.f_file ~pos:e.pos
+ ~message:"cannot resolve the `delete` target's type";
+ put f (ins_abx op_loadk dst (const_int p 0)))
+ | Query q -> emit_query p f v ~dst e q
| Interp inner -> (
(* haxe-parity Task 2: the type-directed half of the interpolation
desugar (parser.ml's own doc comment on Ast.Interp) — a Text
@@ -2347,6 +2490,324 @@ and emit_ctor (p : pctx) (f : fstate) (v : views) ~(dst : int) (e : Ast.expr) (c
ci.Types.fields);
f.f_temp <- outer
+(* iteration 9 Task 3: `insert Class { ... }` lowers to one DB_INSERT
+ builtin whose window is [class-id const, then one slot per DECLARED
+ field in declaration order] — the executor walks the class table's
+ kinds, so slot order must be the table's, not the literal's. A field
+ the literal omits gets its default (same emit_default_value the ctor
+ uses) or, for a `?` field, its kind's own nil (WO_NIL_SCALAR for a
+ nullable scalar, the zero word otherwise). The engine COPIES every
+ value at the row API, so after the builtin every freshly built
+ argument is still this frame's to drop — same reap as push/set. *)
+and emit_query (p : pctx) (f : fstate) (v : views) ~(dst : int) (e : Ast.expr)
+ (q : Ast.query) : unit =
+ (* iteration 9b slice: from/where/select over a table scan. group/order/
+ take/navigation are diagnosed in types.ml, so a written image never
+ reaches this with them set. Lowered to an ordinary bytecode loop over
+ DB_SCAN's materialized id list — no plan tree, no text. *)
+ let cn =
+ match q.Ast.q_src with
+ | Ast.QTable cn -> cn
+ | Ast.QNav nav -> (
+ (* the source's element type is the range var's class *)
+ match ty_of_expr p f nav with
+ | Some t -> ( match unwrap t with Multi c -> c | Scalar c -> c | _ -> "")
+ | None -> "")
+ in
+ match class_of_name p cn with
+ | None ->
+ err p ~code:cannot_lower_code ~file:f.f_file ~pos:e.pos
+ ~message:(Printf.sprintf "query over `%s`, which is not a declared table class" cn);
+ put f (ins_abx op_loadk dst (const_int p 0))
+ | Some cid ->
+ let elem_name = query_elem_scalar p q ~src:cn in
+ let elem = Scalar elem_name in
+ (* a table-class element is a row ID (a scalar), not a heap pointer — so
+ the result container is SCALAR-kinded even though the element TYPES as
+ the class; getting this wrong drops an id as a pointer (ASan SEGV) *)
+ let elem_kind =
+ match class_of_name p elem_name with
+ | Some ecid when is_table_class p ecid -> 0 (* WO_K_SCALAR *)
+ | _ -> field_kind p elem
+ in
+ (* reserve dst past the loop's working registers (same guard emit_ctor
+ uses): dst holds the result multi every push writes into *)
+ let outer = f.f_temp in
+ if f.f_temp <= dst then f.f_temp <- dst + 1;
+ (* loop-carried registers, allocated once above dst, never reset *)
+ let scan = alloc_temp p f e.pos in
+ let idx = alloc_temp p f e.pos in
+ let len = alloc_temp p f e.pos in
+ let idreg = alloc_temp p f e.pos in
+ let body_base = f.f_temp in
+ (* scan -> multi of ids; result multi -> dst *)
+ sync_mask p f v e.id;
+ f.f_cur_line <- e.pos.line;
+ (match q.Ast.q_src with
+ | Ast.QTable _ ->
+ put f (ins_abx op_loadk scan (check_bx p f e.pos "constant" (const_int p cid)));
+ put f (ins_abc op_builtin scan scan b_db_scan)
+ | Ast.QNav nav ->
+ (* the navigation (a backlink) already yields a multi of source ids *)
+ let save = f.f_temp in
+ f.f_temp <- scan + 1;
+ emit_expr p f v ~dst:scan nav;
+ f.f_temp <- save);
+ put f (ins_abc op_builtin dst elem_kind b_multi_new);
+ put f (ins_abc op_builtin len scan b_len);
+ put f (ins_abx op_loadk idx (check_bx p f e.pos "constant" (const_int p 0)));
+ (* bind the range var to the current id (typed as the class), so field
+ access inside where/select routes through DB_GET_FIELD *)
+ let saved_env = f.f_env in
+ f.f_env <- (q.Ast.q_var, (idreg, Scalar cn)) :: f.f_env;
+ ignore body_base;
+ let top = here f in
+ f.f_temp <- body_base;
+ let tc = alloc_temp p f e.pos in
+ put f (ins_abc op_lt tc idx len);
+ let jz_exit = here f in
+ put f (ins_asbx op_jz tc 0);
+ (* id = multi_get(scan, idx) *)
+ let w = alloc_temps p f e.pos 2 in
+ put f (ins_abc op_move w scan 0);
+ put f (ins_abc op_move (w + 1) idx 0);
+ put f (ins_abc op_builtin idreg w b_multi_get);
+ (* where guards: any false skips the push *)
+ let skips = ref [] in
+ List.iter
+ (fun w_expr ->
+ let save = f.f_temp in
+ let wr = emit_operand p f v w_expr in
+ skips := here f :: !skips;
+ put f (ins_asbx op_jz wr 0);
+ f.f_temp <- save)
+ q.Ast.q_wheres;
+ (* select -> push into dst (copying a Text element the container owns) *)
+ let save = f.f_temp in
+ let sel = alloc_temp p f e.pos in
+ emit_expr p f v ~dst:sel q.Ast.q_select;
+ if elem_kind = 3 then put f (ins_abc op_builtin sel sel b_text_copy);
+ let pw = alloc_temps p f e.pos 2 in
+ put f (ins_abc op_move pw dst 0);
+ put f (ins_abc op_move (pw + 1) sel 0);
+ put f (ins_abc op_builtin pw pw b_multi_push);
+ f.f_temp <- save;
+ (* skip target: increment and loop *)
+ let cont = here f in
+ List.iter (fun pc -> patch_jump p f ~file:f.f_file ~pos:e.pos pc cont) !skips;
+ f.f_temp <- body_base;
+ let one = alloc_temp p f e.pos in
+ put f (ins_abx op_loadk one (check_bx p f e.pos "constant" (const_int p 1)));
+ put f (ins_abc op_add idx idx one);
+ let back = here f in
+ put f (ins_asbx op_jmp 0 0);
+ patch_jump p f ~file:f.f_file ~pos:e.pos back top;
+ let exit_pc = here f in
+ patch_jump p f ~file:f.f_file ~pos:e.pos jz_exit exit_pc;
+ f.f_env <- saved_env;
+ (* the scan's id list was this query's own, dropped now *)
+ put f (ins_abc op_drop scan 0 0);
+ (* ---- order by (whole-row selection sort) ----------------------------
+ Elements of dst are row ids; the key re-reads a field through the
+ range var. Selection sort is O(n^2) but the result sets here are
+ small and this is KISS by design (no cost planner). Only the
+ whole-row + field-key shape is supported; grouped/projection ordering
+ lands with group-by. *)
+ (match q.Ast.q_order with
+ | Some (key, desc) ->
+ f.f_temp <- body_base;
+ let n = alloc_temp p f e.pos in
+ put f (ins_abc op_builtin n dst b_count);
+ let i = alloc_temp p f e.pos in
+ let j = alloc_temp p f e.pos in
+ let best = alloc_temp p f e.pos in
+ let elem_j = alloc_temp p f e.pos in
+ let elem_b = alloc_temp p f e.pos in
+ let sort_scratch = f.f_temp in
+ put f (ins_abx op_loadk i (check_bx p f e.pos "constant" (const_int p 0)));
+ let oi = here f in (* outer: while i < n *)
+ let oc = alloc_temp p f e.pos in
+ put f (ins_abc op_lt oc i n);
+ let ojz = here f in
+ put f (ins_asbx op_jz oc 0);
+ put f (ins_abc op_move best i 0);
+ let oneA = alloc_temp p f e.pos in
+ put f (ins_abx op_loadk oneA (check_bx p f e.pos "constant" (const_int p 1)));
+ put f (ins_abc op_add j i oneA);
+ let ij = here f in (* inner: while j < n *)
+ let ic = alloc_temp p f e.pos in
+ put f (ins_abc op_lt ic j n);
+ let ijz = here f in
+ put f (ins_asbx op_jz ic 0);
+ (* elem_j = multi_get(dst,j); elem_b = multi_get(dst,best) *)
+ let gw = alloc_temps p f e.pos 2 in
+ put f (ins_abc op_move gw dst 0);
+ put f (ins_abc op_move (gw + 1) j 0);
+ put f (ins_abc op_builtin elem_j gw b_multi_get);
+ put f (ins_abc op_move (gw + 1) best 0);
+ put f (ins_abc op_builtin elem_b gw b_multi_get);
+ (* keys: bind range var to elem_j / elem_b, eval key expr *)
+ let saved_env2 = f.f_env in
+ f.f_temp <- sort_scratch;
+ f.f_env <- (q.Ast.q_var, (elem_j, Scalar cn)) :: saved_env2;
+ (* key kind must be read with the range var BOUND — else ty_of_expr of
+ `x.name` sees x unbound, returns None, and a Text key silently falls
+ to the pointer-comparing op_lt (the wrong-order bug) *)
+ let key_is_text =
+ match ty_of_expr p f key with Some t -> field_kind p t = 3 | None -> false
+ in
+ let kj = alloc_temp p f e.pos in
+ emit_expr p f v ~dst:kj key;
+ f.f_env <- (q.Ast.q_var, (elem_b, Scalar cn)) :: saved_env2;
+ let kb = alloc_temp p f e.pos in
+ emit_expr p f v ~dst:kb key;
+ f.f_env <- saved_env2;
+ (* cmp: for asc, kj < kb -> best=j; for desc, kj > kb (== kb < kj). *)
+ let cmp = alloc_temp p f e.pos in
+ let lt a b =
+ if key_is_text then begin
+ let save = f.f_temp in
+ let w = alloc_temps p f e.pos 2 in
+ put f (ins_abc op_move w a 0);
+ put f (ins_abc op_move (w + 1) b 0);
+ put f (ins_abc op_builtin cmp w b_str_lt);
+ f.f_temp <- save
+ end
+ else put f (ins_abc op_lt cmp a b)
+ in
+ if desc then lt kb kj else lt kj kb;
+ let cjz = here f in
+ put f (ins_asbx op_jz cmp 0);
+ put f (ins_abc op_move best j 0);
+ let after = here f in
+ patch_jump p f ~file:f.f_file ~pos:e.pos cjz after;
+ f.f_temp <- sort_scratch;
+ let oneB = alloc_temp p f e.pos in
+ put f (ins_abx op_loadk oneB (check_bx p f e.pos "constant" (const_int p 1)));
+ put f (ins_abc op_add j j oneB);
+ let iback = here f in
+ put f (ins_asbx op_jmp 0 0);
+ patch_jump p f ~file:f.f_file ~pos:e.pos iback ij;
+ let iexit = here f in
+ patch_jump p f ~file:f.f_file ~pos:e.pos ijz iexit;
+ (* swap dst[i], dst[best]: read both, multi_set both *)
+ f.f_temp <- sort_scratch;
+ let vi = alloc_temp p f e.pos in
+ let vb = alloc_temp p f e.pos in
+ let sw = alloc_temps p f e.pos 3 in
+ put f (ins_abc op_move sw dst 0);
+ put f (ins_abc op_move (sw + 1) i 0);
+ put f (ins_abc op_builtin vi sw b_multi_get);
+ put f (ins_abc op_move (sw + 1) best 0);
+ put f (ins_abc op_builtin vb sw b_multi_get);
+ (* dst[i] = vb *)
+ put f (ins_abc op_move sw dst 0);
+ put f (ins_abc op_move (sw + 1) i 0);
+ put f (ins_abc op_move (sw + 2) vb 0);
+ put f (ins_abc op_builtin sw sw b_multi_set);
+ (* dst[best] = vi *)
+ put f (ins_abc op_move sw dst 0);
+ put f (ins_abc op_move (sw + 1) best 0);
+ put f (ins_abc op_move (sw + 2) vi 0);
+ put f (ins_abc op_builtin sw sw b_multi_set);
+ f.f_temp <- sort_scratch;
+ let oneC = alloc_temp p f e.pos in
+ put f (ins_abx op_loadk oneC (check_bx p f e.pos "constant" (const_int p 1)));
+ put f (ins_abc op_add i i oneC);
+ let oback = here f in
+ put f (ins_asbx op_jmp 0 0);
+ patch_jump p f ~file:f.f_file ~pos:e.pos oback oi;
+ let oexit = here f in
+ patch_jump p f ~file:f.f_file ~pos:e.pos ojz oexit
+ | None -> ());
+ (* ---- take N: slice dst to [0, N) --------------------------------- *)
+ (match q.Ast.q_take with
+ | Some tk ->
+ f.f_temp <- body_base;
+ let nreg = alloc_temp p f e.pos in
+ emit_expr p f v ~dst:nreg tk;
+ (* clamp N to count(dst) so slice never runs past the end *)
+ let cnt = alloc_temp p f e.pos in
+ put f (ins_abc op_builtin cnt dst b_count);
+ let over = alloc_temp p f e.pos in
+ put f (ins_abc op_lt over cnt nreg); (* count < N ? use count *)
+ let jz2 = here f in
+ put f (ins_asbx op_jz over 0);
+ put f (ins_abc op_move nreg cnt 0);
+ let aft = here f in
+ patch_jump p f ~file:f.f_file ~pos:e.pos jz2 aft;
+ let sw = alloc_temps p f e.pos 3 in
+ let zero = alloc_temp p f e.pos in
+ put f (ins_abx op_loadk zero (check_bx p f e.pos "constant" (const_int p 0)));
+ put f (ins_abc op_move sw dst 0);
+ put f (ins_abc op_move (sw + 1) zero 0);
+ put f (ins_abc op_move (sw + 2) nreg 0);
+ let sliced = alloc_temp p f e.pos in
+ put f (ins_abc op_builtin sliced sw b_slice);
+ put f (ins_abc op_drop dst 0 0); (* the pre-slice multi is discarded *)
+ put f (ins_abc op_move dst sliced 0)
+ | None -> ());
+ f.f_temp <- outer
+
+and emit_insert (p : pctx) (f : fstate) (v : views) ~(dst : int) (e : Ast.expr) (cn : string)
+ (fields : (string * Ast.expr) list) : unit =
+ match class_of_name p cn with
+ | None ->
+ err p ~code:cannot_lower_code ~file:f.f_file ~pos:e.pos
+ ~message:(Printf.sprintf "insert into `%s`, which is not a declared class" cn);
+ put f (ins_abx op_loadk dst (const_int p 0))
+ | Some cid ->
+ let fcnt = Array.length p.p_classes.(cid).cr_fields in
+ let base = alloc_temps p f e.pos (fcnt + 1) in
+ put f (ins_abx op_loadk base (check_bx p f e.pos "constant" (const_int p cid)));
+ (* every field the literal names lands in ITS declared slot *)
+ List.iter
+ (fun ((fname : string), (fe : Ast.expr)) ->
+ match field_of p cid fname with
+ | None ->
+ err p ~code:cannot_lower_code ~file:f.f_file ~pos:e.pos
+ ~message:(Printf.sprintf "`%s` has no field `%s`" cn fname)
+ | Some (idx, fty) ->
+ let save = f.f_temp in
+ emit_expr p f v ~dst:(base + 1 + idx) ~expected:fty fe;
+ f.f_temp <- save)
+ fields;
+ (* omitted fields: declared default, else the kind's own nil *)
+ let provided = List.map fst fields in
+ (match Types.StringMap.find_opt cn p.p_syms.Types.classes with
+ | None -> ()
+ | Some (ci : Types.class_info) ->
+ List.iter
+ (fun (fname, fty, fdefault, _) ->
+ if not (List.mem fname provided) then
+ match field_of p cid fname with
+ | None -> ()
+ | Some (idx, dfty) -> (
+ match fdefault with
+ | Some d ->
+ let save = f.f_temp in
+ emit_default_value p f ~dst:(base + 1 + idx) ~fty:dfty ~pos:e.pos d;
+ f.f_temp <- save
+ | None ->
+ let nil_word =
+ if is_nullable_scalar p fty then const_int p nil_scalar_word
+ else const_int p 0
+ in
+ put f (ins_abx op_loadk (base + 1 + idx) (check_bx p f e.pos "constant" nil_word))))
+ ci.Types.fields);
+ sync_mask p f v e.id;
+ f.f_cur_line <- e.pos.line;
+ put f (ins_abc op_builtin dst base b_db_insert);
+ (* the engine copied: fresh argument values die here *)
+ List.iter
+ (fun ((fname : string), (fe : Ast.expr)) ->
+ match field_of p cid fname with
+ | None -> ()
+ | Some (idx, _) ->
+ drop_fresh_owned ~keep:dst p f (base + 1 + idx) fe;
+ drop_fresh_text ~keep:dst p f (base + 1 + idx) fe)
+ fields
+
(* The default expressions the emitter can lower (haxe-parity Task 4):
the literal shapes the sample's own typedefs use — Int (optionally
negated), Text, Bool, `now()` (parse_default_expr's own recognized
@@ -3224,6 +3685,27 @@ and emit_assign (p : pctx) (f : fstate) (v : views) (s : Ast.stmt) (target : Ast
| None ->
err p ~code:cannot_lower_code ~file:f.f_file ~pos:target.pos
~message:(Printf.sprintf "assignment into `%s`, which is not a declared class" cn)
+ | Some cid when is_table_class p cid -> (
+ (* iteration 9b: `e.salary = v` where e is a table row updates the
+ engine (DB_UPDATE_FIELD: class, id, field, value) — the row's
+ own indexes are maintained at the choke point *)
+ match field_of p cid fname with
+ | None ->
+ err p ~code:cannot_lower_code ~file:f.f_file ~pos:target.pos
+ ~message:(Printf.sprintf "`%s` has no field `%s`" cn fname)
+ | Some (idx, fty) ->
+ let w = alloc_temps p f target.pos 4 in
+ put f (ins_abx op_loadk w (check_bx p f target.pos "constant" (const_int p cid)));
+ let save = f.f_temp in
+ emit_expr p f v ~dst:(w + 1) base;
+ f.f_temp <- save;
+ put f (ins_abx op_loadk (w + 2) (check_bx p f target.pos "constant" (const_int p idx)));
+ let save = f.f_temp in
+ emit_expr p f v ~dst:(w + 3) ~expected:fty value;
+ f.f_temp <- save;
+ sync_mask p f v s.s_id;
+ f.f_cur_line <- s.s_pos.line;
+ put f (ins_abc op_builtin w w b_db_update_field))
| Some cid -> (
match field_of p cid fname with
| None ->
@@ -3852,6 +4334,18 @@ let emit ~(syms : Types.symbols) ~(module_of : string -> string)
~(module_syms : (string, Types.symbols) Hashtbl.t) (coll : Diag.Collector.t) (units : input list) :
string =
let colliding = compute_colliding_fn_names ~module_of units in
+ (* iteration 9 Task 4: index-declaration problems found while building
+ clsrecs — reported once a file/pos-bearing context exists below *)
+ let index_col_err : (Ast.pos * string) option ref = ref None in
+ let index_err_file = ref "" in
+ let ref_index_of_name (fnames : string list) (n : string) : int =
+ let rec go i = function
+ | [] -> 0 (* unknown column: the caller records the diagnostic *)
+ | x :: tl -> if x = n then i else go (i + 1) tl
+ in
+ go 0 fnames
+ in
+ let p_syms_for_indexes = syms in
(* ---- pass 1: declarations, in discovery then declaration order ---- *)
let classes = ref [] and class_id = ref SM.empty and nclasses = ref 0 in
let ifaces = ref [] and iface_id = ref SM.empty and nifaces = ref 0 and nslots = ref 0 in
@@ -3890,6 +4384,7 @@ let emit ~(syms : Types.symbols) ~(module_of : string -> string)
(function
| Ast.Class (c : Ast.class_decl) ->
if not (SM.mem c.name !class_id) then begin
+ (if !index_col_err = None then index_err_file := u.file);
let shape = if c.is_record then Some (record_shape_key c) else None in
let alias_of =
match shape with Some key -> Hashtbl.find_opt record_shape key | None -> None
@@ -3907,10 +4402,77 @@ let emit ~(syms : Types.symbols) ~(module_of : string -> string)
| Some key -> Hashtbl.replace record_shape key cid
| None -> ());
classes :=
- { cr_name = c.name; cr_gc = c.is_gc;
- cr_fields =
- Array.of_list (List.map (fun (fl : Ast.field) -> (fl.name, fl.ty)) c.fields);
- cr_methods = List.map (fun (m : Ast.method_decl) -> m.name) c.methods }
+ (let fnames =
+ List.filter_map
+ (fun (fl : Ast.field) ->
+ match fl.Ast.ty with Ast.Backlink _ -> None | _ -> Some fl.Ast.name)
+ c.fields
+ in
+ let col_of n = ref_index_of_name fnames n in
+ let is_indexable (fl : Ast.field) =
+ match Types.wob_kind_of_typ p_syms_for_indexes (Types.typ_of_field_ty (unwrap fl.Ast.ty)) with
+ | Types.WO_K_SCALAR | Types.WO_K_TEXT -> true
+ | _ -> false
+ in
+ let table_indexes =
+ match c.Ast.table with
+ | None -> []
+ | Some cfg ->
+ List.map
+ (fun cols -> (false, Array.of_list (List.map col_of cols)))
+ cfg.Ast.indexes
+ in
+ let unique_indexes =
+ List.concat_map
+ (fun (fl : Ast.field) ->
+ if List.mem "unique" fl.Ast.annotations then begin
+ if not (is_indexable fl) then
+ index_col_err := Some (c.Ast.pos, Printf.sprintf
+ "`@unique` on `%s.%s`: only scalar and Text fields can be indexed"
+ c.Ast.name fl.Ast.name);
+ [ (true, [| col_of fl.Ast.name |]) ]
+ end
+ else [])
+ c.fields
+ in
+ (match c.Ast.table with
+ | Some cfg ->
+ List.iter
+ (fun cols ->
+ List.iter
+ (fun cn ->
+ match List.find_opt (fun (fl : Ast.field) -> fl.Ast.name = cn) c.fields with
+ | None ->
+ index_col_err := Some (c.Ast.pos, Printf.sprintf
+ "`@table(index: ...)` on `%s` names `%s`, which is not a field"
+ c.Ast.name cn)
+ | Some fl ->
+ if not (is_indexable fl) then
+ index_col_err := Some (c.Ast.pos, Printf.sprintf
+ "`@table(index: ...)` on `%s`: `%s` is not a scalar or Text field"
+ c.Ast.name cn))
+ cols)
+ cfg.Ast.indexes
+ | None -> ());
+ { cr_name = c.name; cr_gc = c.is_gc;
+ cr_fields =
+ Array.of_list
+ (List.filter_map
+ (fun (fl : Ast.field) ->
+ match fl.Ast.ty with
+ | Ast.Backlink _ -> None (* virtual: no stored column *)
+ | _ -> Some (fl.Ast.name, fl.Ast.ty))
+ c.fields);
+ cr_methods = List.map (fun (m : Ast.method_decl) -> m.name) c.methods;
+ cr_indexes = table_indexes @ unique_indexes;
+ cr_is_table = (c.Ast.table <> None);
+ cr_backlinks =
+ List.filter_map
+ (fun (fl : Ast.field) ->
+ match fl.Ast.ty with
+ | Ast.Backlink (sc, sf) -> Some (fl.Ast.name, (sc, sf))
+ | _ -> None)
+ c.fields })
:: !classes
end
| Ast.Union (ud : Ast.union_decl) ->
@@ -3930,7 +4492,8 @@ let emit ~(syms : Types.symbols) ~(module_of : string -> string)
class_id := SM.add key cid !class_id;
incr nclasses;
classes :=
- { cr_name = key; cr_gc = false;
+ { cr_name = key; cr_gc = false; cr_indexes = []; cr_is_table = false;
+ cr_backlinks = [];
cr_fields = Array.of_list vd.Ast.v_fields;
cr_methods = [] }
:: !classes
@@ -3973,11 +4536,18 @@ let emit ~(syms : Types.symbols) ~(module_of : string -> string)
class_id := SM.add name cid !class_id;
incr nclasses;
classes :=
- { cr_name = name; cr_gc = false; cr_fields = Array.of_list fields; cr_methods = [] }
+ { cr_name = name; cr_gc = false; cr_fields = Array.of_list fields; cr_methods = [];
+ cr_indexes = []; cr_is_table = false; cr_backlinks = [] }
:: !classes
end)
Types.predeclared_records;
let class_id = !class_id in
+ (match !index_col_err with
+ | Some (pos, msg) ->
+ Diag.Collector.add coll
+ (Diag.error ~code:cannot_lower_code ~file:!index_err_file ~line:pos.Ast.line
+ ~col:pos.Ast.col ~message:msg ())
+ | None -> ());
let p_classes = Array.of_list (List.rev !classes) in
let p_ifaces = Array.of_list (List.rev !ifaces) in
List.iter
@@ -4149,7 +4719,16 @@ let emit ~(syms : Types.symbols) ~(module_of : string -> string)
needs when decode creates one. *)
Array.iter (fun kidx -> Buf.u32 cls kidx) class_field_names.(cid);
Array.iter (fun (_, ty) -> Buf.u32 cls (field_class_meta p ty)) c.cr_fields;
- Array.iter (fun (_, ty) -> Buf.u32 cls (field_elem_meta p ty)) c.cr_fields)
+ Array.iter (fun (_, ty) -> Buf.u32 cls (field_elem_meta p ty)) c.cr_fields;
+ (* v3 tail (iteration 9 Task 4): the class's secondary indexes —
+ index_cnt, then per index: flags (bit0 unique), col_cnt, cols *)
+ Buf.u32 cls (List.length c.cr_indexes);
+ List.iter
+ (fun (uniq, cols) ->
+ Buf.u32 cls (if uniq then 1 else 0);
+ Buf.u32 cls (Array.length cols);
+ Array.iter (fun ci -> Buf.u32 cls ci) cols)
+ c.cr_indexes)
p_classes;
let ifs = Buf.create () in
Array.iteri
diff --git a/compiler/src/owner.ml b/compiler/src/owner.ml
index c268bf3..c61e1db 100644
--- a/compiler/src/owner.ml
+++ b/compiler/src/owner.ml
@@ -439,6 +439,7 @@ let oclass_of (ctx : ctx) (ft : Ast.field_ty) : oclass =
| Some u -> if u.Types.u_has_payload then Owned else Copy
| None -> Copy (* unknown type: WO-E225 already reported by types.ml *))
| Ref _ -> Copy
+ | Backlink _ -> Copy (* a virtual collection of row ids read on demand *)
| Multi _ | Map _ -> Owned
| Nullable _ -> Copy (* unreachable: unwrapped above *)
@@ -560,6 +561,9 @@ let rec expr_ty (ctx : ctx) (e : Ast.expr) : Ast.field_ty option =
| Binary (Concat, _, _) -> Some (Scalar "Text")
| Binary _ -> None (* arithmetic/comparison: Copy either way *)
| Ctor (cn, _) -> Some (Scalar cn)
+ | Insert _ -> Some (Scalar "Int") (* the new row's id — Copy, nothing to drop *)
+ | Query _ -> Some (Multi "Int") (* a query yields a fresh multi of ids — owned *)
+ | Delete _ -> Some (Scalar "Int") (* the deleted id — Copy *)
| Interp _ -> Some (Scalar "Text") (* an interpolation always produces Text *)
| DbStub _ -> None
| Switch (subject, arms) ->
@@ -1148,6 +1152,14 @@ let rec read_expr (ctx : ctx) (e : Ast.expr) : unit =
read_place_parts ctx e
| Call (callee, args) -> analyze_call ctx e callee args
| Ctor (cn, fields) -> analyze_ctor ctx cn fields
+ | Insert (_, fields) ->
+ (* iteration 9 Task 3: the engine copies every field value at the row
+ API (the two-worlds bulkhead), so an insert BORROWS its values —
+ no transfer, no E304, the source keeps what it had. Trap-capable
+ (unique violations arrive with Task 4), so the drop map is
+ recorded exactly like DbStub's. *)
+ List.iter (fun (_, fe) -> read_expr ctx fe) fields;
+ record_drop ctx ~node:e.id ~pos:e.pos ~kind:DLiveMask ~items:(mask_items (live_holders ctx))
| Unary (_, o) -> read_expr ctx o
| Binary (_, a, b) ->
read_expr ctx a;
@@ -1170,6 +1182,20 @@ let rec read_expr (ctx : ctx) (e : Ast.expr) : unit =
| DbStub _ ->
(* trap-capable: the frame needs its drop map here *)
record_drop ctx ~node:e.id ~pos:e.pos ~kind:DLiveMask ~items:(mask_items (live_holders ctx))
+ | Delete t ->
+ read_expr ctx t;
+ record_drop ctx ~node:e.id ~pos:e.pos ~kind:DLiveMask ~items:(mask_items (live_holders ctx))
+ | Query q ->
+ (* iteration 9b: the sub-expressions only READ (engine field-reads copy
+ out at the boundary); the query is trap-capable (engine faults), so
+ the frame needs its drop map here, exactly like DbStub. *)
+ (match q.q_src with QNav e2 -> read_expr ctx e2 | QTable _ -> ());
+ List.iter (read_expr ctx) q.q_wheres;
+ (match q.q_group with Some (_, k) -> read_expr ctx k | None -> ());
+ (match q.q_order with Some (k, _) -> read_expr ctx k | None -> ());
+ (match q.q_take with Some t -> read_expr ctx t | None -> ());
+ read_expr ctx q.q_select;
+ record_drop ctx ~node:e.id ~pos:e.pos ~kind:DLiveMask ~items:(mask_items (live_holders ctx))
| Switch (subject, arms) -> analyze_switch ctx e.id subject arms
(* The root of a place expression is already accounted for by use_place;
diff --git a/compiler/src/parser.ml b/compiler/src/parser.ml
index c4b1695..ed64b95 100644
--- a/compiler/src/parser.ml
+++ b/compiler/src/parser.ml
@@ -356,6 +356,12 @@ let parse_field_ty (st : state) : Ast.field_ty =
| Token.Ident "multi" ->
ignore (advance st);
Ast.Multi (expect_ident st "multi target type")
+ | Token.Ident "backlink" ->
+ ignore (advance st);
+ let cls = expect_ident st "backlink source class" in
+ expect st Token.Dot "'.'";
+ let fld = expect_ident st "backlink source field" in
+ Ast.Backlink (cls, fld)
| Token.Ident "map" ->
ignore (advance st);
expect st Token.Lt "'<'";
@@ -1009,9 +1015,121 @@ and parse_switch_expr (st : state) : Ast.expr =
done;
{ Ast.id; pos; kind = Ast.Switch (subject, List.rev !arms) }
+and parse_insert_expr (st : state) : Ast.expr =
+ (* `insert` + a constructor literal, sharing parse_ctor_literal so the
+ field-list grammar (trailing commas, newlines) can never drift from the
+ ctor's. The literal's node is unwrapped into Insert — its id is reused,
+ which is safe because the Ctor node itself is discarded whole. *)
+ let pos = peek_pos st in
+ ignore (advance st) (* the `insert` trigger token *);
+ skip_newlines st;
+ let lit = parse_ctor_literal st in
+ (match lit.Ast.kind with
+ | Ast.Ctor (cn, fields) -> { lit with Ast.pos; kind = Ast.Insert (cn, fields) }
+ | _ -> lit (* unreachable: parse_ctor_literal only builds Ctor *))
+
+and is_query_trigger (st : state) : bool =
+ (* `from in` — positional, so `from` stays a usable identifier
+ everywhere else (same discipline as insert/select) *)
+ (match peek st with Token.Ident "from" -> true | _ -> false)
+ && (match (tok_at st (st.pos + 1)).kind with Token.Ident _ -> true | _ -> false)
+ && (tok_at st (st.pos + 2)).kind = Token.KwIn
+
+and parse_query_expr (st : state) : Ast.expr =
+ let pos = peek_pos st in
+ let id = fresh_id st in
+ ignore (advance st) (* from *);
+ let var = expect_ident st "query range variable" in
+ expect st Token.KwIn "`in`";
+ (* source: a bare class name is a table scan; any other expression is a
+ navigation (`d.staff`). One token of lookahead: Ident not followed by a
+ `.`/`(`/`[` and sitting where a clause keyword follows is a table name. *)
+ let src =
+ match peek st with
+ | Token.Ident cn
+ when (match (tok_at st (st.pos + 1)).kind with
+ | Token.Dot | Token.LParen | Token.LBracket -> false
+ | _ -> true) ->
+ ignore (advance st);
+ Ast.QTable cn
+ | _ -> Ast.QNav (parse_expr_no_brace st)
+ in
+ (* clauses may sit on their own lines; skip the separating newlines when
+ looking for the next clause keyword (the query is one expression) *)
+ let clause name =
+ skip_newlines st;
+ match peek st with Token.Ident n when n = name -> true | _ -> false
+ in
+ let wheres = ref [] in
+ while clause "where" do
+ ignore (advance st);
+ wheres := parse_expr_no_brace st :: !wheres
+ done;
+ let group =
+ if clause "group" then begin
+ ignore (advance st);
+ let key_elem = parse_expr_no_brace st in
+ ignore key_elem (* the grouped element is the range var; `group e by k` *);
+ if not (clause "by") then fail st (peek_pos st) syntax_code "expected `by` in a group clause";
+ ignore (advance st);
+ let key = parse_expr_no_brace st in
+ if not (clause "into") then fail st (peek_pos st) syntax_code "expected `into` in a group clause";
+ ignore (advance st);
+ let gvar = expect_ident st "group variable" in
+ Some (gvar, key)
+ end
+ else None
+ in
+ let order =
+ if clause "order" then begin
+ ignore (advance st);
+ if not (clause "by") then fail st (peek_pos st) syntax_code "expected `by` after `order`";
+ ignore (advance st);
+ let key = parse_expr_no_brace st in
+ let desc = clause "desc" in
+ if desc then ignore (advance st);
+ Some (key, desc)
+ end
+ else None
+ in
+ (* `take` is a reserved keyword (KwTake, the param convention), not an
+ Ident — so match the token, not the name *)
+ skip_newlines st;
+ let take =
+ if peek st = Token.KwTake then (ignore (advance st); Some (parse_expr_no_brace st)) else None
+ in
+ if not (clause "select") then fail st (peek_pos st) syntax_code "a query must end in `select`";
+ ignore (advance st);
+ let sel = parse_expr st in
+ {
+ Ast.id;
+ pos;
+ kind =
+ Ast.Query
+ {
+ Ast.q_var = var;
+ q_src = src;
+ q_wheres = List.rev !wheres;
+ q_group = group;
+ q_order = order;
+ q_take = take;
+ q_select = sel;
+ q_pos = pos;
+ };
+ }
+
and parse_primary (st : state) : Ast.expr =
match peek st with
+ | _ when is_query_trigger st -> parse_query_expr st
+ | Token.Ident "delete" when (match (tok_at st (st.pos + 1)).kind with
+ | Token.Newline | Token.Semicolon | Token.Eof -> false | _ -> true) ->
+ let pos = peek_pos st in
+ let id = fresh_id st in
+ ignore (advance st);
+ let target = parse_expr st in
+ { Ast.id; pos; kind = Ast.Delete target }
| k when is_select_trigger k -> parse_dbstub_expr st
+ | k when is_insert_trigger k -> parse_insert_expr st
| Token.KwSwitch -> parse_switch_expr st
| Token.Int n ->
let pos = peek_pos st in
@@ -1286,7 +1404,7 @@ and parse_stmt (st : state) : Ast.stmt =
| k when is_insert_trigger k ->
let pos = peek_pos st in
let id = fresh_id st in
- let e = parse_dbstub_expr st in
+ let e = parse_insert_expr st in
end_of_stmt st;
{ Ast.s_id = id; s_pos = pos; s_kind = Ast.ExprStmt e }
| Token.KwLet -> parse_let_stmt st
@@ -1785,6 +1903,23 @@ let rec subst_expr (consts : Ast.expr StringMap.t) (bound : StringSet.t) (e : As
{ e with Ast.kind = Ast.Binary (op, subst_expr consts bound l, subst_expr consts bound r) }
| Ast.Ctor (cn, fields) ->
{ e with Ast.kind = Ast.Ctor (cn, List.map (fun (n, v) -> (n, subst_expr consts bound v)) fields) }
+ | Ast.Insert (cn, fields) ->
+ { e with Ast.kind = Ast.Insert (cn, List.map (fun (n, v) -> (n, subst_expr consts bound v)) fields) }
+ | Ast.Delete t -> { e with Ast.kind = Ast.Delete (subst_expr consts bound t) }
+ | Ast.Query q ->
+ (* the range/group vars shadow consts inside the query body *)
+ let bound' = StringSet.add q.Ast.q_var bound in
+ let bound' = match q.Ast.q_group with Some (g, _) -> StringSet.add g bound' | None -> bound' in
+ let sub = subst_expr consts bound' in
+ { e with Ast.kind = Ast.Query {
+ q with Ast.q_src = (match q.Ast.q_src with
+ | Ast.QTable cn -> Ast.QTable cn
+ | Ast.QNav e2 -> Ast.QNav (subst_expr consts bound e2));
+ q_wheres = List.map sub q.Ast.q_wheres;
+ q_group = (match q.Ast.q_group with Some (g, k) -> Some (g, sub k) | None -> None);
+ q_order = (match q.Ast.q_order with Some (k, d) -> Some (sub k, d) | None -> None);
+ q_take = (match q.Ast.q_take with Some t -> Some (sub t) | None -> None);
+ q_select = sub q.Ast.q_select } }
| Ast.Interp inner -> { e with Ast.kind = Ast.Interp (subst_expr consts bound inner) }
| Ast.ListLit items -> { e with Ast.kind = Ast.ListLit (List.map (subst_expr consts bound) items) }
| Ast.MapLit | Ast.NilLit -> e
diff --git a/compiler/src/types.ml b/compiler/src/types.ml
index 3a5395f..4b904ec 100644
--- a/compiler/src/types.ml
+++ b/compiler/src/types.ml
@@ -306,6 +306,7 @@ let rec has_recursive_structure (cls : class_info) : bool =
| Ast.Ref name -> name = cls.name
| Ast.Multi name -> name = cls.name (* multi Self *)
| Ast.Map (k, v) -> k = cls.name || v = cls.name (* map<_, Self> / map *)
+ | Ast.Backlink _ -> false
| Ast.Nullable inner -> has_recursive_structure_type inner cls.name
) cls.fields
@@ -315,6 +316,7 @@ and has_recursive_structure_type (ty : Ast.field_ty) (cls_name : string) : bool
| Ast.Ref name -> name = cls_name
| Ast.Multi name -> name = cls_name
| Ast.Map (k, v) -> k = cls_name || v = cls_name
+ | Ast.Backlink _ -> false (* a computed inverse holds no owned structure *)
| Ast.Nullable inner -> has_recursive_structure_type inner cls_name
(* @unique field -> persistent identity (plan's "When NOT to emit": a
@@ -363,6 +365,7 @@ let rec typ_of_field_ty (ft : field_ty) : typ =
| Ref name -> TRef name
| Multi inner_name -> TMulti (TScalar inner_name)
| Map (k_name, v_name) -> TMap (TScalar k_name, TScalar v_name)
+ | Backlink (c, _) -> TMulti (TScalar c) (* reads as a collection of C *)
| Nullable inner -> TNullable (typ_of_field_ty inner)
(* wob_kind_of_typ: maps internal typ to .wob field kind *)
@@ -412,6 +415,7 @@ let unknown_fn_code = Diag.types_prefix ^ "04"
let unsatisfied_interface_code = Diag.types_prefix ^ "05"
let incomplete_ctor_code = Diag.types_prefix ^ "06"
let unknown_type_code = Diag.types_prefix ^ "07"
+let query_code = Diag.types_prefix ^ "50" (* WO-E250: query surface (iteration 9b) *)
let non_exhaustive_switch_code = Diag.types_prefix ^ "08"
let invalid_builtin_code = Diag.types_prefix ^ "09"
let module_not_imported_code = Diag.types_prefix ^ "10"
@@ -632,7 +636,7 @@ let rec scalar_name_of (ft : field_ty) : string option =
match ft with
| Scalar name -> Some name
| Nullable inner -> scalar_name_of inner
- | Ref _ | Multi _ | Map _ -> None
+ | Ref _ | Multi _ | Map _ | Backlink _ -> None
(* Checked once per field declaration (not at every access/use site), so
the diagnostic lands at the field's own declaration position and
@@ -1165,6 +1169,11 @@ let typecheck_program ~file ~(module_of : string -> string)
needs `int_to_text` first) -- unlike the placeholders below,
this is a fact, not a guess. *)
Some (TScalar "Text")
+ | Insert _ ->
+ (* the new row's id — the one thing an insert produces *)
+ Some (TScalar "Int")
+ | Query _ -> None (* a query's type is chased only by typecheck_expr *)
+ | Delete _ -> Some (TScalar "Int")
| Unary _ | Binary _ | DbStub _ ->
(* Not chased: the arithmetic-ladder `Binary` ops have no reliable
per-node type in this pass at all (see above); `Unary`/`DbStub`
@@ -1198,7 +1207,7 @@ let typecheck_program ~file ~(module_of : string -> string)
with Not_found -> { typ = TScalar "Int"; is_nil = false })
| Field (base, field_name) ->
let base_res = typecheck_expr env cenv base in
- (match base_res.typ with
+ (match (match base_res.typ with TRef c -> TScalar c | other -> other) with
| TScalar class_name ->
(* Only a *declared* class can be checked for a missing field.
typecheck_expr falls back to `TScalar "Int"` for everything
@@ -1222,9 +1231,14 @@ let typecheck_program ~file ~(module_of : string -> string)
{ typ = TScalar "Int"; is_nil = false }))
| _ -> { typ = TScalar "Int"; is_nil = false })
| Index (base, idx) ->
- let _ = typecheck_expr env cenv base in
+ let base_res = typecheck_expr env cenv base in
let _ = typecheck_expr env cenv idx in
- { typ = TScalar "Int"; is_nil = false }
+ (* `xs[i]` yields the container's element type — a `multi C` indexed
+ is a C (iteration 9b: query results are indexed to pick a row) *)
+ (match base_res.typ with
+ | TMulti et -> { typ = et; is_nil = false }
+ | TMap (_, vt) -> { typ = vt; is_nil = false }
+ | _ -> { typ = TScalar "Int"; is_nil = false })
| Call (callee, args) ->
List.iter (fun arg -> ignore (typecheck_expr env cenv arg)) args;
(match callee.kind with
@@ -1391,7 +1405,8 @@ let typecheck_program ~file ~(module_of : string -> string)
the zero word NEW already leaves there). Everything else
stays WO-E206, classes and records alike. *)
let omittable (default : default_expr option) (fty : field_ty) : bool =
- Option.is_some default || (match fty with Nullable _ -> true | _ -> false)
+ Option.is_some default
+ || (match fty with Nullable _ | Backlink _ -> true | _ -> false)
in
List.iter (fun (fname, fty, fdefault, _) ->
if not (List.mem fname provided) && not (omittable fdefault fty) then
@@ -1405,6 +1420,100 @@ let typecheck_program ~file ~(module_of : string -> string)
(Diag.error ~code:unknown_type_code ~file ~line:e.pos.line ~col:e.pos.col
~message:(Printf.sprintf "unknown type `%s` in constructor" class_name) ());
{ typ = TScalar "Int"; is_nil = false })
+ | Insert (class_name, fields) ->
+ (* iteration 9 Task 3: typed exactly like a constructor literal —
+ same missing-field rule (defaults and `?` fields omittable),
+ same unknown-class diagnostic — but the VALUE is the new row's
+ id. The engine copies every field at the choke point, so field
+ values keep their owners (owner.ml's stores_by_copy). *)
+ (try
+ let cls = StringMap.find class_name syms.classes in
+ let provided = List.map (fun (n, _) -> n) fields in
+ let omittable (default : default_expr option) (fty : field_ty) : bool =
+ Option.is_some default
+ || (match fty with Nullable _ | Backlink _ -> true | _ -> false)
+ in
+ List.iter (fun (fname, fty, fdefault, _) ->
+ if not (List.mem fname provided) && not (omittable fdefault fty) then
+ Diag.Collector.add collector
+ (Diag.error ~code:incomplete_ctor_code ~file ~line:e.pos.line ~col:e.pos.col
+ ~message:(Printf.sprintf "missing field `%s` in insert of `%s`" fname class_name) ())
+ ) cls.fields;
+ { typ = TScalar "Int"; is_nil = false }
+ with Not_found ->
+ Diag.Collector.add collector
+ (Diag.error ~code:unknown_type_code ~file ~line:e.pos.line ~col:e.pos.col
+ ~message:(Printf.sprintf "unknown type `%s` in insert" class_name) ());
+ { typ = TScalar "Int"; is_nil = false })
+ | Delete target ->
+ let tr = typecheck_expr env cenv target in
+ (match (match tr.typ with TRef c -> TScalar c | o -> o) with
+ | TScalar cn when StringMap.mem cn syms.classes -> ()
+ | _ ->
+ Diag.Collector.add collector
+ (Diag.error ~code:query_code ~file ~line:e.pos.line ~col:e.pos.col
+ ~message:"`delete` takes a table-row value" ()));
+ { typ = TScalar "Int"; is_nil = false }
+ | Query q ->
+ (* iteration 9b slice: from/where/select over a table class. The
+ range variable is bound to the class type; a table-class value is
+ its row id at runtime but types AS the class, so `e.field` checks
+ against the class's fields exactly like a heap instance. group /
+ order / take / navigation sources are diagnosed as not-yet so the
+ surface is honest about its edge. *)
+ let elem_err () =
+ { typ = TMulti (TScalar "Int"); is_nil = false }
+ in
+ (match q.q_src with
+ | Ast.QNav nav ->
+ (* `from s in d.staff`: the navigation yields `multi C`, so the
+ range var is a C. Reuse the QTable body by resolving C. *)
+ let nav_res = typecheck_expr env cenv nav in
+ let cn =
+ match nav_res.typ with
+ | TMulti (TScalar c) -> c
+ | _ -> ""
+ in
+ if not (StringMap.mem cn syms.classes) then begin
+ Diag.Collector.add collector
+ (Diag.error ~code:query_code ~file ~line:q.q_pos.line ~col:q.q_pos.col
+ ~message:"query navigation source must be a `backlink`/`multi` of a table class" ());
+ elem_err ()
+ end
+ else begin
+ (if q.q_group <> None then
+ Diag.Collector.add collector
+ (Diag.error ~code:query_code ~file ~line:q.q_pos.line ~col:q.q_pos.col
+ ~message:"group-by on a navigation query is not supported yet" ()));
+ let env' = StringMap.add q.q_var (TScalar cn) env in
+ let cenv' = StringMap.add q.q_var (TScalar cn) cenv in
+ List.iter (fun w -> ignore (typecheck_expr env' cenv' w)) q.q_wheres;
+ (match q.q_order with Some (k, _) -> ignore (typecheck_expr env' cenv' k) | None -> ());
+ (match q.q_take with Some t -> ignore (typecheck_expr env cenv t) | None -> ());
+ let sel = typecheck_expr env' cenv' q.q_select in
+ { typ = TMulti sel.typ; is_nil = false }
+ end
+ | Ast.QTable cn ->
+ if not (StringMap.mem cn syms.classes) then begin
+ Diag.Collector.add collector
+ (Diag.error ~code:query_code ~file ~line:q.q_pos.line ~col:q.q_pos.col
+ ~message:(Printf.sprintf "`from %s in %s`: `%s` is not a declared table class"
+ q.q_var cn cn) ());
+ elem_err ()
+ end
+ else begin
+ (if q.q_group <> None then
+ Diag.Collector.add collector
+ (Diag.error ~code:query_code ~file ~line:q.q_pos.line ~col:q.q_pos.col
+ ~message:"group-by aggregation is not supported yet" ()));
+ let env' = StringMap.add q.q_var (TScalar cn) env in
+ let cenv' = StringMap.add q.q_var (TScalar cn) cenv in
+ List.iter (fun w -> ignore (typecheck_expr env' cenv' w)) q.q_wheres;
+ (match q.q_order with Some (k, _) -> ignore (typecheck_expr env' cenv' k) | None -> ());
+ (match q.q_take with Some t -> ignore (typecheck_expr env cenv t) | None -> ());
+ let sel = typecheck_expr env' cenv' q.q_select in
+ { typ = TMulti sel.typ; is_nil = false }
+ end)
| DbStub _ -> { typ = TVoid; is_nil = false }
| Switch (subject, arms) -> typecheck_switch ~want_value:true env cenv subject arms
| ListLit items ->
@@ -2102,7 +2211,8 @@ and walk_expr (bound : StringSet.t) (visit : StringSet.t -> expr -> unit) (e : e
| Binary (_, l, r) ->
walk_expr bound visit l;
walk_expr bound visit r
- | Ctor (_, fields) -> List.iter (fun (_, v) -> walk_expr bound visit v) fields
+ | Ctor (_, fields) | Insert (_, fields) ->
+ List.iter (fun (_, v) -> walk_expr bound visit v) fields
| Interp inner -> walk_expr bound visit inner
| ListLit items -> List.iter (walk_expr bound visit) items
| MapLit | NilLit -> ()
@@ -2111,6 +2221,16 @@ and walk_expr (bound : StringSet.t) (visit : StringSet.t -> expr -> unit) (e : e
walk_expr bound visit body;
walk_block (StringSet.add ename bound) visit handler
| DbStub _ -> ()
+ | Delete t -> walk_expr bound visit t
+ | Query q ->
+ (match q.q_src with QNav e -> walk_expr bound visit e | QTable _ -> ());
+ let b = StringSet.add q.q_var bound in
+ let b = match q.q_group with Some (g, _) -> StringSet.add g b | None -> b in
+ List.iter (walk_expr b visit) q.q_wheres;
+ (match q.q_group with Some (_, k) -> walk_expr b visit k | None -> ());
+ (match q.q_order with Some (k, _) -> walk_expr b visit k | None -> ());
+ (match q.q_take with Some t -> walk_expr b visit t | None -> ());
+ walk_expr b visit q.q_select
| Switch (subject, arms) ->
walk_expr bound visit subject;
List.iter
@@ -2370,6 +2490,7 @@ let rec field_ty_str (ft : field_ty) : string =
| Ref s -> "ref " ^ s
| Multi s -> "multi " ^ s
| Map (k, v) -> "map<" ^ k ^ ", " ^ v ^ ">"
+ | Backlink (c, f) -> "backlink " ^ c ^ "." ^ f
| Nullable t -> "?" ^ field_ty_str t
let dump_symbols (syms : symbols) : string =
diff --git a/compiler/test/golden/ast/db-stub.expected b/compiler/test/golden/ast/db-stub.expected
index 17db105..3494df8 100644
--- a/compiler/test/golden/ast/db-stub.expected
+++ b/compiler/test/golden/ast/db-stub.expected
@@ -1,6 +1,6 @@
1:1 METHOD sync()
- 2:3 DB_STUB IDENT(insert) IDENT(Product) LBRACE IDENT(sku) COLON STR(A1) COMMA IDENT(price) COLON INT(10) RBRACE
+ 2:3 EXPR INSERT Product { sku: "A1", price: 10 }
3:3 DB_STUB IDENT(select) IDENT(Product) LBRACE IDENT(sku) EQEQ STR(A1) RBRACE
4:3 LET rows = DB_STUB(IDENT(select) IDENT(Product) LBRACE IDENT(price) GT INT(5) RBRACE)
- 5:3 DB_STUB KW_INSERT IDENT(Product) LBRACE IDENT(sku) COLON STR(A2) RBRACE
+ 5:3 EXPR INSERT Product { sku: "A2" }
6:3 DB_STUB KW_SELECT IDENT(Product) LBRACE IDENT(sku) EQEQ STR(A2) RBRACE
diff --git a/compiler/test/golden/owner/drops.wo b/compiler/test/golden/owner/drops.wo
index f50111f..e487997 100644
--- a/compiler/test/golden/owner/drops.wo
+++ b/compiler/test/golden/owner/drops.wo
@@ -38,7 +38,7 @@ fn pick(take a: Item, take b: Item, flag: Bool) -> Int {
}
fn store(take r: Item) -> Int {
- insert into rows values (1)
+ insert Row { n: 1 }
return 0
}
@@ -63,3 +63,7 @@ fn reinit_after_move(take a: Item) -> Int {
a = Item { n: 7 }
return 0
}
+
+class Row {
+ n: Int
+}
diff --git a/compiler/test/runner.ml b/compiler/test/runner.ml
index 28635c7..8d644a5 100644
--- a/compiler/test/runner.ml
+++ b/compiler/test/runner.ml
@@ -531,35 +531,35 @@ let () =
| _ -> check "ctor literal: exactly one free fn" false
let () =
- (* The brief's stated asymmetry: `insert` is a statement-only trigger
- (parser.ml's is_insert_trigger, checked only in parse_stmt) — a
- bare `insert` reached from parse_primary is just an ordinary
- identifier reference, exactly like self/me/on/service/policy's own
- "recognized positionally, not a reserved word" rule (this task's
- own keyword-discipline note). `select` (is_select_trigger) is
- checked unconditionally *inside* parse_primary, so the same
- position always builds a DbStub instead. Neither is an error on
- its own — the difference shows up in which Ast.expr_kind comes
- back. *)
+ (* Iteration 9 Task 3 retired the old asymmetry: `insert` is grammar-owned
+ in BOTH positions now — a typed Insert node validated like a ctor,
+ returning the id — while `select` stays the opaque DbStub until
+ Task 5. The old contract ("bare insert is a plain Ident") is gone
+ with the stub that motivated it. *)
let prog, collector =
- parse_str ~file:"insert-vs-select.wo" "fn f() {\n let a = insert\n let b = select\n}\n"
+ parse_str ~file:"insert-vs-select.wo"
+ "fn f() {\n let a = insert Product { sku: \"A1\" }\n let b = select\n}\n"
in
- check_eq "insert vs. select as bare expressions: no diagnostics" ~expected:0
+ check_eq "typed insert + stub select: no diagnostics" ~expected:0
~actual:(List.length (Diag.Collector.diagnostics collector))
string_of_int;
- match prog.Ast.decls with
+ (match prog.Ast.decls with
| [ Ast.Fn m ] -> (
match m.body with
| [
{ Ast.s_kind = Ast.Let { name = "a"; value = a_val; _ }; _ };
{ Ast.s_kind = Ast.Let { name = "b"; value = b_val; _ }; _ };
] ->
- check "bare `insert` in expression position is a plain Ident"
- (match a_val.Ast.kind with Ast.Ident "insert" -> true | _ -> false);
+ check "`insert` in expression position is a typed Insert node"
+ (match a_val.Ast.kind with Ast.Insert ("Product", [ ("sku", _) ]) -> true | _ -> false);
check "bare `select` in expression position always becomes a DbStub"
(match b_val.Ast.kind with Ast.DbStub _ -> true | _ -> false)
| _ -> check "insert vs. select: exactly two `let` statements" false)
- | _ -> check "insert vs. select: exactly one free fn" false
+ | _ -> check "insert vs. select: exactly one free fn" false);
+ (* and a bare `insert` with no literal is a parse error now, not an Ident *)
+ let _, c2 = parse_str ~file:"bare-insert.wo" "fn f() {\n let a = insert\n}\n" in
+ check "bare `insert` with no constructor literal is a diagnostic"
+ (List.length (Diag.Collector.diagnostics c2) > 0)
let () =
(* The no_brace guard (parser.ml's state.no_brace / looks_like_ctor):
@@ -2382,7 +2382,7 @@ let validate_image (img : string) : string list =
let u64 o = if ok 8 o then String.get_int64_le img o else 0L in
let none = 0xFFFFFFFF in
if u32 0 <> 0x31424F57 then fail "bad magic";
- if u32 4 <> 2 then fail "unsupported version";
+ if u32 4 <> 3 then fail "unsupported version";
let coff = u32 8 and ccnt = u32 12 in
let koff = u32 16 and kcnt = u32 20 in
let ioff = u32 24 and icnt = u32 28 in
@@ -2419,6 +2419,7 @@ let validate_image (img : string) : string list =
if flags land lnot 0x01 <> 0 then fail (Printf.sprintf "class %d: unknown flags" i);
if fcnt > 65535 then fail (Printf.sprintf "class %d: too many fields" i);
class_fields.(i) <- fcnt;
+ let kco = !o in (* the kind bytes' offset: the v3 index walk re-reads them *)
for j = 0 to fcnt - 1 do
if u8 (!o + j) > 5 then fail (Printf.sprintf "class %d field %d: bad kind" i j)
done;
@@ -2435,6 +2436,25 @@ let validate_image (img : string) : string list =
fail (Printf.sprintf "class %d field %d: field class out of range" i j)
done;
o := !o + (fcnt * 12);
+ (* v3: the index tail — flags (bit0 only), col_cnt 1..8, columns in
+ range and scalar/Text-kinded. Mirrors loader.c's checks. *)
+ let icnt_x = u32 !o in
+ o := !o + 4;
+ if icnt_x > 64 then fail (Printf.sprintf "class %d: too many indexes" i);
+ for x = 0 to icnt_x - 1 do
+ let ifl = u32 !o and ccnt = u32 (!o + 4) in
+ o := !o + 8;
+ if ifl land lnot 1 <> 0 then fail (Printf.sprintf "class %d index %d: unknown flags" i x);
+ if ccnt = 0 || ccnt > 8 then fail (Printf.sprintf "class %d index %d: bad column count" i x);
+ for c = 0 to ccnt - 1 do
+ let col = u32 !o in
+ o := !o + 4;
+ if col >= fcnt then fail (Printf.sprintf "class %d index %d: column out of range" i x);
+ let kind = u8 (kco + col) in
+ if kind <> 0 && kind <> 3 then
+ fail (Printf.sprintf "class %d index %d: column %d is not scalar or Text" i x c)
+ done
+ done;
if !o > len then fail (Printf.sprintf "class %d: truncated" i)
done;
(* interfaces + vtable rows *)
@@ -2580,7 +2600,12 @@ let validate_image (img : string) : string list =
| 22 | 23 | 24 | 25 | 26 | 27 | 28 -> rchk pc a
| 29 ->
rchk pc a;
- if c > 12 then fail (Printf.sprintf "method %d pc %d: builtin out of range" i pc)
+ (* the mirror's ceiling tracks wob.h's WO_B_MAX only for ids the
+ golden lowering suite actually emits; 61 = DB_INSERT (arity 1:
+ the class-id slot — field slots are runtime-validated, same as
+ the C loader) *)
+ if c > 12 && (c < 61 || c > 67) then
+ fail (Printf.sprintf "method %d pc %d: builtin out of range" i pc)
else if c = 4 then begin
if b > 5 then fail (Printf.sprintf "method %d pc %d: bad element kind" i pc)
end
@@ -2595,6 +2620,13 @@ let validate_image (img : string) : string list =
| 1 | 2 | 3 | 7 | 8 -> 1
| 5 | 6 | 11 | 12 -> 2
| 10 -> 3
+ | 61 -> 1
+ | 62 -> 4
+ | 63 -> 2
+ | 64 -> 1
+ | 65 -> 3
+ | 66 -> 3
+ | 67 -> 2
| _ -> 0
in
if arity > 0 then begin
@@ -2951,7 +2983,8 @@ let () =
( "text: concat, equality, words",
"fn f(a: Text, b: Text) -> Int {\n let joined = a .. b\n\
\ if joined == a {\n return 1\n }\n return words(joined)\n}\n" );
- ("db stub statement", "fn f() -> Int {\n insert into rows values (1)\n return 0\n}\n");
+ ( "db insert statement",
+ "class Row {\n n: Int\n}\n\nfn f() -> Int {\n insert Row { n: 1 }\n return 0\n}\n" );
( "nested calls in arguments",
"fn one() -> Int {\n return 1\n}\n\nfn add(a: Int, b: Int) -> Int {\n\
\ return a + b\n}\n\nfn f() -> Int {\n return add(add(one(), one()), one())\n}\n" );
diff --git a/database/src/CODE-LOGIC.md b/database/src/CODE-LOGIC.md
new file mode 100644
index 0000000..ad5987e
--- /dev/null
+++ b/database/src/CODE-LOGIC.md
@@ -0,0 +1,63 @@
+# database/src — how the engine hangs together
+
+The database engine is its own top-level directory, statically linked into
+every `wovm` and every runtime test binary (`runtime/Makefile`'s `DBSRC`).
+One binary, unchanged. Format doc: `docs/plan/oop-vm/04-db-binding.md`.
+Memory-safety doctrine: the 9b design's section 6.
+
+## table.c — rows (iteration 9, Task 1)
+
+```
+VM values ──copy──▶ row slots (engine-owned malloc) ──copy──▶ fresh VM values
+ wo_row_insert wo_row_read
+```
+
+- **No VM pointer ever enters a slab; no slab pointer ever leaves.** Encode
+ copies per kind (Texts to `db_text`, owned objects flattened recursively to
+ `db_rec`, containers element-wise); decode allocates fresh VM values from
+ the caller's `wo_rt`. The GCREF kind is refused at encode — the compiler
+ should have made that impossible (the GC bulkhead), the engine refuses it
+ anyway.
+- **Rows never move.** Slabs of 256 are malloc'd and kept for the table's
+ life; the free-slot list recycles removed slots before any slab grows;
+ the id hash maps id → slot. Ids are never reused (per-table counter,
+ shard-interleaved `S+1, S+1+N, …`), which is also what makes the hash's
+ tombstone sentinel safe.
+- **Choke points**: `wo_row_insert` / `wo_row_remove` carry the `INDEX HOOK`
+ comments where Task 4's secondary indexes attach and Task 2's WAL stages
+ its record. Nothing else may mutate storage.
+- One deliberate file-static: `g_classes` for recursive frees (`db_val_free`
+ has no context parameter). One process, one class table; revisit at
+ iteration 8 (shards share the same immutable table).
+
+## wal.c — durability (iteration 9, Task 2)
+
+The commit order IS the module: RAM apply → stage → one pwrite + one
+fdatasync → ack. `wo_wal_commit` returning 0 is the only thing "durable"
+means. Replay never touches the VM heap — payloads decode straight into
+engine-owned values and re-enter through the row API, so whatever hooks the
+choke points (indexes, Task 4) applies to replayed rows identically. Torn
+tails end the intact prefix and get overwritten by the next commit;
+CRC-valid-but-undecodable records fail replay loudly (corruption is not a
+tear). The crash battery in `runtime/test/test_wal.c` is the module's
+meaning proven: acked-over-a-pipe after commit, SIGKILL mid-stream, replay,
+zero acked-but-missing.
+
+## db.c — statement executors (iteration 9, Task 3)
+
+One dispatcher, the builtin contract (0 ok, else WO_T_* + msg). The engine
+handles ride `wo_rt.db` / `wo_rt.wal` as opaque pointers set by main.c —
+NULL db traps WO_T_DB, NULL wal means RAM-only (the corpus's mode; WO_DATA
+opts into durability). Insert's contract: RAM apply through the row API,
+then stage + commit BEFORE returning — the builtin's return is the
+acknowledgment, so a failed commit un-applies the row and traps WO_T_IO
+rather than acknowledging what disk never got.
+
+## Verifying a change
+
+- `make -C runtime test` — `test_table` is this directory's suite (round
+ trips across kinds, nil encodings, shard interleave, slab growth, slot
+ reuse, misuse), ASan+UBSan like every runtime test.
+- `just oop-e2e`, `just log-watcher` — regression that linking the engine
+ into wovm changed nothing observable (it is dead code until Task 3 wires
+ the first builtin).
diff --git a/database/src/db.c b/database/src/db.c
new file mode 100644
index 0000000..3f40d88
--- /dev/null
+++ b/database/src/db.c
@@ -0,0 +1,163 @@
+#include "db.h"
+
+#include
+
+#include "cont.h"
+#include "table.h"
+#include "wal.h"
+
+int wo_builtin_db(wo_vm *vm, uint64_t *R, uint32_t ins, const char **msg) {
+ uint32_t A = wo_ins_a(ins), B = wo_ins_b(ins), C = wo_ins_c(ins);
+ wo_db *db = (wo_db *)vm->rt.db;
+ if (!db) {
+ *msg = "database engine not initialized";
+ return WO_T_DB;
+ }
+ switch (C) {
+ case WO_B_DB_INSERT: {
+ uint32_t cid = (uint32_t)R[B];
+ int ek = 0;
+ uint64_t id = wo_row_insert(db, cid, &R[B + 1], msg, &ek);
+ if (!id)
+ return ek == DB_ERR_UNIQUE ? WO_T_UNIQUE
+ : ek == DB_ERR_OOM ? WO_T_OOM
+ : WO_T_DB;
+ wo_wal *w = (wo_wal *)vm->rt.wal;
+ if (w) {
+ /* RAM applied, record staged, ONE commit before the ack (the
+ * builtin's return). A failed commit is a failed write: the
+ * row is removed again so RAM never claims what disk never
+ * acknowledged, and the statement traps. */
+ if (wo_wal_append_insert(w, db, cid, id) != 0 || wo_wal_commit(w) != 0) {
+ wo_row_remove(db, cid, id);
+ *msg = "wal commit failed";
+ return WO_T_IO;
+ }
+ }
+ R[A] = id;
+ return 0;
+ }
+ case WO_B_DB_UPDATE_FIELD: {
+ uint32_t cid = (uint32_t)R[B];
+ uint64_t id = R[B + 1];
+ uint32_t field = (uint32_t)R[B + 2];
+ int ek = 0;
+ if (wo_row_update_field(db, cid, id, field, R[B + 3], msg, &ek) != 0)
+ return ek == DB_ERR_UNIQUE ? WO_T_UNIQUE : ek == DB_ERR_OOM ? WO_T_OOM : WO_T_DB;
+ wo_wal *w = (wo_wal *)vm->rt.wal;
+ if (w) {
+ if (wo_wal_append_update(w, db, cid, id) != 0 || wo_wal_commit(w) != 0) {
+ *msg = "wal commit failed"; /* RAM ahead of disk: trap, do not ack */
+ return WO_T_IO;
+ }
+ }
+ R[A] = 0;
+ return 0;
+ }
+ case WO_B_DB_DELETE: {
+ uint32_t cid = (uint32_t)R[B];
+ uint64_t id = R[B + 1];
+ /* FK restrict: refuse if another row still references this one
+ (iteration 9b) — nothing is removed, the statement traps */
+ if (wo_row_has_referrers(db, cid, id)) {
+ *msg = "row is still referenced (restrict)";
+ return WO_T_FK;
+ }
+ if (wo_row_remove(db, cid, id) != 0) {
+ *msg = "no such row";
+ return WO_T_DB;
+ }
+ wo_wal *w = (wo_wal *)vm->rt.wal;
+ if (w) {
+ if (wo_wal_append_remove(w, cid, id) != 0 || wo_wal_commit(w) != 0) {
+ *msg = "wal commit failed";
+ return WO_T_IO;
+ }
+ }
+ R[A] = 0;
+ return 0;
+ }
+ case WO_B_DB_SCAN: {
+ uint32_t cid = (uint32_t)R[B];
+ if (cid >= db->class_cnt) {
+ *msg = "no such class";
+ return WO_T_DB;
+ }
+ wo_multi *ids = wo_multi_new(&vm->rt, WO_K_SCALAR);
+ if (!ids) return WO_T_OOM;
+ /* materialize the id list up front — the 9b cursor-stability rule:
+ * the loop body then point-reads each id, so a row updated mid-loop
+ * (even an indexed column) cannot disturb the iteration */
+ db_table *t = &db->tables[cid];
+ if (t->row_size) {
+ uint32_t total = t->slab_cnt * DB_SLAB_ROWS;
+ for (uint32_t g = 0; g < total; g++) {
+ if (!(t->bitmap[g >> 6] & (1ull << (g & 63)))) continue;
+ db_row *row =
+ (db_row *)(t->slabs[g / DB_SLAB_ROWS] + (size_t)(g % DB_SLAB_ROWS) * t->row_size);
+ if (wo_multi_push(ids, row->id) != 0) return WO_T_OOM;
+ }
+ }
+ R[A] = (uint64_t)(uintptr_t)ids;
+ return 0;
+ }
+ case WO_B_DB_GET_FIELD: {
+ uint32_t cid = (uint32_t)R[B];
+ uint64_t id = R[B + 1];
+ uint32_t field = (uint32_t)R[B + 2];
+ if (cid >= db->class_cnt || field >= db->classes[cid].field_cnt) {
+ *msg = "no such field";
+ return WO_T_DB;
+ }
+ db_row *row = wo_row_ptr(db, cid, id);
+ if (!row) {
+ *msg = "no such row";
+ return WO_T_DB;
+ }
+ int ok = 1;
+ uint64_t v = wo_val_decode_vm(db, &vm->rt, db->classes[cid].kinds[field],
+ row->slots[field], &ok, msg);
+ if (!ok) return WO_T_OOM;
+ R[A] = v;
+ return 0;
+ }
+ case WO_B_DB_PROBE: {
+ uint32_t cid = (uint32_t)R[B];
+ uint32_t index = (uint32_t)R[B + 1];
+ if (cid >= db->class_cnt) {
+ *msg = "no such class";
+ return WO_T_DB;
+ }
+ wo_multi *ids = wo_multi_new(&vm->rt, WO_K_SCALAR);
+ if (!ids) return WO_T_OOM;
+ db_table *t = &db->tables[cid];
+ if (t->row_size && index < t->index_cnt) {
+ db_index *ix = &t->indexes[index];
+ uint32_t col = ix->cols[0];
+ uint8_t kind = db->classes[cid].kinds[col];
+ uint64_t key = R[B + 2];
+ uint32_t total = t->slab_cnt * DB_SLAB_ROWS;
+ for (uint32_t g = 0; g < total; g++) {
+ if (!(t->bitmap[g >> 6] & (1ull << (g & 63)))) continue;
+ db_row *row =
+ (db_row *)(t->slabs[g / DB_SLAB_ROWS] + (size_t)(g % DB_SLAB_ROWS) * t->row_size);
+ int eq;
+ if (kind == WO_K_TEXT) {
+ const wo_str *want = (const wo_str *)(uintptr_t)key;
+ const db_text *have = (const db_text *)(uintptr_t)row->slots[col];
+ eq = (!want && !have) ||
+ (want && have && want->len == have->len &&
+ memcmp(want->data, have->bytes, have->len) == 0);
+ } else
+ eq = row->slots[col] == key;
+ if (eq && wo_multi_push(ids, row->id) != 0) return WO_T_OOM;
+ }
+ }
+ R[A] = (uint64_t)(uintptr_t)ids;
+ return 0;
+ }
+ default:
+ *msg = "unknown db builtin";
+ return WO_T_DB;
+ }
+}
diff --git a/database/src/db.h b/database/src/db.h
new file mode 100644
index 0000000..95eb97d
--- /dev/null
+++ b/database/src/db.h
@@ -0,0 +1,24 @@
+/* db.h — DB statement executors (iteration 9, Task 3+).
+ *
+ * The VM reaches the engine through one dispatcher with the same contract
+ * as every builtin family: 0 = ok, else a WO_T_* code with *msg set. The
+ * engine and WAL handles ride the runtime context as opaque pointers
+ * (obj.h's rt.db / rt.wal) — set by main.c at boot, NULL in test binaries
+ * that never touch DB statements (a DB builtin with rt.db == NULL traps
+ * WO_T_DB "engine not initialized").
+ *
+ * Commit contract per statement (until iteration 8 brings ticks): the
+ * insert applies to RAM, stages its WAL record, and COMMITS before the
+ * builtin returns — the builtin returning IS the acknowledgment, so the
+ * ack-after-fsync doctrine holds at statement granularity. No WAL
+ * (rt.wal == NULL, no WO_DATA) means RAM-only: every test and every
+ * corpus fixture runs that way; durability is opt-in by pointing WO_DATA
+ * at a directory. */
+#ifndef WO_DB_H
+#define WO_DB_H
+
+#include "vm.h"
+
+int wo_builtin_db(wo_vm *vm, uint64_t *R, uint32_t ins, const char **msg);
+
+#endif /* WO_DB_H */
diff --git a/database/src/table.c b/database/src/table.c
new file mode 100644
index 0000000..a8dc324
--- /dev/null
+++ b/database/src/table.c
@@ -0,0 +1,753 @@
+#include "table.h"
+
+#include
+#include
+
+#include "cont.h"
+
+/* ---- engine-owned value encode / free / decode ------------------------- */
+
+/* Free one encoded slot value of [kind]. Recursion mirrors encoding. */
+static void db_val_free(uint8_t kind, uint64_t v);
+
+static void db_rec_free(db_rec *r, const wo_classdesc *classes) {
+ const wo_classdesc *c = &classes[r->class_id];
+ for (uint32_t i = 0; i < c->field_cnt; i++) db_val_free(c->kinds[i], r->slots[i]);
+ free(r);
+}
+
+/* db_val_free needs the class table for nested records; a file-static is
+ * the honest signature here — one engine per process today (N=1), and the
+ * pointer is set once at init. Revisit when iteration 8 brings N>1 shards
+ * (each shard's wo_db shares the same immutable class table anyway). */
+static const wo_classdesc *g_classes;
+
+static void db_val_free(uint8_t kind, uint64_t v) {
+ if (!v) return;
+ switch (kind) {
+ case WO_K_SCALAR: return;
+ case WO_K_TEXT: free((db_text *)(uintptr_t)v); return;
+ case WO_K_OWNED: db_rec_free((db_rec *)(uintptr_t)v, g_classes); return;
+ case WO_K_MULTI: {
+ db_multi *m = (db_multi *)(uintptr_t)v;
+ for (uint32_t i = 0; i < m->len; i++) db_val_free(m->elem_kind, m->items[i]);
+ free(m);
+ return;
+ }
+ case WO_K_MAP: {
+ db_map *m = (db_map *)(uintptr_t)v;
+ for (uint32_t i = 0; i < m->len; i++) {
+ db_val_free(m->key_kind, m->kv[2 * i]);
+ db_val_free(m->val_kind, m->kv[2 * i + 1]);
+ }
+ free(m);
+ return;
+ }
+ default: return; /* GCREF never stored */
+ }
+}
+
+/* Encode one VM value into an engine-owned slot value. 0-with-*ok=0 means
+ * failure (OOM or a GCREF); a genuine nil encodes as 0 with *ok=1. */
+static uint64_t db_val_encode(const wo_classdesc *classes, uint8_t kind, uint64_t v,
+ int *ok, const char **msg) {
+ *ok = 1;
+ switch (kind) {
+ case WO_K_SCALAR: return v;
+ case WO_K_TEXT: {
+ if (!v) return 0;
+ const wo_str *s = (const wo_str *)(uintptr_t)v;
+ db_text *t = malloc(sizeof(db_text) + s->len);
+ if (!t) goto oom;
+ t->len = s->len;
+ memcpy(t->bytes, s->data, s->len);
+ return (uint64_t)(uintptr_t)t;
+ }
+ case WO_K_OWNED: {
+ if (!v) return 0;
+ const wo_hdr *o = (const wo_hdr *)(uintptr_t)v;
+ const wo_classdesc *c = &classes[o->class_id];
+ db_rec *r = malloc(sizeof(db_rec) + (size_t)c->field_cnt * 8u);
+ if (!r) goto oom;
+ r->class_id = o->class_id;
+ r->_pad = 0;
+ const uint64_t *f = (const uint64_t *)(const void *)(o + 1);
+ for (uint32_t i = 0; i < c->field_cnt; i++) {
+ r->slots[i] = db_val_encode(classes, c->kinds[i], f[i], ok, msg);
+ if (!*ok) { /* free what we built so far, then fail upward */
+ for (uint32_t j = 0; j < i; j++) db_val_free(c->kinds[j], r->slots[j]);
+ free(r);
+ return 0;
+ }
+ }
+ return (uint64_t)(uintptr_t)r;
+ }
+ case WO_K_MULTI: {
+ if (!v) return 0;
+ const wo_multi *m = (const wo_multi *)(uintptr_t)v;
+ db_multi *d = malloc(sizeof(db_multi) + (size_t)m->len * 8u);
+ if (!d) goto oom;
+ d->elem_kind = m->elem_kind;
+ d->len = m->len;
+ for (uint32_t i = 0; i < m->len; i++) {
+ d->items[i] = db_val_encode(classes, m->elem_kind, m->items[i], ok, msg);
+ if (!*ok) {
+ for (uint32_t j = 0; j < i; j++) db_val_free(d->elem_kind, d->items[j]);
+ free(d);
+ return 0;
+ }
+ }
+ return (uint64_t)(uintptr_t)d;
+ }
+ case WO_K_MAP: {
+ if (!v) return 0;
+ const wo_map *m = (const wo_map *)(uintptr_t)v;
+ db_map *d = malloc(sizeof(db_map) + (size_t)m->len * 16u);
+ if (!d) goto oom;
+ d->key_kind = m->key_kind;
+ d->val_kind = m->val_kind;
+ d->len = m->len;
+ for (uint32_t i = 0; i < m->len; i++) {
+ d->kv[2 * i] = db_val_encode(classes, m->key_kind, m->keys[i], ok, msg);
+ uint64_t dv = 0;
+ if (*ok) dv = db_val_encode(classes, m->val_kind, m->vals[i], ok, msg);
+ d->kv[2 * i + 1] = dv;
+ if (!*ok) {
+ for (uint32_t j = 0; j <= i; j++) {
+ db_val_free(d->key_kind, d->kv[2 * j]);
+ db_val_free(d->val_kind, d->kv[2 * j + 1]);
+ }
+ free(d);
+ return 0;
+ }
+ }
+ return (uint64_t)(uintptr_t)d;
+ }
+ default:
+ *ok = 0;
+ *msg = "a garbage-collected value cannot be stored in a table field";
+ return 0;
+ }
+oom:
+ *ok = 0;
+ *msg = "out of memory encoding a row";
+ return 0;
+}
+
+/* Decode one engine slot back into a fresh VM value (the out-gate: always
+ * a copy). 0-with-*ok=0 = OOM; nil decodes as 0 with *ok=1. */
+static uint64_t db_val_decode(wo_rt *rt, uint8_t kind, uint64_t v, int *ok,
+ const char **msg) {
+ *ok = 1;
+ switch (kind) {
+ case WO_K_SCALAR: return v;
+ case WO_K_TEXT: {
+ if (!v) return 0;
+ const db_text *t = (const db_text *)(uintptr_t)v;
+ wo_str *s = wo_str_new(rt, t->bytes, t->len);
+ if (!s) goto oom;
+ return (uint64_t)(uintptr_t)s;
+ }
+ case WO_K_OWNED: {
+ if (!v) return 0;
+ const db_rec *r = (const db_rec *)(uintptr_t)v;
+ wo_hdr *o = wo_obj_new(rt, r->class_id);
+ if (!o) goto oom;
+ const wo_classdesc *c = &rt->classes[r->class_id];
+ uint64_t *f = wo_fields(o);
+ for (uint32_t i = 0; i < c->field_cnt; i++) {
+ f[i] = db_val_decode(rt, c->kinds[i], r->slots[i], ok, msg);
+ if (!*ok) return 0; /* partial object: rt teardown reclaims (test scope) */
+ }
+ return (uint64_t)(uintptr_t)o;
+ }
+ case WO_K_MULTI: {
+ if (!v) return 0;
+ const db_multi *d = (const db_multi *)(uintptr_t)v;
+ wo_multi *m = wo_multi_new(rt, d->elem_kind);
+ if (!m) goto oom;
+ for (uint32_t i = 0; i < d->len; i++) {
+ uint64_t ev = db_val_decode(rt, d->elem_kind, d->items[i], ok, msg);
+ if (!*ok || wo_multi_push(m, ev) != 0) goto oom;
+ }
+ return (uint64_t)(uintptr_t)m;
+ }
+ case WO_K_MAP: {
+ if (!v) return 0;
+ const db_map *d = (const db_map *)(uintptr_t)v;
+ wo_map *m = wo_map_new(rt, d->key_kind, d->val_kind);
+ if (!m) goto oom;
+ for (uint32_t i = 0; i < d->len; i++) {
+ uint64_t kv = db_val_decode(rt, d->key_kind, d->kv[2 * i], ok, msg);
+ uint64_t vv = 0;
+ if (*ok) vv = db_val_decode(rt, d->val_kind, d->kv[2 * i + 1], ok, msg);
+ uint64_t old;
+ if (!*ok || wo_map_set(m, kv, vv, &old) < 0) goto oom;
+ }
+ return (uint64_t)(uintptr_t)m;
+ }
+ default: return 0; /* GCREF never stored, so never decoded */
+ }
+oom:
+ *ok = 0;
+ *msg = "out of memory decoding a row";
+ return 0;
+}
+
+/* ---- id hash (open addressing, pow2, id -> global slot + 1) ----------- */
+
+static uint64_t hmix(uint64_t x) { /* splitmix64 finalizer */
+ x += 0x9e3779b97f4a7c15ull;
+ x = (x ^ (x >> 30)) * 0xbf58476d1ce4e5b9ull;
+ x = (x ^ (x >> 27)) * 0x94d049bb133111ebull;
+ return x ^ (x >> 31);
+}
+
+/* Ids are never 0 (0 spells "empty bucket") and never reused, so all-ones
+ * can never collide with a live id — it marks a deleted bucket that probes
+ * walk straight past. */
+#define H_DELETED ((uint64_t)-1)
+
+static int hgrow(db_table *t) {
+ size_t ncap = t->hcap ? t->hcap * 2 : 64;
+ uint64_t *nk = calloc(ncap, 8), *nv = calloc(ncap, 8);
+ if (!nk || !nv) {
+ free(nk);
+ free(nv);
+ return -1;
+ }
+ for (size_t i = 0; i < t->hcap; i++) {
+ if (!t->hkeys[i] || t->hkeys[i] == H_DELETED) continue;
+ size_t j = hmix(t->hkeys[i]) & (ncap - 1);
+ while (nk[j]) j = (j + 1) & (ncap - 1);
+ nk[j] = t->hkeys[i];
+ nv[j] = t->hvals[i];
+ }
+ free(t->hkeys);
+ free(t->hvals);
+ t->hkeys = nk;
+ t->hvals = nv;
+ t->hcap = ncap;
+ return 0;
+}
+
+static int hput(db_table *t, uint64_t id, uint64_t slot1) {
+ if (t->hlen * 10 >= t->hcap * 7 && hgrow(t) != 0) return -1;
+ size_t j = hmix(id) & (t->hcap - 1);
+ while (t->hkeys[j] && t->hkeys[j] != id) j = (j + 1) & (t->hcap - 1);
+ if (!t->hkeys[j]) t->hlen++;
+ t->hkeys[j] = id;
+ t->hvals[j] = slot1;
+ return 0;
+}
+
+static uint64_t hget(const db_table *t, uint64_t id) {
+ if (!t->hcap) return 0;
+ size_t j = hmix(id) & (t->hcap - 1);
+ while (t->hkeys[j]) {
+ if (t->hkeys[j] == id) return t->hvals[j];
+ j = (j + 1) & (t->hcap - 1);
+ }
+ return 0;
+}
+
+static void hdel(db_table *t, uint64_t id) {
+ if (!t->hcap) return;
+ size_t j = hmix(id) & (t->hcap - 1);
+ while (t->hkeys[j]) {
+ if (t->hkeys[j] == id) {
+ t->hkeys[j] = H_DELETED;
+ t->hvals[j] = 0;
+ return;
+ }
+ j = (j + 1) & (t->hcap - 1);
+ }
+}
+
+/* ---- tables and rows ---------------------------------------------------- */
+
+/* ---- secondary indexes (Task 4) ---------------------------------------- */
+
+/* hash of one row's index columns: kind-driven, never trusted for equality */
+static uint64_t idx_hash(const wo_classdesc *c, const db_index *ix, const db_row *r) {
+ uint64_t h = 0x9e3779b97f4a7c15ull;
+ for (uint32_t i = 0; i < ix->col_cnt; i++) {
+ uint32_t col = ix->cols[i];
+ uint64_t v = r->slots[col];
+ if (c->kinds[col] == WO_K_TEXT) {
+ const db_text *t = (const db_text *)(uintptr_t)v;
+ uint64_t th = 1469598103934665603ull; /* FNV-1a over bytes; nil = 0 */
+ if (t)
+ for (uint32_t b = 0; b < t->len; b++) th = (th ^ (uint8_t)t->bytes[b]) * 1099511628211ull;
+ else th = 0;
+ v = th;
+ }
+ h ^= hmix(v + i);
+ }
+ return h ? h : 1; /* 0 marks an empty bucket */
+}
+
+static int idx_cols_equal(const wo_classdesc *c, const db_index *ix, const db_row *a,
+ const db_row *b) {
+ for (uint32_t i = 0; i < ix->col_cnt; i++) {
+ uint32_t col = ix->cols[i];
+ if (c->kinds[col] == WO_K_TEXT) {
+ const db_text *x = (const db_text *)(uintptr_t)a->slots[col];
+ const db_text *y = (const db_text *)(uintptr_t)b->slots[col];
+ if (!x || !y) {
+ if (x != y) return 0;
+ } else if (x->len != y->len || memcmp(x->bytes, y->bytes, x->len) != 0)
+ return 0;
+ } else if (a->slots[col] != b->slots[col])
+ return 0;
+ }
+ return 1;
+}
+
+static db_ibucket *idx_bucket(db_index *ix, uint64_t h, int create) {
+ if (ix->bcap == 0) {
+ if (!create) return NULL;
+ ix->buckets = calloc(64, sizeof(db_ibucket));
+ if (!ix->buckets) return NULL;
+ ix->bcap = 64;
+ }
+ if (create && ix->blen * 10 >= ix->bcap * 7) {
+ size_t ncap = ix->bcap * 2;
+ db_ibucket *nb = calloc(ncap, sizeof(db_ibucket));
+ if (!nb) return NULL;
+ for (size_t i = 0; i < ix->bcap; i++) {
+ if (!ix->buckets[i].hash) continue;
+ size_t j = ix->buckets[i].hash & (ncap - 1);
+ while (nb[j].hash) j = (j + 1) & (ncap - 1);
+ nb[j] = ix->buckets[i];
+ }
+ free(ix->buckets);
+ ix->buckets = nb;
+ ix->bcap = ncap;
+ }
+ size_t j = h & (ix->bcap - 1);
+ while (ix->buckets[j].hash) {
+ if (ix->buckets[j].hash == h) return &ix->buckets[j];
+ j = (j + 1) & (ix->bcap - 1);
+ }
+ if (!create) return NULL;
+ ix->buckets[j].hash = h;
+ ix->blen++;
+ return &ix->buckets[j];
+}
+
+/* Add [r] to every index; unique violation reports which without mutating
+ * anything (checks run before any add). 0 ok, DB_ERR_* otherwise. */
+static int idx_add_row(wo_db *db, db_table *t, db_row *r) {
+ const wo_classdesc *c = &db->classes[t->class_id];
+ for (uint32_t x = 0; x < t->index_cnt; x++) {
+ db_index *ix = &t->indexes[x];
+ if (!(ix->flags & 1u)) continue;
+ db_ibucket *b = idx_bucket(ix, idx_hash(c, ix, r), 0);
+ if (!b) continue;
+ for (uint32_t i = 0; i < b->len; i++) {
+ db_row *other = wo_row_ptr(db, t->class_id, b->ids[i]);
+ if (other && idx_cols_equal(c, ix, r, other)) return DB_ERR_UNIQUE;
+ }
+ }
+ for (uint32_t x = 0; x < t->index_cnt; x++) {
+ db_index *ix = &t->indexes[x];
+ db_ibucket *b = idx_bucket(ix, idx_hash(c, ix, r), 1);
+ if (!b) return DB_ERR_OOM;
+ if (b->len == b->cap) {
+ uint32_t ncap = b->cap ? b->cap * 2 : 4;
+ uint64_t *ni = realloc(b->ids, (size_t)ncap * 8u);
+ if (!ni) return DB_ERR_OOM;
+ b->ids = ni;
+ b->cap = ncap;
+ }
+ b->ids[b->len++] = r->id;
+ }
+ return 0;
+}
+
+static void idx_remove_row(wo_db *db, db_table *t, db_row *r) {
+ const wo_classdesc *c = &db->classes[t->class_id];
+ for (uint32_t x = 0; x < t->index_cnt; x++) {
+ db_index *ix = &t->indexes[x];
+ db_ibucket *b = idx_bucket(ix, idx_hash(c, ix, r), 0);
+ if (!b) continue;
+ for (uint32_t i = 0; i < b->len; i++)
+ if (b->ids[i] == r->id) {
+ b->ids[i] = b->ids[--b->len];
+ break;
+ }
+ }
+}
+
+int wo_db_init(wo_db *db, const wo_classdesc *classes, uint32_t class_cnt,
+ uint32_t shard, uint32_t nshards) {
+ if (!nshards || shard >= nshards) return -1;
+ memset(db, 0, sizeof(*db));
+ db->classes = classes;
+ db->class_cnt = class_cnt;
+ db->shard = shard;
+ db->nshards = nshards;
+ db->tables = calloc(class_cnt ? class_cnt : 1, sizeof(db_table));
+ if (!db->tables) return -1;
+ g_classes = classes;
+ return 0;
+}
+
+static void table_destroy(wo_db *db, db_table *t) {
+ /* free every live row's engine-owned values, then the slabs */
+ const wo_classdesc *c = &db->classes[t->class_id];
+ for (uint32_t s = 0; s < t->slab_cnt; s++) {
+ for (uint32_t i = 0; i < DB_SLAB_ROWS; i++) {
+ uint32_t g = s * DB_SLAB_ROWS + i;
+ if (!(t->bitmap[g >> 6] & (1ull << (g & 63)))) continue;
+ db_row *r = (db_row *)(t->slabs[s] + (size_t)i * t->row_size);
+ for (uint32_t f = 0; f < c->field_cnt; f++)
+ db_val_free(c->kinds[f], r->slots[f]);
+ }
+ free(t->slabs[s]);
+ }
+ free(t->slabs);
+ free(t->bitmap);
+ free(t->free_slots);
+ free(t->hkeys);
+ free(t->hvals);
+ for (uint32_t x = 0; x < t->index_cnt; x++) {
+ for (size_t b = 0; b < t->indexes[x].bcap; b++) free(t->indexes[x].buckets[b].ids);
+ free(t->indexes[x].buckets);
+ }
+ free(t->indexes);
+}
+
+void wo_db_destroy(wo_db *db) {
+ if (!db->tables) return;
+ for (uint32_t i = 0; i < db->class_cnt; i++)
+ if (db->tables[i].slab_cnt || db->tables[i].hkeys) table_destroy(db, &db->tables[i]);
+ free(db->tables);
+ db->tables = NULL;
+}
+
+static db_table *table_of(wo_db *db, uint32_t class_id) {
+ if (class_id >= db->class_cnt) return NULL;
+ db_table *t = &db->tables[class_id];
+ if (!t->row_size) { /* lazy init on first touch */
+ const wo_classdesc *c = &db->classes[class_id];
+ t->class_id = class_id;
+ t->row_size = sizeof(db_row) + (size_t)c->field_cnt * 8u;
+ t->next_id = db->shard + 1; /* S+1, then += N: interleaved, local-only */
+ if (c->idx_cnt) {
+ t->indexes = calloc(c->idx_cnt, sizeof(db_index));
+ if (!t->indexes) return NULL;
+ const uint32_t *im = c->idx_meta;
+ for (uint32_t x = 0; x < c->idx_cnt; x++) {
+ t->indexes[x].flags = im[0];
+ t->indexes[x].col_cnt = im[1];
+ t->indexes[x].cols = im + 2;
+ im += 2 + im[1];
+ }
+ t->index_cnt = c->idx_cnt;
+ }
+ }
+ return t;
+}
+
+static db_row *slot_row(db_table *t, uint32_t g) {
+ return (db_row *)(t->slabs[g / DB_SLAB_ROWS] + (size_t)(g % DB_SLAB_ROWS) * t->row_size);
+}
+
+/* Pick the slot a new row lands in: recycled first, else the next free bit,
+ * else grow a slab. Returns the global slot or UINT32_MAX on OOM. */
+static uint32_t slot_alloc(db_table *t) {
+ if (t->free_cnt) return t->free_slots[--t->free_cnt];
+ uint32_t total = t->slab_cnt * DB_SLAB_ROWS;
+ for (uint32_t g = 0; g < total; g++) /* cheap at slab granularity: only
+ reached when free list is empty, and the bitmap scan is bounded by
+ one word test per 64 slots */
+ if (!(t->bitmap[g >> 6] & (1ull << (g & 63)))) return g;
+ /* grow */
+ if (t->slab_cnt == t->slab_cap) {
+ uint32_t ncap = t->slab_cap ? t->slab_cap * 2 : 4;
+ uint8_t **ns = realloc(t->slabs, (size_t)ncap * sizeof(uint8_t *));
+ if (!ns) return UINT32_MAX;
+ t->slabs = ns;
+ t->slab_cap = ncap;
+ }
+ uint8_t *slab = malloc((size_t)DB_SLAB_ROWS * t->row_size);
+ if (!slab) return UINT32_MAX;
+ size_t nwords = ((size_t)(t->slab_cnt + 1) * DB_SLAB_ROWS + 63) / 64;
+ uint64_t *nb = realloc(t->bitmap, nwords * 8);
+ if (!nb) {
+ free(slab);
+ return UINT32_MAX;
+ }
+ memset(nb + ((size_t)t->slab_cnt * DB_SLAB_ROWS) / 64, 0,
+ (nwords - ((size_t)t->slab_cnt * DB_SLAB_ROWS) / 64) * 8);
+ t->bitmap = nb;
+ t->slabs[t->slab_cnt] = slab;
+ return t->slab_cnt++ * DB_SLAB_ROWS;
+}
+
+uint64_t wo_row_insert(wo_db *db, uint32_t class_id, const uint64_t *vals,
+ const char **msg, int *err_kind) {
+ if (err_kind) *err_kind = DB_ERR_MISC;
+ db_table *t = table_of(db, class_id);
+ if (!t) {
+ *msg = "no such class";
+ return 0;
+ }
+ const wo_classdesc *c = &db->classes[class_id];
+ uint32_t g = slot_alloc(t);
+ if (g == UINT32_MAX) {
+ *msg = "out of memory growing a table";
+ return 0;
+ }
+ db_row *r = slot_row(t, g);
+ r->class_id = class_id;
+ r->flags = 0;
+ int ok = 1;
+ uint32_t i = 0;
+ for (; i < c->field_cnt; i++) {
+ r->slots[i] = db_val_encode(db->classes, c->kinds[i], vals[i], &ok, msg);
+ if (!ok) {
+ if (err_kind) *err_kind = DB_ERR_BADKIND;
+ break;
+ }
+ }
+ if (!ok) {
+ for (uint32_t j = 0; j < i; j++) db_val_free(c->kinds[j], r->slots[j]);
+ /* slot never became live: recycle it (bitmap bit was never set) */
+ if (t->free_cnt == t->free_cap) {
+ uint32_t ncap = t->free_cap ? t->free_cap * 2 : 16;
+ uint32_t *nf = realloc(t->free_slots, (size_t)ncap * 4);
+ if (nf) {
+ t->free_slots = nf;
+ t->free_cap = ncap;
+ }
+ }
+ if (t->free_cnt < t->free_cap) t->free_slots[t->free_cnt++] = g;
+ return 0;
+ }
+ r->id = t->next_id;
+ t->next_id += db->nshards;
+ if (hput(t, r->id, (uint64_t)g + 1) != 0) {
+ for (uint32_t j = 0; j < c->field_cnt; j++) db_val_free(c->kinds[j], r->slots[j]);
+ if (err_kind) *err_kind = DB_ERR_OOM;
+ *msg = "out of memory indexing a row";
+ return 0;
+ }
+ t->bitmap[g >> 6] |= 1ull << (g & 63);
+ t->count++;
+ /* THE index hook (Task 4): inside the choke point, never anywhere else.
+ A unique violation un-applies the row entirely — id never handed out
+ twice matters less than the row never having existed. */
+ int irc = idx_add_row(db, t, r);
+ if (irc != 0) {
+ t->bitmap[g >> 6] &= ~(1ull << (g & 63));
+ hdel(t, r->id);
+ t->count--;
+ t->next_id -= db->nshards; /* the id was never observable: reclaim it */
+ for (uint32_t j = 0; j < c->field_cnt; j++) db_val_free(c->kinds[j], r->slots[j]);
+ if (t->free_cnt < t->free_cap) t->free_slots[t->free_cnt++] = g;
+ if (err_kind) *err_kind = irc;
+ *msg = irc == DB_ERR_UNIQUE ? "unique index violation" : "out of memory indexing a row";
+ return 0;
+ }
+ if (err_kind) *err_kind = DB_ERR_NONE;
+ return r->id;
+}
+
+db_row *wo_row_ptr(wo_db *db, uint32_t class_id, uint64_t id) {
+ if (class_id >= db->class_cnt) return NULL;
+ db_table *t = &db->tables[class_id];
+ if (!t->row_size) return NULL;
+ uint64_t s1 = hget(t, id);
+ if (!s1) return NULL;
+ return slot_row(t, (uint32_t)(s1 - 1));
+}
+
+int wo_row_read(wo_db *db, wo_rt *rt, uint32_t class_id, uint64_t id,
+ uint64_t *out_vals, const char **msg) {
+ db_row *r = wo_row_ptr(db, class_id, id);
+ if (!r) return -1;
+ const wo_classdesc *c = &db->classes[class_id];
+ int ok = 1;
+ for (uint32_t i = 0; i < c->field_cnt; i++) {
+ out_vals[i] = db_val_decode(rt, c->kinds[i], r->slots[i], &ok, msg);
+ if (!ok) return -2;
+ }
+ return 0;
+}
+
+db_row *wo_row_create_raw(wo_db *db, uint32_t class_id, uint64_t id) {
+ db_table *t = table_of(db, class_id);
+ if (!t || !id) return NULL;
+ if (hget(t, id)) return NULL; /* duplicate id: corruption, not a tear */
+ uint32_t g = slot_alloc(t);
+ if (g == UINT32_MAX) return NULL;
+ db_row *r = slot_row(t, g);
+ r->id = id;
+ r->class_id = class_id;
+ r->flags = 0;
+ memset(r->slots, 0, t->row_size - sizeof(db_row));
+ if (hput(t, id, (uint64_t)g + 1) != 0) return NULL;
+ t->bitmap[g >> 6] |= 1ull << (g & 63);
+ t->count++;
+ /* keep the interleave: only ids this shard owns move its counter */
+ if ((id - 1) % db->nshards == db->shard && id >= t->next_id)
+ t->next_id = id + db->nshards;
+ /* indexes: NOT here — the slots are still zero. wal.c fills them and
+ then calls wo_row_raw_commit, which is where replayed rows re-index. */
+ return r;
+}
+
+int wo_row_raw_commit(wo_db *db, uint32_t class_id, db_row *r) {
+ db_table *t = &db->tables[class_id];
+ return idx_add_row(db, t, r) == 0 ? 0 : -1;
+}
+
+void wo_db_val_free(wo_db *db, uint8_t kind, uint64_t v) {
+ (void)db;
+ db_val_free(kind, v);
+}
+
+uint64_t wo_val_decode_vm(wo_db *db, wo_rt *rt, uint8_t kind, uint64_t engine_val,
+ int *ok, const char **msg) {
+ (void)db;
+ return db_val_decode(rt, kind, engine_val, ok, msg);
+}
+
+int wo_row_update_field(wo_db *db, uint32_t class_id, uint64_t id, uint32_t field,
+ uint64_t vm_val, const char **msg, int *err_kind) {
+ if (err_kind) *err_kind = DB_ERR_MISC;
+ db_row *r = wo_row_ptr(db, class_id, id);
+ if (!r) {
+ *msg = "no such row";
+ return -1;
+ }
+ const wo_classdesc *c = &db->classes[class_id];
+ if (field >= c->field_cnt) {
+ *msg = "no such field";
+ return -1;
+ }
+ db_table *t = &db->tables[class_id];
+ int ok = 1;
+ uint64_t nv = db_val_encode(db->classes, c->kinds[field], vm_val, &ok, msg);
+ if (!ok) {
+ if (err_kind) *err_kind = DB_ERR_BADKIND;
+ return -1;
+ }
+ /* indexes containing this column: unique checks against the NEW value
+ run first, against a shadow of the row, before anything mutates */
+ uint64_t old = r->slots[field];
+ r->slots[field] = nv;
+ for (uint32_t x = 0; x < t->index_cnt; x++) {
+ db_index *ix = &t->indexes[x];
+ if (!(ix->flags & 1u)) continue;
+ int touches = 0;
+ for (uint32_t i = 0; i < ix->col_cnt; i++)
+ if (ix->cols[i] == field) touches = 1;
+ if (!touches) continue;
+ db_ibucket *b = idx_bucket(ix, idx_hash(c, ix, r), 0);
+ if (!b) continue;
+ for (uint32_t i = 0; i < b->len; i++) {
+ if (b->ids[i] == id) continue;
+ db_row *other = wo_row_ptr(db, class_id, b->ids[i]);
+ if (other && idx_cols_equal(c, ix, r, other)) {
+ r->slots[field] = old; /* untouched, promised */
+ db_val_free(c->kinds[field], nv);
+ if (err_kind) *err_kind = DB_ERR_UNIQUE;
+ *msg = "unique index violation";
+ return -1;
+ }
+ }
+ }
+ /* commit: fix every index containing the column (old entry out under
+ the OLD value's hash, new entry in), then free the old value */
+ r->slots[field] = old;
+ for (uint32_t x = 0; x < t->index_cnt; x++) {
+ db_index *ix = &t->indexes[x];
+ int touches = 0;
+ for (uint32_t i = 0; i < ix->col_cnt; i++)
+ if (ix->cols[i] == field) touches = 1;
+ if (!touches) continue;
+ db_ibucket *b = idx_bucket(ix, idx_hash(c, ix, r), 0);
+ if (b)
+ for (uint32_t i = 0; i < b->len; i++)
+ if (b->ids[i] == id) {
+ b->ids[i] = b->ids[--b->len];
+ break;
+ }
+ }
+ r->slots[field] = nv;
+ for (uint32_t x = 0; x < t->index_cnt; x++) {
+ db_index *ix = &t->indexes[x];
+ int touches = 0;
+ for (uint32_t i = 0; i < ix->col_cnt; i++)
+ if (ix->cols[i] == field) touches = 1;
+ if (!touches) continue;
+ db_ibucket *b = idx_bucket(ix, idx_hash(c, ix, r), 1);
+ if (b) {
+ if (b->len == b->cap) {
+ uint32_t ncap = b->cap ? b->cap * 2 : 4;
+ uint64_t *ni = realloc(b->ids, (size_t)ncap * 8u);
+ if (ni) {
+ b->ids = ni;
+ b->cap = ncap;
+ }
+ }
+ if (b->len < b->cap) b->ids[b->len++] = id;
+ }
+ }
+ db_val_free(c->kinds[field], old);
+ if (err_kind) *err_kind = DB_ERR_NONE;
+ return 0;
+}
+
+int wo_row_has_referrers(wo_db *db, uint32_t class_id, uint64_t id) {
+ if (!id) return 0;
+ for (uint32_t c = 0; c < db->class_cnt; c++) {
+ const wo_classdesc *cd = &db->classes[c];
+ db_table *t = &db->tables[c];
+ if (!t->row_size || !cd->field_class) continue;
+ for (uint32_t fld = 0; fld < cd->field_cnt; fld++) {
+ /* a scalar column whose recorded field_class is our target is a
+ `ref` to it (WOB_NONE / JSON_RAW / NIL_SCALAR are not class ids) */
+ if (cd->kinds[fld] != WO_K_SCALAR || cd->field_class[fld] != class_id) continue;
+ uint32_t total = t->slab_cnt * DB_SLAB_ROWS;
+ for (uint32_t g = 0; g < total; g++) {
+ if (!(t->bitmap[g >> 6] & (1ull << (g & 63)))) continue;
+ db_row *r = (db_row *)(t->slabs[g / DB_SLAB_ROWS] +
+ (size_t)(g % DB_SLAB_ROWS) * t->row_size);
+ if (r->slots[fld] == id) return 1;
+ }
+ }
+ }
+ return 0;
+}
+
+int wo_row_remove(wo_db *db, uint32_t class_id, uint64_t id) {
+ if (class_id >= db->class_cnt) return -1;
+ db_table *t = &db->tables[class_id];
+ if (!t->row_size) return -1;
+ uint64_t s1 = hget(t, id);
+ if (!s1) return -1;
+ uint32_t g = (uint32_t)(s1 - 1);
+ db_row *r = slot_row(t, g);
+ /* the index hook's remove side: before the row's values die, while the
+ columns are still comparable */
+ idx_remove_row(db, t, r);
+ const wo_classdesc *c = &db->classes[class_id];
+ for (uint32_t i = 0; i < c->field_cnt; i++) db_val_free(c->kinds[i], r->slots[i]);
+ t->bitmap[g >> 6] &= ~(1ull << (g & 63));
+ hdel(t, id);
+ t->count--;
+ if (t->free_cnt == t->free_cap) {
+ uint32_t ncap = t->free_cap ? t->free_cap * 2 : 16;
+ uint32_t *nf = realloc(t->free_slots, (size_t)ncap * 4);
+ if (!nf) return 0; /* slot simply not recycled; bitmap still frees it */
+ t->free_slots = nf;
+ t->free_cap = ncap;
+ }
+ t->free_slots[t->free_cnt++] = g;
+ return 0;
+}
diff --git a/database/src/table.h b/database/src/table.h
new file mode 100644
index 0000000..cf0dd62
--- /dev/null
+++ b/database/src/table.h
@@ -0,0 +1,196 @@
+/* table.h — class-shaped row storage (iteration 9, Task 1).
+ *
+ * The engine and the VM heap are two memory worlds crossed only by copy
+ * (the 9b design's section 6): a row stores NO VM pointer. Every field
+ * lands in one 8-byte slot, kind-driven:
+ *
+ * SCALAR the 8 bytes themselves (WO_NIL_SCALAR spells a ?scalar's nil)
+ * TEXT engine-owned db_text* (0 = nil)
+ * OWNED engine-owned db_rec* — the object flattened by value,
+ * recursively, through these same rules (0 = nil)
+ * MULTI engine-owned db_multi* — elements encoded element-wise
+ * MAP engine-owned db_map* — keys and values encoded pair-wise
+ * GCREF never stored: the compiler rejects it (the GC bulkhead);
+ * the engine refuses it defensively as an encode error
+ *
+ * `ref T` is a SCALAR at this layer — the target row's id, an ordinary
+ * number the compiler produced; the engine learns nothing about it until
+ * the FK checks (9b plan, Task 3).
+ *
+ * Row layout: a 16-byte header (id, class, flags) then field_cnt 8-byte
+ * slots — deliberately the VM object layout's shape, so encode/decode walk
+ * the same class-table kinds the VM walks. Rows live in per-class SLABS
+ * (fixed-count, malloc'd, never moved: a row's address is stable for its
+ * lifetime, which is what lets 9b hand out loop-scoped row views). A
+ * per-table bitmap tracks occupancy; removed slots go on a free list and
+ * are reused before any slab grows. The id->row map is an open-addressing
+ * hash owned by the table.
+ *
+ * Id discipline (the c-runtime plan's shipped behavior): per table, per
+ * shard, ids interleave — shard S of N allocates S+1, S+1+N, S+1+2N, … —
+ * so creation is coordination-free and a row's owner shard is (id-1) % N.
+ * Milestone runs at N=1 (iteration 8 not yet landed); everything here is
+ * N-parametric and degenerates cleanly.
+ *
+ * CHOKE POINT DOCTRINE: wo_row_insert / wo_row_remove are the only paths
+ * that touch storage. Task 4's secondary indexes hook exactly these two
+ * functions; anything else mutating a slab is a defect by definition.
+ */
+#ifndef WO_TABLE_H
+#define WO_TABLE_H
+
+#include "obj.h" /* wo_rt, wo_classdesc, kinds, wo_str, containers */
+
+/* ---- engine-owned value shapes (all malloc'd, all reachable only from
+ * row slots, all freed through db_val_free) ---- */
+
+typedef struct db_text {
+ uint32_t len;
+ char bytes[]; /* len bytes, no NUL */
+} db_text;
+
+typedef struct db_rec { /* an owned object flattened by value */
+ uint32_t class_id; /* index into the SAME class table the VM uses */
+ uint32_t _pad;
+ uint64_t slots[]; /* field_cnt slots, encoded by these rules */
+} db_rec;
+
+typedef struct db_multi {
+ uint8_t elem_kind;
+ uint32_t len;
+ uint64_t items[];
+} db_multi;
+
+typedef struct db_map {
+ uint8_t key_kind, val_kind;
+ uint32_t len;
+ uint64_t kv[]; /* len pairs: k0 v0 k1 v1 … */
+} db_map;
+
+/* ---- rows and tables ---- */
+
+typedef struct db_row {
+ uint64_t id;
+ uint32_t class_id;
+ uint32_t flags; /* reserved (0) */
+ uint64_t slots[];
+} db_row;
+
+#define DB_SLAB_ROWS 256u
+
+/* Secondary index (iteration 9, Task 4): built from the class table's v3
+ * metadata at first touch, maintained ONLY inside the row choke points.
+ * Hash multimap: bucket per column-value hash, ids within; equality is
+ * re-checked against the actual rows on the unique path (a hash is a hint,
+ * never an answer). */
+typedef struct db_ibucket {
+ uint64_t hash;
+ uint64_t *ids;
+ uint32_t len, cap;
+} db_ibucket;
+
+typedef struct db_index {
+ uint32_t flags; /* bit0 = unique */
+ uint32_t col_cnt;
+ const uint32_t *cols; /* into the loader's idx pool */
+ db_ibucket *buckets; /* open addressing by hash; hash==0 stored as 1 */
+ size_t bcap, blen;
+} db_index;
+
+/* wo_row_insert failure classes — *msg carries the sentence, this carries
+ * the machine-readable kind so db.c maps to the right trap. */
+enum { DB_ERR_NONE = 0, DB_ERR_OOM = 1, DB_ERR_BADKIND = 2, DB_ERR_UNIQUE = 3, DB_ERR_MISC = 4 };
+
+typedef struct db_table {
+ uint32_t class_id;
+ size_t row_size; /* 16 + field_cnt * 8 */
+ /* slabs of DB_SLAB_ROWS rows each; addresses stable forever */
+ uint8_t **slabs;
+ uint32_t slab_cnt, slab_cap;
+ uint64_t *bitmap; /* one bit per slot, slab-major */
+ /* removed slots, reused LIFO before any slab grows */
+ uint32_t *free_slots;
+ uint32_t free_cnt, free_cap;
+ uint64_t next_id; /* next id THIS shard hands out for this table */
+ uint64_t count; /* live rows */
+ /* id -> (global slot + 1); 0 = empty. Open addressing, pow2. */
+ uint64_t *hkeys;
+ uint64_t *hvals;
+ size_t hcap, hlen;
+ /* secondary indexes, from the class table's v3 metadata */
+ db_index *indexes;
+ uint32_t index_cnt;
+} db_table;
+
+typedef struct wo_db {
+ const wo_classdesc *classes;
+ uint32_t class_cnt;
+ uint32_t shard, nshards; /* S of N; ids interleave S+1, S+1+N, … */
+ db_table *tables; /* class_cnt entries, created lazily on first insert */
+} wo_db;
+
+/* 0 ok, -1 alloc failure. nshards >= 1, shard < nshards. */
+int wo_db_init(wo_db *db, const wo_classdesc *classes, uint32_t class_cnt,
+ uint32_t shard, uint32_t nshards);
+void wo_db_destroy(wo_db *db);
+
+/* Insert: encode field_cnt VM values (register words, kinds from the class
+ * table) into a fresh row. Returns the new id, or 0 with *msg set (OOM, or
+ * a GCREF field — which the compiler should have refused upstream). */
+uint64_t wo_row_insert(wo_db *db, uint32_t class_id, const uint64_t *vals,
+ const char **msg, int *err_kind);
+
+/* Read: decode the row's fields into VM values freshly allocated from
+ * [rt] — always copies, never a pointer into the slab (the out-gate).
+ * 0 ok, -1 no such row, -2 OOM (*msg set). */
+int wo_row_read(wo_db *db, wo_rt *rt, uint32_t class_id, uint64_t id,
+ uint64_t *out_vals, const char **msg);
+
+/* Remove: free the row's engine-owned field values, clear the slot, recycle
+ * it. 0 ok, -1 no such row. */
+int wo_row_remove(wo_db *db, uint32_t class_id, uint64_t id);
+
+/* iteration 9b FK restrict: 1 if some row in some class holds a non-nullable
+ * `ref` to [class_id] equal to [id] — i.e. deleting this row would dangle a
+ * reference. The compiler records a ref field's target class in the class
+ * table's field_class metadata; this scans those columns. Correctness-first
+ * (a full scan of referencing tables); the backlink index is the later
+ * optimization the spec records. */
+int wo_row_has_referrers(wo_db *db, uint32_t class_id, uint64_t id);
+
+/* Update one field in place (iteration 9 Task 5): encode the VM value,
+ * swap it into the slot, keep every index containing that column honest —
+ * remove-old/add-new with the unique re-check running BEFORE anything
+ * mutates, so a violating update leaves the row untouched. 0 ok, -1 no
+ * such row / bad field, DB_ERR_* codes via *err_kind like insert. */
+int wo_row_update_field(wo_db *db, uint32_t class_id, uint64_t id, uint32_t field,
+ uint64_t vm_val, const char **msg, int *err_kind);
+
+/* Borrowed row pointer for engine-internal callers (the WAL writes a row's
+ * encoded bytes; indexes read key slots). NULL = no such row. NEVER handed
+ * to the VM. */
+db_row *wo_row_ptr(wo_db *db, uint32_t class_id, uint64_t id);
+
+/* Engine-internal, for WAL replay only: create a row with a FIXED id,
+ * slots zeroed — the caller (wal.c) fills them with engine-encoded values
+ * it built while decoding. Advances the table's next_id past [id] when the
+ * id belongs to this shard, so post-replay inserts never collide. NULL =
+ * OOM or duplicate id (corruption beyond a torn tail). */
+db_row *wo_row_create_raw(wo_db *db, uint32_t class_id, uint64_t id);
+
+/* Engine-internal: free one engine-encoded slot value of [kind] (wal.c's
+ * decode error paths). */
+void wo_db_val_free(wo_db *db, uint8_t kind, uint64_t v);
+
+/* Decode one engine slot value to a FRESH VM value in [rt] (the out-gate:
+ * always a copy). The query builtins' field reads go through this. */
+uint64_t wo_val_decode_vm(wo_db *db, wo_rt *rt, uint8_t kind, uint64_t engine_val,
+ int *ok, const char **msg);
+
+/* Engine-internal, replay only: after wal.c fills a raw row's slots, this
+ * runs the index maintenance the normal insert runs inline — including the
+ * unique check, whose violation during replay is corruption, not data
+ * (0 ok, -1). */
+int wo_row_raw_commit(wo_db *db, uint32_t class_id, db_row *r);
+
+#endif /* WO_TABLE_H */
diff --git a/database/src/wal.c b/database/src/wal.c
new file mode 100644
index 0000000..4f58161
--- /dev/null
+++ b/database/src/wal.c
@@ -0,0 +1,474 @@
+/* pread/pwrite/fdatasync/posix_fallocate under -std=c11 */
+#define _POSIX_C_SOURCE 200809L
+
+#include "wal.h"
+
+#include
+#include
+#include
+#include
+#include
+
+/* ---- crc32 (poly 0xEDB88320) — ported from runtime/wo-rt.c ------------- */
+
+static uint32_t crc_table[256];
+static int crc_ready;
+
+static void crc32_init(void) {
+ for (uint32_t i = 0; i < 256; i++) {
+ uint32_t c = i;
+ for (int k = 0; k < 8; k++) c = (c & 1) ? 0xEDB88320u ^ (c >> 1) : c >> 1;
+ crc_table[i] = c;
+ }
+ crc_ready = 1;
+}
+
+static uint32_t crc32(const void *buf, size_t len) {
+ if (!crc_ready) crc32_init();
+ const uint8_t *p = buf;
+ uint32_t c = 0xFFFFFFFFu;
+ while (len--) c = crc_table[(c ^ *p++) & 0xFF] ^ (c >> 8);
+ return c ^ 0xFFFFFFFFu;
+}
+
+/* ---- byte buffer -------------------------------------------------------- */
+
+typedef struct {
+ uint8_t *b;
+ size_t len, cap;
+ int oom;
+} wbuf;
+
+static void wput(wbuf *w, const void *p, size_t n) {
+ if (w->oom) return;
+ if (w->len + n > w->cap) {
+ size_t nc = w->cap ? w->cap * 2 : 256;
+ while (nc < w->len + n) nc *= 2;
+ uint8_t *nb = realloc(w->b, nc);
+ if (!nb) {
+ w->oom = 1;
+ return;
+ }
+ w->b = nb;
+ w->cap = nc;
+ }
+ memcpy(w->b + w->len, p, n);
+ w->len += n;
+}
+
+static void wput_u8(wbuf *w, uint8_t v) { wput(w, &v, 1); }
+static void wput_u32(wbuf *w, uint32_t v) { wput(w, &v, 4); }
+static void wput_u64(wbuf *w, uint64_t v) { wput(w, &v, 8); }
+
+/* bounds-checked reader */
+typedef struct {
+ const uint8_t *p, *end;
+ int bad;
+} rbuf;
+
+static int rtake(rbuf *r, void *out, size_t n) {
+ if (r->bad || (size_t)(r->end - r->p) < n) {
+ r->bad = 1;
+ return -1;
+ }
+ memcpy(out, r->p, n);
+ r->p += n;
+ return 0;
+}
+
+static uint8_t rd_u8(rbuf *r) {
+ uint8_t v = 0;
+ rtake(r, &v, 1);
+ return v;
+}
+static uint32_t rd_u32(rbuf *r) {
+ uint32_t v = 0;
+ rtake(r, &v, 4);
+ return v;
+}
+static uint64_t rd_u64(rbuf *r) {
+ uint64_t v = 0;
+ rtake(r, &v, 8);
+ return v;
+}
+
+#define WAL_NIL_TEXT 0xFFFFFFFFu
+
+/* ---- engine-value <-> bytes (kind-driven, mirrors table.c's encoding) --- */
+
+static void enc_val(wbuf *w, const wo_classdesc *classes, uint8_t kind, uint64_t v) {
+ switch (kind) {
+ case WO_K_SCALAR: wput_u64(w, v); return;
+ case WO_K_TEXT: {
+ if (!v) {
+ wput_u32(w, WAL_NIL_TEXT);
+ return;
+ }
+ const db_text *t = (const db_text *)(uintptr_t)v;
+ wput_u32(w, t->len);
+ wput(w, t->bytes, t->len);
+ return;
+ }
+ case WO_K_OWNED: {
+ if (!v) {
+ wput_u8(w, 0);
+ return;
+ }
+ const db_rec *r = (const db_rec *)(uintptr_t)v;
+ wput_u8(w, 1);
+ wput_u32(w, r->class_id);
+ const wo_classdesc *c = &classes[r->class_id];
+ for (uint32_t i = 0; i < c->field_cnt; i++)
+ enc_val(w, classes, c->kinds[i], r->slots[i]);
+ return;
+ }
+ case WO_K_MULTI: {
+ if (!v) {
+ wput_u8(w, 0);
+ return;
+ }
+ const db_multi *m = (const db_multi *)(uintptr_t)v;
+ wput_u8(w, 1);
+ wput_u8(w, m->elem_kind);
+ wput_u32(w, m->len);
+ for (uint32_t i = 0; i < m->len; i++) enc_val(w, classes, m->elem_kind, m->items[i]);
+ return;
+ }
+ case WO_K_MAP: {
+ if (!v) {
+ wput_u8(w, 0);
+ return;
+ }
+ const db_map *m = (const db_map *)(uintptr_t)v;
+ wput_u8(w, 1);
+ wput_u8(w, m->key_kind);
+ wput_u8(w, m->val_kind);
+ wput_u32(w, m->len);
+ for (uint32_t i = 0; i < m->len; i++) {
+ enc_val(w, classes, m->key_kind, m->kv[2 * i]);
+ enc_val(w, classes, m->val_kind, m->kv[2 * i + 1]);
+ }
+ return;
+ }
+ default: return; /* GCREF never stored, so never logged */
+ }
+}
+
+/* Decode one value into an engine-owned allocation. Returns 0 on success
+ * with *out set (0 = genuine nil); -1 on truncation/corruption/OOM — the
+ * caller frees what it already built. */
+static int dec_val(rbuf *r, wo_db *db, uint8_t kind, uint64_t *out) {
+ *out = 0;
+ switch (kind) {
+ case WO_K_SCALAR: {
+ uint64_t v = rd_u64(r);
+ if (r->bad) return -1;
+ *out = v;
+ return 0;
+ }
+ case WO_K_TEXT: {
+ uint32_t len = rd_u32(r);
+ if (r->bad) return -1;
+ if (len == WAL_NIL_TEXT) return 0;
+ if ((size_t)(r->end - r->p) < len) return -1;
+ db_text *t = malloc(sizeof(db_text) + len);
+ if (!t) return -1;
+ t->len = len;
+ memcpy(t->bytes, r->p, len);
+ r->p += len;
+ *out = (uint64_t)(uintptr_t)t;
+ return 0;
+ }
+ case WO_K_OWNED: {
+ uint8_t tag = rd_u8(r);
+ if (r->bad) return -1;
+ if (!tag) return 0;
+ uint32_t cid = rd_u32(r);
+ if (r->bad || cid >= db->class_cnt) return -1;
+ const wo_classdesc *c = &db->classes[cid];
+ db_rec *rec = malloc(sizeof(db_rec) + (size_t)c->field_cnt * 8u);
+ if (!rec) return -1;
+ rec->class_id = cid;
+ rec->_pad = 0;
+ for (uint32_t i = 0; i < c->field_cnt; i++) {
+ if (dec_val(r, db, c->kinds[i], &rec->slots[i]) != 0) {
+ for (uint32_t j = 0; j < i; j++) wo_db_val_free(db, c->kinds[j], rec->slots[j]);
+ free(rec);
+ return -1;
+ }
+ }
+ *out = (uint64_t)(uintptr_t)rec;
+ return 0;
+ }
+ case WO_K_MULTI: {
+ uint8_t tag = rd_u8(r);
+ if (r->bad) return -1;
+ if (!tag) return 0;
+ uint8_t ek = rd_u8(r);
+ uint32_t len = rd_u32(r);
+ if (r->bad || ek > WO_K_MAX) return -1;
+ if (len > (size_t)(r->end - r->p)) return -1; /* each elem >= 1 byte */
+ db_multi *m = malloc(sizeof(db_multi) + (size_t)len * 8u);
+ if (!m) return -1;
+ m->elem_kind = ek;
+ m->len = len;
+ for (uint32_t i = 0; i < len; i++) {
+ if (dec_val(r, db, ek, &m->items[i]) != 0) {
+ for (uint32_t j = 0; j < i; j++) wo_db_val_free(db, ek, m->items[j]);
+ free(m);
+ return -1;
+ }
+ }
+ *out = (uint64_t)(uintptr_t)m;
+ return 0;
+ }
+ case WO_K_MAP: {
+ uint8_t tag = rd_u8(r);
+ if (r->bad) return -1;
+ if (!tag) return 0;
+ uint8_t kk = rd_u8(r), vk = rd_u8(r);
+ uint32_t len = rd_u32(r);
+ if (r->bad || kk > WO_K_MAX || vk > WO_K_MAX) return -1;
+ if (len > (size_t)(r->end - r->p)) return -1;
+ db_map *m = malloc(sizeof(db_map) + (size_t)len * 16u);
+ if (!m) return -1;
+ m->key_kind = kk;
+ m->val_kind = vk;
+ m->len = len;
+ for (uint32_t i = 0; i < len; i++) {
+ if (dec_val(r, db, kk, &m->kv[2 * i]) != 0 ||
+ dec_val(r, db, vk, &m->kv[2 * i + 1]) != 0) {
+ m->len = i; /* free only the fully-built pairs plus a possible key */
+ for (uint32_t j = 0; j < i; j++) {
+ wo_db_val_free(db, kk, m->kv[2 * j]);
+ wo_db_val_free(db, vk, m->kv[2 * j + 1]);
+ }
+ wo_db_val_free(db, kk, m->kv[2 * i]); /* 0 if the key failed */
+ free(m);
+ return -1;
+ }
+ }
+ *out = (uint64_t)(uintptr_t)m;
+ return 0;
+ }
+ default: return -1;
+ }
+}
+
+/* ---- record scan (shared by open, replay, check) ------------------------ */
+
+/* Read the record at [off]. 0 = intact (*len_out = payload length, payload
+ * malloc'd into *payload_out if non-NULL); 1 = end of intact prefix (zero
+ * length, short read, bad crc, missing mark). */
+static int scan_record(int fd, uint64_t off, uint32_t *len_out, uint8_t **payload_out) {
+ uint8_t hdr[8];
+ ssize_t n = pread(fd, hdr, 8, (off_t)off);
+ if (n != 8) return 1;
+ uint32_t len, crc;
+ memcpy(&len, hdr, 4);
+ memcpy(&crc, hdr + 4, 4);
+ if (len == 0 || len > (64u << 20)) return 1; /* preallocated tail or garbage */
+ uint8_t *payload = malloc(len + 4);
+ if (!payload) return 1;
+ n = pread(fd, payload, len + 4, (off_t)(off + 8));
+ if (n != (ssize_t)(len + 4)) {
+ free(payload);
+ return 1;
+ }
+ uint32_t mark;
+ memcpy(&mark, payload + len, 4);
+ if (mark != WO_WAL_MARK || crc32(payload, len) != crc) {
+ free(payload);
+ return 1;
+ }
+ *len_out = len;
+ if (payload_out) *payload_out = payload;
+ else free(payload);
+ return 0;
+}
+
+/* ---- public API ---------------------------------------------------------- */
+
+int wo_wal_open(wo_wal *w, const char *path, uint64_t prealloc) {
+ memset(w, 0, sizeof(*w));
+ w->fd = open(path, O_RDWR | O_CREAT, 0644);
+ if (w->fd < 0) return -1;
+ if (prealloc) {
+ /* best-effort: a filesystem without fallocate still works */
+ (void)posix_fallocate(w->fd, 0, (off_t)prealloc);
+ }
+ /* position after the intact prefix: a torn tail is OVERWRITTEN by the
+ * next append, never appended after */
+ uint64_t off = 0;
+ uint32_t len;
+ while (scan_record(w->fd, off, &len, NULL) == 0) off += 8u + len + 4u;
+ w->off = off;
+ return 0;
+}
+
+void wo_wal_close(wo_wal *w) {
+ if (w->fd >= 0) close(w->fd);
+ free(w->buf);
+ memset(w, 0, sizeof(*w));
+ w->fd = -1;
+}
+
+/* frame one payload into the staged batch */
+static int stage(wo_wal *w, const wbuf *payload) {
+ if (payload->oom) return -1;
+ wbuf rec = {0};
+ wput_u32(&rec, (uint32_t)payload->len);
+ wput_u32(&rec, crc32(payload->b, payload->len));
+ wput(&rec, payload->b, payload->len);
+ wput_u32(&rec, WO_WAL_MARK);
+ if (rec.oom) {
+ free(rec.b);
+ return -1;
+ }
+ if (w->len + rec.len > w->cap) {
+ size_t nc = w->cap ? w->cap * 2 : 4096;
+ while (nc < w->len + rec.len) nc *= 2;
+ uint8_t *nb = realloc(w->buf, nc);
+ if (!nb) {
+ free(rec.b);
+ return -1;
+ }
+ w->buf = nb;
+ w->cap = nc;
+ }
+ memcpy(w->buf + w->len, rec.b, rec.len);
+ w->len += rec.len;
+ free(rec.b);
+ return 0;
+}
+
+int wo_wal_append_insert(wo_wal *w, wo_db *db, uint32_t class_id, uint64_t id) {
+ db_row *r = wo_row_ptr(db, class_id, id);
+ if (!r) return -1; /* commit order: RAM apply comes FIRST */
+ wbuf p = {0};
+ wput_u8(&p, WO_WAL_INSERT);
+ wput_u32(&p, class_id);
+ wput_u64(&p, id);
+ const wo_classdesc *c = &db->classes[class_id];
+ for (uint32_t i = 0; i < c->field_cnt; i++) enc_val(&p, db->classes, c->kinds[i], r->slots[i]);
+ int rc = stage(w, &p);
+ free(p.b);
+ return rc;
+}
+
+int wo_wal_append_update(wo_wal *w, wo_db *db, uint32_t class_id, uint64_t id) {
+ db_row *r = wo_row_ptr(db, class_id, id);
+ if (!r) return -1;
+ wbuf p = {0};
+ wput_u8(&p, WO_WAL_UPDATE);
+ wput_u32(&p, class_id);
+ wput_u64(&p, id);
+ const wo_classdesc *c = &db->classes[class_id];
+ for (uint32_t i = 0; i < c->field_cnt; i++) enc_val(&p, db->classes, c->kinds[i], r->slots[i]);
+ int rc = stage(w, &p);
+ free(p.b);
+ return rc;
+}
+
+int wo_wal_append_remove(wo_wal *w, uint32_t class_id, uint64_t id) {
+ wbuf p = {0};
+ wput_u8(&p, WO_WAL_REMOVE);
+ wput_u32(&p, class_id);
+ wput_u64(&p, id);
+ int rc = stage(w, &p);
+ free(p.b);
+ return rc;
+}
+
+int wo_wal_commit(wo_wal *w) {
+ if (!w->len) return 0;
+ size_t at = 0;
+ while (at < w->len) {
+ ssize_t n = pwrite(w->fd, w->buf + at, w->len - at, (off_t)(w->off + at));
+ if (n < 0) {
+ if (errno == EINTR) continue;
+ return -1;
+ }
+ at += (size_t)n;
+ }
+ if (fdatasync(w->fd) != 0) return -1;
+ w->off += w->len;
+ w->len = 0; /* acked: the batch is durable */
+ return 0;
+}
+
+static int apply_record(wo_db *db, const uint8_t *payload, uint32_t len) {
+ rbuf r = {payload, payload + len, 0};
+ uint8_t kind = rd_u8(&r);
+ uint32_t cid = rd_u32(&r);
+ uint64_t id = rd_u64(&r);
+ if (r.bad || cid >= db->class_cnt) return -1;
+ if (kind == WO_WAL_REMOVE) return wo_row_remove(db, cid, id);
+ if (kind != WO_WAL_INSERT && kind != WO_WAL_UPDATE) return -1;
+ if (kind == WO_WAL_UPDATE) {
+ /* replace: the row must exist (its insert precedes its update in a
+ correct log); anything else is corruption */
+ if (wo_row_remove(db, cid, id) != 0) return -1;
+ }
+ db_row *row = wo_row_create_raw(db, cid, id);
+ if (!row) return -1;
+ const wo_classdesc *c = &db->classes[cid];
+ for (uint32_t i = 0; i < c->field_cnt; i++) {
+ if (dec_val(&r, db, c->kinds[i], &row->slots[i]) != 0) {
+ /* a record that CRC-passed but does not decode is corruption,
+ * not a tear: fail loudly (the row's built slots are freed by
+ * wo_row_remove, which also unregisters the id) */
+ wo_row_remove(db, cid, id);
+ return -1;
+ }
+ }
+ if ((size_t)(r.end - r.p) != 0) { /* trailing bytes = corrupt */
+ wo_row_remove(db, cid, id);
+ return -1;
+ }
+ /* slots are real now: re-index (Task 4). A unique violation during
+ * replay is corruption — the live insert would have refused it. */
+ if (wo_row_raw_commit(db, cid, row) != 0) {
+ wo_row_remove(db, cid, id);
+ return -1;
+ }
+ return 0;
+}
+
+int64_t wo_wal_replay(const char *path, wo_db *db) {
+ int fd = open(path, O_RDONLY);
+ if (fd < 0) return errno == ENOENT ? 0 : -1; /* no WAL yet = fresh boot */
+ uint64_t off = 0;
+ int64_t applied = 0;
+ for (;;) {
+ uint32_t len;
+ uint8_t *payload;
+ if (scan_record(fd, off, &len, &payload) != 0) break; /* intact prefix ends */
+ int rc = apply_record(db, payload, len);
+ free(payload);
+ if (rc != 0) {
+ close(fd);
+ return -1;
+ }
+ off += 8u + len + 4u;
+ applied++;
+ }
+ close(fd);
+ return applied;
+}
+
+int64_t wo_wal_check(const char *path, uint64_t *intact_bytes) {
+ int fd = open(path, O_RDONLY);
+ if (fd < 0) return -1;
+ uint64_t off = 0;
+ int64_t records = 0;
+ for (;;) {
+ uint32_t len;
+ if (scan_record(fd, off, &len, NULL) != 0) break;
+ off += 8u + len + 4u;
+ records++;
+ }
+ if (intact_bytes) *intact_bytes = off;
+ close(fd);
+ return records;
+}
diff --git a/database/src/wal.h b/database/src/wal.h
new file mode 100644
index 0000000..4acbaf2
--- /dev/null
+++ b/database/src/wal.h
@@ -0,0 +1,91 @@
+/* wal.h — typed-row write-ahead log + boot replay (iteration 9, Task 2).
+ *
+ * The c-runtime plan's shipped pattern (phases D/E), generalized to typed
+ * rows. The commit order is doctrine, verbatim:
+ *
+ * RAM apply → wal_append (staged) → wal_commit (write + fdatasync)
+ * → only then is the write ACKNOWLEDGED
+ *
+ * Record framing — replay-whole-or-not-at-all:
+ *
+ * record := len u32 | crc u32 | payload | mark u32
+ * len = payload byte count (never 0; 0 = preallocated tail, stop)
+ * crc = CRC32 of payload
+ * mark = 0x574F4C31 "WOL1" — written LAST, so a record without its
+ * mark is torn by definition
+ * payload := kind u8 | class_id u32 | row_id u64 | body
+ * kind : 1 insert (body = the row's fields, engine encoding below)
+ * 2 remove (no body)
+ * 3 update (reserved for Task 5)
+ *
+ * Field encoding in a body walks the class table's kinds:
+ * SCALAR 8 bytes
+ * TEXT u32 len | bytes (0xFFFFFFFF = nil)
+ * OWNED u8 0 = nil, or u8 1 | u32 class_id | fields recursively
+ * MULTI u8 0 = nil, or u8 1 | u8 elem_kind | u32 len | elements
+ * MAP u8 0 = nil, or u8 1 | u8 kk | u8 vk | u32 len | k v pairs
+ *
+ * Replay decodes payloads STRAIGHT into engine-owned values — the VM heap
+ * is never involved (boot must not depend on a VM existing yet), and rows
+ * re-enter through the same choke-point row API, so Task 4's indexes are
+ * rebuilt for free. A torn tail (short record, bad CRC, missing mark) drops
+ * everything from the tear onward — never a partial record, never a record
+ * after a tear. Little-endian on-disk, matching the .wob loader's platform
+ * note.
+ *
+ * wo_wal_check is the offline oracle the crash battery verifies with: it
+ * walks a WAL file with no engine at all and reports how many records are
+ * intact and where the intact prefix ends. */
+#ifndef WO_WAL_H
+#define WO_WAL_H
+
+#include "table.h"
+
+#define WO_WAL_MARK 0x574F4C31u /* "WOL1" LE */
+
+enum { WO_WAL_INSERT = 1, WO_WAL_REMOVE = 2, WO_WAL_UPDATE = 3 };
+
+typedef struct wo_wal {
+ int fd;
+ uint64_t off; /* next write offset (the intact tail) */
+ /* staged batch: appended by wal_append_*, flushed by wal_commit */
+ uint8_t *buf;
+ size_t len, cap;
+} wo_wal;
+
+/* Open (create if missing) and preallocate [prealloc] bytes (best-effort;
+ * a filesystem without fallocate still works). Positions the write offset
+ * at the end of the INTACT record prefix — an existing file is scanned the
+ * same way replay scans it, so a torn tail is overwritten, not appended
+ * after. 0 ok, -1 errno-style failure. */
+int wo_wal_open(wo_wal *w, const char *path, uint64_t prealloc);
+void wo_wal_close(wo_wal *w);
+
+/* Stage a record for the row that MUST already be applied to RAM (the
+ * commit-order doctrine). Insert/update read the row via wo_row_ptr.
+ * 0 ok, -1 OOM / no such row. */
+int wo_wal_append_insert(wo_wal *w, wo_db *db, uint32_t class_id, uint64_t id);
+int wo_wal_append_remove(wo_wal *w, uint32_t class_id, uint64_t id);
+/* UPDATE re-logs the whole row (KISS: replay replaces — remove + re-create
+ * with the same id; the prefix/suffix delta trick from the survey is a
+ * later optimization, recorded). Call AFTER the RAM update. */
+int wo_wal_append_update(wo_wal *w, wo_db *db, uint32_t class_id, uint64_t id);
+
+/* Write the staged batch and fdatasync — the ack line. Empty batch = ok,
+ * no syscall. 0 ok, -1 write/sync failure (the batch stays staged). */
+int wo_wal_commit(wo_wal *w);
+
+/* Boot replay: apply every intact record to [db] in order. Ids re-enter
+ * exactly as logged; each table's next_id advances past the replayed ids
+ * that belong to this shard. Returns the number of records applied, or -1
+ * on open failure / a record naming an unknown class (corruption beyond
+ * what a torn tail explains). A torn tail is NOT an error: replay applies
+ * the intact prefix and reports it. */
+int64_t wo_wal_replay(const char *path, wo_db *db);
+
+/* Offline verification (no engine): scan [path], count intact records.
+ * *intact_bytes (optional) = where the intact prefix ends. -1 = open
+ * failure. */
+int64_t wo_wal_check(const char *path, uint64_t *intact_bytes);
+
+#endif /* WO_WAL_H */
diff --git a/docs/00-status.md b/docs/00-status.md
index 49c7d2d..5e41ca6 100644
--- a/docs/00-status.md
+++ b/docs/00-status.md
@@ -115,11 +115,18 @@ that sequences its tasks. Read one, approve, then the next starts.
| 7 | [log-watcher proof](stories/language-runtime-database/07-logwatcher-proof.md) | 🔄 **runs; executable in progress** |
| 7b | [Inferred GC + mark-sweep](stories/language-runtime-database/07b-inferred-gc-mark-sweep.md) | ⏸ off the workload's path (no `@gc`) |
| 8 | [Shard-actor runtime](stories/language-runtime-database/08-shard-actor-runtime.md) | ⬜ |
-| 9 | [Database engine](stories/language-runtime-database/09-database-engine.md) | ⬜ |
-| 9b | [`@table`, relations, query](stories/language-runtime-database/09b-table-relations-query.md) | ⬜ needs a spec first |
+| 9 | [Database engine](stories/language-runtime-database/09-database-engine.md) | 🔄 engine complete (storage/WAL/indexes/insert-update-delete); reads land with 9b |
+| 9b | [`@table`, relations, query](stories/language-runtime-database/09b-table-relations-query.md) | 🔄 query surface + relations + FK done (branch query-surface); group-by parked |
+| 9c | [Cross-program tables](stories/language-runtime-database/09c-cross-program-tables.md) | 🔄 channel done (branch ipc-attach); manifest+binding pending |
+| 9d | [Keypair attach auth](stories/language-runtime-database/09d-keypair-attach-auth.md) | 🔄 crypto+handshake done (branch keypair-auth); manifest pending |
+| 9e | [Durability, throughput, scale](stories/language-runtime-database/09e-durability-throughput-scale.md) | ⬜ needs a spec first |
+| 9f | [io_uring group-commit](stories/language-runtime-database/09f-io-uring-commit.md) | ⬜ after 8 + 9e |
+| 9g | [Query grammar corpus](stories/language-runtime-database/09g-query-grammar-corpus.md) | ⬜ needs a spec first |
| 10 | [HTTP service layer](stories/language-runtime-database/10-http-service.md) | ⬜ | Hold |
| 11 | [Fibers](stories/language-runtime-database/11-fibers.md) | ⬜ | Hold |
| 12 | [Blue-green deploy](stories/language-runtime-database/12-blue-green-deploy.md) | ⬜ | Hold |
+| 13 | [Compile-time metaprogramming](stories/language-runtime-database/13-compile-time-metaprogramming.md) | ⬜ needs a spec first |
+| 14 | [skillhost host workload](stories/language-runtime-database/14-skillhost-host-workload.md) | ⬜ gaps recorded (branch query-grammar found skillhost needs no new query grammar); each gap a candidate iteration |
---
@@ -322,7 +329,13 @@ Ecommerce sample (verified 2026-06-13): `api.rest` 17/17 expected statuses pass.
| 7b | Inferred GC + incremental mark-sweep — `@gc` removed, GC-ness inferred, RC retired | [spec](superpowers/specs/2026-08-11-inferred-gc-mark-sweep-design.md) — plan to be written |
| 8 | Shard-actor runtime | [plan 4](superpowers/plans/2026-08-01-shard-actor-vm-runtime.md) |
| 9 | Database engine binding | [plan 5](superpowers/plans/2026-08-01-db-engine-binding.md) |
-| 9b | `@table` + relations + language-integrated query | **no spec yet** — three open forks recorded in the iteration; brainstorm before planning |
+| 9b | `@table` + relations + language-integrated query — comprehension queries, `ref`/`backlink` navigation, GroupBy aggregates; acceptance: new `docs/examples/employee` sample | [spec](superpowers/specs/2026-08-15-table-relations-query-design.md) · [plan](plan/compiler/2026-08-15-employee-relations-query.md) |
+| 9c | Cross-program tables — attach to a running program's database (IPC string in wo.toml, manifest-granted rights, owner stays the single writer) | **no spec yet** — four open forks recorded in the iteration; brainstorm before planning |
+| 9d | Keypair attach auth — mutual challenge–response, grants name public keys, uid superseded | **no spec yet** — four forks recorded; plan folds into 9c's |
+| 9e | Durability + throughput + scale — restart-persistence, read/write benchmark, ~1M rows; the gate every later optimization re-runs | **no spec yet** — four forks recorded; the measurement backbone |
+| 9f | io_uring group-commit write path — batched durability overlapped on shard threads, fsync fallback | **no spec yet** — brainstorm after iterations 8 + 9e |
+| 9g | Query grammar from real embedded-DB corpora — whole-query count + correlated exists, driven by the skillhost SQL catalogue; add only what a corpus uses | **no spec yet** — three forks; may collapse to "confirm len(query) + add exists" |
+| 14 | skillhost host workload — port skillhost (MCP host + confined script runner) to writeonce; drives the missing host capabilities into the open (bounded subprocess, stdin/stdout transport, fs metadata, FFI-vs-out-of-process) | **no spec yet** — gaps recorded in the iteration; each gap brainstormed on demand, bounded-subprocess first |
| 10 | HTTP service layer | [plan 6](superpowers/plans/2026-08-01-http-service-layer.md) |
| 11 | Fibers | vision §3, [blue-green exploration](plan/exploration/blue-green-vm/00-vision.md) |
| 12 | Blue-green deploy | [spec](superpowers/specs/2026-08-03-blue-green-vm-design.md) — plan authored after iterations 9–10 |
@@ -361,13 +374,14 @@ log-watcher proof.
| ⬜ | 15a–15e MCP over streamable HTTP | [15](plan/15-mcp-streamable-http.md) | 15e needs 13c + 09d |
| ⬜ | 16c–16f typed columns, lossless resync, restore, SCRAM | [16](plan/16-postgres-mirror.md) | |
-### Frontend — parked
+### Frontend — removed as stale (2026-08-17)
-| Status | Phase | Doc |
-| ------ | -------------------------------- | ------------------------------------------------------------------- |
-| ⏸ | 13d pricing UI | [13](plan/13-class-model-live-pricing.md) |
-| ⏸ | 14 MVC UI implementation (14a–f) | [14](plan/14-mvc-ui-implementation.md) |
-| ⏸ | UI exploration track | [exploration/ui/00-overview.md](plan/exploration/ui/00-overview.md) |
+The `##ui` / `.htmlx` LiveView frontend track — 13d pricing UI, the 14-MVC-UI
+implementation plan, the 7-of-7 `ui-htmlx-live` plan, and the 9-doc
+`plan/exploration/ui/` design set — was **removed**. It was built entirely on
+the non-advancing Rust runtime (`.dev/reference/crates/wo-htmlx`, `cargo run`,
+WebSocket live-patches) and contradicts the current woc/wovm direction. Recorded
+in [`discarded.md`](plan/discarded.md).
---
diff --git a/docs/01-problem.md b/docs/01-problem.md
index 584e2a8..afee068 100644
--- a/docs/01-problem.md
+++ b/docs/01-problem.md
@@ -1,6 +1,6 @@
# Problem Statement
-The current writeonce architecture works, but it carries weight that the project doesn't need. This document identifies the structural problems that motivate the redesign described in [02-recovery.md](./02-recovery.md).
+The current writeonce architecture works, but it carries weight that the project doesn't need. This document identifies the structural problems that motivate the redesign described in 02-recovery.md.
## Too Many Moving Parts
diff --git a/docs/02-recovery.md b/docs/02-recovery.md
deleted file mode 100644
index fd7b770..0000000
--- a/docs/02-recovery.md
+++ /dev/null
@@ -1,172 +0,0 @@
-# Recovery — The Target Architecture
-
-This document describes where writeonce is going: a single, self-contained binary that owns its own storage, serves its own content, and pushes updates to connected clients in real-time — with no external database, no cloud pipeline, and no separate API server.
-
-## Guiding Principle
-
-**Everything in one process.** The database, the server logic, and the client-facing interface all live in a single codebase and ship as a single executable. If you can run the binary, you have the full platform.
-
-## Own Database
-
-The current PostgreSQL instance is a derived cache — it stores JSONB copies of files that already exist as the source of truth. The recovery architecture eliminates this indirection entirely.
-
-### What Changes
-
-- **No external database.** No PostgreSQL, no Diesel ORM, no connection pooling, no migrations.
-- **Local file storage.** Markdown files and JSON metadata files are stored in a local directory, just as they are today in `writeonce-articles-s3/`. The file system *is* the database.
-- **Custom storage segments (.seg files).** Research area: segment files that provide efficient read access, indexing, and potentially append-only writes for content. Think of these as a lightweight, purpose-built storage layer — not a general-purpose database engine, but enough to support indexed lookups by `blog-title` and ordered listing by date.
-- **Indexed by blog-title.** The `sys_title` / blog-title field remains the primary key for content retrieval. The embedded storage must support O(1) or O(log n) lookups by this field.
-
-### What Stays the Same
-
-- Articles are still structured as JSON metadata + Markdown content pairs.
-- The `sys_title`, `published`, `tags`, `author`, and section structure remain the content model.
-- Content is still the source of truth — but now it's read directly from local storage instead of being derived through a sync pipeline.
-
-## No AWS Infrastructure
-
-The current architecture uses S3 as a file host and Lambda as a sync trigger. In the target architecture, there is nothing to sync *to* — the files are already where they need to be.
-
-### What Gets Removed
-
-| Current Component | Why It Existed | Why It's No Longer Needed |
-|---|---|---|
-| S3 bucket | Remote file storage | Files live locally alongside the binary |
-| Lambda function (Go) | Watch S3 for changes, call API | No remote store to watch — file changes are local |
-| aws-infra service (Rust) | Bridge to AWS S3/EC2 APIs | No AWS dependency |
-| Pulumi IaC | Manage Lambda + S3 resources | No cloud resources to manage |
-
-### What Replaces It
-
-The binary watches its own content directory. When a file changes (new article, updated metadata), the embedded database re-indexes and notifies subscribers. The deployment model becomes:
-
-```
-1. Place the binary on a server
-2. Point it at a content directory
-3. It serves
-```
-
-No credentials, no IAM roles, no SDK configuration.
-
-## No Separate API
-
-Today, `writeonce-api` is a standalone Actix-web server that mediates between the frontend and the database. In the target architecture, the server logic is embedded in the same process as the database and the content renderer.
-
-### What This Means
-
-- **No HTTP hop between database and server.** Queries go directly from the request handler to the storage engine in-process. No network serialization, no connection pool, no ORM layer.
-- **Single codebase.** No multi-repo coordination. A new article field is added once — in the content model — and it flows through storage, indexing, and rendering in the same compilation unit.
-- **Single deployment.** One binary, one container, one process. No docker-compose orchestrating API + database + infra services.
-
-The binary still exposes HTTP endpoints — it's still a web server. But it's a web server with an embedded database, not a web server that talks to an external one.
-
-## Real-Time Subscriptions Without WebSocket
-
-The current architecture has no mechanism for pushing content updates to connected clients. The target architecture adds real-time subscriptions, but explicitly without WebSocket.
-
-### Why Not WebSocket
-
-WebSocket adds connection state management, heartbeat logic, reconnection handling, and protocol upgrade complexity. For a content platform where updates are infrequent (articles are published, not streamed), the overhead isn't justified.
-
-### Subscription Model
-
-The target is a subscription mechanism where:
-
-- A client subscribes to a content query (e.g., "all published articles" or "article with sys_title X")
-- When the underlying data changes, the server pushes the relevant diff to the subscriber
-- No polling from the client side
-
-Candidate approaches to research:
-
-- **Server-Sent Events (SSE)** — unidirectional push over HTTP. Simple, well-supported, no protocol upgrade. Natural fit for infrequent content updates.
-- **SpacetimeDB-style subscriptions** — clients register queries, the engine tracks which rows match, and only sends diffs when the result set changes. This is the aspirational model.
-- **Long polling** — fallback option. Simple but less efficient than SSE for multiple subscribers.
-
-The key constraint: the subscription mechanism must work without requiring clients to maintain persistent bidirectional connections.
-
-## Target Architecture
-
-```
- content directory
- (JSON + MD files, .seg index)
- |
- | file watch + re-index
- v
- +---------------------------+
- | writeonce binary |
- | |
- | +-------------------+ |
- | | embedded storage | | .seg files, blog-title index
- | | (read/write/index)| |
- | +-------------------+ |
- | | |
- | +-------------------+ |
- | | server logic | | route handlers, content queries
- | | (HTTP endpoints) | |
- | +-------------------+ |
- | | |
- | +-------------------+ |
- | | subscription mgr | | SSE / query-based push
- | | (real-time push) | |
- | +-------------------+ |
- | |
- +---------------------------+
- |
- HTTP / SSE
- |
- v
- +-------------------+
- | frontend app | Angular or successor
- | (browser client) |
- +-------------------+
-```
-
-## Single Repository
-
-The five current repos collapse into one:
-
-```
-writeonce/
- content/ # articles (JSON + MD), images, assets
- storage/ # embedded database engine (.seg files, indexing)
- server/ # HTTP handlers, subscription manager
- frontend/ # client application
- writeonce.toml # configuration (port, content dir, index settings)
-```
-
-One repo. One build. One deploy artifact.
-
-## What Needs Research
-
-| Area | Question | Notes |
-|------|----------|-------|
-| **.seg file format** | What storage format gives efficient indexed reads over JSON+MD content? | Look at LSM trees, append-only logs, SQLite's page format for inspiration |
-| **File watching** | How to efficiently detect content changes on Linux/macOS? | `inotify` on Linux, `kqueue` on macOS, or cross-platform via `notify` crate |
-| **SSE vs alternatives** | Is SSE sufficient for the subscription model, or is something custom needed? | SSE handles the "push diffs to subscribers" case well for low-frequency updates |
-| **Index structure** | What index structure supports `blog-title` lookup + date-ordered listing? | B-tree or hash index for title, sorted set for date ordering |
-| **Language choice** | Continue with Rust for the unified binary? | Rust fits: single binary output, no runtime, strong typing, existing team knowledge |
-| **Frontend coupling** | Should the frontend be embedded in the binary (serve static assets) or remain separate? | Embedding simplifies deployment; separate allows independent frontend iteration |
-
-## Migration Path
-
-The transition from current to target doesn't have to be all-or-nothing:
-
-1. **Phase 1** — Build the embedded storage engine. Read JSON+MD files from a local directory, index by `blog-title`, serve via HTTP. No AWS, no PostgreSQL. This alone replaces `writeonce-api` + `aws-infra` + `lambda-function` + PostgreSQL.
-2. **Phase 2** — Add real-time subscriptions (SSE). Clients subscribe to content queries and receive push updates when files change.
-3. **Phase 3** — Collapse repositories. Move frontend into the unified codebase. Ship as a single binary that serves both API and static assets.
-
-Each phase produces a working system. The current architecture can run in parallel until the new one is ready.
-
-## Implementation phases
-
-The "embedded storage engine" of Phase 1 above lands in three numbered plan docs under [`docs/plan/`](./plan/):
-
-| Phase | Doc | What it ships |
-| --- | --- | --- |
-| 10 | [`plan/10-storage-foundations.md`](./plan/10-storage-foundations.md) | On-disk row codec (length-prefix + flags + LSN + CRC32C); per-type segment files (`data/.seg`); `posix_fallocate` preallocation; `pwrite`-only append path. Reads still in-memory. |
-| 11 | [`plan/11-wal-and-recovery.md`](./plan/11-wal-and-recovery.md) | WAL log with `fdatasync` at commit; group commit per loop tick; control file with `last_durable_lsn` (rename-on-write); replay loop on startup. `kill -9` mid-write loses nothing acknowledged. |
-| 12 | [`plan/12-engine-disk-cutover.md`](./plan/12-engine-disk-cutover.md) | `Engine`'s row payload moves to disk; in-memory map becomes `BTreeMap`. Periodic checkpoint flushes segments + advances the control file. RAM bounded by id-count, not row size. |
-
-Postgres' storage subsystem is the design reference — see [`docs/plan/exploration/postgresql/`](./plan/exploration/postgresql/) for which Postgres modules informed which decision and what writeonce skips (multi-process IPC, latches, separate writer processes).
-
-The durability syscalls themselves live in [`docs/plan/exploration/linux/12-pwrite-fsync.md`](./plan/exploration/linux/12-pwrite-fsync.md).
diff --git a/docs/03-data.md b/docs/03-data.md
deleted file mode 100644
index d444392..0000000
--- a/docs/03-data.md
+++ /dev/null
@@ -1,184 +0,0 @@
-# Data Layer — Local Storage with Subscriptions
-
-This document describes the embedded data layer that replaces PostgreSQL: local `.seg` files with indexing, and a subscription model where clients register queries and receive diffs on route visit — no polling required.
-
-## .seg File Storage
-
-The `.seg` (segment) format is the on-disk representation of article data. Each segment file holds serialized article content with positional indexing for fast lookups.
-
-### Design Goals
-
-- **No external database process.** The binary reads and writes `.seg` files directly. No socket connections, no protocol negotiation, no separate daemon.
-- **Indexed by blog-title.** The primary access pattern is `GET /blog/:sys_title`. The storage layer must resolve a `sys_title` to its article content without scanning all files.
-- **Append-friendly.** New articles and updates append to the segment. Deletes are tombstoned and compacted later.
-- **Human-readable source.** The JSON + Markdown files remain the authoring format. `.seg` files are a derived index — if they're deleted, they can be rebuilt from the content directory.
-
-### Proposed Structure
-
-```
-content/
- linux-misc/
- linux-misc.json # authored metadata (source of truth)
- linux-misc.md # authored content (source of truth)
- aws-lambda-pulumi/
- aws-lambda-pulumi.json
- aws-lambda-pulumi.md
-
-data/
- articles.seg # serialized article records
- index/
- title.idx # blog-title -> offset mapping
- date.idx # publish date -> offset (sorted)
- tags.idx # tag -> [offsets] (inverted index)
-```
-
-The `content/` directory is what the author edits. The `data/` directory is what the engine builds and queries. Losing `data/` is a cold start, not data loss.
-
-### Segment File Internals
-
-```
-+------------------+
-| segment header | magic bytes, version, record count
-+------------------+
-| record 0 | length-prefixed serialized article
-+------------------+
-| record 1 |
-+------------------+
-| ... |
-+------------------+
-| record N |
-+------------------+
-```
-
-Each record is a length-prefixed byte sequence containing the full article (metadata + content merged). Records are addressed by byte offset from the start of the file.
-
-### Index Files
-
-**title.idx** — Hash map serialized to disk. Maps `sys_title` (string) to byte offset in `articles.seg`. Loaded into memory at startup for O(1) lookups.
-
-**date.idx** — Sorted array of `(timestamp, offset)` pairs. Supports range queries for "articles published between X and Y" and ordered listing for the homepage.
-
-**tags.idx** — Inverted index. Maps each tag string to a list of offsets. Supports "all articles tagged with X" queries.
-
-On startup, index files are memory-mapped or loaded into heap. On content change, affected indexes are rebuilt incrementally.
-
-## Subscription Model
-
-The subscription model is inspired by SpacetimeDB: clients register queries, and the engine tracks which results match. When underlying data changes, only the relevant diffs are pushed to subscribers.
-
-### How It Works
-
-```
- Client A Server Content Dir
- | | |
- |--- GET /blog/linux-misc -| |
- | |-- read from .seg index ---->|
- |<-- article + SSE stream -| |
- | | |
- | (subscribed to | |
- | sys_title=linux-misc) | |
- | | |
- | |<-- file change detected ----|
- | | |
- | |-- re-index article -------->|
- | |-- diff against last push -->|
- | | |
- |<-- SSE: updated content -| |
- | | |
-```
-
-### Route-Based Subscription
-
-When a user visits a route, the response includes both the current content and an SSE stream. The client is automatically subscribed to changes for that query — no explicit subscription handshake needed.
-
-```
-GET /blog/linux-misc
-```
-
-Response:
-```
-HTTP/1.1 200 OK
-Content-Type: text/html
-
-
-
-
-
-```
-
-The subscription lives as long as the browser tab is open. When the user navigates away, the EventSource closes and the server drops the subscription. No heartbeat management, no reconnection logic beyond what SSE provides natively (automatic reconnect is built into the EventSource API).
-
-### Query Registration
-
-Subscriptions are not limited to single-article lookups. The engine supports registering arbitrary content queries:
-
-| Query Type | Example | Subscription Behavior |
-|---|---|---|
-| Single article | `sys_title = "linux-misc"` | Push when this specific article changes |
-| All published | `published = true` | Push when any article is published or unpublished |
-| By tag | `tags contains "rust"` | Push when a rust-tagged article is added, removed, or updated |
-| Homepage list | `published = true ORDER BY date DESC LIMIT 10` | Push when the top-10 list changes |
-
-The server maintains a registry of active subscriptions. On each content change, it evaluates which subscriptions are affected and pushes diffs only to those clients.
-
-### Diff Format
-
-When content changes, the server doesn't resend the full article. It sends a minimal diff:
-
-```json
-{
- "type": "update",
- "sys_title": "linux-misc",
- "changes": {
- "content.sections[2].paragraphs[0]": "Updated paragraph text...",
- "content.tags": ["linux", "kernel", "new-tag"]
- },
- "version": 42
-}
-```
-
-The `version` field enables clients to detect missed updates and request a full resync if needed.
-
-## Sample Dataset
-
-To validate the storage engine and subscription model, a sample dataset should exercise the core access patterns:
-
-### Articles
-
-| sys_title | tags | published | purpose |
-|---|---|---|---|
-| `sample-getting-started` | `[tutorial, beginner]` | true | Basic article, tests single-article subscription |
-| `sample-rust-patterns` | `[rust, patterns]` | true | Tests tag-based queries |
-| `sample-draft-wip` | `[draft]` | false | Tests published filter — should not appear in public queries |
-| `sample-long-form` | `[deep-dive, rust]` | true | Multiple sections, images, code snippets — tests complex content rendering |
-| `sample-frequently-updated` | `[changelog]` | true | Updated often — tests subscription diff delivery |
-
-### Test Scenarios
-
-1. **Cold start** — Delete `data/`, start the binary. It should rebuild `.seg` and index files from `content/` and serve all articles.
-2. **Single article query** — `GET /blog/sample-getting-started` returns the article and opens an SSE subscription.
-3. **Live update** — Edit `sample-frequently-updated.json` while a client is subscribed. The client should receive an SSE event with the diff.
-4. **Tag query** — Subscribe to `tags contains "rust"`. Both `sample-rust-patterns` and `sample-long-form` should be in the result set. Adding a new article tagged `rust` should trigger a push.
-5. **Publish toggle** — Change `sample-draft-wip` from `published: false` to `true`. Clients subscribed to the homepage list should receive a push with the new article added.
-
-## SpacetimeDB Reference
-
-SpacetimeDB is the primary architectural inspiration for the subscription model. Key concepts to study:
-
-- **Modules** — server logic that runs inside the database, not beside it
-- **Subscription queries** — clients register SQL-like queries; the engine evaluates them incrementally on each transaction
-- **Incremental view maintenance** — only recompute the parts of a query result that changed
-- **Client SDK generation** — type-safe client code generated from the server schema
-
-Add SpacetimeDB as a reference submodule for quick access to their implementation patterns:
-
-```bash
-git submodule add https://github.com/clockworklabs/SpacetimeDB.git references/spacetimedb
-```
-
-The goal is not to replicate SpacetimeDB — it's to take its subscription semantics and apply them to a much narrower domain (blog content), where the simplicity of the problem allows a simpler implementation.
diff --git a/docs/04-ui.md b/docs/04-ui.md
deleted file mode 100644
index 74dfd6f..0000000
--- a/docs/04-ui.md
+++ /dev/null
@@ -1,216 +0,0 @@
-# User Interface — Server-Rendered HTMLX
-
-No Angular. No React. No frontend framework. The UI is a set of `.htmlx` template files that the server parses, populates with content from the embedded database, and serves as plain HTML. Real-time updates arrive via SSE and are applied with minimal client-side scripting.
-
-## Why Not Angular
-
-The current `writeonce-app` is an Angular 18 SPA with Tailwind, PrismJS, ngx-markdown, and FontAwesome. It works, but it's a heavy delivery mechanism for what is fundamentally a read-heavy content site:
-
-- **~200MB of `node_modules`** for a site that renders markdown articles
-- **Client-side routing** for content that doesn't need it — every article is a distinct URL, not an interactive application
-- **JavaScript-dependent rendering** — content doesn't exist until Angular boots, hydrates, and fetches from the API
-- **Separate build pipeline** — `npm run build` produces static assets that must be deployed to nginx independently of the API
-
-The content is static between updates. The interactivity is limited to navigation and code highlighting. A server-rendered approach matches the actual requirements.
-
-## HTMLX Templates
-
-The author defines the site layout using `.htmlx` files — HTML with embedded data bindings that the server resolves at render time.
-
-### Template Structure
-
-```
-templates/
- layout.htmlx # outer shell: , ,
- header.htmlx # site header, navigation
- footer.htmlx # site footer
- home.htmlx # homepage: article list
- article.htmlx # single article view
- about.htmlx # static page
- contact.htmlx # static page
- components/
- article-card.htmlx # summary card for article listings
- code-snippet.htmlx # code block with language + title
- img-caption.htmlx # image with caption
- section.htmlx # article section (heading + paragraphs)
-```
-
-### Template Syntax
-
-Templates use a binding syntax that references content from the database. The server parses these bindings, resolves them against the current content, and outputs plain HTML.
-
-```html
-
-
-
-
-```
-
-```html
-
-
-
- {{#each articles}}
- {{> article-card article=this}}
- {{/each}}
-
-```
-
-The `{{> partial}}` syntax includes another `.htmlx` file as a component. The server resolves these at render time — no client-side component tree.
-
-### Content Subscription in Templates
-
-Templates declare what data they need. The server resolves these declarations against the embedded database and subscribes the client to changes:
-
-```html
-
-
-
-
-
{{article.title}}
- ...
-
-```
-
-```html
-
-
-
-
- {{#each articles}}
- {{> article-card article=this}}
- {{/each}}
-
-```
-
-The `` comment is a directive to the server. It declares the query that populates the template's data context. The same query is used to register an SSE subscription for live updates (as described in [03-data.md](./03-data.md)).
-
-## Rendering Pipeline
-
-```
- Browser request
- |
- v
- Route match (/blog/linux-misc)
- |
- v
- Load template (article.htmlx)
- |
- v
- Parse subscribe directive
- (article WHERE sys_title = "linux-misc")
- |
- v
- Query embedded database (.seg index)
- |
- v
- Resolve template bindings ({{article.title}}, etc.)
- |
- v
- Compose with layout.htmlx + header.htmlx + footer.htmlx
- |
- v
- Inject SSE subscription script
- |
- v
- Send complete HTML response
-```
-
-The browser receives a fully rendered page on first load. No JavaScript framework boots. No API call fires. The content is already in the HTML.
-
-## Live Updates via SSE
-
-After the initial HTML is delivered, a small inline script opens an SSE connection for the page's subscription query:
-
-```html
-
-```
-
-The `applyDiff` function is a lightweight client-side updater — it targets DOM elements by data attribute and patches their content. No virtual DOM, no reconciliation, no framework. For a content site where updates are infrequent and localized (a paragraph changed, a tag was added), direct DOM manipulation is sufficient.
-
-```html
-
Linux Misc
-
First paragraph...
-```
-
-When a diff arrives for `article.title`, the script finds the element with `data-bind="article.title"` and replaces its text content. This is the minimal client-side code the architecture requires.
-
-## Code Highlighting
-
-The current frontend uses PrismJS for syntax highlighting. In the server-rendered model, highlighting can happen at either layer:
-
-**Server-side (preferred):** The server parses code blocks during template rendering and emits pre-highlighted HTML with CSS classes. The browser only needs the PrismJS CSS theme, not the JavaScript library. This eliminates client-side parsing entirely.
-
-**Client-side (fallback):** Include PrismJS as a small script that runs on page load and on SSE update. Simpler to implement initially but adds a JavaScript dependency.
-
-## Markdown Rendering
-
-The current frontend uses `ngx-markdown` and `marked` to parse markdown in the browser. In the target architecture, markdown is rendered to HTML on the server during template composition. The browser never sees raw markdown.
-
-This aligns with the content model: the JSON metadata already defines the article structure (sections, paragraphs, code snippets, images). The markdown file provides prose content. The server combines both into final HTML — the template just places the pre-rendered blocks.
-
-## What Gets Removed
-
-| Current (Angular) | Target (HTMLX) |
-|---|---|
-| `writeonce-app/` (full Angular project) | `templates/` (handful of .htmlx files) |
-| `node_modules/` (~200MB) | None |
-| `angular.json`, `tsconfig.json`, `karma.conf.js` | None |
-| npm build pipeline | Template parsed at request time |
-| Nginx static file serving | Binary serves its own HTML |
-| Client-side routing | Server-side route matching |
-| Client-side markdown parsing | Server-side rendering |
-| Client-side code highlighting | Server-side or minimal JS |
-
-## Styling
-
-Templates use plain CSS. Tailwind can optionally be used as a build-time utility (generating a static CSS file), but there is no runtime CSS framework. The author writes styles in a `styles.css` file that the server serves as a static asset.
-
-```
-templates/
- styles/
- main.css # site-wide styles
- article.css # article-specific styles
- code-theme.css # syntax highlighting theme (PrismJS compatible)
-```
-
-## Template Authoring Experience
-
-The `.htmlx` files are editable by the same author who writes articles. The template syntax is intentionally close to HTML — there's no JSX, no TypeScript, no build step. An author who knows HTML can modify the site layout.
-
-This closes the loop on the writeonce philosophy: the author writes content (markdown + JSON) and layout (`.htmlx` + CSS) as files, and the binary turns them into a live site.
diff --git a/docs/05-datalayer.md b/docs/05-datalayer.md
deleted file mode 100644
index 97150e3..0000000
--- a/docs/05-datalayer.md
+++ /dev/null
@@ -1,139 +0,0 @@
-# Data Layer — Implementation Status
-
-The embedded data layer described in [02-recovery.md](./02-recovery.md) and [03-data.md](./03-data.md) has been implemented as a Cargo workspace with 8 crates. All 44 tests pass. No external database, no AWS, no tokio — direct Linux syscalls on a custom event loop.
-
-## Workspace Structure
-
-```
-writeonce-all/
- Cargo.toml # workspace root
- docs/ # architecture documentation
- sample-content/ # 5 test articles for validation
- crates/
- wo-model/ # content model
- wo-seg/ # .seg file format
- wo-index/ # index files
- wo-store/ # unified storage engine
- wo-watch/ # inotify file watcher
- wo-event/ # epoll event loop
- wo-sub/ # subscription system
- wo-rt/ # custom runtime
-```
-
-## Crate Summary
-
-| Crate | Purpose | Tests | Key Types |
-|-------|---------|-------|-----------|
-| **wo-model** | Article structs matching existing JSON schema, `ContentLoader` for directory walking | 8 | `Article`, `ArticleContent`, `ArticleBody`, `Section`, `CodeSnippet`, `ContentLoader` |
-| **wo-seg** | Binary `.seg` file format — length-prefixed records, tombstoning, positional I/O | 6 | `SegWriter`, `SegReader`, `SegHeader` |
-| **wo-index** | Three index types for O(1) and O(log n) access patterns | 8 | `TitleIndex`, `DateIndex`, `TagIndex` |
-| **wo-store** | Unified storage engine composing seg + indexes, cold-start rebuild | 3 | `Store` |
-| **wo-watch** | Content directory watcher using inotify | 4 | `ContentWatcher`, `ContentChange` |
-| **wo-event** | Custom event loop on epoll with eventfd, timerfd, signalfd | 5 | `EventLoop`, `EventFd`, `TimerFd`, `SignalFd` |
-| **wo-sub** | Subscription manager with fd-based notifications, `register!` macro | 6 | `SubscriptionManager`, `Subscription`, `Notification` |
-| **wo-rt** | Runtime tying all crates together — single process, single event loop | 4 | `Runtime`, `RuntimeHandle`, `Config` |
-
-## Linux Kernel Syscalls Used
-
-| Syscall | Crate | Purpose |
-|---------|-------|---------|
-| `pread` / `pwrite` | wo-seg | Positional read/write for .seg records without seeking |
-| `fallocate` | wo-seg | Pre-allocate .seg file space to reduce fragmentation |
-| `epoll_create1` / `epoll_ctl` / `epoll_wait` | wo-event | Event-driven I/O multiplexing for the main loop |
-| `eventfd` | wo-event, wo-sub | Lightweight signaling between watcher and subscription manager |
-| `timerfd_create` / `timerfd_settime` | wo-event | Periodic tasks (compaction, keepalive) as file descriptors |
-| `signalfd` | wo-event | SIGINT/SIGTERM delivered as fd events for graceful shutdown |
-| `inotify_init1` / `inotify_add_watch` | wo-watch | File system change detection on the content directory |
-| `pipe2` | wo-sub (tests) | Mock subscriber fds for testing notification delivery |
-
-## .seg File Format
-
-```
-Offset Size Field
-0 4 Magic: b"WOSF"
-4 2 Version: u16 LE (1)
-6 2 Flags: u16 LE (reserved)
-8 8 Record count: u64 LE
-16 8 Data start offset: u64 LE
-24 8 Reserved
-32+ variable Records: [u32 length][u8 flags][bincode payload]...
-```
-
-- Records are addressed by byte offset from file start
-- Flags: `0x00` = active, `0x01` = tombstoned
-- Payload: bincode-serialized `Article` struct
-
-## Index Files
-
-| File | Format | Access Pattern |
-|------|--------|----------------|
-| `title.idx` | On-disk hash table (Robin Hood, load factor 0.5), 138 bytes/slot | O(1) lookup by `sys_title` |
-| `date.idx` | Sorted `(i64 timestamp, u64 offset)` array, 16 bytes/entry | Binary search for date ranges, latest N |
-| `tags.idx` | Bincode-serialized `HashMap>` | Tag-to-offsets inverted index |
-
-All indexes are derived from `.seg` and rebuildable from `content/` on cold start.
-
-## Subscription Model
-
-No SSE. No WebSocket. Notifications are written directly to subscriber file descriptors.
-
-- **Subscribe**: `SubscriptionManager::subscribe(fd, Subscription::ByTitle("linux-misc"))`
-- **Notify**: on content change, length-prefixed `Notification` written to matching fds
-- **Cleanup**: `EPOLLHUP` on epoll triggers automatic `unsubscribe(fd)`
-- **Dedup**: if a fd matches multiple patterns (title + tag), it receives only one notification
-
-Subscription patterns:
-- `Subscription::ByTitle(sys_title)` — single article
-- `Subscription::ByTag(tag)` — all articles with tag
-- `Subscription::All` — all content changes
-
-## Store Query API
-
-```rust
-store.get_by_title("linux-misc") -> Option
-store.list_published(skip, limit) -> Vec
-store.list_by_tag("rust") -> Vec
-store.list_by_date_range(start, end) -> Vec
-store.count_published() -> usize
-store.article_version("linux-misc") -> Option
-store.rebuild() // full rebuild from content/
-```
-
-## Runtime Event Loop
-
-Single `epoll` instance multiplexing all file descriptors:
-
-| Token | Fd | Handler |
-|-------|----|---------|
-| `WATCHER` | inotify fd | Process file changes → update store → notify subscribers |
-| `SIGNAL` | signalfd | SIGINT/SIGTERM → graceful shutdown |
-| `TIMER` | timerfd | Periodic tasks (compaction, stats) |
-| `NOTIFY` | eventfd | Subscription notification signal |
-| `1000+` | subscriber fds | Hangup detection → unsubscribe + cleanup |
-
-## External Dependencies
-
-| Crate | Version | Purpose |
-|-------|---------|---------|
-| `serde` | 1.x | Serialization derives |
-| `serde_json` | 1.x | JSON parsing for article files |
-| `bincode` | 1.x | Compact binary serialization for .seg records and notifications |
-| `libc` | 0.2.x | Raw Linux syscall bindings |
-
-No tokio. No async-std. No database driver. No HTTP framework (yet).
-
-## What Comes Next
-
-The data layer delivers everything the HTTP server and UI layers need:
-
-1. **`Store` with zero-copy query access** — all article queries resolve in-process
-2. **Subscription system accepting raw fds** — HTTP layer hands socket fds to `subscribe()`
-3. **Shared event loop** — HTTP listener socket registers on the same epoll
-4. **Automatic cold-start** — if `data/` is missing, rebuilds from `content/` on startup
-5. **Graceful shutdown** — SIGTERM triggers clean fd cleanup
-
-Next phases per [02-recovery.md](./02-recovery.md):
-- **HTTP server** — route handlers using the `Store` query API, embedded in the same binary
-- **HTMLX templates** — server-rendered HTML with `{{bindings}}` per [04-ui.md](./04-ui.md)
-- **Frontend collapse** — serve static assets from the binary, eliminate the Angular app
-3
\ No newline at end of file
diff --git a/docs/06-markdown-render.md b/docs/06-markdown-render.md
deleted file mode 100644
index 224b84f..0000000
--- a/docs/06-markdown-render.md
+++ /dev/null
@@ -1,246 +0,0 @@
-# Markdown File Rendering
-
-## Current State (writeonce-articles-s3)
-
-Each article is a directory containing a JSON metadata file and one or more `.md` files:
-
-```
-auto-scale-gitlab-runner-using-aws-spot-instance/
- docker-machine-test-with-t2.md
- gitlab-runner-config.md
- stop-test-gitlab-docker-machine.md
-
-gitlab-runner-with-kubernetes-executor/
- gitlab-runner-with-kubernetes-executor.json
- deploy.md
- permission.md
- role-binding.md
- role-defination.md
- gitlab-runnergitlab-runner-deploy.md
-```
-
-The JSON metadata currently defines the full article structure — sections, headings, paragraphs, and code snippet references. Markdown files are limited to code blocks referenced via the `codes[].snippet` field.
-
-## Problem
-
-The JSON metadata carries too much content. Headings, paragraphs, prose — all of this is duplicated as JSON strings inside `content.content.sections`. The markdown files only hold code snippets, referenced by `sectionIndex` and `paragraphIndex`.
-
-This is backwards. The markdown file should be the content. The JSON should be minimal metadata.
-
-## Target: Markdown-First Content Model
-
-**The markdown file is the article.** All prose, headings, code blocks, and inline formatting live in the `.md` file. The JSON metadata file holds only what markdown cannot express: system fields, tags, publication state, and author.
-
-### Minimal JSON Metadata
-
-```json
-{
- "sys_title": "gitlab-runner-with-kubernetes-executor",
- "title": "Gitlab Runner with Kubernetes Executor",
- "published": true,
- "author": "Shoney Arickathil",
- "tags": ["kubernetes", "gitlab", "ci-cd"],
- "published_on": 1740950884
-}
-```
-
-No `content.content.sections`. No `content.content.codes`. No `paragraphs[]` arrays. No `sectionIndex`/`paragraphIndex` mapping.
-
-### Markdown File = Full Article Content
-
-````markdown
-# Introduction
-
-Deploying a Gitlab runner using kubernetes is a great option to overcome
-the limitations of other gitlab runner executor such as docker and docker machine.
-
-## Running Gitlab Runner in gitlab namespace
-
-Create the namespace and apply the deployment:
-
-```yaml
-apiVersion: apps/v1
-kind: Deployment
-metadata:
- name: gitlab-runner
- namespace: gitlab
-```
-````
-
-## Permissions
-
-The runner needs RBAC permissions to create pods:
-
-```yaml
-apiVersion: rbac.authorization.k8s.io/v1
-kind: Role
-metadata:
- name: gitlab-runner
-```
-
-Everything is in the markdown — headings, paragraphs, code blocks with language hints, links, images. The rendering pipeline parses the markdown directly.
-
-### Directory Structure
-
-```
-content/
- gitlab-runner-with-kubernetes-executor/
- gitlab-runner-with-kubernetes-executor.json # minimal metadata
- gitlab-runner-with-kubernetes-executor.md # full article content
- linux-misc/
- linux-misc.json
- linux-misc.md
-```
-
-One JSON for metadata. One markdown for content. No scattered `.md` files per code snippet.
-
-## What Changes
-
-| Before | After |
-| -------------------------------------------------------------- | ----------------------------------------------------------------------- |
-| JSON holds sections, headings, paragraphs as structured arrays | JSON holds only sys_title, title, published, author, tags, published_on |
-| Markdown files hold only code snippets | Markdown file holds the entire article |
-| `codes[].snippet` maps filename to sectionIndex/paragraphIndex | No mapping needed — headings and code blocks are inline in markdown |
-| Renderer reads JSON structure, injects code from .md files | Renderer parses markdown directly into HTML |
-| Multiple .md files per article (one per code snippet) | One .md file per article |
-
-## Impact on the Data Layer
-
-### wo-model
-
-The `Article` struct simplifies:
-
-```rust
-pub struct Article {
- pub sys_title: String,
- pub title: String,
- pub published: bool,
- pub author: String,
- pub tags: Vec,
- pub published_on: Option,
-}
-```
-
-The nested `ArticleContent` / `ArticleBody` / `Section` / `CodeSnippet` hierarchy is no longer needed. Article content comes from parsing the `.md` file at render time, not from the JSON.
-
-### wo-md
-
-Currently handles only inline markdown (`**bold**`, `` `code` ``, links). Needs to become a full markdown-to-HTML renderer:
-
-- Block elements: headings (`#`, `##`), paragraphs, code fences (` `lang ```), lists, blockquotes
-- Inline elements: bold, italic, code, links, images
-- Code fence language extraction for `wo-md::highlight()`
-- The renderer reads `{sys_title}/{sys_title}.md`, parses it, and returns HTML
-
-### wo-htmlx
-
-The `article.htmlx` template simplifies. Instead of iterating `{{#each article.content.content.sections}}`, it renders the pre-parsed markdown HTML:
-
-```html
-
-