Merge pull request #1 from shoneyJ/database-engine

Database engine + query surface + docs cleanup
This commit is contained in:
Shoney Arickathil 2026-08-17 23:59:20 +02:00 committed by GitHub
commit 2c2a3bb7cd
118 changed files with 6542 additions and 3617 deletions

1
.gitignore vendored
View file

@ -3,6 +3,7 @@
# `woc .` manifest builds (wo.toml [build] target) # `woc .` manifest builds (wo.toml [build] target)
/docs/examples/log-watcher/target /docs/examples/log-watcher/target
/docs/examples/employee/target
# Rust runtime (crates/rt/): compiled binary + build artifacts # Rust runtime (crates/rt/): compiled binary + build artifacts
/crates/rt/target /crates/rt/target

346
README.md
View file

@ -1,77 +1,323 @@
# writeonce # writeonce
A declarative full-stack programming language. You write `.wo` files; the runtime compiles them into a binary that owns the database, serves REST, and pushes live subscriptions — no external database, no external web server, no frontend framework. **A small compiled language with a database built in.** You write `.wo`
files; one command turns them into a single native binary that carries its
own storage engine — a typed, WAL-durable, crash-recoverable database — with
no server to install, no ORM, and no query strings. Tables are just classes,
queries are written in the language and checked by the compiler, and the whole
program ships as one file that depends only on the system C library.
Think **Go + Postgres + `net/http` + Phoenix LiveView, folded into one language and one binary.** > **Status: early, honest.** Everything documented on this page compiles and
> runs today and is exercised by the acceptance tests in this repository.
> Features that are planned but **not yet available** are listed separately
> under [Roadmap](#roadmap) — they are not described as if they work. Nothing
> here is API-stable yet.
# persistant database ---
- reads and writes database to RAM, persist data to postgres SQL. ## Why writeonce
- The entire database lives in RAM; every committed write is mirrored to PostgreSQL **as a backup** — asynchronously, behind the runtime's own WAL, never in the read or ack path. Set `WO_PG=postgres://user@host:5432/db` and every type's rows appear as a Postgres table (named by its `@table(name: ...)` annotation) that you can query with plain `psql`. Plan and phases: [`docs/plan/16-postgres-mirror.md`](docs/plan/16-postgres-mirror.md); try it: `just pricing-pg-demo`.
## Quickstart - **The database is part of the language.** A `class` marked `@table` *is* a
table. Its rows persist through a write-ahead log, survive a restart, and are
reached by navigating typed relations — not by assembling SQL text.
- **Queries are compiled, not interpreted.** `from e in Employee where
e.salary > 90000 select e` lowers to bytecode loops over the engine. A
mistyped field name is a **compile error**, not a runtime surprise. There is
no SQL string anywhere in the shipped binary.
- **One binary, no runtime dependencies.** `woc .` produces a self-contained
executable (~100 KB for the sample programs) that links only libc. Copy it to
a server and run it.
- **Small on purpose.** No FFI, no package manager, no framework. The standard
library is a handful of OS modules. The language is designed to be read.
writeonce is **not** a web framework and does not (yet) serve HTTP, WebSockets,
or a UI. It is a systems language whose distinguishing feature is the embedded
database. If you have seen an older "writeonce" that served REST from `cargo
run`, that was a separate, earlier runtime; this page documents the current
`woc`/`wovm` toolchain.
---
## System requirements
**To run a compiled writeonce program:**
- Linux on x86-64. The produced binary is a native executable that links only
the system C library (`libc`); nothing else is required at runtime.
**To build programs from source (the toolchain), you need:**
| Tool | Version tested | Purpose |
| --- | --- | --- |
| OCaml | 4.14+ | builds `woc`, the compiler front end |
| dune | 3.14+ | OCaml build driver |
| A C11 compiler | gcc 13 / clang | builds `wovm`, the runtime VM |
| just | 1.x | task runner for the build/test recipes |
| make | any | drives the runtime build |
Other POSIX platforms (macOS, BSD) are untested. The toolchain itself has no
network or package-download step — it builds entirely from the checked-in
source.
---
## Getting the toolchain
Two artifacts make up the toolchain:
- **`woc`** — the compiler (OCaml). Reads `.wo` source, type-checks it, runs
the ownership pass, and emits a `.wob` image or a standalone binary.
- **`wovm`** — the runtime (C11). Loads a `.wob` image and executes it. When
`woc` builds a standalone binary, it embeds the image into a copy of `wovm`.
Build both from the repository root:
```bash ```bash
git clone https://github.com/shoneyJ/writeonce just woc-build # builds compiler/_build/default/bin/woc
cd writeonce just wovm-build # builds runtime/wovm
cargo run --bin wo -- run docs/examples/blog # serve the sample blog on :8080
curl http://127.0.0.1:8080/api/articles # it's a real REST API now # gate them (optional but recommended)
just woc-test # compiler unit + golden suites
just wovm-test # runtime unit suites, both dispatch flavors, ASan-clean
``` ```
See [`.dev/reference/rest/blog.rest`](.dev/reference/rest/blog.rest) for a preconfigured HTTP-request file that drives the whole sample — open it in VS Code (with the REST Client extension) or JetBrains and click "Send Request" on each block. ---
## What this repository contains ## Your first program
| Path | What it is | A writeonce project is a directory with a `wo.toml` manifest and one or more
| ------------------------------------------------------------------------------------------ | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | `.wo` files. Every program has an entry point:
| [`crates/rt/`](crates/rt/) | The new `.wo` language runtime — lexer, type-DSL parser, in-memory engine, axum REST server. Produces the `wo` binary. |
| [`crates/{ql,value,engine,txn,db,wal,sub,http,gen,policy,logic,service,ui,app}/`](crates/) | 14 empty placeholder crates scaffolded for Phases 2–6. Real code extracts from `rt/` as each phase activates. |
| [`docs/runtime/wo-language.md`](docs/runtime/wo-language.md) | **Start here.** The language overview: toolchain, hello-world, stdlib, client model. |
| [`docs/runtime/database.md`](docs/runtime/database.md) | The 7-phase engineering series that drives the runtime's design. |
| [`docs/examples/blog/`](docs/examples/blog/) | Sample `.wo` project: blog with articles, authors, tags, comments. ~200 lines. |
| [`docs/examples/ecommerce/`](docs/examples/ecommerce/) | Sample `.wo` project: storefront + live order-ops table + cross-paradigm checkout. ~300 lines. |
| [`prototypes/wo-db/`](prototypes/wo-db/) | C++ prototype of the query-layer engine (SQL + Cypher + document paths, `RETURNING` aliases, `LIVE` stub). ~2k lines, smoke tests pass. Reference implementation the Rust port follows. |
| [`.dev/reference/rest/`](.dev/reference/rest/) | `.rest` files (VS Code REST Client / JetBrains HTTP format) for manually testing the running prototype. |
| [`.dev/reference/crates/`](.dev/reference/crates/) | The v1 writeonce blog — 13 Rust crates implementing the original `.seg` + sidecar-index storage engine and `.htmlx` templating. Preserved as a nested workspace; see [`.dev/reference/README.md`](.dev/reference/README.md). |
## Current stage ```
-- hello/main.wo
fn main(args: multi Text) -> Int {
print("hello, writeonce");
return 0;
}
```
The runtime is under active development. Each stage lands as an independently shippable cut: ```toml
# hello/wo.toml
name = "hello"
version = "0.1.0"
| Stage | What works | Status | [runtime]
| ------ | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------------------------------------------------------- | wo = ">= 0.1"
| **1** | `wo run <dir>` discovers every `.wo` file under a directory | ✅ shipped | ```
| **2** | Type-DSL parser, in-memory engine, REST CRUD (`list` / `get` / `create` / `update` / `delete`) generated from `service rest` blocks, JSON bodies with auto-id, default-value seeding, partial-update PATCH | ✅ shipped — `cargo run -- run docs/examples/blog` |
| **3** | LIVE subscriptions over WebSocket, delta frames on commit, `me` / session layer | pending |
| **4+** | Transactional fns (`fn checkout in txn snapshot`), row-level policies, type-attached triggers, `##ui` SSR, WAL durability, codegen | see [docs/runtime/database.md](docs/runtime/database.md) |
`cargo test --lib` at the root runs 14 unit tests covering the lexer, parser, compiler, and engine. Stage-3 endpoints respond `501 Not Implemented` until they land. Compile the directory into a single binary and run it:
## Build & test
```bash ```bash
cargo build # builds all 15 crates (only `rt` has real code) woc hello/ # produces hello/target/hello
cargo test --lib # 14 unit tests ./hello/target/hello
# hello, writeonce
cargo run --bin wo -- run docs/examples/blog # serve the blog sample
cargo run --bin wo -- run docs/examples/ecommerce # serve the ecommerce sample
# Override the listen address
WO_LISTEN=127.0.0.1:9000 cargo run --bin wo -- run docs/examples/blog
``` ```
## The v1 codebase (reference) `main` returns an `Int` — that value is the process **exit code**. `args` is
the command-line arguments (the program name is not included).
The original writeonce blog engine — 13 crates, flat-file `.seg` storage, sidecar indexes, `.htmlx` templates, hand-rolled `epoll` event loop — moved to [`.dev/reference/crates/`](.dev/reference/crates/) when the new runtime was scaffolded. It's a nested Cargo workspace: ### The two build paths
```bash ```bash
cd .dev/reference/crates # 1. standalone binary (what you ship): woc reads wo.toml, emits target/<name>
cargo build # all 13 v1 crates still compile woc myproject/
cargo test # 12 unit tests, 1 ignored integration test
# 2. image + VM (handy while developing): emit a .wob, run it with wovm
woc --emit myproject/ -o app.wob
wovm app.wob arg1 arg2
``` ```
V1 crates keep the `wo-` prefix (`wo-seg`, `wo-store`, …). The new runtime crates dropped it (`ql`, `value`, `engine`, …). [`docs/runtime/database/07-wo-seg-migration.md`](docs/runtime/database/07-wo-seg-migration.md) is the phased coexistence plan for replacing v1 with the new runtime — abstract behind a trait, dual-write, cut over, decommission. Both paths run the same program. The standalone binary is the release artifact;
the image path lets you inspect or move the image around.
## License & status ---
Work in progress. Nothing here is stable. Read the language overview in [`docs/runtime/wo-language.md`](docs/runtime/wo-language.md) if you want to know the shape; read the phase docs if you want to see the engineering plan; look in [`docs/examples/`](docs/examples/) if you want to see what the end product feels like. ## Language at a glance
writeonce is statically typed with a compile-time ownership model — every value
has a known owner, memory is freed deterministically, and values that form
cycles are collected by an inferred garbage collector (you never annotate GC-
ness; the compiler infers it). The surface will look familiar:
- **Types:** `Int`, `Text`, `Bool`, and user `class` types. `?T` marks an
optional (nullable) value; `nil` is the empty case.
- **Containers:** `multi T` (a growable list) and `map<K, V>`. Literals:
`[]`, `[a, b]`, `{}`.
- **Classes & records:** classes with fields and methods, `static const` /
`static fn` members, module-scoped across files.
- **Control flow:** `if`/`else`, `for x in xs`, `for k, v in m`, `switch`
expressions, and `try { … } catch (e) { … }` (also an expression form).
- **Strings:** interpolation with `${expr}` inside a `"…"` literal.
- **Functions:** free functions and methods; arguments and returns are typed.
```
fn classify(n: Int) -> Text {
if n < 0 { return "negative"; }
return switch n {
case 0: "zero";
default: "positive";
};
}
```
### Standard library
A compact set of OS modules, reached by their reserved names — no imports:
| Module | What it does |
| --- | --- |
| `fs` | `exists`, `list`, `stat`, `read_all`, `read_at`, `append` |
| `time` | `sleep`, `now`, `local`, `iso` |
| `env` | `get`, `stopping` (a cooperative shutdown flag) |
| `net` | TCP `listen` / `accept` / `read` / `write` / `close` (host + port) |
| `proc` | `run` a child process, capture stdout/stderr/exit |
| `json` | `encode` / `decode` (`json.decode(t) as T` yields `?T`) |
These are deliberately minimal — the surface a real program needs, and no more.
---
## The database
This is the point of the language. Declaring storage is declaring a class:
```
@table(name: "departments", index: [name])
class Department {
name: Text @unique
staff: backlink Employee.dept -- reverse relation, not a stored column
}
@table(name: "employees", index: [dept], index: [dept, salary])
class Employee {
name: Text
salary: Int
hired: Int
dept: ref Department -- foreign key: stored as the row id
}
```
- **`@table`** makes a class persistent — named storage plus declared secondary
indexes. Every instance you `insert` is written to a write-ahead log **before**
it is acknowledged, so an acked write survives a crash; on the next start the
log is replayed.
- **`ref T`** is a typed foreign key (a forward relation). **`backlink T.f`** is
its inverse — a virtual field, no stored column, resolved by an index scan.
- **`@unique`** enforces uniqueness at insert/update; a violation is a
**catchable** trap.
- **Foreign keys restrict deletes**: deleting a row that another row still
references traps rather than orphaning it.
### Writing and reading data
Mutation is direct; queries are a comprehension the compiler lowers to engine
operations:
```
-- insert (WAL-durable); @unique makes a re-insert trap, and try/catch it:
let eng = try insert Department { name: "Engineering" } catch (e) nil;
insert Employee { name: "Asha", salary: 9200000, hired: 1704067200000, dept: eng };
-- query: filter, order, limit, project — checked at compile time
for e in from s in Employee where s.salary > 8000000 order by s.salary desc select s {
print("${e.name} ${e.salary} (${e.dept.name})"); -- ref navigation
}
-- navigate a backlink (the department's staff), update through the result
for e in from s in dept.staff select s {
e.salary = e.salary + e.salary * 5 / 100; -- update-through-row
}
-- delete (restricted if still referenced)
let ok = try delete row catch (e) nil;
```
The query surface available today is **`from v in <table | relation> where …
[order by k [desc]] [take n] select v | v.field`**, plus `insert`, delete, and
update-through-a-row. It is proven end to end by the `employee` sample, whose
data survives a process restart via log replay.
---
## Project layout & the manifest
```
myproject/
├── wo.toml # manifest: name, version, [runtime], [build]
├── main.wo # entry point (fn main)
├── types.wo # your @table classes, other types
└── target/ # build output (the standalone binary lands here)
```
```toml
name = "myproject"
version = "0.1.0"
[runtime]
wo = ">= 0.1"
[build]
runtime = "../../../runtime/wovm" # path to the wovm the binary is built from
```
`woc myproject/` compiles every `.wo` file under the directory as one program.
Programs that create tables read their data directory from the `WO_DATA`
environment variable at run time:
```bash
WO_DATA=./data ./target/myproject seed
WO_DATA=./data ./target/myproject report # a fresh process still sees the data
```
---
## Worked examples
Two complete sample programs live in the repository and double as the language's
acceptance tests:
- **`docs/examples/employee/`** — departments and employees related by
`ref`/`backlink`, `@unique`, foreign-key restrict on delete, per-department
reports, and persistence across a restart. Run it:
```bash
just employee # compile + run every mode against a durable database
```
- **`docs/examples/log-watcher/`** — a long-running daemon that watches log
files for silent death, using the `fs`/`time`/`net`/`proc` stdlib. Run it:
```bash
just log-watcher
```
Read either program's `main.wo` for idiomatic, working writeonce.
---
## Roadmap
Planned, **not yet available** — listed so the shipped surface above stays
honest. These exist as design iterations and/or work-in-progress branches, not
as features you can use today:
- **Query aggregates** — `group … by … into g` with `count`/`avg`/`min`/`max`
and projection records. (Today the same result is written by hand from the
shipped primitives.)
- **HTTP service layer** — `service` blocks that route requests to methods.
- **Concurrency** — a shard-actor runtime and green-threaded fibers.
- **Cross-program database access** — one program attaching to another's
database over a local channel, with keypair authentication and per-client
rights.
- **Blue-green deployment** — in-process recompile and atomic version switch.
- **Compile-time metaprogramming** — `@derive(Json/Csv/Eq/…)` generated from a
class's own metadata, no reflection.
Known current limits worth naming: `net` is TCP host+port only; `proc.run` has
no timeout or signal control; there is no stdin/stdout byte I/O and no FFI.
---
*writeonce is a work in progress. Interfaces will change. If you build
something with it, pin to a commit.*

View file

@ -600,10 +600,23 @@ let manifest_parse (path : string) : (string * string) list =
if line = "" || (String.length line >= 1 && line.[0] = '#') then () if line = "" || (String.length line >= 1 && line.[0] = '#') then ()
else if line.[0] = '[' then begin else if line.[0] = '[' then begin
if line.[String.length line - 1] <> ']' then fail !lineno "malformed section header"; if line.[String.length line - 1] <> ']' then fail !lineno "malformed section header";
section := String.sub line 1 (String.length line - 2); (* accept `[[table.array]]` headers too (iteration 9c's
if !section <> "runtime" && !section <> "build" then [[share.clients]]) by trimming the doubled brackets *)
fail !lineno (Printf.sprintf "unknown section [%s] (runtime and build exist)" !section) let inner = String.sub line 1 (String.length line - 2) in
let inner =
if String.length inner >= 2 && inner.[0] = '[' && inner.[String.length inner - 1] = ']'
then String.sub inner 1 (String.length inner - 2)
else inner
in
section := inner;
if !section <> "runtime" && !section <> "build" && !section <> "share"
&& !section <> "share.clients"
then
fail !lineno
(Printf.sprintf "unknown section [%s] (runtime and build exist)" !section)
end end
else if !section = "share" || !section = "share.clients" then
() (* iteration 9c manifest keys — parsed by the attach feature, ignored here *)
else else
match String.index_opt line '=' with match String.index_opt line '=' with
| None -> fail !lineno "expected `key = \"value\"`" | None -> fail !lineno "expected `key = \"value\"`"

View file

@ -69,6 +69,9 @@ type field_ty =
| Ref of string | Ref of string
| Multi of string | Multi of string
| Map of string * string (* key type, value type: map<K, V> *) | Map of string * string (* key type, value type: map<K, V> *)
| Backlink of string * string (* backlink C.f: the computed inverse of a
`ref` — NOT a stored column; reading it
scans C's index on f. Types as multi C. *)
| Nullable of field_ty (* ?T wrapper *) | Nullable of field_ty (* ?T wrapper *)
(* Parameter passing convention (spec section 3, rule 2): default is an (* Parameter passing convention (spec section 3, rule 2): default is an
@ -211,6 +214,15 @@ and expr_kind =
| Binary of binop * expr * expr | Binary of binop * expr * expr
| Ctor of string * (string * expr) list | Ctor of string * (string * expr) list
| DbStub of Token.t list | DbStub of Token.t list
(* `insert Class { field: expr, ... }` — the FIRST DB statement to leave
the stub behind (iteration 9, Task 3). Typed like a constructor
literal, returns the new row's id (Int), legal in statement and
expression position both. `select` stays a DbStub until Task 5. *)
| Insert of string * (string * expr) list
(* `delete <row>` (iteration 9b): removes the row a table-class value
names; an expression yielding the deleted id (restrict/trap surfaces
through the engine like any DB fault, catchable). *)
| Delete of expr
(* haxe-parity Task 2: one `${expr}` interpolation site, produced only (* haxe-parity Task 2: one `${expr}` interpolation site, produced only
by the string-interpolation desugar (parser.ml) — never written by the string-interpolation desugar (parser.ml) — never written
directly by a parse rule the way every other expr_kind is. Its directly by a parse rule the way every other expr_kind is. Its
@ -274,6 +286,31 @@ and expr_kind =
ename : string; ename : string;
handler : stmt list; handler : stmt list;
} }
(* iteration 9b: a language-integrated query. `from <var> in <source>
where <e>* [group <e> by <k> into <g>] [order by <e> [desc]] [take <e>]
select <e>` — lowered to a bytecode loop over engine cursor builtins,
never SQL text. A table-class value is its row id at runtime, so field
access on a range variable reads through the engine. Slice scope today:
from/where/order/take/select and group-by aggregation; join is later. *)
| Query of query
and query_source =
| QTable of string (* a table class by name: `from e in Employee` *)
| QNav of expr (* a backlink/multi navigation: `from s in d.staff` *)
and query = {
q_var : string;
q_src : query_source;
q_wheres : expr list;
(* group <key_expr> by ... into <gvar>: present iff this is an aggregating
query. q_group_key is the whole grouped element (`e`), q_group_by the
key, q_gvar the group binding whose `.f` columns feed aggregates. *)
q_group : (string * expr) option; (* (gvar, key_expr) *)
q_order : (expr * bool) option; (* (key, desc?) *)
q_take : expr option;
q_select : expr;
q_pos : pos;
}
(* ---- statements (Task 5) --------------------------------------------- (* ---- statements (Task 5) ---------------------------------------------

View file

@ -164,7 +164,7 @@ let dump (img : string) : string =
let line fmt = Buffer.add_string out (fmt ^ "\n") in let line fmt = Buffer.add_string out (fmt ^ "\n") in
if u32 img 0 <> magic then raise (Bad "bad magic"); if u32 img 0 <> magic then raise (Bad "bad magic");
let ver = u32 img 4 in let ver = u32 img 4 in
if ver <> 2 then raise (Bad (Printf.sprintf "unsupported version %d" ver)); if ver <> 3 then raise (Bad (Printf.sprintf "unsupported version %d" ver));
let coff = u32 img 8 and ccnt = u32 img 12 in let coff = u32 img 8 and ccnt = u32 img 12 in
let koff = u32 img 16 and kcnt = u32 img 20 in let koff = u32 img 16 and kcnt = u32 img 20 in
let ioff = u32 img 24 and icnt = u32 img 28 in let ioff = u32 img 24 and icnt = u32 img 28 in
@ -213,6 +213,14 @@ let dump (img : string) : string =
renders as keys, so a wrong one is worth seeing. *) renders as keys, so a wrong one is worth seeing. *)
let names = List.init fcnt (fun j -> u32 img (!o + (j * 4))) in let names = List.init fcnt (fun j -> u32 img (!o + (j * 4))) in
o := !o + (fcnt * 12); o := !o + (fcnt * 12);
(* v3 index tail: walk past (the disassembly prints class shape, not
indexes — dump goldens stay byte-stable across the version bump) *)
let icnt = u32 img !o in
o := !o + 4;
for _ = 1 to icnt do
let ccnt = u32 img (!o + 4) in
o := !o + 8 + (ccnt * 4)
done;
let fields = let fields =
List.map2 List.map2
(fun nmk k -> if nmk = 0xFFFFFFFF then k else Printf.sprintf "%s:%s" (kname nmk) k) (fun nmk k -> if nmk = 0xFFFFFFFF then k else Printf.sprintf "%s:%s" (kname nmk) k)

View file

@ -150,6 +150,7 @@ let rec field_ty_str : Ast.field_ty -> string = function
| Ast.Ref s -> Printf.sprintf "ref %s" s | Ast.Ref s -> Printf.sprintf "ref %s" s
| Ast.Multi s -> Printf.sprintf "multi %s" s | Ast.Multi s -> Printf.sprintf "multi %s" s
| Ast.Map (k, v) -> Printf.sprintf "map<%s, %s>" k v | Ast.Map (k, v) -> Printf.sprintf "map<%s, %s>" k v
| Ast.Backlink (c, f) -> Printf.sprintf "backlink %s.%s" c f
| Ast.Nullable t -> "?" ^ field_ty_str t | Ast.Nullable t -> "?" ^ field_ty_str t
let param_str (p : Ast.param) : string = Printf.sprintf "%s%s: %s" (conv_str p.conv) p.name (field_ty_str p.ty) let param_str (p : Ast.param) : string = Printf.sprintf "%s%s: %s" (conv_str p.conv) p.name (field_ty_str p.ty)
@ -222,6 +223,19 @@ let rec expr_str (e : Ast.expr) : string =
Printf.sprintf "%s { %s }" name Printf.sprintf "%s { %s }" name
(String.concat ", " (String.concat ", "
(List.map (fun (fname, fval) -> Printf.sprintf "%s: %s" fname (expr_str fval)) fields)) (List.map (fun (fname, fval) -> Printf.sprintf "%s: %s" fname (expr_str fval)) fields))
| Ast.Insert (name, fields) ->
Printf.sprintf "INSERT %s { %s }" name
(String.concat ", "
(List.map (fun (fname, fval) -> Printf.sprintf "%s: %s" fname (expr_str fval)) fields))
| Ast.Query q ->
let src = match q.Ast.q_src with Ast.QTable cn -> cn | Ast.QNav e -> expr_str e in
Printf.sprintf "QUERY from %s in %s%s%s select %s" q.Ast.q_var src
(String.concat "" (List.map (fun w -> " where " ^ expr_str w) q.Ast.q_wheres))
(match q.Ast.q_group with
| Some (g, k) -> Printf.sprintf " group by %s into %s" (expr_str k) g
| None -> "")
(expr_str q.Ast.q_select)
| Ast.Delete t -> Printf.sprintf "DELETE %s" (expr_str t)
| Ast.DbStub toks -> Printf.sprintf "DB_STUB(%s)" (dbstub_tokens_str toks) | Ast.DbStub toks -> Printf.sprintf "DB_STUB(%s)" (dbstub_tokens_str toks)
| Ast.Interp inner -> Printf.sprintf "INTERP(%s)" (expr_str inner) | Ast.Interp inner -> Printf.sprintf "INTERP(%s)" (expr_str inner)
| Ast.ListLit items -> Printf.sprintf "[%s]" (String.concat ", " (List.map expr_str items)) | Ast.ListLit items -> Printf.sprintf "[%s]" (String.concat ", " (List.map expr_str items))

View file

@ -152,7 +152,7 @@ let stdlib_not_linked_code = Diag.emitter_prefix ^ "06"
============================================================ *) ============================================================ *)
let wob_magic = 0x31424F57 (* "WOB1" read as an LE u32 *) let wob_magic = 0x31424F57 (* "WOB1" read as an LE u32 *)
let wob_version = 2 (* v2: per-field class-table metadata *) let wob_version = 3 (* v3: v2 + per-class secondary-index metadata *)
let wob_hdr_size = 44 let wob_hdr_size = 44
let wob_none = 0xFFFFFFFF let wob_none = 0xFFFFFFFF
let k_int = 0 let k_int = 0
@ -257,6 +257,7 @@ let b_map_val_at = 38
let b_multi_set = 39 let b_multi_set = 39
let b_map_get_opt = 59 let b_map_get_opt = 59
let b_text_copy = 60 let b_text_copy = 60
let b_db_insert = 61
(* json (runtime/src/json.c): encode takes the value's static kind as its (* json (runtime/src/json.c): encode takes the value's static kind as its
second argument, decode the class id to build as its second. *) second argument, decode the class id to build as its second. *)
@ -348,6 +349,16 @@ type clsrec = {
cr_gc : bool; cr_gc : bool;
cr_fields : (string * Ast.field_ty) array; cr_fields : (string * Ast.field_ty) array;
cr_methods : string list; (* method names, declaration order *) cr_methods : string list; (* method names, declaration order *)
(* iteration 9 Task 4: (unique, column indices) per secondary index —
`@table(index: [a, b])` entries (non-unique, composite) plus one
unique single-column entry per `@unique` field. Serialized as the v3
class-record tail; the engine builds its runtime indexes from this. *)
cr_indexes : (bool * int array) list;
cr_is_table : bool; (* has @table — its instances are row ids (iteration 9b) *)
(* backlink fields (iteration 9b): name -> (source class, source field).
Virtual — not in cr_fields, no stored column; `d.staff` reads them by
probing the source class's index on the source field. *)
cr_backlinks : (string * (string * string)) list;
} }
type ifacerec = { type ifacerec = {
@ -772,6 +783,44 @@ let field_kind (p : pctx) (ft : Ast.field_ty) : int =
let class_of_name (p : pctx) (n : string) : int option = SM.find_opt n p.p_class_id let class_of_name (p : pctx) (n : string) : int option = SM.find_opt n p.p_class_id
(* iteration 9b: a @table class's instances are row ids, so field access on
one reads through the engine (DB_GET_FIELD) rather than GETF. *)
let is_table_class (p : pctx) (cid : int) : bool =
cid >= 0 && cid < Array.length p.p_classes && p.p_classes.(cid).cr_is_table
let b_str_lt = 67
let b_db_update_field = 62
let b_db_delete = 63
let b_db_scan = 64
let b_db_get_field = 65
let b_db_probe = 66
(* iteration 9b: `d.staff` where staff is `backlink Employee.dept` reads by
probing Employee's index on its `dept` column. Resolve to (source cid,
index number) — None if the source field is not a declared index (a
backlink without a backing index has no efficient read and is rejected). *)
let backlink_target (p : pctx) (base_cid : int) (fname : string) : (int * int) option =
match List.assoc_opt fname p.p_classes.(base_cid).cr_backlinks with
| None -> None
| Some (src_class, src_field) -> (
match class_of_name p src_class with
| None -> None
| Some scid ->
let sc = p.p_classes.(scid) in
(* stored column index of the source field *)
let col = ref (-1) in
Array.iteri (fun i (n, _) -> if n = src_field then col := i) sc.cr_fields;
if !col < 0 then None
else
(* the index whose single column is that field *)
let rec find n = function
| [] -> None
| (_, cols) :: tl ->
if Array.length cols = 1 && cols.(0) = !col then Some (scid, n)
else find (n + 1) tl
in
find 0 sc.cr_indexes)
let field_of (p : pctx) (cid : int) (fname : string) : (int * Ast.field_ty) option = let field_of (p : pctx) (cid : int) (fname : string) : (int * Ast.field_ty) option =
let fs = p.p_classes.(cid).cr_fields in let fs = p.p_classes.(cid).cr_fields in
let rec go i = if i >= Array.length fs then None else let rec go i = if i >= Array.length fs then None else
@ -941,6 +990,22 @@ let variant_tag_value (p : pctx) (u : Types.union_info) (vi : Types.variant_info
| None -> 0 (* unreachable: pass 1 registers every payload-union variant *) | None -> 0 (* unreachable: pass 1 registers every payload-union variant *)
else vi.Types.vi_tag else vi.Types.vi_tag
(* iteration 9b: a query's element type, as the name a `Multi` carries.
`select x` yields the source class (a row id typed as the class);
`select x.field` yields that field's type; anything else falls back to
Int (the slice's shapes are these two). *)
let query_elem_scalar (p : pctx) (q : Ast.query) ~(src : string) : string =
match q.Ast.q_select.Ast.kind with
| Ast.Ident v when v = q.Ast.q_var -> src (* select the whole row: element = source class *)
| Ast.Field ({ Ast.kind = Ast.Ident v; _ }, fname) when v = q.Ast.q_var -> (
match class_of_name p src with
| Some cid -> (
match field_of p cid fname with
| Some (_, ty) -> ( match unwrap ty with Scalar n -> n | _ -> "Int")
| None -> "Int")
| None -> "Int")
| _ -> "Int"
let rec ty_of_expr (p : pctx) (f : fstate) (e : Ast.expr) : Ast.field_ty option = let rec ty_of_expr (p : pctx) (f : fstate) (e : Ast.expr) : Ast.field_ty option =
match e.kind with match e.kind with
| IntLit _ -> Some (Scalar "Int") | IntLit _ -> Some (Scalar "Int")
@ -970,10 +1035,17 @@ let rec ty_of_expr (p : pctx) (f : fstate) (e : Ast.expr) : Ast.field_ty option
| Field (base, fname) -> ( | Field (base, fname) -> (
match ty_of_expr p f base with match ty_of_expr p f base with
| Some bt -> ( | Some bt -> (
match unwrap bt with (* a `ref C` navigates into C: the target is a table row id *)
match (match unwrap bt with Ref c -> Scalar c | other -> other) with
| Scalar cn -> ( | Scalar cn -> (
match class_of_name p cn with match class_of_name p cn with
| Some cid -> ( match field_of p cid fname with Some (_, t) -> Some t | None -> None) | Some cid -> (
match field_of p cid fname with
| Some (_, t) -> Some t
| None -> (
match List.assoc_opt fname p.p_classes.(cid).cr_backlinks with
| Some (sc, _) -> Some (Multi sc)
| None -> None))
| None -> None) | None -> None)
| _ -> None) | _ -> None)
| None -> None) | None -> None)
@ -1054,6 +1126,16 @@ let rec ty_of_expr (p : pctx) (f : fstate) (e : Ast.expr) : Ast.field_ty option
| Eq | Ne | Lt | Le | Gt | Ge | And | Or -> Some (Scalar "Bool") | Eq | Ne | Lt | Le | Gt | Ge | And | Or -> Some (Scalar "Bool")
| Add | Sub | Mul | Div | Mod -> ( match ty_of_expr p f l with Some t -> Some t | None -> Some (Scalar "Int"))) | Add | Sub | Mul | Div | Mod -> ( match ty_of_expr p f l with Some t -> Some t | None -> Some (Scalar "Int")))
| Ctor (cn, _) -> Some (Scalar cn) | Ctor (cn, _) -> Some (Scalar cn)
| Insert _ -> Some (Scalar "Int")
| Delete _ -> Some (Scalar "Int")
| Query q ->
let src =
match q.Ast.q_src with
| Ast.QTable cn -> cn
| Ast.QNav nav -> (
match ty_of_expr p f nav with Some t -> (match unwrap t with Multi c -> c | Scalar c -> c | _ -> "") | None -> "")
in
Some (Multi (query_elem_scalar p q ~src))
| Interp _ -> Some (Scalar "Text") | Interp _ -> Some (Scalar "Text")
| DbStub _ -> None | DbStub _ -> None
| Switch (subject, arms) -> ( | Switch (subject, arms) -> (
@ -1344,7 +1426,8 @@ let field_class_meta (p : pctx) (ty : Ast.field_ty) : int =
match name_of (Ast.Scalar e) with match name_of (Ast.Scalar e) with
| Some n -> ( match class_of_name p n with Some cid -> cid | None -> wob_none) | Some n -> ( match class_of_name p n with Some cid -> cid | None -> wob_none)
| None -> wob_none) | None -> wob_none)
| Ast.Ref _ | Ast.Nullable _ -> wob_none | Ast.Ref n -> ( match class_of_name p n with Some cid -> cid | None -> wob_none)
| Ast.Backlink _ | Ast.Nullable _ -> wob_none
let field_elem_meta (p : pctx) (ty : Ast.field_ty) : int = let field_elem_meta (p : pctx) (ty : Ast.field_ty) : int =
match unwrap ty with match unwrap ty with
@ -1583,9 +1666,22 @@ let rec emit_expr (p : pctx) (f : fstate) (v : views) ~(dst : int) ?expected (e
| Field (base, fname) -> ( | Field (base, fname) -> (
match ty_of_expr p f base with match ty_of_expr p f base with
| Some bt -> ( | Some bt -> (
match unwrap bt with match (match unwrap bt with Ref c -> Scalar c | other -> other) with
| Scalar cn -> ( | Scalar cn -> (
match class_of_name p cn with match class_of_name p cn with
| Some cid when is_table_class p cid && backlink_target p cid fname <> None -> (
(* `d.staff`: probe the source class's index for rows referencing
this row's id. Window: [class, index, key(=base id)]. *)
match backlink_target p cid fname with
| Some (scid, ino) ->
let b = emit_operand p f v base in
let w = alloc_temps p f e.pos 3 in
put f (ins_abx op_loadk w (check_bx p f e.pos "constant" (const_int p scid)));
put f (ins_abx op_loadk (w + 1) (check_bx p f e.pos "constant" (const_int p ino)));
put f (ins_abc op_move (w + 2) b 0);
sync_mask p f v e.id;
put f (ins_abc op_builtin dst w b_db_probe)
| None -> ())
| Some cid -> ( | Some cid -> (
match field_of p cid fname with match field_of p cid fname with
| Some (idx, _) -> | Some (idx, _) ->
@ -1603,7 +1699,17 @@ let rec emit_expr (p : pctx) (f : fstate) (v : views) ~(dst : int) ?expected (e
f.f_stmt_drops <- g :: f.f_stmt_drops; f.f_stmt_drops <- g :: f.f_stmt_drops;
f.f_esc_drops <- g :: f.f_esc_drops f.f_esc_drops <- g :: f.f_esc_drops
end; end;
put f (ins_abc op_getf dst b (check_field_idx p f e.pos idx)) if is_table_class p cid then begin
(* a table-class value is its row id; read the column from the
engine. Window: [class-id, id, field-idx]. *)
let w = alloc_temps p f e.pos 3 in
put f (ins_abx op_loadk w (check_bx p f e.pos "constant" (const_int p cid)));
put f (ins_abc op_move (w + 1) b 0);
put f (ins_abx op_loadk (w + 2) (check_bx p f e.pos "constant" (const_int p idx)));
sync_mask p f v e.id;
put f (ins_abc op_builtin dst w b_db_get_field)
end
else put f (ins_abc op_getf dst b (check_field_idx p f e.pos idx))
| None -> | None ->
err p ~code:cannot_lower_code ~file:f.f_file ~pos:e.pos err p ~code:cannot_lower_code ~file:f.f_file ~pos:e.pos
~message:(Printf.sprintf "`%s` has no field `%s`" cn fname); ~message:(Printf.sprintf "`%s` has no field `%s`" cn fname);
@ -1654,6 +1760,43 @@ let rec emit_expr (p : pctx) (f : fstate) (v : views) ~(dst : int) ?expected (e
put f (ins_abc op_neg dst b 0) put f (ins_abc op_neg dst b 0)
| Binary (op, l, r) -> emit_binary p f v ~dst op l r | Binary (op, l, r) -> emit_binary p f v ~dst op l r
| Ctor (cn, fields) -> emit_ctor p f v ~dst e cn fields | Ctor (cn, fields) -> emit_ctor p f v ~dst e cn fields
| Insert (cn, fields) -> emit_insert p f v ~dst e cn fields
| Delete target -> (
match ty_of_expr p f target with
| Some bt -> (
match (match unwrap bt with Ref c -> Scalar c | o -> o) with
| Scalar cn -> (
match class_of_name p cn with
| Some cid when is_table_class p cid ->
(* reserve dst past the window: in tail position dst == the first
window reg, and moving the id into dst would clobber the class
id — the disassembly-caught bug *)
let outer = f.f_temp in
if f.f_temp <= dst then f.f_temp <- dst + 1;
let w = alloc_temps p f e.pos 2 in
put f (ins_abx op_loadk w (check_bx p f e.pos "constant" (const_int p cid)));
let save = f.f_temp in
emit_expr p f v ~dst:(w + 1) target;
f.f_temp <- save;
(* keep the id so `delete x` can be used as an expression *)
put f (ins_abc op_move dst (w + 1) 0);
sync_mask p f v e.id;
f.f_cur_line <- e.pos.line;
put f (ins_abc op_builtin w w b_db_delete);
f.f_temp <- outer
| _ ->
err p ~code:cannot_lower_code ~file:f.f_file ~pos:e.pos
~message:"`delete` target is not a table row";
put f (ins_abx op_loadk dst (const_int p 0)))
| _ ->
err p ~code:cannot_lower_code ~file:f.f_file ~pos:e.pos
~message:"`delete` target is not a table row";
put f (ins_abx op_loadk dst (const_int p 0)))
| None ->
err p ~code:cannot_lower_code ~file:f.f_file ~pos:e.pos
~message:"cannot resolve the `delete` target's type";
put f (ins_abx op_loadk dst (const_int p 0)))
| Query q -> emit_query p f v ~dst e q
| Interp inner -> ( | Interp inner -> (
(* haxe-parity Task 2: the type-directed half of the interpolation (* haxe-parity Task 2: the type-directed half of the interpolation
desugar (parser.ml's own doc comment on Ast.Interp) — a Text desugar (parser.ml's own doc comment on Ast.Interp) — a Text
@ -2347,6 +2490,324 @@ and emit_ctor (p : pctx) (f : fstate) (v : views) ~(dst : int) (e : Ast.expr) (c
ci.Types.fields); ci.Types.fields);
f.f_temp <- outer f.f_temp <- outer
(* iteration 9 Task 3: `insert Class { ... }` lowers to one DB_INSERT
builtin whose window is [class-id const, then one slot per DECLARED
field in declaration order] — the executor walks the class table's
kinds, so slot order must be the table's, not the literal's. A field
the literal omits gets its default (same emit_default_value the ctor
uses) or, for a `?` field, its kind's own nil (WO_NIL_SCALAR for a
nullable scalar, the zero word otherwise). The engine COPIES every
value at the row API, so after the builtin every freshly built
argument is still this frame's to drop — same reap as push/set. *)
and emit_query (p : pctx) (f : fstate) (v : views) ~(dst : int) (e : Ast.expr)
(q : Ast.query) : unit =
(* iteration 9b slice: from/where/select over a table scan. group/order/
take/navigation are diagnosed in types.ml, so a written image never
reaches this with them set. Lowered to an ordinary bytecode loop over
DB_SCAN's materialized id list — no plan tree, no text. *)
let cn =
match q.Ast.q_src with
| Ast.QTable cn -> cn
| Ast.QNav nav -> (
(* the source's element type is the range var's class *)
match ty_of_expr p f nav with
| Some t -> ( match unwrap t with Multi c -> c | Scalar c -> c | _ -> "")
| None -> "")
in
match class_of_name p cn with
| None ->
err p ~code:cannot_lower_code ~file:f.f_file ~pos:e.pos
~message:(Printf.sprintf "query over `%s`, which is not a declared table class" cn);
put f (ins_abx op_loadk dst (const_int p 0))
| Some cid ->
let elem_name = query_elem_scalar p q ~src:cn in
let elem = Scalar elem_name in
(* a table-class element is a row ID (a scalar), not a heap pointer — so
the result container is SCALAR-kinded even though the element TYPES as
the class; getting this wrong drops an id as a pointer (ASan SEGV) *)
let elem_kind =
match class_of_name p elem_name with
| Some ecid when is_table_class p ecid -> 0 (* WO_K_SCALAR *)
| _ -> field_kind p elem
in
(* reserve dst past the loop's working registers (same guard emit_ctor
uses): dst holds the result multi every push writes into *)
let outer = f.f_temp in
if f.f_temp <= dst then f.f_temp <- dst + 1;
(* loop-carried registers, allocated once above dst, never reset *)
let scan = alloc_temp p f e.pos in
let idx = alloc_temp p f e.pos in
let len = alloc_temp p f e.pos in
let idreg = alloc_temp p f e.pos in
let body_base = f.f_temp in
(* scan -> multi of ids; result multi -> dst *)
sync_mask p f v e.id;
f.f_cur_line <- e.pos.line;
(match q.Ast.q_src with
| Ast.QTable _ ->
put f (ins_abx op_loadk scan (check_bx p f e.pos "constant" (const_int p cid)));
put f (ins_abc op_builtin scan scan b_db_scan)
| Ast.QNav nav ->
(* the navigation (a backlink) already yields a multi of source ids *)
let save = f.f_temp in
f.f_temp <- scan + 1;
emit_expr p f v ~dst:scan nav;
f.f_temp <- save);
put f (ins_abc op_builtin dst elem_kind b_multi_new);
put f (ins_abc op_builtin len scan b_len);
put f (ins_abx op_loadk idx (check_bx p f e.pos "constant" (const_int p 0)));
(* bind the range var to the current id (typed as the class), so field
access inside where/select routes through DB_GET_FIELD *)
let saved_env = f.f_env in
f.f_env <- (q.Ast.q_var, (idreg, Scalar cn)) :: f.f_env;
ignore body_base;
let top = here f in
f.f_temp <- body_base;
let tc = alloc_temp p f e.pos in
put f (ins_abc op_lt tc idx len);
let jz_exit = here f in
put f (ins_asbx op_jz tc 0);
(* id = multi_get(scan, idx) *)
let w = alloc_temps p f e.pos 2 in
put f (ins_abc op_move w scan 0);
put f (ins_abc op_move (w + 1) idx 0);
put f (ins_abc op_builtin idreg w b_multi_get);
(* where guards: any false skips the push *)
let skips = ref [] in
List.iter
(fun w_expr ->
let save = f.f_temp in
let wr = emit_operand p f v w_expr in
skips := here f :: !skips;
put f (ins_asbx op_jz wr 0);
f.f_temp <- save)
q.Ast.q_wheres;
(* select -> push into dst (copying a Text element the container owns) *)
let save = f.f_temp in
let sel = alloc_temp p f e.pos in
emit_expr p f v ~dst:sel q.Ast.q_select;
if elem_kind = 3 then put f (ins_abc op_builtin sel sel b_text_copy);
let pw = alloc_temps p f e.pos 2 in
put f (ins_abc op_move pw dst 0);
put f (ins_abc op_move (pw + 1) sel 0);
put f (ins_abc op_builtin pw pw b_multi_push);
f.f_temp <- save;
(* skip target: increment and loop *)
let cont = here f in
List.iter (fun pc -> patch_jump p f ~file:f.f_file ~pos:e.pos pc cont) !skips;
f.f_temp <- body_base;
let one = alloc_temp p f e.pos in
put f (ins_abx op_loadk one (check_bx p f e.pos "constant" (const_int p 1)));
put f (ins_abc op_add idx idx one);
let back = here f in
put f (ins_asbx op_jmp 0 0);
patch_jump p f ~file:f.f_file ~pos:e.pos back top;
let exit_pc = here f in
patch_jump p f ~file:f.f_file ~pos:e.pos jz_exit exit_pc;
f.f_env <- saved_env;
(* the scan's id list was this query's own, dropped now *)
put f (ins_abc op_drop scan 0 0);
(* ---- order by (whole-row selection sort) ----------------------------
Elements of dst are row ids; the key re-reads a field through the
range var. Selection sort is O(n^2) but the result sets here are
small and this is KISS by design (no cost planner). Only the
whole-row + field-key shape is supported; grouped/projection ordering
lands with group-by. *)
(match q.Ast.q_order with
| Some (key, desc) ->
f.f_temp <- body_base;
let n = alloc_temp p f e.pos in
put f (ins_abc op_builtin n dst b_count);
let i = alloc_temp p f e.pos in
let j = alloc_temp p f e.pos in
let best = alloc_temp p f e.pos in
let elem_j = alloc_temp p f e.pos in
let elem_b = alloc_temp p f e.pos in
let sort_scratch = f.f_temp in
put f (ins_abx op_loadk i (check_bx p f e.pos "constant" (const_int p 0)));
let oi = here f in (* outer: while i < n *)
let oc = alloc_temp p f e.pos in
put f (ins_abc op_lt oc i n);
let ojz = here f in
put f (ins_asbx op_jz oc 0);
put f (ins_abc op_move best i 0);
let oneA = alloc_temp p f e.pos in
put f (ins_abx op_loadk oneA (check_bx p f e.pos "constant" (const_int p 1)));
put f (ins_abc op_add j i oneA);
let ij = here f in (* inner: while j < n *)
let ic = alloc_temp p f e.pos in
put f (ins_abc op_lt ic j n);
let ijz = here f in
put f (ins_asbx op_jz ic 0);
(* elem_j = multi_get(dst,j); elem_b = multi_get(dst,best) *)
let gw = alloc_temps p f e.pos 2 in
put f (ins_abc op_move gw dst 0);
put f (ins_abc op_move (gw + 1) j 0);
put f (ins_abc op_builtin elem_j gw b_multi_get);
put f (ins_abc op_move (gw + 1) best 0);
put f (ins_abc op_builtin elem_b gw b_multi_get);
(* keys: bind range var to elem_j / elem_b, eval key expr *)
let saved_env2 = f.f_env in
f.f_temp <- sort_scratch;
f.f_env <- (q.Ast.q_var, (elem_j, Scalar cn)) :: saved_env2;
(* key kind must be read with the range var BOUND — else ty_of_expr of
`x.name` sees x unbound, returns None, and a Text key silently falls
to the pointer-comparing op_lt (the wrong-order bug) *)
let key_is_text =
match ty_of_expr p f key with Some t -> field_kind p t = 3 | None -> false
in
let kj = alloc_temp p f e.pos in
emit_expr p f v ~dst:kj key;
f.f_env <- (q.Ast.q_var, (elem_b, Scalar cn)) :: saved_env2;
let kb = alloc_temp p f e.pos in
emit_expr p f v ~dst:kb key;
f.f_env <- saved_env2;
(* cmp: for asc, kj < kb -> best=j; for desc, kj > kb (== kb < kj). *)
let cmp = alloc_temp p f e.pos in
let lt a b =
if key_is_text then begin
let save = f.f_temp in
let w = alloc_temps p f e.pos 2 in
put f (ins_abc op_move w a 0);
put f (ins_abc op_move (w + 1) b 0);
put f (ins_abc op_builtin cmp w b_str_lt);
f.f_temp <- save
end
else put f (ins_abc op_lt cmp a b)
in
if desc then lt kb kj else lt kj kb;
let cjz = here f in
put f (ins_asbx op_jz cmp 0);
put f (ins_abc op_move best j 0);
let after = here f in
patch_jump p f ~file:f.f_file ~pos:e.pos cjz after;
f.f_temp <- sort_scratch;
let oneB = alloc_temp p f e.pos in
put f (ins_abx op_loadk oneB (check_bx p f e.pos "constant" (const_int p 1)));
put f (ins_abc op_add j j oneB);
let iback = here f in
put f (ins_asbx op_jmp 0 0);
patch_jump p f ~file:f.f_file ~pos:e.pos iback ij;
let iexit = here f in
patch_jump p f ~file:f.f_file ~pos:e.pos ijz iexit;
(* swap dst[i], dst[best]: read both, multi_set both *)
f.f_temp <- sort_scratch;
let vi = alloc_temp p f e.pos in
let vb = alloc_temp p f e.pos in
let sw = alloc_temps p f e.pos 3 in
put f (ins_abc op_move sw dst 0);
put f (ins_abc op_move (sw + 1) i 0);
put f (ins_abc op_builtin vi sw b_multi_get);
put f (ins_abc op_move (sw + 1) best 0);
put f (ins_abc op_builtin vb sw b_multi_get);
(* dst[i] = vb *)
put f (ins_abc op_move sw dst 0);
put f (ins_abc op_move (sw + 1) i 0);
put f (ins_abc op_move (sw + 2) vb 0);
put f (ins_abc op_builtin sw sw b_multi_set);
(* dst[best] = vi *)
put f (ins_abc op_move sw dst 0);
put f (ins_abc op_move (sw + 1) best 0);
put f (ins_abc op_move (sw + 2) vi 0);
put f (ins_abc op_builtin sw sw b_multi_set);
f.f_temp <- sort_scratch;
let oneC = alloc_temp p f e.pos in
put f (ins_abx op_loadk oneC (check_bx p f e.pos "constant" (const_int p 1)));
put f (ins_abc op_add i i oneC);
let oback = here f in
put f (ins_asbx op_jmp 0 0);
patch_jump p f ~file:f.f_file ~pos:e.pos oback oi;
let oexit = here f in
patch_jump p f ~file:f.f_file ~pos:e.pos ojz oexit
| None -> ());
(* ---- take N: slice dst to [0, N) --------------------------------- *)
(match q.Ast.q_take with
| Some tk ->
f.f_temp <- body_base;
let nreg = alloc_temp p f e.pos in
emit_expr p f v ~dst:nreg tk;
(* clamp N to count(dst) so slice never runs past the end *)
let cnt = alloc_temp p f e.pos in
put f (ins_abc op_builtin cnt dst b_count);
let over = alloc_temp p f e.pos in
put f (ins_abc op_lt over cnt nreg); (* count < N ? use count *)
let jz2 = here f in
put f (ins_asbx op_jz over 0);
put f (ins_abc op_move nreg cnt 0);
let aft = here f in
patch_jump p f ~file:f.f_file ~pos:e.pos jz2 aft;
let sw = alloc_temps p f e.pos 3 in
let zero = alloc_temp p f e.pos in
put f (ins_abx op_loadk zero (check_bx p f e.pos "constant" (const_int p 0)));
put f (ins_abc op_move sw dst 0);
put f (ins_abc op_move (sw + 1) zero 0);
put f (ins_abc op_move (sw + 2) nreg 0);
let sliced = alloc_temp p f e.pos in
put f (ins_abc op_builtin sliced sw b_slice);
put f (ins_abc op_drop dst 0 0); (* the pre-slice multi is discarded *)
put f (ins_abc op_move dst sliced 0)
| None -> ());
f.f_temp <- outer
and emit_insert (p : pctx) (f : fstate) (v : views) ~(dst : int) (e : Ast.expr) (cn : string)
(fields : (string * Ast.expr) list) : unit =
match class_of_name p cn with
| None ->
err p ~code:cannot_lower_code ~file:f.f_file ~pos:e.pos
~message:(Printf.sprintf "insert into `%s`, which is not a declared class" cn);
put f (ins_abx op_loadk dst (const_int p 0))
| Some cid ->
let fcnt = Array.length p.p_classes.(cid).cr_fields in
let base = alloc_temps p f e.pos (fcnt + 1) in
put f (ins_abx op_loadk base (check_bx p f e.pos "constant" (const_int p cid)));
(* every field the literal names lands in ITS declared slot *)
List.iter
(fun ((fname : string), (fe : Ast.expr)) ->
match field_of p cid fname with
| None ->
err p ~code:cannot_lower_code ~file:f.f_file ~pos:e.pos
~message:(Printf.sprintf "`%s` has no field `%s`" cn fname)
| Some (idx, fty) ->
let save = f.f_temp in
emit_expr p f v ~dst:(base + 1 + idx) ~expected:fty fe;
f.f_temp <- save)
fields;
(* omitted fields: declared default, else the kind's own nil *)
let provided = List.map fst fields in
(match Types.StringMap.find_opt cn p.p_syms.Types.classes with
| None -> ()
| Some (ci : Types.class_info) ->
List.iter
(fun (fname, fty, fdefault, _) ->
if not (List.mem fname provided) then
match field_of p cid fname with
| None -> ()
| Some (idx, dfty) -> (
match fdefault with
| Some d ->
let save = f.f_temp in
emit_default_value p f ~dst:(base + 1 + idx) ~fty:dfty ~pos:e.pos d;
f.f_temp <- save
| None ->
let nil_word =
if is_nullable_scalar p fty then const_int p nil_scalar_word
else const_int p 0
in
put f (ins_abx op_loadk (base + 1 + idx) (check_bx p f e.pos "constant" nil_word))))
ci.Types.fields);
sync_mask p f v e.id;
f.f_cur_line <- e.pos.line;
put f (ins_abc op_builtin dst base b_db_insert);
(* the engine copied: fresh argument values die here *)
List.iter
(fun ((fname : string), (fe : Ast.expr)) ->
match field_of p cid fname with
| None -> ()
| Some (idx, _) ->
drop_fresh_owned ~keep:dst p f (base + 1 + idx) fe;
drop_fresh_text ~keep:dst p f (base + 1 + idx) fe)
fields
(* The default expressions the emitter can lower (haxe-parity Task 4): (* The default expressions the emitter can lower (haxe-parity Task 4):
the literal shapes the sample's own typedefs use — Int (optionally the literal shapes the sample's own typedefs use — Int (optionally
negated), Text, Bool, `now()` (parse_default_expr's own recognized negated), Text, Bool, `now()` (parse_default_expr's own recognized
@ -3224,6 +3685,27 @@ and emit_assign (p : pctx) (f : fstate) (v : views) (s : Ast.stmt) (target : Ast
| None -> | None ->
err p ~code:cannot_lower_code ~file:f.f_file ~pos:target.pos err p ~code:cannot_lower_code ~file:f.f_file ~pos:target.pos
~message:(Printf.sprintf "assignment into `%s`, which is not a declared class" cn) ~message:(Printf.sprintf "assignment into `%s`, which is not a declared class" cn)
| Some cid when is_table_class p cid -> (
(* iteration 9b: `e.salary = v` where e is a table row updates the
engine (DB_UPDATE_FIELD: class, id, field, value) — the row's
own indexes are maintained at the choke point *)
match field_of p cid fname with
| None ->
err p ~code:cannot_lower_code ~file:f.f_file ~pos:target.pos
~message:(Printf.sprintf "`%s` has no field `%s`" cn fname)
| Some (idx, fty) ->
let w = alloc_temps p f target.pos 4 in
put f (ins_abx op_loadk w (check_bx p f target.pos "constant" (const_int p cid)));
let save = f.f_temp in
emit_expr p f v ~dst:(w + 1) base;
f.f_temp <- save;
put f (ins_abx op_loadk (w + 2) (check_bx p f target.pos "constant" (const_int p idx)));
let save = f.f_temp in
emit_expr p f v ~dst:(w + 3) ~expected:fty value;
f.f_temp <- save;
sync_mask p f v s.s_id;
f.f_cur_line <- s.s_pos.line;
put f (ins_abc op_builtin w w b_db_update_field))
| Some cid -> ( | Some cid -> (
match field_of p cid fname with match field_of p cid fname with
| None -> | None ->
@ -3852,6 +4334,18 @@ let emit ~(syms : Types.symbols) ~(module_of : string -> string)
~(module_syms : (string, Types.symbols) Hashtbl.t) (coll : Diag.Collector.t) (units : input list) : ~(module_syms : (string, Types.symbols) Hashtbl.t) (coll : Diag.Collector.t) (units : input list) :
string = string =
let colliding = compute_colliding_fn_names ~module_of units in let colliding = compute_colliding_fn_names ~module_of units in
(* iteration 9 Task 4: index-declaration problems found while building
clsrecs — reported once a file/pos-bearing context exists below *)
let index_col_err : (Ast.pos * string) option ref = ref None in
let index_err_file = ref "" in
let ref_index_of_name (fnames : string list) (n : string) : int =
let rec go i = function
| [] -> 0 (* unknown column: the caller records the diagnostic *)
| x :: tl -> if x = n then i else go (i + 1) tl
in
go 0 fnames
in
let p_syms_for_indexes = syms in
(* ---- pass 1: declarations, in discovery then declaration order ---- *) (* ---- pass 1: declarations, in discovery then declaration order ---- *)
let classes = ref [] and class_id = ref SM.empty and nclasses = ref 0 in let classes = ref [] and class_id = ref SM.empty and nclasses = ref 0 in
let ifaces = ref [] and iface_id = ref SM.empty and nifaces = ref 0 and nslots = ref 0 in let ifaces = ref [] and iface_id = ref SM.empty and nifaces = ref 0 and nslots = ref 0 in
@ -3890,6 +4384,7 @@ let emit ~(syms : Types.symbols) ~(module_of : string -> string)
(function (function
| Ast.Class (c : Ast.class_decl) -> | Ast.Class (c : Ast.class_decl) ->
if not (SM.mem c.name !class_id) then begin if not (SM.mem c.name !class_id) then begin
(if !index_col_err = None then index_err_file := u.file);
let shape = if c.is_record then Some (record_shape_key c) else None in let shape = if c.is_record then Some (record_shape_key c) else None in
let alias_of = let alias_of =
match shape with Some key -> Hashtbl.find_opt record_shape key | None -> None match shape with Some key -> Hashtbl.find_opt record_shape key | None -> None
@ -3907,10 +4402,77 @@ let emit ~(syms : Types.symbols) ~(module_of : string -> string)
| Some key -> Hashtbl.replace record_shape key cid | Some key -> Hashtbl.replace record_shape key cid
| None -> ()); | None -> ());
classes := classes :=
{ cr_name = c.name; cr_gc = c.is_gc; (let fnames =
cr_fields = List.filter_map
Array.of_list (List.map (fun (fl : Ast.field) -> (fl.name, fl.ty)) c.fields); (fun (fl : Ast.field) ->
cr_methods = List.map (fun (m : Ast.method_decl) -> m.name) c.methods } match fl.Ast.ty with Ast.Backlink _ -> None | _ -> Some fl.Ast.name)
c.fields
in
let col_of n = ref_index_of_name fnames n in
let is_indexable (fl : Ast.field) =
match Types.wob_kind_of_typ p_syms_for_indexes (Types.typ_of_field_ty (unwrap fl.Ast.ty)) with
| Types.WO_K_SCALAR | Types.WO_K_TEXT -> true
| _ -> false
in
let table_indexes =
match c.Ast.table with
| None -> []
| Some cfg ->
List.map
(fun cols -> (false, Array.of_list (List.map col_of cols)))
cfg.Ast.indexes
in
let unique_indexes =
List.concat_map
(fun (fl : Ast.field) ->
if List.mem "unique" fl.Ast.annotations then begin
if not (is_indexable fl) then
index_col_err := Some (c.Ast.pos, Printf.sprintf
"`@unique` on `%s.%s`: only scalar and Text fields can be indexed"
c.Ast.name fl.Ast.name);
[ (true, [| col_of fl.Ast.name |]) ]
end
else [])
c.fields
in
(match c.Ast.table with
| Some cfg ->
List.iter
(fun cols ->
List.iter
(fun cn ->
match List.find_opt (fun (fl : Ast.field) -> fl.Ast.name = cn) c.fields with
| None ->
index_col_err := Some (c.Ast.pos, Printf.sprintf
"`@table(index: ...)` on `%s` names `%s`, which is not a field"
c.Ast.name cn)
| Some fl ->
if not (is_indexable fl) then
index_col_err := Some (c.Ast.pos, Printf.sprintf
"`@table(index: ...)` on `%s`: `%s` is not a scalar or Text field"
c.Ast.name cn))
cols)
cfg.Ast.indexes
| None -> ());
{ cr_name = c.name; cr_gc = c.is_gc;
cr_fields =
Array.of_list
(List.filter_map
(fun (fl : Ast.field) ->
match fl.Ast.ty with
| Ast.Backlink _ -> None (* virtual: no stored column *)
| _ -> Some (fl.Ast.name, fl.Ast.ty))
c.fields);
cr_methods = List.map (fun (m : Ast.method_decl) -> m.name) c.methods;
cr_indexes = table_indexes @ unique_indexes;
cr_is_table = (c.Ast.table <> None);
cr_backlinks =
List.filter_map
(fun (fl : Ast.field) ->
match fl.Ast.ty with
| Ast.Backlink (sc, sf) -> Some (fl.Ast.name, (sc, sf))
| _ -> None)
c.fields })
:: !classes :: !classes
end end
| Ast.Union (ud : Ast.union_decl) -> | Ast.Union (ud : Ast.union_decl) ->
@ -3930,7 +4492,8 @@ let emit ~(syms : Types.symbols) ~(module_of : string -> string)
class_id := SM.add key cid !class_id; class_id := SM.add key cid !class_id;
incr nclasses; incr nclasses;
classes := classes :=
{ cr_name = key; cr_gc = false; { cr_name = key; cr_gc = false; cr_indexes = []; cr_is_table = false;
cr_backlinks = [];
cr_fields = Array.of_list vd.Ast.v_fields; cr_fields = Array.of_list vd.Ast.v_fields;
cr_methods = [] } cr_methods = [] }
:: !classes :: !classes
@ -3973,11 +4536,18 @@ let emit ~(syms : Types.symbols) ~(module_of : string -> string)
class_id := SM.add name cid !class_id; class_id := SM.add name cid !class_id;
incr nclasses; incr nclasses;
classes := classes :=
{ cr_name = name; cr_gc = false; cr_fields = Array.of_list fields; cr_methods = [] } { cr_name = name; cr_gc = false; cr_fields = Array.of_list fields; cr_methods = [];
cr_indexes = []; cr_is_table = false; cr_backlinks = [] }
:: !classes :: !classes
end) end)
Types.predeclared_records; Types.predeclared_records;
let class_id = !class_id in let class_id = !class_id in
(match !index_col_err with
| Some (pos, msg) ->
Diag.Collector.add coll
(Diag.error ~code:cannot_lower_code ~file:!index_err_file ~line:pos.Ast.line
~col:pos.Ast.col ~message:msg ())
| None -> ());
let p_classes = Array.of_list (List.rev !classes) in let p_classes = Array.of_list (List.rev !classes) in
let p_ifaces = Array.of_list (List.rev !ifaces) in let p_ifaces = Array.of_list (List.rev !ifaces) in
List.iter List.iter
@ -4149,7 +4719,16 @@ let emit ~(syms : Types.symbols) ~(module_of : string -> string)
needs when decode creates one. *) needs when decode creates one. *)
Array.iter (fun kidx -> Buf.u32 cls kidx) class_field_names.(cid); Array.iter (fun kidx -> Buf.u32 cls kidx) class_field_names.(cid);
Array.iter (fun (_, ty) -> Buf.u32 cls (field_class_meta p ty)) c.cr_fields; Array.iter (fun (_, ty) -> Buf.u32 cls (field_class_meta p ty)) c.cr_fields;
Array.iter (fun (_, ty) -> Buf.u32 cls (field_elem_meta p ty)) c.cr_fields) Array.iter (fun (_, ty) -> Buf.u32 cls (field_elem_meta p ty)) c.cr_fields;
(* v3 tail (iteration 9 Task 4): the class's secondary indexes —
index_cnt, then per index: flags (bit0 unique), col_cnt, cols *)
Buf.u32 cls (List.length c.cr_indexes);
List.iter
(fun (uniq, cols) ->
Buf.u32 cls (if uniq then 1 else 0);
Buf.u32 cls (Array.length cols);
Array.iter (fun ci -> Buf.u32 cls ci) cols)
c.cr_indexes)
p_classes; p_classes;
let ifs = Buf.create () in let ifs = Buf.create () in
Array.iteri Array.iteri

View file

@ -439,6 +439,7 @@ let oclass_of (ctx : ctx) (ft : Ast.field_ty) : oclass =
| Some u -> if u.Types.u_has_payload then Owned else Copy | Some u -> if u.Types.u_has_payload then Owned else Copy
| None -> Copy (* unknown type: WO-E225 already reported by types.ml *)) | None -> Copy (* unknown type: WO-E225 already reported by types.ml *))
| Ref _ -> Copy | Ref _ -> Copy
| Backlink _ -> Copy (* a virtual collection of row ids read on demand *)
| Multi _ | Map _ -> Owned | Multi _ | Map _ -> Owned
| Nullable _ -> Copy (* unreachable: unwrapped above *) | Nullable _ -> Copy (* unreachable: unwrapped above *)
@ -560,6 +561,9 @@ let rec expr_ty (ctx : ctx) (e : Ast.expr) : Ast.field_ty option =
| Binary (Concat, _, _) -> Some (Scalar "Text") | Binary (Concat, _, _) -> Some (Scalar "Text")
| Binary _ -> None (* arithmetic/comparison: Copy either way *) | Binary _ -> None (* arithmetic/comparison: Copy either way *)
| Ctor (cn, _) -> Some (Scalar cn) | Ctor (cn, _) -> Some (Scalar cn)
| Insert _ -> Some (Scalar "Int") (* the new row's id — Copy, nothing to drop *)
| Query _ -> Some (Multi "Int") (* a query yields a fresh multi of ids — owned *)
| Delete _ -> Some (Scalar "Int") (* the deleted id — Copy *)
| Interp _ -> Some (Scalar "Text") (* an interpolation always produces Text *) | Interp _ -> Some (Scalar "Text") (* an interpolation always produces Text *)
| DbStub _ -> None | DbStub _ -> None
| Switch (subject, arms) -> | Switch (subject, arms) ->
@ -1148,6 +1152,14 @@ let rec read_expr (ctx : ctx) (e : Ast.expr) : unit =
read_place_parts ctx e read_place_parts ctx e
| Call (callee, args) -> analyze_call ctx e callee args | Call (callee, args) -> analyze_call ctx e callee args
| Ctor (cn, fields) -> analyze_ctor ctx cn fields | Ctor (cn, fields) -> analyze_ctor ctx cn fields
| Insert (_, fields) ->
(* iteration 9 Task 3: the engine copies every field value at the row
API (the two-worlds bulkhead), so an insert BORROWS its values —
no transfer, no E304, the source keeps what it had. Trap-capable
(unique violations arrive with Task 4), so the drop map is
recorded exactly like DbStub's. *)
List.iter (fun (_, fe) -> read_expr ctx fe) fields;
record_drop ctx ~node:e.id ~pos:e.pos ~kind:DLiveMask ~items:(mask_items (live_holders ctx))
| Unary (_, o) -> read_expr ctx o | Unary (_, o) -> read_expr ctx o
| Binary (_, a, b) -> | Binary (_, a, b) ->
read_expr ctx a; read_expr ctx a;
@ -1170,6 +1182,20 @@ let rec read_expr (ctx : ctx) (e : Ast.expr) : unit =
| DbStub _ -> | DbStub _ ->
(* trap-capable: the frame needs its drop map here *) (* trap-capable: the frame needs its drop map here *)
record_drop ctx ~node:e.id ~pos:e.pos ~kind:DLiveMask ~items:(mask_items (live_holders ctx)) record_drop ctx ~node:e.id ~pos:e.pos ~kind:DLiveMask ~items:(mask_items (live_holders ctx))
| Delete t ->
read_expr ctx t;
record_drop ctx ~node:e.id ~pos:e.pos ~kind:DLiveMask ~items:(mask_items (live_holders ctx))
| Query q ->
(* iteration 9b: the sub-expressions only READ (engine field-reads copy
out at the boundary); the query is trap-capable (engine faults), so
the frame needs its drop map here, exactly like DbStub. *)
(match q.q_src with QNav e2 -> read_expr ctx e2 | QTable _ -> ());
List.iter (read_expr ctx) q.q_wheres;
(match q.q_group with Some (_, k) -> read_expr ctx k | None -> ());
(match q.q_order with Some (k, _) -> read_expr ctx k | None -> ());
(match q.q_take with Some t -> read_expr ctx t | None -> ());
read_expr ctx q.q_select;
record_drop ctx ~node:e.id ~pos:e.pos ~kind:DLiveMask ~items:(mask_items (live_holders ctx))
| Switch (subject, arms) -> analyze_switch ctx e.id subject arms | Switch (subject, arms) -> analyze_switch ctx e.id subject arms
(* The root of a place expression is already accounted for by use_place; (* The root of a place expression is already accounted for by use_place;

View file

@ -356,6 +356,12 @@ let parse_field_ty (st : state) : Ast.field_ty =
| Token.Ident "multi" -> | Token.Ident "multi" ->
ignore (advance st); ignore (advance st);
Ast.Multi (expect_ident st "multi target type") Ast.Multi (expect_ident st "multi target type")
| Token.Ident "backlink" ->
ignore (advance st);
let cls = expect_ident st "backlink source class" in
expect st Token.Dot "'.'";
let fld = expect_ident st "backlink source field" in
Ast.Backlink (cls, fld)
| Token.Ident "map" -> | Token.Ident "map" ->
ignore (advance st); ignore (advance st);
expect st Token.Lt "'<'"; expect st Token.Lt "'<'";
@ -1009,9 +1015,121 @@ and parse_switch_expr (st : state) : Ast.expr =
done; done;
{ Ast.id; pos; kind = Ast.Switch (subject, List.rev !arms) } { Ast.id; pos; kind = Ast.Switch (subject, List.rev !arms) }
and parse_insert_expr (st : state) : Ast.expr =
(* `insert` + a constructor literal, sharing parse_ctor_literal so the
field-list grammar (trailing commas, newlines) can never drift from the
ctor's. The literal's node is unwrapped into Insert — its id is reused,
which is safe because the Ctor node itself is discarded whole. *)
let pos = peek_pos st in
ignore (advance st) (* the `insert` trigger token *);
skip_newlines st;
let lit = parse_ctor_literal st in
(match lit.Ast.kind with
| Ast.Ctor (cn, fields) -> { lit with Ast.pos; kind = Ast.Insert (cn, fields) }
| _ -> lit (* unreachable: parse_ctor_literal only builds Ctor *))
and is_query_trigger (st : state) : bool =
(* `from <ident> in` — positional, so `from` stays a usable identifier
everywhere else (same discipline as insert/select) *)
(match peek st with Token.Ident "from" -> true | _ -> false)
&& (match (tok_at st (st.pos + 1)).kind with Token.Ident _ -> true | _ -> false)
&& (tok_at st (st.pos + 2)).kind = Token.KwIn
and parse_query_expr (st : state) : Ast.expr =
let pos = peek_pos st in
let id = fresh_id st in
ignore (advance st) (* from *);
let var = expect_ident st "query range variable" in
expect st Token.KwIn "`in`";
(* source: a bare class name is a table scan; any other expression is a
navigation (`d.staff`). One token of lookahead: Ident not followed by a
`.`/`(`/`[` and sitting where a clause keyword follows is a table name. *)
let src =
match peek st with
| Token.Ident cn
when (match (tok_at st (st.pos + 1)).kind with
| Token.Dot | Token.LParen | Token.LBracket -> false
| _ -> true) ->
ignore (advance st);
Ast.QTable cn
| _ -> Ast.QNav (parse_expr_no_brace st)
in
(* clauses may sit on their own lines; skip the separating newlines when
looking for the next clause keyword (the query is one expression) *)
let clause name =
skip_newlines st;
match peek st with Token.Ident n when n = name -> true | _ -> false
in
let wheres = ref [] in
while clause "where" do
ignore (advance st);
wheres := parse_expr_no_brace st :: !wheres
done;
let group =
if clause "group" then begin
ignore (advance st);
let key_elem = parse_expr_no_brace st in
ignore key_elem (* the grouped element is the range var; `group e by k` *);
if not (clause "by") then fail st (peek_pos st) syntax_code "expected `by` in a group clause";
ignore (advance st);
let key = parse_expr_no_brace st in
if not (clause "into") then fail st (peek_pos st) syntax_code "expected `into` in a group clause";
ignore (advance st);
let gvar = expect_ident st "group variable" in
Some (gvar, key)
end
else None
in
let order =
if clause "order" then begin
ignore (advance st);
if not (clause "by") then fail st (peek_pos st) syntax_code "expected `by` after `order`";
ignore (advance st);
let key = parse_expr_no_brace st in
let desc = clause "desc" in
if desc then ignore (advance st);
Some (key, desc)
end
else None
in
(* `take` is a reserved keyword (KwTake, the param convention), not an
Ident — so match the token, not the name *)
skip_newlines st;
let take =
if peek st = Token.KwTake then (ignore (advance st); Some (parse_expr_no_brace st)) else None
in
if not (clause "select") then fail st (peek_pos st) syntax_code "a query must end in `select`";
ignore (advance st);
let sel = parse_expr st in
{
Ast.id;
pos;
kind =
Ast.Query
{
Ast.q_var = var;
q_src = src;
q_wheres = List.rev !wheres;
q_group = group;
q_order = order;
q_take = take;
q_select = sel;
q_pos = pos;
};
}
and parse_primary (st : state) : Ast.expr = and parse_primary (st : state) : Ast.expr =
match peek st with match peek st with
| _ when is_query_trigger st -> parse_query_expr st
| Token.Ident "delete" when (match (tok_at st (st.pos + 1)).kind with
| Token.Newline | Token.Semicolon | Token.Eof -> false | _ -> true) ->
let pos = peek_pos st in
let id = fresh_id st in
ignore (advance st);
let target = parse_expr st in
{ Ast.id; pos; kind = Ast.Delete target }
| k when is_select_trigger k -> parse_dbstub_expr st | k when is_select_trigger k -> parse_dbstub_expr st
| k when is_insert_trigger k -> parse_insert_expr st
| Token.KwSwitch -> parse_switch_expr st | Token.KwSwitch -> parse_switch_expr st
| Token.Int n -> | Token.Int n ->
let pos = peek_pos st in let pos = peek_pos st in
@ -1286,7 +1404,7 @@ and parse_stmt (st : state) : Ast.stmt =
| k when is_insert_trigger k -> | k when is_insert_trigger k ->
let pos = peek_pos st in let pos = peek_pos st in
let id = fresh_id st in let id = fresh_id st in
let e = parse_dbstub_expr st in let e = parse_insert_expr st in
end_of_stmt st; end_of_stmt st;
{ Ast.s_id = id; s_pos = pos; s_kind = Ast.ExprStmt e } { Ast.s_id = id; s_pos = pos; s_kind = Ast.ExprStmt e }
| Token.KwLet -> parse_let_stmt st | Token.KwLet -> parse_let_stmt st
@ -1785,6 +1903,23 @@ let rec subst_expr (consts : Ast.expr StringMap.t) (bound : StringSet.t) (e : As
{ e with Ast.kind = Ast.Binary (op, subst_expr consts bound l, subst_expr consts bound r) } { e with Ast.kind = Ast.Binary (op, subst_expr consts bound l, subst_expr consts bound r) }
| Ast.Ctor (cn, fields) -> | Ast.Ctor (cn, fields) ->
{ e with Ast.kind = Ast.Ctor (cn, List.map (fun (n, v) -> (n, subst_expr consts bound v)) fields) } { e with Ast.kind = Ast.Ctor (cn, List.map (fun (n, v) -> (n, subst_expr consts bound v)) fields) }
| Ast.Insert (cn, fields) ->
{ e with Ast.kind = Ast.Insert (cn, List.map (fun (n, v) -> (n, subst_expr consts bound v)) fields) }
| Ast.Delete t -> { e with Ast.kind = Ast.Delete (subst_expr consts bound t) }
| Ast.Query q ->
(* the range/group vars shadow consts inside the query body *)
let bound' = StringSet.add q.Ast.q_var bound in
let bound' = match q.Ast.q_group with Some (g, _) -> StringSet.add g bound' | None -> bound' in
let sub = subst_expr consts bound' in
{ e with Ast.kind = Ast.Query {
q with Ast.q_src = (match q.Ast.q_src with
| Ast.QTable cn -> Ast.QTable cn
| Ast.QNav e2 -> Ast.QNav (subst_expr consts bound e2));
q_wheres = List.map sub q.Ast.q_wheres;
q_group = (match q.Ast.q_group with Some (g, k) -> Some (g, sub k) | None -> None);
q_order = (match q.Ast.q_order with Some (k, d) -> Some (sub k, d) | None -> None);
q_take = (match q.Ast.q_take with Some t -> Some (sub t) | None -> None);
q_select = sub q.Ast.q_select } }
| Ast.Interp inner -> { e with Ast.kind = Ast.Interp (subst_expr consts bound inner) } | Ast.Interp inner -> { e with Ast.kind = Ast.Interp (subst_expr consts bound inner) }
| Ast.ListLit items -> { e with Ast.kind = Ast.ListLit (List.map (subst_expr consts bound) items) } | Ast.ListLit items -> { e with Ast.kind = Ast.ListLit (List.map (subst_expr consts bound) items) }
| Ast.MapLit | Ast.NilLit -> e | Ast.MapLit | Ast.NilLit -> e

View file

@ -306,6 +306,7 @@ let rec has_recursive_structure (cls : class_info) : bool =
| Ast.Ref name -> name = cls.name | Ast.Ref name -> name = cls.name
| Ast.Multi name -> name = cls.name (* multi Self *) | Ast.Multi name -> name = cls.name (* multi Self *)
| Ast.Map (k, v) -> k = cls.name || v = cls.name (* map<_, Self> / map<Self, _> *) | Ast.Map (k, v) -> k = cls.name || v = cls.name (* map<_, Self> / map<Self, _> *)
| Ast.Backlink _ -> false
| Ast.Nullable inner -> has_recursive_structure_type inner cls.name | Ast.Nullable inner -> has_recursive_structure_type inner cls.name
) cls.fields ) cls.fields
@ -315,6 +316,7 @@ and has_recursive_structure_type (ty : Ast.field_ty) (cls_name : string) : bool
| Ast.Ref name -> name = cls_name | Ast.Ref name -> name = cls_name
| Ast.Multi name -> name = cls_name | Ast.Multi name -> name = cls_name
| Ast.Map (k, v) -> k = cls_name || v = cls_name | Ast.Map (k, v) -> k = cls_name || v = cls_name
| Ast.Backlink _ -> false (* a computed inverse holds no owned structure *)
| Ast.Nullable inner -> has_recursive_structure_type inner cls_name | Ast.Nullable inner -> has_recursive_structure_type inner cls_name
(* @unique field -> persistent identity (plan's "When NOT to emit": a (* @unique field -> persistent identity (plan's "When NOT to emit": a
@ -363,6 +365,7 @@ let rec typ_of_field_ty (ft : field_ty) : typ =
| Ref name -> TRef name | Ref name -> TRef name
| Multi inner_name -> TMulti (TScalar inner_name) | Multi inner_name -> TMulti (TScalar inner_name)
| Map (k_name, v_name) -> TMap (TScalar k_name, TScalar v_name) | Map (k_name, v_name) -> TMap (TScalar k_name, TScalar v_name)
| Backlink (c, _) -> TMulti (TScalar c) (* reads as a collection of C *)
| Nullable inner -> TNullable (typ_of_field_ty inner) | Nullable inner -> TNullable (typ_of_field_ty inner)
(* wob_kind_of_typ: maps internal typ to .wob field kind *) (* wob_kind_of_typ: maps internal typ to .wob field kind *)
@ -412,6 +415,7 @@ let unknown_fn_code = Diag.types_prefix ^ "04"
let unsatisfied_interface_code = Diag.types_prefix ^ "05" let unsatisfied_interface_code = Diag.types_prefix ^ "05"
let incomplete_ctor_code = Diag.types_prefix ^ "06" let incomplete_ctor_code = Diag.types_prefix ^ "06"
let unknown_type_code = Diag.types_prefix ^ "07" let unknown_type_code = Diag.types_prefix ^ "07"
let query_code = Diag.types_prefix ^ "50" (* WO-E250: query surface (iteration 9b) *)
let non_exhaustive_switch_code = Diag.types_prefix ^ "08" let non_exhaustive_switch_code = Diag.types_prefix ^ "08"
let invalid_builtin_code = Diag.types_prefix ^ "09" let invalid_builtin_code = Diag.types_prefix ^ "09"
let module_not_imported_code = Diag.types_prefix ^ "10" let module_not_imported_code = Diag.types_prefix ^ "10"
@ -632,7 +636,7 @@ let rec scalar_name_of (ft : field_ty) : string option =
match ft with match ft with
| Scalar name -> Some name | Scalar name -> Some name
| Nullable inner -> scalar_name_of inner | Nullable inner -> scalar_name_of inner
| Ref _ | Multi _ | Map _ -> None | Ref _ | Multi _ | Map _ | Backlink _ -> None
(* Checked once per field declaration (not at every access/use site), so (* Checked once per field declaration (not at every access/use site), so
the diagnostic lands at the field's own declaration position and the diagnostic lands at the field's own declaration position and
@ -1165,6 +1169,11 @@ let typecheck_program ~file ~(module_of : string -> string)
needs `int_to_text` first) -- unlike the placeholders below, needs `int_to_text` first) -- unlike the placeholders below,
this is a fact, not a guess. *) this is a fact, not a guess. *)
Some (TScalar "Text") Some (TScalar "Text")
| Insert _ ->
(* the new row's id — the one thing an insert produces *)
Some (TScalar "Int")
| Query _ -> None (* a query's type is chased only by typecheck_expr *)
| Delete _ -> Some (TScalar "Int")
| Unary _ | Binary _ | DbStub _ -> | Unary _ | Binary _ | DbStub _ ->
(* Not chased: the arithmetic-ladder `Binary` ops have no reliable (* Not chased: the arithmetic-ladder `Binary` ops have no reliable
per-node type in this pass at all (see above); `Unary`/`DbStub` per-node type in this pass at all (see above); `Unary`/`DbStub`
@ -1198,7 +1207,7 @@ let typecheck_program ~file ~(module_of : string -> string)
with Not_found -> { typ = TScalar "Int"; is_nil = false }) with Not_found -> { typ = TScalar "Int"; is_nil = false })
| Field (base, field_name) -> | Field (base, field_name) ->
let base_res = typecheck_expr env cenv base in let base_res = typecheck_expr env cenv base in
(match base_res.typ with (match (match base_res.typ with TRef c -> TScalar c | other -> other) with
| TScalar class_name -> | TScalar class_name ->
(* Only a *declared* class can be checked for a missing field. (* Only a *declared* class can be checked for a missing field.
typecheck_expr falls back to `TScalar "Int"` for everything typecheck_expr falls back to `TScalar "Int"` for everything
@ -1222,9 +1231,14 @@ let typecheck_program ~file ~(module_of : string -> string)
{ typ = TScalar "Int"; is_nil = false })) { typ = TScalar "Int"; is_nil = false }))
| _ -> { typ = TScalar "Int"; is_nil = false }) | _ -> { typ = TScalar "Int"; is_nil = false })
| Index (base, idx) -> | Index (base, idx) ->
let _ = typecheck_expr env cenv base in let base_res = typecheck_expr env cenv base in
let _ = typecheck_expr env cenv idx in let _ = typecheck_expr env cenv idx in
{ typ = TScalar "Int"; is_nil = false } (* `xs[i]` yields the container's element type — a `multi C` indexed
is a C (iteration 9b: query results are indexed to pick a row) *)
(match base_res.typ with
| TMulti et -> { typ = et; is_nil = false }
| TMap (_, vt) -> { typ = vt; is_nil = false }
| _ -> { typ = TScalar "Int"; is_nil = false })
| Call (callee, args) -> | Call (callee, args) ->
List.iter (fun arg -> ignore (typecheck_expr env cenv arg)) args; List.iter (fun arg -> ignore (typecheck_expr env cenv arg)) args;
(match callee.kind with (match callee.kind with
@ -1391,7 +1405,8 @@ let typecheck_program ~file ~(module_of : string -> string)
the zero word NEW already leaves there). Everything else the zero word NEW already leaves there). Everything else
stays WO-E206, classes and records alike. *) stays WO-E206, classes and records alike. *)
let omittable (default : default_expr option) (fty : field_ty) : bool = let omittable (default : default_expr option) (fty : field_ty) : bool =
Option.is_some default || (match fty with Nullable _ -> true | _ -> false) Option.is_some default
|| (match fty with Nullable _ | Backlink _ -> true | _ -> false)
in in
List.iter (fun (fname, fty, fdefault, _) -> List.iter (fun (fname, fty, fdefault, _) ->
if not (List.mem fname provided) && not (omittable fdefault fty) then if not (List.mem fname provided) && not (omittable fdefault fty) then
@ -1405,6 +1420,100 @@ let typecheck_program ~file ~(module_of : string -> string)
(Diag.error ~code:unknown_type_code ~file ~line:e.pos.line ~col:e.pos.col (Diag.error ~code:unknown_type_code ~file ~line:e.pos.line ~col:e.pos.col
~message:(Printf.sprintf "unknown type `%s` in constructor" class_name) ()); ~message:(Printf.sprintf "unknown type `%s` in constructor" class_name) ());
{ typ = TScalar "Int"; is_nil = false }) { typ = TScalar "Int"; is_nil = false })
| Insert (class_name, fields) ->
(* iteration 9 Task 3: typed exactly like a constructor literal —
same missing-field rule (defaults and `?` fields omittable),
same unknown-class diagnostic — but the VALUE is the new row's
id. The engine copies every field at the choke point, so field
values keep their owners (owner.ml's stores_by_copy). *)
(try
let cls = StringMap.find class_name syms.classes in
let provided = List.map (fun (n, _) -> n) fields in
let omittable (default : default_expr option) (fty : field_ty) : bool =
Option.is_some default
|| (match fty with Nullable _ | Backlink _ -> true | _ -> false)
in
List.iter (fun (fname, fty, fdefault, _) ->
if not (List.mem fname provided) && not (omittable fdefault fty) then
Diag.Collector.add collector
(Diag.error ~code:incomplete_ctor_code ~file ~line:e.pos.line ~col:e.pos.col
~message:(Printf.sprintf "missing field `%s` in insert of `%s`" fname class_name) ())
) cls.fields;
{ typ = TScalar "Int"; is_nil = false }
with Not_found ->
Diag.Collector.add collector
(Diag.error ~code:unknown_type_code ~file ~line:e.pos.line ~col:e.pos.col
~message:(Printf.sprintf "unknown type `%s` in insert" class_name) ());
{ typ = TScalar "Int"; is_nil = false })
| Delete target ->
let tr = typecheck_expr env cenv target in
(match (match tr.typ with TRef c -> TScalar c | o -> o) with
| TScalar cn when StringMap.mem cn syms.classes -> ()
| _ ->
Diag.Collector.add collector
(Diag.error ~code:query_code ~file ~line:e.pos.line ~col:e.pos.col
~message:"`delete` takes a table-row value" ()));
{ typ = TScalar "Int"; is_nil = false }
| Query q ->
(* iteration 9b slice: from/where/select over a table class. The
range variable is bound to the class type; a table-class value is
its row id at runtime but types AS the class, so `e.field` checks
against the class's fields exactly like a heap instance. group /
order / take / navigation sources are diagnosed as not-yet so the
surface is honest about its edge. *)
let elem_err () =
{ typ = TMulti (TScalar "Int"); is_nil = false }
in
(match q.q_src with
| Ast.QNav nav ->
(* `from s in d.staff`: the navigation yields `multi C`, so the
range var is a C. Reuse the QTable body by resolving C. *)
let nav_res = typecheck_expr env cenv nav in
let cn =
match nav_res.typ with
| TMulti (TScalar c) -> c
| _ -> ""
in
if not (StringMap.mem cn syms.classes) then begin
Diag.Collector.add collector
(Diag.error ~code:query_code ~file ~line:q.q_pos.line ~col:q.q_pos.col
~message:"query navigation source must be a `backlink`/`multi` of a table class" ());
elem_err ()
end
else begin
(if q.q_group <> None then
Diag.Collector.add collector
(Diag.error ~code:query_code ~file ~line:q.q_pos.line ~col:q.q_pos.col
~message:"group-by on a navigation query is not supported yet" ()));
let env' = StringMap.add q.q_var (TScalar cn) env in
let cenv' = StringMap.add q.q_var (TScalar cn) cenv in
List.iter (fun w -> ignore (typecheck_expr env' cenv' w)) q.q_wheres;
(match q.q_order with Some (k, _) -> ignore (typecheck_expr env' cenv' k) | None -> ());
(match q.q_take with Some t -> ignore (typecheck_expr env cenv t) | None -> ());
let sel = typecheck_expr env' cenv' q.q_select in
{ typ = TMulti sel.typ; is_nil = false }
end
| Ast.QTable cn ->
if not (StringMap.mem cn syms.classes) then begin
Diag.Collector.add collector
(Diag.error ~code:query_code ~file ~line:q.q_pos.line ~col:q.q_pos.col
~message:(Printf.sprintf "`from %s in %s`: `%s` is not a declared table class"
q.q_var cn cn) ());
elem_err ()
end
else begin
(if q.q_group <> None then
Diag.Collector.add collector
(Diag.error ~code:query_code ~file ~line:q.q_pos.line ~col:q.q_pos.col
~message:"group-by aggregation is not supported yet" ()));
let env' = StringMap.add q.q_var (TScalar cn) env in
let cenv' = StringMap.add q.q_var (TScalar cn) cenv in
List.iter (fun w -> ignore (typecheck_expr env' cenv' w)) q.q_wheres;
(match q.q_order with Some (k, _) -> ignore (typecheck_expr env' cenv' k) | None -> ());
(match q.q_take with Some t -> ignore (typecheck_expr env cenv t) | None -> ());
let sel = typecheck_expr env' cenv' q.q_select in
{ typ = TMulti sel.typ; is_nil = false }
end)
| DbStub _ -> { typ = TVoid; is_nil = false } | DbStub _ -> { typ = TVoid; is_nil = false }
| Switch (subject, arms) -> typecheck_switch ~want_value:true env cenv subject arms | Switch (subject, arms) -> typecheck_switch ~want_value:true env cenv subject arms
| ListLit items -> | ListLit items ->
@ -2102,7 +2211,8 @@ and walk_expr (bound : StringSet.t) (visit : StringSet.t -> expr -> unit) (e : e
| Binary (_, l, r) -> | Binary (_, l, r) ->
walk_expr bound visit l; walk_expr bound visit l;
walk_expr bound visit r walk_expr bound visit r
| Ctor (_, fields) -> List.iter (fun (_, v) -> walk_expr bound visit v) fields | Ctor (_, fields) | Insert (_, fields) ->
List.iter (fun (_, v) -> walk_expr bound visit v) fields
| Interp inner -> walk_expr bound visit inner | Interp inner -> walk_expr bound visit inner
| ListLit items -> List.iter (walk_expr bound visit) items | ListLit items -> List.iter (walk_expr bound visit) items
| MapLit | NilLit -> () | MapLit | NilLit -> ()
@ -2111,6 +2221,16 @@ and walk_expr (bound : StringSet.t) (visit : StringSet.t -> expr -> unit) (e : e
walk_expr bound visit body; walk_expr bound visit body;
walk_block (StringSet.add ename bound) visit handler walk_block (StringSet.add ename bound) visit handler
| DbStub _ -> () | DbStub _ -> ()
| Delete t -> walk_expr bound visit t
| Query q ->
(match q.q_src with QNav e -> walk_expr bound visit e | QTable _ -> ());
let b = StringSet.add q.q_var bound in
let b = match q.q_group with Some (g, _) -> StringSet.add g b | None -> b in
List.iter (walk_expr b visit) q.q_wheres;
(match q.q_group with Some (_, k) -> walk_expr b visit k | None -> ());
(match q.q_order with Some (k, _) -> walk_expr b visit k | None -> ());
(match q.q_take with Some t -> walk_expr b visit t | None -> ());
walk_expr b visit q.q_select
| Switch (subject, arms) -> | Switch (subject, arms) ->
walk_expr bound visit subject; walk_expr bound visit subject;
List.iter List.iter
@ -2370,6 +2490,7 @@ let rec field_ty_str (ft : field_ty) : string =
| Ref s -> "ref " ^ s | Ref s -> "ref " ^ s
| Multi s -> "multi " ^ s | Multi s -> "multi " ^ s
| Map (k, v) -> "map<" ^ k ^ ", " ^ v ^ ">" | Map (k, v) -> "map<" ^ k ^ ", " ^ v ^ ">"
| Backlink (c, f) -> "backlink " ^ c ^ "." ^ f
| Nullable t -> "?" ^ field_ty_str t | Nullable t -> "?" ^ field_ty_str t
let dump_symbols (syms : symbols) : string = let dump_symbols (syms : symbols) : string =

View file

@ -1,6 +1,6 @@
1:1 METHOD sync() 1:1 METHOD sync()
2:3 DB_STUB IDENT(insert) IDENT(Product) LBRACE IDENT(sku) COLON STR(A1) COMMA IDENT(price) COLON INT(10) RBRACE 2:3 EXPR INSERT Product { sku: "A1", price: 10 }
3:3 DB_STUB IDENT(select) IDENT(Product) LBRACE IDENT(sku) EQEQ STR(A1) RBRACE 3:3 DB_STUB IDENT(select) IDENT(Product) LBRACE IDENT(sku) EQEQ STR(A1) RBRACE
4:3 LET rows = DB_STUB(IDENT(select) IDENT(Product) LBRACE IDENT(price) GT INT(5) RBRACE) 4:3 LET rows = DB_STUB(IDENT(select) IDENT(Product) LBRACE IDENT(price) GT INT(5) RBRACE)
5:3 DB_STUB KW_INSERT IDENT(Product) LBRACE IDENT(sku) COLON STR(A2) RBRACE 5:3 EXPR INSERT Product { sku: "A2" }
6:3 DB_STUB KW_SELECT IDENT(Product) LBRACE IDENT(sku) EQEQ STR(A2) RBRACE 6:3 DB_STUB KW_SELECT IDENT(Product) LBRACE IDENT(sku) EQEQ STR(A2) RBRACE

View file

@ -38,7 +38,7 @@ fn pick(take a: Item, take b: Item, flag: Bool) -> Int {
} }
fn store(take r: Item) -> Int { fn store(take r: Item) -> Int {
insert into rows values (1) insert Row { n: 1 }
return 0 return 0
} }
@ -63,3 +63,7 @@ fn reinit_after_move(take a: Item) -> Int {
a = Item { n: 7 } a = Item { n: 7 }
return 0 return 0
} }
class Row {
n: Int
}

View file

@ -531,35 +531,35 @@ let () =
| _ -> check "ctor literal: exactly one free fn" false | _ -> check "ctor literal: exactly one free fn" false
let () = let () =
(* The brief's stated asymmetry: `insert` is a statement-only trigger (* Iteration 9 Task 3 retired the old asymmetry: `insert` is grammar-owned
(parser.ml's is_insert_trigger, checked only in parse_stmt) — a in BOTH positions now — a typed Insert node validated like a ctor,
bare `insert` reached from parse_primary is just an ordinary returning the id — while `select` stays the opaque DbStub until
identifier reference, exactly like self/me/on/service/policy's own Task 5. The old contract ("bare insert is a plain Ident") is gone
"recognized positionally, not a reserved word" rule (this task's with the stub that motivated it. *)
own keyword-discipline note). `select` (is_select_trigger) is
checked unconditionally *inside* parse_primary, so the same
position always builds a DbStub instead. Neither is an error on
its own — the difference shows up in which Ast.expr_kind comes
back. *)
let prog, collector = let prog, collector =
parse_str ~file:"insert-vs-select.wo" "fn f() {\n let a = insert\n let b = select\n}\n" parse_str ~file:"insert-vs-select.wo"
"fn f() {\n let a = insert Product { sku: \"A1\" }\n let b = select\n}\n"
in in
check_eq "insert vs. select as bare expressions: no diagnostics" ~expected:0 check_eq "typed insert + stub select: no diagnostics" ~expected:0
~actual:(List.length (Diag.Collector.diagnostics collector)) ~actual:(List.length (Diag.Collector.diagnostics collector))
string_of_int; string_of_int;
match prog.Ast.decls with (match prog.Ast.decls with
| [ Ast.Fn m ] -> ( | [ Ast.Fn m ] -> (
match m.body with match m.body with
| [ | [
{ Ast.s_kind = Ast.Let { name = "a"; value = a_val; _ }; _ }; { Ast.s_kind = Ast.Let { name = "a"; value = a_val; _ }; _ };
{ Ast.s_kind = Ast.Let { name = "b"; value = b_val; _ }; _ }; { Ast.s_kind = Ast.Let { name = "b"; value = b_val; _ }; _ };
] -> ] ->
check "bare `insert` in expression position is a plain Ident" check "`insert` in expression position is a typed Insert node"
(match a_val.Ast.kind with Ast.Ident "insert" -> true | _ -> false); (match a_val.Ast.kind with Ast.Insert ("Product", [ ("sku", _) ]) -> true | _ -> false);
check "bare `select` in expression position always becomes a DbStub" check "bare `select` in expression position always becomes a DbStub"
(match b_val.Ast.kind with Ast.DbStub _ -> true | _ -> false) (match b_val.Ast.kind with Ast.DbStub _ -> true | _ -> false)
| _ -> check "insert vs. select: exactly two `let` statements" false) | _ -> check "insert vs. select: exactly two `let` statements" false)
| _ -> check "insert vs. select: exactly one free fn" false | _ -> check "insert vs. select: exactly one free fn" false);
(* and a bare `insert` with no literal is a parse error now, not an Ident *)
let _, c2 = parse_str ~file:"bare-insert.wo" "fn f() {\n let a = insert\n}\n" in
check "bare `insert` with no constructor literal is a diagnostic"
(List.length (Diag.Collector.diagnostics c2) > 0)
let () = let () =
(* The no_brace guard (parser.ml's state.no_brace / looks_like_ctor): (* The no_brace guard (parser.ml's state.no_brace / looks_like_ctor):
@ -2382,7 +2382,7 @@ let validate_image (img : string) : string list =
let u64 o = if ok 8 o then String.get_int64_le img o else 0L in let u64 o = if ok 8 o then String.get_int64_le img o else 0L in
let none = 0xFFFFFFFF in let none = 0xFFFFFFFF in
if u32 0 <> 0x31424F57 then fail "bad magic"; if u32 0 <> 0x31424F57 then fail "bad magic";
if u32 4 <> 2 then fail "unsupported version"; if u32 4 <> 3 then fail "unsupported version";
let coff = u32 8 and ccnt = u32 12 in let coff = u32 8 and ccnt = u32 12 in
let koff = u32 16 and kcnt = u32 20 in let koff = u32 16 and kcnt = u32 20 in
let ioff = u32 24 and icnt = u32 28 in let ioff = u32 24 and icnt = u32 28 in
@ -2419,6 +2419,7 @@ let validate_image (img : string) : string list =
if flags land lnot 0x01 <> 0 then fail (Printf.sprintf "class %d: unknown flags" i); if flags land lnot 0x01 <> 0 then fail (Printf.sprintf "class %d: unknown flags" i);
if fcnt > 65535 then fail (Printf.sprintf "class %d: too many fields" i); if fcnt > 65535 then fail (Printf.sprintf "class %d: too many fields" i);
class_fields.(i) <- fcnt; class_fields.(i) <- fcnt;
let kco = !o in (* the kind bytes' offset: the v3 index walk re-reads them *)
for j = 0 to fcnt - 1 do for j = 0 to fcnt - 1 do
if u8 (!o + j) > 5 then fail (Printf.sprintf "class %d field %d: bad kind" i j) if u8 (!o + j) > 5 then fail (Printf.sprintf "class %d field %d: bad kind" i j)
done; done;
@ -2435,6 +2436,25 @@ let validate_image (img : string) : string list =
fail (Printf.sprintf "class %d field %d: field class out of range" i j) fail (Printf.sprintf "class %d field %d: field class out of range" i j)
done; done;
o := !o + (fcnt * 12); o := !o + (fcnt * 12);
(* v3: the index tail — flags (bit0 only), col_cnt 1..8, columns in
range and scalar/Text-kinded. Mirrors loader.c's checks. *)
let icnt_x = u32 !o in
o := !o + 4;
if icnt_x > 64 then fail (Printf.sprintf "class %d: too many indexes" i);
for x = 0 to icnt_x - 1 do
let ifl = u32 !o and ccnt = u32 (!o + 4) in
o := !o + 8;
if ifl land lnot 1 <> 0 then fail (Printf.sprintf "class %d index %d: unknown flags" i x);
if ccnt = 0 || ccnt > 8 then fail (Printf.sprintf "class %d index %d: bad column count" i x);
for c = 0 to ccnt - 1 do
let col = u32 !o in
o := !o + 4;
if col >= fcnt then fail (Printf.sprintf "class %d index %d: column out of range" i x);
let kind = u8 (kco + col) in
if kind <> 0 && kind <> 3 then
fail (Printf.sprintf "class %d index %d: column %d is not scalar or Text" i x c)
done
done;
if !o > len then fail (Printf.sprintf "class %d: truncated" i) if !o > len then fail (Printf.sprintf "class %d: truncated" i)
done; done;
(* interfaces + vtable rows *) (* interfaces + vtable rows *)
@ -2580,7 +2600,12 @@ let validate_image (img : string) : string list =
| 22 | 23 | 24 | 25 | 26 | 27 | 28 -> rchk pc a | 22 | 23 | 24 | 25 | 26 | 27 | 28 -> rchk pc a
| 29 -> | 29 ->
rchk pc a; rchk pc a;
if c > 12 then fail (Printf.sprintf "method %d pc %d: builtin out of range" i pc) (* the mirror's ceiling tracks wob.h's WO_B_MAX only for ids the
golden lowering suite actually emits; 61 = DB_INSERT (arity 1:
the class-id slot — field slots are runtime-validated, same as
the C loader) *)
if c > 12 && (c < 61 || c > 67) then
fail (Printf.sprintf "method %d pc %d: builtin out of range" i pc)
else if c = 4 then begin else if c = 4 then begin
if b > 5 then fail (Printf.sprintf "method %d pc %d: bad element kind" i pc) if b > 5 then fail (Printf.sprintf "method %d pc %d: bad element kind" i pc)
end end
@ -2595,6 +2620,13 @@ let validate_image (img : string) : string list =
| 1 | 2 | 3 | 7 | 8 -> 1 | 1 | 2 | 3 | 7 | 8 -> 1
| 5 | 6 | 11 | 12 -> 2 | 5 | 6 | 11 | 12 -> 2
| 10 -> 3 | 10 -> 3
| 61 -> 1
| 62 -> 4
| 63 -> 2
| 64 -> 1
| 65 -> 3
| 66 -> 3
| 67 -> 2
| _ -> 0 | _ -> 0
in in
if arity > 0 then begin if arity > 0 then begin
@ -2951,7 +2983,8 @@ let () =
( "text: concat, equality, words", ( "text: concat, equality, words",
"fn f(a: Text, b: Text) -> Int {\n let joined = a .. b\n\ "fn f(a: Text, b: Text) -> Int {\n let joined = a .. b\n\
\ if joined == a {\n return 1\n }\n return words(joined)\n}\n" ); \ if joined == a {\n return 1\n }\n return words(joined)\n}\n" );
("db stub statement", "fn f() -> Int {\n insert into rows values (1)\n return 0\n}\n"); ( "db insert statement",
"class Row {\n n: Int\n}\n\nfn f() -> Int {\n insert Row { n: 1 }\n return 0\n}\n" );
( "nested calls in arguments", ( "nested calls in arguments",
"fn one() -> Int {\n return 1\n}\n\nfn add(a: Int, b: Int) -> Int {\n\ "fn one() -> Int {\n return 1\n}\n\nfn add(a: Int, b: Int) -> Int {\n\
\ return a + b\n}\n\nfn f() -> Int {\n return add(add(one(), one()), one())\n}\n" ); \ return a + b\n}\n\nfn f() -> Int {\n return add(add(one(), one()), one())\n}\n" );

View file

@ -0,0 +1,63 @@
# database/src — how the engine hangs together
The database engine is its own top-level directory, statically linked into
every `wovm` and every runtime test binary (`runtime/Makefile`'s `DBSRC`).
One binary, unchanged. Format doc: `docs/plan/oop-vm/04-db-binding.md`.
Memory-safety doctrine: the 9b design's section 6.
## table.c — rows (iteration 9, Task 1)
```
VM values ──copy──▶ row slots (engine-owned malloc) ──copy──▶ fresh VM values
wo_row_insert wo_row_read
```
- **No VM pointer ever enters a slab; no slab pointer ever leaves.** Encode
copies per kind (Texts to `db_text`, owned objects flattened recursively to
`db_rec`, containers element-wise); decode allocates fresh VM values from
the caller's `wo_rt`. The GCREF kind is refused at encode — the compiler
should have made that impossible (the GC bulkhead), the engine refuses it
anyway.
- **Rows never move.** Slabs of 256 are malloc'd and kept for the table's
life; the free-slot list recycles removed slots before any slab grows;
the id hash maps id → slot. Ids are never reused (per-table counter,
shard-interleaved `S+1, S+1+N, …`), which is also what makes the hash's
tombstone sentinel safe.
- **Choke points**: `wo_row_insert` / `wo_row_remove` carry the `INDEX HOOK`
comments where Task 4's secondary indexes attach and Task 2's WAL stages
its record. Nothing else may mutate storage.
- One deliberate file-static: `g_classes` for recursive frees (`db_val_free`
has no context parameter). One process, one class table; revisit at
iteration 8 (shards share the same immutable table).
## wal.c — durability (iteration 9, Task 2)
The commit order IS the module: RAM apply → stage → one pwrite + one
fdatasync → ack. `wo_wal_commit` returning 0 is the only thing "durable"
means. Replay never touches the VM heap — payloads decode straight into
engine-owned values and re-enter through the row API, so whatever hooks the
choke points (indexes, Task 4) applies to replayed rows identically. Torn
tails end the intact prefix and get overwritten by the next commit;
CRC-valid-but-undecodable records fail replay loudly (corruption is not a
tear). The crash battery in `runtime/test/test_wal.c` is the module's
meaning proven: acked-over-a-pipe after commit, SIGKILL mid-stream, replay,
zero acked-but-missing.
## db.c — statement executors (iteration 9, Task 3)
One dispatcher, the builtin contract (0 ok, else WO_T_* + msg). The engine
handles ride `wo_rt.db` / `wo_rt.wal` as opaque pointers set by main.c —
NULL db traps WO_T_DB, NULL wal means RAM-only (the corpus's mode; WO_DATA
opts into durability). Insert's contract: RAM apply through the row API,
then stage + commit BEFORE returning — the builtin's return is the
acknowledgment, so a failed commit un-applies the row and traps WO_T_IO
rather than acknowledging what disk never got.
## Verifying a change
- `make -C runtime test` — `test_table` is this directory's suite (round
trips across kinds, nil encodings, shard interleave, slab growth, slot
reuse, misuse), ASan+UBSan like every runtime test.
- `just oop-e2e`, `just log-watcher` — regression that linking the engine
into wovm changed nothing observable (it is dead code until Task 3 wires
the first builtin).

163
database/src/db.c Normal file
View file

@ -0,0 +1,163 @@
#include "db.h"
#include <string.h>
#include "cont.h"
#include "table.h"
#include "wal.h"
int wo_builtin_db(wo_vm *vm, uint64_t *R, uint32_t ins, const char **msg) {
uint32_t A = wo_ins_a(ins), B = wo_ins_b(ins), C = wo_ins_c(ins);
wo_db *db = (wo_db *)vm->rt.db;
if (!db) {
*msg = "database engine not initialized";
return WO_T_DB;
}
switch (C) {
case WO_B_DB_INSERT: {
uint32_t cid = (uint32_t)R[B];
int ek = 0;
uint64_t id = wo_row_insert(db, cid, &R[B + 1], msg, &ek);
if (!id)
return ek == DB_ERR_UNIQUE ? WO_T_UNIQUE
: ek == DB_ERR_OOM ? WO_T_OOM
: WO_T_DB;
wo_wal *w = (wo_wal *)vm->rt.wal;
if (w) {
/* RAM applied, record staged, ONE commit before the ack (the
* builtin's return). A failed commit is a failed write: the
* row is removed again so RAM never claims what disk never
* acknowledged, and the statement traps. */
if (wo_wal_append_insert(w, db, cid, id) != 0 || wo_wal_commit(w) != 0) {
wo_row_remove(db, cid, id);
*msg = "wal commit failed";
return WO_T_IO;
}
}
R[A] = id;
return 0;
}
case WO_B_DB_UPDATE_FIELD: {
uint32_t cid = (uint32_t)R[B];
uint64_t id = R[B + 1];
uint32_t field = (uint32_t)R[B + 2];
int ek = 0;
if (wo_row_update_field(db, cid, id, field, R[B + 3], msg, &ek) != 0)
return ek == DB_ERR_UNIQUE ? WO_T_UNIQUE : ek == DB_ERR_OOM ? WO_T_OOM : WO_T_DB;
wo_wal *w = (wo_wal *)vm->rt.wal;
if (w) {
if (wo_wal_append_update(w, db, cid, id) != 0 || wo_wal_commit(w) != 0) {
*msg = "wal commit failed"; /* RAM ahead of disk: trap, do not ack */
return WO_T_IO;
}
}
R[A] = 0;
return 0;
}
case WO_B_DB_DELETE: {
uint32_t cid = (uint32_t)R[B];
uint64_t id = R[B + 1];
/* FK restrict: refuse if another row still references this one
(iteration 9b) — nothing is removed, the statement traps */
if (wo_row_has_referrers(db, cid, id)) {
*msg = "row is still referenced (restrict)";
return WO_T_FK;
}
if (wo_row_remove(db, cid, id) != 0) {
*msg = "no such row";
return WO_T_DB;
}
wo_wal *w = (wo_wal *)vm->rt.wal;
if (w) {
if (wo_wal_append_remove(w, cid, id) != 0 || wo_wal_commit(w) != 0) {
*msg = "wal commit failed";
return WO_T_IO;
}
}
R[A] = 0;
return 0;
}
case WO_B_DB_SCAN: {
uint32_t cid = (uint32_t)R[B];
if (cid >= db->class_cnt) {
*msg = "no such class";
return WO_T_DB;
}
wo_multi *ids = wo_multi_new(&vm->rt, WO_K_SCALAR);
if (!ids) return WO_T_OOM;
/* materialize the id list up front — the 9b cursor-stability rule:
* the loop body then point-reads each id, so a row updated mid-loop
* (even an indexed column) cannot disturb the iteration */
db_table *t = &db->tables[cid];
if (t->row_size) {
uint32_t total = t->slab_cnt * DB_SLAB_ROWS;
for (uint32_t g = 0; g < total; g++) {
if (!(t->bitmap[g >> 6] & (1ull << (g & 63)))) continue;
db_row *row =
(db_row *)(t->slabs[g / DB_SLAB_ROWS] + (size_t)(g % DB_SLAB_ROWS) * t->row_size);
if (wo_multi_push(ids, row->id) != 0) return WO_T_OOM;
}
}
R[A] = (uint64_t)(uintptr_t)ids;
return 0;
}
case WO_B_DB_GET_FIELD: {
uint32_t cid = (uint32_t)R[B];
uint64_t id = R[B + 1];
uint32_t field = (uint32_t)R[B + 2];
if (cid >= db->class_cnt || field >= db->classes[cid].field_cnt) {
*msg = "no such field";
return WO_T_DB;
}
db_row *row = wo_row_ptr(db, cid, id);
if (!row) {
*msg = "no such row";
return WO_T_DB;
}
int ok = 1;
uint64_t v = wo_val_decode_vm(db, &vm->rt, db->classes[cid].kinds[field],
row->slots[field], &ok, msg);
if (!ok) return WO_T_OOM;
R[A] = v;
return 0;
}
case WO_B_DB_PROBE: {
uint32_t cid = (uint32_t)R[B];
uint32_t index = (uint32_t)R[B + 1];
if (cid >= db->class_cnt) {
*msg = "no such class";
return WO_T_DB;
}
wo_multi *ids = wo_multi_new(&vm->rt, WO_K_SCALAR);
if (!ids) return WO_T_OOM;
db_table *t = &db->tables[cid];
if (t->row_size && index < t->index_cnt) {
db_index *ix = &t->indexes[index];
uint32_t col = ix->cols[0];
uint8_t kind = db->classes[cid].kinds[col];
uint64_t key = R[B + 2];
uint32_t total = t->slab_cnt * DB_SLAB_ROWS;
for (uint32_t g = 0; g < total; g++) {
if (!(t->bitmap[g >> 6] & (1ull << (g & 63)))) continue;
db_row *row =
(db_row *)(t->slabs[g / DB_SLAB_ROWS] + (size_t)(g % DB_SLAB_ROWS) * t->row_size);
int eq;
if (kind == WO_K_TEXT) {
const wo_str *want = (const wo_str *)(uintptr_t)key;
const db_text *have = (const db_text *)(uintptr_t)row->slots[col];
eq = (!want && !have) ||
(want && have && want->len == have->len &&
memcmp(want->data, have->bytes, have->len) == 0);
} else
eq = row->slots[col] == key;
if (eq && wo_multi_push(ids, row->id) != 0) return WO_T_OOM;
}
}
R[A] = (uint64_t)(uintptr_t)ids;
return 0;
}
default:
*msg = "unknown db builtin";
return WO_T_DB;
}
}

24
database/src/db.h Normal file
View file

@ -0,0 +1,24 @@
/* db.h — DB statement executors (iteration 9, Task 3+).
*
* The VM reaches the engine through one dispatcher with the same contract
* as every builtin family: 0 = ok, else a WO_T_* code with *msg set. The
* engine and WAL handles ride the runtime context as opaque pointers
* (obj.h's rt.db / rt.wal) — set by main.c at boot, NULL in test binaries
* that never touch DB statements (a DB builtin with rt.db == NULL traps
* WO_T_DB "engine not initialized").
*
* Commit contract per statement (until iteration 8 brings ticks): the
* insert applies to RAM, stages its WAL record, and COMMITS before the
* builtin returns — the builtin returning IS the acknowledgment, so the
* ack-after-fsync doctrine holds at statement granularity. No WAL
* (rt.wal == NULL, no WO_DATA) means RAM-only: every test and every
* corpus fixture runs that way; durability is opt-in by pointing WO_DATA
* at a directory. */
#ifndef WO_DB_H
#define WO_DB_H
#include "vm.h"
int wo_builtin_db(wo_vm *vm, uint64_t *R, uint32_t ins, const char **msg);
#endif /* WO_DB_H */

753
database/src/table.c Normal file
View file

@ -0,0 +1,753 @@
#include "table.h"
#include <stdlib.h>
#include <string.h>
#include "cont.h"
/* ---- engine-owned value encode / free / decode ------------------------- */
/* Free one encoded slot value of [kind]. Recursion mirrors encoding. */
static void db_val_free(uint8_t kind, uint64_t v);
static void db_rec_free(db_rec *r, const wo_classdesc *classes) {
const wo_classdesc *c = &classes[r->class_id];
for (uint32_t i = 0; i < c->field_cnt; i++) db_val_free(c->kinds[i], r->slots[i]);
free(r);
}
/* db_val_free needs the class table for nested records; a file-static is
* the honest signature here — one engine per process today (N=1), and the
* pointer is set once at init. Revisit when iteration 8 brings N>1 shards
* (each shard's wo_db shares the same immutable class table anyway). */
static const wo_classdesc *g_classes;
static void db_val_free(uint8_t kind, uint64_t v) {
if (!v) return;
switch (kind) {
case WO_K_SCALAR: return;
case WO_K_TEXT: free((db_text *)(uintptr_t)v); return;
case WO_K_OWNED: db_rec_free((db_rec *)(uintptr_t)v, g_classes); return;
case WO_K_MULTI: {
db_multi *m = (db_multi *)(uintptr_t)v;
for (uint32_t i = 0; i < m->len; i++) db_val_free(m->elem_kind, m->items[i]);
free(m);
return;
}
case WO_K_MAP: {
db_map *m = (db_map *)(uintptr_t)v;
for (uint32_t i = 0; i < m->len; i++) {
db_val_free(m->key_kind, m->kv[2 * i]);
db_val_free(m->val_kind, m->kv[2 * i + 1]);
}
free(m);
return;
}
default: return; /* GCREF never stored */
}
}
/* Encode one VM value into an engine-owned slot value. 0-with-*ok=0 means
* failure (OOM or a GCREF); a genuine nil encodes as 0 with *ok=1. */
static uint64_t db_val_encode(const wo_classdesc *classes, uint8_t kind, uint64_t v,
int *ok, const char **msg) {
*ok = 1;
switch (kind) {
case WO_K_SCALAR: return v;
case WO_K_TEXT: {
if (!v) return 0;
const wo_str *s = (const wo_str *)(uintptr_t)v;
db_text *t = malloc(sizeof(db_text) + s->len);
if (!t) goto oom;
t->len = s->len;
memcpy(t->bytes, s->data, s->len);
return (uint64_t)(uintptr_t)t;
}
case WO_K_OWNED: {
if (!v) return 0;
const wo_hdr *o = (const wo_hdr *)(uintptr_t)v;
const wo_classdesc *c = &classes[o->class_id];
db_rec *r = malloc(sizeof(db_rec) + (size_t)c->field_cnt * 8u);
if (!r) goto oom;
r->class_id = o->class_id;
r->_pad = 0;
const uint64_t *f = (const uint64_t *)(const void *)(o + 1);
for (uint32_t i = 0; i < c->field_cnt; i++) {
r->slots[i] = db_val_encode(classes, c->kinds[i], f[i], ok, msg);
if (!*ok) { /* free what we built so far, then fail upward */
for (uint32_t j = 0; j < i; j++) db_val_free(c->kinds[j], r->slots[j]);
free(r);
return 0;
}
}
return (uint64_t)(uintptr_t)r;
}
case WO_K_MULTI: {
if (!v) return 0;
const wo_multi *m = (const wo_multi *)(uintptr_t)v;
db_multi *d = malloc(sizeof(db_multi) + (size_t)m->len * 8u);
if (!d) goto oom;
d->elem_kind = m->elem_kind;
d->len = m->len;
for (uint32_t i = 0; i < m->len; i++) {
d->items[i] = db_val_encode(classes, m->elem_kind, m->items[i], ok, msg);
if (!*ok) {
for (uint32_t j = 0; j < i; j++) db_val_free(d->elem_kind, d->items[j]);
free(d);
return 0;
}
}
return (uint64_t)(uintptr_t)d;
}
case WO_K_MAP: {
if (!v) return 0;
const wo_map *m = (const wo_map *)(uintptr_t)v;
db_map *d = malloc(sizeof(db_map) + (size_t)m->len * 16u);
if (!d) goto oom;
d->key_kind = m->key_kind;
d->val_kind = m->val_kind;
d->len = m->len;
for (uint32_t i = 0; i < m->len; i++) {
d->kv[2 * i] = db_val_encode(classes, m->key_kind, m->keys[i], ok, msg);
uint64_t dv = 0;
if (*ok) dv = db_val_encode(classes, m->val_kind, m->vals[i], ok, msg);
d->kv[2 * i + 1] = dv;
if (!*ok) {
for (uint32_t j = 0; j <= i; j++) {
db_val_free(d->key_kind, d->kv[2 * j]);
db_val_free(d->val_kind, d->kv[2 * j + 1]);
}
free(d);
return 0;
}
}
return (uint64_t)(uintptr_t)d;
}
default:
*ok = 0;
*msg = "a garbage-collected value cannot be stored in a table field";
return 0;
}
oom:
*ok = 0;
*msg = "out of memory encoding a row";
return 0;
}
/* Decode one engine slot back into a fresh VM value (the out-gate: always
* a copy). 0-with-*ok=0 = OOM; nil decodes as 0 with *ok=1. */
static uint64_t db_val_decode(wo_rt *rt, uint8_t kind, uint64_t v, int *ok,
const char **msg) {
*ok = 1;
switch (kind) {
case WO_K_SCALAR: return v;
case WO_K_TEXT: {
if (!v) return 0;
const db_text *t = (const db_text *)(uintptr_t)v;
wo_str *s = wo_str_new(rt, t->bytes, t->len);
if (!s) goto oom;
return (uint64_t)(uintptr_t)s;
}
case WO_K_OWNED: {
if (!v) return 0;
const db_rec *r = (const db_rec *)(uintptr_t)v;
wo_hdr *o = wo_obj_new(rt, r->class_id);
if (!o) goto oom;
const wo_classdesc *c = &rt->classes[r->class_id];
uint64_t *f = wo_fields(o);
for (uint32_t i = 0; i < c->field_cnt; i++) {
f[i] = db_val_decode(rt, c->kinds[i], r->slots[i], ok, msg);
if (!*ok) return 0; /* partial object: rt teardown reclaims (test scope) */
}
return (uint64_t)(uintptr_t)o;
}
case WO_K_MULTI: {
if (!v) return 0;
const db_multi *d = (const db_multi *)(uintptr_t)v;
wo_multi *m = wo_multi_new(rt, d->elem_kind);
if (!m) goto oom;
for (uint32_t i = 0; i < d->len; i++) {
uint64_t ev = db_val_decode(rt, d->elem_kind, d->items[i], ok, msg);
if (!*ok || wo_multi_push(m, ev) != 0) goto oom;
}
return (uint64_t)(uintptr_t)m;
}
case WO_K_MAP: {
if (!v) return 0;
const db_map *d = (const db_map *)(uintptr_t)v;
wo_map *m = wo_map_new(rt, d->key_kind, d->val_kind);
if (!m) goto oom;
for (uint32_t i = 0; i < d->len; i++) {
uint64_t kv = db_val_decode(rt, d->key_kind, d->kv[2 * i], ok, msg);
uint64_t vv = 0;
if (*ok) vv = db_val_decode(rt, d->val_kind, d->kv[2 * i + 1], ok, msg);
uint64_t old;
if (!*ok || wo_map_set(m, kv, vv, &old) < 0) goto oom;
}
return (uint64_t)(uintptr_t)m;
}
default: return 0; /* GCREF never stored, so never decoded */
}
oom:
*ok = 0;
*msg = "out of memory decoding a row";
return 0;
}
/* ---- id hash (open addressing, pow2, id -> global slot + 1) ----------- */
static uint64_t hmix(uint64_t x) { /* splitmix64 finalizer */
x += 0x9e3779b97f4a7c15ull;
x = (x ^ (x >> 30)) * 0xbf58476d1ce4e5b9ull;
x = (x ^ (x >> 27)) * 0x94d049bb133111ebull;
return x ^ (x >> 31);
}
/* Ids are never 0 (0 spells "empty bucket") and never reused, so all-ones
* can never collide with a live id — it marks a deleted bucket that probes
* walk straight past. */
#define H_DELETED ((uint64_t)-1)
static int hgrow(db_table *t) {
size_t ncap = t->hcap ? t->hcap * 2 : 64;
uint64_t *nk = calloc(ncap, 8), *nv = calloc(ncap, 8);
if (!nk || !nv) {
free(nk);
free(nv);
return -1;
}
for (size_t i = 0; i < t->hcap; i++) {
if (!t->hkeys[i] || t->hkeys[i] == H_DELETED) continue;
size_t j = hmix(t->hkeys[i]) & (ncap - 1);
while (nk[j]) j = (j + 1) & (ncap - 1);
nk[j] = t->hkeys[i];
nv[j] = t->hvals[i];
}
free(t->hkeys);
free(t->hvals);
t->hkeys = nk;
t->hvals = nv;
t->hcap = ncap;
return 0;
}
static int hput(db_table *t, uint64_t id, uint64_t slot1) {
if (t->hlen * 10 >= t->hcap * 7 && hgrow(t) != 0) return -1;
size_t j = hmix(id) & (t->hcap - 1);
while (t->hkeys[j] && t->hkeys[j] != id) j = (j + 1) & (t->hcap - 1);
if (!t->hkeys[j]) t->hlen++;
t->hkeys[j] = id;
t->hvals[j] = slot1;
return 0;
}
static uint64_t hget(const db_table *t, uint64_t id) {
if (!t->hcap) return 0;
size_t j = hmix(id) & (t->hcap - 1);
while (t->hkeys[j]) {
if (t->hkeys[j] == id) return t->hvals[j];
j = (j + 1) & (t->hcap - 1);
}
return 0;
}
static void hdel(db_table *t, uint64_t id) {
if (!t->hcap) return;
size_t j = hmix(id) & (t->hcap - 1);
while (t->hkeys[j]) {
if (t->hkeys[j] == id) {
t->hkeys[j] = H_DELETED;
t->hvals[j] = 0;
return;
}
j = (j + 1) & (t->hcap - 1);
}
}
/* ---- tables and rows ---------------------------------------------------- */
/* ---- secondary indexes (Task 4) ---------------------------------------- */
/* hash of one row's index columns: kind-driven, never trusted for equality */
static uint64_t idx_hash(const wo_classdesc *c, const db_index *ix, const db_row *r) {
uint64_t h = 0x9e3779b97f4a7c15ull;
for (uint32_t i = 0; i < ix->col_cnt; i++) {
uint32_t col = ix->cols[i];
uint64_t v = r->slots[col];
if (c->kinds[col] == WO_K_TEXT) {
const db_text *t = (const db_text *)(uintptr_t)v;
uint64_t th = 1469598103934665603ull; /* FNV-1a over bytes; nil = 0 */
if (t)
for (uint32_t b = 0; b < t->len; b++) th = (th ^ (uint8_t)t->bytes[b]) * 1099511628211ull;
else th = 0;
v = th;
}
h ^= hmix(v + i);
}
return h ? h : 1; /* 0 marks an empty bucket */
}
static int idx_cols_equal(const wo_classdesc *c, const db_index *ix, const db_row *a,
const db_row *b) {
for (uint32_t i = 0; i < ix->col_cnt; i++) {
uint32_t col = ix->cols[i];
if (c->kinds[col] == WO_K_TEXT) {
const db_text *x = (const db_text *)(uintptr_t)a->slots[col];
const db_text *y = (const db_text *)(uintptr_t)b->slots[col];
if (!x || !y) {
if (x != y) return 0;
} else if (x->len != y->len || memcmp(x->bytes, y->bytes, x->len) != 0)
return 0;
} else if (a->slots[col] != b->slots[col])
return 0;
}
return 1;
}
static db_ibucket *idx_bucket(db_index *ix, uint64_t h, int create) {
if (ix->bcap == 0) {
if (!create) return NULL;
ix->buckets = calloc(64, sizeof(db_ibucket));
if (!ix->buckets) return NULL;
ix->bcap = 64;
}
if (create && ix->blen * 10 >= ix->bcap * 7) {
size_t ncap = ix->bcap * 2;
db_ibucket *nb = calloc(ncap, sizeof(db_ibucket));
if (!nb) return NULL;
for (size_t i = 0; i < ix->bcap; i++) {
if (!ix->buckets[i].hash) continue;
size_t j = ix->buckets[i].hash & (ncap - 1);
while (nb[j].hash) j = (j + 1) & (ncap - 1);
nb[j] = ix->buckets[i];
}
free(ix->buckets);
ix->buckets = nb;
ix->bcap = ncap;
}
size_t j = h & (ix->bcap - 1);
while (ix->buckets[j].hash) {
if (ix->buckets[j].hash == h) return &ix->buckets[j];
j = (j + 1) & (ix->bcap - 1);
}
if (!create) return NULL;
ix->buckets[j].hash = h;
ix->blen++;
return &ix->buckets[j];
}
/* Add [r] to every index; unique violation reports which without mutating
* anything (checks run before any add). 0 ok, DB_ERR_* otherwise. */
static int idx_add_row(wo_db *db, db_table *t, db_row *r) {
const wo_classdesc *c = &db->classes[t->class_id];
for (uint32_t x = 0; x < t->index_cnt; x++) {
db_index *ix = &t->indexes[x];
if (!(ix->flags & 1u)) continue;
db_ibucket *b = idx_bucket(ix, idx_hash(c, ix, r), 0);
if (!b) continue;
for (uint32_t i = 0; i < b->len; i++) {
db_row *other = wo_row_ptr(db, t->class_id, b->ids[i]);
if (other && idx_cols_equal(c, ix, r, other)) return DB_ERR_UNIQUE;
}
}
for (uint32_t x = 0; x < t->index_cnt; x++) {
db_index *ix = &t->indexes[x];
db_ibucket *b = idx_bucket(ix, idx_hash(c, ix, r), 1);
if (!b) return DB_ERR_OOM;
if (b->len == b->cap) {
uint32_t ncap = b->cap ? b->cap * 2 : 4;
uint64_t *ni = realloc(b->ids, (size_t)ncap * 8u);
if (!ni) return DB_ERR_OOM;
b->ids = ni;
b->cap = ncap;
}
b->ids[b->len++] = r->id;
}
return 0;
}
static void idx_remove_row(wo_db *db, db_table *t, db_row *r) {
const wo_classdesc *c = &db->classes[t->class_id];
for (uint32_t x = 0; x < t->index_cnt; x++) {
db_index *ix = &t->indexes[x];
db_ibucket *b = idx_bucket(ix, idx_hash(c, ix, r), 0);
if (!b) continue;
for (uint32_t i = 0; i < b->len; i++)
if (b->ids[i] == r->id) {
b->ids[i] = b->ids[--b->len];
break;
}
}
}
int wo_db_init(wo_db *db, const wo_classdesc *classes, uint32_t class_cnt,
uint32_t shard, uint32_t nshards) {
if (!nshards || shard >= nshards) return -1;
memset(db, 0, sizeof(*db));
db->classes = classes;
db->class_cnt = class_cnt;
db->shard = shard;
db->nshards = nshards;
db->tables = calloc(class_cnt ? class_cnt : 1, sizeof(db_table));
if (!db->tables) return -1;
g_classes = classes;
return 0;
}
static void table_destroy(wo_db *db, db_table *t) {
/* free every live row's engine-owned values, then the slabs */
const wo_classdesc *c = &db->classes[t->class_id];
for (uint32_t s = 0; s < t->slab_cnt; s++) {
for (uint32_t i = 0; i < DB_SLAB_ROWS; i++) {
uint32_t g = s * DB_SLAB_ROWS + i;
if (!(t->bitmap[g >> 6] & (1ull << (g & 63)))) continue;
db_row *r = (db_row *)(t->slabs[s] + (size_t)i * t->row_size);
for (uint32_t f = 0; f < c->field_cnt; f++)
db_val_free(c->kinds[f], r->slots[f]);
}
free(t->slabs[s]);
}
free(t->slabs);
free(t->bitmap);
free(t->free_slots);
free(t->hkeys);
free(t->hvals);
for (uint32_t x = 0; x < t->index_cnt; x++) {
for (size_t b = 0; b < t->indexes[x].bcap; b++) free(t->indexes[x].buckets[b].ids);
free(t->indexes[x].buckets);
}
free(t->indexes);
}
void wo_db_destroy(wo_db *db) {
if (!db->tables) return;
for (uint32_t i = 0; i < db->class_cnt; i++)
if (db->tables[i].slab_cnt || db->tables[i].hkeys) table_destroy(db, &db->tables[i]);
free(db->tables);
db->tables = NULL;
}
static db_table *table_of(wo_db *db, uint32_t class_id) {
if (class_id >= db->class_cnt) return NULL;
db_table *t = &db->tables[class_id];
if (!t->row_size) { /* lazy init on first touch */
const wo_classdesc *c = &db->classes[class_id];
t->class_id = class_id;
t->row_size = sizeof(db_row) + (size_t)c->field_cnt * 8u;
t->next_id = db->shard + 1; /* S+1, then += N: interleaved, local-only */
if (c->idx_cnt) {
t->indexes = calloc(c->idx_cnt, sizeof(db_index));
if (!t->indexes) return NULL;
const uint32_t *im = c->idx_meta;
for (uint32_t x = 0; x < c->idx_cnt; x++) {
t->indexes[x].flags = im[0];
t->indexes[x].col_cnt = im[1];
t->indexes[x].cols = im + 2;
im += 2 + im[1];
}
t->index_cnt = c->idx_cnt;
}
}
return t;
}
static db_row *slot_row(db_table *t, uint32_t g) {
return (db_row *)(t->slabs[g / DB_SLAB_ROWS] + (size_t)(g % DB_SLAB_ROWS) * t->row_size);
}
/* Pick the slot a new row lands in: recycled first, else the next free bit,
* else grow a slab. Returns the global slot or UINT32_MAX on OOM. */
static uint32_t slot_alloc(db_table *t) {
if (t->free_cnt) return t->free_slots[--t->free_cnt];
uint32_t total = t->slab_cnt * DB_SLAB_ROWS;
for (uint32_t g = 0; g < total; g++) /* cheap at slab granularity: only
reached when free list is empty, and the bitmap scan is bounded by
one word test per 64 slots */
if (!(t->bitmap[g >> 6] & (1ull << (g & 63)))) return g;
/* grow */
if (t->slab_cnt == t->slab_cap) {
uint32_t ncap = t->slab_cap ? t->slab_cap * 2 : 4;
uint8_t **ns = realloc(t->slabs, (size_t)ncap * sizeof(uint8_t *));
if (!ns) return UINT32_MAX;
t->slabs = ns;
t->slab_cap = ncap;
}
uint8_t *slab = malloc((size_t)DB_SLAB_ROWS * t->row_size);
if (!slab) return UINT32_MAX;
size_t nwords = ((size_t)(t->slab_cnt + 1) * DB_SLAB_ROWS + 63) / 64;
uint64_t *nb = realloc(t->bitmap, nwords * 8);
if (!nb) {
free(slab);
return UINT32_MAX;
}
memset(nb + ((size_t)t->slab_cnt * DB_SLAB_ROWS) / 64, 0,
(nwords - ((size_t)t->slab_cnt * DB_SLAB_ROWS) / 64) * 8);
t->bitmap = nb;
t->slabs[t->slab_cnt] = slab;
return t->slab_cnt++ * DB_SLAB_ROWS;
}
uint64_t wo_row_insert(wo_db *db, uint32_t class_id, const uint64_t *vals,
const char **msg, int *err_kind) {
if (err_kind) *err_kind = DB_ERR_MISC;
db_table *t = table_of(db, class_id);
if (!t) {
*msg = "no such class";
return 0;
}
const wo_classdesc *c = &db->classes[class_id];
uint32_t g = slot_alloc(t);
if (g == UINT32_MAX) {
*msg = "out of memory growing a table";
return 0;
}
db_row *r = slot_row(t, g);
r->class_id = class_id;
r->flags = 0;
int ok = 1;
uint32_t i = 0;
for (; i < c->field_cnt; i++) {
r->slots[i] = db_val_encode(db->classes, c->kinds[i], vals[i], &ok, msg);
if (!ok) {
if (err_kind) *err_kind = DB_ERR_BADKIND;
break;
}
}
if (!ok) {
for (uint32_t j = 0; j < i; j++) db_val_free(c->kinds[j], r->slots[j]);
/* slot never became live: recycle it (bitmap bit was never set) */
if (t->free_cnt == t->free_cap) {
uint32_t ncap = t->free_cap ? t->free_cap * 2 : 16;
uint32_t *nf = realloc(t->free_slots, (size_t)ncap * 4);
if (nf) {
t->free_slots = nf;
t->free_cap = ncap;
}
}
if (t->free_cnt < t->free_cap) t->free_slots[t->free_cnt++] = g;
return 0;
}
r->id = t->next_id;
t->next_id += db->nshards;
if (hput(t, r->id, (uint64_t)g + 1) != 0) {
for (uint32_t j = 0; j < c->field_cnt; j++) db_val_free(c->kinds[j], r->slots[j]);
if (err_kind) *err_kind = DB_ERR_OOM;
*msg = "out of memory indexing a row";
return 0;
}
t->bitmap[g >> 6] |= 1ull << (g & 63);
t->count++;
/* THE index hook (Task 4): inside the choke point, never anywhere else.
A unique violation un-applies the row entirely — id never handed out
twice matters less than the row never having existed. */
int irc = idx_add_row(db, t, r);
if (irc != 0) {
t->bitmap[g >> 6] &= ~(1ull << (g & 63));
hdel(t, r->id);
t->count--;
t->next_id -= db->nshards; /* the id was never observable: reclaim it */
for (uint32_t j = 0; j < c->field_cnt; j++) db_val_free(c->kinds[j], r->slots[j]);
if (t->free_cnt < t->free_cap) t->free_slots[t->free_cnt++] = g;
if (err_kind) *err_kind = irc;
*msg = irc == DB_ERR_UNIQUE ? "unique index violation" : "out of memory indexing a row";
return 0;
}
if (err_kind) *err_kind = DB_ERR_NONE;
return r->id;
}
db_row *wo_row_ptr(wo_db *db, uint32_t class_id, uint64_t id) {
if (class_id >= db->class_cnt) return NULL;
db_table *t = &db->tables[class_id];
if (!t->row_size) return NULL;
uint64_t s1 = hget(t, id);
if (!s1) return NULL;
return slot_row(t, (uint32_t)(s1 - 1));
}
int wo_row_read(wo_db *db, wo_rt *rt, uint32_t class_id, uint64_t id,
uint64_t *out_vals, const char **msg) {
db_row *r = wo_row_ptr(db, class_id, id);
if (!r) return -1;
const wo_classdesc *c = &db->classes[class_id];
int ok = 1;
for (uint32_t i = 0; i < c->field_cnt; i++) {
out_vals[i] = db_val_decode(rt, c->kinds[i], r->slots[i], &ok, msg);
if (!ok) return -2;
}
return 0;
}
db_row *wo_row_create_raw(wo_db *db, uint32_t class_id, uint64_t id) {
db_table *t = table_of(db, class_id);
if (!t || !id) return NULL;
if (hget(t, id)) return NULL; /* duplicate id: corruption, not a tear */
uint32_t g = slot_alloc(t);
if (g == UINT32_MAX) return NULL;
db_row *r = slot_row(t, g);
r->id = id;
r->class_id = class_id;
r->flags = 0;
memset(r->slots, 0, t->row_size - sizeof(db_row));
if (hput(t, id, (uint64_t)g + 1) != 0) return NULL;
t->bitmap[g >> 6] |= 1ull << (g & 63);
t->count++;
/* keep the interleave: only ids this shard owns move its counter */
if ((id - 1) % db->nshards == db->shard && id >= t->next_id)
t->next_id = id + db->nshards;
/* indexes: NOT here — the slots are still zero. wal.c fills them and
then calls wo_row_raw_commit, which is where replayed rows re-index. */
return r;
}
int wo_row_raw_commit(wo_db *db, uint32_t class_id, db_row *r) {
db_table *t = &db->tables[class_id];
return idx_add_row(db, t, r) == 0 ? 0 : -1;
}
void wo_db_val_free(wo_db *db, uint8_t kind, uint64_t v) {
(void)db;
db_val_free(kind, v);
}
uint64_t wo_val_decode_vm(wo_db *db, wo_rt *rt, uint8_t kind, uint64_t engine_val,
int *ok, const char **msg) {
(void)db;
return db_val_decode(rt, kind, engine_val, ok, msg);
}
int wo_row_update_field(wo_db *db, uint32_t class_id, uint64_t id, uint32_t field,
uint64_t vm_val, const char **msg, int *err_kind) {
if (err_kind) *err_kind = DB_ERR_MISC;
db_row *r = wo_row_ptr(db, class_id, id);
if (!r) {
*msg = "no such row";
return -1;
}
const wo_classdesc *c = &db->classes[class_id];
if (field >= c->field_cnt) {
*msg = "no such field";
return -1;
}
db_table *t = &db->tables[class_id];
int ok = 1;
uint64_t nv = db_val_encode(db->classes, c->kinds[field], vm_val, &ok, msg);
if (!ok) {
if (err_kind) *err_kind = DB_ERR_BADKIND;
return -1;
}
/* indexes containing this column: unique checks against the NEW value
run first, against a shadow of the row, before anything mutates */
uint64_t old = r->slots[field];
r->slots[field] = nv;
for (uint32_t x = 0; x < t->index_cnt; x++) {
db_index *ix = &t->indexes[x];
if (!(ix->flags & 1u)) continue;
int touches = 0;
for (uint32_t i = 0; i < ix->col_cnt; i++)
if (ix->cols[i] == field) touches = 1;
if (!touches) continue;
db_ibucket *b = idx_bucket(ix, idx_hash(c, ix, r), 0);
if (!b) continue;
for (uint32_t i = 0; i < b->len; i++) {
if (b->ids[i] == id) continue;
db_row *other = wo_row_ptr(db, class_id, b->ids[i]);
if (other && idx_cols_equal(c, ix, r, other)) {
r->slots[field] = old; /* untouched, promised */
db_val_free(c->kinds[field], nv);
if (err_kind) *err_kind = DB_ERR_UNIQUE;
*msg = "unique index violation";
return -1;
}
}
}
/* commit: fix every index containing the column (old entry out under
the OLD value's hash, new entry in), then free the old value */
r->slots[field] = old;
for (uint32_t x = 0; x < t->index_cnt; x++) {
db_index *ix = &t->indexes[x];
int touches = 0;
for (uint32_t i = 0; i < ix->col_cnt; i++)
if (ix->cols[i] == field) touches = 1;
if (!touches) continue;
db_ibucket *b = idx_bucket(ix, idx_hash(c, ix, r), 0);
if (b)
for (uint32_t i = 0; i < b->len; i++)
if (b->ids[i] == id) {
b->ids[i] = b->ids[--b->len];
break;
}
}
r->slots[field] = nv;
for (uint32_t x = 0; x < t->index_cnt; x++) {
db_index *ix = &t->indexes[x];
int touches = 0;
for (uint32_t i = 0; i < ix->col_cnt; i++)
if (ix->cols[i] == field) touches = 1;
if (!touches) continue;
db_ibucket *b = idx_bucket(ix, idx_hash(c, ix, r), 1);
if (b) {
if (b->len == b->cap) {
uint32_t ncap = b->cap ? b->cap * 2 : 4;
uint64_t *ni = realloc(b->ids, (size_t)ncap * 8u);
if (ni) {
b->ids = ni;
b->cap = ncap;
}
}
if (b->len < b->cap) b->ids[b->len++] = id;
}
}
db_val_free(c->kinds[field], old);
if (err_kind) *err_kind = DB_ERR_NONE;
return 0;
}
int wo_row_has_referrers(wo_db *db, uint32_t class_id, uint64_t id) {
if (!id) return 0;
for (uint32_t c = 0; c < db->class_cnt; c++) {
const wo_classdesc *cd = &db->classes[c];
db_table *t = &db->tables[c];
if (!t->row_size || !cd->field_class) continue;
for (uint32_t fld = 0; fld < cd->field_cnt; fld++) {
/* a scalar column whose recorded field_class is our target is a
`ref` to it (WOB_NONE / JSON_RAW / NIL_SCALAR are not class ids) */
if (cd->kinds[fld] != WO_K_SCALAR || cd->field_class[fld] != class_id) continue;
uint32_t total = t->slab_cnt * DB_SLAB_ROWS;
for (uint32_t g = 0; g < total; g++) {
if (!(t->bitmap[g >> 6] & (1ull << (g & 63)))) continue;
db_row *r = (db_row *)(t->slabs[g / DB_SLAB_ROWS] +
(size_t)(g % DB_SLAB_ROWS) * t->row_size);
if (r->slots[fld] == id) return 1;
}
}
}
return 0;
}
int wo_row_remove(wo_db *db, uint32_t class_id, uint64_t id) {
if (class_id >= db->class_cnt) return -1;
db_table *t = &db->tables[class_id];
if (!t->row_size) return -1;
uint64_t s1 = hget(t, id);
if (!s1) return -1;
uint32_t g = (uint32_t)(s1 - 1);
db_row *r = slot_row(t, g);
/* the index hook's remove side: before the row's values die, while the
columns are still comparable */
idx_remove_row(db, t, r);
const wo_classdesc *c = &db->classes[class_id];
for (uint32_t i = 0; i < c->field_cnt; i++) db_val_free(c->kinds[i], r->slots[i]);
t->bitmap[g >> 6] &= ~(1ull << (g & 63));
hdel(t, id);
t->count--;
if (t->free_cnt == t->free_cap) {
uint32_t ncap = t->free_cap ? t->free_cap * 2 : 16;
uint32_t *nf = realloc(t->free_slots, (size_t)ncap * 4);
if (!nf) return 0; /* slot simply not recycled; bitmap still frees it */
t->free_slots = nf;
t->free_cap = ncap;
}
t->free_slots[t->free_cnt++] = g;
return 0;
}

196
database/src/table.h Normal file
View file

@ -0,0 +1,196 @@
/* table.h — class-shaped row storage (iteration 9, Task 1).
*
* The engine and the VM heap are two memory worlds crossed only by copy
* (the 9b design's section 6): a row stores NO VM pointer. Every field
* lands in one 8-byte slot, kind-driven:
*
* SCALAR the 8 bytes themselves (WO_NIL_SCALAR spells a ?scalar's nil)
* TEXT engine-owned db_text* (0 = nil)
* OWNED engine-owned db_rec* — the object flattened by value,
* recursively, through these same rules (0 = nil)
* MULTI engine-owned db_multi* — elements encoded element-wise
* MAP engine-owned db_map* — keys and values encoded pair-wise
* GCREF never stored: the compiler rejects it (the GC bulkhead);
* the engine refuses it defensively as an encode error
*
* `ref T` is a SCALAR at this layer — the target row's id, an ordinary
* number the compiler produced; the engine learns nothing about it until
* the FK checks (9b plan, Task 3).
*
* Row layout: a 16-byte header (id, class, flags) then field_cnt 8-byte
* slots — deliberately the VM object layout's shape, so encode/decode walk
* the same class-table kinds the VM walks. Rows live in per-class SLABS
* (fixed-count, malloc'd, never moved: a row's address is stable for its
* lifetime, which is what lets 9b hand out loop-scoped row views). A
* per-table bitmap tracks occupancy; removed slots go on a free list and
* are reused before any slab grows. The id->row map is an open-addressing
* hash owned by the table.
*
* Id discipline (the c-runtime plan's shipped behavior): per table, per
* shard, ids interleave — shard S of N allocates S+1, S+1+N, S+1+2N, … —
* so creation is coordination-free and a row's owner shard is (id-1) % N.
* Milestone runs at N=1 (iteration 8 not yet landed); everything here is
* N-parametric and degenerates cleanly.
*
* CHOKE POINT DOCTRINE: wo_row_insert / wo_row_remove are the only paths
* that touch storage. Task 4's secondary indexes hook exactly these two
* functions; anything else mutating a slab is a defect by definition.
*/
#ifndef WO_TABLE_H
#define WO_TABLE_H
#include "obj.h" /* wo_rt, wo_classdesc, kinds, wo_str, containers */
/* ---- engine-owned value shapes (all malloc'd, all reachable only from
* row slots, all freed through db_val_free) ---- */
typedef struct db_text {
uint32_t len;
char bytes[]; /* len bytes, no NUL */
} db_text;
typedef struct db_rec { /* an owned object flattened by value */
uint32_t class_id; /* index into the SAME class table the VM uses */
uint32_t _pad;
uint64_t slots[]; /* field_cnt slots, encoded by these rules */
} db_rec;
typedef struct db_multi {
uint8_t elem_kind;
uint32_t len;
uint64_t items[];
} db_multi;
typedef struct db_map {
uint8_t key_kind, val_kind;
uint32_t len;
uint64_t kv[]; /* len pairs: k0 v0 k1 v1 … */
} db_map;
/* ---- rows and tables ---- */
typedef struct db_row {
uint64_t id;
uint32_t class_id;
uint32_t flags; /* reserved (0) */
uint64_t slots[];
} db_row;
#define DB_SLAB_ROWS 256u
/* Secondary index (iteration 9, Task 4): built from the class table's v3
* metadata at first touch, maintained ONLY inside the row choke points.
* Hash multimap: bucket per column-value hash, ids within; equality is
* re-checked against the actual rows on the unique path (a hash is a hint,
* never an answer). */
typedef struct db_ibucket {
uint64_t hash;
uint64_t *ids;
uint32_t len, cap;
} db_ibucket;
typedef struct db_index {
uint32_t flags; /* bit0 = unique */
uint32_t col_cnt;
const uint32_t *cols; /* into the loader's idx pool */
db_ibucket *buckets; /* open addressing by hash; hash==0 stored as 1 */
size_t bcap, blen;
} db_index;
/* wo_row_insert failure classes — *msg carries the sentence, this carries
* the machine-readable kind so db.c maps to the right trap. */
enum { DB_ERR_NONE = 0, DB_ERR_OOM = 1, DB_ERR_BADKIND = 2, DB_ERR_UNIQUE = 3, DB_ERR_MISC = 4 };
typedef struct db_table {
uint32_t class_id;
size_t row_size; /* 16 + field_cnt * 8 */
/* slabs of DB_SLAB_ROWS rows each; addresses stable forever */
uint8_t **slabs;
uint32_t slab_cnt, slab_cap;
uint64_t *bitmap; /* one bit per slot, slab-major */
/* removed slots, reused LIFO before any slab grows */
uint32_t *free_slots;
uint32_t free_cnt, free_cap;
uint64_t next_id; /* next id THIS shard hands out for this table */
uint64_t count; /* live rows */
/* id -> (global slot + 1); 0 = empty. Open addressing, pow2. */
uint64_t *hkeys;
uint64_t *hvals;
size_t hcap, hlen;
/* secondary indexes, from the class table's v3 metadata */
db_index *indexes;
uint32_t index_cnt;
} db_table;
typedef struct wo_db {
const wo_classdesc *classes;
uint32_t class_cnt;
uint32_t shard, nshards; /* S of N; ids interleave S+1, S+1+N, … */
db_table *tables; /* class_cnt entries, created lazily on first insert */
} wo_db;
/* 0 ok, -1 alloc failure. nshards >= 1, shard < nshards. */
int wo_db_init(wo_db *db, const wo_classdesc *classes, uint32_t class_cnt,
uint32_t shard, uint32_t nshards);
void wo_db_destroy(wo_db *db);
/* Insert: encode field_cnt VM values (register words, kinds from the class
* table) into a fresh row. Returns the new id, or 0 with *msg set (OOM, or
* a GCREF field — which the compiler should have refused upstream). */
uint64_t wo_row_insert(wo_db *db, uint32_t class_id, const uint64_t *vals,
const char **msg, int *err_kind);
/* Read: decode the row's fields into VM values freshly allocated from
* [rt] — always copies, never a pointer into the slab (the out-gate).
* 0 ok, -1 no such row, -2 OOM (*msg set). */
int wo_row_read(wo_db *db, wo_rt *rt, uint32_t class_id, uint64_t id,
uint64_t *out_vals, const char **msg);
/* Remove: free the row's engine-owned field values, clear the slot, recycle
* it. 0 ok, -1 no such row. */
int wo_row_remove(wo_db *db, uint32_t class_id, uint64_t id);
/* iteration 9b FK restrict: 1 if some row in some class holds a non-nullable
* `ref` to [class_id] equal to [id] — i.e. deleting this row would dangle a
* reference. The compiler records a ref field's target class in the class
* table's field_class metadata; this scans those columns. Correctness-first
* (a full scan of referencing tables); the backlink index is the later
* optimization the spec records. */
int wo_row_has_referrers(wo_db *db, uint32_t class_id, uint64_t id);
/* Update one field in place (iteration 9 Task 5): encode the VM value,
* swap it into the slot, keep every index containing that column honest —
* remove-old/add-new with the unique re-check running BEFORE anything
* mutates, so a violating update leaves the row untouched. 0 ok, -1 no
* such row / bad field, DB_ERR_* codes via *err_kind like insert. */
int wo_row_update_field(wo_db *db, uint32_t class_id, uint64_t id, uint32_t field,
uint64_t vm_val, const char **msg, int *err_kind);
/* Borrowed row pointer for engine-internal callers (the WAL writes a row's
* encoded bytes; indexes read key slots). NULL = no such row. NEVER handed
* to the VM. */
db_row *wo_row_ptr(wo_db *db, uint32_t class_id, uint64_t id);
/* Engine-internal, for WAL replay only: create a row with a FIXED id,
* slots zeroed — the caller (wal.c) fills them with engine-encoded values
* it built while decoding. Advances the table's next_id past [id] when the
* id belongs to this shard, so post-replay inserts never collide. NULL =
* OOM or duplicate id (corruption beyond a torn tail). */
db_row *wo_row_create_raw(wo_db *db, uint32_t class_id, uint64_t id);
/* Engine-internal: free one engine-encoded slot value of [kind] (wal.c's
* decode error paths). */
void wo_db_val_free(wo_db *db, uint8_t kind, uint64_t v);
/* Decode one engine slot value to a FRESH VM value in [rt] (the out-gate:
* always a copy). The query builtins' field reads go through this. */
uint64_t wo_val_decode_vm(wo_db *db, wo_rt *rt, uint8_t kind, uint64_t engine_val,
int *ok, const char **msg);
/* Engine-internal, replay only: after wal.c fills a raw row's slots, this
* runs the index maintenance the normal insert runs inline — including the
* unique check, whose violation during replay is corruption, not data
* (0 ok, -1). */
int wo_row_raw_commit(wo_db *db, uint32_t class_id, db_row *r);
#endif /* WO_TABLE_H */

474
database/src/wal.c Normal file
View file

@ -0,0 +1,474 @@
/* pread/pwrite/fdatasync/posix_fallocate under -std=c11 */
#define _POSIX_C_SOURCE 200809L
#include "wal.h"
#include <errno.h>
#include <fcntl.h>
#include <stdlib.h>
#include <string.h>
#include <unistd.h>
/* ---- crc32 (poly 0xEDB88320) — ported from runtime/wo-rt.c ------------- */
static uint32_t crc_table[256];
static int crc_ready;
static void crc32_init(void) {
for (uint32_t i = 0; i < 256; i++) {
uint32_t c = i;
for (int k = 0; k < 8; k++) c = (c & 1) ? 0xEDB88320u ^ (c >> 1) : c >> 1;
crc_table[i] = c;
}
crc_ready = 1;
}
static uint32_t crc32(const void *buf, size_t len) {
if (!crc_ready) crc32_init();
const uint8_t *p = buf;
uint32_t c = 0xFFFFFFFFu;
while (len--) c = crc_table[(c ^ *p++) & 0xFF] ^ (c >> 8);
return c ^ 0xFFFFFFFFu;
}
/* ---- byte buffer -------------------------------------------------------- */
typedef struct {
uint8_t *b;
size_t len, cap;
int oom;
} wbuf;
static void wput(wbuf *w, const void *p, size_t n) {
if (w->oom) return;
if (w->len + n > w->cap) {
size_t nc = w->cap ? w->cap * 2 : 256;
while (nc < w->len + n) nc *= 2;
uint8_t *nb = realloc(w->b, nc);
if (!nb) {
w->oom = 1;
return;
}
w->b = nb;
w->cap = nc;
}
memcpy(w->b + w->len, p, n);
w->len += n;
}
static void wput_u8(wbuf *w, uint8_t v) { wput(w, &v, 1); }
static void wput_u32(wbuf *w, uint32_t v) { wput(w, &v, 4); }
static void wput_u64(wbuf *w, uint64_t v) { wput(w, &v, 8); }
/* bounds-checked reader */
typedef struct {
const uint8_t *p, *end;
int bad;
} rbuf;
static int rtake(rbuf *r, void *out, size_t n) {
if (r->bad || (size_t)(r->end - r->p) < n) {
r->bad = 1;
return -1;
}
memcpy(out, r->p, n);
r->p += n;
return 0;
}
static uint8_t rd_u8(rbuf *r) {
uint8_t v = 0;
rtake(r, &v, 1);
return v;
}
static uint32_t rd_u32(rbuf *r) {
uint32_t v = 0;
rtake(r, &v, 4);
return v;
}
static uint64_t rd_u64(rbuf *r) {
uint64_t v = 0;
rtake(r, &v, 8);
return v;
}
#define WAL_NIL_TEXT 0xFFFFFFFFu
/* ---- engine-value <-> bytes (kind-driven, mirrors table.c's encoding) --- */
static void enc_val(wbuf *w, const wo_classdesc *classes, uint8_t kind, uint64_t v) {
switch (kind) {
case WO_K_SCALAR: wput_u64(w, v); return;
case WO_K_TEXT: {
if (!v) {
wput_u32(w, WAL_NIL_TEXT);
return;
}
const db_text *t = (const db_text *)(uintptr_t)v;
wput_u32(w, t->len);
wput(w, t->bytes, t->len);
return;
}
case WO_K_OWNED: {
if (!v) {
wput_u8(w, 0);
return;
}
const db_rec *r = (const db_rec *)(uintptr_t)v;
wput_u8(w, 1);
wput_u32(w, r->class_id);
const wo_classdesc *c = &classes[r->class_id];
for (uint32_t i = 0; i < c->field_cnt; i++)
enc_val(w, classes, c->kinds[i], r->slots[i]);
return;
}
case WO_K_MULTI: {
if (!v) {
wput_u8(w, 0);
return;
}
const db_multi *m = (const db_multi *)(uintptr_t)v;
wput_u8(w, 1);
wput_u8(w, m->elem_kind);
wput_u32(w, m->len);
for (uint32_t i = 0; i < m->len; i++) enc_val(w, classes, m->elem_kind, m->items[i]);
return;
}
case WO_K_MAP: {
if (!v) {
wput_u8(w, 0);
return;
}
const db_map *m = (const db_map *)(uintptr_t)v;
wput_u8(w, 1);
wput_u8(w, m->key_kind);
wput_u8(w, m->val_kind);
wput_u32(w, m->len);
for (uint32_t i = 0; i < m->len; i++) {
enc_val(w, classes, m->key_kind, m->kv[2 * i]);
enc_val(w, classes, m->val_kind, m->kv[2 * i + 1]);
}
return;
}
default: return; /* GCREF never stored, so never logged */
}
}
/* Decode one value into an engine-owned allocation. Returns 0 on success
* with *out set (0 = genuine nil); -1 on truncation/corruption/OOM — the
* caller frees what it already built. */
static int dec_val(rbuf *r, wo_db *db, uint8_t kind, uint64_t *out) {
*out = 0;
switch (kind) {
case WO_K_SCALAR: {
uint64_t v = rd_u64(r);
if (r->bad) return -1;
*out = v;
return 0;
}
case WO_K_TEXT: {
uint32_t len = rd_u32(r);
if (r->bad) return -1;
if (len == WAL_NIL_TEXT) return 0;
if ((size_t)(r->end - r->p) < len) return -1;
db_text *t = malloc(sizeof(db_text) + len);
if (!t) return -1;
t->len = len;
memcpy(t->bytes, r->p, len);
r->p += len;
*out = (uint64_t)(uintptr_t)t;
return 0;
}
case WO_K_OWNED: {
uint8_t tag = rd_u8(r);
if (r->bad) return -1;
if (!tag) return 0;
uint32_t cid = rd_u32(r);
if (r->bad || cid >= db->class_cnt) return -1;
const wo_classdesc *c = &db->classes[cid];
db_rec *rec = malloc(sizeof(db_rec) + (size_t)c->field_cnt * 8u);
if (!rec) return -1;
rec->class_id = cid;
rec->_pad = 0;
for (uint32_t i = 0; i < c->field_cnt; i++) {
if (dec_val(r, db, c->kinds[i], &rec->slots[i]) != 0) {
for (uint32_t j = 0; j < i; j++) wo_db_val_free(db, c->kinds[j], rec->slots[j]);
free(rec);
return -1;
}
}
*out = (uint64_t)(uintptr_t)rec;
return 0;
}
case WO_K_MULTI: {
uint8_t tag = rd_u8(r);
if (r->bad) return -1;
if (!tag) return 0;
uint8_t ek = rd_u8(r);
uint32_t len = rd_u32(r);
if (r->bad || ek > WO_K_MAX) return -1;
if (len > (size_t)(r->end - r->p)) return -1; /* each elem >= 1 byte */
db_multi *m = malloc(sizeof(db_multi) + (size_t)len * 8u);
if (!m) return -1;
m->elem_kind = ek;
m->len = len;
for (uint32_t i = 0; i < len; i++) {
if (dec_val(r, db, ek, &m->items[i]) != 0) {
for (uint32_t j = 0; j < i; j++) wo_db_val_free(db, ek, m->items[j]);
free(m);
return -1;
}
}
*out = (uint64_t)(uintptr_t)m;
return 0;
}
case WO_K_MAP: {
uint8_t tag = rd_u8(r);
if (r->bad) return -1;
if (!tag) return 0;
uint8_t kk = rd_u8(r), vk = rd_u8(r);
uint32_t len = rd_u32(r);
if (r->bad || kk > WO_K_MAX || vk > WO_K_MAX) return -1;
if (len > (size_t)(r->end - r->p)) return -1;
db_map *m = malloc(sizeof(db_map) + (size_t)len * 16u);
if (!m) return -1;
m->key_kind = kk;
m->val_kind = vk;
m->len = len;
for (uint32_t i = 0; i < len; i++) {
if (dec_val(r, db, kk, &m->kv[2 * i]) != 0 ||
dec_val(r, db, vk, &m->kv[2 * i + 1]) != 0) {
m->len = i; /* free only the fully-built pairs plus a possible key */
for (uint32_t j = 0; j < i; j++) {
wo_db_val_free(db, kk, m->kv[2 * j]);
wo_db_val_free(db, vk, m->kv[2 * j + 1]);
}
wo_db_val_free(db, kk, m->kv[2 * i]); /* 0 if the key failed */
free(m);
return -1;
}
}
*out = (uint64_t)(uintptr_t)m;
return 0;
}
default: return -1;
}
}
/* ---- record scan (shared by open, replay, check) ------------------------ */
/* Read the record at [off]. 0 = intact (*len_out = payload length, payload
* malloc'd into *payload_out if non-NULL); 1 = end of intact prefix (zero
* length, short read, bad crc, missing mark). */
static int scan_record(int fd, uint64_t off, uint32_t *len_out, uint8_t **payload_out) {
uint8_t hdr[8];
ssize_t n = pread(fd, hdr, 8, (off_t)off);
if (n != 8) return 1;
uint32_t len, crc;
memcpy(&len, hdr, 4);
memcpy(&crc, hdr + 4, 4);
if (len == 0 || len > (64u << 20)) return 1; /* preallocated tail or garbage */
uint8_t *payload = malloc(len + 4);
if (!payload) return 1;
n = pread(fd, payload, len + 4, (off_t)(off + 8));
if (n != (ssize_t)(len + 4)) {
free(payload);
return 1;
}
uint32_t mark;
memcpy(&mark, payload + len, 4);
if (mark != WO_WAL_MARK || crc32(payload, len) != crc) {
free(payload);
return 1;
}
*len_out = len;
if (payload_out) *payload_out = payload;
else free(payload);
return 0;
}
/* ---- public API ---------------------------------------------------------- */
int wo_wal_open(wo_wal *w, const char *path, uint64_t prealloc) {
memset(w, 0, sizeof(*w));
w->fd = open(path, O_RDWR | O_CREAT, 0644);
if (w->fd < 0) return -1;
if (prealloc) {
/* best-effort: a filesystem without fallocate still works */
(void)posix_fallocate(w->fd, 0, (off_t)prealloc);
}
/* position after the intact prefix: a torn tail is OVERWRITTEN by the
* next append, never appended after */
uint64_t off = 0;
uint32_t len;
while (scan_record(w->fd, off, &len, NULL) == 0) off += 8u + len + 4u;
w->off = off;
return 0;
}
void wo_wal_close(wo_wal *w) {
if (w->fd >= 0) close(w->fd);
free(w->buf);
memset(w, 0, sizeof(*w));
w->fd = -1;
}
/* frame one payload into the staged batch */
static int stage(wo_wal *w, const wbuf *payload) {
if (payload->oom) return -1;
wbuf rec = {0};
wput_u32(&rec, (uint32_t)payload->len);
wput_u32(&rec, crc32(payload->b, payload->len));
wput(&rec, payload->b, payload->len);
wput_u32(&rec, WO_WAL_MARK);
if (rec.oom) {
free(rec.b);
return -1;
}
if (w->len + rec.len > w->cap) {
size_t nc = w->cap ? w->cap * 2 : 4096;
while (nc < w->len + rec.len) nc *= 2;
uint8_t *nb = realloc(w->buf, nc);
if (!nb) {
free(rec.b);
return -1;
}
w->buf = nb;
w->cap = nc;
}
memcpy(w->buf + w->len, rec.b, rec.len);
w->len += rec.len;
free(rec.b);
return 0;
}
int wo_wal_append_insert(wo_wal *w, wo_db *db, uint32_t class_id, uint64_t id) {
db_row *r = wo_row_ptr(db, class_id, id);
if (!r) return -1; /* commit order: RAM apply comes FIRST */
wbuf p = {0};
wput_u8(&p, WO_WAL_INSERT);
wput_u32(&p, class_id);
wput_u64(&p, id);
const wo_classdesc *c = &db->classes[class_id];
for (uint32_t i = 0; i < c->field_cnt; i++) enc_val(&p, db->classes, c->kinds[i], r->slots[i]);
int rc = stage(w, &p);
free(p.b);
return rc;
}
int wo_wal_append_update(wo_wal *w, wo_db *db, uint32_t class_id, uint64_t id) {
db_row *r = wo_row_ptr(db, class_id, id);
if (!r) return -1;
wbuf p = {0};
wput_u8(&p, WO_WAL_UPDATE);
wput_u32(&p, class_id);
wput_u64(&p, id);
const wo_classdesc *c = &db->classes[class_id];
for (uint32_t i = 0; i < c->field_cnt; i++) enc_val(&p, db->classes, c->kinds[i], r->slots[i]);
int rc = stage(w, &p);
free(p.b);
return rc;
}
int wo_wal_append_remove(wo_wal *w, uint32_t class_id, uint64_t id) {
wbuf p = {0};
wput_u8(&p, WO_WAL_REMOVE);
wput_u32(&p, class_id);
wput_u64(&p, id);
int rc = stage(w, &p);
free(p.b);
return rc;
}
int wo_wal_commit(wo_wal *w) {
if (!w->len) return 0;
size_t at = 0;
while (at < w->len) {
ssize_t n = pwrite(w->fd, w->buf + at, w->len - at, (off_t)(w->off + at));
if (n < 0) {
if (errno == EINTR) continue;
return -1;
}
at += (size_t)n;
}
if (fdatasync(w->fd) != 0) return -1;
w->off += w->len;
w->len = 0; /* acked: the batch is durable */
return 0;
}
static int apply_record(wo_db *db, const uint8_t *payload, uint32_t len) {
rbuf r = {payload, payload + len, 0};
uint8_t kind = rd_u8(&r);
uint32_t cid = rd_u32(&r);
uint64_t id = rd_u64(&r);
if (r.bad || cid >= db->class_cnt) return -1;
if (kind == WO_WAL_REMOVE) return wo_row_remove(db, cid, id);
if (kind != WO_WAL_INSERT && kind != WO_WAL_UPDATE) return -1;
if (kind == WO_WAL_UPDATE) {
/* replace: the row must exist (its insert precedes its update in a
correct log); anything else is corruption */
if (wo_row_remove(db, cid, id) != 0) return -1;
}
db_row *row = wo_row_create_raw(db, cid, id);
if (!row) return -1;
const wo_classdesc *c = &db->classes[cid];
for (uint32_t i = 0; i < c->field_cnt; i++) {
if (dec_val(&r, db, c->kinds[i], &row->slots[i]) != 0) {
/* a record that CRC-passed but does not decode is corruption,
* not a tear: fail loudly (the row's built slots are freed by
* wo_row_remove, which also unregisters the id) */
wo_row_remove(db, cid, id);
return -1;
}
}
if ((size_t)(r.end - r.p) != 0) { /* trailing bytes = corrupt */
wo_row_remove(db, cid, id);
return -1;
}
/* slots are real now: re-index (Task 4). A unique violation during
* replay is corruption — the live insert would have refused it. */
if (wo_row_raw_commit(db, cid, row) != 0) {
wo_row_remove(db, cid, id);
return -1;
}
return 0;
}
int64_t wo_wal_replay(const char *path, wo_db *db) {
int fd = open(path, O_RDONLY);
if (fd < 0) return errno == ENOENT ? 0 : -1; /* no WAL yet = fresh boot */
uint64_t off = 0;
int64_t applied = 0;
for (;;) {
uint32_t len;
uint8_t *payload;
if (scan_record(fd, off, &len, &payload) != 0) break; /* intact prefix ends */
int rc = apply_record(db, payload, len);
free(payload);
if (rc != 0) {
close(fd);
return -1;
}
off += 8u + len + 4u;
applied++;
}
close(fd);
return applied;
}
int64_t wo_wal_check(const char *path, uint64_t *intact_bytes) {
int fd = open(path, O_RDONLY);
if (fd < 0) return -1;
uint64_t off = 0;
int64_t records = 0;
for (;;) {
uint32_t len;
if (scan_record(fd, off, &len, NULL) != 0) break;
off += 8u + len + 4u;
records++;
}
if (intact_bytes) *intact_bytes = off;
close(fd);
return records;
}

91
database/src/wal.h Normal file
View file

@ -0,0 +1,91 @@
/* wal.h — typed-row write-ahead log + boot replay (iteration 9, Task 2).
*
* The c-runtime plan's shipped pattern (phases D/E), generalized to typed
* rows. The commit order is doctrine, verbatim:
*
* RAM apply → wal_append (staged) → wal_commit (write + fdatasync)
* → only then is the write ACKNOWLEDGED
*
* Record framing — replay-whole-or-not-at-all:
*
* record := len u32 | crc u32 | payload | mark u32
* len = payload byte count (never 0; 0 = preallocated tail, stop)
* crc = CRC32 of payload
* mark = 0x574F4C31 "WOL1" — written LAST, so a record without its
* mark is torn by definition
* payload := kind u8 | class_id u32 | row_id u64 | body
* kind : 1 insert (body = the row's fields, engine encoding below)
* 2 remove (no body)
* 3 update (reserved for Task 5)
*
* Field encoding in a body walks the class table's kinds:
* SCALAR 8 bytes
* TEXT u32 len | bytes (0xFFFFFFFF = nil)
* OWNED u8 0 = nil, or u8 1 | u32 class_id | fields recursively
* MULTI u8 0 = nil, or u8 1 | u8 elem_kind | u32 len | elements
* MAP u8 0 = nil, or u8 1 | u8 kk | u8 vk | u32 len | k v pairs
*
* Replay decodes payloads STRAIGHT into engine-owned values — the VM heap
* is never involved (boot must not depend on a VM existing yet), and rows
* re-enter through the same choke-point row API, so Task 4's indexes are
* rebuilt for free. A torn tail (short record, bad CRC, missing mark) drops
* everything from the tear onward — never a partial record, never a record
* after a tear. Little-endian on-disk, matching the .wob loader's platform
* note.
*
* wo_wal_check is the offline oracle the crash battery verifies with: it
* walks a WAL file with no engine at all and reports how many records are
* intact and where the intact prefix ends. */
#ifndef WO_WAL_H
#define WO_WAL_H
#include "table.h"
#define WO_WAL_MARK 0x574F4C31u /* "WOL1" LE */
enum { WO_WAL_INSERT = 1, WO_WAL_REMOVE = 2, WO_WAL_UPDATE = 3 };
typedef struct wo_wal {
int fd;
uint64_t off; /* next write offset (the intact tail) */
/* staged batch: appended by wal_append_*, flushed by wal_commit */
uint8_t *buf;
size_t len, cap;
} wo_wal;
/* Open (create if missing) and preallocate [prealloc] bytes (best-effort;
* a filesystem without fallocate still works). Positions the write offset
* at the end of the INTACT record prefix — an existing file is scanned the
* same way replay scans it, so a torn tail is overwritten, not appended
* after. 0 ok, -1 errno-style failure. */
int wo_wal_open(wo_wal *w, const char *path, uint64_t prealloc);
void wo_wal_close(wo_wal *w);
/* Stage a record for the row that MUST already be applied to RAM (the
* commit-order doctrine). Insert/update read the row via wo_row_ptr.
* 0 ok, -1 OOM / no such row. */
int wo_wal_append_insert(wo_wal *w, wo_db *db, uint32_t class_id, uint64_t id);
int wo_wal_append_remove(wo_wal *w, uint32_t class_id, uint64_t id);
/* UPDATE re-logs the whole row (KISS: replay replaces — remove + re-create
* with the same id; the prefix/suffix delta trick from the survey is a
* later optimization, recorded). Call AFTER the RAM update. */
int wo_wal_append_update(wo_wal *w, wo_db *db, uint32_t class_id, uint64_t id);
/* Write the staged batch and fdatasync — the ack line. Empty batch = ok,
* no syscall. 0 ok, -1 write/sync failure (the batch stays staged). */
int wo_wal_commit(wo_wal *w);
/* Boot replay: apply every intact record to [db] in order. Ids re-enter
* exactly as logged; each table's next_id advances past the replayed ids
* that belong to this shard. Returns the number of records applied, or -1
* on open failure / a record naming an unknown class (corruption beyond
* what a torn tail explains). A torn tail is NOT an error: replay applies
* the intact prefix and reports it. */
int64_t wo_wal_replay(const char *path, wo_db *db);
/* Offline verification (no engine): scan [path], count intact records.
* *intact_bytes (optional) = where the intact prefix ends. -1 = open
* failure. */
int64_t wo_wal_check(const char *path, uint64_t *intact_bytes);
#endif /* WO_WAL_H */

View file

@ -115,11 +115,18 @@ that sequences its tasks. Read one, approve, then the next starts.
| 7 | [log-watcher proof](stories/language-runtime-database/07-logwatcher-proof.md) | 🔄 **runs; executable in progress** | | 7 | [log-watcher proof](stories/language-runtime-database/07-logwatcher-proof.md) | 🔄 **runs; executable in progress** |
| 7b | [Inferred GC + mark-sweep](stories/language-runtime-database/07b-inferred-gc-mark-sweep.md) | ⏸ off the workload's path (no `@gc`) | | 7b | [Inferred GC + mark-sweep](stories/language-runtime-database/07b-inferred-gc-mark-sweep.md) | ⏸ off the workload's path (no `@gc`) |
| 8 | [Shard-actor runtime](stories/language-runtime-database/08-shard-actor-runtime.md) | ⬜ | | 8 | [Shard-actor runtime](stories/language-runtime-database/08-shard-actor-runtime.md) | ⬜ |
| 9 | [Database engine](stories/language-runtime-database/09-database-engine.md) | ⬜ | | 9 | [Database engine](stories/language-runtime-database/09-database-engine.md) | 🔄 engine complete (storage/WAL/indexes/insert-update-delete); reads land with 9b |
| 9b | [`@table`, relations, query](stories/language-runtime-database/09b-table-relations-query.md) | ⬜ needs a spec first | | 9b | [`@table`, relations, query](stories/language-runtime-database/09b-table-relations-query.md) | 🔄 query surface + relations + FK done (branch query-surface); group-by parked |
| 9c | [Cross-program tables](stories/language-runtime-database/09c-cross-program-tables.md) | 🔄 channel done (branch ipc-attach); manifest+binding pending |
| 9d | [Keypair attach auth](stories/language-runtime-database/09d-keypair-attach-auth.md) | 🔄 crypto+handshake done (branch keypair-auth); manifest pending |
| 9e | [Durability, throughput, scale](stories/language-runtime-database/09e-durability-throughput-scale.md) | ⬜ needs a spec first |
| 9f | [io_uring group-commit](stories/language-runtime-database/09f-io-uring-commit.md) | ⬜ after 8 + 9e |
| 9g | [Query grammar corpus](stories/language-runtime-database/09g-query-grammar-corpus.md) | ⬜ needs a spec first |
| 10 | [HTTP service layer](stories/language-runtime-database/10-http-service.md) | ⬜ | Hold | | 10 | [HTTP service layer](stories/language-runtime-database/10-http-service.md) | ⬜ | Hold |
| 11 | [Fibers](stories/language-runtime-database/11-fibers.md) | ⬜ | Hold | | 11 | [Fibers](stories/language-runtime-database/11-fibers.md) | ⬜ | Hold |
| 12 | [Blue-green deploy](stories/language-runtime-database/12-blue-green-deploy.md) | ⬜ | Hold | | 12 | [Blue-green deploy](stories/language-runtime-database/12-blue-green-deploy.md) | ⬜ | Hold |
| 13 | [Compile-time metaprogramming](stories/language-runtime-database/13-compile-time-metaprogramming.md) | ⬜ needs a spec first |
| 14 | [skillhost host workload](stories/language-runtime-database/14-skillhost-host-workload.md) | ⬜ gaps recorded (branch query-grammar found skillhost needs no new query grammar); each gap a candidate iteration |
--- ---
@ -322,7 +329,13 @@ Ecommerce sample (verified 2026-06-13): `api.rest` 17/17 expected statuses pass.
| 7b | Inferred GC + incremental mark-sweep — `@gc` removed, GC-ness inferred, RC retired | [spec](superpowers/specs/2026-08-11-inferred-gc-mark-sweep-design.md) — plan to be written | | 7b | Inferred GC + incremental mark-sweep — `@gc` removed, GC-ness inferred, RC retired | [spec](superpowers/specs/2026-08-11-inferred-gc-mark-sweep-design.md) — plan to be written |
| 8 | Shard-actor runtime | [plan 4](superpowers/plans/2026-08-01-shard-actor-vm-runtime.md) | | 8 | Shard-actor runtime | [plan 4](superpowers/plans/2026-08-01-shard-actor-vm-runtime.md) |
| 9 | Database engine binding | [plan 5](superpowers/plans/2026-08-01-db-engine-binding.md) | | 9 | Database engine binding | [plan 5](superpowers/plans/2026-08-01-db-engine-binding.md) |
| 9b | `@table` + relations + language-integrated query | **no spec yet** — three open forks recorded in the iteration; brainstorm before planning | | 9b | `@table` + relations + language-integrated query — comprehension queries, `ref`/`backlink` navigation, GroupBy aggregates; acceptance: new `docs/examples/employee` sample | [spec](superpowers/specs/2026-08-15-table-relations-query-design.md) · [plan](plan/compiler/2026-08-15-employee-relations-query.md) |
| 9c | Cross-program tables — attach to a running program's database (IPC string in wo.toml, manifest-granted rights, owner stays the single writer) | **no spec yet** — four open forks recorded in the iteration; brainstorm before planning |
| 9d | Keypair attach auth — mutual challenge–response, grants name public keys, uid superseded | **no spec yet** — four forks recorded; plan folds into 9c's |
| 9e | Durability + throughput + scale — restart-persistence, read/write benchmark, ~1M rows; the gate every later optimization re-runs | **no spec yet** — four forks recorded; the measurement backbone |
| 9f | io_uring group-commit write path — batched durability overlapped on shard threads, fsync fallback | **no spec yet** — brainstorm after iterations 8 + 9e |
| 9g | Query grammar from real embedded-DB corpora — whole-query count + correlated exists, driven by the skillhost SQL catalogue; add only what a corpus uses | **no spec yet** — three forks; may collapse to "confirm len(query) + add exists" |
| 14 | skillhost host workload — port skillhost (MCP host + confined script runner) to writeonce; drives the missing host capabilities into the open (bounded subprocess, stdin/stdout transport, fs metadata, FFI-vs-out-of-process) | **no spec yet** — gaps recorded in the iteration; each gap brainstormed on demand, bounded-subprocess first |
| 10 | HTTP service layer | [plan 6](superpowers/plans/2026-08-01-http-service-layer.md) | | 10 | HTTP service layer | [plan 6](superpowers/plans/2026-08-01-http-service-layer.md) |
| 11 | Fibers | vision §3, [blue-green exploration](plan/exploration/blue-green-vm/00-vision.md) | | 11 | Fibers | vision §3, [blue-green exploration](plan/exploration/blue-green-vm/00-vision.md) |
| 12 | Blue-green deploy | [spec](superpowers/specs/2026-08-03-blue-green-vm-design.md) — plan authored after iterations 9–10 | | 12 | Blue-green deploy | [spec](superpowers/specs/2026-08-03-blue-green-vm-design.md) — plan authored after iterations 9–10 |
@ -361,13 +374,14 @@ log-watcher proof.
| ⬜ | 15a–15e MCP over streamable HTTP | [15](plan/15-mcp-streamable-http.md) | 15e needs 13c + 09d | | ⬜ | 15a–15e MCP over streamable HTTP | [15](plan/15-mcp-streamable-http.md) | 15e needs 13c + 09d |
| ⬜ | 16c–16f typed columns, lossless resync, restore, SCRAM | [16](plan/16-postgres-mirror.md) | | | ⬜ | 16c–16f typed columns, lossless resync, restore, SCRAM | [16](plan/16-postgres-mirror.md) | |
### Frontend — parked ### Frontend — removed as stale (2026-08-17)
| Status | Phase | Doc | The `##ui` / `.htmlx` LiveView frontend track — 13d pricing UI, the 14-MVC-UI
| ------ | -------------------------------- | ------------------------------------------------------------------- | implementation plan, the 7-of-7 `ui-htmlx-live` plan, and the 9-doc
| ⏸ | 13d pricing UI | [13](plan/13-class-model-live-pricing.md) | `plan/exploration/ui/` design set — was **removed**. It was built entirely on
| ⏸ | 14 MVC UI implementation (14a–f) | [14](plan/14-mvc-ui-implementation.md) | the non-advancing Rust runtime (`.dev/reference/crates/wo-htmlx`, `cargo run`,
| ⏸ | UI exploration track | [exploration/ui/00-overview.md](plan/exploration/ui/00-overview.md) | WebSocket live-patches) and contradicts the current woc/wovm direction. Recorded
in [`discarded.md`](plan/discarded.md).
--- ---

View file

@ -1,6 +1,6 @@
# Problem Statement # Problem Statement
The current writeonce architecture works, but it carries weight that the project doesn't need. This document identifies the structural problems that motivate the redesign described in [02-recovery.md](./02-recovery.md). The current writeonce architecture works, but it carries weight that the project doesn't need. This document identifies the structural problems that motivate the redesign described in 02-recovery.md.
## Too Many Moving Parts ## Too Many Moving Parts

View file

@ -1,172 +0,0 @@
# Recovery — The Target Architecture
This document describes where writeonce is going: a single, self-contained binary that owns its own storage, serves its own content, and pushes updates to connected clients in real-time — with no external database, no cloud pipeline, and no separate API server.
## Guiding Principle
**Everything in one process.** The database, the server logic, and the client-facing interface all live in a single codebase and ship as a single executable. If you can run the binary, you have the full platform.
## Own Database
The current PostgreSQL instance is a derived cache — it stores JSONB copies of files that already exist as the source of truth. The recovery architecture eliminates this indirection entirely.
### What Changes
- **No external database.** No PostgreSQL, no Diesel ORM, no connection pooling, no migrations.
- **Local file storage.** Markdown files and JSON metadata files are stored in a local directory, just as they are today in `writeonce-articles-s3/`. The file system *is* the database.
- **Custom storage segments (.seg files).** Research area: segment files that provide efficient read access, indexing, and potentially append-only writes for content. Think of these as a lightweight, purpose-built storage layer — not a general-purpose database engine, but enough to support indexed lookups by `blog-title` and ordered listing by date.
- **Indexed by blog-title.** The `sys_title` / blog-title field remains the primary key for content retrieval. The embedded storage must support O(1) or O(log n) lookups by this field.
### What Stays the Same
- Articles are still structured as JSON metadata + Markdown content pairs.
- The `sys_title`, `published`, `tags`, `author`, and section structure remain the content model.
- Content is still the source of truth — but now it's read directly from local storage instead of being derived through a sync pipeline.
## No AWS Infrastructure
The current architecture uses S3 as a file host and Lambda as a sync trigger. In the target architecture, there is nothing to sync *to* — the files are already where they need to be.
### What Gets Removed
| Current Component | Why It Existed | Why It's No Longer Needed |
|---|---|---|
| S3 bucket | Remote file storage | Files live locally alongside the binary |
| Lambda function (Go) | Watch S3 for changes, call API | No remote store to watch — file changes are local |
| aws-infra service (Rust) | Bridge to AWS S3/EC2 APIs | No AWS dependency |
| Pulumi IaC | Manage Lambda + S3 resources | No cloud resources to manage |
### What Replaces It
The binary watches its own content directory. When a file changes (new article, updated metadata), the embedded database re-indexes and notifies subscribers. The deployment model becomes:
```
1. Place the binary on a server
2. Point it at a content directory
3. It serves
```
No credentials, no IAM roles, no SDK configuration.
## No Separate API
Today, `writeonce-api` is a standalone Actix-web server that mediates between the frontend and the database. In the target architecture, the server logic is embedded in the same process as the database and the content renderer.
### What This Means
- **No HTTP hop between database and server.** Queries go directly from the request handler to the storage engine in-process. No network serialization, no connection pool, no ORM layer.
- **Single codebase.** No multi-repo coordination. A new article field is added once — in the content model — and it flows through storage, indexing, and rendering in the same compilation unit.
- **Single deployment.** One binary, one container, one process. No docker-compose orchestrating API + database + infra services.
The binary still exposes HTTP endpoints — it's still a web server. But it's a web server with an embedded database, not a web server that talks to an external one.
## Real-Time Subscriptions Without WebSocket
The current architecture has no mechanism for pushing content updates to connected clients. The target architecture adds real-time subscriptions, but explicitly without WebSocket.
### Why Not WebSocket
WebSocket adds connection state management, heartbeat logic, reconnection handling, and protocol upgrade complexity. For a content platform where updates are infrequent (articles are published, not streamed), the overhead isn't justified.
### Subscription Model
The target is a subscription mechanism where:
- A client subscribes to a content query (e.g., "all published articles" or "article with sys_title X")
- When the underlying data changes, the server pushes the relevant diff to the subscriber
- No polling from the client side
Candidate approaches to research:
- **Server-Sent Events (SSE)** — unidirectional push over HTTP. Simple, well-supported, no protocol upgrade. Natural fit for infrequent content updates.
- **SpacetimeDB-style subscriptions** — clients register queries, the engine tracks which rows match, and only sends diffs when the result set changes. This is the aspirational model.
- **Long polling** — fallback option. Simple but less efficient than SSE for multiple subscribers.
The key constraint: the subscription mechanism must work without requiring clients to maintain persistent bidirectional connections.
## Target Architecture
```
content directory
(JSON + MD files, .seg index)
|
| file watch + re-index
v
+---------------------------+
| writeonce binary |
| |
| +-------------------+ |
| | embedded storage | | .seg files, blog-title index
| | (read/write/index)| |
| +-------------------+ |
| | |
| +-------------------+ |
| | server logic | | route handlers, content queries
| | (HTTP endpoints) | |
| +-------------------+ |
| | |
| +-------------------+ |
| | subscription mgr | | SSE / query-based push
| | (real-time push) | |
| +-------------------+ |
| |
+---------------------------+
|
HTTP / SSE
|
v
+-------------------+
| frontend app | Angular or successor
| (browser client) |
+-------------------+
```
## Single Repository
The five current repos collapse into one:
```
writeonce/
content/ # articles (JSON + MD), images, assets
storage/ # embedded database engine (.seg files, indexing)
server/ # HTTP handlers, subscription manager
frontend/ # client application
writeonce.toml # configuration (port, content dir, index settings)
```
One repo. One build. One deploy artifact.
## What Needs Research
| Area | Question | Notes |
|------|----------|-------|
| **.seg file format** | What storage format gives efficient indexed reads over JSON+MD content? | Look at LSM trees, append-only logs, SQLite's page format for inspiration |
| **File watching** | How to efficiently detect content changes on Linux/macOS? | `inotify` on Linux, `kqueue` on macOS, or cross-platform via `notify` crate |
| **SSE vs alternatives** | Is SSE sufficient for the subscription model, or is something custom needed? | SSE handles the "push diffs to subscribers" case well for low-frequency updates |
| **Index structure** | What index structure supports `blog-title` lookup + date-ordered listing? | B-tree or hash index for title, sorted set for date ordering |
| **Language choice** | Continue with Rust for the unified binary? | Rust fits: single binary output, no runtime, strong typing, existing team knowledge |
| **Frontend coupling** | Should the frontend be embedded in the binary (serve static assets) or remain separate? | Embedding simplifies deployment; separate allows independent frontend iteration |
## Migration Path
The transition from current to target doesn't have to be all-or-nothing:
1. **Phase 1** — Build the embedded storage engine. Read JSON+MD files from a local directory, index by `blog-title`, serve via HTTP. No AWS, no PostgreSQL. This alone replaces `writeonce-api` + `aws-infra` + `lambda-function` + PostgreSQL.
2. **Phase 2** — Add real-time subscriptions (SSE). Clients subscribe to content queries and receive push updates when files change.
3. **Phase 3** — Collapse repositories. Move frontend into the unified codebase. Ship as a single binary that serves both API and static assets.
Each phase produces a working system. The current architecture can run in parallel until the new one is ready.
## Implementation phases
The "embedded storage engine" of Phase 1 above lands in three numbered plan docs under [`docs/plan/`](./plan/):
| Phase | Doc | What it ships |
| --- | --- | --- |
| 10 | [`plan/10-storage-foundations.md`](./plan/10-storage-foundations.md) | On-disk row codec (length-prefix + flags + LSN + CRC32C); per-type segment files (`data/<TypeName>.seg`); `posix_fallocate` preallocation; `pwrite`-only append path. Reads still in-memory. |
| 11 | [`plan/11-wal-and-recovery.md`](./plan/11-wal-and-recovery.md) | WAL log with `fdatasync` at commit; group commit per loop tick; control file with `last_durable_lsn` (rename-on-write); replay loop on startup. `kill -9` mid-write loses nothing acknowledged. |
| 12 | [`plan/12-engine-disk-cutover.md`](./plan/12-engine-disk-cutover.md) | `Engine`'s row payload moves to disk; in-memory map becomes `BTreeMap<i64, SegmentOffset>`. Periodic checkpoint flushes segments + advances the control file. RAM bounded by id-count, not row size. |
Postgres' storage subsystem is the design reference — see [`docs/plan/exploration/postgresql/`](./plan/exploration/postgresql/) for which Postgres modules informed which decision and what writeonce skips (multi-process IPC, latches, separate writer processes).
The durability syscalls themselves live in [`docs/plan/exploration/linux/12-pwrite-fsync.md`](./plan/exploration/linux/12-pwrite-fsync.md).

View file

@ -1,184 +0,0 @@
# Data Layer — Local Storage with Subscriptions
This document describes the embedded data layer that replaces PostgreSQL: local `.seg` files with indexing, and a subscription model where clients register queries and receive diffs on route visit — no polling required.
## .seg File Storage
The `.seg` (segment) format is the on-disk representation of article data. Each segment file holds serialized article content with positional indexing for fast lookups.
### Design Goals
- **No external database process.** The binary reads and writes `.seg` files directly. No socket connections, no protocol negotiation, no separate daemon.
- **Indexed by blog-title.** The primary access pattern is `GET /blog/:sys_title`. The storage layer must resolve a `sys_title` to its article content without scanning all files.
- **Append-friendly.** New articles and updates append to the segment. Deletes are tombstoned and compacted later.
- **Human-readable source.** The JSON + Markdown files remain the authoring format. `.seg` files are a derived index — if they're deleted, they can be rebuilt from the content directory.
### Proposed Structure
```
content/
linux-misc/
linux-misc.json # authored metadata (source of truth)
linux-misc.md # authored content (source of truth)
aws-lambda-pulumi/
aws-lambda-pulumi.json
aws-lambda-pulumi.md
data/
articles.seg # serialized article records
index/
title.idx # blog-title -> offset mapping
date.idx # publish date -> offset (sorted)
tags.idx # tag -> [offsets] (inverted index)
```
The `content/` directory is what the author edits. The `data/` directory is what the engine builds and queries. Losing `data/` is a cold start, not data loss.
### Segment File Internals
```
+------------------+
| segment header | magic bytes, version, record count
+------------------+
| record 0 | length-prefixed serialized article
+------------------+
| record 1 |
+------------------+
| ... |
+------------------+
| record N |
+------------------+
```
Each record is a length-prefixed byte sequence containing the full article (metadata + content merged). Records are addressed by byte offset from the start of the file.
### Index Files
**title.idx** — Hash map serialized to disk. Maps `sys_title` (string) to byte offset in `articles.seg`. Loaded into memory at startup for O(1) lookups.
**date.idx** — Sorted array of `(timestamp, offset)` pairs. Supports range queries for "articles published between X and Y" and ordered listing for the homepage.
**tags.idx** — Inverted index. Maps each tag string to a list of offsets. Supports "all articles tagged with X" queries.
On startup, index files are memory-mapped or loaded into heap. On content change, affected indexes are rebuilt incrementally.
## Subscription Model
The subscription model is inspired by SpacetimeDB: clients register queries, and the engine tracks which results match. When underlying data changes, only the relevant diffs are pushed to subscribers.
### How It Works
```
Client A Server Content Dir
| | |
|--- GET /blog/linux-misc -| |
| |-- read from .seg index ---->|
|<-- article + SSE stream -| |
| | |
| (subscribed to | |
| sys_title=linux-misc) | |
| | |
| |<-- file change detected ----|
| | |
| |-- re-index article -------->|
| |-- diff against last push -->|
| | |
|<-- SSE: updated content -| |
| | |
```
### Route-Based Subscription
When a user visits a route, the response includes both the current content and an SSE stream. The client is automatically subscribed to changes for that query — no explicit subscription handshake needed.
```
GET /blog/linux-misc
```
Response:
```
HTTP/1.1 200 OK
Content-Type: text/html
<!-- full article content rendered -->
<!-- SSE connection opened for this query -->
<script>
const source = new EventSource('/subscribe/blog/linux-misc');
source.onmessage = (event) => {
// apply diff to current content
};
</script>
```
The subscription lives as long as the browser tab is open. When the user navigates away, the EventSource closes and the server drops the subscription. No heartbeat management, no reconnection logic beyond what SSE provides natively (automatic reconnect is built into the EventSource API).
### Query Registration
Subscriptions are not limited to single-article lookups. The engine supports registering arbitrary content queries:
| Query Type | Example | Subscription Behavior |
|---|---|---|
| Single article | `sys_title = "linux-misc"` | Push when this specific article changes |
| All published | `published = true` | Push when any article is published or unpublished |
| By tag | `tags contains "rust"` | Push when a rust-tagged article is added, removed, or updated |
| Homepage list | `published = true ORDER BY date DESC LIMIT 10` | Push when the top-10 list changes |
The server maintains a registry of active subscriptions. On each content change, it evaluates which subscriptions are affected and pushes diffs only to those clients.
### Diff Format
When content changes, the server doesn't resend the full article. It sends a minimal diff:
```json
{
"type": "update",
"sys_title": "linux-misc",
"changes": {
"content.sections[2].paragraphs[0]": "Updated paragraph text...",
"content.tags": ["linux", "kernel", "new-tag"]
},
"version": 42
}
```
The `version` field enables clients to detect missed updates and request a full resync if needed.
## Sample Dataset
To validate the storage engine and subscription model, a sample dataset should exercise the core access patterns:
### Articles
| sys_title | tags | published | purpose |
|---|---|---|---|
| `sample-getting-started` | `[tutorial, beginner]` | true | Basic article, tests single-article subscription |
| `sample-rust-patterns` | `[rust, patterns]` | true | Tests tag-based queries |
| `sample-draft-wip` | `[draft]` | false | Tests published filter — should not appear in public queries |
| `sample-long-form` | `[deep-dive, rust]` | true | Multiple sections, images, code snippets — tests complex content rendering |
| `sample-frequently-updated` | `[changelog]` | true | Updated often — tests subscription diff delivery |
### Test Scenarios
1. **Cold start** — Delete `data/`, start the binary. It should rebuild `.seg` and index files from `content/` and serve all articles.
2. **Single article query** — `GET /blog/sample-getting-started` returns the article and opens an SSE subscription.
3. **Live update** — Edit `sample-frequently-updated.json` while a client is subscribed. The client should receive an SSE event with the diff.
4. **Tag query** — Subscribe to `tags contains "rust"`. Both `sample-rust-patterns` and `sample-long-form` should be in the result set. Adding a new article tagged `rust` should trigger a push.
5. **Publish toggle** — Change `sample-draft-wip` from `published: false` to `true`. Clients subscribed to the homepage list should receive a push with the new article added.
## SpacetimeDB Reference
SpacetimeDB is the primary architectural inspiration for the subscription model. Key concepts to study:
- **Modules** — server logic that runs inside the database, not beside it
- **Subscription queries** — clients register SQL-like queries; the engine evaluates them incrementally on each transaction
- **Incremental view maintenance** — only recompute the parts of a query result that changed
- **Client SDK generation** — type-safe client code generated from the server schema
Add SpacetimeDB as a reference submodule for quick access to their implementation patterns:
```bash
git submodule add https://github.com/clockworklabs/SpacetimeDB.git references/spacetimedb
```
The goal is not to replicate SpacetimeDB — it's to take its subscription semantics and apply them to a much narrower domain (blog content), where the simplicity of the problem allows a simpler implementation.

View file

@ -1,216 +0,0 @@
# User Interface — Server-Rendered HTMLX
No Angular. No React. No frontend framework. The UI is a set of `.htmlx` template files that the server parses, populates with content from the embedded database, and serves as plain HTML. Real-time updates arrive via SSE and are applied with minimal client-side scripting.
## Why Not Angular
The current `writeonce-app` is an Angular 18 SPA with Tailwind, PrismJS, ngx-markdown, and FontAwesome. It works, but it's a heavy delivery mechanism for what is fundamentally a read-heavy content site:
- **~200MB of `node_modules`** for a site that renders markdown articles
- **Client-side routing** for content that doesn't need it — every article is a distinct URL, not an interactive application
- **JavaScript-dependent rendering** — content doesn't exist until Angular boots, hydrates, and fetches from the API
- **Separate build pipeline** — `npm run build` produces static assets that must be deployed to nginx independently of the API
The content is static between updates. The interactivity is limited to navigation and code highlighting. A server-rendered approach matches the actual requirements.
## HTMLX Templates
The author defines the site layout using `.htmlx` files — HTML with embedded data bindings that the server resolves at render time.
### Template Structure
```
templates/
layout.htmlx # outer shell: <html>, <head>, <body>
header.htmlx # site header, navigation
footer.htmlx # site footer
home.htmlx # homepage: article list
article.htmlx # single article view
about.htmlx # static page
contact.htmlx # static page
components/
article-card.htmlx # summary card for article listings
code-snippet.htmlx # code block with language + title
img-caption.htmlx # image with caption
section.htmlx # article section (heading + paragraphs)
```
### Template Syntax
Templates use a binding syntax that references content from the database. The server parses these bindings, resolves them against the current content, and outputs plain HTML.
```html
<!-- header.htmlx -->
<header>
<nav>
<a href="/">writeonce</a>
<a href="/about">about</a>
<a href="/contact">contact</a>
</nav>
</header>
```
```html
<!-- article.htmlx -->
<article>
<h1>{{article.title}}</h1>
<p class="meta">by {{article.author}} &middot; {{article.tags}}</p>
{{#each article.sections}}
<section>
<h2>{{heading}}</h2>
{{#each paragraphs}}
<p>{{this}}</p>
{{/each}}
</section>
{{/each}}
{{#each article.codes}}
{{> code-snippet snippet=this}}
{{/each}}
{{#each article.images}}
{{> img-caption image=this}}
{{/each}}
</article>
```
```html
<!-- home.htmlx -->
<main>
<h1>articles</h1>
{{#each articles}}
{{> article-card article=this}}
{{/each}}
</main>
```
The `{{> partial}}` syntax includes another `.htmlx` file as a component. The server resolves these at render time — no client-side component tree.
### Content Subscription in Templates
Templates declare what data they need. The server resolves these declarations against the embedded database and subscribes the client to changes:
```html
<!-- article.htmlx -->
<!-- subscribe: article WHERE sys_title = :route_param -->
<article>
<h1>{{article.title}}</h1>
...
</article>
```
```html
<!-- home.htmlx -->
<!-- subscribe: articles WHERE published = true ORDER BY date DESC LIMIT 10 -->
<main>
{{#each articles}}
{{> article-card article=this}}
{{/each}}
</main>
```
The `<!-- subscribe: ... -->` comment is a directive to the server. It declares the query that populates the template's data context. The same query is used to register an SSE subscription for live updates (as described in [03-data.md](./03-data.md)).
## Rendering Pipeline
```
Browser request
|
v
Route match (/blog/linux-misc)
|
v
Load template (article.htmlx)
|
v
Parse subscribe directive
(article WHERE sys_title = "linux-misc")
|
v
Query embedded database (.seg index)
|
v
Resolve template bindings ({{article.title}}, etc.)
|
v
Compose with layout.htmlx + header.htmlx + footer.htmlx
|
v
Inject SSE subscription script
|
v
Send complete HTML response
```
The browser receives a fully rendered page on first load. No JavaScript framework boots. No API call fires. The content is already in the HTML.
## Live Updates via SSE
After the initial HTML is delivered, a small inline script opens an SSE connection for the page's subscription query:
```html
<script>
const source = new EventSource('/subscribe/blog/linux-misc');
source.onmessage = (event) => {
const diff = JSON.parse(event.data);
applyDiff(diff);
};
</script>
```
The `applyDiff` function is a lightweight client-side updater — it targets DOM elements by data attribute and patches their content. No virtual DOM, no reconciliation, no framework. For a content site where updates are infrequent and localized (a paragraph changed, a tag was added), direct DOM manipulation is sufficient.
```html
<h1 data-bind="article.title">Linux Misc</h1>
<p data-bind="article.sections[0].paragraphs[0]">First paragraph...</p>
```
When a diff arrives for `article.title`, the script finds the element with `data-bind="article.title"` and replaces its text content. This is the minimal client-side code the architecture requires.
## Code Highlighting
The current frontend uses PrismJS for syntax highlighting. In the server-rendered model, highlighting can happen at either layer:
**Server-side (preferred):** The server parses code blocks during template rendering and emits pre-highlighted HTML with CSS classes. The browser only needs the PrismJS CSS theme, not the JavaScript library. This eliminates client-side parsing entirely.
**Client-side (fallback):** Include PrismJS as a small script that runs on page load and on SSE update. Simpler to implement initially but adds a JavaScript dependency.
## Markdown Rendering
The current frontend uses `ngx-markdown` and `marked` to parse markdown in the browser. In the target architecture, markdown is rendered to HTML on the server during template composition. The browser never sees raw markdown.
This aligns with the content model: the JSON metadata already defines the article structure (sections, paragraphs, code snippets, images). The markdown file provides prose content. The server combines both into final HTML — the template just places the pre-rendered blocks.
## What Gets Removed
| Current (Angular) | Target (HTMLX) |
|---|---|
| `writeonce-app/` (full Angular project) | `templates/` (handful of .htmlx files) |
| `node_modules/` (~200MB) | None |
| `angular.json`, `tsconfig.json`, `karma.conf.js` | None |
| npm build pipeline | Template parsed at request time |
| Nginx static file serving | Binary serves its own HTML |
| Client-side routing | Server-side route matching |
| Client-side markdown parsing | Server-side rendering |
| Client-side code highlighting | Server-side or minimal JS |
## Styling
Templates use plain CSS. Tailwind can optionally be used as a build-time utility (generating a static CSS file), but there is no runtime CSS framework. The author writes styles in a `styles.css` file that the server serves as a static asset.
```
templates/
styles/
main.css # site-wide styles
article.css # article-specific styles
code-theme.css # syntax highlighting theme (PrismJS compatible)
```
## Template Authoring Experience
The `.htmlx` files are editable by the same author who writes articles. The template syntax is intentionally close to HTML — there's no JSX, no TypeScript, no build step. An author who knows HTML can modify the site layout.
This closes the loop on the writeonce philosophy: the author writes content (markdown + JSON) and layout (`.htmlx` + CSS) as files, and the binary turns them into a live site.

View file

@ -1,139 +0,0 @@
# Data Layer — Implementation Status
The embedded data layer described in [02-recovery.md](./02-recovery.md) and [03-data.md](./03-data.md) has been implemented as a Cargo workspace with 8 crates. All 44 tests pass. No external database, no AWS, no tokio — direct Linux syscalls on a custom event loop.
## Workspace Structure
```
writeonce-all/
Cargo.toml # workspace root
docs/ # architecture documentation
sample-content/ # 5 test articles for validation
crates/
wo-model/ # content model
wo-seg/ # .seg file format
wo-index/ # index files
wo-store/ # unified storage engine
wo-watch/ # inotify file watcher
wo-event/ # epoll event loop
wo-sub/ # subscription system
wo-rt/ # custom runtime
```
## Crate Summary
| Crate | Purpose | Tests | Key Types |
|-------|---------|-------|-----------|
| **wo-model** | Article structs matching existing JSON schema, `ContentLoader` for directory walking | 8 | `Article`, `ArticleContent`, `ArticleBody`, `Section`, `CodeSnippet`, `ContentLoader` |
| **wo-seg** | Binary `.seg` file format — length-prefixed records, tombstoning, positional I/O | 6 | `SegWriter`, `SegReader`, `SegHeader` |
| **wo-index** | Three index types for O(1) and O(log n) access patterns | 8 | `TitleIndex`, `DateIndex`, `TagIndex` |
| **wo-store** | Unified storage engine composing seg + indexes, cold-start rebuild | 3 | `Store` |
| **wo-watch** | Content directory watcher using inotify | 4 | `ContentWatcher`, `ContentChange` |
| **wo-event** | Custom event loop on epoll with eventfd, timerfd, signalfd | 5 | `EventLoop`, `EventFd`, `TimerFd`, `SignalFd` |
| **wo-sub** | Subscription manager with fd-based notifications, `register!` macro | 6 | `SubscriptionManager`, `Subscription`, `Notification` |
| **wo-rt** | Runtime tying all crates together — single process, single event loop | 4 | `Runtime`, `RuntimeHandle`, `Config` |
## Linux Kernel Syscalls Used
| Syscall | Crate | Purpose |
|---------|-------|---------|
| `pread` / `pwrite` | wo-seg | Positional read/write for .seg records without seeking |
| `fallocate` | wo-seg | Pre-allocate .seg file space to reduce fragmentation |
| `epoll_create1` / `epoll_ctl` / `epoll_wait` | wo-event | Event-driven I/O multiplexing for the main loop |
| `eventfd` | wo-event, wo-sub | Lightweight signaling between watcher and subscription manager |
| `timerfd_create` / `timerfd_settime` | wo-event | Periodic tasks (compaction, keepalive) as file descriptors |
| `signalfd` | wo-event | SIGINT/SIGTERM delivered as fd events for graceful shutdown |
| `inotify_init1` / `inotify_add_watch` | wo-watch | File system change detection on the content directory |
| `pipe2` | wo-sub (tests) | Mock subscriber fds for testing notification delivery |
## .seg File Format
```
Offset Size Field
0 4 Magic: b"WOSF"
4 2 Version: u16 LE (1)
6 2 Flags: u16 LE (reserved)
8 8 Record count: u64 LE
16 8 Data start offset: u64 LE
24 8 Reserved
32+ variable Records: [u32 length][u8 flags][bincode payload]...
```
- Records are addressed by byte offset from file start
- Flags: `0x00` = active, `0x01` = tombstoned
- Payload: bincode-serialized `Article` struct
## Index Files
| File | Format | Access Pattern |
|------|--------|----------------|
| `title.idx` | On-disk hash table (Robin Hood, load factor 0.5), 138 bytes/slot | O(1) lookup by `sys_title` |
| `date.idx` | Sorted `(i64 timestamp, u64 offset)` array, 16 bytes/entry | Binary search for date ranges, latest N |
| `tags.idx` | Bincode-serialized `HashMap<String, Vec<u64>>` | Tag-to-offsets inverted index |
All indexes are derived from `.seg` and rebuildable from `content/` on cold start.
## Subscription Model
No SSE. No WebSocket. Notifications are written directly to subscriber file descriptors.
- **Subscribe**: `SubscriptionManager::subscribe(fd, Subscription::ByTitle("linux-misc"))`
- **Notify**: on content change, length-prefixed `Notification` written to matching fds
- **Cleanup**: `EPOLLHUP` on epoll triggers automatic `unsubscribe(fd)`
- **Dedup**: if a fd matches multiple patterns (title + tag), it receives only one notification
Subscription patterns:
- `Subscription::ByTitle(sys_title)` — single article
- `Subscription::ByTag(tag)` — all articles with tag
- `Subscription::All` — all content changes
## Store Query API
```rust
store.get_by_title("linux-misc") -> Option<Article>
store.list_published(skip, limit) -> Vec<Article>
store.list_by_tag("rust") -> Vec<Article>
store.list_by_date_range(start, end) -> Vec<Article>
store.count_published() -> usize
store.article_version("linux-misc") -> Option<u64>
store.rebuild() // full rebuild from content/
```
## Runtime Event Loop
Single `epoll` instance multiplexing all file descriptors:
| Token | Fd | Handler |
|-------|----|---------|
| `WATCHER` | inotify fd | Process file changes → update store → notify subscribers |
| `SIGNAL` | signalfd | SIGINT/SIGTERM → graceful shutdown |
| `TIMER` | timerfd | Periodic tasks (compaction, stats) |
| `NOTIFY` | eventfd | Subscription notification signal |
| `1000+` | subscriber fds | Hangup detection → unsubscribe + cleanup |
## External Dependencies
| Crate | Version | Purpose |
|-------|---------|---------|
| `serde` | 1.x | Serialization derives |
| `serde_json` | 1.x | JSON parsing for article files |
| `bincode` | 1.x | Compact binary serialization for .seg records and notifications |
| `libc` | 0.2.x | Raw Linux syscall bindings |
No tokio. No async-std. No database driver. No HTTP framework (yet).
## What Comes Next
The data layer delivers everything the HTTP server and UI layers need:
1. **`Store` with zero-copy query access** — all article queries resolve in-process
2. **Subscription system accepting raw fds** — HTTP layer hands socket fds to `subscribe()`
3. **Shared event loop** — HTTP listener socket registers on the same epoll
4. **Automatic cold-start** — if `data/` is missing, rebuilds from `content/` on startup
5. **Graceful shutdown** — SIGTERM triggers clean fd cleanup
Next phases per [02-recovery.md](./02-recovery.md):
- **HTTP server** — route handlers using the `Store` query API, embedded in the same binary
- **HTMLX templates** — server-rendered HTML with `{{bindings}}` per [04-ui.md](./04-ui.md)
- **Frontend collapse** — serve static assets from the binary, eliminate the Angular app
3

View file

@ -1,246 +0,0 @@
# Markdown File Rendering
## Current State (writeonce-articles-s3)
Each article is a directory containing a JSON metadata file and one or more `.md` files:
```
auto-scale-gitlab-runner-using-aws-spot-instance/
docker-machine-test-with-t2.md
gitlab-runner-config.md
stop-test-gitlab-docker-machine.md
gitlab-runner-with-kubernetes-executor/
gitlab-runner-with-kubernetes-executor.json
deploy.md
permission.md
role-binding.md
role-defination.md
gitlab-runnergitlab-runner-deploy.md
```
The JSON metadata currently defines the full article structure — sections, headings, paragraphs, and code snippet references. Markdown files are limited to code blocks referenced via the `codes[].snippet` field.
## Problem
The JSON metadata carries too much content. Headings, paragraphs, prose — all of this is duplicated as JSON strings inside `content.content.sections`. The markdown files only hold code snippets, referenced by `sectionIndex` and `paragraphIndex`.
This is backwards. The markdown file should be the content. The JSON should be minimal metadata.
## Target: Markdown-First Content Model
**The markdown file is the article.** All prose, headings, code blocks, and inline formatting live in the `.md` file. The JSON metadata file holds only what markdown cannot express: system fields, tags, publication state, and author.
### Minimal JSON Metadata
```json
{
"sys_title": "gitlab-runner-with-kubernetes-executor",
"title": "Gitlab Runner with Kubernetes Executor",
"published": true,
"author": "Shoney Arickathil",
"tags": ["kubernetes", "gitlab", "ci-cd"],
"published_on": 1740950884
}
```
No `content.content.sections`. No `content.content.codes`. No `paragraphs[]` arrays. No `sectionIndex`/`paragraphIndex` mapping.
### Markdown File = Full Article Content
````markdown
# Introduction
Deploying a Gitlab runner using kubernetes is a great option to overcome
the limitations of other gitlab runner executor such as docker and docker machine.
## Running Gitlab Runner in gitlab namespace
Create the namespace and apply the deployment:
```yaml
apiVersion: apps/v1
kind: Deployment
metadata:
name: gitlab-runner
namespace: gitlab
```
````
## Permissions
The runner needs RBAC permissions to create pods:
```yaml
apiVersion: rbac.authorization.k8s.io/v1
kind: Role
metadata:
name: gitlab-runner
```
Everything is in the markdown — headings, paragraphs, code blocks with language hints, links, images. The rendering pipeline parses the markdown directly.
### Directory Structure
```
content/
gitlab-runner-with-kubernetes-executor/
gitlab-runner-with-kubernetes-executor.json # minimal metadata
gitlab-runner-with-kubernetes-executor.md # full article content
linux-misc/
linux-misc.json
linux-misc.md
```
One JSON for metadata. One markdown for content. No scattered `.md` files per code snippet.
## What Changes
| Before | After |
| -------------------------------------------------------------- | ----------------------------------------------------------------------- |
| JSON holds sections, headings, paragraphs as structured arrays | JSON holds only sys_title, title, published, author, tags, published_on |
| Markdown files hold only code snippets | Markdown file holds the entire article |
| `codes[].snippet` maps filename to sectionIndex/paragraphIndex | No mapping needed — headings and code blocks are inline in markdown |
| Renderer reads JSON structure, injects code from .md files | Renderer parses markdown directly into HTML |
| Multiple .md files per article (one per code snippet) | One .md file per article |
## Impact on the Data Layer
### wo-model
The `Article` struct simplifies:
```rust
pub struct Article {
pub sys_title: String,
pub title: String,
pub published: bool,
pub author: String,
pub tags: Vec<String>,
pub published_on: Option<i64>,
}
```
The nested `ArticleContent` / `ArticleBody` / `Section` / `CodeSnippet` hierarchy is no longer needed. Article content comes from parsing the `.md` file at render time, not from the JSON.
### wo-md
Currently handles only inline markdown (`**bold**`, `` `code` ``, links). Needs to become a full markdown-to-HTML renderer:
- Block elements: headings (`#`, `##`), paragraphs, code fences (` `lang ```), lists, blockquotes
- Inline elements: bold, italic, code, links, images
- Code fence language extraction for `wo-md::highlight()`
- The renderer reads `{sys_title}/{sys_title}.md`, parses it, and returns HTML
### wo-htmlx
The `article.htmlx` template simplifies. Instead of iterating `{{#each article.content.content.sections}}`, it renders the pre-parsed markdown HTML:
```html
<article>
<h1>{{article.title}}</h1>
<p class="meta">by {{article.author}} &middot; {{article.tags}}</p>
{{article.content_html}}
</article>
```
Where `content_html` is the full HTML output from the markdown renderer.
### wo-store
`ContentLoader` reads the `.json` for metadata and the `.md` for content. The `.seg` file stores both. At query time, the markdown is either:
- Pre-rendered to HTML during ingestion (stored in .seg alongside metadata)
- Rendered on-demand at request time (read .md from disk)
Pre-rendering is preferred — it avoids parsing markdown on every HTTP request.
## Migration Path
1. Update `wo-model` with the simplified `Article` struct
2. Extend `wo-md` to handle full markdown (block-level parsing, code fences)
3. Update `ContentLoader` to read `.json` + `.md` pairs
4. Update `wo-store` to store pre-rendered HTML in the .seg file
5. Simplify `article.htmlx` template
6. Migrate existing articles: extract prose from JSON into `.md` files
Existing articles with the old JSON format can coexist during migration — `ContentLoader` checks for a `.md` file and falls back to the JSON structure if none exists.
## Blog Subscription — Live Content Reload
When a user visits `http://localhost:3000/blog/sample-rust-patterns`, the content should stay live. Any edit to `sample-content/sample-rust-patterns/sample-rust-patterns.md` must auto-reflect in the browser without a page refresh.
### How It Works
```
Browser visits /blog/sample-rust-patterns
│
▼
1. Server renders article HTML from .seg (pre-rendered from .md)
2. Server writes HTML response to socket fd
3. Server registers socket fd in subscription table:
register!(sub_manager, socket_fd, ByTitle("sample-rust-patterns"))
4. Connection transitions to Subscribed state (stays open)
│
│ (user edits sample-rust-patterns.md)
│
▼
5. inotify fires IN_MODIFY on sample-rust-patterns.md
6. ContentWatcher maps file → sys_title "sample-rust-patterns"
7. Store rebuilds: re-reads .json + .md, re-renders markdown to HTML, updates .seg + indexes
8. SubscriptionManager::notify("sample-rust-patterns", ...) fires
9. For each subscribed fd: write(fd, diff_payload)
│
▼
10. Browser receives payload on the open connection
11. Client-side script applies the update to the DOM
```
### What Needs to Work
| Component | Requirement |
|-----------|-------------|
| **inotify** (wo-watch) | Already watches `content/` directory. `.md` file changes must trigger `ContentChange::Modified(sys_title)` |
| **Store rebuild** (wo-store) | On `.md` change: re-read file, re-render markdown to HTML, update `.seg` and indexes |
| **Subscription table** (wo-sub) | Route handler registers the browser's socket fd via `register!` after sending initial HTML |
| **Notification** (wo-sub) | On content change, write updated `content_html` to all subscribed fds as JSON payload |
| **Event loop** (wo-rt) | After writing initial response, transition connection to `Subscribed` state. Keep fd on epoll for hangup detection. |
| **Client script** | Injected in the HTML. Reads payloads from the open connection. Replaces article content in the DOM. |
### Client-Side Script
Injected by the template renderer into every article page:
```html
<script>
// Connection stays open after initial HTML.
// Server writes length-prefixed JSON payloads when content changes.
const decoder = new TextDecoder();
const articleEl = document.querySelector('article');
fetch(window.location.href, { headers: { 'X-Subscribe': '1' } })
.then(r => r.body.getReader())
.then(reader => {
(function read() {
reader.read().then(({ done, value }) => {
if (done) return;
try {
const payload = JSON.parse(decoder.decode(value));
if (payload.content_html) {
articleEl.innerHTML = payload.content_html;
}
} catch (e) {}
read();
});
})();
});
</script>
```
### inotify and .md Files
The current `ContentWatcher` watches for `.json` file changes. It must also trigger on `.md` file changes:
- `IN_MODIFY` on `*.md` → `ContentChange::Modified(sys_title)`
- The sys_title is derived from the parent directory name (same as for JSON)
- Both `.json` and `.md` changes trigger a store rebuild and subscriber notification

View file

@ -1,252 +0,0 @@
# SSL and Deployment
## Problem
In [01-problem.md](./01-problem.md), the infrastructure overhead was identified — multiple repos, AWS dependencies, separate deployment pipelines. But one problem went unaddressed: the server-side infrastructure that sits in front of the application — nginx reverse proxy, SSL certificates, systemd service management, and deployment to the production host.
Currently this requires manual SSH, manual nginx config, manual certbot runs. For a single-binary platform, the deployment should be as simple as the architecture.
## Target
Given:
- SSH access to `writeonce.de` is configured
- nginx exists at the default path `/etc/nginx/`
- The writeonce binary listens on a local port (e.g., `127.0.0.1:3000`)
The deployment pipeline should:
1. Build the binary
2. Copy it to the server
3. Create/update the systemd service
4. Restart the service
5. Configure nginx as a reverse proxy
6. Obtain and auto-renew SSL certificates via Let's Encrypt
## Systemd Service
The writeonce binary runs as a systemd service for automatic restart, logging, and boot-start.
### Service File
```ini
# /etc/systemd/system/writeonce.service
[Unit]
Description=writeonce content platform
After=network.target
[Service]
Type=simple
User=writeonce
Group=writeonce
WorkingDirectory=/opt/writeonce
ExecStart=/opt/writeonce/writeonce
Restart=on-failure
RestartSec=5
StandardOutput=journal
StandardError=journal
# Security hardening
NoNewPrivileges=true
ProtectSystem=strict
ProtectHome=true
ReadWritePaths=/opt/writeonce/data
PrivateTmp=true
[Install]
WantedBy=multi-user.target
```
### Directory Layout on Server
```
/opt/writeonce/
writeonce # the binary
content/ # article .json + .md files
data/ # derived .seg + .idx (rebuilt on start)
templates/ # .htmlx templates
static/ # CSS, images
```
### Service Management
```bash
# Install / update
sudo systemctl daemon-reload
sudo systemctl enable writeonce
sudo systemctl restart writeonce
# Check status
sudo systemctl status writeonce
journalctl -u writeonce -f
```
## Nginx Reverse Proxy
Nginx sits in front of the writeonce binary, handling SSL termination and proxying requests to `127.0.0.1:3000`.
### Nginx Config
```nginx
# /etc/nginx/sites-available/writeonce.de
server {
listen 80;
server_name writeonce.de www.writeonce.de;
return 301 https://$server_name$request_uri;
}
server {
listen 443 ssl http2;
server_name writeonce.de www.writeonce.de;
ssl_certificate /etc/letsencrypt/live/writeonce.de/fullchain.pem;
ssl_certificate_key /etc/letsencrypt/live/writeonce.de/privkey.pem;
ssl_protocols TLSv1.2 TLSv1.3;
ssl_ciphers HIGH:!aNULL:!MD5;
ssl_prefer_server_ciphers on;
# HSTS
add_header Strict-Transport-Security "max-age=31536000; includeSubDomains" always;
location / {
proxy_pass http://127.0.0.1:3000;
proxy_set_header Host $host;
proxy_set_header X-Real-IP $remote_addr;
proxy_set_header X-Forwarded-For $proxy_add_x_forwarded_for;
proxy_set_header X-Forwarded-Proto $scheme;
# Keep connections open for database subscriptions
proxy_http_version 1.1;
proxy_set_header Connection "";
proxy_read_timeout 86400s;
proxy_send_timeout 86400s;
}
# Static assets — let nginx serve directly for better caching
location /static/ {
alias /opt/writeonce/static/;
expires 1y;
add_header Cache-Control "public, immutable";
}
}
```
### Enable Site
```bash
sudo ln -sf /etc/nginx/sites-available/writeonce.de /etc/nginx/sites-enabled/
sudo nginx -t
sudo systemctl reload nginx
```
## SSL with Let's Encrypt
### Initial Certificate
```bash
sudo apt install certbot python3-certbot-nginx
sudo certbot --nginx -d writeonce.de -d www.writeonce.de
```
Certbot modifies the nginx config to add SSL directives and obtains the certificate.
### Auto-Renewal
Certbot installs a systemd timer that runs twice daily:
```bash
# Check timer
systemctl list-timers | grep certbot
# Manual test
sudo certbot renew --dry-run
```
Certificates auto-renew before expiry. Nginx reloads automatically via certbot's deploy hook.
### Deploy Hook for Nginx Reload
```bash
# /etc/letsencrypt/renewal-hooks/deploy/reload-nginx.sh
#!/bin/bash
systemctl reload nginx
```
## Deployment Script
A single script that builds, copies, and restarts:
```bash
#!/bin/bash
# deploy.sh — run from the development machine
set -e
SERVER="writeonce.de"
REMOTE_DIR="/opt/writeonce"
echo "Building release binary..."
cargo build --release -p wo-rt --bin writeonce
echo "Copying binary to server..."
scp target/release/writeonce $SERVER:$REMOTE_DIR/writeonce.new
echo "Syncing content and templates..."
rsync -az --delete content/ $SERVER:$REMOTE_DIR/content/
rsync -az --delete templates/ $SERVER:$REMOTE_DIR/templates/
rsync -az --delete static/ $SERVER:$REMOTE_DIR/static/
echo "Swapping binary and restarting..."
ssh $SERVER "
sudo mv $REMOTE_DIR/writeonce.new $REMOTE_DIR/writeonce
sudo systemctl restart writeonce
"
echo "Deployed. Checking status..."
ssh $SERVER "sudo systemctl status writeonce --no-pager"
```
### First-Time Setup
Run once on the server to create the user, directory, and service:
```bash
#!/bin/bash
# setup.sh — run on the server
set -e
# Create user
sudo useradd -r -s /bin/false writeonce
# Create directory
sudo mkdir -p /opt/writeonce/{content,data,templates,static}
sudo chown -R writeonce:writeonce /opt/writeonce
# Install service
sudo cp writeonce.service /etc/systemd/system/
sudo systemctl daemon-reload
sudo systemctl enable writeonce
# Configure nginx
sudo cp writeonce.de.nginx /etc/nginx/sites-available/writeonce.de
sudo ln -sf /etc/nginx/sites-available/writeonce.de /etc/nginx/sites-enabled/
sudo nginx -t
sudo systemctl reload nginx
# SSL
sudo certbot --nginx -d writeonce.de -d www.writeonce.de
```
## What This Replaces
| Before | After |
|--------|-------|
| Pulumi IaC managing Lambda + S3 | `deploy.sh` with scp + rsync |
| AWS Lambda deployment pipeline | `systemctl restart writeonce` |
| S3 bucket for content hosting | `rsync content/` to server |
| Docker Compose for API + DB | Single binary, one systemd service |
| Multiple nginx configs for API + frontend | One nginx config, one proxy_pass |
| Manual SSL setup | `certbot --nginx` with auto-renewal |
## Connection Keepalive for Subscriptions
The nginx config sets `proxy_read_timeout 86400s` (24 hours) to keep persistent connections open for the database subscription model. When a browser visits an article page and the connection transitions to `Subscribed` state, nginx must not timeout and close the upstream connection.
If nginx is removed in the future (the binary handles TLS directly via `rustls`), this concern disappears — the binary owns the socket end-to-end.

View file

@ -90,19 +90,18 @@ Unchanged from CLAUDE.md's description: `rt` is the monolithic Stage-2 runtime p
``` ```
docs/ docs/
├── 01…08-*.md numbered design docs (this file is 08) ├── 00-*,01,08-*.md status / principles / problem / structure docs
├── writeonce-pl.md language positioning ├── runtime/ the 7-phase database design series + runtime concept refs
├── runtime/ user-facing language overview + the 7-phase database series ├── examples/ log-watcher/, employee/, employee-list/ samples
├── examples/ blog/, ecommerce/, pricing/ samples; ⏳ log-watcher/ (plan 10)
├── plan/ numbered engineering plans 00–16, linux/ cards, assembly/, ├── plan/ numbered engineering plans 00–16, linux/ cards, assembly/,
│ ├── exploration/ c-runtime/ (A–F, done), ui/ (htmlx track), colibri/ │ ├── exploration/ c-runtime/ (A–F, done), linux/, postgresql/, assembly/
│ └── oop-vm/ ⏳ the OOP-track contracts: 00-wob-format, 01-error-catalog, │ └── oop-vm/ ⏳ the OOP-track contracts: 00-wob-format, 01-error-catalog,
│ 02-corpus, 03-shard-actor, 04-db-binding, 05-http-service, │ 02-corpus, 03-shard-actor, 04-db-binding, 05-http-service,
│ 06-ui-live, 07-systems-stdlib │ 06-ui-live, 07-systems-stdlib
├── superpowers/ ├── superpowers/
│ ├── specs/ the two approved track specs (2026-08-01) │ ├── specs/ the two approved track specs (2026-08-01)
│ └── plans/ implementation plans 1–10 (2026-08-01, prose-only) │ └── plans/ implementation plans 1–10 (2026-08-01, prose-only)
└── future-scope/, cm.md legacy notes └── cm.md legacy notes
``` ```
Repo rule restated: documentation belongs here; code directories keep one orientation README each. Repo rule restated: documentation belongs here; code directories keep one orientation README each.

View file

@ -0,0 +1,42 @@
# employee-list — program B: attach, authenticate, read
> **Status: target workload — does not compile on today's toolchain.**
> Written ahead of iterations
> [9c (cross-program tables)](../../stories/language-runtime-database/09c-cross-program-tables.md)
> and [9d (keypair attach auth)](../../stories/language-runtime-database/09d-keypair-attach-auth.md),
> the way every acceptance sample here precedes its features. It also leans
> on 9/9b (the [employee sample](../employee/) it attaches to must run
> first).
Two programs, one database, one writer:
```
employee (A) employee-list (B)
owns WO_DATA + WAL no database of its own
[share] listen = unix:...sock [connect.employee] ipc = unix:...sock
[[share.clients]] public_key = <A's fingerprint, pinned>
public_key = <B's fingerprint> project = ../employee (shapes)
rights = "read"
▲ │
└── every statement executes here ◄───┘ (typed, over the wire)
```
The manifests are the design: **A grants, B pins.** A's `[share]` names B's
public-key fingerprint with rights (`read` here); B's `[connect.employee]`
names A's IPC string AND A's fingerprint, so neither side talks to an
impostor. Fingerprints are printed by each program's `--identity` after
first boot (keys are generated into `WO_DATA`, never written into a toml)
and pasted — the `PASTE-…-HERE` placeholders mark exactly where. The
connect-section name is the code's namespace: `[connect.employee]` is why
the source says `employee.Employee`.
| Mode | What it proves |
| --- | --- |
| `employee-list list` | typed reads over the wire, `e.dept.name` ref navigation executing inside A |
| `employee-list report` | **byte-identical output to A's own `report`** — attach + GroupBy compose, the wire changes nothing |
| `employee-list staff <dept>` | unique-name index probe + `staff` backlink scan, both in A |
| `employee-list probe-write` | the rights matrix: registered read-only, so the insert traps with access-denied (caught, `DENIED …`, exit 4) and A's row count is unchanged |
The 9d acceptance drives the rest from the outside: wrong key, no key,
same-uid-wrong-key, impostor socket, handshake replay, key rotation — see
the iteration's criteria; this sample is the workload they run against.

View file

@ -0,0 +1,73 @@
-- employee-list — program B of the cross-program story (iterations 9c/9d).
-- Attaches to the RUNNING employee program (A) named by [connect.employee]
-- in wo.toml: keypair handshake first (9d), then typed statements over the
-- wire (9c). A registered this program read-only, so every mode here reads —
-- except probe-write, which exists to prove the rights matrix refuses.
--
-- The `employee.` prefix is the manifest's connect-section name: these are
-- A's tables, checked against A's shapes at compile time and re-verified by
-- the schema handshake at attach. A stays the single writer; every statement
-- below executes inside A.
fn main(args: multi Text) -> Int {
if len(args) >= 1 and args[0] == "list" { return list(); }
if len(args) >= 1 and args[0] == "report" { return report(); }
if len(args) >= 2 and args[0] == "staff" { return staff(args[1]); }
if len(args) >= 1 and args[0] == "probe-write" { return probe_write(); }
print_err("usage:");
print_err(" employee-list list every employee, with department");
print_err(" employee-list report aggregates by department (A's own report, over the wire)");
print_err(" employee-list staff <dept> one department's staff");
print_err(" employee-list probe-write prove read-only: the insert must be DENIED");
return 1;
}
-- Every employee with forward ref navigation — each `e.dept.name` is a
-- point read executing inside A.
fn list() -> Int {
for e in from x in employee.Employee order by x.name select x {
print("EMP ${e.name} ${e.salary} ${e.dept.name}");
}
return 0;
}
-- Byte-identical output to A's own `employee report` — the 9c acceptance
-- line: attach + query compose, and the wire changes nothing.
fn report() -> Int {
let rows = from e in employee.Employee
group e by e.dept into g
order by avg(g.salary) desc
select { dept: g.key.name, headcount: count(g),
avg_salary: avg(g.salary), min_salary: min(g.salary),
max_salary: max(g.salary) };
for r in rows {
print("DEPT ${r.dept} headcount=${r.headcount} avg=${r.avg_salary} min=${r.min_salary} max=${r.max_salary}");
}
let payroll = sum(from e in employee.Employee select e.salary);
print("PAYROLL ${payroll}");
return 0;
}
-- Unique-name index probe + backlink scan, both executing in A.
fn staff(name: Text) -> Int {
let ds = from d in employee.Department where d.name == name take 1 select d;
if len(ds) == 0 { print_err("no such department: ${name}"); return 1; }
for e in from s in ds[0].staff order by s.salary desc select s {
print("STAFF ${e.name} ${e.salary}");
}
return 0;
}
-- The rights matrix, exercised: this program is registered READ-ONLY, so
-- the insert must trap with the access-denied code — caught here, printed,
-- and nothing was applied or WAL-logged in A (the acceptance asserts A's
-- row count is unchanged).
fn probe_write() -> Int {
let id = try insert employee.Department { name: "Intruders" } catch (e) nil;
if id == nil {
print("DENIED write to employee.Department (registered read-only)");
return 4;
}
print_err("UNEXPECTED: write succeeded with id ${id} — rights not enforced");
return 1;
}

View file

@ -0,0 +1,29 @@
name = "employee-list"
version = "0.1.0"
description = "Program B of the cross-program story: attaches read-only to the running employee program (A) over its IPC channel after keypair authentication, and lists/aggregates A's tables over the wire"
[runtime]
wo = ">= 0.1"
[build]
runtime = "../../../runtime/wovm"
# Iteration 9c/9d surface (target — compiles once those land).
# The section NAME is the namespace this program's code uses: [connect.employee]
# makes A's tables reachable as `employee.Department` / `employee.Employee`.
[connect.employee]
# A's IPC string (9c fork 1: unix socket, listening beside A's WO_DATA;
# A declares the same path in its [share] section).
ipc = "unix:../employee/target/data/employee.sock"
# A's PINNED public-key fingerprint (9d): B refuses to speak to anything on
# that socket path that cannot sign A's challenge with this key — the
# impostor-socket acceptance line. Printed by `employee --identity` after
# A's first boot; paste it here. Never a private key, never a secret.
public_key = "ed25519:PASTE-EMPLOYEE-FINGERPRINT-HERE"
# Where B's compiler reads A's table shapes at compile time (9c fork 2's
# milestone lean: project-directory reference). The runtime handshake
# re-verifies the shapes against A's live class table at attach — a stale
# checkout refuses the attachment with both shapes named.
project = "../employee"

View file

@ -0,0 +1,32 @@
# employee — the database track's acceptance workload
> **Status: target workload — does not compile on today's toolchain.**
> This sample is written *ahead of* the features it exercises, exactly as
> log-watcher was written ahead of iterations 5–7: the sample is the test,
> and the plans compile toward it. It becomes buildable when iteration 9
> (engine: [`2026-08-01-db-engine-binding.md`](../../superpowers/plans/2026-08-01-db-engine-binding.md))
> and iteration 9b (query surface:
> [`2026-08-15-employee-relations-query.md`](../../plan/compiler/2026-08-15-employee-relations-query.md))
> land. Normative semantics:
> [the 9b spec](../../superpowers/specs/2026-08-15-table-relations-query-design.md).
Two `@table` classes and every 9b feature load-bearing:
| Mode | What it proves |
| --- | --- |
| `employee seed` | `insert` + WAL-before-ack; a second run catches the `departments.name` `@unique` trap (`SEED-DUP`, exit 3) |
| `employee report` | `group … by … into g` lowered as one hash pass; `count`/`avg`/`min`/`max` per department; `order by avg(g.salary) desc`; whole-query `sum` for payroll |
| `employee staff <dept>` | unique-name **index probe** (asserted via the engine's probe counter, not assumed), `backlink` scan one way, `e.dept.name` ref navigation the other |
| `employee raise <dept> <pct>` | update through a query result; a missing department takes the empty-query path |
| `employee drop <dept>` | delete **restrict**: a department with staff traps (exit 4); the trap code is the acceptance's assertion |
The engine lives in `database/` (statically linked into `wovm` — still one
binary). Rows are RAM-authoritative, WAL-durable; a kill -9 between `seed`
and `report` followed by an identical `report` is part of the acceptance
script (iteration 9's replay, proven on this workload).
Data shape: `Department { name @unique, staff: backlink Employee.dept }`,
`Employee { name, salary (cents), hired (epoch ms), dept: ref Department }`,
indexes `[name]`, `[dept]`, `[dept, salary]`. Salaries wrap like all language
arithmetic; `avg`/`min`/`max` are `?Int` because an empty group is data, not
a fault.

View file

@ -0,0 +1,14 @@
# docs/examples/employee — the database track's acceptance workload.
# `just employee::<recipe>` from the repo root, or plain `just <recipe>` here.
ROOT := source_directory() / "../../.."
default: accept
# build the standalone binary from wo.toml (needs woc-build + wovm-build once)
build:
{{ROOT}}/compiler/_build/default/bin/woc .
@ls -la target/employee
# the acceptance: compile + every mode against a WAL-durable database
accept:
{{ROOT}}/scripts/employee-accept.sh

View file

@ -0,0 +1,124 @@
-- Employee management — iteration 9/9b's acceptance workload.
-- Four modes, each existing to make one feature load-bearing:
-- seed inserts (WAL, @unique trap on the second run)
-- report the GroupBy showcase (headcount/avg/min/max, payroll)
-- staff <dept> relation navigation both directions + index probe
-- raise <dept> <p> update through a query result; empty query => nil path
-- drop <dept> delete restrict: a department with staff must trap
fn main(args: multi Text) -> Int {
if len(args) >= 1 and args[0] == "seed" { return seed(); }
if len(args) >= 1 and args[0] == "report" { return report(); }
if len(args) >= 2 and args[0] == "staff" { return staff(args[1]); }
if len(args) >= 3 and args[0] == "raise" {
let pct = parse_int(args[2]);
if pct == nil { print_err("raise: <pct> must be a number"); return 2; }
return raise(args[1], pct);
}
if len(args) >= 2 and args[0] == "drop" { return drop_dept(args[1]); }
print_err("usage:");
print_err(" employee seed seed departments and employees");
print_err(" employee report aggregates by department");
print_err(" employee staff <department> list a department's staff");
print_err(" employee raise <department> <p> raise a department's salaries p%");
print_err(" employee drop <department> delete a department (restrict demo)");
return 1;
}
-- Inserts are WAL-logged before acknowledgment (iteration 9). A second run
-- hits the departments.name @unique index and traps; catching it here is the
-- sample's unique-violation acceptance line.
fn seed() -> Int {
let eng = try insert Department { name: "Engineering" } catch (e) nil;
if eng == nil {
print("SEED-DUP departments already seeded (unique violation caught)");
return 3;
}
let ops = insert Department { name: "Operations" };
let sales = insert Department { name: "Sales" };
insert Employee { name: "Asha", salary: 9200000, hired: 1704067200000, dept: eng };
insert Employee { name: "Bram", salary: 8100000, hired: 1706745600000, dept: eng };
insert Employee { name: "Chidi", salary: 7300000, hired: 1709251200000, dept: eng };
insert Employee { name: "Dora", salary: 6400000, hired: 1711929600000, dept: ops };
insert Employee { name: "Emil", salary: 5900000, hired: 1714521600000, dept: ops };
insert Employee { name: "Farah", salary: 8800000, hired: 1717200000000, dept: sales };
print("SEEDED 3 departments, 6 employees");
return 0;
}
-- Per-department aggregates. The group-by SYNTAX
-- from e in Employee group e by e.dept into g order by avg(g.salary) desc
-- select { dept: g.key.name, headcount: count(g), avg_salary: avg(g.salary), ... }
-- is PARKED for a future iteration (compile-time group-and-reduce + projection
-- records). Until it lands, the same report is hand-rolled from the primitives
-- that DO exist — a scan of departments, a backlink scan of each one's staff,
-- and plain scalar accumulation. Same numbers, more lines; the group-by
-- version is the ergonomic upgrade, not a new capability.
fn report() -> Int {
let payroll = 0;
for d in from x in Department order by x.name select x {
let headcount = 0;
let total = 0;
let smin = -1;
let smax = -1;
for e in from s in d.staff select s {
headcount = headcount + 1;
total = total + e.salary;
payroll = payroll + e.salary;
if smin == -1 or e.salary < smin { smin = e.salary; }
if smax == -1 or e.salary > smax { smax = e.salary; }
}
if headcount == 0 {
print("DEPT ${d.name} headcount=0 avg=nil min=nil max=nil");
} else {
print("DEPT ${d.name} headcount=${headcount} avg=${total / headcount} min=${smin} max=${smax}");
}
}
print("PAYROLL ${payroll}");
return 0;
}
-- Both navigation directions on one screen: the department found by its
-- unique-name index (a probe, not a scan — the acceptance asserts the probe
-- counter), its `staff` backlink scanned, and each employee's forward
-- `e.dept.name` printed to prove ref navigation.
fn staff(name: Text) -> Int {
let ds = from d in Department where d.name == name take 1 select d;
if len(ds) == 0 { print_err("no such department: ${name}"); return 1; }
let d = ds[0];
for e in from s in d.staff order by s.salary desc select s {
print("STAFF ${e.name} ${e.salary} (${e.dept.name})");
}
return 0;
}
-- Update through a query result; a missing department exercises the
-- empty-query path (the loop body never runs, nothing prints but the DONE).
fn raise(name: Text, pct: Int) -> Int {
let ds = from d in Department where d.name == name take 1 select d;
if len(ds) == 0 { print("RAISE ${name}: no such department (0 rows)"); return 0; }
let d = ds[0];
let n = 0;
for e in from s in d.staff select s {
e.salary = e.salary + e.salary * pct / 100;
n = n + 1;
}
print("RAISE ${name} ${pct}% applied to ${n} employees");
return 0;
}
-- Restrict is the only FK action: deleting a department that employees still
-- reference traps, and the trap code is the acceptance's assertion.
fn drop_dept(name: Text) -> Int {
let ds = from d in Department where d.name == name take 1 select d;
if len(ds) == 0 { print_err("no such department: ${name}"); return 1; }
let ok = try delete ds[0] catch (e) nil;
if ok == nil {
print("DROP ${name}: restricted (staff still reference it)");
return 4;
}
print("DROP ${name}: deleted");
return 0;
}

View file

@ -0,0 +1,24 @@
-- The two tables the whole sample exists to relate. Every class IS a table;
-- @table only configures storage (name, indexes) — the language's oldest
-- doctrine. The [dept] index serves the `staff` backlink and the restrict
-- check; [dept, salary] serves the per-department salary queries and the
-- report's ordering inside a department.
@table(name: "departments", index: [name])
class Department {
name: Text @unique
-- Not a stored column: the declared inverse of Employee.dept. Reading
-- `d.staff` is a secondary-index scan of employees.dept and yields
-- `multi Employee`.
staff: backlink Employee.dept
}
@table(name: "employees", index: [dept], index: [dept, salary])
class Employee {
name: Text
salary: Int -- cents; sum wraps like all language arithmetic
hired: Int -- epoch ms
dept: ref Department -- FK: stored as the department's row id, checked
-- by a primary-index probe on insert/update
}

View file

@ -0,0 +1,28 @@
name = "employee"
version = "0.1.0"
description = "Employee management — iteration 9/9b acceptance workload: @table storage, ref/backlink relations, compiler-checked GroupBy queries"
[runtime]
wo = ">= 0.1"
# `woc .` builds target/employee once iterations 9 (engine) and 9b (query
# surface) land; until then this sample is the target the plans compile
# toward, sample-first like log-watcher was.
[build]
runtime = "../../../runtime/wovm"
# Iteration 9c/9d surface (target): this program OWNS its database and
# shares it. The runtime listens on `listen` beside WO_DATA; every client
# below is a grant — no registration, no attach, same uid included.
[share]
listen = "unix:target/data/employee.sock"
# Program B (docs/examples/employee-list), granted read-only. Identity is
# B's public-key fingerprint (9d): printed by `employee-list --identity`
# after B's first boot; paste it here. Rights: "read" or "rw" — B's
# probe-write mode exists to prove "read" refuses. Rotation and revocation
# are edits to this table plus a restart, never an API.
[[share.clients]]
name = "employee-list"
public_key = "ed25519:PASTE-EMPLOYEE-LIST-FINGERPRINT-HERE"
rights = "read"

View file

@ -1,182 +0,0 @@
# AI Agents and Content Management
## Context
AI agents (Claude Code, Copilot, Cursor, custom agents) work within a specific project or working directory. Their sessions, context, and understanding are scoped to that directory. This works well when projects are completely different domains.
But writeonce content is not isolated — articles reference each other, share tags, build on concepts from other articles. An agent editing `gitlab-runner-with-kubernetes-executor.md` would benefit from knowing that `auto-scale-gitlab-runner-using-aws-spot-instance.md` exists and covers related infrastructure. Without explicit mappings, the agent treats each article as an island.
## Problem
1. **Agents lack cross-article awareness.** When asked to write or update an article about Kubernetes, the agent doesn't know that related articles about Docker, CI/CD, or AWS already exist in the content directory — unless it manually searches.
2. **No semantic grouping.** Tags provide flat categorization (`kubernetes`, `ci-cd`), but they don't express relationships: "this article is a prerequisite for that one", "these three articles form a series", "this article supersedes that one."
3. **Context window waste.** Without mappings, the agent must scan all articles to find related content. With explicit mappings, it can load exactly the relevant files.
## Solution: Metadata-Driven Content Mappings
Users define relationships between articles in the JSON metadata. These mappings serve two purposes:
1. **Human navigation** — rendered as "related articles" links on the site
2. **Agent context** — when an agent works on an article, it loads the mapped articles into its context for cross-referencing
### Mapping Fields in JSON Metadata
Per [06-markdown-render.md](../06-markdown-render.md), the JSON metadata is minimal. Add a `mappings` field:
```json
{
"sys_title": "gitlab-runner-with-kubernetes-executor",
"title": "Gitlab Runner with Kubernetes Executor",
"published": true,
"author": "Shoney Arickathil",
"tags": ["kubernetes", "gitlab", "ci-cd"],
"published_on": 1740950884,
"mappings": {
"related": ["auto-scale-gitlab-runner-using-aws-spot-instance"],
"prerequisite": ["linux-misc"],
"series": {
"name": "gitlab-runner",
"order": 2
}
}
}
```
### Mapping Types
| Type | Meaning | Agent Use |
| -------------- | ----------------------------------------------- | --------------------------------------------------------------------------------------- |
| `related` | Topically related articles | Agent loads these for cross-reference when editing |
| `prerequisite` | Articles the reader should read first | Agent ensures no concept duplication, references prerequisites instead of re-explaining |
| `series` | Articles that form an ordered sequence | Agent maintains narrative continuity across the series |
| `supersedes` | This article replaces an older one | Agent can mark the old article as outdated or unpublished |
| `references` | External articles or URLs the content builds on | Agent checks links are still valid, cites them properly |
### Directory Structure with Mappings
```
content/
gitlab-runner-with-kubernetes-executor/
gitlab-runner-with-kubernetes-executor.json # metadata + mappings
gitlab-runner-with-kubernetes-executor.md # full article
auto-scale-gitlab-runner-using-aws-spot-instance/
auto-scale-gitlab-runner-using-aws-spot-instance.json
auto-scale-gitlab-runner-using-aws-spot-instance.md
linux-misc/
linux-misc.json
linux-misc.md
```
## Agent Workflows
### 1. Writing a New Article
The author asks an agent: "Write an article about deploying GitLab Runner on ECS."
The agent:
1. Scans the content directory for existing articles with tags `gitlab`, `ci-cd`, `aws`
2. Finds `gitlab-runner-with-kubernetes-executor` and `auto-scale-gitlab-runner-using-aws-spot-instance`
3. Reads their `.md` files to understand what's already covered
4. Writes the new article, referencing existing articles rather than re-explaining shared concepts
5. Suggests `mappings.related` entries for the new article's JSON
### 2. Updating an Existing Article
The author asks: "Update the Kubernetes executor article with the new runner token format."
The agent:
1. Reads the article's JSON metadata and `.md` content
2. Reads the `mappings.related` articles to check for consistency
3. Makes the update in the `.md` file
4. Checks if the change affects any prerequisite or series articles
5. inotify detects the `.md` change → store rebuilds → subscribers notified
### 3. Content Audit
The author asks: "Which articles reference outdated AWS configurations?"
The agent:
1. Loads all article metadata (the Store already indexes everything)
2. Follows `mappings` to build a dependency graph
3. Reads the `.md` files of articles tagged with `aws`
4. Identifies outdated patterns (old SDK versions, deprecated services)
5. Reports findings with links to specific articles and line numbers
### 4. Series Management
The author asks: "Add a new part to the gitlab-runner series."
The agent:
1. Finds all articles with `mappings.series.name == "gitlab-runner"`
2. Reads them in order to understand the narrative arc
3. Writes the new article continuing from where the series left off
4. Sets `mappings.series.order` to the next number
5. Updates the previous article's mappings to reference the new one
## Integration with writeonce Architecture
### Store Index
Add a mappings index alongside the existing title, date, and tag indexes:
```
data/
articles.seg
index/
title.idx
date.idx
tags.idx
mappings.idx # sys_title → related sys_titles
```
The mappings index allows efficient traversal: "give me all articles related to X" without scanning every article's JSON.
### Template Rendering
The `article.htmlx` template can render related articles:
```html
<article>
<h1>{{article.title}}</h1>
{{article.content_html}} {{#each article.related}}
<aside class="related">
<h3>Related</h3>
<ul>
<li><a href="/blog/{{sys_title}}">{{title}}</a></li>
</ul>
</aside>
{{/each}}
</article>
```
### Subscription
When a mapped article changes, subscribers to related articles can optionally be notified. If article A lists article B in `mappings.related`, and article B is updated, subscribers to article A can receive a notification that related content changed.
## Agent Configuration
For agents to use the mappings effectively, the project can include an agent instruction file (e.g., `CLAUDE.md` or `.agent/instructions.md`):
```markdown
## Content Management
- Articles are in `content/{sys_title}/{sys_title}.md`
- Metadata is in `content/{sys_title}/{sys_title}.json`
- Before writing or editing an article, read its `mappings` field and load related articles for context
- When creating a new article, suggest appropriate `mappings` based on tags and content overlap
- Maintain narrative continuity within `series` mappings
- Do not duplicate explanations that exist in `prerequisite` articles — reference them instead
```
This turns the content directory into an agent-navigable knowledge graph where the metadata provides the edges and the markdown files provide the nodes.
## queryable graph database
- traversable knowledge graphs available on RAM.
- which linux kernels, develop in C ++. User wants to learn it.

View file

@ -2,7 +2,7 @@
> **Status: ⬜ not started** — Track 1 (runtime foundations). Board: [00-status.md](../00-status.md) > **Status: ⬜ not started** — Track 1 (runtime foundations). Board: [00-status.md](../00-status.md)
**Context sources:** [`./02-event-loop-epoll.md`](./done/02-event-loop-epoll.md), [`./linux/00-linux.md`](./exploration/linux/00-linux.md) § File Watching, [`../02-recovery.md`](../02-recovery.md) § No AWS Infrastructure. **Context sources:** [`./02-event-loop-epoll.md`](./done/02-event-loop-epoll.md), [`./linux/00-linux.md`](./exploration/linux/00-linux.md) § File Watching, `../02-recovery.md` § No AWS Infrastructure.
## Goal ## Goal

View file

@ -2,7 +2,7 @@
> **Status: ⬜ not started** — Track 1 (runtime foundations); also a prerequisite of the parked UI track. Board: [00-status.md](../00-status.md) > **Status: ⬜ not started** — Track 1 (runtime foundations); also a prerequisite of the parked UI track. Board: [00-status.md](../00-status.md)
**Context sources:** [`./03-hand-rolled-http.md`](./done/03-hand-rolled-http.md), [`./linux/00-linux.md`](./exploration/linux/00-linux.md) § Efficient File Serving, [`../02-recovery.md`](../02-recovery.md). **Context sources:** [`./03-hand-rolled-http.md`](./done/03-hand-rolled-http.md), [`./linux/00-linux.md`](./exploration/linux/00-linux.md) § Efficient File Serving, `../02-recovery.md`.
## Goal ## Goal

View file

@ -2,7 +2,7 @@
> **Status: ⬜ not started (scope reduced)** — replay + ack-after-fsync + group commit landed via 09c and its follow-ups; remaining here: snapshots (`.data`), compaction, WAL rotation. Board: [00-status.md](../00-status.md) > **Status: ⬜ not started (scope reduced)** — replay + ack-after-fsync + group commit landed via 09c and its follow-ups; remaining here: snapshots (`.data`), compaction, WAL rotation. Board: [00-status.md](../00-status.md)
**Context sources:** [`./10-storage-foundations.md`](./10-storage-foundations.md), [`../runtime/database/02-wo-language.md#concurrency-model`](../runtime/database/02-wo-language.md#concurrency-model), [`../runtime/database/03-inmemory-engine.md`](../runtime/database/03-inmemory-engine.md), [`./exploration/postgresql/wal.md`](./exploration/postgresql/wal.md), [`./exploration/postgresql/buffer-and-checkpoint.md`](./exploration/postgresql/buffer-and-checkpoint.md), [`./exploration/linux/12-pwrite-fsync.md`](./exploration/linux/12-pwrite-fsync.md), [`../02-recovery.md`](../02-recovery.md). **Context sources:** [`./10-storage-foundations.md`](./10-storage-foundations.md), [`../runtime/database/02-wo-language.md#concurrency-model`](../runtime/database/02-wo-language.md#concurrency-model), [`../runtime/database/03-inmemory-engine.md`](../runtime/database/03-inmemory-engine.md), [`./exploration/postgresql/wal.md`](./exploration/postgresql/wal.md), [`./exploration/postgresql/buffer-and-checkpoint.md`](./exploration/postgresql/buffer-and-checkpoint.md), [`./exploration/linux/12-pwrite-fsync.md`](./exploration/linux/12-pwrite-fsync.md), `../02-recovery.md`.
## Goal ## Goal

View file

@ -2,7 +2,7 @@
> **Status: 🔄 in progress** — 13a ✅ shipped; 13b ✅ shipped (methods execute over RPC); 13c (LIVE push) is next; 13d ⏸ parked (frontend); 13e ⬜. Board: [00-status.md](../00-status.md) > **Status: 🔄 in progress** — 13a ✅ shipped; 13b ✅ shipped (methods execute over RPC); 13c (LIVE push) is next; 13d ⏸ parked (frontend); 13e ⬜. Board: [00-status.md](../00-status.md)
**Context sources:** [`../runtime/database/02-wo-language.md`](../runtime/database/02-wo-language.md) (schema layer, § Schema-Layer DML brace disambiguation, § Cross-Paradigm Transaction Coordinator), [`../runtime/database/04-client-api.md`](../runtime/database/04-client-api.md) (subscription engine), [`./09-concurrency-scaleout.md`](./09-concurrency-scaleout.md) (thread-per-core scale-out), [`./exploration/ui/00-overview.md`](./exploration/ui/00-overview.md) + [`./exploration/ui/01-htmlx-format-spec.md`](./exploration/ui/01-htmlx-format-spec.md) (live UI), [`../examples/pricing/`](../examples/pricing/) (the demo this phase makes real), [`../examples/ecommerce/shared/logic/checkout.wo`](../examples/ecommerce/shared/logic/checkout.wo) (the existing `fn … in txn snapshot` signature style methods reuse). **Context sources:** [`../runtime/database/02-wo-language.md`](../runtime/database/02-wo-language.md) (schema layer, § Schema-Layer DML brace disambiguation, § Cross-Paradigm Transaction Coordinator), [`../runtime/database/04-client-api.md`](../runtime/database/04-client-api.md) (subscription engine), [`./09-concurrency-scaleout.md`](./09-concurrency-scaleout.md) (thread-per-core scale-out), `./exploration/ui/00-overview.md` + `./exploration/ui/01-htmlx-format-spec.md` (live UI), [`../examples/pricing/`](../examples/pricing/) (the demo this phase makes real), [`../examples/ecommerce/shared/logic/checkout.wo`](../examples/ecommerce/shared/logic/checkout.wo) (the existing `fn … in txn snapshot` signature style methods reuse).
## Context ## Context
@ -73,7 +73,7 @@ Each lands as its own numbered plan doc (`13a-…`, `13b-…`) when ready. The b
### `13a-class-surface.md` — lexer, parser, AST, spec amendments — ✅ shipped ### `13a-class-surface.md` — lexer, parser, AST, spec amendments — ✅ shipped
`class` joins the keyword map (`crates/rt/src/lexer.rs` keyword match, ~line 202 — note `self` stays an ident per decision 4). `parse_type` (`crates/rt/src/parser.rs:120`) takes the leading keyword as a parameter and serves both constructs; `fn` members inside the body parse-and-discard through the existing brace-depth skip — the same mechanism that already swallows `on update … do { … }` triggers. `ast::TypeDecl` gains `is_class: bool`; `Catalog::from_schemas` ignores it (decision 5), so REST CRUD works the moment parsing does. Docs amended in the same change: a "Class Model" subsection in [`02-wo-language.md`](../runtime/database/02-wo-language.md) next to § Schema-Layer DML, the "Isn't OO" paragraph in [`wo-language.md`](../runtime/wo-language.md), the class line in [`writeonce-pl.md`](../writeonce-pl.md), and `just pricing` / `just pricing-demo` recipes. `class` joins the keyword map (`crates/rt/src/lexer.rs` keyword match, ~line 202 — note `self` stays an ident per decision 4). `parse_type` (`crates/rt/src/parser.rs:120`) takes the leading keyword as a parameter and serves both constructs; `fn` members inside the body parse-and-discard through the existing brace-depth skip — the same mechanism that already swallows `on update … do { … }` triggers. `ast::TypeDecl` gains `is_class: bool`; `Catalog::from_schemas` ignores it (decision 5), so REST CRUD works the moment parsing does. Docs amended in the same change: a "Class Model" subsection in [`02-wo-language.md`](../runtime/database/02-wo-language.md) next to § Schema-Layer DML, the "Isn't OO" paragraph in `wo-language.md`, the class line in `writeonce-pl.md`, and `just pricing` / `just pricing-demo` recipes.
**Exit (met):** `wo run docs/examples/pricing` parses 2 classes, serves `/api/products` CRUD; parser unit tests (`parses_class_with_methods`, `class_method_braces_do_not_truncate_body`) green; blog/ecommerce/hello unchanged. **Exit (met):** `wo run docs/examples/pricing` parses 2 classes, serves `/api/products` CRUD; parser unit tests (`parses_class_with_methods`, `class_method_braces_do_not_truncate_body`) green; blog/ecommerce/hello unchanged.
### `13b-method-execution.md` — methods over RPC — ✅ shipped ### `13b-method-execution.md` — methods over RPC — ✅ shipped
@ -88,7 +88,7 @@ The subscription registry from [`04-client-api.md`](../runtime/database/04-clien
### `13d-pricing-ui.md` — the `/pricing` screen, MVC ### `13d-pricing-ui.md` — the `/pricing` screen, MVC
The screen ships as an **MVC triplet** per [`exploration/ui/08-mvc-structure.md`](./exploration/ui/08-mvc-structure.md), built in the sub-phase sequence of [`14-mvc-ui-implementation.md`](./14-mvc-ui-implementation.md): model = the classes themselves, view = [`pricing.htmlx`](../examples/pricing/ui/pricing/pricing.htmlx) (plain htmlx, logic-free) + external [`pricing.scss`](../examples/pricing/ui/pricing/pricing.scss) (strict SCSS subset compiled at `wo build`, no external deps), controller = [`pricing.wo`](../examples/pricing/ui/pricing/pricing.wo) (`route:`/`view:`/`styles:`, `model:` bindings, `actions:` calling the 13b class methods). SSR per [`exploration/ui/01-htmlx-format-spec.md`](./exploration/ui/01-htmlx-format-spec.md), compiler glue per [`02-ui-compiler.md`](./exploration/ui/02-ui-compiler.md), and the vanilla-JS client runtime ([`03-client-runtime.md`](./exploration/ui/03-client-runtime.md)) patches the price cell when the 13c delta lands. The controller's `model:` block is the M→V binding; the watchlist narrows the subscription predicate server-side. The screen ships as an **MVC triplet** per `exploration/ui/08-mvc-structure.md`, built in the sub-phase sequence of `14-mvc-ui-implementation.md`: model = the classes themselves, view = [`pricing.htmlx`](../examples/pricing/ui/pricing/pricing.htmlx) (plain htmlx, logic-free) + external [`pricing.scss`](../examples/pricing/ui/pricing/pricing.scss) (strict SCSS subset compiled at `wo build`, no external deps), controller = [`pricing.wo`](../examples/pricing/ui/pricing/pricing.wo) (`route:`/`view:`/`styles:`, `model:` bindings, `actions:` calling the 13b class methods). SSR per `exploration/ui/01-htmlx-format-spec.md`, compiler glue per `02-ui-compiler.md`, and the vanilla-JS client runtime (`03-client-runtime.md`) patches the price cell when the 13c delta lands. The controller's `model:` block is the M→V binding; the watchlist narrows the subscription predicate server-side.
**Exit:** browser at `/pricing` shows selected products; a `set_price` commit from curl changes the price cell in every open browser without reload. **Exit:** browser at `/pricing` shows selected products; a `set_price` commit from curl changes the price cell in every open browser without reload.
### `13e-pricing-at-scale.md` — millions of readers, millions of live updates ### `13e-pricing-at-scale.md` — millions of readers, millions of live updates
@ -122,5 +122,4 @@ No new architecture — this sub-phase wires the demo to [`09-concurrency-scaleo
- [`../runtime/database/02-wo-language.md`](../runtime/database/02-wo-language.md) — schema layer the class grammar extends; transaction coordinator methods reuse. - [`../runtime/database/02-wo-language.md`](../runtime/database/02-wo-language.md) — schema layer the class grammar extends; transaction coordinator methods reuse.
- [`../runtime/database/04-client-api.md`](../runtime/database/04-client-api.md) — subscription engine 13c scopes down. - [`../runtime/database/04-client-api.md`](../runtime/database/04-client-api.md) — subscription engine 13c scopes down.
- [`./09-concurrency-scaleout.md`](./09-concurrency-scaleout.md) — the scale architecture 13e instantiates. - [`./09-concurrency-scaleout.md`](./09-concurrency-scaleout.md) — the scale architecture 13e instantiates.
- [`./exploration/ui/00-overview.md`](./exploration/ui/00-overview.md) — UI track 13d draws on.
- [`../examples/hello/main.wo`](../examples/hello/main.wo) — the minimal example whose `Revision`-trigger pattern is the declarative ancestor of methods. - [`../examples/hello/main.wo`](../examples/hello/main.wo) — the minimal example whose `Revision`-trigger pattern is the declarative ancestor of methods.

View file

@ -1,90 +0,0 @@
# 14 — MVC UI implementation: model = class, view = htmlx + scss, controller = .wo
> **Status: ⏸ parked (frontend)** — backend focus first; design stays current. Board: [00-status.md](../00-status.md)
**Context sources:** [`./exploration/ui/08-mvc-structure.md`](./exploration/ui/08-mvc-structure.md) (the design this plan implements), [`./exploration/ui/01-htmlx-format-spec.md`](./exploration/ui/01-htmlx-format-spec.md) / [`02-ui-compiler.md`](./exploration/ui/02-ui-compiler.md) / [`03-client-runtime.md`](./exploration/ui/03-client-runtime.md) (the three UI-track pieces this plan sequences, each with port sources and LOC budgets), [`./13-class-model-live-pricing.md`](./13-class-model-live-pricing.md) (the class methods controllers call: 13a/13b; the LIVE deltas views consume: 13c), [`../examples/pricing/ui/pricing/`](../examples/pricing/ui/pricing/) (the reference MVC triplet), [`reference/crates/wo-htmlx/`](../../.dev/reference/crates/wo-htmlx/) (the v1 template engine, primary port source).
## Context
[`exploration/ui/08-mvc-structure.md`](./exploration/ui/08-mvc-structure.md) locks the screen anatomy: **model** = the `class`/`type` itself, **view** = plain `.htmlx` + external `.scss`, **controller** = a `.wo` file (`route:`/`view:`/`styles:`/`model:`/`actions:`) that binds the model into the view and is the only place UI may call class methods. The UI exploration docs 01–03 already specify the htmlx engine, the `##ui` compiler, and the client runtime in implementable detail. What's missing is the build order, the two genuinely new pieces (the controller format and the SCSS subset compiler), and the wiring into the `13` class-model track. This doc is that sequence.
Everything lands in **`crates/ui`** (currently a placeholder) and small deltas to `crates/rt` — consistent with the crate inventory in [`crates/README.md`](../../crates/README.md). The deployment shape never changes: one binary serving SSR + database + API on the kernel-primitive runtime.
## Goal
`cargo run --bin wo -- run docs/examples/pricing` (after plan 13a–13c land) serves `GET /pricing` as styled SSR HTML; clicking ☆ dispatches a controller action; an Ops `set-price` action calls `Product.set_price`, the commit pushes a delta, and the price cell patches in every open browser without reload — the [plan 13d exit criterion](./13-class-model-live-pricing.md), implemented MVC-shaped.
## Dependency graph
```
14a htmlx engine ──────┬─→ 14c controller format ─→ 14d SSR routes ─→ 14e actions ─→ 14f live patch
14b scss compiler ─────┘ │ │ │
(14a ∥ 14b — no shared code) needs engine needs 13a+13b needs 13c
(exists today)
```
Phases 05/06 (hand-rolled JSON / bespoke error) are orthogonal: `crates/ui` adopts `serde`/`serde_json` per the ui/01 decision and migrates when 05 lands. Phase 08 (`sendfile`) upgrades static-asset serving in 14d when it arrives; 14d ships with plain buffered writes first.
## Sub-phase sequence
### `14a-htmlx-engine.md` — port the view engine into `crates/ui`
Execute [`exploration/ui/01-htmlx-format-spec.md`](./exploration/ui/01-htmlx-format-spec.md) as written: port `reference/crates/wo-htmlx` (585 LOC — `parser.rs`, `ast.rs`, `value.rs`, `registry.rs`, `render.rs` carried over per its table) into `crates/ui/src/htmlx/`, extend with `<wo:live>` structured nodes, `wo:bind` capture, and the `data-wo-manifest` JSON emitter (~250 LOC new). One addition beyond the 01 spec, from the MVC design: `<wo:live source="…">` records whether `source` is a bare name (controller model binding, resolved in 14d) or an inline query — a one-field change to `LiveSubscription`.
**Exit:** the 01 spec's criteria — `cargo build -p ui` green, golden parse+render for every `.htmlx` under `docs/examples/{blog,ecommerce}` **plus** [`pricing/ui/pricing/pricing.htmlx`](../examples/pricing/ui/pricing/pricing.htmlx), manifest matches the 01 schema.
### `14b-scss-subset.md` — the stylesheet compiler
New, no port source (~400 LOC at `crates/ui/src/scss/`): scanner → rule tree → flattener. Exactly the subset locked in [08-mvc decision 3](./exploration/ui/08-mvc-structure.md): `$variables`, nesting (including `&`-less descendant flattening), `@use "partials"` (`ui/styles/_*.scss`), comments. **No mixins, functions, `@extend`, or color math** — `rgba($accent, 0.06)` in the reference file compiles by literal substitution of `$accent` and is the only function-form supported. Output is one flat `.css` per screen, written to `target/wo/<app>/static/`.
**Exit:** [`pricing.scss`](../examples/pricing/ui/pricing/pricing.scss) → golden-file CSS; unknown construct = compile error naming file:line (never silent passthrough); runs standalone (`wo build` integration is 14d).
### `14c-controller-format.md` — parse the controller, keep the shorthand
The `Kind::HashHash("ui")` skip arm in `crates/rt/src/parser.rs` (L80–88) parses for real, into one of two IRs by key-shape dispatch:
- **Controller form** (`route:`/`view:`/`styles:`/`model:`/`actions:` — the MVC triplet's `pricing.wo`) → new `Controller` IR: route pattern, view/styles paths resolved relative to the screen directory, `model:` entries as named query strings (`LIVE` flag captured, execution deferred), `actions:` entries as `(name, params, target-method-or-fn, role-set)`.
- **Shorthand form** (`source:`/`columns:`/… — the existing ecommerce/blog screens) → the `Screen` IR of [`exploration/ui/02-ui-compiler.md`](./exploration/ui/02-ui-compiler.md), whose codegen emits a generated view + a synthesized `Controller` — [08-mvc decision 5](./exploration/ui/08-mvc-structure.md): the shorthand is sugar over the triplet, one downstream path.
`##app` and `##component` keep their current skip behaviour (owned by ui/05 and ui/02 respectively).
**Exit:** `pricing.wo` parses to a `Controller` with 2 model bindings + 3 actions; every existing `##ui` screen in blog/ecommerce parses to `Screen` and compiles to an `.htmlx` that 14a round-trips; a controller naming a missing view file is a compile error.
### `14d-ssr-routes.md` — the single binary serves the screen
Wire controllers into `crates/rt`'s router (`crates/rt/src/server.rs`): each `Controller.route` becomes a GET route; the handler resolves `model:` bindings against the in-process engine (snapshot `select` now — the `LIVE` flag additionally registers a 13c subscription when that phase is live), renders the view via 14a with the model names as root scope, and emits HTML + manifest + `<link href="/static/<screen>.css">`. `wo build`/`wo run` gain the asset step: compile SCSS (14b), bake `/_wo/runtime.js` (`include_bytes!`, per ui/03 decision 6), serve `target/wo/.../static/` with buffered writes (upgraded to `sendfile` when [phase 08](./08-sendfile-static-assets.md) lands).
**Exit:** `GET /pricing` returns styled SSR HTML with a valid manifest and resolvable CSS/JS links; `GET /` on the blog sample is unaffected; route table printed at boot includes UI routes alongside REST.
### `14e-action-dispatch.md` — controller actions call class methods
`wo:action` buttons POST to `/_wo/action/<screen>/<action>` with `wo:args` + form payload. The dispatcher looks up the controller's action table, enforces the `role:` set server-side (per-app policy model of [`exploration/ui/07-per-app-policies.md`](./exploration/ui/07-per-app-policies.md)), and invokes the target: a class method via the 13b row-scoped RPC path (`Product{ id == id }.set_price(amount)`) or a free `fn`. Response is 204 — the UI never re-renders from the action response; the visible change arrives as a 13c delta, keeping one update path.
**Requires:** 13a + 13b. **Exit:** the ☆/★ watch toggle round-trips; `set-price` with an Ops session commits a Price; without the role it's 403 and no transaction starts.
### `14f-live-patching.md` — the browser follows commits
Execute [`exploration/ui/03-client-runtime.md`](./exploration/ui/03-client-runtime.md) as written (~500 LOC vanilla JS at `crates/ui/assets/wo-runtime.js`, JSON frames over one WebSocket, targeted DOM patching by `data-key` + `wo:bind`, coalescing backpressure, snapshot resync on reconnect), pointed at the 13c subscription endpoint.
**Requires:** 13c. **Exit:** the plan-13d criterion — `set_price` via curl in one terminal, the price cell changes in every open `/pricing` browser without reload; kill the server, restart, the page resyncs on reconnect.
## Verification targets (after 14f)
| Check | Target | How |
| --- | --- | --- |
| Golden corpus | every `.htmlx` in blog/ecommerce/pricing parses + renders byte-stable | `cargo test -p ui` golden files |
| SCSS | `pricing.scss` → golden CSS; errors carry file:line | `cargo test -p ui scss` |
| SSR | `GET /pricing` < 5 ms p99 on the dev box (RAM engine, no I/O on read path) | scripted curl loop |
| End-to-end | 13d criterion green | two-terminal demo, scripted in `just pricing-demo` |
| Single binary | UI + DB + API + WS in one `wo build` output, no Node anywhere | `ldd` shows libc only; no build-step JS |
| Dep budget | `crates/ui`: `serde`/`serde_json` only (dropped when phase 05 lands) | `Cargo.toml` review |
## Non-scope
- **No SPA router, no client-side templates.** Navigation is full-page SSR; only `wo:bind` cells and `<wo:live>` subtrees mutate in place. (ui/00 decision; unchanged.)
- **No SCSS mixins/functions/`@extend`/color math** beyond literal variable substitution — the subset is a floor, widened only by demonstrated need in the sample corpus.
- **No theme system / design tokens.** Shared partials under `ui/styles/_*.scss` are the only sharing mechanism for now.
- **No component framework.** `##component` partials render server-side via the 14a engine; they have no client behaviour beyond inherited `wo:bind` sites.
- **No changes to the REST API surface.** UI routes live beside `/api/*`; nothing under `/api` changes shape in this plan.
## Cross-references
- [`./exploration/ui/08-mvc-structure.md`](./exploration/ui/08-mvc-structure.md) — the design; its exit criteria are satisfied by 14c/14b/14f respectively.
- [`./13-class-model-live-pricing.md`](./13-class-model-live-pricing.md) — 13a/13b gate 14e; 13c gates 14f; 13d's exit criterion is this plan's end-to-end target.
- [`./exploration/ui/00-overview.md`](./exploration/ui/00-overview.md) — the UI track's master frame (per-app binaries, shared DB daemon) that 14d's asset/serving choices stay compatible with.
- [`reference/crates/wo-htmlx/`](../../.dev/reference/crates/wo-htmlx/) — primary port source (585 LOC), per ui/01.
- [`../examples/pricing/ui/pricing/`](../examples/pricing/ui/pricing/) — the reference triplet every sub-phase tests against.

View file

@ -0,0 +1,226 @@
# `@table`, Relations, Query — Implementation Plan (employee sample)
> **Status: ⬜ pending** (story iteration 9b) — blocked on iteration 9's engine
> plan ([`2026-08-01-db-engine-binding.md`](../../superpowers/plans/2026-08-01-db-engine-binding.md)):
> Tasks 3–6 below consume its row storage, WAL, indexes and select subset.
> Board: [00-status.md](../../00-status.md)
> **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (`- [ ]`) syntax for tracking.
>
> **Style rule (user convention):** concept, reason, and required behavior in words only; the executor writes the code.
**Spec:** [`docs/superpowers/specs/2026-08-15-table-relations-query-design.md`](../../superpowers/specs/2026-08-15-table-relations-query-design.md) (normative: the settled forks, the clause grammar, the aggregate semantics table, the lowering model), amended by the systems-track spec's program-mode contract for the sample's CLI.
**Goal:** `@table` classes queried **in the language**: comprehension queries
with `where`/`join`/`group`/`order`/`take`/`select`, typed `ref`/`backlink`
navigation with restrict integrity, and the LINQ aggregate vocabulary
(`count`/`sum`/`avg`/`min`/`max`, `GroupBy` as group-and-reduce) — all lowered
to bytecode loops over engine cursor builtins, proven by a new
`docs/examples/employee` sample whose acceptance script is the gate.
**Architecture:** compiler front (`compiler/src/{lexer,parser,ast,types,owner,emit}.ml`)
for the query surface; the database engine (`database/src/` — its own
top-level directory per the 2026-08-15 decision, statically linked into wovm;
`table.c`/`db.c` from iteration 9, plus a new `query.c` for cursors and group
hashing). The compiler
is the planner: index selection happens at lowering, the VM never sees a plan
tree. Reference semantics: System.Linq for operator meaning, PostgreSQL's
nodeAgg/ri_triggers for execution and integrity vocabulary (both surveyed in
the spec, fork 3).
**Tech Stack:** OCaml stdlib (compiler), C11 libc (runtime). No new opcodes;
new builtin ids appended to the format doc.
## Global Constraints
- **The sample is the test.** `docs/examples/employee` plus its acceptance
script is 9b's gate; the only corpus additions are the db-corpus fixtures
iteration 9 already plans. `just oop-e2e`, `just woc-test`, `just wovm-test`
stay green as regression after every task.
- **Index doctrine** (iteration 9, verbatim): secondary indexes are maintained
only through the engine's row choke points; FK checks and backlink reads are
index probes, never storage walks.
- **No SQL text anywhere** — the 9b acceptance's disassembly criterion.
- **No function values.** Every predicate/projection is expression syntax; a
query value is never deferred or passed around.
- Plans/specs in words; commits local only, never push; docs under `docs/`;
CODE-LOGIC.md updated beside changed code.
## File Structure
```
compiler/src/lexer.ml parser.ml ast.ml query-expression grammar, clause AST (Task 1)
compiler/src/types.ml range/group scopes, navigation, projection synthesis (Task 2)
compiler/src/owner.ml emit.ml query ownership + lowering to cursor builtins (Task 5)
database/src/query.c query.h cursors, group hash, transition/finalize aggregates (Tasks 3, 4)
database/src/db.c table.c FK restrict checks at the row choke points (Task 3)
docs/examples/employee/ the acceptance workload (Task 6)
scripts/employee-accept.sh the gate (Task 6)
docs/plan/oop-vm/00-wob-format.md appended builtin ids (Tasks 3-5)
```
---
### Task 1: Query expressions parse
**Concept & reason:** the comprehension grammar from spec section 3 —
`from`/`where`/`join`/`group…into`/`order by`/`take`/`skip`/`select`, clause
order fixed, aggregates as clause functions — becomes lexer keywords (contextual,
so `from`/`group`/`order` stay legal identifiers outside a query), a clause AST,
and parse diagnostics in a new WO-E5xx range (clause out of order, missing
`select`, aggregate named outside a query). Fixed clause order is a parse rule,
not a type rule, so the error lands on the exact token.
- [ ] Grammar and AST for every clause; contextual keywords verified against
the existing samples (no `.wo` file in the tree breaks).
- [ ] WO-E5xx diagnostics with golden coverage in the existing `woc-test`
suite (dump-ast goldens for well-formed queries, diagnostic goldens for
each malformed shape).
- [ ] Gates green; commit locally.
### Task 2: Queries typecheck
**Concept & reason:** the surface's whole promise is compile-time checking.
`from e in Employee` opens a scope where `e`'s fields are the class's; `ref`
navigation substitutes the target class's field set (`e.dept.name`);
`backlink` reads type as `multi` of the source class; `group … into g` closes
the range scope and opens the group scope, where `g.key` has the key's type
and `g.f` is legal **only** inside an aggregate call; `select { … }`
synthesizes an anonymous record type so later use of a dropped column is a
compile error. Aggregate result types follow the spec's table exactly —
`count` is Int, `sum` Int, `avg`/`min`/`max` are `?T` because an empty group
is data. Unknown table, unknown column (naming the class), navigation through
a non-`ref`, bare group-member column, join sides not separable: each is its
own WO-E5xx with a golden.
- [ ] Scope machinery for range/group variables; navigation typing both
directions; projection record synthesis interned like other class
shapes.
- [ ] Aggregate typing per the spec table; `?T` results force the existing
nil-handling style at use sites.
- [ ] Diagnostic goldens for every error class named above; gates green;
commit locally.
### Task 3: Cursors and integrity in the engine
**Concept & reason:** the runtime side queries need, built on iteration 9's
storage: a table-scan cursor (open by class id, advance, borrowed row view),
a primary-index point read (`ref` navigation), a secondary-index cursor with
longest-prefix probe (backlink reads, indexed `where`), all as builtins
appended to the format doc. Integrity lands at the row choke points where the
indexes already live: inserting/updating a non-nil `ref` probes the referenced
primary index and traps on a miss (foreign-key violation, sibling of the
unique trap); deleting a row still referenced traps (restrict) via the same
secondary index a `backlink` requires — one probe, no new structure. Nil
`ref` never probes (the MATCH SIMPLE rule); an update that leaves the key
unchanged skips the check (both from the PostgreSQL RI survey, mechanism
discarded — a check is an index probe, never query text).
- [ ] Cursor builtins + borrowed-row-view lifetime rules written into the
binding doc (a row view never escapes the loop that opened the cursor —
the ownership pass enforces it, mirror of the container-read borrow).
**Rows have no borrow word** (they share the VM's field encoding, not
its header), so unlike every VM-heap borrow there is no runtime trap
behind this rule — the compile-time check is load-bearing alone, which
is why it gets its own diagnostic and goldens rather than riding on
E30x.
- [ ] Cursor stability per spec section 6: scans **materialize their id list**
before the body runs and point-read per iteration; updates through the
row view stay legal (exclusive row borrow, index maintenance at the row
API); `insert`/`delete` targeting a table with an open cursor is a new
WO-E5xx (the ownership pass carries the open-cursor table set through
the loop body); read-only nested queries over the same table stay legal.
The `raise` mode — updating an indexed column mid-scan — is the fixture
that proves the materialized-id semantics.
- [ ] FK trap + restrict trap wired through the row API; a debug-build probe
counter exposed for Task 5's index-selection proof.
- [ ] The `delete` statement (point delete of a row value) lands here too:
iteration 9's subset is insert/select/update-point, and restrict has
nothing to restrict without it — same typed-AST-plus-builtin shape as
insert, WAL remove record already specified by iteration 9's Task 2.
- [ ] db-corpus fixtures from iteration 9's plan extended with one FK-violation
and one restrict fixture (trap-code exact); ASan green; commit locally.
### Task 4: Group hash + aggregate execution
**Concept & reason:** the `AggregateBy` shape from the spec — group-and-reduce
in one pass, no group objects. A group hash table builtin set: create (keyed
by the group key's kind, nil legal), upsert-advance (locate-or-create the
group, advance each aggregate's transition slot), drain (iterate groups,
finalize, hand key + finals to the compiled projection loop). Transition/
finalize is PostgreSQL's split: `avg` carries sum+count in its transition
state and divides at finalize; the "no row seen yet" state is distinct from
"transition value is nil" so nil-skipping aggregates need no first-row special
case. Aggregate semantics are the spec's table: empty ⇒ nil (or 0 for
count/sum), nil elements skipped, sum wraps, avg truncates.
- [ ] Builtins implemented; the input projected to only the columns the
aggregates read before hashing (the nodeAgg memory lesson).
- [ ] Fixture-level verification through iteration 9's db corpus (one grouped
query, exact output) plus the sample in Task 6; ASan green; commit
locally.
### Task 5: Lowering — the compiler is the planner
**Concept & reason:** a query desugars to the bytecode loops the language
already has. Scan or probe chosen at compile time: leading `where` equality
conjuncts matched against declared indexes longest-prefix-first, probe emitted
on a hit, scan otherwise — and the choice is **demonstrated** via Task 3's
probe counter in the acceptance, not assumed. `join` builds a transient hash
on the inner side and probes with the outer (nil never inserted, never
probes). `group` emits the two-phase hash-aggregation loop from Task 4.
`order by` materializes and stable-sorts (original-index tiebreak — the LINQ
stability guarantee); `take`/`skip` slice the result. Ownership: row views
stay loop-bound borrows; everything `select` emits is copied/built at the
boundary under the established Text-copy and fresh-value rules — queries add
no new ownership classes, and the ownership pass's existing drop machinery
covers the query's temporaries because the lowering IS ordinary loops.
- [ ] Desugar + lowering for every clause; disassembly of the sample's report
mode shows loops and builtins, no plan tree, no text. The select
boundary is the ownership bulkhead (spec section 6): everything a query
returns is copied or freshly built, so no value anywhere points into a
row slab after the query ends — asserted under ASan by mutating rows
after a query and re-reading the query's results.
- [ ] Index selection proven: the acceptance asserts probe-counter deltas for
the indexed `staff <department>` path versus a full-scan query.
- [ ] `oop-e2e`, `woc-test`, ASan-corpus gates green; commit locally.
### Task 6: The employee sample is the acceptance
**Concept & reason:** spec section 7 verbatim — `docs/examples/employee` with
`Department`/`Employee` (`@table`, `@unique` name, composite `[dept, salary]`
index, `ref`/`backlink` pair), wo.toml manifest so `woc .` builds it, a `just`
module beside it (the log-watcher convention), and `scripts/employee-accept.sh`
as the gate: `seed` (insert + WAL + duplicate-department trap on the second
run), `report` (headcount/avg/min/max by department, total payroll, ordered by
average salary — byte-exact lines), `staff` (both navigation directions +
index-probe proof), `raise` (update through a query; missing department ⇒ the
empty-query nil path), the restrict-trap demo, and a kill -9 between `seed`
and `report` proving replay on this workload.
- [ ] Sample source pre-authored 2026-08-15 (`docs/examples/employee/` —
the target workload, sample-first like log-watcher was); this task makes
`woc docs/examples/employee` compile it with zero diagnostics and the
binary's modes run. The sample is authoritative: divergence between it
and the spec is resolved in the spec's favor and committed.
- [ ] `scripts/employee-accept.sh` + `just employee` module: every mode
checked with exact expectations, trap codes asserted, crash step
included; soak-style RSS/fd sampling reused from the log-watcher
script's pattern for the `report` loop.
- [ ] ASan run of the full acceptance: zero leaks (the log-watcher bar).
- [ ] Docs: README beside the sample, CODE-LOGIC.md updates beside changed
compiler/runtime code, format-doc builtin table final, board + story
rows updated with measured numbers; commit locally.
## Out of scope — deferred by name
- Everything spec section 8 lists: set operators, outer joins, subqueries,
composite group keys, groups as values, FK cascade/set-nil, deferred
checks, SQL text in any role, cross-shard queries, `LIVE`, migrations,
cost-based planning.
- Sorted-grouping and partial-sort optimizations (recorded LINQ/postgres
precedents; hash + full stable sort are this plan's only strategies).
- The ecommerce sample's query rewrite (9b's fifth acceptance criterion) —
it lands as its own follow-up once the employee gate is green, so this
plan's blast radius stays one new sample.

View file

@ -52,3 +52,5 @@ Status board: [`00-status.md`](../00-status.md) · Doctrine: [`../00-principles.
| **Per-example `principle.md` files** | 2026-08-08: one canonical repo-level [`docs/00-principles.md`](../00-principles.md) instead; examples link to it. | | **Per-example `principle.md` files** | 2026-08-08: one canonical repo-level [`docs/00-principles.md`](../00-principles.md) instead; examples link to it. |
| **Minimal 3-file log-watcher sample** | Breaks the file-for-file `.hx` → `.wo` mapping and leaves the "could not express" column unproven — which is the sample's entire acceptance criterion. | | **Minimal 3-file log-watcher sample** | Breaks the file-for-file `.hx` → `.wo` mapping and leaves the "could not express" column unproven — which is the sample's entire acceptance criterion. |
| **Raw code in plan documents** | Plans carry concept, reason, and required behavior in words; the executor writes the code. | | **Raw code in plan documents** | Plans carry concept, reason, and required behavior in words; the executor writes the code. |
| **`##ui` / `.htmlx` LiveView frontend track** | 2026-08-17: removed the 9-doc `exploration/ui/` design set, the `14-mvc-ui-implementation` plan, and the `ui-htmlx-live` plan. All were built on the non-advancing Rust runtime (`.dev/reference/crates/wo-htmlx`, `cargo run`, WebSocket live-patches) and contradict the current woc/wovm direction. The 13d pricing-UI row went with them. Revisit only if a UI story is re-opened on the woc/wovm stack. |
| **Old-runtime "front door" + v1 design docs** | 2026-08-17: removed `writeonce-pl.md`, `runtime/wo-language.md`, `future-scope/ai-agents-content-management.md`, the numbered v1 set `02-recovery`/`03-data`/`04-ui`/`05-datalayer`/`06-markdown-render`/`07-ssl`, and `runtime/database/05-go-sdk.md`. They pitched the old Rust `wo` runtime (REST + LiveView + SQL/Cypher) as the current language and contradicted the shipped woc/wovm toolchain. The `runtime/database/` design series is kept as cited design history; the Rust-track plans/`done` are kept per the status board. |

View file

@ -2,7 +2,7 @@
> **Status: ✅ done** (Rust Stage 2 — shipped, maintained, not advancing) — `runtime/netpoll_epoll.rs`: the hand-rolled `epoll` loop that replaced the async runtime. Board: [00-status.md](../../00-status.md) > **Status: ✅ done** (Rust Stage 2 — shipped, maintained, not advancing) — `runtime/netpoll_epoll.rs`: the hand-rolled `epoll` loop that replaced the async runtime. Board: [00-status.md](../../00-status.md)
**Context sources:** [`../01-problem.md`](../../01-problem.md), [`../02-recovery.md`](../../02-recovery.md), [`./linux/00-linux.md`](../exploration/linux/00-linux.md), [`./done/01-scafolding-crates.md`](01-scafolding-crates.md). **Context sources:** [`../01-problem.md`](../../01-problem.md), `../02-recovery.md`, [`./linux/00-linux.md`](../exploration/linux/00-linux.md), [`./done/01-scafolding-crates.md`](01-scafolding-crates.md).
## Goal ## Goal

View file

@ -2,7 +2,7 @@
> **Status: ✅ done** (Rust Stage 2 — shipped, maintained, not advancing) — hand-rolled HTTP/1.1, plus keep-alive and pipelining. Board: [00-status.md](../../00-status.md) > **Status: ✅ done** (Rust Stage 2 — shipped, maintained, not advancing) — hand-rolled HTTP/1.1, plus keep-alive and pipelining. Board: [00-status.md](../../00-status.md)
**Context sources:** [`./02-event-loop-epoll.md`](./02-event-loop-epoll.md), [`./linux/00-linux.md`](../exploration/linux/00-linux.md), [`../02-recovery.md`](../../02-recovery.md). **Context sources:** [`./02-event-loop-epoll.md`](./02-event-loop-epoll.md), [`./linux/00-linux.md`](../exploration/linux/00-linux.md), `../02-recovery.md`.
## Goal ## Goal

View file

@ -146,4 +146,4 @@ Work stealing (breaks single-writer ACID), shared-heap locking (the doctrine exi
- [`../../docs/plan/09-concurrency-scaleout.md`](../../09-concurrency-scaleout.md) — the thread-per-core doctrine. - [`../../docs/plan/09-concurrency-scaleout.md`](../../09-concurrency-scaleout.md) — the thread-per-core doctrine.
- [`../../docs/plan/exploration/linux/07-io_uring.md`](../linux/07-io_uring.md), [`08-mmap.md`](../linux/08-mmap.md) — the two shared-page mechanisms. - [`../../docs/plan/exploration/linux/07-io_uring.md`](../linux/07-io_uring.md), [`08-mmap.md`](../linux/08-mmap.md) — the two shared-page mechanisms.
- [`../../docs/plan/13-class-model-live-pricing.md`](../../13-class-model-live-pricing.md) — 13e's read-replica alternative, contrasted in improvement 1. - [`../../docs/plan/13-class-model-live-pricing.md`](../../13-class-model-live-pricing.md) — 13e's read-replica alternative, contrasted in improvement 1.
- [`../../docs/writeonce-pl.md`](../../../writeonce-pl.md) — the C/assembly "one address" pedagogy this doc extends to a full runtime. - [`README.md`](../../../../README.md) — the C/assembly "one address" pedagogy the single-binary story extends to a full runtime.

View file

@ -1,6 +1,6 @@
# 02 — The end goal: the writeonce single binary on this runtime environment # 02 — The end goal: the writeonce single binary on this runtime environment
**Context sources:** [`00-plan.md`](./00-plan.md) (the runtime-environment phases, A–B ✅), [`01-architecture.md`](./01-architecture.md) (the one-address trace), [`../../../runtime/wo-language.md`](../../../runtime/wo-language.md) ("one binary per project; no runtime to install on the target host"; `.wo` has "its own lexer, parser, analyzer, and bytecode"), [`../../09-concurrency-scaleout.md`](../../09-concurrency-scaleout.md)–[`12`](../../12-engine-disk-cutover.md) (the Rust product track this proves out), [`../../../runtime/database/02-wo-language.md`](../../../runtime/database/02-wo-language.md) (catalog + transaction semantics the payload carries). **Context sources:** [`00-plan.md`](./00-plan.md) (the runtime-environment phases, A–B ✅), [`01-architecture.md`](./01-architecture.md) (the one-address trace), `../../../runtime/wo-language.md` ("one binary per project; no runtime to install on the target host"; `.wo` has "its own lexer, parser, analyzer, and bytecode"), [`../../09-concurrency-scaleout.md`](../../09-concurrency-scaleout.md)–[`12`](../../12-engine-disk-cutover.md) (the Rust product track this proves out), [`../../../runtime/database/02-wo-language.md`](../../../runtime/database/02-wo-language.md) (catalog + transaction semantics the payload carries).
## The end goal, stated once ## The end goal, stated once
@ -85,6 +85,6 @@ Steps 1–2 and 4–6 exist in `wo-rt-c` today with the notes store standing in
## Cross-references ## Cross-references
- [`00-plan.md`](./00-plan.md) — the kernel phases; [`01-architecture.md`](./01-architecture.md) — the one-address trace through the same stack. - [`00-plan.md`](./00-plan.md) — the kernel phases; [`01-architecture.md`](./01-architecture.md) — the one-address trace through the same stack.
- [`../../../runtime/wo-language.md`](../../../runtime/wo-language.md) — the user-facing single-binary promise this document implements. - [`README.md`](../../../../README.md) — the user-facing single-binary promise this document implements.
- [`../../13-class-model-live-pricing.md`](../../13-class-model-live-pricing.md) (13b methods), [`../../14-mvc-ui-implementation.md`](../../14-mvc-ui-implementation.md) (UI assets) — the payload-side tracks. - [`../../13-class-model-live-pricing.md`](../../13-class-model-live-pricing.md) (13b methods) — the payload-side track.
- [`../../09-concurrency-scaleout.md`](../../09-concurrency-scaleout.md) — shard-key routing and 2PC the contract defers to. - [`../../09-concurrency-scaleout.md`](../../09-concurrency-scaleout.md) — shard-key routing and 2PC the contract defers to.

View file

@ -1,6 +1,6 @@
## Linux Kernel Features ## Linux Kernel Features
Kernel primitives that the writeonce binary can leverage, mapped to the architectural needs identified in [docs/01-problem.md](../../../01-problem.md) and [docs/02-recovery.md](../../../02-recovery.md). Kernel primitives that the writeonce binary can leverage, mapped to the architectural needs identified in [docs/01-problem.md](../../../01-problem.md) and docs/02-recovery.md.
### Per-primitive reference cards ### Per-primitive reference cards

View file

@ -1,213 +0,0 @@
# UI track — `.htmlx` live templates + Angular-style monorepo
**Context sources:** [`docs/examples/ecommerce/ui/`](../../examples/ecommerce/ui/) (current `##ui` screens — storefront, order_tracker, admin_orders), [`docs/examples/ecommerce/types/`](../../examples/ecommerce/types/) + [`docs/examples/ecommerce/logic/`](../../examples/ecommerce/logic/) (the shared-schema + shared-fn anchor), [`reference/crates/wo-htmlx/`](../../../../.dev/reference/crates/wo-htmlx/) (v1 template engine — `{{path}}`, `{{#each}}`, `{{> partial}}`, `data-bind` attributes), [`templates/`](../../../../templates/) (v1 blog's concrete `.htmlx` usage), [`docs/runtime/database/06-lowcode-fullstack.md`](../../../runtime/database/06-lowcode-fullstack.md) (Phase 6's `##ui` + `##app` block spec).
## Context
Three threads converge into one plan:
1. **`##ui` needs a concrete output format.** Phase 6's spec says screens "compile to a render tree" served as SSR HTML with a thin client runtime, but the actual template format isn't named. The v1 `.htmlx` engine at [`reference/crates/wo-htmlx/`](../../../../.dev/reference/crates/wo-htmlx/) already speaks `{{bindings}}`, `{{#each}}`, `{{> partials}}`, and `data-bind` attributes — it's 90% of what the new runtime needs and already has a working parser + renderer. Adopting it (and extending it with live-subscription semantics) is cheaper than inventing a new format.
2. **The samples want a home that matches how real frontends are organised.** The ecommerce sample today is one flat directory with `types/`, `logic/`, and `ui/` beside each other. A real deployment has *multiple apps* against the same data: a customer storefront, an admin dashboard, a fulfillment console, maybe a read-only analytics viewer. Each has its own routes, its own policies, its own ideal binary shape. Angular (via Nx / Angular CLI workspaces) solved this with `apps/*` + `libs/*` on top of a shared root config — writeonce adopts the same shape.
3. **Each app wants to be its own binary but share a database.** Running storefront and admin as one monolith conflates concerns: a CPU spike in admin stalls customer checkout; an admin auth bug opens customer data paths. Splitting into per-app binaries that share a single database backend (via the Phase-4 native wire protocol) gives blast-radius isolation without duplicating data.
Intended outcome: `docs/examples/ecommerce/` refactors into a workspace with `shared/` + `apps/storefront/` + `apps/admin/`. `wo build apps/storefront` produces a `storefront` binary. `wo db serve` runs the shared database. The apps connect over `wo://…` and serve `.htmlx` SSR pages that subscribe to LIVE queries without a page reload.
## Goal
After this track's sub-phases land:
- A **monorepo workspace** at `docs/examples/ecommerce/` with `shared/{types,logic,components}` + `apps/{storefront,admin}` structure.
- A **per-app binary** for each app under `apps/`: `wo build apps/storefront` → `./target/wo/storefront`, `wo build apps/admin` → `./target/wo/admin`. Each binary includes only its own `##ui` / `##app` / app-local types and logic; shared code compiles in by path reference.
- A **shared DB daemon** (`wo db serve`) — one process, no UI, just the engine and wire-protocol server. Each app binary connects as a client via the Phase-4 native protocol.
- **`.htmlx` as the compiled UI output.** Every `##ui` screen compiles into an `.htmlx` template file that the app binary serves; hand-written `.htmlx` files under `apps/X/ui/*.htmlx` are accepted as a first-class authoring alternative.
- **`.htmlx` subscribes.** A `<wo:live source="...">` subtree in the template registers a LIVE query at page load; a ~20 KB client JS runtime patches DOM nodes on delta frames without reloading the page.
- **Per-app users + policies.** Each app declares its role set in `apps/X/app.wo`; row-level policies on shared types stay global, app-scoped policies layer on top per route.
## Design decisions (locked)
1. **`.htmlx` is the template format; `##ui` is the DSL that emits it.** Authors choose per-screen: declare `##ui #home { source: Product, columns: [...] }` in `.wo` and let the compiler produce `home.htmlx`; OR hand-write `home.htmlx` for a custom page. Both flow through the same `wo-htmlx` renderer.
2. **Extend v1 `.htmlx` with two new constructs** — `<wo:live source="..." key="...">...</wo:live>` (subscription subtree) and `wo:bind="field"` (field-level live binding). The rest of the v1 syntax (`{{path}}`, `{{#each}}`, `{{> partial}}`) carries through unchanged.
3. **One binary per app, shared database process.** Not a shared library, not a monolith. Apps connect via the Phase-4 wire protocol (`wo://host:port`) — the same connection a Go/TS client would use. No in-process shared state between apps; their isolation is enforced by the OS process boundary.
4. **Angular-parallel workspace layout.** `apps/*` for deployable binaries, `shared/*` for libs shared across apps (types, logic, UI components), `wo.toml` at the workspace root. Each app also has its own `wo.toml` that names which `shared/` directories it depends on.
5. **File structure mirrors Angular component organisation.** Each UI screen lives in its own directory: `apps/storefront/ui/home/{home.wo, home.htmlx, home.css}`. Tests go in `home.test.wo`. Matches the Angular component pattern (`home.component.ts`, `home.component.html`, `home.component.scss`).
6. **Per-app routes, not per-screen routes.** `apps/storefront/app.wo` declares route table; each route maps to a `ui.<screen>` declared under `apps/storefront/ui/*/`. Cross-app navigation is an external redirect, not an internal route.
7. **Policy composition.** Global policies live in `shared/types/<type>.wo` next to the `type` declaration (today). App-scoped policies live in `apps/X/app.wo` and AND with the global set — an admin app might relax a storefront policy for ops roles but can never relax beyond what the type's own policy permits.
## Angular parallels — what writeonce copies, what it doesn't
| Angular feature | Writeonce translation | Notes |
| --- | --- | --- |
| `nx workspace` / `angular.json` | Root `wo.toml` with `[workspace] apps = [...], shared = [...]` | Path references, not package registry |
| `apps/<app>/` | `apps/<app>/` with `app.wo` + `ui/` + local `types/` + local `logic/` | 1:1 naming |
| `libs/<lib>/` | `shared/<lib>/` | Used `shared/` instead of `libs/` — matches the more common monorepo idiom (Nx defaults to `libs`, but `shared` is clearer for this audience) |
| `<component>.ts` + `.html` + `.scss` | `<screen>.wo` + `.htmlx` + `.css` under `ui/<screen>/` | One-directory-per-screen |
| `ng build <app>` | `wo build apps/<app>` | Per-app binary output |
| Dependency injection | Service-block resolution — `service rest` blocks in shared types are callable from any app by import | No runtime DI container; bindings are resolved at compile time |
| RxJS observables | LIVE subscription frames on a WebSocket | Declarative `live` attribute instead of imperative `.subscribe(...)` |
| Zone.js change detection | Per-row delta dispatch + field-level `wo:bind` | No full-tree change detection — only the rows/fields the delta names get repainted |
| `HttpClient` | Built-in wire-protocol client inside the app binary | No separate library to import; always present |
### What we don't copy
- **No TypeScript.** Authoring is `.wo` (for logic) + `.htmlx` (for templates) + `.css`. If a page needs bespoke JS interactivity beyond what `wo:bind` covers, a `<script>` tag inside `.htmlx` is fine — but the client runtime itself is vanilla JS, not a framework.
- **No component library split.** Angular's `@angular/core`, `@angular/common`, etc. are a package hierarchy. Writeonce's runtime is one binary; "components" are just shared `.htmlx` partials under `shared/components/`.
- **No decorator metadata / reflect-metadata.** Rust macros + compile-time codegen do the same work.
## Reference materials
Read before writing each sub-phase:
| Source | Why |
| --- | --- |
| [`reference/crates/wo-htmlx/src/parser.rs`](../../../../.dev/reference/crates/wo-htmlx/src/parser.rs) + [`render.rs`](../../../../.dev/reference/crates/wo-htmlx/src/render.rs) | The v1 template engine's exact surface — what parses, what renders, what the AST looks like. ~500 LOC total. |
| [`templates/article.htmlx`](../../../../templates/article.htmlx), [`templates/home.htmlx`](../../../../templates/home.htmlx) | Concrete usage of the v1 format — how `{{path}}` and `data-bind` actually read in real templates. |
| [`docs/examples/ecommerce/ui/{storefront,order_tracker,admin_orders}.wo`](../../examples/ecommerce/ui/) | The `##ui` side — what the declarative DSL promises to produce. These screens are the target of the first compiler pass. |
| [`docs/runtime/database/06-lowcode-fullstack.md`](../../../runtime/database/06-lowcode-fullstack.md) | Phase 6's full-stack block spec — `##ui`, `##app`, `##policy`, `##service`, `##logic` — already designed but not yet compiled. |
| [`docs/runtime/database/04-client-api.md`](../../../runtime/database/04-client-api.md) | Phase 4's wire protocol — what the per-app binary speaks to the shared DB over. |
| [Nx monorepo docs](https://nx.dev/concepts/more-concepts/why-monorepos) | Background on the apps/libs split pattern; shape of `nx.json`. |
## Target layout
```
docs/examples/ecommerce/
├── wo.toml # workspace — lists apps, names the shared DB port
├── shared/ # imported by any apps that need it
│ ├── types/
│ │ ├── customer.wo # Customer + role union + policy read/write
│ │ ├── product.wo # Product + inventory + similar_to graph
│ │ ├── order.wo # Order + line_items + lifecycle triggers
│ │ └── purchase.wo # link Customer -> Product
│ ├── logic/
│ │ └── checkout.wo # fn checkout / mark_paid / mark_shipped
│ └── components/ # reusable .htmlx partials
│ ├── layout.htmlx # top-level page chrome
│ ├── header.htmlx
│ ├── money.htmlx # {{> money amount=x}} → $x.xx
│ └── order-row.htmlx # used by both storefront and admin
├── apps/
│ ├── storefront/ # customer-facing; no admin routes
│ │ ├── wo.toml # declares `shared = ["../shared"]`
│ │ ├── app.wo # routes: / → home, /product/:sku → product-detail
│ │ ├── logic/
│ │ │ └── cart.wo # app-local: fn add_to_cart, fn remove_from_cart
│ │ ├── types/
│ │ │ └── cart.wo # type Cart { lines: [CartLine], ... } — not shared
│ │ └── ui/
│ │ ├── home/
│ │ │ ├── home.wo # ##ui #home — declarative spec
│ │ │ ├── home.htmlx # optional hand-written override
│ │ │ └── home.css
│ │ └── product-detail/
│ │ └── product-detail.wo
│ └── admin/
│ ├── wo.toml
│ ├── app.wo # routes: /orders → orders, /inventory → inventory; gated role Admin|Ops
│ ├── logic/
│ │ └── fulfillment.wo # app-local: fn ship_order calls shared.mark_shipped
│ └── ui/
│ ├── orders/
│ │ ├── orders.wo # ##ui #admin-orders, live: true
│ │ └── orders.htmlx # hand-tuned layout overrides the auto-generated
│ └── inventory/
│ └── inventory.wo
└── tests/ # workspace-level integration
└── cross-app.test.wo # a checkout from storefront visible in admin live feed
```
Compile outputs:
```
target/wo/
├── storefront # ~12 MB static binary — app.wo compiled + ui/ templates baked in
├── admin # ~12 MB static binary
└── db # the `wo db` server (shared by all apps)
```
## `.htmlx` with live subscriptions — target format
v1 carries forward unchanged:
```htmlx
<h1>{{article.title}}</h1>
<ul>
{{#each articles}}
<li><a href="/article/{{slug}}">{{title}}</a></li>
{{/each}}
</ul>
{{> layout.footer}}
```
Two new constructs for live:
```htmlx
<!-- Subtree bound to a LIVE query; client subscribes at page load -->
<wo:live source="Order{ status != Cancelled }" sort="placed_at desc" key="id">
<table class="orders">
<thead><tr><th>#</th><th>Status</th><th>Customer</th><th>Total</th></tr></thead>
<tbody>
{{#each rows}}
<tr data-key="{{id}}">
<td>{{id}}</td>
<td wo:bind="status" class="status-{{status}}">{{status}}</td>
<td>{{customer.name}}</td>
<td>{{> money amount=total}}</td>
</tr>
{{/each}}
</tbody>
</table>
</wo:live>
```
Semantics:
- `<wo:live source="...">` emits a `LIVE <source>` query registration at SSR time. The initial result renders the `{{#each rows}}` body.
- The compiler also emits a JSON manifest (in a `<script data-wo-manifest>` tag) telling the client runtime which DOM id maps to which row key and what fields are `wo:bind`-ed.
- On page load, the client runtime opens a WebSocket back to the app, subscribes, and processes delta frames: `Insert` appends a row, `Update` finds `[data-key="<id>"]` and replaces `wo:bind`-ed cells, `Delete` removes the row.
- `wo:bind="field"` on any element tells the runtime "this element's text content reflects `row.field`"; delta Updates patch it in place.
## Sub-phase sequence
Each lands as its own plan doc under `docs/plan/ui/`. No code yet — this master plan outlines the order.
| # | File | Goal |
| --- | --- | --- |
| `01` | `01-htmlx-format-spec.md` | Nail down the exact `.htmlx` grammar — everything v1 has plus `<wo:live>` and `wo:bind`. Includes a manifest-emission spec so the client knows what to subscribe to. |
| `02` | `02-ui-compiler.md` | `##ui` → `.htmlx` compiler. Walks the parsed Phase-6 block and emits the template with the right `<wo:live>` / `{{#each}}` / `wo:bind` skeleton. Falls back gracefully when a hand-written `.htmlx` exists beside the `.wo`. |
| `03` | `03-client-runtime.md` | ~20 KB vanilla-JS runtime bundled with the app binary. Parses `<script data-wo-manifest>`, opens WebSocket, handles `snapshot`/`insert`/`update`/`delete` frames, patches DOM by `data-key` + `wo:bind`. |
| `04` | `04-workspace-layout.md` | Concrete refactor of `docs/examples/ecommerce/` from the current flat shape into `shared/` + `apps/*`. Defines `wo.toml` workspace grammar. |
| `05` | `05-per-app-binaries.md` | `wo build apps/X` produces one static binary per app. Each contains its own types/logic/ui + the shared dirs it imports. Shared DB connection is configured via `WO_DB` env var. |
| `06` | `06-shared-db-daemon.md` | `wo db serve` — headless database daemon. Per-app authn (API key per app), per-app connection scope. Apps see only types their policy allows. |
| `07` | `07-per-app-policies.md` | App-scope policy composition — global `policy read ...` on a type AND app-local `policy` in `app.wo` ⇒ effective policy = AND of both. Admin app's relaxations, storefront's restrictions. |
## Verification
After all seven sub-phases land:
| Target | Measure |
| --- | --- |
| `wo build apps/storefront` succeeds, produces one static binary | `file target/wo/storefront` → ELF, `ldd` shows only libc |
| `wo build apps/admin` succeeds | same |
| `wo db serve` + `storefront --db wo://localhost:5555` + `admin --db wo://localhost:5555` all run simultaneously | three processes, three ports, one data directory |
| Admin live orders table updates within 100 ms of a checkout on the storefront | Open `/admin/orders` in a browser, fire `POST /api/fn/checkout` against storefront's wire port, observe DOM patch |
| Customer's Order visible in their storefront order tracker but not to other customers; admin sees all | Policy round-trip |
| `wo run apps/storefront` serves hand-written `home.htmlx` if present, falls back to `##ui #home` generation if not | File-presence-based dispatch |
| `curl http://localhost:8080/healthz` from each app process | `200 ok` — standard liveness across the tracks |
## Non-scope
- **No TypeScript, no JSX.** `.htmlx` is HTML + Mustache + two `wo:` tags. The client runtime is 500 lines of vanilla JS.
- **No build-time Angular-style bundling.** No Webpack, no esbuild, no tree-shaking. The client JS is a pre-compiled static artifact inside each app binary.
- **No hot module reload in production.** `wo dev` reloads in development (inotify watches `apps/*/ui/`); production binaries don't self-reload.
- **No cross-app shared session state.** Each app authenticates independently. Shared identity is the customer row in the shared DB — both apps see the same user, but each app issues its own session token.
- **No dynamic shared-library linking between apps.** Sharing happens at source level (`shared/` dirs imported by path). Each binary is a fully-static blob.
- **No React/Vue compatibility layer.** If a downstream app wants those, they sit outside the writeonce runtime and talk to the shared DB over the wire protocol — same as any other client.
## Cross-references
- [`../09-concurrency-scaleout.md`](../../09-concurrency-scaleout.md) — when the shared DB daemon needs to handle 10k connections across multiple apps, that plan's thread-per-core model applies to the daemon process.
- [`../assembly/02-writeonce-stance.md`](../assembly/02-writeonce-stance.md) — still no asm. The client runtime is vanilla JS, no WASM.
- [`../../runtime/database/06-lowcode-fullstack.md`](../../../runtime/database/06-lowcode-fullstack.md) — Phase 6's full-stack block spec that this track implements.
- [`../../runtime/database/04-client-api.md`](../../../runtime/database/04-client-api.md) — the wire protocol per-app binaries speak to the shared DB over.
- [`../../examples/ecommerce/ui/admin_orders.wo`](../../examples/ecommerce/ui/admin_orders.wo) — the motivating workload: a live ops table bound to the order stream.
- [`reference/crates/wo-htmlx/`](../../../../.dev/reference/crates/wo-htmlx/) — the template engine ~90% of this track will reuse.
- [`templates/`](../../../../templates/) — v1 blog's actual `.htmlx` files; the format this track extends.

View file

@ -1,111 +0,0 @@
# 01 — `.htmlx` format spec
**Context sources:** [`./00-overview.md`](./00-overview.md) §§ "`.htmlx` with live subscriptions — target format" (L127–166), "Design decisions" (L28–37), [`reference/crates/wo-htmlx/`](../../../../.dev/reference/crates/wo-htmlx/) (the v1 template engine that 90% of this phase ports), [`templates/article.htmlx`](../../../../templates/article.htmlx) and [`templates/home.htmlx`](../../../../templates/home.htmlx) (v1 concrete usage), [`docs/examples/ecommerce/shared/components/order-row.htmlx`](../../../examples/ecommerce/shared/components/order-row.htmlx) (the live-binding workload this format must serve).
## Goal
Lock the exact `.htmlx` grammar — every v1 Mustache construct unchanged plus two new constructs: a `<wo:live source=… key=…>…</wo:live>` subscription subtree and a `wo:bind="field"` field-level live attribute — and define the JSON manifest emitted in a `<script data-wo-manifest>` so the client runtime (phase 03) knows which DOM nodes map to which subscription rows and fields.
## Design decisions (locked)
1. **Mustache constructs carry through unchanged.** `{{path}}`, `{{#each xs as y}}…{{/each}}`, `{{#if cond}}…{{/if}}`, `{{#when cond}}…{{/when}}`, `{{> partial arg=val}}`. The v1 parser already handles all of these; the new parser inherits them verbatim. See [`reference/crates/wo-htmlx/src/parser.rs`](../../../../.dev/reference/crates/wo-htmlx/src/parser.rs) (173 LOC) and the AST in [`ast.rs`](../../../../.dev/reference/crates/wo-htmlx/src/ast.rs) (18 LOC).
2. **`<wo:live>` is a parsed structured node, not HTML passthrough.** The parser recognises the `<wo:` prefix, captures attributes (`source`, `key`, optional `sort`, `filter`), and recursively parses the body as a normal `.htmlx` subtree. No nesting in this phase — error at parse if a `<wo:live>` contains another `<wo:live>`.
3. **`wo:bind="field"` is an HTML attribute, parsed but emitted verbatim.** SSR writes the attribute through; the consumer is the client runtime. The parser records each `(element, field)` pair into the manifest; nothing else changes about element rendering.
4. **Helpers are a closed Rust enum.** v1 invocation forms (`{{relative ts}}`, `{{#if (eq for "ops")}}`, `{{> money amount=x}}`) carry through. The registered set is fixed for this phase: `relative`, `eq`, `markdown`, `code`, `money`, `tag-chips`, `pill`, `image`, `stock-badge`, `list`. No author extensibility.
5. **Manifest is one JSON object per page.** Emitted as `<script type="application/json" data-wo-manifest>{ … }</script>` and consumed only by phase 03's client runtime. Schema below; `version: 1` is a constant for this phase.
## Scope
### New files inside `crates/ui/src/htmlx/`
| File | Responsibility | Port source |
| --- | --- | --- |
| `mod.rs` | Re-exports `Template`, `Manifest`, `LiveSubscription`, `BindSite`, `ParseError`, `RenderError` | [`reference/crates/wo-htmlx/src/lib.rs`](../../../../.dev/reference/crates/wo-htmlx/src/lib.rs) (11 LOC) |
| `ast.rs` | Adds `Node::Live { attrs, body }` and `wo_bind: Option<String>` on element nodes | [`reference/crates/wo-htmlx/src/ast.rs`](../../../../.dev/reference/crates/wo-htmlx/src/ast.rs) (18 LOC) — extend by ~50 LOC |
| `parser.rs` | Adds `<wo:` prefix recognition + attribute capture; rest unchanged | [`reference/crates/wo-htmlx/src/parser.rs`](../../../../.dev/reference/crates/wo-htmlx/src/parser.rs) (173 LOC) — extend by ~90 LOC |
| `value.rs` | Path resolution against a context Value | [`reference/crates/wo-htmlx/src/value.rs`](../../../../.dev/reference/crates/wo-htmlx/src/value.rs) (122 LOC) — copied verbatim |
| `registry.rs` | Closed helper-fn registry | [`reference/crates/wo-htmlx/src/registry.rs`](../../../../.dev/reference/crates/wo-htmlx/src/registry.rs) (121 LOC) — extend by ~60 LOC for new helpers |
| `render.rs` | Emits HTML; wraps `<wo:live>` body in `<div data-wo-subscription="…">` for the runtime | [`reference/crates/wo-htmlx/src/render.rs`](../../../../.dev/reference/crates/wo-htmlx/src/render.rs) (140 LOC) — extend by ~70 LOC |
| `manifest.rs` | Walks the AST, collects subscriptions + bind sites, serialises JSON | new (~150 LOC) |
Total: ~835 LOC (585 ported + ~250 new).
### `Cargo.toml` change
```toml
[dependencies]
serde = { version = "1", features = ["derive"] }
serde_json = "1"
```
(Both already in `crates/rt`; `crates/ui` adopts them rather than hand-rolling JSON for the manifest at this phase.)
### Manifest schema
```json
{
"version": 1,
"subscriptions": [
{
"id": "orders-live-0",
"source": "Order{ status != Cancelled }",
"key": "id",
"sort": "placed_at desc",
"filter": null,
"root_selector": "[data-wo-subscription=\"orders-live-0\"]"
}
],
"bind_sites": [
{ "subscription_id": "orders-live-0", "key": "id", "field": "status" },
{ "subscription_id": "orders-live-0", "key": "id", "field": "total" }
]
}
```
## API shape (target)
```rust
use ui::htmlx::{Template, Manifest, HelperRegistry};
let tmpl = Template::parse(src)?; // Result<Template, ParseError>
let html = tmpl.render(&ctx, &registry)?; // Result<String, RenderError>
let mani = tmpl.manifest(); // Manifest
// SSR pattern: page = head + html + "<script data-wo-manifest>" + mani.to_json() + "</script>"
```
## Exit criteria
1. `cargo build -p ui` green; no new top-level dependencies beyond `serde`/`serde_json`.
2. **Golden parse + render** for every `.htmlx` file under [`docs/examples/blog/ui/components/`](../../examples/blog/ui/components/) and [`docs/examples/ecommerce/shared/components/`](../../examples/ecommerce/shared/components/) — output is HTML and parses back into an isomorphic AST.
3. **Manifest emission** for `<wo:live source="Order{ status != Cancelled }" key="id" sort="placed_at desc">…</wo:live>` produces a `LiveSubscription` with the source string preserved verbatim and the body wrapped under `data-wo-subscription="orders-live-0"`.
4. **Bind-site collection** for `<td wo:bind="status">{{status}}</td>` inside `<wo:live>` records `(subscription_id, key="id", field="status")` once and only once.
5. **v1 regression**: `cargo run --bin wo -- run docs/examples/blog` continues to start without parser errors. The blog sample has no `<wo:live>` or `wo:bind` today; nothing should regress.
6. All 14 existing `crates/rt` tests pass.
## Non-scope
- **No SSR.** Phase 02 emits templates; phase 03 ships the runtime; serving them is a downstream concern that the per-app binary in phase 05 wires together.
- **No nested `<wo:live>`.** Parse error in this phase. Author can compose live subtrees by partial inclusion (`{{> child}}`).
- **No author-extensible helpers.** The closed enum is the contract for this phase.
- **No streaming render.** Templates are rendered to a single `String`.
- **No manifest version negotiation.** Wire format is `version: 1` always.
## Verification
```bash
cargo build -p ui
cargo test -p ui --test golden_v1 # blog/ecommerce templates byte-identical
cargo test -p ui --test golden_extensions # <wo:live> + wo:bind cases
cargo test -p ui --test manifest # manifest emission
# v1 regression
cargo run --bin wo -- run docs/examples/blog &
PID=$!; sleep 1; curl -fsS http://127.0.0.1:8080/ >/dev/null; kill $PID
cd reference/crates && cargo build && cargo test
```
## After this phase
Phase 02 (`02-ui-compiler.md`) consumes the `Template` + `Manifest` types defined here as its emission target — every `##ui` block compiles down to an `.htmlx` file that parses cleanly under this phase's parser. Phase 03 (`03-client-runtime.md`) consumes the manifest JSON schema as its wire input.

View file

@ -1,95 +0,0 @@
# 02 — `##ui` → `.htmlx` compiler
**Context sources:** [`./00-overview.md`](./00-overview.md) §§ "Sub-phase sequence" (L172–180), "Design decisions" 1–6, [`./01-htmlx-format-spec.md`](./01-htmlx-format-spec.md) (the emission target), [`../../runtime/database/06-lowcode-fullstack.md`](../../../runtime/database/06-lowcode-fullstack.md) (the `##ui` block spec), [`docs/examples/blog/ui/article_list.wo`](../../../examples/blog/ui/article_list.wo), [`docs/examples/blog/ui/article_detail.wo`](../../../examples/blog/ui/article_detail.wo), [`docs/examples/ecommerce/apps/admin/ui/orders/orders.wo`](../../examples/ecommerce/apps/admin/ui/orders/orders.wo) (the test corpus), [`crates/rt/src/parser.rs:80–116`](../../../../crates/rt/src/parser.rs) (the parse-and-discard call site to replace).
## Goal
Walk a parsed `##ui` block — every key listed in 00-overview's locked grammar (`title`, `source`, `live`, `role`, `filter`, `quick-filters`, `columns`, `sort`, `actions`, `pagination`, `refresh`, `highlight-new`, `key`, `sections`, `inputs`, `use`, `with`, `template`, `styles`) — and emit a complete `.htmlx` template that parses under phase 01's grammar. When a hand-written `<screen>.htmlx` sits beside the `.wo`, the compiler honours it and only validates that the manifest still aligns.
## Design decisions (locked)
1. **File-presence dispatch.** If `apps/<X>/ui/<screen>/<screen>.htmlx` exists alongside `<screen>.wo`, the hand-written file wins. The compiler still emits a manifest; it errors if the manifest's `bind_sites` reference fields the hand-written template doesn't expose. Anchored in [`./00-overview.md`](./00-overview.md) L24, L30, L46.
2. **`live: true` ⇒ `<wo:live>` wrapper.** The compiler emits `<wo:live source="<source>" key="<key|id>" sort="<sort.default|nothing>" filter="<filter|nothing>">` around the auto-generated `<table>`. `key` defaults to `id` when not declared.
3. **`renderer: <name>` ⇒ helper invocation by rule table.** A static rule table maps each renderer name to its emission form: `markdown` → `{{markdown <field>}}`, `money` → `{{> money amount=<field>}}`, `relative-date` → `{{relative <field>}}`, `code` / `tag-chips` / `pill` / `image` / `stock-badge` / `list` similarly. The set is closed and matches phase 01's helper registry.
4. **`actions: row-* / bulk-*` ⇒ `data-action` + `data-role` attributes.** Click → POST `/api/fn/<fn>`. Role gating is a `data-role="<set>"` attribute the runtime hides on; phase 07 wires the server-side check.
5. **Generated templates land in `target/wo/<app>/ui/<screen>.htmlx`.** Same path the per-app binary in phase 05 reads from at startup. Build artefact, not committed.
6. **`crates/rt/src/parser.rs:80–88` no longer skips `##ui`.** The `Kind::HashHash` arm parses into a `Screen` IR (this phase's new type). All other `##` blocks (`##app`, `##component`) keep their current `skip_top_level_chunk` behaviour for now — `##app` is owned by phase 05 and `##component` parses inline as a sibling of `##ui` but emits no template (it's a partial that other screens reference).
## Scope
### New files inside `crates/ui/src/compiler/`
| File | Responsibility | Port source |
| --- | --- | --- |
| `mod.rs` | Re-exports `compile_screen`, `Screen`, `Column`, `Action`, `CompileError` | new (~30 LOC) |
| `screen.rs` | `Screen` IR — every key listed in the goal section above | new (~150 LOC) |
| `codegen.rs` | Walk `Screen` → emit `.htmlx` source string | new (~280 LOC) |
| `renderers.rs` | Closed `renderer:` → helper-emission rule table | new (~120 LOC) |
| `fallback.rs` | File-presence dispatch + manifest cross-check | new (~80 LOC) |
### Modified file
| File | Change | Notes |
| --- | --- | --- |
| `crates/rt/src/parser.rs` | Replace the `Kind::HashHash(_)` skip arm at L80–88 with a real parse into `Screen` when the tag is `ui` | +60 LOC delta |
Total: ~720 LOC (all new — the v1 codebase has no `##ui` precedent to port).
### `Cargo.toml` change
`crates/ui` already depends on `serde`/`serde_json` from phase 01. No new deps.
## API shape (target)
```rust
use ui::compiler::{compile_screen, compile_app, Screen};
let screen: Screen = ql::parse_ui_block(src)?;
let template: String = compile_screen(&screen, &ctx)?; // an .htmlx string
let mani = ui::htmlx::Template::parse(&template)?.manifest();
// Whole-app pipeline used by phase 05's `wo build`:
let outputs: Vec<(PathBuf, String)> = compile_app(&app_dir)?;
for (path, src) in outputs { fs::write(path, src)?; }
```
## Exit criteria
1. `cargo build -p ui` and `cargo build -p rt` green.
2. **Compile every sample `##ui`** in the test corpus: blog `article_list.wo`, blog `article_detail.wo`, ecommerce `apps/storefront/ui/home/home.wo`, `apps/storefront/ui/orders/orders.wo`, `apps/admin/ui/orders/orders.wo`. Output template parses cleanly under phase 01.
3. **Hand-written fallback honoured.** With a hand-written `apps/admin/ui/orders/orders.htmlx` present, the compiler returns its source unchanged but still emits the manifest.
4. **Manifest cross-check fires.** Renaming `body` to `text` in a hand-written template that the `##ui` block expects under `wo:bind="body"` produces a `CompileError::HandWrittenMissingField` diagnostic.
5. **Parser change is non-breaking.** `crates/rt`'s 14 unit tests still pass; `cargo run --bin wo -- run docs/examples/blog` boots and serves REST as before.
6. `cd reference/crates && cargo build && cargo test`.
## Non-scope
- **No SSR.** This phase emits files only — the runtime in phase 05 reads them at startup.
- **No author-defined renderers.** The closed table from phase 01's helper registry is the contract.
- **No build-output caching.** The compiler runs on every `wo build`. Caching is a future concern.
- **No partial recovery.** First parse or codegen error aborts compilation for the screen; whole-app compilation reports per-screen status.
- **No `##component` codegen in this phase.** Components remain their existing partial-include shape (`{{> name args}}`) — only `##ui` screens drive new emission.
## Verification
```bash
cargo build -p ui -p rt
cargo test -p ui --test compile_blog
cargo test -p ui --test compile_ecommerce
cargo test -p ui --test fallback_handwritten
# end-to-end: emit templates for the storefront and inspect them
cargo run --bin wo -- build docs/examples/ecommerce/apps/storefront --emit-templates-only
ls target/wo/storefront/ui/ # home.htmlx, orders.htmlx
head -1 target/wo/storefront/ui/orders.htmlx # starts with <wo:live source="Order{ … }">
# v1 regression
cargo run --bin wo -- run docs/examples/blog &
PID=$!; sleep 1; curl -fsS http://127.0.0.1:8080/ >/dev/null; kill $PID
cd reference/crates && cargo build && cargo test
```
## After this phase
Phase 03 (`03-client-runtime.md`) is now unblocked: every screen has a manifest the client runtime can read. Phase 04 (`04-workspace-layout.md`) places the compiled outputs under `target/wo/<app>/ui/`, which phase 05 then bakes into the per-app binary.

View file

@ -1,109 +0,0 @@
# 03 — Client runtime
**Context sources:** [`./00-overview.md`](./00-overview.md) §§ "`.htmlx` with live subscriptions — target format" (L127–166) and decisions 1–2, [`./01-htmlx-format-spec.md`](./01-htmlx-format-spec.md) (the manifest schema this runtime consumes), [`reference/crates/wo-sub/src/lib.rs`](../../../../.dev/reference/crates/wo-sub/src/lib.rs) (the v1 frame model the wire format mirrors), [`docs/examples/ecommerce/shared/components/order-row.htmlx`](../../../examples/ecommerce/shared/components/order-row.htmlx) (the live workload the runtime must update without reload).
## Goal
Ship a ~500-line vanilla-JS client at `crates/ui/assets/wo-runtime.js` that, on page load, reads the `<script data-wo-manifest>` JSON, opens one WebSocket back to the originating app, subscribes to every declared `LiveSubscription`, and patches DOM in place on `Insert`/`Update`/`Delete` frames — without a framework, without a build step, served as a static asset at `/_wo/runtime.js`.
## Design decisions (locked)
1. **Vanilla JS, no transpiler.** The file shipped is the file written. Anchored in [`./00-overview.md`](./00-overview.md) L25, L198–199.
2. **JSON over WebSocket.** Frame schema mirrors `reference/crates/wo-sub` semantics evolved into this phase's predicate-subscription model. `{ subscription_id, kind: "snapshot"|"insert"|"update"|"delete", key, row|fields }`.
3. **Targeted DOM patching, not virtual-DOM.** `update` ⇒ `document.querySelectorAll('[data-wo-subscription="<id>"] [data-key="<k>"] [wo\\:bind="<f>"]')` ⇒ `el.textContent = row[f]`. Matches the Zone-less Angular note in 00-overview decision 9.
4. **Reconnect = full snapshot resync.** On reconnect the runtime re-subscribes and replaces each `<wo:live>` body with the fresh snapshot. No diff, no replay buffer.
5. **Backpressure = drop all but latest update per `data-key`.** A coalescing queue keyed by `(subscription_id, key)` collapses queued `update` frames; the latest wins. New frames of other kinds (`insert`/`delete`) flush the queue.
6. **Asset baked into binary, served at `/_wo/runtime.js`.** `crates/ui/src/runtime/asset.rs` does `pub const RUNTIME_JS: &[u8] = include_bytes!("../../assets/wo-runtime.js");`. Phase 05 mounts the route in the per-app binary.
## Scope
### New files
| File | Responsibility | Port source |
| --- | --- | --- |
| `crates/ui/assets/wo-runtime.js` | The runtime — reads manifest, opens WS, dispatches frames, patches DOM | new (~500 LOC JS) |
| `crates/ui/src/runtime/mod.rs` | Re-exports `RUNTIME_JS`, `runtime_etag()`, `Frame` | new (~20 LOC) |
| `crates/ui/src/runtime/asset.rs` | `include_bytes!` of the JS + sha256 ETag | new (~30 LOC) |
| `crates/ui/src/runtime/frame.rs` | Wire-frame `enum Frame` mirroring the JS schema | new (~120 LOC) |
Total: ~170 LOC Rust + ~500 LOC JS.
### `Cargo.toml` change
```toml
[dependencies]
serde = { version = "1", features = ["derive"] }
serde_json = "1"
sha2 = "0.10" # ETag for the runtime asset
```
`serde` / `serde_json` were already pulled in by phase 01. `sha2` is new for the asset ETag.
### Wire frame schema (mirrored in JS and Rust)
```json
{ "kind": "snapshot", "subscription_id": "orders-live-0", "rows": [ {…}, … ] }
{ "kind": "insert", "subscription_id": "orders-live-0", "key": "42", "row": {…} }
{ "kind": "update", "subscription_id": "orders-live-0", "key": "42", "fields": { "status": "Paid" } }
{ "kind": "delete", "subscription_id": "orders-live-0", "key": "42" }
```
## API shape (target)
```rust
use ui::runtime::{RUNTIME_JS, runtime_etag, Frame};
// Phase-05 router mounts:
router.route(Method::GET, "/_wo/runtime.js", |_req| {
Response::ok()
.header("Content-Type", "application/javascript")
.header("ETag", runtime_etag())
.body(RUNTIME_JS.to_vec())
});
// Phase-06 wire-protocol handler emits frames:
let frame = Frame::Update { subscription_id: id.into(), key: k.into(), fields: patch };
ws.send_text(serde_json::to_string(&frame)?)?;
```
## Exit criteria
1. `cargo build -p ui` green; `wc -c crates/ui/assets/wo-runtime.js` ≤ 25 600 bytes (25 KB cap, slack on the 20 KB target).
2. **Frame round-trip test.** A Rust unit test serialises one of each `Frame` variant; a Node-driven test (`node --test`) parses the same JSON, asserts shape.
3. **DOM-patch test (jsdom).** `node crates/ui/runtime-tests/run.mjs` loads a stub HTML containing one `<wo:live>` block and a manifest, fakes a WebSocket emitting `snapshot` → `insert` → `update` → `delete` frames, and asserts each patch hits the right element.
4. **Reconnect test.** Killing the fake WS triggers exponential backoff; on resume the runtime re-issues subscriptions and replaces the body with the new snapshot.
5. **Asset served.** Once phase 05 lands, `curl http://127.0.0.1:8080/_wo/runtime.js` returns the file with a stable `ETag` matching `sha256(RUNTIME_JS)`.
6. `cd reference/crates && cargo build && cargo test`.
## Non-scope
- **No browser test matrix.** jsdom is the only target; Playwright/headless-Chrome are deferred.
- **No optimistic UI.** Action buttons (`data-action="…"`) POST and wait — no client-side state mutation before the server confirms.
- **No client-side routing.** Page navigation is full-page reload.
- **No framework integration.** No React/Vue/Solid bindings.
- **No gzip / Brotli.** The JS ships uncompressed; HTTP-level compression is a phase-08-style concern.
- **No `<noscript>` fallback.** Pages with `<wo:live>` without JS show the SSR snapshot frozen.
## Verification
```bash
cargo build -p ui
cargo test -p ui --test wire_frames
# JS unit test (jsdom) — repo will need node ≥ 20
node crates/ui/runtime-tests/run.mjs
# Size budget
wc -c crates/ui/assets/wo-runtime.js
test "$(wc -c < crates/ui/assets/wo-runtime.js)" -le 25600
# v1 regression
cargo run --bin wo -- run docs/examples/blog &
PID=$!; sleep 1; curl -fsS http://127.0.0.1:8080/ >/dev/null; kill $PID
cd reference/crates && cargo build && cargo test
```
## After this phase
Phases 01 + 02 + 03 together cover the prototype demo path: a `##ui` block compiles to `.htmlx`, the SSR pass renders it with a manifest, the runtime opens a WebSocket and patches the DOM on commit. Phase 05 (`05-per-app-binaries.md`) bakes `RUNTIME_JS` into each app binary; phase 06 (`06-shared-db-daemon.md`) is the WebSocket origin that emits the frames defined here.

View file

@ -1,139 +0,0 @@
# 04 — Workspace layout + `wo.toml` grammar
**Context sources:** [`./00-overview.md`](./00-overview.md) §§ "Target layout" (L73–117), "Design decisions" 4–5 (L33–34), [`docs/examples/ecommerce/wo.toml`](../../../examples/ecommerce/wo.toml) (the workspace manifest already in the repo), [`docs/examples/ecommerce/apps/storefront/wo.toml`](../../../examples/ecommerce/apps/storefront/wo.toml) (the per-app manifest already in the repo), [`docs/examples/blog/`](../../examples/blog/) (the degenerate single-app form).
## Goal
Lock the `wo.toml` grammar at both workspace and per-app scope, fill any structural gaps in `docs/examples/ecommerce/` so its layout matches 00-overview's "Target layout" exactly, and ship a Rust workspace loader (`crates/app/src/workspace.rs` + sibling) that walks a workspace root, reads each app's manifest, and resolves `{{> partial}}` lookups across `shared = […]` directories first-match-wins.
## Design decisions (locked)
1. **`wo.toml` is TOML.** Not `.wo`. The workspace + app manifests already exist in the repo using TOML; this phase formalises the schema and adds a parser. Anchored in [`docs/examples/ecommerce/wo.toml`](../../../examples/ecommerce/wo.toml).
2. **Path-based shared dependencies, no registry.** `shared = ["../../shared/types", …]` resolves at parse time relative to the per-app `wo.toml`. No semver, no fetch. Anchored in [`./00-overview.md`](./00-overview.md) decision 4.
3. **Workspace root vs. single app, by `kind`/`app_kind` field.** A `wo.toml` with `kind = "workspace"` triggers workspace loading and reads `[workspace]`. A `wo.toml` with `app_kind = "app"` is a leaf app. A `wo.toml` with neither is a degenerate single-app workspace (the blog example) — loaded as if it were `apps/<itself>`.
4. **One screen per directory.** `apps/<X>/ui/<screen>/{<screen>.wo, <screen>.htmlx, <screen>.css}` is locked layout. The loader walks `apps/<X>/ui/*/` and registers each subdirectory as a screen. Anchored in [`./00-overview.md`](./00-overview.md) decision 5 + Target Layout (L98–103).
5. **Component resolution: app-local first, then shared in declared order.** `{{> money amount=x}}` looks in `apps/<X>/ui/components/` first, then each path in `[dependencies] shared = […]` in order. First match wins. Errors at compile time if the partial cannot be resolved.
## Scope
### New files inside `crates/app/src/`
| File | Responsibility | Port source |
| --- | --- | --- |
| `workspace.rs` | Parse workspace `wo.toml`, walk and load each app | new (~200 LOC) |
| `manifest.rs` | Per-app `wo.toml` schema (`AppManifest`) | new (~150 LOC) |
| `resolver.rs` | Component / type / logic path resolution across `shared = […]` | new (~120 LOC) |
| `mod.rs` | Re-exports `Workspace`, `App`, `AppManifest`, `WorkspaceError`, `ResolverError` | new (~30 LOC) |
Total: ~500 LOC.
### Existing files filled in (not new code)
`docs/examples/ecommerce/` — gap-fill any directories named in 00-overview's "Target layout" that don't yet exist (e.g. `shared/components/header.htmlx`, `apps/storefront/types/cart.wo`). Pure structural work; no runtime change.
### `Cargo.toml` change
```toml
[dependencies]
toml = "0.8"
```
`toml` is the explicit cost of skipping a hand-rolled TOML parser at this phase. The dependency-removal track will revisit when relevant; for prototype, parsing TOML by hand isn't worth the LOC.
### Workspace `wo.toml` schema
```toml
name = "ecommerce-workspace"
version = "0.1.0"
kind = "workspace"
[runtime]
wo = ">= 0.1"
[workspace]
apps = ["apps/storefront", "apps/admin"]
shared = ["shared/types", "shared/logic", "shared/components"]
[database]
listen = "127.0.0.1:5555"
data_dir = "./data"
isolation = "snapshot"
```
### Per-app `wo.toml` schema
```toml
name = "storefront"
version = "0.1.0"
app_kind = "app"
[dependencies]
shared = ["../../shared/types", "../../shared/logic", "../../shared/components"]
[server]
listen = ":8080"
[database]
url = "wo://127.0.0.1:5555"
api_key_env = "STOREFRONT_DB_KEY"
```
## API shape (target)
```rust
use app::{Workspace, App};
let ws = Workspace::load(Path::new("docs/examples/ecommerce"))?;
assert_eq!(ws.apps().len(), 2);
let storefront = ws.app("storefront").unwrap();
let money_partial = storefront.resolve_component("money")?; // shared/components/money.htmlx
let cart_type = storefront.resolve_type("Cart")?; // apps/storefront/types/cart.wo
// Degenerate form:
let blog = Workspace::load(Path::new("docs/examples/blog"))?; // single-app workspace
assert_eq!(blog.apps().len(), 1);
```
## Exit criteria
1. `cargo build -p app` green; the new `toml` dep compiles.
2. **Workspace load.** `Workspace::load(docs/examples/ecommerce)` returns 2 apps + 3 shared dirs and no errors.
3. **Component resolution.** `storefront.resolve_component("money")` returns the path to `shared/components/money.htmlx`. `storefront.resolve_component("nonsense")` errors as `ResolverError::NotFound`.
4. **App-local override.** Adding `apps/storefront/ui/components/money.htmlx` makes `resolve_component("money")` return the app-local path; removing it falls back to the shared one.
5. **Degenerate form.** `Workspace::load(docs/examples/blog)` loads as a one-app workspace; `wo run docs/examples/blog` continues to start unchanged.
6. `cd reference/crates && cargo build && cargo test`.
## Non-scope
- **No semver, no registry, no lockfile.** Path references only.
- **No `wo dev` hot-reload.** File watching against `apps/*/ui/` is deferred (would consume `reference/crates/wo-watch/`).
- **No cross-workspace symlinks.** `shared = […]` paths must resolve under the workspace root.
- **No build-time enforcement that an app touches only its declared shared dirs.** That's an integrity check for a later hardening phase.
- **No env-var interpolation in `wo.toml`.** `${VAR}` syntax stays out; runtime config comes through env vars at startup, not manifest time.
## Verification
```bash
cargo build -p app
cargo test -p app --test workspace_load
cargo test -p app --test resolver
# end-to-end inspection
cargo run --bin wo -- ls-apps docs/examples/ecommerce
# storefront apps/storefront
# admin apps/admin
cargo run --bin wo -- ls-apps docs/examples/blog
# blog .
# v1 regression
cargo run --bin wo -- run docs/examples/blog &
PID=$!; sleep 1; curl -fsS http://127.0.0.1:8080/ >/dev/null; kill $PID
cd reference/crates && cargo build && cargo test
```
## After this phase
The workspace layout is the input that phase 05 (`05-per-app-binaries.md`) compiles into a binary per app, and the namespace within which phase 06 (`06-shared-db-daemon.md`) issues per-app API keys. Phase 07 (`07-per-app-policies.md`) reads `apps/<X>/app.wo` for app-scope policy declarations whose location this phase locks.

View file

@ -1,124 +0,0 @@
# 05 — Per-app static binaries (`wo build apps/X`)
**Context sources:** [`./00-overview.md`](./00-overview.md) §§ "Goal" (L21–23), "Design decisions" 3–4 (L32–33), [`./01-htmlx-format-spec.md`](./01-htmlx-format-spec.md), [`./02-ui-compiler.md`](./02-ui-compiler.md), [`./03-client-runtime.md`](./03-client-runtime.md), [`./04-workspace-layout.md`](./04-workspace-layout.md), [`./06-shared-db-daemon.md`](./06-shared-db-daemon.md) (the wire URL contract this phase consumes), [`docs/examples/ecommerce/apps/storefront/wo.toml`](../../../examples/ecommerce/apps/storefront/wo.toml).
## Goal
`wo build apps/<X>` produces a single static binary at `target/wo/<X>` that contains the app's `##ui` / `##app` blocks compiled to `.htmlx`, the imported `shared/` dirs the app's `wo.toml` names, the vanilla-JS client runtime from phase 03, and a thin `main()` that reads `WO_DB`, opens a wire connection to the shared daemon, and serves the public HTTP listener declared in `[server] listen`.
## Design decisions (locked)
1. **One Cargo build per app, dynamic Cargo project templating.** `wo build apps/<X>` materialises a Cargo project under `target/wo-build/<X>/`, fills `[bin] name = "<X>"`, copies/generates `app_config.rs`, and invokes `cargo build --release`. The resulting binary is copied to `target/wo/<X>`.
2. **No per-app Rust source generation beyond config.** The same `crates/app` is linked into every app binary. The only generated Rust file is `app_config.rs` containing the route table, embedded templates, and embedded runtime asset. Avoids exploding cargo metadata across N apps.
3. **`include_bytes!` bakes templates + runtime + CSS at compile time.** A Cargo `build.rs` writes `app_config.rs` enumerating every compiled `.htmlx`, every `.css` from `apps/<X>/ui/<screen>/<screen>.css` and `shared/components/*.css`, plus the runtime JS via `RUNTIME_JS` from phase 03.
4. **Connection target precedence: `WO_DB` env > `[database].url` from manifest > error.** The app refuses to start if neither is set. Anchored in [`./00-overview.md`](./00-overview.md) "Goal" (L23) and [`docs/examples/ecommerce/apps/storefront/wo.toml`](../../../examples/ecommerce/apps/storefront/wo.toml) L21–26.
5. **HTTP listener address comes from `[server] listen`.** Different from the database URL — the database URL is what this binary connects *to*; `[server].listen` is what the binary itself exposes to browsers. `WO_LISTEN` env var overrides for ops.
## Scope
### New files
| File | Responsibility | Port source |
| --- | --- | --- |
| `crates/app/src/build.rs` | `wo build apps/<X>` driver: template Cargo project, run cargo, copy binary | new (~250 LOC) |
| `crates/app/build.rs` (Cargo build script) | Generates `app_config.rs` with embedded templates + runtime + routes | new (~100 LOC) |
| `crates/app/src/main.rs` | App-binary entrypoint: reads `WO_DB`, opens wire, starts HTTP | new (~150 LOC) |
| `crates/app/src/route_table.rs` | Compile-time route table from `app.wo` routes block | new (~120 LOC) |
| `crates/rt/src/bin/wo.rs` | Wire `wo build <path>` subcommand | modify (+40 LOC) |
Total: ~660 LOC new + ~40 LOC modified.
### `Cargo.toml` change
`crates/app` adds itself as a workspace member that produces a binary. No new external deps beyond what phases 01–04 already brought in.
```toml
[[bin]]
name = "wo-app"
path = "src/main.rs"
```
### Generated `app_config.rs` shape
```rust
// Generated by crates/app/build.rs — do not edit.
pub const APP_NAME: &str = "storefront";
pub const APP_LISTEN: &str = ":8080";
pub const APP_DB_URL: Option<&str> = Some("wo://127.0.0.1:5555");
pub const APP_API_KEY_ENV: &str = "STOREFRONT_DB_KEY";
pub const TEMPLATES: &[(&str, &[u8])] = &[
("home", include_bytes!("../target/wo/storefront/ui/home.htmlx")),
("orders", include_bytes!("../target/wo/storefront/ui/orders.htmlx")),
];
pub const ROUTES: &[(&str, &str, &str)] = &[
("GET", "/", "ui.home"),
("GET", "/orders", "ui.orders"),
];
```
## API shape (target)
```rust
// build-time API used by the wo CLI:
use app::build;
let binary_path: PathBuf = build::build(Path::new("docs/examples/ecommerce/apps/storefront"),
Path::new("target/wo"))?;
// runtime: every app binary's main() looks like this
fn main() -> Result<()> {
let db_url = env::var("WO_DB").ok()
.or(app_config::APP_DB_URL.map(str::to_owned))
.ok_or(BootError::NoDatabase)?;
let api_key = env::var(app_config::APP_API_KEY_ENV)?;
let db = wire::connect(&db_url, &api_key)?;
let r = build_router(&app_config::ROUTES, &app_config::TEMPLATES, db);
rt::http::serve(app_config::APP_LISTEN, r)
}
```
## Exit criteria
1. `cargo build -p app` green; `crates/app` produces the `wo-app` library + the `wo build` driver.
2. **Storefront builds.** `cargo run --bin wo -- build docs/examples/ecommerce/apps/storefront` produces `target/wo/storefront`. `file target/wo/storefront` reports an ELF executable.
3. **Storefront boots.** `WO_DB=wo://127.0.0.1:5555 STOREFRONT_DB_KEY=test ./target/wo/storefront &` then `curl -fsS http://127.0.0.1:8080/healthz` returns `200`. (The DB daemon from phase 06 is mocked or stubbed for this test if 06 hasn't landed yet — refuse-to-start without DB is the contract; the test verifies refuse-to-start when `WO_DB` is unset.)
4. **Admin builds separately.** `wo build apps/admin` produces a *different* binary with a disjoint route table. Diffing the two `app_config.rs` files shows different route lists.
5. **Refuse-to-start without DB.** `./target/wo/storefront` with no `WO_DB` and no manifest URL exits non-zero with a clear error.
6. `cd reference/crates && cargo build && cargo test`.
## Non-scope
- **No cross-compilation.** Linux x86_64 only this phase. `--target` flags are passed through but untested.
- **No musl static linking.** glibc-linked binaries are fine for prototype.
- **No container packaging, no systemd unit generation.**
- **No binary-size optimisation** beyond `--release`. `wo build --strip` is a future flag.
- **No incremental compile cache management.** `target/wo-build/<X>/` is reused across builds but not pruned.
## Verification
```bash
cargo build -p app
# storefront build + boot
cargo run --bin wo -- build docs/examples/ecommerce/apps/storefront
file target/wo/storefront # ELF 64-bit
ls -la target/wo/storefront/ui/ # home.htmlx, orders.htmlx (compiled in phase 02)
# refuse-to-start without WO_DB
./target/wo/storefront 2>&1 | grep -q "WO_DB"
# admin builds independently
cargo run --bin wo -- build docs/examples/ecommerce/apps/admin
test -x target/wo/admin
# v1 regression
cargo run --bin wo -- run docs/examples/blog &
PID=$!; sleep 1; curl -fsS http://127.0.0.1:8080/ >/dev/null; kill $PID
cd reference/crates && cargo build && cargo test
```
## After this phase
`wo build` produces shippable per-app binaries; what they connect to is owned by phase 06 (`06-shared-db-daemon.md`), and the policy gate they enforce on every query is owned by phase 07 (`07-per-app-policies.md`). After 06 + 05 + 07 land together, the prototype demo path closes: `wo db serve` + `wo build apps/storefront` + `wo build apps/admin` running side by side, sharing one data dir, with admin live updates triggered by storefront commits.

View file

@ -1,120 +0,0 @@
# 06 — Shared database daemon (`wo db serve`)
**Context sources:** [`./00-overview.md`](./00-overview.md) §§ "Goal" (L23), "Design decisions" 3 (L32), "Non-scope" (L201–203), [`./03-client-runtime.md`](./03-client-runtime.md) (the wire frames this daemon emits), [`./05-per-app-binaries.md`](./05-per-app-binaries.md) (the apps that connect), [`reference/crates/wo-sub/src/lib.rs`](../../../../.dev/reference/crates/wo-sub/src/lib.rs) (the v1 subscription registry, 470 LOC, that needs generalising past `ByTitle`/`ByTag`/`All`), [`../../runtime/database/04-client-api.md`](../../../runtime/database/04-client-api.md) (the wire-protocol owner).
## Goal
Stand up a headless daemon — `wo db serve` — that runs the engine + WAL + subscription registry behind the Phase-4 native wire protocol on `127.0.0.1:5555`, with no HTTP and no template rendering. Each connection presents an API key from a static table, binds a `Principal { app, roles }` for downstream policy evaluation, and can register `Subscription::ByPredicate` against any type — a generalisation of v1's article-only subscription model that this phase ports and broadens.
## Design decisions (locked)
1. **Daemon = `crates/db` thin entrypoint + `crates/engine` + the wire acceptor.** No HTTP, no `.htmlx`, no `##ui`. The shared DB process knows nothing about the UI layer.
2. **API-key table is in-memory, env-seeded.** On startup the daemon reads `WO_DB_KEY_<APP>=<hex>` for each app declared in the workspace and builds an `AuthTable: HashMap<ApiKey, Principal>`. A `--keys <file>` flag is accepted but treated as a future hook.
3. **Generalise `wo-sub`** from `Subscription::ByTitle/ByTag/All` to `Subscription::ByPredicate(TypeRef, Expr, SortKey)`. The v1 variants stay as legacy aliases (`ByTitle(t)` ⇒ `ByPredicate(Article, sys_title == t, _)`) for the blog regression test. Anchored in [`reference/crates/wo-sub/src/lib.rs`](../../../../.dev/reference/crates/wo-sub/src/lib.rs) L8–17.
4. **Connection scope = `Principal { app, roles }` stored on the connection.** Every query evaluator reads it; phase 07 wires it into policy AND-composition.
5. **One data dir, one engine, many connections.** Snapshot isolation by default (per `[database].isolation = "snapshot"` in the workspace `wo.toml`).
6. **Foreground-only this phase.** No daemonisation, no PID file, no signal handling beyond `SIGTERM` graceful shutdown. A future ops doc can add `wo db daemonize`.
## Scope
### New files
| File | Responsibility | Port source |
| --- | --- | --- |
| `crates/db/src/main.rs` | Entrypoint, arg parsing, env-key loading | new (~100 LOC) |
| `crates/db/src/server.rs` | Wire-protocol acceptor (TCP listener + per-conn handler) | new (~250 LOC) |
| `crates/db/src/auth.rs` | `AuthTable`, `Principal`, key handshake | new (~120 LOC) |
| `crates/sub/src/lib.rs` | Generalised subscription manager | port [`reference/crates/wo-sub/src/lib.rs`](../../../../.dev/reference/crates/wo-sub/src/lib.rs) (470 LOC) + ~150 new |
| `crates/sub/src/predicate.rs` | Predicate evaluation against a row (uses `crates/ql` if available, else minimal subset) | new (~150 LOC) |
Total: ~1240 LOC (470 ported + ~770 new).
### `Cargo.toml` change
`crates/db` becomes a binary target:
```toml
[[bin]]
name = "wo-db"
path = "src/main.rs"
```
Plus deps already in the workspace: `serde`, `serde_json`, optionally `libc` for the listener (matching `crates/rt`'s direction).
### Wire handshake (added to Phase 4 protocol)
```
client → server: HELLO app="storefront" api_key="<hex>"
server → client: WELCOME principal={ app, roles } | ERROR "unauthorised"
```
After `WELCOME`, frames follow the Phase-4 native protocol. Subscription-registration frames carry `ByPredicate(TypeRef, Expr, SortKey)`.
## API shape (target)
```rust
use db::{DbServer, AuthTable, Principal};
use sub::{SubscriptionManager, Subscription};
let mut auth = AuthTable::new();
auth.insert(ApiKey::from_env("WO_DB_KEY_STOREFRONT")?,
Principal { app: "storefront".into(), roles: roles!("Customer") });
auth.insert(ApiKey::from_env("WO_DB_KEY_ADMIN")?,
Principal { app: "admin".into(), roles: roles!("Admin", "Ops") });
let server = DbServer::bind("127.0.0.1:5555", Path::new("./data"), auth)?;
server.run()?; // foreground; SIGTERM exits cleanly
// inside a connection handler:
let sub = Subscription::ByPredicate(
TypeRef::new("Order"),
parse_expr("status != Cancelled")?,
SortKey::new("placed_at", SortDir::Desc),
);
let id = subs.register(conn_fd, sub)?;
```
## Exit criteria
1. `cargo build -p db -p sub` green; `wo-db` binary produced under `target/release/`.
2. **Daemon starts.** `WO_DB_KEY_STOREFRONT=aaaa WO_DB_KEY_ADMIN=bbbb cargo run --bin wo-db -- --listen 127.0.0.1:5555 --data-dir /tmp/wo-test` runs foreground and accepts `SIGTERM`.
3. **Two principals.** Two clients connect, one with each API key; each receives a distinct `Principal` in the `WELCOME` frame.
4. **Predicate subscription.** Client registers `Subscription::ByPredicate(Order, "status != Cancelled", "placed_at desc")`; the manager returns a fresh `subscription_id`; on a stub `Order` insert, the matching client receives an `insert` frame.
5. **v1 regression.** A connection running the legacy `Subscription::ByTitle("hello-world")` against the blog corpus still produces notifications via the legacy alias.
6. `cd reference/crates && cargo build && cargo test`.
## Non-scope
- **TLS.** `wo://` is plaintext this phase. TLS is its own phase later.
- **API-key rotation, revocation, expiry.** Static map only. JWT, mTLS, etc. — out of scope.
- **Multi-data-dir, replication, sharding.** One process, one data dir.
- **WAL changes.** Engine + WAL semantics inherit from the Stage-2 in-memory engine; persistent storage and crash recovery belong to the database series, not this phase.
- **Daemonisation, PID file, systemd integration.** Foreground-only.
- **Metric / structured-log emission.** Plain `eprintln!` traces only.
### Risk to flag in the doc
If `crates/ql` is too thin to evaluate `status != Cancelled` end-to-end at the time this phase lands, lock a minimal predicate subset — `==`, `!=`, `>`, `<`, `&&`, `||` against scalar fields — and document the gap explicitly. Phases that need richer predicates (graph traversals, computed fields) wait for `ql` to mature.
## Verification
```bash
cargo build -p db -p sub
# foreground daemon + two connections
WO_DB_KEY_STOREFRONT=aaaa WO_DB_KEY_ADMIN=bbbb \
cargo run --bin wo-db -- --listen 127.0.0.1:5555 --data-dir /tmp/wo-test &
DB_PID=$!
cargo test -p db --test multi_app_principals
cargo test -p sub --test predicate_subscription
kill $DB_PID
# legacy v1 path
cargo test -p sub --test legacy_by_title
cd reference/crates && cargo build && cargo test
```
## After this phase
The wire URL contract is now stable, which unblocks phase 05 (per-app binaries) connecting to `wo://127.0.0.1:5555`. Phase 07 (`07-per-app-policies.md`) hooks its `EffectivePolicy` resolver into the query path inside this daemon — every query the daemon executes carries the connection's `Principal`, which is exactly what 07's AND-composition needs.

View file

@ -1,105 +0,0 @@
# 07 — Per-app policy composition
**Context sources:** [`./00-overview.md`](./00-overview.md) decision 7 (L36) and "Goal" (L26), [`./04-workspace-layout.md`](./04-workspace-layout.md) (where app-scope policies live), [`./06-shared-db-daemon.md`](./06-shared-db-daemon.md) (where composition is evaluated), [`docs/examples/blog/types/article.wo`](../../../examples/blog/types/article.wo) and the other type files (existing global `policy read/write` blocks), [`docs/examples/ecommerce/apps/admin/ui/orders/orders.wo`](../../examples/ecommerce/apps/admin/ui/orders/orders.wo) L18 (a `role: Admin | Ops` set expression).
## Goal
Layer per-app policy scopes (declared in `apps/<X>/app.wo` and per-`##ui` `role:` clauses) on top of the global per-type policies that already live next to type declarations, enforcing `effective = global AND app_scope`. An app can narrow what its connection sees, never broaden. A static check at `wo build` time refuses any app-scope that names a role outside the type's own policy domain.
## Design decisions (locked)
1. **AND-only composition.** App-scope can narrow, never broaden. Anchored in [`./00-overview.md`](./00-overview.md) decision 7 (L36).
2. **Three carrier locations.** (a) Global `policy read/write` blocks beside `type` declarations in `shared/types/<type>.wo` — already in the language. (b) `policy` blocks inside `apps/<X>/app.wo` — new structured form parsed in this phase. (c) `role: <RoleSet>` clauses inside `##ui` blocks and `actions:` rows — already in samples (e.g. [`docs/examples/ecommerce/apps/admin/ui/orders/orders.wo`](../../examples/ecommerce/apps/admin/ui/orders/orders.wo) L18, L50–53). All three feed the same `EffectivePolicy` composer.
3. **Evaluation point = wire-protocol query handler in phase 06.** When a connection issues a query, the daemon reads `connection.principal`, fetches global + app-scope policies for the touched types, AND-composes, applies. No client-side enforcement; no compile-time inlining.
4. **`role:` is a set expression, not a single string.** `Admin | Ops` parses to `RoleSet::union(Admin, Ops)`. Anchored in [`docs/examples/ecommerce/apps/admin/ui/orders/orders.wo`](../../examples/ecommerce/apps/admin/ui/orders/orders.wo) L18.
5. **Build-time domain check.** `wo build apps/<X>` walks every `role: <set>` referenced by the app and verifies each role is one the global type policy actually defines for that type. An app cannot mention `role: Anonymous` against a type whose global policy never permits anonymous access — that's a build error, not a runtime one.
## Scope
### New files
| File | Responsibility | Port source |
| --- | --- | --- |
| `crates/policy/src/scope.rs` | `AppScope`, `EffectivePolicy`, AND composer | new (~150 LOC) |
| `crates/policy/src/parse.rs` | Parse `role: A \| B` set expressions and `policy` blocks in `app.wo` | new (~100 LOC) |
| `crates/policy/src/check.rs` | Static build-time domain check | new (~100 LOC) |
| `crates/policy/src/mod.rs` | Re-exports | new (~30 LOC) |
### Modified files
| File | Change |
| --- | --- |
| `crates/db/src/server.rs` | Plug `EffectivePolicy::compose(global, app_scope, type)` into the per-query path; pass `connection.principal` through. (+50 LOC) |
| `crates/app/src/build.rs` | Run `policy::check::domain_check(app)` before invoking cargo. (+30 LOC) |
Total: ~460 LOC (all new — no v1 precedent for app-scope policy composition).
### `Cargo.toml` change
None beyond what phases 01–06 already added.
## API shape (target)
```rust
use policy::{GlobalPolicy, AppScope, EffectivePolicy, RoleSet};
// Composition (called per query in the daemon)
let effective = EffectivePolicy::compose(&global_for(&order_type),
&app_scope_for("admin"),
&order_type);
let allowed_rows = effective.filter(rows, &principal);
// Build-time domain check (called by `wo build apps/X`)
policy::check::domain_check(&app)?; // errors if any role: in app references undefined role
// Role set parser (used by both compiler and daemon)
let rs: RoleSet = "Admin | Ops".parse()?;
assert!(rs.contains(Role::Admin));
assert!(rs.contains(Role::Ops));
```
## Exit criteria
1. `cargo build -p policy -p db -p app` green.
2. **AND composition.** Unit test `compose(global = "published == true", app_scope = "owner == $session.id", t = Article)` returns an `EffectivePolicy` whose `allows_read` is true only when *both* clauses hold.
3. **Build-time domain check fires.** A test workspace where `apps/storefront/app.wo` declares `role: Anonymous` against a type whose global policy does not define `Anonymous` — `wo build apps/storefront` exits non-zero with `PolicyDomainError`.
4. **Cross-app integration.** Two storefront customers issue the same `GET /api/orders` against the daemon; each sees only their own rows (storefront app-scope narrows global). Admin sees both. Test runs against the phase-06 daemon.
5. **v1 regression.** `wo run docs/examples/blog` boots; the global `policy read for anyone when published == true` on the blog `Article` type continues to gate anonymous reads as it does today.
6. `cd reference/crates && cargo build && cargo test`.
## Non-scope
- **Row-level write policies.** Read only this phase. Phase-6 spec mentions write composition; defer.
- **Field-level (column-level) policies.** All-or-nothing per row.
- **Audit logging of policy decisions.** No structured emission in this phase.
- **Policy versioning, migration, schema changes.** Whatever the type currently declares is the one definition.
- **Dynamic role assignment.** Roles are static per session; promotion / impersonation flows are out.
## Verification
```bash
cargo build -p policy -p db -p app
cargo test -p policy --test compose
cargo test -p policy --test domain_check
cargo test -p policy --test role_set_parse
# integration: two customers + one admin against the daemon
WO_DB_KEY_STOREFRONT=aaaa WO_DB_KEY_ADMIN=bbbb \
cargo run --bin wo-db -- --listen 127.0.0.1:5555 --data-dir /tmp/wo-pol &
DB_PID=$!
cargo test --test cross_app_policy
kill $DB_PID
# v1 regression — blog policy unchanged
cargo run --bin wo -- run docs/examples/blog &
PID=$!; sleep 1
curl -fsS http://127.0.0.1:8080/api/articles # only published rows
test -z "$(curl -fsS http://127.0.0.1:8080/api/articles | grep '"published":false')"
kill $PID
cd reference/crates && cargo build && cargo test
```
## After this phase
The seven-doc UI track is complete. With phases 01–07 implemented end-to-end, the prototype demo path closes: `wo db serve` runs the shared daemon; `wo build apps/storefront` and `wo build apps/admin` produce two binaries with disjoint route tables and separate API keys; a checkout on the storefront fires a commit that the daemon broadcasts to admin's open `<wo:live>` subscription, the admin client runtime patches the orders table in place, and the same row never crosses storefront's narrower policy back to a different customer's session. After this phase, future hardening — TLS, key rotation, write-side composition, hot reload — gets its own track.

View file

@ -1,81 +0,0 @@
# 08 — MVC structure: model = class, view = htmlx + scss, controller = .wo
**Context sources:** [`reference/writeonce-app/src/app/`](../../../../.dev/reference/writeonce-app/src/app/) (the v1 Angular app whose component anatomy this formalizes), [`./00-overview.md`](./00-overview.md) ("Angular-component-style layout" — `home/{home.wo, home.htmlx, home.css}`), [`./01-htmlx-format-spec.md`](./01-htmlx-format-spec.md) (the view grammar: Mustache + `<wo:live>` + `wo:bind`), [`./02-ui-compiler.md`](./02-ui-compiler.md), [`./03-client-runtime.md`](./03-client-runtime.md), [`../../13-class-model-live-pricing.md`](../../13-class-model-live-pricing.md) (the class methods controllers call), [`../../../examples/pricing/ui/pricing/`](../../../examples/pricing/ui/pricing/) (the reference screen).
## Goal
Every writeonce UI screen follows **Model–View–Controller**, with the same file anatomy the v1 Angular app used — but collapsed into the single binary. One screen = one directory with three files:
```
ui/pricing/
├── pricing.wo # Controller — binds the model into the view, exposes actions
├── pricing.htmlx # View — plain htmlx (Mustache + wo:live/wo:bind), no logic
└── pricing.scss # View styles — external, compiled at `wo build`
```
## The mapping, against the v1 Angular app
| MVC role | v1 Angular (`reference/writeonce-app/src/app/`) | writeonce |
| --- | --- | --- |
| **Model** | `models/article.ts` (interface) + `services/article.service.ts` (HTTP fetch) | the `class` / `type` declaration itself (`types/product.wo`). No service layer: the database is in-process, and a model binding **is** a query — `LIVE select` for push, `select` for snapshot |
| **View** | `article.component.html` + `article.component.css` | `pricing.htmlx` + `pricing.scss`. Plain markup; the only dynamic constructs are Mustache paths and `<wo:live>` / `wo:bind` from [`01-htmlx-format-spec.md`](./01-htmlx-format-spec.md) |
| **Controller** | `article.component.ts` (`@Component({templateUrl, styleUrl})`, fields, methods, `service.subscribe(...)`) | `pricing.wo` — declares `view:` / `styles:` (≈ `templateUrl` / `styleUrl`), a `model:` block (≈ component fields), and an `actions:` block whose handlers **call class methods** |
What Angular needed four layers for (interface, service, component class, template) writeonce does in three files against one runtime — there is no HTTP client between controller and model because there is no network between them.
## Design decisions
1. **The controller is declarative, like everything else in `.wo`.** It does not contain imperative rendering code; it declares *what* is bound and *which* method each action invokes. Shape:
```wo
##ui
#pricing
route: /pricing
view: pricing.htmlx -- ≈ Angular templateUrl
styles: pricing.scss -- ≈ Angular styleUrl
-- Model → View binding. Names declared here are the root scope of
-- the .htmlx file; LIVE bindings re-patch the view on every commit.
model:
products: LIVE select Product{ name, sku, prices }
watchlist: $session.watchlist
-- Controller actions: the only place UI may invoke class methods.
actions:
set-price(id, amount): Product{ id == id }.set_price(amount) role: Ops | Admin
watch(id): session.watchlist += id
```
2. **The view is plain `.htmlx`, logic-free.** Mustache paths, `{{#each}}`/`{{#if}}`, partials, helpers, `<wo:live>` subtrees, `wo:bind` attributes — nothing else. A `<wo:live source="products">` whose `source` is a bare name resolves against the controller's `model:` block (the M→V binding); an inline query in `source` remains legal for controller-less partials. Views never call methods — they raise actions (`wo:action="set-price"`), the controller dispatches.
3. **Styles are external SCSS, compiled at `wo build`.** No `<style>` blocks in views, no inline styles, one `.scss` per screen plus shared partials (`ui/styles/_*.scss`). `wo build` compiles a **strict SCSS subset** — variables, nesting, `@use` of partials; no mixins/functions in the first cut — to flat CSS served as a static asset via `sendfile` ([`../../08-sendfile-static-assets.md`](../../08-sendfile-static-assets.md)). Hand-rolled in `crates/ui` (~400 LOC scanner + nesting flattener), zero external dependencies — same stance as every other phase.
4. **One binary, unchanged.** `wo build` links the SSR renderer, the compiled views + manifest, the flattened CSS, the database engine, the REST/WS API, and the kernel-primitive concurrency runtime (epoll today, thread-per-core io_uring per [`../../09-concurrency-scaleout.md`](../../09-concurrency-scaleout.md)) into the single output. MVC changes the *source layout*, not the deployment shape.
5. **The `##ui` table shorthand survives as sugar.** The earlier declarative screen spec (`columns:` / `sort:` / `pagination:`, as in the ecommerce `home.wo`) compiles to a generated view + controller pair. Writing the triplet by hand is the general form; the shorthand is the 80% case. Either way the compiler output is identical: SSR HTML + manifest + bindings.
## Request flow
```
GET /pricing
→ router (controller route:) [crates/http]
→ controller resolves model: bindings against the engine (in-process)
→ SSR renders pricing.htmlx with model scope [crates/ui, per 01/02]
→ emits HTML + <script data-wo-manifest> + <link pricing.css>
browser action wo:action="set-price"
→ POST dispatched to the controller action
→ action calls Product.set_price(amount) [class method, plan 13b]
→ commit → delta → every <wo:live> subscriber [plan 13c]
→ client runtime patches wo:bind cells [03-client-runtime.md]
```
## Migration note
The two existing screen specs (`docs/examples/ecommerce/apps/*/ui/*/`, single-file `##ui` shorthand) stay valid under decision 5. New screens — starting with [`docs/examples/pricing/ui/pricing/`](../../../examples/pricing/ui/pricing/) — use the triplet. The v1 Angular app stays archived; its components are the *shape* reference, not a port source (the htmlx port source remains `reference/crates/wo-htmlx`).
## Exit criteria (implementation sequenced in [plan 14](../../14-mvc-ui-implementation.md), landing with plan 13d)
1. `crates/ui` resolves a controller file: `route:`/`view:`/`styles:`/`model:`/`actions:` parsed, view rendered with model scope, actions dispatched to class methods.
2. SCSS subset compiler: `pricing.scss` → flat CSS at build, golden-file tested.
3. The pricing screen works end-to-end per [plan 13d's exit criterion](../../13-class-model-live-pricing.md): a `set_price` commit patches the price cell in every open browser without reload.

View file

@ -61,6 +61,8 @@ The metadata exists for exactly one reason: `json.encode`/`json.decode` are runt
- **text/containers** — len, byte_at, print_err, starts_with, ends_with, index_of, last_index_of, substr, trim, to_lower, char_of, parse_int, split, split_ws, join, slice, pop, shift, sort, reverse, remove, key_at, val_at, multi_set. Ids 16–39; `runtime/src/builtin.c`. - **text/containers** — len, byte_at, print_err, starts_with, ends_with, index_of, last_index_of, substr, trim, to_lower, char_of, parse_int, split, split_ws, join, slice, pop, shift, sort, reverse, remove, key_at, val_at, multi_set. Ids 16–39; `runtime/src/builtin.c`.
- **the OS half** — fs.exists/list/stat/read_all/read_at/append, time.sleep/local/iso, env.get/stopping, net.listen/accept/read/write/close, proc.run. Ids 40–56; `runtime/src/sysio.c`. A member that returns a record takes that record's **class id as its last argument**, so the VM allocates what it fills without knowing any source type name. - **the OS half** — fs.exists/list/stat/read_all/read_at/append, time.sleep/local/iso, env.get/stopping, net.listen/accept/read/write/close, proc.run. Ids 40–56; `runtime/src/sysio.c`. A member that returns a record takes that record's **class id as its last argument**, so the VM allocates what it fills without knowing any source type name.
- **json** — encode (value + the value's static kind), decode (text + the class id to build). Ids 57–58; `runtime/src/json.c`. Decode yields the zero word on malformed input rather than trapping, which is what makes `json.decode(t) as T` a checked decode. - **json** — encode (value + the value's static kind), decode (text + the class id to build). Ids 57–58; `runtime/src/json.c`. Decode yields the zero word on malformed input rather than trapping, which is what makes `json.decode(t) as T` a checked decode.
- **59 `map_get_opt`** (`m[k]`'s optional read), **60 `text_copy`** (Text's ownership-boundary copy — Task 1 of the executable plan).
- **database** — **61 `db_insert`** (iteration 9, Task 3): window is R[B] = class id, R[B+1..] = one slot per **declared** field in declaration order; result R[A] = the new row's id. The loader validates the class-id slot statically (variable window: the field slots are validated at runtime by the engine against the class table). Engine failure traps `WO_T_DB`; a failed WAL commit traps `WO_T_IO` after un-applying the row. `database/src/db.c`.
**`?T` and nil.** A heap-shaped optional (`?Text`, `?Rec`, `?multi`, `?map`, `?@gc`) stores what `T` stores and spells nil as the **zero word** — every per-kind drop plan already ignores a zero slot, so `?T`'s field kind is `T`'s. A **nullable scalar** (`?Int`, `?Bool`, `?Timestamp`, `?Id`) cannot: `0` is a perfectly good `Int`, and real programs store it in a `?Int`. Its nil is therefore `WO_NIL_SCALAR` = −2^62 (not `INT64_MIN`: the compiler's own integers are 63-bit, so that value is not expressible on the emitting side). Such a field is marked `WOB_FIELD_NIL_SCALAR` in `field_class[i]`, which is how the runtime knows to write that word where it must produce absence itself — today only `json.decode` leaving a key absent, and `parse_int` on unparseable input. **`?T` and nil.** A heap-shaped optional (`?Text`, `?Rec`, `?multi`, `?map`, `?@gc`) stores what `T` stores and spells nil as the **zero word** — every per-kind drop plan already ignores a zero slot, so `?T`'s field kind is `T`'s. A **nullable scalar** (`?Int`, `?Bool`, `?Timestamp`, `?Id`) cannot: `0` is a perfectly good `Int`, and real programs store it in a `?Int`. Its nil is therefore `WO_NIL_SCALAR` = −2^62 (not `INT64_MIN`: the compiler's own integers are 63-bit, so that value is not expressible on the emitting side). Such a field is marked `WOB_FIELD_NIL_SCALAR` in `field_class[i]`, which is how the runtime knows to write that word where it must produce absence itself — today only `json.decode` leaving a key absent, and `parse_int` on unparseable input.

View file

@ -0,0 +1,137 @@
# DB binding — row format, id discipline, WAL layout, query subset
> Normative companion to the engine plan
> ([`2026-08-01-db-engine-binding.md`](../../superpowers/plans/2026-08-01-db-engine-binding.md)),
> the way `00-wob-format.md` is normative for the image. Grows with the
> plan's tasks; this revision covers **Task 1 (row storage)**. Memory-safety
> doctrine lives in the 9b design's section 6 (the copy bulkhead) — this doc
> is the *format*.
## Two memory worlds, one crossing rule
Rows store **no VM pointer**, ever. Values cross from VM heap to row storage
by copy on insert, and back by copy on read (`wo_row_read` allocates fresh VM
values from the shard's runtime). The engine's own allocations are plain
malloc — never the VM arena, so table growth cannot eat the program's heap
cap, and a heap-exhausted program can still read its data.
## Row format
```
row := header slots
header := id u64 | class_id u32 | flags u32 (16 bytes)
slots := field_cnt × u64, declaration order (the VM object shape)
```
One 8-byte slot per field, kind-driven — the same kind bytes the `.wob`
class table carries, walked the same way the VM walks them:
| kind | slot holds | engine-owned shape |
| --- | --- | --- |
| `SCALAR` | the 8 bytes themselves | — (`WO_NIL_SCALAR` spells a `?scalar` nil) |
| `TEXT` | pointer, 0 = nil | `db_text { len u32; bytes[] }` |
| `OWNED` | pointer, 0 = nil | `db_rec { class_id u32; slots[] }` — flattened by value, recursively through these same rules |
| `MULTI` | pointer, 0 = nil | `db_multi { elem_kind u8; len u32; items[] }`, elements encoded element-wise |
| `MAP` | pointer, 0 = nil | `db_map { key_kind, val_kind u8; len u32; kv pairs }` |
| `GCREF` | **never stored** | compile error upstream (the GC bulkhead); the engine refuses it defensively as an encode error |
`ref T` is a `SCALAR` at this layer — the target row's id. The engine learns
what it references only when the FK checks land (9b plan, Task 3).
## Storage
Per shard, per class, created lazily on first insert:
- **Slabs** of 256 rows (`DB_SLAB_ROWS`), malloc'd, **never moved or freed
while the table lives** — a row's address is stable for its lifetime,
which is the property 9b's loop-scoped row views stand on.
- An **occupancy bitmap** (one bit per slot, slab-major) and a LIFO
**free-slot list**: removal recycles the slot; a recycled slot is always
used before a new slab grows. Ids are never reused; slots are.
- The **primary index**: an open-addressing hash, id → slot, splitmix64
finalizer, power-of-two capacity, 0.7 load, tombstoned deletes (ids are
never 0 and never reused, so the all-ones sentinel cannot collide).
## Id discipline
Per table, per shard: shard S of N allocates `S+1, S+1+N, S+1+2N, …` — the
c-runtime plan's shipped interleave. Creation is coordination-free; a row's
owner shard is `(id-1) % N`. Milestone 1 runs at N=1 and everything
degenerates to `1, 2, 3, …`. Id 0 does not exist (it is the hash's "empty"
and the `?ref`'s nil).
## Choke points
`wo_row_insert` and `wo_row_remove` are the only functions that mutate a
table. Task 4's secondary indexes hook exactly these two sites (marked
`INDEX HOOK` in `database/src/table.c`); the WAL (Task 2) stages its record
beside the same calls. Anything else touching a slab is a defect by
definition — the doctrine the Rust engine learned and this engine enforces.
## WAL (Task 2) — `database/src/wal.{c,h}`
Record framing, replay-whole-or-not-at-all:
```
record := len u32 | crc u32 | payload | mark u32
len = payload bytes (never 0: a zero length is the preallocated tail)
crc = CRC32 (poly 0xEDB88320) of the payload
mark = 0x574F4C31 "WOL1", the last bytes of the record — a record
without its mark is torn by definition
payload := kind u8 | class_id u32 | row_id u64 | body
kind = 1 insert (body = fields), 2 remove (no body), 3 update (Task 5)
```
Body fields walk the class table's kinds: `SCALAR` 8 bytes; `TEXT` u32 len +
bytes (`0xFFFFFFFF` = nil); `OWNED` presence u8 then class id + fields
recursively; `MULTI` presence + elem kind + len + elements; `MAP` presence +
both kinds + len + pairs. Little-endian, same platform note as the loader.
**Commit order (doctrine, verbatim from the shipped phase-D pattern):** RAM
apply → stage record → `wo_wal_commit` (one pwrite of the batch + one
fdatasync) → only then acknowledge. Group commit = everything staged since
the last commit rides one sync.
**Replay** decodes payloads straight into engine-owned values — no VM heap
involved, boot cannot depend on a VM existing — and rows re-enter through
the choke-point row API, so Task 4's indexes rebuild for free. A torn tail
(short record, bad CRC, missing mark, zero length) ends the intact prefix:
everything from the tear on is dropped whole, and `wo_wal_open` positions
its write offset AT the tear so the next commit overwrites it. A record that
CRC-passes but does not decode is corruption, not a tear — replay fails
loudly. A missing file is a fresh boot, not an error. After replay each
table's `next_id` sits past every replayed id this shard owns.
**Oracle:** `wo_wal_check(path)` walks a file with no engine and reports the
intact record count and prefix end — the crash battery's verifier
(`runtime/test/test_wal.c`: five rounds of insert/commit/ack-over-pipe with
SIGKILL mid-stream; every acked row present and exact after replay).
## Insert (Task 3) — builtin 61, `database/src/db.c`
`insert Class { field: expr, … }` is a typed expression (statement position
included): fields validate like a constructor literal (defaults and `?`
fields omittable — an omitted `?scalar` gets `WO_NIL_SCALAR`, other omitted
optionals the zero word, declared defaults their value), and the result is
the new row's id. Lowering emits builtin **61**: R[B] = class-id constant,
R[B+1..] = one slot per declared field in declaration order (the literal's
order is irrelevant — slots are the class table's).
Execution: `wo_row_insert` (RAM, engine copies every value), then — when
durability is on — stage + **commit before the builtin returns**: the
builtin's return IS the acknowledgment, so ack-after-fsync holds at
statement granularity until iteration 8 brings tick-scoped group commit. A
failed commit un-applies the row and traps `WO_T_IO`; engine failures trap
`WO_T_DB`. Durability is opt-in: `WO_DATA=<dir>` makes the CLI replay
`<dir>/shard-0.wal` before the entry runs and commit every insert; without
it the engine is RAM-only (every corpus fixture runs that way).
Ownership: the engine copies at the row API, so an insert **borrows** its
field values — no transfer, no E304; freshly built values are dropped at the
site (emit.ml mirrors the push/set reap). The insert node is trap-capable
(unique violations arrive with Task 4) and carries a live-mask drop entry.
## Still to come in this document
- **Task 4**: secondary-index format, `@unique` trap code.
- **Task 5**: the select subset, update record semantics, and its builtins.

View file

@ -2,7 +2,7 @@
A seven-phase design series that starts with "should writeonce use a document or graph database?" and arrives at a full-stack declarative application platform — then plans the migration from writeonce's current flat-file store to that platform. A seven-phase design series that starts with "should writeonce use a document or graph database?" and arrives at a full-stack declarative application platform — then plans the migration from writeonce's current flat-file store to that platform.
> **Start here if you're new:** [wo-language.md](./wo-language.md) — the user-facing overview of what writeonce *is* (a programming language with DB + HTTP in its runtime, Go-style toolchain). This series is the engineering plan that gets you there. > **Start here if you're new:** wo-language.md — the user-facing overview of what writeonce *is* (a programming language with DB + HTTP in its runtime, Go-style toolchain). This series is the engineering plan that gets you there.
Each phase is self-contained and shippable on its own. Every phase after Phase 1 builds on the previous ones. Each phase is self-contained and shippable on its own. Every phase after Phase 1 builds on the previous ones.
@ -14,7 +14,7 @@ Each phase is self-contained and shippable on its own. Every phase after Phase 1
| **2** | [The `.wo` Language & ACID Engine](./database/02-wo-language.md) | Design a two-layer `.wo` language (unified `type` schema layer + hybrid SQL/Cypher query layer with fixed glue) for an e-commerce platform with ACID transactions across relational, document, and graph storage. | | **2** | [The `.wo` Language & ACID Engine](./database/02-wo-language.md) | Design a two-layer `.wo` language (unified `type` schema layer + hybrid SQL/Cypher query layer with fixed glue) for an e-commerce platform with ACID transactions across relational, document, and graph storage. |
| **3** | [In-Memory Engine](./database/03-inmemory-engine.md) | RAM-primary, SSD-durable storage engine using `io_uring`, `mlockall`, group commit. 64 GB Linux machine. | | **3** | [In-Memory Engine](./database/03-inmemory-engine.md) | RAM-primary, SSD-durable storage engine using `io_uring`, `mlockall`, group commit. 64 GB Linux machine. |
| **4** | [Client API: Wire Protocol & Subscriptions](./database/04-client-api.md) | Native binary protocol + GraphQL over WebSocket. Subscription engine inside the transaction coordinator — no polling anywhere. | | **4** | [Client API: Wire Protocol & Subscriptions](./database/04-client-api.md) | Native binary protocol + GraphQL over WebSocket. Subscription engine inside the transaction coordinator — no polling anywhere. |
| **5** | [Go Client SDK](./database/05-go-sdk.md) | Typed Go client with subscription-first design. `gen` codegen from `.wo` schema. Subscribe to a live query in 5 lines. | | **5** | Go Client SDK | Typed Go client with subscription-first design. `gen` codegen from `.wo` schema. Subscribe to a live query in 5 lines. |
| **6** | [Low-Code Full-Stack](./database/06-lowcode-fullstack.md) | Expand `.wo` into an application language (like SAP CDS): `##ui`, `##logic`, `##policy`, `##service` blocks compiled into a single binary. | | **6** | [Low-Code Full-Stack](./database/06-lowcode-fullstack.md) | Expand `.wo` into an application language (like SAP CDS): `##ui`, `##logic`, `##policy`, `##service` blocks compiled into a single binary. |
| **7** | [Replacing `wo-seg`](./database/07-wo-seg-migration.md) | Phased coexistence plan: abstract the article store behind a trait, stand up the `.wo` engine as a second impl, dual-run, cut over, decommission `wo-seg`. | | **7** | [Replacing `wo-seg`](./database/07-wo-seg-migration.md) | Phased coexistence plan: abstract the article store behind a trait, stand up the `.wo` engine as a second impl, dual-run, cut over, decommission `wo-seg`. |
@ -56,10 +56,5 @@ Phase 7: Replace wo-seg ← trait abstraction, dual-run, cutover, decommis
## Related Documents ## Related Documents
- [wo-language.md](./wo-language.md) — writeonce as a programming language: toolchain, hello-world, stdlib, client model
- [surreal-case-study.md](./surreal-case-study.md) — SurrealDB runtime analysis; why writeonce does not use a multi-model DB for live queries - [surreal-case-study.md](./surreal-case-study.md) — SurrealDB runtime analysis; why writeonce does not use a multi-model DB for live queries
- [async.md](./async.md) — custom async runtime using Linux kernel primitives - [async.md](./async.md) — custom async runtime using Linux kernel primitives
- [05-datalayer.md](../05-datalayer.md) — current `.seg` + `.idx` implementation (8 crates, 44 tests)
- [03-data.md](../03-data.md) — data layer design with subscription model
- [06-markdown-render.md](../06-markdown-render.md) — markdown-first content model
- [ai-agents-content-management.md](../future-scope/ai-agents-content-management.md) — the `mappings` feature that started this series

View file

@ -15,9 +15,9 @@ Reference repositories:
## The Question ## The Question
writeonce articles have two shapes at once: they are **documents** (per-article JSON metadata + markdown body, per [06-markdown-render.md](../../06-markdown-render.md)) and they form a **graph** (the `mappings` field — `related`, `prerequisite`, `series`, `supersedes`, `references` — per [ai-agents-content-management.md](../../future-scope/ai-agents-content-management.md)). writeonce articles have two shapes at once: they are **documents** (per-article JSON metadata + markdown body, per 06-markdown-render.md) and they form a **graph** (the `mappings` field — `related`, `prerequisite`, `series`, `supersedes`, `references` — per ai-agents-content-management.md).
Should writeonce adopt an off-the-shelf document or graph database to back these two shapes, or keep the flat-file `.seg` + `.idx` storage already implemented in [05-datalayer.md](../../05-datalayer.md)? Should writeonce adopt an off-the-shelf document or graph database to back these two shapes, or keep the flat-file `.seg` + `.idx` storage already implemented in 05-datalayer.md?
**Short answer: no external DB.** The dataset is small (hundreds of articles, not millions of rows), single-writer (author commits), and read-heavy. A full rebuild on change is cheap. The `mappings` graph fits entirely in RAM. External databases would add a process, a protocol, a driver, and a failure mode — none of which writeonce needs. **Short answer: no external DB.** The dataset is small (hundreds of articles, not millions of rows), single-writer (author commits), and read-heavy. A full rebuild on change is cheap. The `mappings` graph fits entirely in RAM. External databases would add a process, a protocol, a driver, and a failure mode — none of which writeonce needs.
@ -45,7 +45,7 @@ Every article is a pair of files in `content/{sys_title}/`:
Plus `{sys_title}.md` with the full article body. Plus `{sys_title}.md` with the full article body.
The access patterns, per [05-datalayer.md](../../05-datalayer.md): The access patterns, per 05-datalayer.md:
| Pattern | Frequency | Current Implementation | | Pattern | Frequency | Current Implementation |
| --- | --- | --- | | --- | --- | --- |
@ -112,7 +112,7 @@ What it buys:
What it costs: What it costs:
- A Postgres process, a driver (tokio-postgres or raw libpq), connection pooling — writeonce's [05-datalayer.md](../../05-datalayer.md) explicitly removed all of this - A Postgres process, a driver (tokio-postgres or raw libpq), connection pooling — writeonce's 05-datalayer.md explicitly removed all of this
- JSONB query planning is excellent but still pays per-query cost that an in-process hash does not - JSONB query planning is excellent but still pays per-query cost that an in-process hash does not
- Recursive CTEs on deep mapping chains are slower than a RAM graph walk - Recursive CTEs on deep mapping chains are slower than a RAM graph walk
@ -161,7 +161,7 @@ What it costs:
## Option 4: In-Memory Graph — NetworkX / petgraph ## Option 4: In-Memory Graph — NetworkX / petgraph
NetworkX (Python) and petgraph (Rust) are _libraries_, not databases. You load the graph into process memory and traverse it directly. This is the model gestured at in [ai-agents-content-management.md line 181](../../future-scope/ai-agents-content-management.md) — "traversable knowledge graphs available on RAM." NetworkX (Python) and petgraph (Rust) are _libraries_, not databases. You load the graph into process memory and traverse it directly. This is the model gestured at in ai-agents-content-management.md line 181 — "traversable knowledge graphs available on RAM."
For writeonce, petgraph is the right shape: For writeonce, petgraph is the right shape:
@ -208,7 +208,7 @@ For writeonce's hundreds of articles, petgraph is the correct answer. It fits th
## Proposed Addition: `mappings.idx` Backed by petgraph ## Proposed Addition: `mappings.idx` Backed by petgraph
Per [05-datalayer.md](../../05-datalayer.md), indexes live alongside `.seg`. Add a fourth index: Per 05-datalayer.md, indexes live alongside `.seg`. Add a fourth index:
``` ```
data/ data/
@ -268,5 +268,3 @@ Key files to study:
See also: See also:
- [surreal-case-study.md](../surreal-case-study.md) — why writeonce does not use a multi-model DB for live queries - [surreal-case-study.md](../surreal-case-study.md) — why writeonce does not use a multi-model DB for live queries
- [05-datalayer.md](../../05-datalayer.md) — current `.seg` + `.idx` implementation
- [ai-agents-content-management.md](../../future-scope/ai-agents-content-management.md) — the `mappings` feature this index supports

View file

@ -2,6 +2,14 @@
> A two-layer multi-paradigm language for an e-commerce platform with ACID transactions across relational, document, and graph storage. > A two-layer multi-paradigm language for an e-commerce platform with ACID transactions across relational, document, and graph storage.
> **Query layer superseded (2026-08-15):** the SQL + Cypher query layer below
> is design history — on the C stack, programs query their tables through the
> language-integrated surface specified in
> [`2026-08-15-table-relations-query-design.md`](../../superpowers/specs/2026-08-15-table-relations-query-design.md)
> (fork 1's decision record). This document remains the reference for the
> schema layer's vocabulary and for the `wo-db` prototype's engine semantics;
> SQL text is at most a future export format, never a program surface.
**Previous**: [Phase 1 — Database Evaluation](./01-evaluation.md) | **Next**: [Phase 3 — In-Memory Engine](./03-inmemory-engine.md) | **Index**: [database.md](../database.md) **Previous**: [Phase 1 — Database Evaluation](./01-evaluation.md) | **Next**: [Phase 3 — In-Memory Engine](./03-inmemory-engine.md) | **Index**: [database.md](../database.md)
--- ---
@ -65,7 +73,7 @@ That three-paradigm sketch is **the execution substrate** — the thing the engi
Both layers are `.wo` files — same extension, same tooling, same parser front-end. They differ in role: Both layers are `.wo` files — same extension, same tooling, same parser front-end. They differ in role:
- The **schema layer** names the data model once. One `type` declaration per entity covers what the three paradigm blocks cover today (relational columns, embedded documents, graph edges) plus constraints, computed fields, policies, and triggers. It is the source of truth for codegen ([Phase 5](./05-go-sdk.md)) and for the full-stack blocks ([Phase 6](./06-lowcode-fullstack.md)). - The **schema layer** names the data model once. One `type` declaration per entity covers what the three paradigm blocks cover today (relational columns, embedded documents, graph edges) plus constraints, computed fields, policies, and triggers. It is the source of truth for codegen (Phase 5) and for the full-stack blocks ([Phase 6](./06-lowcode-fullstack.md)).
- The **query layer** is the operational surface. SQL and Cypher stay as-is — they are universally legible, every backend developer already reads them — but five things are tightened so the three grammars share semantics (parameters, `RETURNING`, dotted paths, transactions, `LIVE`). - The **query layer** is the operational surface. SQL and Cypher stay as-is — they are universally legible, every backend developer already reads them — but five things are tightened so the three grammars share semantics (parameters, `RETURNING`, dotted paths, transactions, `LIVE`).
The two layers ship on different timelines. The query layer is Phase 2 (already prototyped at [`prototypes/wo-db/`](../../../prototypes/wo-db/)). The schema layer enters when Phase 5 codegen needs a single authoritative input. The two layers ship on different timelines. The query layer is Phase 2 (already prototyped at [`prototypes/wo-db/`](../../../prototypes/wo-db/)). The schema layer enters when Phase 5 codegen needs a single authoritative input.

View file

@ -2,7 +2,7 @@
> How remote clients connect, query, and subscribe to live changes — no polling anywhere in the chain. > How remote clients connect, query, and subscribe to live changes — no polling anywhere in the chain.
**Previous**: [Phase 3 — In-Memory Engine](./03-inmemory-engine.md) | **Next**: [Phase 5 — Go Client SDK](./05-go-sdk.md) | **Index**: [database.md](../database.md) **Previous**: [Phase 3 — In-Memory Engine](./03-inmemory-engine.md) | **Next**: Phase 5 — Go Client SDK | **Index**: [database.md](../database.md)
--- ---
@ -27,7 +27,7 @@ Every naive realtime system reaches for polling first. For an e-commerce platfor
| Inventory display lies for up to N ms | Inventory display reflects the commit | | Inventory display lies for up to N ms | Inventory display reflects the commit |
| "Order shipped" email triggered by cron | Fired by a committed status-change | | "Order shipped" email triggered by cron | Fired by a committed status-change |
Every subscription-based design in this section follows the same rule already set by writeonce in [05-datalayer.md](../../05-datalayer.md) and [03-data.md](../../03-data.md): **the client registers a query once, the server pushes deltas on commit, the client never asks again.** Every subscription-based design in this section follows the same rule already set by writeonce in 05-datalayer.md and 03-data.md: **the client registers a query once, the server pushes deltas on commit, the client never asks again.**
## Protocol Layer — Pick One or Both ## Protocol Layer — Pick One or Both
@ -193,7 +193,7 @@ Each connected client has server-side state:
| Role / RBAC context | Session | Feeds row-level policies into the planner | | Role / RBAC context | Session | Feeds row-level policies into the planner |
| Back-pressure credits | Per-subscription | Client advertises how many outstanding `DELTA` frames it can buffer | | Back-pressure credits | Per-subscription | Client advertises how many outstanding `DELTA` frames it can buffer |
On disconnect (TCP close, keepalive failure, `EPOLLHUP`-equivalent from io_uring completion): all sessions state is freed, all subscriptions unregistered. Same philosophy as `wo-sub`'s `EPOLLHUP` → automatic `unsubscribe(fd)` from [05-datalayer.md](../../05-datalayer.md), scaled up to a real server. On disconnect (TCP close, keepalive failure, `EPOLLHUP`-equivalent from io_uring completion): all sessions state is freed, all subscriptions unregistered. Same philosophy as `wo-sub`'s `EPOLLHUP` → automatic `unsubscribe(fd)` from 05-datalayer.md, scaled up to a real server.
## Back-Pressure ## Back-Pressure

View file

@ -1,371 +0,0 @@
# Phase 5 — Go Client SDK
> A typed Go client with subscription-first design — subscribe to a live query in 5 lines, deltas arrive on a channel.
**Previous**: [Phase 4 — Client API](./04-client-api.md) | **Next**: [Phase 6 — Low-Code Full-Stack](./06-lowcode-fullstack.md) | **Index**: [database.md](../database.md)
---
Concrete scenario: a developer writes a Go backend that serves a web frontend, and uses the `.wo` database as the store. They must be able to **subscribe to a query in ~5 lines of idiomatic Go** and have deltas arrive on a channel.
Everything else the SDK does — connect, query, mutate, transact — is table stakes covered by every existing Go DB driver. Subscriptions are what this SDK has to get right.
## Target API Surface
The engine speaks `.wo` on the wire. Every SDK method — typed or untyped — is a thin wrapper over one primitive: **send a `.wo` source string with `$name` parameters, get back a uniform `Result`**.
```go
import "go.writeonce.dev/wo"
// 1. Connect
client, err := wo.Connect(ctx, "wo://db.example.com:5555",
wo.WithAPIKey(os.Getenv("WO_KEY")),
wo.WithTLS(tlsConfig),
)
defer client.Close()
// 2. Wo — the primitive: run ANY .wo source (one statement or a whole
// BEGIN...COMMIT block mixing SQL, Cypher, and document updates). Params
// use the $name rule from the .wo language spec.
result, err := client.Wo(ctx, `
UPDATE products
SET inventory.on_hand -= $qty
WHERE id = $pid AND inventory.on_hand >= $qty
RETURNING id AS pid;
INSERT INTO orders (user_id, total_cents, status)
VALUES ($uid, $total, 'pending')
RETURNING id AS oid;
MATCH (u:user {id: $uid}), (p:product {id: $pid})
CREATE (u)-[:PURCHASED {order_id: $oid, qty: $qty, at: now()}]->(p);
`, wo.Params{
"uid": uid, "pid": 42, "qty": 2, "total": 9800,
})
// Result layout:
// result.Rows — rows from trailing SELECT/MATCH/RETURNING statements
// result.Aliases — the RETURNING alias table: {"pid": 42, "oid": 17}
// result.Affected — rows touched by INSERT/UPDATE/DELETE, per-statement
// 3. Query — sugar over Wo that scans trailing rows into a typed destination.
var products []Product
err = client.Query(ctx,
"SELECT id, sku, price_cents, meta FROM products WHERE price_cents < $max",
wo.Params{"max": 5000},
).Scan(&products)
// 4. Exec — sugar for writes that don't return rows.
_, err = client.Exec(ctx,
"UPDATE products SET price_cents = $new WHERE id = $id",
wo.Params{"new": 4900, "id": 42},
)
// 5. Tx — wraps Wo/Query/Exec in a server-side BEGIN...COMMIT. The callback's
// return value decides commit vs rollback. Useful when the program needs
// to branch between statements on intermediate results.
err = client.Tx(ctx, func(tx *wo.Tx) error {
r, err := tx.Wo(ctx,
`UPDATE products SET inventory.on_hand -= $qty
WHERE id = $pid AND inventory.on_hand >= $qty
RETURNING id AS pid;`,
wo.Params{"pid": 42, "qty": 2})
if err != nil { return err }
if r.Affected[0] == 0 { return wo.ErrInsufficientInventory }
_, err = tx.Wo(ctx, `
INSERT INTO orders (user_id, status) VALUES ($uid, 'pending') RETURNING id AS oid;
MATCH (u:user {id: $uid}), (p:product {id: $pid})
CREATE (u)-[:PURCHASED {order_id: $oid, qty: $qty}]->(p);
`, wo.Params{"uid": uid, "pid": 42, "qty": 2})
return err
})
```
`Wo` is the primitive; `Query`, `Exec`, and `Subscribe` are typed sugar. If the engine accepts the `.wo` source on disk, `client.Wo` accepts the same string over the wire.
### When to use raw `.wo` vs typed codegen
Both styles coexist in the same program; they share the connection pool.
| Use raw `.wo` (`client.Wo`) when | Use typed codegen (`client.Orders.Create`, etc.) when |
| --- | --- |
| Ad-hoc queries, admin tools, one-off scripts | The app's hot path — compile-time schema checking + IDE autocomplete pay for themselves |
| Cross-cutting queries that join multiple generated types | Per-type CRUD + subscriptions |
| Multi-statement transactions threading `RETURNING` aliases | Single-statement operations |
| DB repair, data migration, ad-hoc analytics | Anything the codegen already covers |
| You want to paste a block from a `.wo` source file straight into Go | You want refactor-safe struct field access |
The canonical rule: **write typed code first, drop to raw `.wo` when the type system gets in the way**. They interleave freely — a typed `client.Orders.Subscribe(...)` can run next to a raw `client.Wo(...)` admin query in the same handler.
## Subscribe — The Primary Use Case
Idiomatic Go for a stream of values is a channel read inside a `for` loop, cancelled by `context.Context`. That is the exact shape a `.wo` subscription should take:
```go
sub, err := client.Subscribe(ctx,
"LIVE SELECT sku, inventory.on_hand FROM products WHERE sku IN $skus",
wo.Params{"skus": cartSkus},
)
if err != nil { return err }
defer sub.Close()
for delta := range sub.Deltas() {
switch d := delta.(type) {
case wo.Insert:
log.Printf("new row: %+v", d.Row)
case wo.Update:
log.Printf("sku=%s on_hand=%d -> %d", d.Key, d.Old["on_hand"], d.New["on_hand"])
case wo.Delete:
log.Printf("removed: %s", d.Key)
case wo.Resync:
// server dropped our queue — refetch and resume
currentState = refetch()
}
}
// loop exits when:
// - ctx cancelled (client shutdown)
// - sub.Close() called (defer)
// - server sent COMPLETE (schema change, permission revoked)
// sub.Err() returns the reason
if err := sub.Err(); err != nil { log.Fatal(err) }
```
**Contract:**
- `sub.Deltas()` returns `<-chan wo.Delta` — standard read-only channel. The SDK closes it when the subscription ends.
- `ctx` cancellation immediately stops deliveries and closes the channel. No leaked goroutines.
- Ordering: deltas arrive in commit order. A `DELTA` on the wire always reflects a committed transaction.
- Back-pressure: the channel has a bounded buffer (default 1024). If it fills, the SDK's policy kicks in (see below).
## Typed SDK via `.wo` Schema Codegen
The `.wo` **schema layer** ([Phase 2](./02-wo-language.md)) declares types. `wo-gen` reads the type DSL — not the underlying `##sql`/`##doc`/`##graph` blocks — as its input; that way one Go struct corresponds to one entity, with embedded documents and graph-edge projections folded in naturally.
```wo
type Product {
id: Id
sku: SKU @unique
price: Money
meta: { title: Text, description: Markdown, images: [Url], reviews: [Review] }
inventory: { on_hand: Int @check(>= 0), reserved: Int = 0, reorder_at: Int }
purchased_by: multi User via Purchase -- inverse graph link
}
```
```bash
wo-gen --schema ./schema.wo --out ./internal/wodb
```
Produces:
```go
package wodb
// from `type Product` — embedded structs compile from the inline `{...}` fields
type Product struct {
ID int64 `wo:"id"`
SKU string `wo:"sku"`
Price Money `wo:"price"`
Meta ProductMeta `wo:"meta"`
Inventory InventoryLvl `wo:"inventory"`
PurchasedBy []PurchaseEdge `wo:"purchased_by"` // link-with-props → edge struct
}
// embedded document inside Product.Meta
type ProductMeta struct {
Title string `wo:"title"`
Description string `wo:"description"`
Images []string `wo:"images"`
Attributes map[string]string `wo:"attributes"`
Reviews []Review `wo:"reviews"`
}
// graph link carrying properties — target + edge props in one struct
type PurchaseEdge struct {
Target User `wo:"target"`
Order int64 `wo:"order"`
Qty int `wo:"qty"`
At time.Time `wo:"at"`
}
// registered live queries become typed helpers
func InventoryChanged(ctx context.Context, c *wo.Client, skus []string) (*wo.TypedSubscription[InventoryLvl], error)
```
One type declaration → one Go struct. Zero-property graph edges (`multi User @edge(:FOLLOWS)`) generate `Friends []User`; link-with-properties types generate `[]EdgeStruct`; computed fields become read-only struct fields populated by the planner.
**Transition path.** While the Phase 2 prototype is still authored directly in `##sql/##doc/##graph` blocks (the query layer), `wo-gen` accepts either — a file of type declarations, or the raw paradigm blocks — and emits the same Go output. The type DSL becomes mandatory only once Phase 6 full-stack blocks (which attach to types) start shipping.
Typed subscription loop loses all `interface{}` ceremony:
```go
sub, err := wodb.InventoryChanged(ctx, client, cartSkus)
if err != nil { return err }
defer sub.Close()
for d := range sub.C {
switch d.Kind {
case wo.DeltaUpdate:
log.Printf("sku=%s now %d in stock", d.Key, d.New.OnHand)
}
}
```
Generics (Go 1.18+) make `TypedSubscription[T]` a single parameterized type — no per-query generated struct. Only the `T` struct itself is generated.
## Connection Lifecycle
```go
type Client struct {
// opaque; holds a connection pool, codec, session registry
}
func Connect(ctx context.Context, dsn string, opts ...Option) (*Client, error)
func (c *Client) Close() error
func (c *Client) Ping(ctx context.Context) error
```
Inside the SDK, one TCP connection per client is fine for native protocol (multiplexed), but a small pool (2–4) helps when one connection's receive goroutine is saturated decoding a large result set. Connection state:
- **Connecting** → `HELLO` sent, waiting for `WELCOME`
- **Ready** → normal operation
- **Reconnecting** → transient network error; automatic exponential backoff; all subscriptions queued for re-registration
- **Closed** → terminal
**Reconnection semantics for subscriptions** (the subtle part): on reconnect, the SDK re-sends every active `SUBSCRIBE` frame. The server replies with a `RESYNC` marker and the current matching state. The app's subscription channel emits a single `wo.Resync{}` value so the consumer knows to rebuild local state. No delta is silently lost, no delta is silently duplicated.
## Options
Fluent options, not a bloated config struct:
```go
wo.WithAPIKey(key string)
wo.WithJWT(token string)
wo.WithMTLS(cert tls.Certificate)
wo.WithTLS(cfg *tls.Config)
wo.WithPoolSize(n int) // default 2
wo.WithSubscriptionBuffer(n int) // default 1024
wo.WithOverflowPolicy(wo.DropAndResync | wo.Coalesce | wo.Disconnect)
wo.WithLogger(l *slog.Logger)
wo.WithRetry(wo.RetryPolicy{...})
wo.WithProtocol(wo.ProtocolNative | wo.ProtocolGraphQL) // native default
```
## Transactions
`client.Tx` maps to [Phase 2's](./02-wo-language.md) `BEGIN ... COMMIT`. The callback's return value decides commit vs rollback:
- `return nil` → `COMMIT` sent, error only if server rejects commit
- `return err` → `ROLLBACK` sent, original `err` surfaced to caller
- `panic` → `ROLLBACK` sent, panic re-raised
- `ctx` cancel → `ROLLBACK` sent, `ctx.Err()` returned
Nested `tx.Wo`/`tx.Query`/`tx.Exec` route to the same server-side transaction — no connection hopping. The SDK enforces this by pinning the transaction to one connection for its lifetime. `RETURNING` aliases bound by one statement in the txn are visible to every later statement in the same txn through the server-side alias table (see [Phase 2 — Transaction Coordinator](./02-wo-language.md#cross-paradigm-transaction-coordinator)), so a Go `tx.Wo` call can leave `$oid` set and the next `tx.Wo` call can use it.
## Back-Pressure Handling in the SDK
The server's back-pressure policy from [Phase 4](./04-client-api.md) is mirrored client-side:
| Client situation | SDK behavior |
| --- | --- |
| Consumer reading channel fast enough | Normal delivery |
| Channel buffer full (1024 unread deltas) | Per `WithOverflowPolicy`: drop buffered + emit `wo.Resync`, or coalesce same-key updates, or close the subscription with `ErrOverflow` |
| Network stalled | `ctx.Deadline` + keepalive `PING` every 10s; stall > 30s → disconnect and reconnect |
| Server closed subscription | Channel closed, `sub.Err()` returns reason (schema change, permission revoked, engine shutdown) |
**Never block the receive goroutine on a full channel.** The SDK drains the socket no matter what; overflow policy decides what to do with the deltas it can't deliver.
## Full Cart-Inventory Example
End-to-end: Go HTTP handler that renders a cart page and keeps its inventory line live via Server-Sent Events to the browser. The SDK drives the upstream subscription to `.wo`:
```go
func (h *Handler) cartInventoryStream(w http.ResponseWriter, r *http.Request) {
skus := parseSkus(r.URL.Query().Get("skus"))
w.Header().Set("Content-Type", "text/event-stream")
w.Header().Set("Cache-Control", "no-cache")
flusher := w.(http.Flusher)
sub, err := wodb.InventoryChanged(r.Context(), h.db, skus)
if err != nil { http.Error(w, err.Error(), 500); return }
defer sub.Close()
for d := range sub.C {
payload, _ := json.Marshal(d)
fmt.Fprintf(w, "event: inventory\ndata: %s\n\n", payload)
flusher.Flush()
}
}
```
The browser connects once with `new EventSource('/cart/inventory?skus=...')`. The Go handler holds one subscription to `.wo`. When inventory commits in the database, the delta flows: engine → subscription registry → Go SDK channel → SSE stream → DOM update. Zero polling anywhere in the chain.
## Go-Specific Design Details
| Go idiom | Application |
| --- | --- |
| `context.Context` threading | Every method takes `ctx` as first arg; cancellation propagates to the wire |
| `io.Closer` | `Client`, `Tx`, `Subscription` all implement `Close() error` |
| Small interfaces | `type Runner interface { Wo(ctx, src, params) (wo.Result, error) }` — `*Client` and `*Tx` both satisfy it, so helper functions compose cleanly |
| `database/sql`-style `Scan` | `client.Query(...).Scan(&dest)` accepts struct, slice of struct, or primitives |
| `sql.Null*` analogues | `wo.NullString`, `wo.NullInt64`, `wo.NullDoc` for optional doc columns |
| Struct tags | `wo:"column_name"` + JSON-style for nested doc fields (`wo:"meta.title"`) |
| `errors.Is` / `errors.As` | `errors.Is(err, wo.ErrConflict)`, `wo.AsError(err, &woErr)` |
| No goroutine leaks | Every background goroutine tied to ctx or a sync.WaitGroup closed in `Client.Close()` |
| Testing via interfaces | `wo.DB` interface; provide `wotest.NewMock()` for unit tests; real embedded engine for integration |
## Comparison With Existing Go DB SDKs
| SDK | Query style | Subscriptions | Transactions | Typed results |
| --- | --- | --- | --- | --- |
| `database/sql` + `pq` | SQL strings | No | Yes | Manual `Scan` |
| `pgx` | SQL strings | `LISTEN/NOTIFY` only (no row-level) | Yes | Manual or `pgxscan` |
| `sqlc` | Generated Go funcs from `.sql` | No | Yes | Generated structs |
| `ent` | ORM | No | Yes | Generated |
| `go-redis` | Commands | Pub/sub + keyspace notifications (no query) | Multi/Exec | Manual |
| `surrealdb/surrealdb.go` | Raw queries | `Live()` returning channel | Yes | Manual |
| `gqlgen` / `machinebox/graphql` | GraphQL docs | WebSocket subscriptions | N/A | Generated |
| **`sa`** (this design) | `.wo` queries | **Native `LIVE` → typed channel** | Yes | Codegen from schema |
The reference points are `sqlc` (for the codegen pipeline) and `surrealdb-go` (for the subscription channel API). Combining their best ideas and tightening the subscription contract is what this SDK is.
## Reference Implementations To Steal From
- **surrealdb/surrealdb.go** — `Live()` returns a channel; closest API precedent. <https://github.com/surrealdb/surrealdb.go>
- **jackc/pgx** — reference quality for a Go database driver. Connection pool, copy protocol, prepared statements all done right. <https://github.com/jackc/pgx>
- **sqlc-dev/sqlc** — codegen from SQL to typed Go. The model for `.wo` → Go. <https://github.com/sqlc-dev/sqlc>
- **Khan/genqlient** — generated typed GraphQL client. Ergonomic precedent for typed query helpers. <https://github.com/Khan/genqlient>
- **nats-io/nats.go** — subscription-first API, back-pressure handled well. `sub.NextMsg(ctx)` and channel-based `ChanSubscribe` both supported. <https://github.com/nats-io/nats.go>
- **hasura/go-graphql-client** — GraphQL subscriptions over WebSocket in Go. <https://github.com/hasura/go-graphql-client>
## SDK Delivery
| Artifact | Purpose |
| --- | --- |
| `go.writeonce.dev/wo` | Runtime package: client, query, subscribe |
| `go.writeonce.dev/wo/wotest` | Mock client + in-memory engine for unit tests |
| `wo-gen` binary | Reads `schema.wo`, emits typed Go code |
| Go module example repo | Cart + inventory demo wired end-to-end |
| Generated docs | `go doc` + hosted examples |
Publishing strategy: semantic versioning, `v0.x` while the wire protocol is unstable, `v1.0` only after the protocol is frozen.
## Why The SDK Matters As Much As The Engine
A database with a beautiful engine and a painful client is a database no one uses. The e-commerce Go backends this targets are built under deadline — if `Subscribe` is not as easy as opening a channel, developers will reach for polling (`time.Tick` + `SELECT`) and defeat the whole architecture.
The success metric is blunt: **a developer who has never seen `.wo` before should have a working subscription to a live query inside 15 minutes**, counting install, schema codegen, and the first delta landing on their channel. If the SDK is any harder than that, the rest of this doc is academic.
## Future SDKs
Same shape, other languages:
- **TypeScript / browser** — fetch + WebSocket for GraphQL subscriptions; types via codegen from `.wo`. Highest priority after Go for a web-first product.
- **Rust** — direct native protocol, `tokio`-friendly, `impl Stream<Item = Delta>` for subscriptions.
- **Python** — async/await, `async for delta in sub` idiom.
- **Java / Kotlin** — Flow (Kotlin) or Reactive Streams (Java) for subscriptions.
Each follows the same rule: subscribe-to-query must be the shortest, most obvious thing in the API.

View file

@ -2,7 +2,7 @@
> Expand `.wo` from a query language into a declarative application DSL — schema, services, UI, business logic, and authorization in one language, compiled into a single binary. > Expand `.wo` from a query language into a declarative application DSL — schema, services, UI, business logic, and authorization in one language, compiled into a single binary.
**Previous**: [Phase 5 — Go Client SDK](./05-go-sdk.md) | **Index**: [database.md](../database.md) **Previous**: Phase 5 — Go Client SDK | **Index**: [database.md](../database.md)
--- ---
@ -282,7 +282,7 @@ For the writeonce schema above, `sa build` produces:
| SSR HTML | `app/ui/*.wo` + routes in `app.wo` | Per-route renderers compiled into the server binary | | SSR HTML | `app/ui/*.wo` + routes in `app.wo` | Per-route renderers compiled into the server binary |
| Client runtime | `app/ui/*.wo` | Small JS bundle: subscription client + DOM patcher + form binding | | Client runtime | `app/ui/*.wo` | Small JS bundle: subscription client + DOM patcher + form binding |
| Admin UI | All of the above | Auto-generated CRUD screens for every `##sql`/`##doc` entity (override any with a `##ui` block) | | Admin UI | All of the above | Auto-generated CRUD screens for every `##sql`/`##doc` entity (override any with a `##ui` block) |
| Typed SDKs | `app/database/*.wo` | Go/TypeScript/Rust clients per [Phase 5](./05-go-sdk.md) | | Typed SDKs | `app/database/*.wo` | Go/TypeScript/Rust clients per Phase 5 |
| Migrations | Schema diff vs. current database | Versioned forward/backward migrations in `migrations/` | | Migrations | Schema diff vs. current database | Versioned forward/backward migrations in `migrations/` |
| Observability | Everything | Structured logs, query metrics, subscription lag dashboards | | Observability | Everything | Structured logs, query metrics, subscription lag dashboards |
@ -359,7 +359,7 @@ Rough effort on top of Phases 2–4: **12–24 months** with a small team, most
This section is the endgame, not the next step. The sensible build order: This section is the endgame, not the next step. The sensible build order:
1. Ship the engine ([Phase 2](./02-wo-language.md), in-memory + io_uring durability via [Phase 3](./03-inmemory-engine.md)) — query layer (`##sql`/`##doc`/`##graph`) only. 1. Ship the engine ([Phase 2](./02-wo-language.md), in-memory + io_uring durability via [Phase 3](./03-inmemory-engine.md)) — query layer (`##sql`/`##doc`/`##graph`) only.
2. Ship the wire protocol + Go SDK ([Phase 4](./04-client-api.md) + [Phase 5](./05-go-sdk.md)). 2. Ship the wire protocol + Go SDK ([Phase 4](./04-client-api.md) + Phase 5).
3. Ship the **schema-layer `type` DSL** that compiles to the three paradigm blocks. From this point forward, authoring happens against types; the paradigm blocks become an artifact the compiler emits. 3. Ship the **schema-layer `type` DSL** that compiles to the three paradigm blocks. From this point forward, authoring happens against types; the paradigm blocks become an artifact the compiler emits.
4. Add type-attached `service` (and standalone `##service` for bundles) — declarative endpoints. 4. Add type-attached `service` (and standalone `##service` for bundles) — declarative endpoints.
5. Add type-attached `policy` (and standalone `##policy` for cross-entity rules) — declarative authorization. 5. Add type-attached `policy` (and standalone `##policy` for cross-entity rules) — declarative authorization.

View file

@ -56,7 +56,7 @@ Port the C++ prototype (`prototypes/wo-db/src/*`) to Rust, split along the natur
| `sub` | live subscriptions — delta frames on commit | new ([Phase 4](./04-client-api.md)) | 4 | | `sub` | live subscriptions — delta frames on commit | new ([Phase 4](./04-client-api.md)) | 4 |
| `http` | wire protocol — REST / GraphQL-over-WS / native codec | new ([Phase 4](./04-client-api.md)) | 4 | | `http` | wire protocol — REST / GraphQL-over-WS / native codec | new ([Phase 4](./04-client-api.md)) | 4 |
| `db` | top-level facade: `open()`, `Tx`, `Query`, `Subscribe` — the Rust SDK | integrates the above | 2–4 | | `db` | top-level facade: `open()`, `Tx`, `Query`, `Subscribe` — the Rust SDK | integrates the above | 2–4 |
| `gen` | codegen: `.wo type` → Rust structs, Go structs, TypeScript | `sa-gen`/`wo-gen` in [Phase 5](./05-go-sdk.md) | 5 | | `gen` | codegen: `.wo type` → Rust structs, Go structs, TypeScript | `sa-gen`/`wo-gen` in Phase 5 | 5 |
All 15 crates (these 14 plus the existing `rt` binary crate) now exist as empty skeletons in `crates/`. See [`crates/README.md`](../../../crates/README.md) and [`docs/plan/done/01-scafolding-crates.md`](../../plan/done/01-scafolding-crates.md) for the scaffolding plan that landed them. All 15 crates (these 14 plus the existing `rt` binary crate) now exist as empty skeletons in `crates/`. See [`crates/README.md`](../../../crates/README.md) and [`docs/plan/done/01-scafolding-crates.md`](../../plan/done/01-scafolding-crates.md) for the scaffolding plan that landed them.
@ -203,7 +203,5 @@ Each phase has its own exit criteria above. End-to-end verification for the whol
- [02-wo-language.md](./02-wo-language.md) — the two-layer `.wo` language the engine speaks - [02-wo-language.md](./02-wo-language.md) — the two-layer `.wo` language the engine speaks
- [03-inmemory-engine.md](./03-inmemory-engine.md) — the storage engine behind `wo-db` - [03-inmemory-engine.md](./03-inmemory-engine.md) — the storage engine behind `wo-db`
- [04-client-api.md](./04-client-api.md) — wire protocol and `LIVE` subscriptions - [04-client-api.md](./04-client-api.md) — wire protocol and `LIVE` subscriptions
- [05-go-sdk.md](./05-go-sdk.md) — the Go SDK built from `.wo` types via `wo-gen`
- [01-evaluation.md](./01-evaluation.md) — why writeonce built `wo-seg` in the first place, and why that choice still looks right for the blog even as the platform grows past it - [01-evaluation.md](./01-evaluation.md) — why writeonce built `wo-seg` in the first place, and why that choice still looks right for the blog even as the platform grows past it
- [../05-datalayer.md](../../05-datalayer.md) — current `.seg` + `.idx` implementation details
- `prototypes/wo-db/` — the C++ prototype of the `.wo` engine, the reference implementation the Rust port follows - `prototypes/wo-db/` — the C++ prototype of the `.wo` engine, the reference implementation the Rust port follows

View file

@ -104,7 +104,7 @@ SurrealDB's live query system pushes changes to connected clients in real-time:
| Runtime | Tokio multi-threaded executor | Single-threaded epoll event loop | | Runtime | Tokio multi-threaded executor | Single-threaded epoll event loop |
| Protocol framing | WebSocket frames | Length-prefixed payloads (no protocol) | | Protocol framing | WebSocket frames | Length-prefixed payloads (no protocol) |
SurrealDB's live queries are the architectural inspiration for writeonce's subscription model (as noted in [03-data.md](../03-data.md)), but the implementation is fundamentally different — SurrealDB uses a full async runtime with WebSocket transport, while writeonce uses kernel fd notifications with no protocol layer. SurrealDB's live queries are the architectural inspiration for writeonce's subscription model (as noted in 03-data.md), but the implementation is fundamentally different — SurrealDB uses a full async runtime with WebSocket transport, while writeonce uses kernel fd notifications with no protocol layer.
## Storage Engine Architecture ## Storage Engine Architecture

View file

@ -1,249 +0,0 @@
# writeonce — the `.wo` Language and Runtime
> A declarative programming language with database and subscription-native HTTP in its standard runtime. Like `go run`, you write `.wo` files and execute them — but your program is a full-stack application.
---
## What writeonce is
`writeonce` is a programming language, a standard runtime, and a toolchain. Three layers of one product:
1. **The language** — `.wo` source files. Declarative by default (`type`, `class`, `service`, `policy`, `on <event>`) with a hybrid SQL+Cypher query sublanguage for the imperative parts. Types, queries, transactions, subscriptions, policies, triggers, HTTP endpoints, and UI screens are all first-class language constructs. A `class` is a `type` plus `fn` methods (`self` receiver, transactional) — state and behavior, **no inheritance** ([plan 13](../plan/13-class-model-live-pricing.md)).
2. **The runtime** — an ACID multi-paradigm database (relational + document + graph), an HTTP server, a subscription engine, and a scheduler. All of it links into a single binary with your program. No external Postgres, no external Redis, no separate Node process.
3. **The toolchain** — the `wo` command: `wo run`, `wo build`, `wo test`, `wo fmt`, `wo mod`, `wo gen`. Modelled directly on the Go toolchain. One binary per project; no runtime to install on the target host.
The one-line pitch: **Go + Postgres + `net/http` + Phoenix LiveView, folded into one language and one binary.**
## Hello, world
> Full example projects:
> - [`docs/examples/blog/`](../examples/blog/) — a blog (~200 lines): articles, authors, tags, comments, live subscriptions, row-level policies, typed Go client.
> - [`docs/examples/ecommerce/`](../examples/ecommerce/) — an e-commerce store (~300 lines): cross-paradigm ACID checkout, link types with properties, tagged unions, a **live order-ops table** that delta-updates in place.
A complete `.wo` program that creates a database table, exposes six REST endpoints with live subscriptions, and emits a typed Go client:
```wo
-- article.wo
type Article {
id: Id
title: Text
body: Markdown
author: Text
created_at: Timestamp = now()
service rest "/api/articles"
expose list, get, create, update, delete, subscribe
}
```
Run it:
```bash
$ wo run
[wo] compiling ./article.wo
[wo] schema: 1 type, 0 migrations needed
[wo] listening on :8080
GET /api/articles list
GET /api/articles/:id get
POST /api/articles create
PATCH /api/articles/:id update
DELETE /api/articles/:id delete
WS /api/articles/live subscribe
```
Use it:
```bash
$ curl -X POST localhost:8080/api/articles \
-H "Content-Type: application/json" \
-d '{"title":"Hello","body":"# First post","author":"me"}'
{"id":1,"title":"Hello","body":"# First post","author":"me","created_at":"2026-04-17T..."}
$ curl localhost:8080/api/articles
[{"id":1,"title":"Hello","...":"..."}]
```
Generate a typed client:
```bash
$ wo gen sdk --lang go --out ./client
# produces ./client/sdk.go with typed Article struct and Subscribe helper
```
Subscribe from the client — deltas push on every commit, no polling:
```go
import "myapp.example.com/client"
c, _ := client.Connect("wo://localhost:8080")
sub, _ := c.Articles.Subscribe(ctx, client.Where{Author: "me"})
for delta := range sub.C {
fmt.Printf("%s: %+v\n", delta.Kind, delta.Row)
}
```
Three files, five commands, zero infrastructure. Compare the same thing in Go+Postgres+React: one SQL schema, one migration tool, one ORM, one HTTP router, one subscription layer (polling or Redis pub-sub), one hand-written client, one React hook — roughly 2000 lines before you write any business logic.
## The toolchain
Go-literal. Every command maps to a Go equivalent so the mental model transfers:
| Command | Go equivalent | Purpose |
| --- | --- | --- |
| `wo init <name>` | `go mod init` | scaffold a new project |
| `wo run` | `go run ./...` | compile and execute |
| `wo build` | `go build` | emit a static binary |
| `wo test` | `go test` | run `.wo` tests |
| `wo fmt` | `gofmt` | canonical formatter |
| `wo vet` | `go vet` | lint + type-check without running |
| `wo mod <cmd>` | `go mod` | dependencies |
| `wo doc <sym>` | `go doc` | render docs for a type |
| `wo gen sdk --lang <L>` | `go generate` (codegen) | emit a client SDK |
| `wo migrate [--plan\|--apply]` | no direct equivalent | schema evolution |
| `wo dev` | no direct equivalent | hot-reload dev server |
**`wo run` vs `wo build`.** Same as Go: `wo run` compiles to a temp binary and executes it; `wo build` writes a named binary. No interpreter mode — `.wo` is compiled, always.
**`wo dev` is the one non-Go addition.** Edit a `.wo` file, the runtime hot-swaps the affected module without restarting. Live subscriptions survive the reload. This is the Phoenix LiveView influence.
## Program structure
```
myapp/
├── wo.toml # like go.mod — name, version, dependencies
├── main.wo # optional entry point
├── types/ # `type` declarations (one file per domain concept)
│ ├── article.wo
│ └── user.wo
├── ui/ # ##ui screens (optional)
├── tests/ # *_test.wo files
└── wo.lock # locked dependency graph (like go.sum)
```
Minimum project is one `.wo` file with one `type` declaration. The compiler generates:
- the database schema (relational row, document structures, graph edges) from the type's fields
- HTTP handlers from type-attached `service` blocks
- transactional triggers from `on <event>` blocks
- row-level policies from `policy` blocks
- typed client SDKs from the same type, on demand
No `main()` is required for a pure type-and-service app. The runtime starts the HTTP server, loads the database, and dispatches. If you need procedural entry logic (CLI args, graceful shutdown hooks, cron jobs), add `main.wo` with a `main { ... }` block.
## The runtime — what's in the standard library
Every `wo build` links these in. They're not external packages you import — they're the language.
| Component | Responsibility | Mapped to phase |
| --- | --- | --- |
| **Database** | In-RAM ACID multi-paradigm (relational + doc + graph) with WAL durability | [Phase 2](./database/02-wo-language.md) + [Phase 3](./database/03-inmemory-engine.md) |
| **Transaction coordinator** | MVCC, snapshot isolation, cross-paradigm `RETURNING` alias table | [Phase 2](./database/02-wo-language.md) |
| **HTTP server** | REST + GraphQL dispatch generated from `service` blocks | [Phase 4](./database/04-client-api.md) + `crates/http` |
| **Subscription engine** | `LIVE` queries push deltas on commit, zero polling | [Phase 4](./database/04-client-api.md) |
| **Wire protocol** | Native binary codec for typed clients | [Phase 4](./database/04-client-api.md) |
| **Codegen** | `wo gen sdk` — Go, TypeScript, Rust, Python clients from `type` declarations | [Phase 5](./database/05-go-sdk.md) |
| **UI renderer** | `##ui` screens → SSR HTML + client runtime | [Phase 6](./database/06-lowcode-fullstack.md) |
| **Authorization** | `policy` blocks compiled into planner rewrite rules | [Phase 6](./database/06-lowcode-fullstack.md) |
| **Scheduler** | Single-threaded event loop over io_uring; one core per process (shard to scale) | [async.md](./async.md) + [Phase 2 concurrency](./database/02-wo-language.md#concurrency-model) |
Comparison to Go's stdlib:
| Need | Go | writeonce |
| --- | --- | --- |
| HTTP server | `net/http` | built-in `service rest` |
| Database | none (use `database/sql` + driver + Postgres) | **built-in** |
| Template rendering | `html/template` | `##ui` blocks |
| Concurrency | goroutines + channels | single-threaded event loop (Redis-style); shard to scale past one core |
| Testing | `testing` | `wo test` + `.wo` test syntax |
| Formatting | `gofmt` | `wo fmt` |
| Modules | `go.mod` + `go.sum` | `wo.toml` + `wo.lock` |
## Clients — who consumes your program
The same `.wo` type declarations that define the database also define the wire format. `wo gen sdk` emits:
- **Go** — typed structs, `*Client`, `TypedSubscription[T]` generics over a channel
- **TypeScript / browser** — types + `fetch` + WebSocket subscriptions
- **Rust** — structs, `tokio` async client, `impl Stream<Item = Delta>`
- **Python** — dataclasses, `async for delta in sub`
- **curl / raw REST** — documented via auto-generated OpenAPI spec at `/openapi.json`
- **GraphQL clients** — SDL auto-generated at `/graphql/schema.graphql`
**Raw `.wo` DML is a first-class escape hatch.** Every client SDK exposes a single method — `client.Wo(ctx, src, params)` in Go, equivalents in TypeScript/Rust/Python — that accepts any `.wo` source the server would accept: mixed SQL + Cypher, `BEGIN … COMMIT` blocks with `RETURNING` aliases threading across statements, ad-hoc MATCH-then-SELECT queries that cross multiple generated types. The typed methods are sugar; the engine speaks `.wo` on the wire. A Go program can send a cross-paradigm transaction as a single string and the server parses + executes it exactly like `wo run` would — see [Phase 5: Go Client SDK](./database/05-go-sdk.md) for the full API.
One schema, every protocol. A browser app, a mobile client, and a background worker can all subscribe to the same live query and receive the same delta stream.
## What this is, and isn't
**Is.** A declarative, full-stack, single-binary language for building CRUD apps with live data. A replacement for the "Go backend + Postgres + Redis + React + Prisma + GraphQL server" stack.
**Isn't.**
- Not a general-purpose language like Rust or Go. You can't write a kernel module or a video codec in `.wo`. The scope is data-shaped applications.
- Not a JavaScript meta-framework. No Node, no React. The UI layer (`##ui`) is declarative and compiles to SSR HTML with a small vanilla-JS client.
- Not a DSL that transpiles to another language. `.wo` has its own lexer, parser, analyzer, and bytecode. The [`prototypes/`db`/`](../../prototypes/`db`/) C++ prototype and the planned Rust crates implement the runtime natively.
- Not a hosted service. Your binary owns its own DB file. No managed cloud offering is required.
## How this maps to the design series
This overview is the user-facing frame. The underlying engineering plan is the 7-phase series linked from [database.md](./database.md):
- **[Phase 2](./database/02-wo-language.md)** designs the language and the transaction coordinator.
- **[Phase 3](./database/03-inmemory-engine.md)** builds the storage engine.
- **[Phase 4](./database/04-client-api.md)** builds the wire protocol and subscription engine.
- **[Phase 5](./database/05-go-sdk.md)** builds the first typed client (Go) and `wo gen`.
- **[Phase 6](./database/06-lowcode-fullstack.md)** adds `##ui` and the application-level blocks.
- **[Phase 7](./database/07-wo-seg-migration.md)** migrates writeonce-the-blog from `wo-seg` onto this runtime.
Phase 1 (evaluation) and the case studies in [surreal-case-study.md](./surreal-case-study.md) argue *why* the language exists at all. Read those first if you're skeptical; read the phase docs if you're implementing; read this page if you want to know what it feels like to use.
## Reference points
The design absorbs lessons from several systems. In order of influence:
- **Go** — toolchain shape, single-binary deployment, "the language is the build system"
- **Phoenix LiveView** — subscription-native UI, hot-reloading dev server
- **SAP CDS** — declarative entity/service language, admin UI generation
- **SurrealDB** — multi-paradigm query language, `LIVE` subscriptions over wire
- **PocketBase** — single-binary CRUD backend (the proof of concept that this is shippable)
- **EdgeDB** — unified type system above storage paradigms
- **Elixir / Erlang / OTP** — hot code loading, supervision, subscription semantics
- **Django** — admin UI as a built-in, not a bolt-on
None of these give you all of: a language, a database, a subscription engine, a UI toolkit, a client codegen, and a single-binary output. writeonce is the attempt to fuse the best of each into one thing.
## Minimal "hello, world" as a full program
If you want a pure procedural test, without the server:
```wo
-- hello.wo
main {
print("hello, world")
}
```
```bash
$ wo run hello.wo
hello, world
```
If you want the database without HTTP:
```wo
type Counter {
name: Text @unique
value: Int = 0
}
main {
insert Counter { name: "visits" };
update Counter{ name == "visits" }.value += 1;
let c = select Counter{ name == "visits" };
print(c.value);
}
```
If you want the full app — database, HTTP, subscriptions, clients — it's the article example at the top of this page.
Three progressive shapes, one language, one command to run each.

View file

@ -53,9 +53,16 @@ iterations); no commits by agents — drafts go to `.dev/commit.md`.
| 8 | [Shard-actor runtime](08-shard-actor-runtime.md) | thread-per-core shards, per-shard heaps, ownership-move messaging | | 8 | [Shard-actor runtime](08-shard-actor-runtime.md) | thread-per-core shards, per-shard heaps, ownership-move messaging |
| 9 | [Database engine](09-database-engine.md) | class-shaped tables, typed WAL + recovery, `insert`/`select` execute | | 9 | [Database engine](09-database-engine.md) | class-shaped tables, typed WAL + recovery, `insert`/`select` execute |
| 9b | [`@table`, relations, query](09b-table-relations-query.md) | `@table` becomes real storage; typed `ref`/`backlink`/`multi` relations; compiler-checked LINQ-shaped queries lowered to engine ops | | 9b | [`@table`, relations, query](09b-table-relations-query.md) | `@table` becomes real storage; typed `ref`/`backlink`/`multi` relations; compiler-checked LINQ-shaped queries lowered to engine ops |
| 9c | [Cross-program tables](09c-cross-program-tables.md) | attach to a running program's database over a local IPC channel: manifest-granted read/write rights, typed statements checked against the owner's shapes, owner stays the single writer |
| 9d | [Keypair attach auth](09d-keypair-attach-auth.md) | program identity is a keypair: mutual challenge–response at attach, grants name public keys, replay-proof, rotation is a config change |
| 9e | [Durability, throughput, scale](09e-durability-throughput-scale.md) | restart-persistence proof, read/write benchmark, ~1M-row load; the measurement gate every optimization signs |
| 9f | [io_uring group-commit](09f-io-uring-commit.md) | replace fsync-per-commit with io_uring batched durability, overlapped on the shard threads; fsync fallback kept |
| 9g | [Query grammar corpus](09g-query-grammar-corpus.md) | grow the query grammar from real embedded-DB apps: `count`/existence subqueries from the skillhost corpus; add only what a corpus uses |
| 10 | [HTTP service layer](10-http-service.md) | `service` blocks route to VM methods; REST parity with Stage 2 | | 10 | [HTTP service layer](10-http-service.md) | `service` blocks route to VM methods; REST parity with Stage 2 |
| 11 | [Fibers](11-fibers.md) | green threads on the shard scheduler: reduction-budget preemption, park on I/O | | 11 | [Fibers](11-fibers.md) | green threads on the shard scheduler: reduction-budget preemption, park on I/O |
| 12 | [Blue-green deploy](12-blue-green-deploy.md) | two VM slots, in-runtime compile, atomic switch, resident rollback | | 12 | [Blue-green deploy](12-blue-green-deploy.md) | two VM slots, in-runtime compile, atomic switch, resident rollback |
| 13 | [Compile-time metaprogramming](13-compile-time-metaprogramming.md) | `@derive(Json/Csv/Eq/Hash/Show)` — the compiler generates per-type code from the class-table metadata; generic capabilities within principle 13, no reflection |
| 14 | [skillhost host workload](14-skillhost-host-workload.md) | a host-shaped driving workload (writeonce port of skillhost) that names the runtime gaps it exposes: bounded/killable subprocess, stdin/stdout transport, fs metadata, and FFI-vs-out-of-process model driver — each a candidate iteration |
Review protocol: the developer reads one iteration, approves or amends; Review protocol: the developer reads one iteration, approves or amends;
the next starts only after approval. Each iteration is an unsplittable the next starts only after approval. Each iteration is an unsplittable

View file

@ -80,6 +80,18 @@
## Info ## Info
- Governing spec: [`docs/superpowers/specs/2026-08-11-inferred-gc-mark-sweep-design.md`](../../superpowers/specs/2026-08-11-inferred-gc-mark-sweep-design.md). - Governing spec: [`docs/superpowers/specs/2026-08-11-inferred-gc-mark-sweep-design.md`](../../superpowers/specs/2026-08-11-inferred-gc-mark-sweep-design.md).
- **Gated by the benchmark (2026-08-15):** this is the "implement garbage
collection" lever of the performance arc — tri-color mark-sweep replacing
RC changes the write path's tail latency, so landing it means re-running
iteration [9e](09e-durability-throughput-scale.md) and recording the
delta (does tracing help or hurt p99 under write load?).
- **Constraint added by the database track (2026-08-15):** a GC-managed value
in a `@table` field is a compile error (the engine/heap bulkhead — 9b
design, section 6). Once GC-ness is inferred rather than annotated, the
inference pass must classify every class **before** table-field validation,
and the diagnostic must name the inference reason ("class X is
garbage-collected via Y and cannot be stored in a table field") — otherwise
the error becomes unactionable exactly when it stops being self-evident.
- **Why the annotation was insufficient, not merely inconvenient:** the OOP - **Why the annotation was insufficient, not merely inconvenient:** the OOP
spec's own example, `@gc class PriceCache { entries: map<SKU, Money> }`, is spec's own example, `@gc class PriceCache { entries: map<SKU, Money> }`, is
acyclic. It needs GC because it is shared, and second-class borrows cannot acyclic. It needs GC because it is shared, and second-class borrows cannot

View file

@ -42,6 +42,13 @@
loops, eventfd mail, the machinery this iteration lifts into `wovm`. loops, eventfd mail, the machinery this iteration lifts into `wovm`.
- The VM's object header has carried a shard id since iteration 2 — no - The VM's object header has carried a shard id since iteration 2 — no
relayout. relayout.
- **Gated by the benchmark (2026-08-15):** this is the "optimize
multithreading" lever of the performance arc — thread-per-core is a
throughput/scale claim, so landing it means re-running iteration
[9e](09e-durability-throughput-scale.md) at the connection/concurrency
scale it unlocks and recording the before/after delta. It is also where
the io_uring write path ([9f](09f-io-uring-commit.md)) gets a thread to
overlap durability against.
## Proposed Solution ## Proposed Solution

View file

@ -47,6 +47,14 @@
lesson, the C engine enforces it. lesson, the C engine enforces it.
- The wo-db overlap manifest keeps the C++ prototype and this engine - The wo-db overlap manifest keeps the C++ prototype and this engine
answer-compatible where features overlap. answer-compatible where features overlap.
- **Ownership/GC analysis (2026-08-15):** the engine and the VM heap are two
memory worlds crossed only by copy — rows store no VM pointers, GC-managed
values in stored fields are a compile error, and everything a query returns
is copied out — so the collector never traces rows and the engine never
counts references. The full analysis (row views as borrows without a
runtime net, cursor stability, GC-pause interaction) lives in the 9b
design's section 6:
[`2026-08-15-table-relations-query-design.md`](../../superpowers/specs/2026-08-15-table-relations-query-design.md).
## Proposed Solution ## Proposed Solution

View file

@ -8,9 +8,14 @@
> precedes iteration 10 because `service` blocks will want to return query > precedes iteration 10 because `service` blocks will want to return query
> results. > results.
> >
> **No spec exists yet.** This iteration frames the outcome and records the > **Spec exists (2026-08-15):**
> open questions; the design must be brainstormed before a plan is written. > [`2026-08-15-table-relations-query-design.md`](../../superpowers/specs/2026-08-15-table-relations-query-design.md)
> The three questions in *Info* are genuine forks, not details. > settles the three forks recorded in *Info* below (kept as the decision
> record): the SQL/Cypher layer is superseded as the program surface,
> the syntax is a compiler-desugared comprehension, and the references
> contribute vocabulary + semantics (System.Linq) and execution + integrity
> vocabulary (PostgreSQL, surveyed with the spec). Plan:
> [`2026-08-15-employee-relations-query.md`](../../plan/compiler/2026-08-15-employee-relations-query.md).
## Goals ## Goals
@ -111,18 +116,65 @@ delegates, and `IQueryable`'s runtime expression trees, the last of which
depends on reflection that principle 13 forbids outright. Take the vocabulary depends on reflection that principle 13 forbids outright. Take the vocabulary
and the semantics; leave the plumbing. and the semantics; leave the plumbing.
**4. (Settled with the spec, 2026-08-15) How do queries interact with the
borrow checker and the GC?** Row views are borrows of engine memory with **no
runtime borrow word behind them** — the compile-time escape rule is
load-bearing alone. Scans materialize their id list up front, so updating a
row (even an indexed column) inside the loop is sound, while `insert`/`delete`
on a table with an open cursor is a compile error. The GC never meets the
engine at all: both directions across the boundary are copies, and GC-managed
values cannot be stored — spec section 6 is the full analysis, including the
iteration-7b ordering constraint (inference before table-field validation).
Also relevant: `@table(name:, index:)` already parses today with known-key Also relevant: `@table(name:, index:)` already parses today with known-key
validation (`WO-E102`), the Rust runtime already ships secondary indexes and validation (`WO-E102`), the Rust runtime already ships secondary indexes and
`find_by` behind that annotation, and `ref T` already classifies as a scalar `find_by` behind that annotation, and `ref T` already classifies as a scalar
id rather than a pointer — so the relational vocabulary partly exists and this id rather than a pointer — so the relational vocabulary partly exists and this
iteration makes it mean something in the C stack. iteration makes it mean something in the C stack.
## Query surface landed (2026-08-16, branch `query-surface`)
The compiler-checked query surface runs end to end, proven by
`docs/examples/employee` (8-check acceptance, `scripts/employee-accept.sh`):
- **Queries**: `from <v> in <table|nav> where* [order by <k> [desc]] [take n]
select <v|v.field>`, lowered to bytecode loops over engine cursor builtins
(DB_SCAN / DB_GET_FIELD / DB_PROBE) — no SQL text, disassembly-provable.
- **Relations**: `ref C` forward navigation (`e.dept.name`, a point read),
`backlink C.f` reverse navigation (`d.staff`, an index probe); backlink
fields are virtual (no stored column).
- **Mutation**: update-through-row (`e.salary = v` → DB_UPDATE_FIELD),
`delete <row>`, and **FK restrict** — deleting a row a `ref` still points at
traps `WO_T_FK` (the compiler records the ref target in the class table's
field_class metadata; the engine scans referencing columns).
- **`@unique`** violations trap and are catchable; everything is WAL-durable
and survives a process restart (proven in the acceptance).
**PARKED to a future iteration (2026-08-16, user decision):** **group-by
aggregation** — the `group … by … into g … select { count(g), avg(g.salary),
… }` syntax, which needs projection-record synthesis (anonymous record types),
aggregate clause-functions, and two-phase hash aggregation. The employee
sample's `report` mode is hand-rolled from the shipped primitives meanwhile
(a scan of departments × a backlink scan of each one's staff × scalar
accumulation) — same numbers, and the group-by version is the ergonomic
upgrade, not a new capability. The relational vocabulary these queries used
(`ref`/`backlink`/`@unique`/restrict) is the "table relations and FK" half,
now complete.
## Proposed Solution ## Proposed Solution
- **Brainstorm a spec first**, settling the three forks above; only then write - ~~Brainstorm a spec first~~ — **done 2026-08-15**; the spec settles all
the plan. This iteration deliberately ships no plan pointer, because three forks and the plan exists (pointers in the header note). The fork-1
choosing between "replace the SQL layer" and "sit beside it" changes what outcome for the record: language-integrated query is the only program
the plan contains. surface; `docs/runtime/database/02-wo-language.md`'s SQL/Cypher layer stays
as design history and as the `wo-db` prototype's engine-semantics
reference, never as syntax.
- **The acceptance workload is a new sample**: `docs/examples/employee` —
`Department`/`Employee` with `@unique`, a composite index, a `ref`/
`backlink` pair, and a report mode that is one `GROUP BY` after another
(headcount, avg/min/max salary by department). It is 9b's acceptance the
way log-watcher was iterations 1–7's; the ecommerce query rewrite (the
fifth criterion below) follows as its own step once employee is green.
- Study `.dev/reference/dotnet-runtime`'s `System.Linq` operator set for the - Study `.dev/reference/dotnet-runtime`'s `System.Linq` operator set for the
vocabulary, and `docs/runtime/database/02-wo-language.md` plus vocabulary, and `docs/runtime/database/02-wo-language.md` plus
`prototypes/wo-db/` for the semantics already committed to. `prototypes/wo-db/` for the semantics already committed to.

View file

@ -0,0 +1,171 @@
# Iteration 9c — cross-program tables: attach to a running program's database
> Format: `product/story-iteration-template`. Part of
> [Story — one language, one runtime, one database, one binary](00-story.md).
>
> **Inserted 2026-08-15**, hence `9c`. It follows 9b because a program
> attaching to another's tables wants the same typed statements and queries
> the owner has — a surface that must exist before it can be shared — and
> precedes iteration 10 because HTTP is the *external* face of a program;
> this iteration is the *writeonce-native* face, program to program on the
> same machine.
>
> **No spec exists yet.** This iteration frames the outcome and records the
> open forks; the design must be brainstormed before a plan is written. The
> forks in *Info* are genuine decisions, not details.
## Goals
- A running program **A** with a persistent database (`WO_DATA`, iteration 9)
can be **attached** by a second writeonce program **B**: B names A's
IPC connection string in its own `wo.toml`, and from then on reads and
writes `A.Table` rows with the same typed statements it uses on its own
tables — checked by B's compiler against A's declared table shapes.
- **A stays the single writer.** B never opens A's WAL, never maps A's
slabs: every statement B issues travels the IPC channel and executes
inside A's engine, through the same choke-point row API A's own
statements use. The ownership doctrine survives contact with a second
process because the second process never touches the memory.
- **Access is granted, never assumed.** A's manifest *registers* B by name
with explicit rights (read, or read+write); an unregistered client is
refused at connect, an under-privileged statement is refused at execute
with a trap B can catch. No registration, no access — including on the
same uid.
## Acceptance Criteria
- What to achieve?
- **Given** A running with `[share]` registering client "b" as
read+write, and B's `wo.toml` carrying `[connect.a]` with A's IPC
string,
- **when** B executes `insert a.AuditLog { … }` and a query over
`a.AuditLog`,
- **then** the row exists in A (visible to A's own queries, WAL-logged
before B's insert acknowledges), and B's query returns it — with B's
compiler having checked every field name against A's declared shape.
- What to achieve?
- **Given** A registers client "c" as read-only,
- **when** C executes a query it succeeds, and when C attempts an
insert,
- **then** the insert traps with the access-denied code inside C
(catchable), and A's log records the refusal; nothing was applied,
nothing was WAL-logged.
- What to achieve?
- **Given** a program with no registration in A's manifest,
- **when** it presents A's IPC string and attempts to attach,
- **then** the connect itself is refused — rights are checked at the
door, not per statement only.
- What to achieve?
- **Given** B attached and mid-statement,
- **when** A shuts down cleanly (SIGTERM) or crashes,
- **then** B's in-flight statement traps with a connection error B can
catch (never a hang), and B can re-attach after A reboots and
replays — with every previously acknowledged write still present.
- What to achieve?
- **Given** the employee sample running as A with its departments and
employees tables,
- **when** a second sample program (a thin reporting client) attaches
read-only and runs the GroupBy report over `a.Employee`,
- **then** it prints the same report the owner prints — the
demonstration that attach + query compose.
## Out Of Scope
- **Remote machines.** The IPC string names a local channel; cross-host
access is the HTTP/service layer's job (iteration 10) or a much later
network protocol. Same-machine is what "attach" means here.
- **B caching A's rows.** Every read crosses the channel; a client-side
cache (and its invalidation) is a later performance iteration, if ever.
- **Cross-program transactions.** A statement is atomic inside A exactly as
A's own statements are; B cannot open a transaction spanning its own
tables and A's. That is 2PC territory, recorded with the database track's
deferred items.
- **`LIVE` subscriptions over the channel** — composes with the
subscription registry later (the client-api phase doc already sketches
the wire shape).
- **Schema migration while attached** — a blue-green swap in A while B
holds an attachment is iteration 12's compatibility problem; this
iteration may simply drop attachments on swap.
## Info
Prior art in the tree: `docs/runtime/database/04-client-api.md` already
designs a native binary wire protocol for external clients (length-prefixed,
typed, subscription-ready) — this iteration's channel should be its
same-machine profile, not a new invention. The WAL's typed value encoding
(`database/src/wal.c`, iteration 9 Task 2) is a working engine-value wire
format today: statements and rows can ride the same encoding the log already
uses. The `wo.toml` manifest exists and is compiler-read (`woc <dir>`), so
both ends' declarations have a natural home.
Forks the spec must settle:
**1. What carries the channel — and what does the IPC string name?**
A unix domain socket is the obvious carrier (peer credentials for free,
`net`-stdlib adjacency); the string would be `unix:/path/a.sock` in B's
`[connect.a]` and A would listen beside its `WO_DATA` directory. The
alternatives — a FIFO pair, shared memory + doorbell — buy latency at the
cost of the credential story and the crash-detection story (a dead socket
peer is unambiguous; a dead shm peer is a protocol). Leaning: unix socket,
one connection per attached client, A serving requests on its event loop
(iteration 8's shard-actor loop when it lands; a dedicated accept loop
until then — which is also the fork's dependency question: how much of
iteration 8 does this need?).
**2. How does B's compiler know A's table shapes?** B typechecks
`a.Employee { … }` against A's declarations, so B needs them at compile
time. Options: B's `[connect.a]` names A's **project directory** and `woc`
reads A's types straight from A's source (simple, but couples B's build to
A's checkout); A **exports a schema file** (a `.wob`-adjacent digest of its
class table) that B's manifest points at (decoupled, but a new artifact
with a staleness story); or shared type definitions in a common module both
import (cleanest language story, needs the module system to span projects).
A runtime schema handshake must exist regardless — B's compiled expectation
of `a.Employee`'s shape is verified against A's live class table at attach,
and a mismatch refuses the attachment with both sides' shapes named.
Leaning: project-directory reference for the milestone plus the mandatory
handshake; the export artifact when the staleness story matters.
**3. What exactly does A's registration grant?** The request's shape is
per-client rights: `[share] clients = [{ name = "b", rights = "rw" }]` or
per-table refinement (`tables = ["AuditLog"]`). Identity: the client NAME
must be bound to something a peer cannot fake — unix peer credentials
(uid), a token A mints, or both. Leaning: name + uid via `SO_PEERCRED` for
the milestone (same-machine, same-trust-domain), rights whole-database
read or read+write (per-table refinement deferred until a workload needs
it), and the registration is A's manifest so a grant is a config change +
restart, not an API. **Superseded as the end state (2026-08-15):**
identity is a keypair and grants name public keys — iteration
[9d](09d-keypair-attach-auth.md) owns that; the uid check is only this
iteration's bootstrap and must be flagged pre-9d wherever it ships.
**4. What does B's statement actually block on?** B's insert crosses the
channel, executes in A (RAM + WAL + fsync), and acknowledges back — a
blocking round-trip on B's thread, exactly like B's own `WO_DATA` inserts
block on their own fsync. Queries stream results back whole (materialized;
no cursors over the wire this iteration). The alternative — async
statements with completion callbacks — has no language surface to stand on
(no function values) and waits for fibers (iteration 11). Leaning:
blocking, with the stop-flag rule from the log-watcher work applying (a
SIGTERM'd B parked on a channel read exits cleanly).
## Proposed Solution
- **Brainstorm the spec first**, settling the four forks; then a plan.
Expected shape: A-side — a listener beside the engine, a request
dispatcher that executes through the same statement executors iteration
9 built (`database/src/db.c`), the registration check at accept and per
statement; B-side — `[connect.<name>]` manifest surface, compiler
namespace `<name>.Table` binding table statements/queries to channel
stubs instead of local engine builtins; both — the client-api phase
doc's wire protocol, profiled for unix sockets, values in the WAL's
encoding.
- **The acceptance workload extends the employee sample**: A = the employee
program with `[share]`; B = `docs/examples/employee-list` (pre-authored
2026-08-15, sample-first — both manifests designed as a pair), attaching
read-only for the list/report/staff modes and proving the rights matrix
with its `probe-write` mode. The sample stays the test.
- Depends on iterations 9 (engine, WAL — done through Task 3 as of
2026-08-15) and 9b (typed statements and queries worth sharing); wants
iteration 8's event loop for A's serving side but can prototype on a
dedicated accept loop the way the MCP sample serves today.

View file

@ -0,0 +1,146 @@
# Iteration 9d — keypair authentication for cross-program attach
> Format: `product/story-iteration-template`. Part of
> [Story — one language, one runtime, one database, one binary](00-story.md).
>
> **Inserted 2026-08-15.** Promotes iteration 9c's identity fork (Info,
> fork 3) to its own iteration: the name + unix-uid lean is the milestone
> bootstrap, and THIS is what replaces it — program identity is a keypair,
> and an attachment is granted to a public key, not to a process that
> happens to share a uid. It follows 9c (there is nothing to authenticate
> until attach exists) and stays same-machine; the same handshake is what
> a future remote channel would reuse, which is the point of doing it
> properly now.
>
> **No spec exists yet.** The forks in *Info* are genuine decisions.
## Goals
- **A program's identity is a keypair.** Each writeonce program owns a
private key (generated once, stored beside its data, never in the
manifest) and a public key it can print/export. Identity stops being
"whoever reached the socket first with the right uid".
- **Grants name public keys.** A's `[share]` registers a client by its
public key (fingerprint), with rights exactly as 9c defined them; B's
`[connect.a]` **pins A's public key** beside the IPC string. Both sides
authenticate: A proves it is A before B sends a byte of intent, B proves
it is B before A executes a statement.
- **The handshake is mutual challenge–response, replay-proof**: fresh
nonces each attach, signatures over the nonce + channel binding, no
secret ever crosses the channel. A failed handshake refuses the
attachment with a catchable trap on the connecting side and one log line
naming the offered fingerprint on the listening side.
## Acceptance Criteria
- What to achieve?
- **Given** A's `[share]` registering B's public-key fingerprint with
read+write, and B's `[connect.a]` pinning A's public key,
- **when** B attaches,
- **then** the mutual handshake completes, the attachment carries B's
granted rights, and every 9c acceptance behavior (statements, traps,
refusals) holds unchanged on top of it.
- What to achieve?
- **Given** a client presenting a keypair A never registered,
- **when** it attempts the handshake,
- **then** the attach is refused before any statement is read, the
client sees the catchable authentication trap, and A logs the offered
fingerprint (so granting it is a copy-paste, not an investigation).
- What to achieve?
- **Given** a same-uid process (the 9c bootstrap's whole trust basis)
presenting no key or the wrong key,
- **when** it attempts to attach,
- **then** it is refused — proving the uid check has been superseded,
not merely supplemented.
- What to achieve?
- **Given** an impostor listening on A's socket path (or a swapped
socket file),
- **when** B attaches and the impostor cannot sign A's challenge
response with A's private key,
- **then** B aborts before sending any statement or data, with a trap
that names the fingerprint mismatch — the pinned-key check working in
the B→A direction.
- What to achieve?
- **Given** a recorded handshake transcript from a legitimate attach,
- **when** it is replayed against A,
- **then** the attach is refused — the nonce is fresh per handshake and
a signature over an old nonce proves nothing.
- What to achieve?
- **Given** A rotates B's registered key (manifest update + restart),
- **when** B attaches with the old key and then with the new one,
- **then** the old key is refused and the new one works — rotation is a
config change, exactly like the grant itself.
## Out Of Scope
- **Transport encryption.** Same-machine unix sockets; the kernel is the
wire. Session encryption (and the key exchange it needs) arrives with a
remote channel, if one ever ships — this iteration's handshake is
designed not to preclude it, nothing more.
- **Certificate hierarchies, expiry, revocation lists.** A grant is a
public key in a manifest; revocation is deleting the line and
restarting. CA machinery has no workload here.
- **Key escrow / multi-key identities / agent forwarding.** One program,
one keypair.
- **Protecting the private key from a root attacker or from the program's
own uid.** File permissions (0600) are the boundary this iteration
claims; anything stronger (TPM, keyring) is explicitly not promised.
## Info
Forks the spec must settle:
**1. Where does the crypto come from?** The runtime is libc-only by
doctrine, and hand-rolling signature crypto is the one wheel nobody gets to
reinvent. The realistic options: **vendor a compact, audited Ed25519**
implementation (TweetNaCl-lineage, a few files, no allocation, no OS
dependencies) into `database/src/` or a new `vendor/`; or take libsodium as
the first external dependency and break the doctrine openly. Leaning:
vendored compact Ed25519, recorded as the single sanctioned vendored
component with its provenance pinned in the tree — the doctrine's spirit is
"no dependency sprawl", not "write your own constant-time field
arithmetic".
**2. Key generation and storage.** Options: a `woc keygen` subcommand
(keys are a toolchain concern), or first-boot generation by the runtime
into the data directory (keys are a runtime concern, zero setup). Leaning:
first-boot generation into `WO_DATA` (0600, alongside the WAL — a program
with a persistent database already has the directory), plus a way to print
the public fingerprint (`program --identity` or a stdlib call) so the
operator can paste it into A's `[share]`. A program without `WO_DATA` has
no identity and cannot attach anywhere — which is coherent: attach is a
database feature.
**3. What exactly gets signed.** A bare nonce signature is vulnerable to
cross-protocol reuse; the lean is signing a transcript hash: protocol tag,
both fingerprints, both nonces, and the channel identity — so a signature
from this handshake means nothing in any other context. The spec should
write the exact byte layout down (the WAL encoding conventions apply: the
format is normative, little-endian, versioned by the protocol tag).
**4. Does the uid check survive at all?** Options: keys only (one
mechanism, one story), or keys AND peer-cred as defense in depth. Leaning:
keys only — two mechanisms invite "it worked because of the other one"
confusion in exactly the code that must never be confusing; `SO_PEERCRED`
remains a log-line enrichment (who was that fingerprint), never an
authorization input.
## Proposed Solution
- **Brainstorm the spec** settling the four forks, then fold the plan into
9c's implementation plan as its authentication tasks — one plan, because
9c without 9d ships a placeholder identity and 9d without 9c has nothing
to authenticate. The 9c milestone may still land first with the uid
bootstrap, flagged loudly as pre-9d.
- **Acceptance extends the 9c workload**: the employee-A /
employee-list-B pair (`docs/examples/employee-list`, pre-authored
2026-08-15) carries the key exchange in both manifests — A's
`[[share.clients]]` names B's fingerprint, B's `[connect.employee]` pins
A's; the acceptance script adds the wrong-key, no-key,
same-uid-wrong-key, replay, impostor-socket, and rotation checks above,
each asserting the exact trap/refusal.
- Expected shape: handshake module beside the channel code (both ends),
`[share]`/`[connect]` manifest keys for fingerprints, first-boot keygen
in the runtime's data-directory setup, vendored signature primitive with
its own unit suite (known-answer tests from the algorithm's reference
vectors).

View file

@ -0,0 +1,136 @@
# Iteration 9e — durability proof, throughput, and scale under load
> Format: `product/story-iteration-template`. Part of
> [Story — one language, one runtime, one database, one binary](00-story.md).
>
> **Inserted 2026-08-15.** The measurement backbone. Everything after the
> functional engine (9/9b) is an *optimization*, and an optimization without
> a number is a guess — this iteration is the number. It comes before the
> optimization iterations (7b GC, 8 shard-actor, 9f io_uring) reopen for
> performance work, because each of those must be gated by re-running THIS
> iteration's benchmark and showing the number moved the right way.
>
> **No spec exists yet.** The forks in *Info* are genuine decisions.
## Goals
- **Durability is proven by a restart, not asserted.** The employee program
(iteration 9b) runs, writes rows, is stopped and restarted, and every
acknowledged write is present after replay — the WAL's promise turned into
a scripted acceptance on a real program, not just the unit-level crash
battery.
- **Read and write throughput are measured, published, and defended.** A
repeatable benchmark drives the engine through the language (not the C
API): inserts/sec, point-reads/sec, indexed-query/sec, each with p50/p99
latency, recorded in the tree so a regression is a diff.
- **The scale target is a gate, not a slogan.** "A million users can read and
write" becomes a concrete load: a dataset of ~1M rows across the sample's
tables, a mixed read/write workload at a stated concurrency, sustained for
a stated duration, with throughput and tail latency inside a stated budget
and RSS flat (the log-watcher soak discipline, at database scale).
- **The benchmark is the contract every later optimization signs.** 7b (GC),
8 (shard-actor threads), and 9f (io_uring) each re-run this and record the
before/after — no optimization lands without a measured delta.
## Acceptance Criteria
- What to achieve?
- **Given** the employee program seeded with data and then stopped,
- **when** it is restarted and queried,
- **then** every acknowledged row is present with its exact contents,
the ids continue past the persisted maximum, and a query that used an
index before the restart uses it after (the index was rebuilt on
replay).
- What to achieve?
- **Given** the benchmark harness driving inserts, point reads, and
indexed queries through compiled `.wo`,
- **when** it runs to completion,
- **then** it reports ops/sec and p50/p99 for each operation class, writes
the numbers to a tracked results file, and fails if any number crosses
a recorded regression threshold.
- What to achieve?
- **Given** ~1M rows and a mixed read/write workload at the target
concurrency held for the target duration,
- **when** it runs,
- **then** throughput stays above the floor, p99 stays under the ceiling,
RSS is flat between a warmed baseline and the end (no growth beyond
tolerance), zero descriptors leak, and — for a write-inclusive run under
a durable configuration — a kill mid-load followed by replay loses no
acknowledged write.
- What to achieve?
- **Given** any later optimization iteration (7b, 8, 9f),
- **when** it claims a speedup,
- **then** this benchmark's before/after numbers are in that iteration's
record, and a claim with no measured delta is not accepted.
## Out Of Scope
- **The optimizations themselves.** This iteration MEASURES; 7b/8/9f change.
A single-thread RAM-authoritative baseline is a legitimate first number —
the point is to have one before anyone tunes.
- **Distributed / multi-machine load.** Same-machine, one process (or one
process per shard once iteration 8 lands). Cross-host is the network layer's
concern, much later.
- **Micro-optimizing the benchmark harness.** It must be honest and
repeatable, not itself fast; if the harness is the bottleneck the spec says
so and fixes that, but a perfect load generator is not the deliverable.
- **A cost-based query planner.** Index selection is 9b's; this iteration
measures what 9b lowers, it does not make the planner smarter.
## Info
Forks the spec must settle:
**1. What generates the load, and in what language?** The doctrine is "the
sample is the test", so the honest generator drives compiled `.wo` — a
benchmark mode in the employee program (or a sibling sample) that loops
inserts/reads/queries and times them. The alternative — a C harness calling
the engine API directly — measures the engine but skips the compiler's
lowering, which is exactly the layer a language-integrated query has to pay
for. Leaning: `.wo` benchmark mode for the headline numbers (the number that
matters is end to end), with the C-API microbench kept only to attribute a
regression to engine vs lowering.
**2. What are the actual budgets?** Throughput floors and latency ceilings
have to be numbers, and the first run sets them — but the spec must decide
whether the gate is absolute (">= N ops/sec on the reference machine") or
relative ("no worse than the last recorded run by more than X%"). Absolute
gates rot across machines; relative gates need a committed baseline file.
Leaning: relative gates against a tracked `bench/baseline.json`, refreshed
deliberately with a commit that says why, plus a loud absolute floor so a
catastrophic regression fails even on a slow machine.
**3. What does "1M users read and write" concretely mean?** A million
long-lived idle connections is a different test from a million rows under a
churning read/write mix from a bounded connection pool. The sample's shape
(departments, employees) suggests rows, not connections, as the scale axis
for THIS iteration; the connection-scale test belongs with the shard-actor
runtime (iteration 8) and the eventual network layer. Leaning: ~1M rows +
a bounded concurrent read/write workload here; connection scale deferred to
8 with a cross-reference.
**4. Durable or RAM-only for the throughput headline?** fsync-per-commit
(the current per-statement durability) will dominate write throughput and is
the honest number for a durable workload; RAM-only (no `WO_DATA`) measures
the engine's ceiling. Both matter and mean different things. Leaning:
publish both, labeled — durable is the number an operator plans against, and
the gap between them is precisely what iteration 9f (io_uring group-commit)
exists to close.
## Proposed Solution
- **Brainstorm the spec**, settling the four forks; then a plan whose first
task is the harness and the baseline file, because nothing downstream means
anything without them.
- **Sequence the whole performance arc around this iteration:**
1. 9b lands → employee compiles and runs → **9e restart-persistence** and
**9e baseline benchmark** (single-thread, both durable and RAM-only).
2. **7b** (inferred GC + mark-sweep) → re-run 9e, record the delta (does
tracing change the write path's tail latency?).
3. **8** (shard-actor, thread-per-core) → re-run 9e at the connection/
concurrency scale it unlocks, record the delta.
4. **9f** (io_uring group-commit) → re-run 9e's durable write number, record
the delta against the fsync-per-commit baseline — the payoff.
- The benchmark harness and its baseline live under `bench/` (or the existing
`runtime/bench/`), and `just` gets a `db-bench` recipe kept off the fast
path, exactly like `log-watcher::soak`.

View file

@ -0,0 +1,116 @@
# Iteration 9f — io_uring group-commit write path
> Format: `product/story-iteration-template`. Part of
> [Story — one language, one runtime, one database, one binary](00-story.md).
>
> **Inserted 2026-08-15.** The write-path optimization, and deliberately the
> LAST database performance iteration: it only earns its complexity once
> there is a measured fsync-per-commit baseline to beat (iteration 9e) and a
> multithreaded runtime to overlap against (iteration 8). Doing it earlier
> would optimize a number nobody had measured, against a runtime that
> couldn't use it.
>
> **No spec exists yet.** The forks in *Info* are genuine decisions.
## Goals
- **Replace fsync-per-commit with io_uring group-commit** on the WAL write
path: batch a tick's committed records into one submission, let the kernel
overlap the write and the durability barrier, and acknowledge each writer
only after the barrier its record rode has completed — the same
ack-after-durable contract, at a fraction of the syscall cost.
- **Overlap durability with work.** With the shard-actor runtime
(iteration 8) the shard thread submits its batch and keeps executing ready
statements while the ring drains, instead of blocking one thread on one
fdatasync — the multithreading the throughput number has been waiting for.
- **Keep the durability promise byte-for-byte.** Every guarantee iterations 9
and 9e proved — replay-whole-or-not-at-all, torn-tail drop, no
acknowledged write ever lost — holds identically; io_uring changes HOW the
bytes reach the platter, never WHETHER an ack means durable.
## Acceptance Criteria
- What to achieve?
- **Given** the io_uring write path under the iteration-9e crash battery
(concurrent writers, kill -9 mid-stream, reboot, replay),
- **when** it runs,
- **then** every acknowledged write is present after replay and no
unacknowledged partial write is ever visible — the exact result the
fsync path gives, so durability is provably unchanged.
- What to achieve?
- **Given** the iteration-9e durable write benchmark,
- **when** it is run on the fsync-per-commit path and then the io_uring
group-commit path on the same machine,
- **then** the io_uring path's write throughput is materially higher and
its p99 commit latency lower, with the before/after numbers recorded —
the payoff, measured, not asserted.
- What to achieve?
- **Given** a kernel without io_uring (old, or restricted by seccomp),
- **when** the runtime starts,
- **then** it falls back to the pwrite + fdatasync path automatically and
correctly — io_uring is an accelerator, never a hard dependency, and a
binary that runs everywhere is the whole project's premise.
## Out Of Scope
- **io_uring for the network/accept path.** This iteration is the WAL write
path only; the socket side is the shard-actor runtime's and the network
layer's concern.
- **io_uring for reads.** RAM is authoritative — reads never touch a
descriptor (phase-B doctrine), so there is nothing to accelerate on the
read path. This is a write-durability optimization, full stop.
- **Registered buffers / fixed files / SQPOLL tuning** beyond what the
benchmark shows is worth it. Start with the plain submit/complete model;
add ring features only when 9e's number says a specific one pays.
- **Replacing the WAL format or the commit contract.** The bytes on disk and
the meaning of an ack are iteration 9's; this changes the syscall, not the
format.
## Info
Forks the spec must settle:
**1. How much of the ring model, and behind what abstraction?** The write
path today is `pwrite` + `fdatasync` in `database/src/wal.c`; io_uring adds a
submission/completion queue and a durability barrier op
(`IORING_OP_FSYNC`/`IORING_FSYNC_DATASYNC` or `O_DSYNC` writes). The fork:
wrap it behind the existing `wo_wal_commit` boundary (drop-in, the engine
never learns) or expose an async-commit primitive the shard scheduler drives
(faster overlap, but couples the WAL to iteration 8's loop). Leaning:
drop-in behind `wo_wal_commit` first — it is the correctness-preserving
step and 9e can measure it standalone — then an async variant only if 8's
scheduler shows the blocking boundary is the remaining bottleneck.
**2. liburing or raw syscalls?** liburing is the ergonomic wrapper but is a
new external dependency, against the libc-only doctrine; the raw
`io_uring_setup`/`io_uring_enter` syscalls are a few hundred lines and keep
the doctrine. Leaning: raw syscalls (the doctrine is load-bearing and this is
a bounded surface), with the mmap'd ring setup written down in the binding
doc the way the WAL format is — normative, versioned.
**3. What is the batch boundary?** Per-statement commit (today) is the
simplest correct thing and the slowest; a group commit needs a boundary — a
tick (iteration 8's scheduler quantum), a count, or a short time window.
Leaning: the shard tick once iteration 8 lands (a batch is "everything
committed this tick"), with a single-writer fallback that batches whatever
accumulated between one `wo_wal_commit` call and the ring draining.
**4. How is the fallback chosen and tested?** A kernel probe at startup
(attempt `io_uring_setup`, fall back on ENOSYS/EPERM) is the mechanism; the
question is how CI proves BOTH paths without two kernels. Leaning: an
environment override (`WO_WAL_MODE=fsync|uring`) so the test matrix runs the
crash battery and the benchmark on both on any capable machine, and the
auto-probe is what production uses.
## Proposed Solution
- **Brainstorm the spec** after iterations 8 and 9e exist — this iteration is
meaningless without a multithreaded runtime to overlap against and a
measured baseline to beat, and its plan's acceptance is literally "9e's
durable number improved, 9e's crash battery still green, fsync fallback
still correct".
- Expected shape: a `wo_wal` write-mode switch (fsync vs uring), the raw ring
setup + submit/complete in `database/src/wal.c` (or a `wal_uring.c`
beside it), the startup probe + `WO_WAL_MODE` override, the binding doc's
WAL section extended with the ring layout, and iteration 9e re-run on both
paths with the delta committed.

View file

@ -0,0 +1,161 @@
# Iteration 9g — query grammar, driven by real embedded-DB corpora
> Format: `product/story-iteration-template`. Part of
> [Story — one language, one runtime, one database, one binary](00-story.md).
>
> **Inserted 2026-08-16.** A query-surface iteration in the 9b family: the
> language-integrated query grows to cover the grammar that *real
> applications backed by an embedded SQL database actually use* — measured,
> not guessed, by cataloguing a real app and adding only the constructs it
> depends on. The method is the postgres/System.Linq reference pattern
> applied to a whole application: **an embedded-SQLite app is a grammar
> corpus; each one analysed drives a grammar increment.**
>
> **No spec exists yet.** The forks in *Info* are genuine decisions.
## Why this iteration exists
The 9b query surface ships scan / `where` / `select` / `order by` / `take`,
`ref`/`backlink` navigation, insert / update / delete, `@unique`, and FK
restrict (all running, `docs/examples/employee`). What it does NOT yet cover
is everything past that — and "everything" is unbounded, so the sensible way
to choose the *next* grammar is to point at a real program that uses an
embedded SQL database and add exactly what it needs.
**Corpus #1: `~/projects/skillhost`** (a C++ MCP host, embedded SQLite as its
in-memory skill catalog; surveyed 2026-08-16). Its entire SQL footprint is
one file (`src/catalog/catalog.cpp`, 172 lines): one table + index, a
single-row parameterized `INSERT`, and four `SELECT`s. Mapping each statement
to the writeonce query surface:
| skillhost statement | writeonce today |
| --- | --- |
| `INSERT INTO skills (…) VALUES (?,…)` | ✅ `insert Skill { … }` |
| `SELECT … WHERE name = ?` | ✅ `from s in Skill where s.name == n select s` (unique-index probe) |
| `SELECT … ORDER BY name` | ✅ `from s in Skill order by s.name select s` |
| `SELECT COUNT(*) FROM skills` | ❌ **whole-query `count`** |
| `SELECT … WHERE NOT EXISTS (SELECT 1 FROM skills c WHERE c.parent = s.name) ORDER BY name` | ❌ **correlated `not exists` subquery** |
So the real grammar gap this corpus demands is **two constructs**, and —
importantly — *neither is the parked full group-by/projection machinery*.
Everything else SQLite offers (JOIN, HAVING, LIMIT/OFFSET, DISTINCT, CTE,
window functions, UNION, UPSERT, RETURNING, JSON1, FTS5, triggers, generated
columns) skillhost does not touch, so none of it is in scope here.
## Goals
- **Whole-query `count`**: `count(from s in Table [where …] select …)` yields
an `Int` — the trivial, group-free special case of aggregation (materialize
the query, take its length). It is a stepping stone toward, and independent
of, the parked group-by aggregation.
- **Existence subqueries**: `exists(<query>)` and `not exists(<query>)` as a
boolean, usable in a `where` guard, where the inner query may reference the
outer range variable (a *correlated* subquery — skillhost's roots-of-the-
tree query). Short-circuits: existence needs only the first matching row.
- **Parity, proven by translation**: a new `docs/examples/skill-catalog`
sample mirrors skillhost's schema and expresses all five of its statements
in writeonce, producing results identical to what skillhost's SQLite
returns for the same data.
## Acceptance Criteria
- What to achieve?
- **Given** `count(from s in Skill select s)` and
`count(from s in Skill where s.parent == nil select s)`,
- **when** compiled and run,
- **then** each yields the correct row count as an `Int`, lowered to a
materialize-then-length over the existing scan/where loop — no group
machinery, provable by disassembly.
- What to achieve?
- **Given** `from s in Skill where not exists(from c in Skill where
c.parent == s.name select c) order by s.name select s` — the roots of
the skill tree,
- **when** run over a catalog with parent/child skills,
- **then** it returns exactly the childless skills in name order, and the
inner query correctly sees the outer `s` (correlation), matching
skillhost's `NOT EXISTS` result row-for-row.
- What to achieve?
- **Given** the `skill-catalog` sample seeded with the same rows a
skillhost session would load,
- **when** each of skillhost's five catalog operations is run through the
writeonce translation,
- **then** every result matches, and the sample's README records the one
translation choice made (see fork 1).
## Out Of Scope
- **Full group-by aggregation** (`group … by … into g … select { count(g),
avg(g.f) }`) — still parked (9b's deferral). Whole-query `count` here is the
degenerate case, not the general one; `sum`/`avg`/`min`/`max` as query
aggregates ride with the group-by iteration.
- **Every SQL construct skillhost does not use**: JOIN, HAVING, LIMIT/OFFSET
(writeonce has `take`; `skip`/offset waits for a workload), DISTINCT, CTE /
`WITH RECURSIVE`, window functions, UNION/INTERSECT/EXCEPT, UPSERT /
`ON CONFLICT`, `RETURNING`, multi-row `VALUES`, `INSERT … SELECT`, JSON1
operators, FTS5, triggers, generated columns, explicit collation. Each
enters only when a corpus demands it — that is this iteration's whole
method.
- **A resident SQL parser** — the doctrine stands: skillhost is a grammar
*corpus* to translate against, never a syntax writeonce adopts. No SQL text
in a compiled image.
- **Subqueries in general** beyond correlated `exists`/`not exists` (e.g. a
subquery producing a value, `IN (subquery)`, scalar subqueries) — add when
a corpus uses them.
## Info
Forks the spec must settle:
**1. Does skillhost's `NOT EXISTS` even need a subquery, or does a `backlink`
express it?** skillhost's `skills` table is self-referential (`parent` → a
`name`), and its roots query is "skills no other skill names as parent." In
writeonce that is naturally a **backlink emptiness**: give `Skill` a
`children: backlink Skill.parent` and write `where len(s.children) == 0` — no
subquery at all, using machinery that already exists (backlink probe) plus a
`len` on the result. So the corpus may be fully expressible *today* once
`count`/`len` over a query lands, making `exists` strictly optional for
skillhost. Leaning: ship whole-query `count`/`len` (needed regardless), and
add `exists`/`not exists` as the general construct for correlations a backlink
cannot express (a correlation on a non-relation column) — but let the
`skill-catalog` sample use the idiomatic backlink form for its roots query and
record the subquery form as the alternative. This keeps the new surface
minimal and honest about what the corpus actually forces.
**2. `count` vs `len`.** writeonce already has `len`/`count` builtins on a
`multi`. A query yields a `multi`, so `len(from … select …)` may already work
with no new surface at all — the "gap" could be purely that a bare query in
argument position typechecks and lowers. Leaning: verify `len(<query>)` works
end to end first; if it does, whole-query `count` is a documentation/alias
matter, not new code, and the only real new construct in this iteration is the
existence subquery (fork 1's optional half). The spec must confirm this
against the running compiler before committing scope.
**3. Correlated-subquery execution.** If `exists` lands, the inner query
references the outer row, so it re-evaluates per outer row (a nested loop) or
uses the referenced index. Leaning: nested-loop for correctness first (the
data is small; skillhost's catalog is dozens of skills), index-backed probe as
the optimization the spec records — mirroring how 9b did scans before index
selection.
**Method note (the durable part):** this iteration establishes the pattern for
*all* future query-grammar growth — catalogue a real embedded-DB application,
add only the constructs it uses, translate its statements 1:1 as the
acceptance, and park the rest by name. skillhost is corpus #1 and, tellingly,
needs almost nothing beyond what 9b already shipped — which is itself the
strongest evidence that the 9b surface was scoped right.
## Proposed Solution
- **Brainstorm the spec** settling the three forks — especially forks 1/2,
which may collapse the iteration to "confirm `len(<query>)` works + add
`exists`," a very small increment.
- **Acceptance workload**: `docs/examples/skill-catalog` — `Skill { name:
Text @unique, description: Text, location: Text, root: Text, parent: ?ref
Skill, children: backlink Skill.parent }` and a CLI mirroring skillhost's
catalog operations (add, get-by-name, list, list-roots, count), each a
direct translation of the corresponding SQLite statement, with an
acceptance script asserting the same results skillhost produces.
- Expected shape: small parser/type/emit additions for `exists`/`not exists`
(a query in boolean position, correlated), whole-query `count`/`len`
confirmed or wired, and the sample + script. No engine format change beyond
what 9b already appended.

View file

@ -0,0 +1,167 @@
# Iteration 13 — compile-time metaprogramming (derive from the class table)
> Format: `product/story-iteration-template`. Part of
> [Story — one language, one runtime, one database, one binary](00-story.md).
>
> **Inserted 2026-08-16.** A language-capability iteration, deliberately
> numbered to echo the principle it lives inside:
> [principle 13, "statically typed, all the way to the register"](../../00-principles.md).
> It comes late because it earns its keep only once there are enough
> types worth deriving over (the `@table` classes of iteration 9/9b, the
> records the query surface projects), and it must never be the excuse that
> re-opens a dynamic hole.
>
> **No spec exists yet.** The forks in *Info* are genuine decisions.
## Why this iteration exists
Principle 13 forbids runtime reflection: no `Dynamic`, no runtime type tags,
no walking an unknown value's fields at run time. That ban is correct — the
untagged VM, the borrow checker, and the ORM all stand on it. But it leaves a
real gap: a *generic* capability like "serialize any type to CSV" cannot be a
user-written function, because such a function would need to enumerate a
value's fields at run time, which is exactly what is forbidden. Today the only
escape is a **hand-written function per type**, or a **single compiler
builtin** (`json.encode`) that already does the right thing — it is lowered by
the compiler and walks the class-table metadata (`field_names` / `field_class`
/ `field_elem`, `.wob` v2/v3), never a runtime type tag.
Rust faced the identical ban (it has no runtime field reflection either) and
answered with **compile-time metaprogramming**: `#[derive(Serialize)]` reads a
type's fields *at compile time* and emits per-type field-naming code, so
`serde` serializes any deriving type with zero reflection. `json.encode` is,
in effect, a single hand-built instance of exactly that mechanism. This
iteration **generalizes `json.encode`'s mechanism into a reusable derive
facility**: the compiler generates per-type code from the class-table metadata
it already emits, so generic-feeling capabilities exist *within* principle 13
rather than against it.
## Goals
- **A closed, compiler-known set of derivable capabilities** requestable on a
class — the first set: `Json` (retrofitting the existing `json.encode`),
`Csv`, structural `Eq`, `Hash`, and `Show` (a debug rendering). Each is
generated by the compiler from the class's field names and kinds; none is a
runtime reflective loop.
- **Generation preserves principle 13 exactly.** The emitted code is ordinary
bytecode over statically-known offsets and kinds — monomorphic per type, no
`Dynamic`, no runtime type tag, no dynamic dispatch. Disassembly must show
a plain per-type routine, not a reflection opcode.
- **`json.encode` becomes the `Json` derive**, reimplemented on the framework
so the framework is proven by rebuilding the thing that already works —
byte-identical output, or the change is wrong.
- **The query-result serialization gap closes**: a `multi Employee` whose
`Employee` derives `Csv` can be serialized whole, which is precisely the
`toCSV(from e in Employee where … select e)` case that has no expression
today (a generic serializer can neither be user-written under principle 13
nor attached as a method to a native `multi`).
## Acceptance Criteria
- What to achieve?
- **Given** a class annotated to derive `Csv` (surface per the spec),
- **when** the program is compiled and a value (or a `multi` of values) is
encoded,
- **then** the output is the expected CSV, the encoder is generated from
the class table, and **disassembly shows ordinary bytecode with no
reflection and no dynamic dispatch** — principle 13 provable, not
asserted.
- What to achieve?
- **Given** `json.encode` reimplemented as the `Json` derive,
- **when** the existing db, log-watcher, and json corpus run,
- **then** every output is byte-identical to today — the framework
generalizes the mechanism without changing its result.
- What to achieve?
- **Given** a derive requested on a class one of whose fields the derive
cannot handle (a `@gc` field for a value derive, a kind with no CSV
rendering),
- **when** it is compiled,
- **then** it is a **compile error naming the field and the reason** — no
silent partial output, no runtime failure. A derive's applicability is
decided entirely at compile time.
- What to achieve?
- **Given** two classes deriving `Eq` where one embeds the other,
- **when** structural equality is generated,
- **then** it recurses through the embedded type's own derived `Eq` — the
framework composes across types the way the field kinds nest.
## Out Of Scope
- **A full trait / typeclass system** — bounds like `fn f<T: Serialize>(x: T)`,
generic functions, and the inference they need. That is a large, separate
language iteration; this one ships a **closed, compiler-known derivable
set**, not open generics. The derive facility is the pragmatic 80% without
the type-system weight.
- **User-defined / procedural macros.** Rust lets users write `proc_macro`
derives; writeonce does not, and this iteration keeps it that way — only the
compiler-builtin derive set. A user-macro system is a much larger surface
and likely never wanted (KISS).
- **Monomorphized generics as a general feature.** Per-type generation here is
specific to the derive set, not a general generics engine.
- **Deriving across the attach channel** — a client generating an encoder over
the owner's types (iterations 9c/9d). Composes later; the class-table
metadata already crosses the channel's schema handshake, so the pieces are
in place, but it is not this iteration's problem.
- **Reopening principle 13 in any form.** If a derive appears to need runtime
reflection, the derive is wrong, not the principle — that is a defect report
against this iteration.
## Info
Prior art in the tree:
- **`json.encode`/`decode` (`runtime/src/json.c`) is already this mechanism**,
built once by hand: metadata-driven, compiler-lowered with the class id,
no reflection. This iteration lifts its shape into a reusable framework.
- **The class-table metadata** (`.wob` v2's `field_names`/`field_class`/
`field_elem`, v3's index metadata) is the substrate every derive reads. It
already exists and is already what `json.encode`'s lowering walks.
- **Principle 13 is both the constraint and the enabler**: because every
type is known at compile time, per-type generation needs no runtime
dispatch, so the generated code is as fast and as untagged as hand-written.
Forks the spec must settle:
**1. The request surface.** Options: an annotation in the existing ORM style
(`@derive(Csv, Json, Eq)` on the class, matching `@table`/`@unique`); a
`derive` keyword; or trait-style `impl`-blocks. Leaning: the `@derive(...)`
annotation — smallest surface, consistent with the annotation-driven design
the language already has, and it keeps derives a closed compiler-known set
rather than implying an open trait system.
**2. How a derived capability is invoked.** With no UFCS and no methods on
native containers, `value.to_csv()` cannot be a method on a `multi`. Options:
a compiler-recognized builtin per capability (`csv.encode(x)`, exactly like
`json.encode(x)` is lowered today), or generated free functions named by
convention (`Employee_to_csv`). Leaning: compiler-recognized builtins
(`json.encode`/`csv.encode`/…), so the invocation is uniform and the
collection case (`csv.encode(a_multi)`) is handled by the same lowering that
already special-cases a value's static kind.
**3. Whether `Eq`/`Hash` change what the VM already does.** Structural
equality and hashing over stored/embedded types touch the same metadata the
engine's indexes use — the spec should decide whether derived `Eq`/`Hash`
share code with the engine's key comparison (`database/src/table.c`'s
`idx_cols_equal`/`idx_hash`) or generate independent routines. Leaning: share
where the shapes match (one definition of "these two values are equal"), so a
derived `Eq` and an index's uniqueness check can never disagree.
**4. Applicability checking.** A derive must reject at compile time any field
it cannot handle (a `@gc` field in a by-value derive, a kind with no rendering
for the target format). The spec pins the rule per capability — and this is
the mechanism by which the facility stays inside principle 13: applicability
is a static question with a static answer, never a runtime probe.
## Proposed Solution
- **Brainstorm the spec** settling the four forks, then a plan whose first
task is retrofitting `json.encode` onto the framework — the proof that the
generalization changes nothing observable — before adding `Csv`/`Eq`/`Hash`/
`Show`.
- Expected shape: a `@derive(...)` annotation parsed like `@table`; a
compiler pass that, per derived capability per type, generates a routine
from the class-table metadata (the same metadata `json.encode` walks);
compiler-recognized encode builtins that lower to those routines; and
applicability diagnostics in a new WO-E range. The runtime gains no new
reflective machinery — only, at most, small shared helpers the generated
code calls.

View file

@ -0,0 +1,199 @@
# Iteration 14 — skillhost: a host-shaped workload, and the capability gaps it exposes
> Format: `product/story-iteration-template`. Part of
> [Story — one language, one runtime, one database, one binary](00-story.md).
>
> **Inserted 2026-08-16.** A driving-workload iteration, the way iteration 7's
> log-watcher drove the systems stdlib. The workload is a writeonce port of
> **`~/projects/skillhost`** (a C++ MCP host that links `libllama` in-process,
> discovers filesystem skills, and runs their scripts under confinement). The
> point is not the port for its own sake — it is that a *host-shaped* program
> (an MCP tool server that orchestrates subprocesses) exercises runtime
> capabilities no prior sample needed, and this iteration names each gap so it
> becomes schedulable.
>
> **No spec exists yet.** This iteration frames the workload and records the
> gaps (surveyed 2026-08-16 against the running compiler + the skillhost
> source); each gap below is a candidate iteration of its own.
## Goals
- **A writeonce program that is skillhost's shape**: one MCP tool
(`delegate_task`) exposed to a hosted caller, an on-disk skill catalog
discovered at startup, and a confined subprocess runner that executes a
skill's scripts and returns exit code + captured output. Target sample:
`docs/examples/skillhost`.
- **Drive the runtime's missing host capabilities into the open.** Each thing
the port cannot express today is a named gap with a decision attached — the
same discipline log-watcher used to grow `fs`/`proc`/`net`/`time`.
- **Ship the expressible part now, block honestly on the rest.** Much of
skillhost is already writable (below); the port lands in the pragmatic
variant that avoids the hard blockers, and the blockers become their own
iterations.
## What is already expressible (verified 2026-08-16)
- **The skill catalog** — `@table` with the query surface already exceeds
skillhost's in-memory SQLite `skills` table; iteration 9g's
`docs/examples/skill-catalog` is literally this table, running. (Or a plain
`map`/`multi` would do — the catalog is a lookup cache, not persistence.)
- **Discovery** — `fs.exists`/`fs.list` (one level) + `fs.read_all` walk
`.agents/*/SKILL.md` exactly as skillhost does one level deep.
- **Frontmatter** — skillhost's own parser is a hand-rolled YAML subset
(`--- … key: value …`); that is plain Text parsing in writeonce, no YAML
library needed.
- **Config** — `env.get` (the `SKILLHOST_*` overrides), `fs.read_all` +
`json.decode` (the config file), `fn main(args)` (flags). The "context
gate" is pure Int arithmetic (`tokens + max_tokens >= n_ctx`); skillhost
does **no** runtime VRAM query, so none is needed.
- **Single-threaded serving** — skillhost handles one request at a time (its
mutexes are defensive only); writeonce's blocking single-thread model maps
cleanly, as log-watcher's MCP mode already proved.
## The gaps (each a candidate iteration)
### Blocker A — in-process native library binding (FFI)
skillhost links `libllama`'s C API in-process (`llama_model_load_from_file`,
`llama_decode`, the sampler chain, GBNF grammar, tokenizer, chat template).
writeonce has **no FFI**, so none of it can be linked or called. The
architecture doc is explicit that in-process linkage is *what forces C++* —
so this gap is the whole reason skillhost is not already portable.
Two directions, and they are a real fork, not a detail:
- **FFI as a language capability** — a way to declare and call C functions
from writeonce. This is a large, doctrine-level addition (the runtime is
libc-only by principle; the one sanctioned exception so far is the vendored
Ed25519, 9d). FFI would reopen the dependency-sprawl question the whole
project is built to avoid. Likely its own spec, likely contested.
- **Out-of-process model, no FFI** — drive a llama.cpp binary via `proc`
(`llama-cli`) or `llama-server` over `net` + `json` (it accepts a GBNF
`grammar` param, so grammar-constrained tool calls survive). This needs
**nothing new** and is how a host in any language would do it — the doc
notes `llama-server` "a host in any language could drive." It loses
skillhost's per-turn `llama_memory_clear` + custom sampler control, which
become server request options. **This is the variant the port should use.**
Leaning: build the port on the out-of-process model; record FFI as a separate
iteration to be opened only if a workload genuinely needs in-process linkage,
and expect it to be weighed hard against the libc-only doctrine.
### Blocker B — stdio transport (process stdin/stdout)
skillhost speaks MCP as **newline-delimited JSON-RPC over stdin/stdout**
(`server.serve(std::cin, std::cout)`), which is what a stdio-launching MCP
client expects. writeonce has **no stdin builtin and no stdout-write
builtin** — only `net` sockets and `print`. So a faithful stdio transport is
impossible today.
- The pragmatic port re-hosts MCP on a **TCP socket** (`net.listen/accept/
read/write` + newline framing + `json`), exactly as log-watcher's MCP mode
does — a legitimate but *different* transport than skillhost's stdio.
- The real capability gap: **raw stdin/stdout byte I/O** (read fd 0, write fd
1, flush), so a writeonce program can be a stdio-transport MCP server — the
default an MCP client launches. Small stdlib addition (`io.stdin_read` /
`io.stdout_write`, or an `fd` surface). A candidate iteration.
### Blocker C — bounded, killable subprocesses
skillhost's `run_script` spawns a script, drains stdout+stderr, and — on a
120s deadline — `SIGKILL`s the whole **process group** so backgrounded
grandchildren die too (`poll`-with-deadline + `kill(-pid, SIGKILL)`).
writeonce's `proc.run` captures stdout/stderr/exit but has **no timeout, no
signal control, no process-group kill**, and blocking-only I/O with no
threads means it cannot wait on a child with a wall-clock deadline. A runaway
skill script cannot be bounded — unacceptable for a host that runs untrusted
scripts.
The capability gap: **`proc.run` with a timeout that kills the child (group)
on expiry** — a deadline-bounded subprocess. This composes with the
log-watcher Task-4 stop-flag work (interruptible blocking calls) and is a
clear candidate iteration.
### Partial cluster — filesystem metadata
Smaller gaps, all in `fs`:
- **Recursive directory walk** — skillhost lists a skill's executables/
references recursively; `fs.list` is one level, so this is built by hand
(fine) or `fs` grows a recursive walk.
- **Executable-bit detection** (`access(X_OK)`) — skillhost only offers
runnable scripts to the model; unclear whether `fs.stat` exposes mode bits.
Candidate `fs.stat` extension.
- **Symlink-resolving path confinement** (`weakly_canonical`/`realpath`) —
skillhost resolves symlinks before its prefix check so a link pointing out
of the skill dir is rejected. writeonce has no `realpath`; a string-prefix
check alone is **weaker** (a symlink escape). Candidate `fs` addition —
and a security-relevant one, since confinement is the host's trust boundary.
## Acceptance Criteria
- What to achieve?
- **Given** the expressible subset (catalog + discovery + frontmatter +
config + a TCP-hosted `delegate_task` tool + a subprocess runner),
- **when** `docs/examples/skillhost` is written and compiled,
- **then** it runs as an MCP server over a socket, discovers `.agents/*`
skills, answers `initialize`/`tools/list`/`tools/call`, and runs a
skill's script returning its exit code and captured output — the
skillhost shape, minus the named blockers.
- What to achieve?
- **Given** a skill script that runs forever,
- **when** the runner has no bounded-subprocess capability (Blocker C),
- **then** the port documents that it cannot yet bound it — the gap is
demonstrated, not hidden, and the acceptance records it as the reason
Blocker C's iteration exists.
- What to achieve?
- **Given** each gap iteration lands (bounded subprocess, stdio
transport, fs metadata, and — if ever — FFI),
- **when** the port adopts it,
- **then** the port moves one step closer to skillhost's exact behavior,
and the divergence list in its README shrinks by exactly that gap.
## Out Of Scope
- **In-process `libllama` / CUDA offload** — requires Blocker A (FFI); the
port uses an out-of-process model instead and says so.
- **Runtime VRAM/NVML introspection** — skillhost does none (the context
gate is arithmetic); nothing to build.
- **A resident MCP-over-stdio server** until Blocker B lands — the port uses
a socket transport meanwhile.
- **Reproducing skillhost's exact sampler chain / per-turn memory clear** —
those are in-process libllama specifics; the out-of-process variant
approximates them with server request options.
## Info
The three blockers are independent and each merits its own iteration; the
filesystem partials could ride together as one small `fs`-metadata iteration.
Priority order by leverage:
1. **Bounded subprocess (Blocker C)** — the smallest, most broadly useful, and
a hard requirement for any host that runs untrusted scripts. Do first.
2. **stdio transport (Blocker B)** — small, unblocks writeonce being a
*standard* MCP server (stdio is the default an MCP client launches), useful
far beyond skillhost.
3. **fs metadata (the partials)** — small, security-relevant (symlink
confinement).
4. **FFI (Blocker A)** — largest and most contentious; open only on genuine
demand, and expect the out-of-process variant to make it unnecessary for
this workload.
The port itself is buildable today in the out-of-process / TCP-transport
variant *once Blocker C exists* (a host that cannot bound a runaway script is
not honest to ship) — so the natural first step is Blocker C, then the port,
then B and the partials narrow the gap to skillhost's real behavior.
## Proposed Solution
- **Brainstorm each gap iteration on demand**, starting with the bounded
subprocess (Blocker C) since the port cannot responsibly run scripts
without it.
- **Write `docs/examples/skillhost`** in the out-of-process / socket-transport
variant, with a README whose "divergence from the C++ skillhost" list is
exactly the open gaps (A: out-of-process model, B: socket not stdio, C:
bounded once its iteration lands, plus the fs partials) — the list is the
iteration's own scoreboard.
- Reuse iteration 9g's `skill-catalog` as the catalog layer, log-watcher's
MCP mode as the transport skeleton, and the systems stdlib for discovery
and execution.

View file

@ -1,6 +1,6 @@
# DB Engine Binding Implementation Plan # DB Engine Binding Implementation Plan
> **Status: ⬜ pending** (story iteration 9) — class-shaped tables, typed WAL + recovery, `insert`/`select` execution. Story iteration 9b (`@table` relations + language-integrated query) follows it and needs a spec brainstormed first. Board: [00-status.md](../../00-status.md) > **Status: 🔄 in progress — Tasks 1–4 done; Task 5 engine half done; Task 6 incremental. The language read surface is the 9b plan's (recorded deviation at Task 5); the engine itself is COMPLETE for single-shard: storage, WAL+replay, insert/update/delete execution, indexes, unique.** (story iteration 9) — class-shaped tables, typed WAL + recovery, `insert`/`select` execution. Story iteration 9b (`@table` relations + language-integrated query) follows it and needs a spec brainstormed first. Board: [00-status.md](../../00-status.md)
> **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (`- [ ]`) syntax for tracking. > **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (`- [ ]`) syntax for tracking.
> >
@ -20,6 +20,7 @@
- **Id discipline:** ids interleave per shard (`t+1, t+1+N, …`); a row's owner shard is `(id-1) % N`; point ops on a foreign row hop once via the plan-4 mailbox — creates are always local. - **Id discipline:** ids interleave per shard (`t+1, t+1+N, …`); a row's owner shard is `(id-1) % N`; point ops on a foreign row hop once via the plan-4 mailbox — creates are always local.
- **Index doctrine:** secondary indexes update only inside the engine's insert/remove path; direct storage mutation is a defect by definition. - **Index doctrine:** secondary indexes update only inside the engine's insert/remove path; direct storage mutation is a defect by definition.
- **RAM is authoritative:** reads never touch a file descriptor (phase-B doctrine); disk exists for durability and boot. - **RAM is authoritative:** reads never touch a file descriptor (phase-B doctrine); disk exists for durability and boot.
- **The GC bulkhead (analysis 2026-08-15, spec'd in the 9b design section 6):** values cross between VM heap and row storage only by copy, and a GC-managed value in a stored field is a compile error — so the collector never traces engine memory and the engine never touches reference counts. When iteration 7b makes GC-ness inferred, inference must classify classes **before** table-field validation so this error keeps firing, with the message naming the inference reason.
- **Format changes go through the format doc:** DB operations extend the builtin table (ids appended to `docs/plan/oop-vm/00-wob-format.md`); no opcode-space or version change. - **Format changes go through the format doc:** DB operations extend the builtin table (ids appended to `docs/plan/oop-vm/00-wob-format.md`); no opcode-space or version change.
--- ---
@ -27,7 +28,8 @@
## File Structure ## File Structure
``` ```
runtime/src/ database/src/ the engine is its own top-level directory (user decision
2026-08-15), statically linked into wovm — one binary, unchanged
table.c table.h class-shaped row storage: slabs, slots, id alloc, indexes (Tasks 1, 4) table.c table.h class-shaped row storage: slabs, slots, id alloc, indexes (Tasks 1, 4)
wal.c wal.h typed-row WAL records, group commit, replay (Task 2) wal.c wal.h typed-row WAL records, group commit, replay (Task 2)
db.c db.h statement executors: insert/select/update-point (Tasks 3, 5) db.c db.h statement executors: insert/select/update-point (Tasks 3, 5)
@ -36,31 +38,72 @@ tests/corpus/db/ DB fixtures incl. crash/replay (Task 6)
docs/plan/oop-vm/04-db-binding.md row format, WAL record layout, query subset (Task 1) docs/plan/oop-vm/04-db-binding.md row format, WAL record layout, query subset (Task 1)
``` ```
The engine's headers are included by `runtime/src` (the VM calls the row API);
`runtime/Makefile` compiles `database/src/*.c` into every `wovm` target,
sanitizers included. `database/` gets its own CODE-LOGIC.md as code lands.
--- ---
### Task 1: Class-shaped row storage ### Task 1: Class-shaped row storage
**Concept & reason:** generalize phase B. Per shard, per class: a slab of fixed-size row slots sized from the class's field count (16-byte row header — id, class, flags — plus the same 8-byte slots the VM object layout uses, so a row and an object share their field encoding; text and container fields store engine-owned copies, not VM pointers). An allocation bitmap per slab; slab growth by arena extension. Id allocation interleaved per shard for coordination-free global uniqueness (shipped phase-A behavior). Row create/read/remove go through one API that Task 4's indexes hook — the doctrine choke point. The binding doc pins the row format, the field-encoding rules (what happens to each of the six kinds when a value crosses from VM heap to row storage — scalars copy, texts copy, owned objects flatten by value, `@gc` references are a compile error in stored fields already, `ref` is an id, containers copy element-wise), and the query subset promised by Task 5. **Concept & reason:** generalize phase B. Per shard, per class: a slab of fixed-size row slots sized from the class's field count (16-byte row header — id, class, flags — plus the same 8-byte slots the VM object layout uses, so a row and an object share their field encoding; text and container fields store engine-owned copies, not VM pointers). An allocation bitmap per slab; slab growth by arena extension. Id allocation interleaved per shard for coordination-free global uniqueness (shipped phase-A behavior). Row create/read/remove go through one API that Task 4's indexes hook — the doctrine choke point. The binding doc pins the row format, the field-encoding rules (what happens to each of the six kinds when a value crosses from VM heap to row storage — scalars copy, texts copy, owned objects flatten by value, `@gc` references are a compile error in stored fields already, `ref` is an id, containers copy element-wise), and the query subset promised by Task 5.
- [ ] Failing tests: create/read/remove round-trips across kinds; id interleave across shards; slab growth; removal reuses slots. - [x] Tests first: round-trips across every kind (Text/owned-nested/multi/map
- [ ] Implement; ASan green. Write the binding doc. copies proven by mutating the originals), nil encodings incl.
- [ ] Record commit draft: `feat(runtime): class-shaped row storage — per-shard per-class slabs from .wob class table, VM-compatible field encoding, interleaved id allocation, single choke-point row API; docs/plan/oop-vm/04-db-binding.md.` `WO_NIL_SCALAR`, id interleave at N=3, slab growth past three slabs
with stable addresses, removal reuses the slot while never reusing the
id, misuse (unknown class, double remove). `test_table` 827/0 under
ASan+UBSan.
- [x] Implemented in `database/src/table.{c,h}` (the 2026-08-15 directory
decision), linked into every wovm and test binary via the Makefile's
`DBSRC`. Binding doc written (`docs/plan/oop-vm/04-db-binding.md`:
row format, encoding table, id discipline, choke points). All prior
gates stay green with the engine linked (oop-e2e 71/0, log-watcher
7/0) — it is dead code until Task 3 wires the first builtin.
- [x] Committed locally (2026-08-15). N=1 today: iteration 8 has not landed,
so everything is N-parametric and tested at N=3 through the API.
### Task 2: Typed WAL + boot replay ### Task 2: Typed WAL + boot replay
**Concept & reason:** port phase D/E to typed rows. Records frame `length | crc | payload | commit-mark` (replay-whole-or-not-at-all); payload = record kind (insert/remove/update), class id, row id, encoded fields. Per-shard WAL files, fallocate-preallocated, appended on commit AFTER the RAM apply, one fdatasync covering the tick's batch, ack after the sync completes — the shipped commit order, verbatim. Boot: per-shard parallel replay before listeners open; torn tails detected and dropped whole (phase E behavior). The offline `wal-check` verification mode ports too — it is the crash test's oracle. **Concept & reason:** port phase D/E to typed rows. Records frame `length | crc | payload | commit-mark` (replay-whole-or-not-at-all); payload = record kind (insert/remove/update), class id, row id, encoded fields. Per-shard WAL files, fallocate-preallocated, appended on commit AFTER the RAM apply, one fdatasync covering the tick's batch, ack after the sync completes — the shipped commit order, verbatim. Boot: per-shard parallel replay before listeners open; torn tails detected and dropped whole (phase E behavior). The offline `wal-check` verification mode ports too — it is the crash test's oracle.
- [ ] Failing tests: commit-then-kill crash battery (the phase-D test shape: concurrent writes, SIGKILL mid-stream, offline verification proves every acked write present and CRC-valid, zero acked-but-missing); torn-tail drop; parallel replay rebuilds identical RAM state (deep-compare against pre-crash snapshot dump). - [x] Tests first: replay round-trip with deep compare (Texts included) and
- [ ] Implement; green. next_id advance; torn-tail drop with the oracle agreeing on the intact
- [ ] Record commit draft: `feat(runtime): typed-row WAL + replay — framed CRC records over class rows, group-commit ack-after-fsync, parallel boot replay with torn-tail drop, wal-check oracle; crash battery green.` prefix byte-for-byte; reopen positions AT the tear so the next commit
overwrites it; crash battery — five rounds of fork + insert/commit/
ack-over-pipe + SIGKILL mid-stream, zero acked-but-missing, zero
acked-but-wrong, contents exact. `test_wal` 90/0 under ASan+UBSan.
- [x] Implemented in `database/src/wal.{c,h}`: framed CRC records
(`len|crc|payload|mark`), typed-row payloads decoded straight into
engine values (no VM at boot), group commit as one pwrite + one
fdatasync, replay through the choke-point row API (indexes rebuild for
free when Task 4 lands), `wo_wal_check` offline oracle. "Parallel"
replay is per-shard by construction and runs at N=1 until iteration 8.
- [x] Committed locally (2026-08-15). Binding doc's WAL section written.
### Task 3: `insert` executes ### Task 3: `insert` executes
**Concept & reason:** the compiler's DbStub node for `insert` becomes a typed AST: target class, field initializer list (defaults applied for omitted fields with defaults — the explicit now() form computes at execution), returning the new id. Typechecking validates fields against the class exactly like constructor literals. Lowering emits DB builtins (ids appended to the format doc's builtin table): the executor allocates the id, encodes fields from registers, applies to RAM through the Task-1 API, stages the WAL record; the VM sees the id as the result. Inserts targeting the local shard complete inline; there is no remote insert — creates are always local by the id discipline. The pricing corpus's `set_price` fixture flips from expecting the DB trap to expecting success — the milestone's most satisfying diff. **Concept & reason:** the compiler's DbStub node for `insert` becomes a typed AST: target class, field initializer list (defaults applied for omitted fields with defaults — the explicit now() form computes at execution), returning the new id. Typechecking validates fields against the class exactly like constructor literals. Lowering emits DB builtins (ids appended to the format doc's builtin table): the executor allocates the id, encodes fields from registers, applies to RAM through the Task-1 API, stages the WAL record; the VM sees the id as the result. Inserts targeting the local shard complete inline; there is no remote insert — creates are always local by the id discipline. The pricing corpus's `set_price` fixture flips from expecting the DB trap to expecting success — the milestone's most satisfying diff.
- [ ] Failing tests: compiler goldens (typed insert AST, emitted builtins); runtime fixtures (insert then read back through select-by-id once Task 5 lands — interim: through a test hook on the row API); default-value application; the flipped pricing fixture. - [x] Compiler half: `insert Class { … }` is a typed `Ast.Insert` in BOTH
- [ ] Implement both halves; green. positions (the old "bare insert is an Ident" contract retired, its
- [ ] Record commit draft: `feat: insert executes — DbStub becomes typed insert AST with constructor-grade field checking, DB builtins apply RAM-then-WAL through the row API; pricing set_price fixture flips from trap to green.` unit test rewritten to the new one; `select` stays a DbStub for
Task 5). Typechecked like a ctor (same omittable rule), owner pass
treats field values as borrows (the engine copies — no transfer),
emitter lowers to builtin 61 with declaration-order slots, defaults
and `?`-nils filled, fresh values reaped after. Goldens re-blessed
(`ast/db-stub` shows the INSERT node), 566/0.
- [x] Runtime half: `database/src/db.c` executes through the choke-point
row API; `WO_DATA=<dir>` turns on replay-at-boot + commit-before-ack
per statement (the builtin's return IS the ack until iteration 8's
ticks); failed commit un-applies the row and traps WO_T_IO. The
loader validates the class-id slot (variable window documented).
- [x] **The pricing fixture flipped**: `trap/pricing-set-price-db-stub`
(expected trap 5) is now `run/pricing-set-price-insert` printing the
ids the engine allocated — the milestone's promised diff. Durability
smoke: two consecutive `WO_DATA` runs print 1,2 then 3,4 (replay +
next_id advance). oop-e2e 71/0, all runtime suites green,
log-watcher 7/0. Committed locally (2026-08-15).
### Task 4: Secondary indexes ### Task 4: Secondary indexes
@ -70,7 +113,19 @@ docs/plan/oop-vm/04-db-binding.md row format, WAL record layout, query subset
- [ ] Implement; green. - [ ] Implement; green.
- [ ] Record commit draft: `feat(runtime): secondary indexes — per-shard hash indexes maintained only inside the row API, @unique constraint trap, replay rebuild; doctrine enforced by construction.` - [ ] Record commit draft: `feat(runtime): secondary indexes — per-shard hash indexes maintained only inside the row API, @unique constraint trap, replay rebuild; doctrine enforced by construction.`
### Task 5: `select` subset + cross-shard point reads ### Task 5 🔄: `select` subset + cross-shard point reads — **superseded in part (2026-08-15)**
> **Deviation, recorded:** the 9b design (spec'd after this plan was
> written) makes compiler-lowered comprehension queries THE select
> surface — a brace-select subset built here would be a second grammar
> retired months later. So Task 5's ENGINE half landed now (update-point
> through the choke point with unique re-check, delete, the WAL UPDATE
> record with replace-on-replay, builtins 62/63 — `test_table` 839/0,
> `test_wal` 102/0), and the LANGUAGE half (reads, queries, row views,
> the delete statement) is the 9b plan's Tasks 1–5, where it belongs.
> Cross-shard point reads wait for iteration 8's mailboxes, as written.
**Concept & reason:** the milestone query subset, semantics per the C++ `wo-db` reference where they overlap: select-by-id; select with a WHERE conjunction over indexed fields (index-backed) or a full shard scan (explicitly allowed, explicitly slower); dotted-path field access in the projection; results materialize as VM objects (rows decode back through the field-encoding rules — the Task-1 doc's table read in reverse). Local rows resolve inline; a by-id read of a foreign row hops once via the plan-4 mailbox (point ops hop once; the requesting job pumps its inbox while waiting — never blocks). List queries stay shard-local in this plan; scatter-gather fan-out is future work, stated in the doc. `update` limited to point-by-id field sets (the method-transaction pattern the pricing demo uses); no joins, no aggregations beyond the existing builtins, no RETURNING chaining — all named as out-of-scope in the binding doc. **Concept & reason:** the milestone query subset, semantics per the C++ `wo-db` reference where they overlap: select-by-id; select with a WHERE conjunction over indexed fields (index-backed) or a full shard scan (explicitly allowed, explicitly slower); dotted-path field access in the projection; results materialize as VM objects (rows decode back through the field-encoding rules — the Task-1 doc's table read in reverse). Local rows resolve inline; a by-id read of a foreign row hops once via the plan-4 mailbox (point ops hop once; the requesting job pumps its inbox while waiting — never blocks). List queries stay shard-local in this plan; scatter-gather fan-out is future work, stated in the doc. `update` limited to point-by-id field sets (the method-transaction pattern the pricing demo uses); no joins, no aggregations beyond the existing builtins, no RETURNING chaining — all named as out-of-scope in the binding doc.
@ -78,7 +133,18 @@ docs/plan/oop-vm/04-db-binding.md row format, WAL record layout, query subset
- [ ] Implement; green. - [ ] Implement; green.
- [ ] Record commit draft: `feat: select subset — by-id (cross-shard hop-once), indexed/scan WHERE conjunctions, projection decode to VM objects, point update; scatter-gather and joins explicitly deferred.` - [ ] Record commit draft: `feat: select subset — by-id (cross-shard hop-once), indexed/scan WHERE conjunctions, projection decode to VM objects, point update; scatter-gather and joins explicitly deferred.`
### Task 6: DB corpus + acceptance ### Task 6 🔄: DB corpus + acceptance — **landing incrementally**
> **State 2026-08-15:** db fixtures live in the EXISTING corpus kinds
> (`run/db-unique-catch`, `trap/db-unique-violation`, the flipped
> `run/pricing-set-price-insert`) rather than a new `db/` kind — the
> runner stays one harness. The crash/replay battery is proven at unit
> level (`test_wal`'s five-round SIGKILL battery) and lands as a scripted
> scenario with the employee acceptance (whose script owns the kill -9
> step). The `wo-db` smoke overlap and select-round-trip fixtures wait on
> 9b's query surface — reads have no language spelling until then.
**Concept & reason:** the corpus grows a `db/` kind wired into the standard runner: insert/select round-trip programs with exact stdout; the crash/replay battery as a scripted scenario (run, kill, reboot, assert identical query results); a unique-violation trap fixture; the C++ `wo-db` smoke overlap — where `prototypes/wo-db`'s `.wo` smoke files exercise semantics this subset implements, run both and compare (manifest-scoped like the plan-3 parity harness). `just oop-accept` gains the DB corpus and the crash scenario. The kanban and CLAUDE.md sync: Stage-3-adjacent language ("insert/select execute in the C runtime") replaces the DB_STUB story. **Concept & reason:** the corpus grows a `db/` kind wired into the standard runner: insert/select round-trip programs with exact stdout; the crash/replay battery as a scripted scenario (run, kill, reboot, assert identical query results); a unique-violation trap fixture; the C++ `wo-db` smoke overlap — where `prototypes/wo-db`'s `.wo` smoke files exercise semantics this subset implements, run both and compare (manifest-scoped like the plan-3 parity harness). `just oop-accept` gains the DB corpus and the crash scenario. The kanban and CLAUDE.md sync: Stage-3-adjacent language ("insert/select execute in the C runtime") replaces the DB_STUB story.

View file

@ -1,94 +0,0 @@
# UI — .htmlx SSR + LIVE Subscriptions Implementation Plan
> **Status: ⏸ parked** — the UI track (`.htmlx` SSR + LIVE subscriptions) is off the critical path by the 2026-08-08 scope directive and is deliberately **not sequenced** on the board; the language track runs to the log-watcher proof first. Board: [00-status.md](../../00-status.md)
> **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (`- [ ]`) syntax for tracking.
>
> **Style rule (user convention):** concept, reason, and required behavior in words only; the executor writes the code.
**Goal:** Sub-project 5 of the OOP spec — the single binary serves a user interface: `##ui` screens compile to `.htmlx` templates rendered server-side, LIVE subscriptions push delta frames over WebSocket on commit, and the pricing demo's driving workload runs end to end — a price update on the server patches subscribed browsers' cells in place.
**Architecture:** Plan 7 of 7. Depends on plans 1–6. Direction is already locked by the UI track (`docs/plan/exploration/ui/00-overview.md`): **`.htmlx` is the template format; `##ui` is the DSL that emits it** — the v1 engine at `.dev/reference/crates/wo-htmlx/` (bindings, each-blocks, partials, data-bind attributes) is the semantic reference the C renderer ports. The subscription machinery is the Stage-3 story (`docs/runtime/database/04-client-api.md`, plan 13c): a per-shard registry hooked into the commit path emits delta frames to WebSocket subscribers — replacing the honest 501 that plan 6 preserved. Static assets (the client runtime JS) serve per the sendfile doctrine (`docs/plan/08-sendfile-static-assets.md`). The UI track's larger workspace story (apps/, per-app binaries, shared DB daemon — ui docs 04–06) is explicitly OUT of this plan: one binary, its own screens, first.
**Tech Stack:** C11 + libc (WebSocket framing hand-rolled; SHA-1 + base64 for the upgrade handshake hand-rolled — the only crypto in the runtime, ~150 lines, documented). OCaml (##ui parsing, .htmlx emission). Vanilla JS client runtime (~20 KB target per the UI track), no build tooling.
## Global Constraints
- All prior plans' constraints carry over (libc only, no commits — drafts to `.dev/commit.md`, gates, docs under `docs/`).
- **`.htmlx` semantics follow the v1 engine** where features overlap (`{{path}}` bindings, `{{#each}}`, `{{> partial}}`, `data-bind`); the live-subtree extension (`<wo:live source="...">`) is this plan's addition, specified in the format doc it writes.
- **Deltas ride commits:** subscription taps sit AFTER the WAL accepts, never in the read or ack path — a slow subscriber can never delay a commit (the Postgres-mirror tap discipline, applied to WebSockets: overflow drops the subscriber loudly, never blocks the shard).
- **Subscriptions are shard-local:** a socket subscribes on the shard that owns its connection; queries against rows on that shard push directly. Cross-shard live queries are out of scope, stated in the doc.
- **No JS frameworks, no bundlers** — the client runtime is one hand-written file served as a static asset.
---
## File Structure
```
runtime/src/
ws.c ws.h WebSocket upgrade (SHA-1/base64), frame codec, ping/pong (Task 1)
sub.c sub.h per-shard subscription registry + commit tap + delta encode (Task 2)
htmlx.c htmlx.h template parse + SSR render (Task 4)
assets.c assets.h static asset serving, sendfile path (Task 5)
compiler/src/ ##ui parsing + .htmlx emission (Task 3)
client/wo-live.js the DOM-patching client runtime (Task 5)
tests/corpus/ui/ render goldens + live end-to-end scripts (Task 6)
docs/plan/oop-vm/06-ui-live.md .htmlx subset, wo:live semantics, delta frame format (Task 1)
```
---
### Task 1: WebSocket transport
**Concept & reason:** the push channel, hand-rolled to the zero-dep bar: HTTP upgrade handshake (the fixed-GUID SHA-1/base64 accept key — the two primitives implemented locally with test vectors from their RFCs), frame codec (text frames, masking rules, fragmentation tolerated on receive, close handshake, ping/pong), mounted on plan 6's connection machine so a socket upgrades in place and joins the shard's loop. The doc this task writes pins everything downstream: the delta frame JSON shape (kind: insert/update/delete, class, id, changed fields), the subscribe message a client sends, and the `wo:live` template semantics Task 3–5 implement.
- [ ] Failing tests: RFC test vectors for the accept key; frame codec round-trips incl. masked payloads and close; a socket-level upgrade-then-echo harness on the real loop.
- [ ] Implement; green.
- [ ] Record commit draft: `feat(runtime): hand-rolled WebSocket — upgrade handshake (local SHA-1/base64 with RFC vectors), frame codec, ping/pong/close, mounted on the shard loop; docs/plan/oop-vm/06-ui-live.md pins delta/subscribe/wo:live formats.`
### Task 2: Subscription registry + commit taps
**Concept & reason:** Stage 3's engine, per shard. A registry maps subscription keys (class + optional indexed-field filter, the plan-5 WHERE subset) to subscriber lists (socket + subscription id). The commit path gains a tap after the WAL accept: each committed mutation consults the registry, encodes one delta frame, and enqueues it on matching subscribers' write queues with try-semantics — a full queue drops that subscriber with a loud close frame and a log line (the mirror discipline: RAM-side progress never waits on a consumer). The `/api/<type>/live` route flips from 501 to the upgrade + subscribe flow; unsubscribe and disconnect clean the registry.
- [ ] Failing tests: registry match/miss across filters; commit-to-frame flow on a rigged shard (insert/update/delete each produce the right frame); overflow drops the subscriber and only the subscriber; disconnect cleanup; the 501 fixture from plan 6 flips to expecting an upgrade.
- [ ] Implement; green under ASan/TSan (sockets and registry are shard-local — the tests prove no cross-thread traffic exists).
- [ ] Record commit draft: `feat(runtime): LIVE subscriptions — per-shard registry with filter matching, post-WAL commit taps with try-enqueue drop-loudly discipline, /live flips from 501 to upgrade+subscribe; delta frames per the pinned format.`
### Task 3: `##ui` compiles to `.htmlx`
**Concept & reason:** the compiler's side of the locked UI decision. The parser accepts the `##ui` block form the samples use (screen name, source class, projection/columns, filters) and the typechecker validates it against the class (fields exist, filter fields indexed where the live path requires it). Emission writes an `.htmlx` file per screen into the build output beside the `.wob`: bindings for projected fields, an each-block over the source rows, and the screen's live subtree wrapped in `<wo:live source="...">` carrying the subscription key the runtime will register. Hand-written `.htmlx` files in the project pass through untouched (first-class authoring alternative, per the locked decision) — the compiler only validates their `wo:live` sources against the schema.
- [ ] Failing tests: golden `.htmlx` output for a pricing screen fixture; validation diagnostics (unknown field, unfilterable live source); pass-through of a hand-written template with source validation.
- [ ] Implement; green.
- [ ] Record commit draft: `feat(compiler): ##ui emits .htmlx — screen DSL parse/typecheck, binding+each+wo:live template generation with subscription keys, hand-written .htmlx pass-through with source validation.`
### Task 4: `.htmlx` SSR renderer
**Concept & reason:** the C port of the v1 engine's render semantics, scoped to the compiled subset: parse the template once at boot into a node tree (static chunks, bindings, each-blocks, partials, live-subtree markers); render a screen by walking the tree against query results from the plan-5 select path, HTML-escaping bound values, expanding each-blocks per row, and stamping each live subtree with the ids the client runtime needs to patch later (stable per-row element ids derived from class + row id — the contract the delta patcher relies on). Rendered pages route like any handler; the screen's route comes from the service/UI declarations.
- [ ] Failing tests: render goldens (template + fixture rows → exact HTML) covering escaping, each over rows, partials, and live-subtree id stamping; a malformed-template diagnostic at boot, not at request time.
- [ ] Implement; green.
- [ ] Record commit draft: `feat(runtime): .htmlx SSR renderer — boot-time parse to node tree, escaped binding render over select results, stable per-row ids in live subtrees; render goldens.`
### Task 5: Client runtime + static assets
**Concept & reason:** the last mile. `wo-live.js` (hand-written, one file, ~20 KB budget): on load, find `wo:live` subtrees, open the WebSocket, send subscribe messages from the stamped keys, and patch on frames — update replaces bound cell contents by stable id, insert appends a row rendered from a client-side row template the SSR emitted, delete removes the row's element; reconnect with backoff; a visible stale indicator when the socket is down (honesty over silence). Static serving: the asset module serves the JS (and any project assets) with correct content types and the sendfile-doctrine path for regular files.
- [ ] Failing tests: asset serving (content type, byte-exact body, 404 miss); client runtime exercised by the Task-6 end-to-end (no separate JS test harness — stated tradeoff: the e2e is the test).
- [ ] Implement; green.
- [ ] Record commit draft: `feat: wo-live.js client runtime (subscribe from stamped keys, patch update/insert/delete by stable ids, reconnect+stale indicator) + static asset serving on the sendfile path.`
### Task 6: Pricing live demo end to end + acceptance
**Concept & reason:** the spec's driving workload, closed: the pricing project compiles to one binary; a scripted browser-less client (a test WebSocket client speaking the pinned protocol) loads the SSR page, subscribes, then a method RPC (`set_price` over plan 6's route) commits an insert — the script asserts the delta frame arrives with the new amount, and that a second subscriber sees it too (fan-out). A manual demo recipe (`just pricing-live-demo`) serves it for human eyes. Render goldens + the scripted live scenario + the asset tests join `just oop-accept`. Docs closeout: kanban (13c-equivalent milestone on the C stack), CLAUDE.md stage table amendment, the UI track doc gains a status note that the format/live layer shipped and the workspace/per-app-binary layers (ui docs 04–06) remain open.
- [ ] Wire the scenario + recipe + gate; green; docs synced.
- [ ] Record commit draft: `test: pricing live demo e2e — SSR load, subscribe, set_price RPC commit, delta fan-out asserted by scripted WS clients; just pricing-live-demo; oop-accept gains the UI gate; UI-track status synced.`
---
## Plan self-review notes
- **Spec coverage (sub-project 5 + Stage 3):** `##ui` SSR, live subscriptions with delta frames on commit, the pricing live workload, single binary serving UI + API + DB — the spec's "single binary that is the database, the web API, and the UI" sentence is fully mechanized after this plan. Non-scope, stated: cross-shard live queries, the workspace/per-app-binaries/shared-daemon layers (ui docs 04–06), auth/`me` sessions, TLS.
- **Order rationale:** transport before registry before templates before renderer before client — each is the next one's substrate; the e2e needs all five.
- **Consistency check:** delta frame shape, subscribe message, stable-id contract, and `wo:live` semantics are pinned once (Task 1 doc) and consumed by Tasks 2–5; subscription keys originate in the compiler (Task 3) and terminate in the registry (Task 2) — same key format, one doc section.

View file

@ -0,0 +1,381 @@
# `@table`, Relations, and Language-Integrated Query — Design
> **Status: proposed** (story iteration 9b). Settles the three forks recorded in
> [`09b-table-relations-query.md`](../../stories/language-runtime-database/09b-table-relations-query.md).
> Plan: [`docs/plan/compiler/2026-08-15-employee-relations-query.md`](../../plan/compiler/2026-08-15-employee-relations-query.md).
> Depends on iteration 9's engine plan
> ([`2026-08-01-db-engine-binding.md`](../plans/2026-08-01-db-engine-binding.md))
> for storage, WAL, insert and the select subset.
**One sentence:** queries are written in the language, checked by the compiler,
and lowered to bytecode loops over a handful of engine cursor builtins — no SQL
text exists anywhere in a compiled program — and the acceptance workload is a
new `docs/examples/employee` sample whose report mode is one `GROUP BY` after
another.
---
## 1. The three forks, settled
### Fork 1 — the SQL/Cypher layer is superseded as the program surface
`docs/runtime/database/02-wo-language.md` specifies a query layer of literal
SQL and Cypher with five fixed-glue rules. That layer is **not built on the C
stack**. The language-integrated surface below is the only way a `.wo` program
queries its tables.
Reasons, in order of weight:
- The iteration's own acceptance forbids the alternative: a query must lower to
engine operations "not a string handed to a parser at runtime — provable by
disassembly." A resident SQL parser and a compile-checked surface would be
two grammars for one meaning, and the second grammar would re-introduce at
runtime every error class the first one eliminated at compile time.
- The language's identity is compile-time checking (the WO-E diagnostic
catalogue). "A typo in a field name is a compile error" cannot be delivered
by a text layer.
- A binary that ships a SQL parser it uses only for its own programs pays
image size and attack surface for nothing.
What survives: `02-wo-language.md` stays as design history and as the
specification the `prototypes/wo-db` C++ prototype implements; that prototype
remains the **engine-semantics** reference (what an index probe returns, what
unique violation means), not a syntax reference. SQL text remains a candidate
*export/interop* format for much later (external tools speaking to a writeonce
service), explicitly not part of this iteration.
### Fork 2 — comprehension syntax, desugared at compile time
writeonce has no function values, so LINQ's method-chains-taking-lambdas are
unwritable. The surface is a **query expression** — a comprehension the
compiler desugars — where every predicate and projection is ordinary
expression syntax with the range variable in scope:
```wo
let seniors = from e in Employee
where e.salary > 90000
order by e.salary desc
select e;
let by_dept = from e in Employee
group e by e.dept into g
select { dept: g.key.name, headcount: count(g), avg_salary: avg(g.salary) };
```
(Illustrative; the grammar is normative in prose, section 3.)
Why comprehension and not chaining: a chained `.where(e.salary > 100)` has no
binding site for `e` — the comprehension's `from e in` clause is what
introduces the variable, which is exactly the property that makes every later
clause an ordinary typed expression the existing typechecker can check. C# had
to add query expressions *on top of* lambdas for the same readability reason;
we get to skip the lambda layer entirely. The predicate is compile-time
syntax, never a runtime closure — which is also what lets the whole query
lower to plain bytecode.
### Fork 3 — what each reference contributes
**System.Linq** (surveyed 2026-08-15) contributes the operator vocabulary and
its semantic edge cases, not machinery:
- The lowering target for `group … select aggregate` is the shape of
`AggregateBy`/`CountBy` and the `GroupBy(key, resultSelector)` overload:
**group-and-reduce as one node, no intermediate group object materialized**
(`Grouping.cs:63`, `AggregateBy.cs`). The `IGrouping`-returning overloads
exist to hand groups around as values; we have no delegates to hand them to,
so groups surface only as the `into g` binding inside the query itself.
- Empty-source rules per aggregate (section 4's table) adapt LINQ's — where
C# throws `InvalidOperationException`, writeonce answers `nil` through a
`?T` result, because an empty group is data, not a fault.
- Ordering is **stable**, guaranteed — LINQ enforces it with an original-index
tiebreak (`OrderedEnumerable.cs:432`); we adopt the same guarantee and the
same escape hatch (an unstable sort is legal when the key is a scalar and
rows are not identity-bearing, `OrderBy.cs:144-163` precedent).
- Join semantics: build a hash on the inner side, probe with the outer, and
**nil never joins** — LINQ achieves SQL's NULL-never-matches by refusing to
insert null keys into the build side (`Lookup.cs:119-122`); we do the same.
In grouping, by contrast, nil **is** a legitimate key (`Lookup.cs:207`).
**PostgreSQL** (surveyed 2026-08-15) contributes execution vocabulary for the
engine side:
- Aggregate execution is the transition/finalize split from `nodeAgg.c`: a
per-group transition value advanced once per row, a finalize step converting
state to result (`avg` carries sum+count without the executor knowing). The
`noTransValue` vs `transValueIsNull` distinction — "no row seen yet" is not
"the value is nil" — is adopted verbatim; it is what makes nil-skipping
aggregates correct without special-casing the first row.
- Grouping strategy: hash aggregation (`AGG_HASHED`) is the only strategy this
iteration builds; sorted grouping is an optimization for later. Project the
input to only the columns the aggregate needs before hashing
(`find_hash_columns` precedent).
- Referential integrity semantics from `ri_triggers.c`, mechanism discarded:
an FK check is a **direct probe of the referenced table's primary index**
(never query text); a nil `ref` passes the check (MATCH SIMPLE rule); an
update that does not change the key skips the check
(`RI_FKey_fk_upd_check_required` precedent). Enforcement action is
**restrict only** — deleting a Department that still has Employees traps;
cascade/set-nil are out of scope.
- Everything MVCC, buffer-manager, lock-manager and cost-planner shaped is
explicitly non-transferable: single-writer-per-shard RAM-authoritative
storage designed those problems away (iteration 9's doctrine).
---
## 2. Relations become typed navigation
The declaration vocabulary already exists in the samples and partially in the
compiler; this iteration makes it mean something:
- `dept: ref Department` — a foreign key. Stored as the target's row id (a
scalar column, already the compiler's classification). **Navigating** it in
an expression (`e.dept.name`) typechecks with `Department`'s field set in
scope and lowers to a primary-index point read.
- `staff: backlink Employee.dept` — the declared inverse. Not a stored column;
reading it (`d.staff`) is a secondary-index scan of `Employee`'s `dept`
column and yields `multi Employee`. Declaring a `backlink` whose target
field is not a `ref` to this class is a compile error.
- `?ref Department` — an optional relation; nil stores as id 0, never probes,
never joins.
Integrity, enforced at the engine's row choke points (iteration 9's index
doctrine — nothing touches storage except the row API):
- Insert/update of a non-nil `ref`: probe the referenced primary index; a miss
traps with a foreign-key violation code (sibling of iteration 9's
unique-violation trap).
- Delete of a row that a non-nil `ref` still points at: trap (restrict). The
check is a probe of the secondary index that the `backlink` already
requires, so restrict costs one lookup and no new structure.
---
## 3. The query surface (normative, in prose)
A **query expression** is an expression. Its clauses, in the only order they
may appear:
1. `from <var> in <source>` — required, first. Source is a table (class name),
a `backlink` navigation, or a `multi` value. Introduces the range variable.
2. `where <bool-expr>` — optional, repeatable. Ordinary boolean expression
over range variables; each `where` is a filter.
3. `join <var2> in <source2> on <expr> == <expr2>` — optional. Equi-join only;
one side references only the outer variable, the other only the joined one.
Most joins in practice are spelled as `ref` navigation instead; explicit
`join` exists for joining on non-relation columns.
4. `group <expr> by <key-expr> into <g>` — optional. Ends the scope of the
range variables and opens the scope of `g`. `g.key` is the key's value.
For any field `f` of the grouped element, `g.f` names the **column of
members** — legal only inside an aggregate call. Nil is a legitimate key.
5. `order by <expr> [desc]` — optional, comma-repeatable keys. Stable.
6. `take <int-expr>` / `skip <int-expr>` — optional.
7. `select <expr>` — required, last. The result element: a whole row, a single
field, or a projection literal `{ name: expr, … }` whose type the compiler
synthesizes as an anonymous record. Using a column the projection dropped
is a compile error thereafter (the 9b acceptance's projection criterion).
A query's value is `multi <element>`; a query wrapped directly in a whole-
query aggregate (`count(from …)`) is that aggregate's scalar. Execution is
eager at the point the expression is evaluated — no deferred queries, no query
values passed around (that would be a function value in disguise).
Aggregates are compiler-recognized **clause functions**, legal over a group
binding or a whole query: `count(g)`, `count(from …)`, `sum(g.f)`, `avg(g.f)`,
`min(g.f)`, `max(g.f)`. They are not general functions; naming one outside a
query is the existing unknown-identifier error.
Diagnostics this surface owns (new WO-E5xx range): unknown table, unknown
column (naming the class), relation navigated through a non-`ref` field,
aggregate outside a query, group member column used bare outside an aggregate,
join sides not separable, projection field name collision.
---
## 4. Aggregate semantics (normative table)
| Aggregate | Input | Result type | Empty source | Nil elements (`?T` column) |
|---|---|---|---|---|
| `count(g)` | group/query | `Int` | `0` | counted (row exists) |
| `sum(g.f)` | `Int` column | `Int` | `0` | skipped |
| `avg(g.f)` | `Int` column | `?Int` | `nil` | skipped; all-nil ⇒ `nil` |
| `min(g.f)` / `max(g.f)` | `Int` or `Text` column | `?T` | `nil` | skipped; all-nil ⇒ `nil` |
Decisions behind the table:
- **Empty is data, not a fault.** LINQ throws on empty `Average`/`Min`/`Max`
over non-nullables; writeonce has `?T` and a nil-forcing style already, so
the nullable-column LINQ behavior (`null` on empty, skip nils, all-nil ⇒
`null` — `Average.cs:165`, `Min.cs:47-56`) is the behavior for everyone.
Traps stay reserved for integrity violations.
- **`sum` wraps.** The VM's ADD wraps two's-complement; `sum` is a loop of
ADDs and inherits that. Documented, consistent with language arithmetic,
and the alternative (checked overflow, LINQ's `Sum.cs:40`) would make `sum`
the only trapping arithmetic in the language.
- **`avg` truncates** toward zero (integer division), same as `/`.
- Execution is the transition/finalize ABI (section 1, fork 3): one opaque
transition slot per (group, aggregate), advanced per row, finalized per
group when the hash table drains.
---
## 5. Lowering and execution model
A query compiles to an **ordinary bytecode loop** — the same opcodes every
`for` loop uses — over a small set of new engine cursor builtins. There is no
plan-tree interpreter in the VM: the compiler *is* the planner, at the only
scale this iteration promises (index selection, no cost model).
- **Scan**: table cursor builtins (open by class id, advance, read row into a
register as a borrowed row view). `where` clauses are ordinary compiled
boolean expressions guarding the loop body — the flattened-expression
lesson of PostgreSQL's executor (`execExpr.c`) is in our case simply "the
bytecode we already generate."
- **Index selection**: if the leading `where` conjuncts equality-match a
declared index's prefix (longest-prefix rule, iteration 9's `find_by`), the
compiler emits an index-probe cursor instead of a scan, and the acceptance
demonstrates the difference (a probe counter the engine exposes in debug
builds — "demonstrated, not assumed").
- **`ref` navigation** lowers to a primary-index point-read builtin.
**`backlink`** lowers to a secondary-index cursor over the ref column.
- **`group … into g … select`** lowers to the two-phase hash-aggregation
shape: phase one iterates the source, keying a group hash table and
advancing transition slots (builtins: group-table create / upsert-advance);
phase two drains the table, runs finalize, evaluates the `select`
projection per group. One node, no intermediate group objects — the
`AggregateBy` shape.
- **`join`** lowers to build-inner/probe-outer against a transient hash
keyed on the join column; nil keys are never inserted and never probe.
- **`order by`** materializes the result `multi` and stable-sorts it; `take`/
`skip` slice afterwards (a partial-sort fast path is recorded as a later
optimization, LINQ `SpeedOpt` precedent).
- **Ownership**: rows read from a table are engine-owned; anything a query
*returns* is copied out at the select boundary under Task-2's established
rule — a Text crossing an ownership boundary is copied, a projection record
is freshly built and caller-owned. Queries introduce no new ownership
classes.
Format consequences: new builtin ids appended to the format doc's table (the
cursor/group/probe set), no new opcodes, no version bump beyond iteration 9's.
---
## 6. Ownership, borrows, and GC across the engine boundary
The engine and the VM heap are two memory worlds, and the whole safety story
is that values only ever CROSS between them by copy. Analysis recorded here
because both iterations' correctness hangs on it (2026-08-15).
### The bulkhead: two one-way gates, both already doctrine
- **Into storage:** a row stores no VM pointer — scalars copy, Texts copy,
owned objects flatten by value, containers copy element-wise, `ref` is an
id, and a GC-managed value in a `@table` field is a **compile error**
(iteration 9's field-encoding rules). So no row ever points at a GC object.
- **Out of storage:** everything a `select` emits is copied or freshly built
at the boundary — Texts via the established copy rule, projections as new
records. So no GC root, no local, and no container ever points into a row
slab once the query ends.
Consequence: **the collector never traces engine memory and the engine never
touches reference counts.** Iteration 7b (inferred GC, mark-sweep) does not
change this — it changes only *when* the "into" gate's error fires: GC-ness
becomes inferred, so inference must classify every class **before** table-field
validation runs, and the diagnostic reads "class X is garbage-collected
(inferred via Y) and cannot be stored in a table field." A class stored in a
table is thereby constrained to ownership-expressible shapes; that is a
feature, not a limitation — tables are the language's answer to shared
long-lived data, which is most of what `@gc` exists for.
### Row views are borrows without a runtime net
A cursor yields a **row view**: a borrow of engine-owned memory, valid until
the cursor advances or closes. Two things make this different from every
borrow the language has today:
- VM-heap borrows have a runtime defense (the object header's borrow word,
`WO_T_BORROW` traps). Rows share the VM's field *encoding* but not its
header — there is no borrow word in a row slab, so **the compile-time rule
is load-bearing alone**. The ownership pass enforces: a row view never
escapes the query loop that produced it (the container-read-borrow mirror),
and anything that leaves does so as a copy through `select`.
- The program can mutate the table it is iterating — single-writer per shard
removes concurrent writers, not the program's own hand.
### Cursor stability: materialize ids, allow row updates, forbid structural
The `raise` mode is the honest case: it updates `salary` — an **indexed**
column — while iterating an index scan. Naive cursor-over-index breaks here
(entries move mid-scan). The semantics, chosen for KISS and enforceability:
- **A scan materializes its matching id list before the body runs**, then
point-reads each row per iteration. O(matches) ids of memory, recorded as
the cost; index-order iteration falls out for free.
- **Updates through the row view are allowed** — the view is an exclusive
borrow of that row for the iteration (the `mut` analog); index maintenance
for the changed column happens at the row API as always, and cannot disturb
the already-collected id list.
- **`insert` into or `delete` from a table with an open cursor is a compile
error** (new WO-E5xx): a materialized id list cannot defend a point-read
against a row deleted mid-loop, and silently skipping a vanished id is the
kind of quiet wrongness this language exists to refuse. The ownership pass
carries an open-cursor table set through the loop body, statically — insert
and delete name their target class at compile time. Read-only nested
queries over the same table remain legal (shared borrows).
### Query temporaries and GC pressure
Group hash tables and join build sides are engine-side C allocations scoped
to the statement — freed when the query ends, invisible to both the drop
tables and the collector. Query results are ordinary owned VM values, freed
by the existing drop machinery. Nothing on the query path allocates a GC
object or an RC operation. One accepted interaction: the budgeted collector
runs between statements, so a long full-table scan delays GC slices for its
duration — acceptable at this scale, recorded so nobody rediscovers it as a
latency mystery.
## 7. Acceptance workload — `docs/examples/employee`
A new sample, structured like `log-watcher` (wo.toml manifest, program mode,
`just` module, acceptance script), small enough to read in one sitting and
shaped so every 9b feature is load-bearing:
- **Types**: `Department` (`@table(name: "departments", index: [name])`,
`name: Text @unique`) and `Employee` (`@table(name: "employees",
index: [dept], index: [dept, salary])`, `name: Text`, `salary: Int` cents,
`hired: Int` ms, `dept: ref Department`); `Department.staff: backlink
Employee.dept`.
- **Modes** (CLI, exit codes per the program-mode contract):
- `seed` — inserts departments and employees; proves insert + WAL + unique
trap (second `seed` run reports the duplicate-department trap caught).
- `report` — the GroupBy showcase: headcount by department, average /
min / max salary by department, total payroll, departments ordered by
average salary — each line's expected output is byte-exact in the
acceptance script.
- `staff <department>` — relation navigation both directions: finds the
department by unique name (index probe), lists its `staff` backlink
ordered by salary; also prints each employee's `e.dept.name` to prove
forward navigation.
- `raise <department> <pct>` — update through a query result; re-running
`report` shows the moved averages; a `raise` for a missing department
exercises the empty-query path (`avg` ⇒ nil).
- **Integrity demo**: deleting a department with staff traps (restrict);
the acceptance asserts the trap code.
- **Crash step**: kill -9 between `seed` and `report`, re-run `report` —
iteration 9's WAL replay proven on this workload too.
The sample is 9b's acceptance the way log-watcher was iterations 1–7's: no new
corpus fixtures beyond the db corpus iteration 9 already plans; the sample is
the test.
## 8. Out of scope (inherited and new)
- Everything 9b's story already excludes: cross-shard queries and distributed
joins, `LIVE` subscriptions, migrations, cost-based planning.
- The `IGrouping`-as-value surface (groups escaping their query), `Distinct`/
set operators, outer joins (`LeftJoin` family), subqueries in `where`, and
`group by` composite keys beyond a single expression — each waits for a
workload that demands it.
- FK actions other than restrict (cascade, set-nil), deferred constraint
checking (the after-trigger queue pattern is recorded for when transactions
span statements).
- SQL text in any runtime role.

View file

@ -1,113 +0,0 @@
# writeonce-pl — the `.wo` language at a glance
writeonce is a **declarative full-stack programming language**. You write `.wo` files; the `wo` toolchain compiles them into a single binary that owns its database, serves REST, and pushes live subscriptions. The one-line pitch: **Go + Postgres + `net/http` + Phoenix LiveView, folded into one language and one binary.**
It is declarative by default — programs are built from `type`, `class`, `service`, `policy`, and `on <event>` declarations. A `type` declares a data shape; a `class` is its behavior-bearing sibling — the same fields plus `fn` methods with a `self` receiver. **There is no inheritance**: no `extends`, no overriding, no virtual dispatch — composition via `ref`/`multi`, Go-style. From either declaration the compiler derives the database schema, HTTP endpoints, triggers, and client SDKs. The imperative parts (queries, transactions, method bodies) use a hybrid SQL + Cypher sublanguage.
## The three pillars
| Pillar | Language construct | Status |
| ------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------ | --------------------------------------------------------------------------------------------------------------------- |
| **Live subscriptions** | `LIVE` prefix on queries; `subscribe` in a service's `expose` list. Deltas push on every commit over WebSocket — no polling. | Stage 3 — `/api/<type>/live` is a 501 stub today |
| **Front-end development** | `##ui` blocks compile to server-rendered HTML plus a small vanilla-JS client runtime. No Node, no React. | Phase 6 — design-only ([spec](runtime/database/06-lowcode-fullstack.md)) |
| **Database DML** | Hybrid SQL + Cypher + document paths: `$name` parameters everywhere, cross-paradigm `RETURNING col AS alias`, one `BEGIN … COMMIT` block syntax. | Query layer prototyped in C++ ([`prototypes/wo-db/`](../prototypes/wo-db/README.md)); schema layer shipped in Stage 2 |
## A complete program
One file is a full application — a database table, six REST endpoints, and a live subscription:
```wo
-- article.wo
type Article {
id: Id
title: Text
body: Markdown
author: Text
created_at: Timestamp = now()
service rest "/api/articles"
expose list, get, create, update, delete, subscribe
}
```
```bash
$ wo run
[wo] listening on :8080
```
All three pillars in ~10 lines: the `type` fields are the DML schema, `expose subscribe` is the live subscription, and (from Phase 6 on) a `##ui` block alongside it renders the screen.
## Implementation status
| Stage | What | Status |
| ----- | ------------------------------------------------------------ | ----------------------------------------------------------- |
| 1 | `wo run <dir>` discovers `.wo` files | ✅ shipped |
| 2 | parser + in-memory engine + REST CRUD | ✅ shipped — `cargo run --bin wo -- run docs/examples/blog` |
| 3 | LIVE subscriptions over WebSocket | pending (501 stub) |
| 4+ | transactional `fn`, policies, triggers, `##ui`, WAL, codegen | design-only |
**Class model status:** `class` declarations parse and serve REST CRUD today (plan 13a, shipped); method execution (13b), live pricing push (13c), and the MVC UI ([plan 14](plan/14-mvc-ui-implementation.md)) follow. Demo: [`examples/pricing/`](examples/pricing/); master plan: [plan 13](plan/13-class-model-live-pricing.md).
The runtime itself targets **zero external dependencies** — all I/O driven directly by Linux kernel primitives (`epoll`, `inotify`, `sendfile`, …); see the [kernel-primitive catalogue](plan/exploration/linux/00-linux.md).
## Where to read next
- [`runtime/wo-language.md`](runtime/wo-language.md) — the full user-facing language overview: toolchain, runtime stdlib, client SDKs
- [`runtime/database.md`](runtime/database.md) — the 7-phase engineering series behind the language
- [`runtime/database/02-wo-language.md`](runtime/database/02-wo-language.md) — the two-layer language spec (schema layer + query layer)
- [`examples/blog/README.md`](examples/blog/README.md) — the canonical worked example
- [`../prototypes/wo-db/README.md`](../prototypes/wo-db/README.md) — the C++ prototype of the query-layer engine
## Understanding the basics
Every layer of writeonce ultimately reduces to one primitive operation: **ask the kernel for memory, store a value in it, read it back**. Walking that operation up the abstraction ladder shows what the `.wo` syntax is actually hiding.
### Level 0 — assembly: the kernel gives you a page
```asm
; x86-64 Linux — map one anonymous page, store 42 in it
mov rax, 9 ; syscall number: mmap
xor rdi, rdi ; addr = NULL (kernel picks)
mov rsi, 4096 ; len = one page
mov rdx, 3 ; prot = PROT_READ | PROT_WRITE
mov r10, 0x22 ; flags = MAP_PRIVATE | MAP_ANONYMOUS
mov r8, -1 ; fd = none
xor r9, r9 ; off = 0
syscall ; rax now holds the page address
mov qword [rax], 42 ; store the value
mov rbx, [rax] ; read it back
```
There is no "variable" — only an address the kernel handed back and a `mov` into it.
### Level 1 — C: the libc wrapper names the address
```c
long *p = mmap(NULL, 4096, PROT_READ | PROT_WRITE,
MAP_PRIVATE | MAP_ANONYMOUS, -1, 0);
*p = 42; /* store */
long v = *p; /* read */
```
Same syscall, same page — C just gives the address a typed name and lets the compiler emit the `mov`s. (`malloc` is one more layer: a userland allocator carving up pages obtained exactly this way.)
### Level 2 — `.wo`: the value gets a type, a lifetime, and durability
```wo
type Counter {
name: Text @unique
value: Int = 0
}
main {
insert Counter { name: "visits" }; -- allocate + store
update Counter{ name == "visits" }.value += 1;
let c = select Counter{ name == "visits" };
print(c.value); -- read back
}
```
The `insert` is still, underneath, "obtain memory, write bytes at an offset" — but the declaration has been folded into the language: the `type` decides the layout, the engine owns the allocation (in-RAM rows over `mmap`-backed segments), and Phase 3+ adds what raw memory never had — ACID transactions, WAL durability, and `LIVE` subscribers notified on every store.
This is the whole design in miniature: the runtime is written in Rust against raw kernel primitives (level 0–1, see [`plan/exploration/linux/00-linux.md`](plan/exploration/linux/00-linux.md) and the [assembly stance](plan/exploration/assembly/00-overview.md)), so that the `.wo` author never has to leave level 2.

View file

@ -50,6 +50,11 @@ wovm-test:
# pass; the corpus below gates the individual behaviors underneath it. # pass; the corpus below gates the individual behaviors underneath it.
mod log-watcher "docs/examples/log-watcher" mod log-watcher "docs/examples/log-watcher"
# the database track's acceptance workload (iteration 9/9b): @table storage,
# ref/backlink relations + FK restrict, and the compiler-checked query surface
# (scan/where/select/order/take, update, delete). `just employee` runs it.
mod employee "docs/examples/employee"
# conformance harness (plan 3): walks tests/corpus/{run,compile-fail,trap}, # conformance harness (plan 3): walks tests/corpus/{run,compile-fail,trap},
# exact outcome per fixture kind — see docs/plan/oop-vm/02-corpus.md. # exact outcome per fixture kind — see docs/plan/oop-vm/02-corpus.md.
# Fails loudly (and names the recipe to run) if woc or wovm isn't built. # Fails loudly (and names the recipe to run) if woc or wovm isn't built.
@ -72,7 +77,7 @@ oop-accept:
echo "=== criterion 1: woc compile time, pricing subset (budget: under 100ms) ===" echo "=== criterion 1: woc compile time, pricing subset (budget: under 100ms) ==="
dune build --root compiler || fail "criterion 1: dune build --root compiler" dune build --root compiler || fail "criterion 1: dune build --root compiler"
WOC="$ROOT/compiler/_build/default/bin/woc" WOC="$ROOT/compiler/_build/default/bin/woc"
PRICING="tests/corpus/run/pricing-containers/fixture.wo tests/corpus/run/pricing-current-price/fixture.wo tests/corpus/run/pricing-discounted/fixture.wo tests/corpus/run/pricing-text/fixture.wo tests/corpus/trap/pricing-set-price-db-stub/fixture.wo" PRICING="tests/corpus/run/pricing-containers/fixture.wo tests/corpus/run/pricing-current-price/fixture.wo tests/corpus/run/pricing-discounted/fixture.wo tests/corpus/run/pricing-text/fixture.wo tests/corpus/run/pricing-set-price-insert/fixture.wo"
N=20 N=20
total_ns=0; max_ns=0; min_ns="" total_ns=0; max_ns=0; min_ns=""
for i in $(seq 1 "$N"); do for i in $(seq 1 "$N"); do

View file

@ -8,8 +8,13 @@ wo-rt: wo-rt.c
# ---- wovm VM core (src/) + unit tests (test/) ---- # ---- wovm VM core (src/) + unit tests (test/) ----
# Each test/test_*.c builds into its own ASan+UBSan binary linked against # Each test/test_*.c builds into its own ASan+UBSan binary linked against
# every src/*.c except main.c; `make test` runs them all. # every src/*.c except main.c; `make test` runs them all.
VMSRC := $(filter-out src/main.c,$(wildcard src/*.c)) # the database engine lives in its own top-level directory (iteration 9;
VMHDR := $(wildcard src/*.h) $(wildcard test/*.h) # user decision 2026-08-15) and is statically linked into every wovm and
# every test binary — one binary, unchanged
DBSRC := $(wildcard ../database/src/*.c)
DBHDR := $(wildcard ../database/src/*.h)
VMSRC := $(filter-out src/main.c,$(wildcard src/*.c)) $(DBSRC)
VMHDR := $(wildcard src/*.h) $(wildcard test/*.h) $(DBHDR)
TESTS := $(wildcard test/test_*.c) TESTS := $(wildcard test/test_*.c)
TESTBIN := $(TESTS:test/%.c=build/%) TESTBIN := $(TESTS:test/%.c=build/%)
TCFLAGS := -std=c11 -Wall -Wextra -Werror -g -O1 \ TCFLAGS := -std=c11 -Wall -Wextra -Werror -g -O1 \
@ -21,7 +26,7 @@ build:
TESTHELP := $(wildcard test/wob_build.c) TESTHELP := $(wildcard test/wob_build.c)
build/%: test/%.c $(VMSRC) $(TESTHELP) $(VMHDR) | build build/%: test/%.c $(VMSRC) $(TESTHELP) $(VMHDR) | build
$(CC) $(TCFLAGS) -Isrc -Itest -o $@ $< $(TESTHELP) $(VMSRC) $(CC) $(TCFLAGS) -Isrc -Itest -I../database/src -o $@ $< $(TESTHELP) $(VMSRC)
test: $(TESTBIN) test: $(TESTBIN)
@for t in $(TESTBIN); do echo "== $$t"; ./$$t || exit 1; done @for t in $(TESTBIN); do echo "== $$t"; ./$$t || exit 1; done
@ -31,20 +36,20 @@ test: $(TESTBIN)
ISOBIN := $(TESTS:test/%.c=build/iso_%) ISOBIN := $(TESTS:test/%.c=build/iso_%)
build/iso_%: test/%.c $(VMSRC) $(TESTHELP) $(VMHDR) | build build/iso_%: test/%.c $(VMSRC) $(TESTHELP) $(VMHDR) | build
$(CC) $(TCFLAGS) -DWO_ISO_C -Isrc -Itest -o $@ $< $(TESTHELP) $(VMSRC) $(CC) $(TCFLAGS) -DWO_ISO_C -Isrc -Itest -I../database/src -o $@ $< $(TESTHELP) $(VMSRC)
test-iso: $(ISOBIN) test-iso: $(ISOBIN)
@for t in $(ISOBIN); do echo "== $$t"; ./$$t || exit 1; done @for t in $(ISOBIN); do echo "== $$t"; ./$$t || exit 1; done
# the wovm binary (plain optimized build; the test suite is the ASan gate) # the wovm binary (plain optimized build; the test suite is the ASan gate)
wovm: src/main.c $(VMSRC) $(VMHDR) wovm: src/main.c $(VMSRC) $(VMHDR)
$(CC) $(CFLAGS) -Isrc -o $@ src/main.c $(VMSRC) $(CC) $(CFLAGS) -Isrc -I../database/src -o $@ src/main.c $(VMSRC)
# ASan+UBSan wovm, same flags as the unit tests, for corpus fixtures that # ASan+UBSan wovm, same flags as the unit tests, for corpus fixtures that
# need a sanitizer to prove a free actually happened (gc/ cycle fixtures) — # need a sanitizer to prove a free actually happened (gc/ cycle fixtures) —
# tasks 3/4 hand-built this each time because it didn't exist yet # tasks 3/4 hand-built this each time because it didn't exist yet
build/wovm_asan: src/main.c $(VMSRC) $(VMHDR) | build build/wovm_asan: src/main.c $(VMSRC) $(VMHDR) | build
$(CC) $(TCFLAGS) -Isrc -o $@ src/main.c $(VMSRC) $(CC) $(TCFLAGS) -Isrc -I../database/src -o $@ src/main.c $(VMSRC)
wovm-asan: build/wovm_asan wovm-asan: build/wovm_asan

View file

@ -2,6 +2,8 @@
#include "builtin.h" #include "builtin.h"
#include "db.h" /* database/src — the engine's statement executors */
#include <stdio.h> #include <stdio.h>
#include <string.h> #include <string.h>
#include <time.h> #include <time.h>
@ -68,6 +70,7 @@ int wo_builtin(wo_vm *vm, uint64_t *R, uint32_t ins, const char **msg) {
if (C == WO_B_JSON_ENCODE || C == WO_B_JSON_DECODE) if (C == WO_B_JSON_ENCODE || C == WO_B_JSON_DECODE)
return wo_builtin_json(vm, R, ins, msg); return wo_builtin_json(vm, R, ins, msg);
if (C >= WO_B_SYS_FIRST && C <= WO_B_PROC_RUN) return wo_builtin_sys(vm, R, ins, msg); if (C >= WO_B_SYS_FIRST && C <= WO_B_PROC_RUN) return wo_builtin_sys(vm, R, ins, msg);
if (C >= WO_B_DB_INSERT && C <= WO_B_DB_PROBE) return wo_builtin_db(vm, R, ins, msg);
switch (C) { switch (C) {
case WO_B_NOW: { /* wall-clock milliseconds */ case WO_B_NOW: { /* wall-clock milliseconds */
struct timespec ts; struct timespec ts;
@ -249,6 +252,10 @@ int wo_builtin(wo_vm *vm, uint64_t *R, uint32_t ins, const char **msg) {
} }
return 0; return 0;
} }
case WO_B_STR_LT: {
R[A] = elem_cmp(WO_K_TEXT, R[B], R[B + 1]) < 0 ? 1 : 0;
return 0;
}
case WO_B_TEXT_COPY: { /* nil copies to nil: a `?Text` crosses this boundary case WO_B_TEXT_COPY: { /* nil copies to nil: a `?Text` crosses this boundary
* exactly like a Text does */ * exactly like a Text does */
if (!R[B]) { if (!R[B]) {

View file

@ -36,6 +36,18 @@ static int rd_u64(cur_t *c, uint64_t *v) { return rd(c, v, 8); }
/* per-builtin fixed arity (args at B..B+arity-1); kind-immediate builtins /* per-builtin fixed arity (args at B..B+arity-1); kind-immediate builtins
* (multi_new/map_new) carry kinds in B and take no register args */ * (multi_new/map_new) carry kinds in B and take no register args */
static const uint8_t b_arity[WO_B_MAX + 1] = { static const uint8_t b_arity[WO_B_MAX + 1] = {
/* WO_B_DB_INSERT's window is class-id + one slot per DECLARED field —
variable, so the static table validates only the class-id slot (arity
1); the field slots are validated at runtime by the engine against
the class table (db.c / wo_row_insert). Same trust level as the
kind-immediate builtins' B nibble. */
[WO_B_DB_INSERT] = 1,
[WO_B_DB_UPDATE_FIELD] = 4,
[WO_B_DB_DELETE] = 2,
[WO_B_DB_SCAN] = 1,
[WO_B_DB_GET_FIELD] = 3,
[WO_B_DB_PROBE] = 3,
[WO_B_STR_LT] = 2,
[WO_B_NOW] = 0, [WO_B_PRINT] = 1, [WO_B_PRINT_INT] = 1, [WO_B_NOW] = 0, [WO_B_PRINT] = 1, [WO_B_PRINT_INT] = 1,
[WO_B_WORDS] = 1, [WO_B_MULTI_NEW] = 0, [WO_B_MULTI_PUSH] = 2, [WO_B_WORDS] = 1, [WO_B_MULTI_NEW] = 0, [WO_B_MULTI_PUSH] = 2,
[WO_B_MULTI_GET] = 2, [WO_B_COUNT] = 1, [WO_B_LATEST] = 1, [WO_B_MULTI_GET] = 2, [WO_B_COUNT] = 1, [WO_B_LATEST] = 1,
@ -141,7 +153,7 @@ int wo_load_buf(wo_module *m, const uint8_t *buf, size_t len, char *err,
m->classes = calloc(kcnt, sizeof(wo_classdesc)); m->classes = calloc(kcnt, sizeof(wo_classdesc));
if (!m->classes) BAIL("out of memory"); if (!m->classes) BAIL("out of memory");
} }
size_t pool_len = 0, meta_pool = 0; size_t pool_len = 0, meta_pool = 0, idx_pool = 0;
for (uint32_t i = 0; i < kcnt; i++) { for (uint32_t i = 0; i < kcnt; i++) {
uint32_t name, flags, fcnt; uint32_t name, flags, fcnt;
if (rd_u32(&k, &name) || rd_u32(&k, &flags) || rd_u32(&k, &fcnt)) if (rd_u32(&k, &name) || rd_u32(&k, &flags) || rd_u32(&k, &fcnt))
@ -197,6 +209,36 @@ int wo_load_buf(wo_module *m, const uint8_t *buf, size_t len, char *err,
m->classes[i].field_elem = (const uint32_t *)(uintptr_t)(meta_pool + 2u * (size_t)fcnt); m->classes[i].field_elem = (const uint32_t *)(uintptr_t)(meta_pool + 2u * (size_t)fcnt);
pool_len += fcnt ? fcnt : 1; pool_len += fcnt ? fcnt : 1;
meta_pool += meta_words ? meta_words : 1; meta_pool += meta_words ? meta_words : 1;
/* v3 tail: secondary indexes — flags/col_cnt/cols per index, columns
bounded and scalar/Text-kinded (the only indexable kinds) */
uint32_t icnt;
if (rd_u32(&k, &icnt)) BAIL("class %u: truncated index count", (unsigned)i);
if (icnt > 64) BAIL("class %u: too many indexes", (unsigned)i);
m->classes[i].idx_cnt = icnt;
m->classes[i].idx_meta = (const uint32_t *)(uintptr_t)idx_pool;
for (uint32_t x = 0; x < icnt; x++) {
uint32_t iflags, ccnt;
if (rd_u32(&k, &iflags) || rd_u32(&k, &ccnt))
BAIL("class %u index %u: truncated", (unsigned)i, (unsigned)x);
if (iflags & ~1u) BAIL("class %u index %u: unknown flags", (unsigned)i, (unsigned)x);
if (ccnt == 0 || ccnt > 8)
BAIL("class %u index %u: bad column count", (unsigned)i, (unsigned)x);
uint32_t *ip = realloc(m->idxpool, (idx_pool + 2 + ccnt) * sizeof(uint32_t));
if (!ip) BAIL("out of memory");
m->idxpool = ip;
m->idxpool[idx_pool++] = iflags;
m->idxpool[idx_pool++] = ccnt;
for (uint32_t cix = 0; cix < ccnt; cix++) {
uint32_t col;
if (rd_u32(&k, &col)) BAIL("class %u index %u: truncated column", (unsigned)i, (unsigned)x);
if (col >= fcnt) BAIL("class %u index %u: column out of range", (unsigned)i, (unsigned)x);
uint8_t kind = m->kindpool[(uintptr_t)m->classes[i].kinds + col];
if (kind != WO_K_SCALAR && kind != WO_K_TEXT)
BAIL("class %u index %u: column %u is not scalar or Text", (unsigned)i,
(unsigned)x, (unsigned)col);
m->idxpool[idx_pool++] = col;
}
}
m->class_cnt = i + 1; m->class_cnt = i + 1;
} }
for (uint32_t i = 0; i < m->class_cnt; i++) { for (uint32_t i = 0; i < m->class_cnt; i++) {
@ -204,6 +246,8 @@ int wo_load_buf(wo_module *m, const uint8_t *buf, size_t len, char *err,
m->classes[i].field_names = m->metapool + (uintptr_t)m->classes[i].field_names; m->classes[i].field_names = m->metapool + (uintptr_t)m->classes[i].field_names;
m->classes[i].field_class = m->metapool + (uintptr_t)m->classes[i].field_class; m->classes[i].field_class = m->metapool + (uintptr_t)m->classes[i].field_class;
m->classes[i].field_elem = m->metapool + (uintptr_t)m->classes[i].field_elem; m->classes[i].field_elem = m->metapool + (uintptr_t)m->classes[i].field_elem;
m->classes[i].idx_meta =
m->idxpool ? m->idxpool + (uintptr_t)m->classes[i].idx_meta : NULL;
} }
/* ---- interfaces + vtable rows (expanded to sorted triples) ---- */ /* ---- interfaces + vtable rows (expanded to sorted triples) ---- */
@ -526,6 +570,7 @@ void wo_module_free(wo_module *m) {
free(m->classes); free(m->classes);
free(m->kindpool); free(m->kindpool);
free(m->metapool); free(m->metapool);
free(m->idxpool);
free(m->vtabs); free(m->vtabs);
for (uint32_t i = 0; i < m->method_cnt; i++) { for (uint32_t i = 0; i < m->method_cnt; i++) {
free(m->methods[i].code); free(m->methods[i].code);

View file

@ -54,6 +54,7 @@ typedef struct wo_module {
/* pooled per-field metadata (v2): names, referenced class ids, element /* pooled per-field metadata (v2): names, referenced class ids, element
kinds — see wob.h's "class-table field metadata" note */ kinds — see wob.h's "class-table field metadata" note */
uint32_t *metapool; uint32_t *metapool;
uint32_t *idxpool; /* v3 pooled per-class index metadata (flags/cols) */
uint32_t slot_cnt; /* total interface slots across all interfaces */ uint32_t slot_cnt; /* total interface slots across all interfaces */
wo_vtabent *vtabs; wo_vtabent *vtabs;
uint32_t vtab_cnt; uint32_t vtab_cnt;

View file

@ -13,9 +13,13 @@
#include "cont.h" #include "cont.h"
#include "gc.h" #include "gc.h"
#include "table.h"
#include "vm.h" #include "vm.h"
#include "wal.h"
static wo_vm VM; /* 32K value stack: keep it off the C stack */ static wo_vm VM; /* 32K value stack: keep it off the C stack */
static wo_db DB; /* the per-shard engine (one shard until iteration 8) */
static wo_wal WAL;
/* ---- self-exec detection (Task 6, plan 3) ------------------------------- /* ---- self-exec detection (Task 6, plan 3) -------------------------------
* `woc build` makes a single executable by copying wovm and appending the * `woc build` makes a single executable by copying wovm and appending the
@ -157,6 +161,39 @@ int main(int argc, char **argv) {
wo_module_free(&mod); wo_module_free(&mod);
return 2; return 2;
} }
/* The database engine boots with the VM: every class IS a table.
* Durability is opt-in — WO_DATA=<dir> opens <dir>/shard-0.wal,
* replays it before the entry runs (boot-before-listeners doctrine),
* and every insert commits before it acknowledges. Without WO_DATA
* the engine runs RAM-only, which is what the corpus expects. */
if (wo_db_init(&DB, mod.classes, mod.class_cnt, 0, 1) != 0) {
fprintf(stderr, "wovm: cannot initialize the database engine\n");
wo_vm_destroy(&VM);
wo_module_free(&mod);
return 2;
}
VM.rt.db = &DB;
const char *data_dir = getenv("WO_DATA");
if (data_dir && data_dir[0]) {
char wal_path[512];
snprintf(wal_path, sizeof wal_path, "%s/shard-0.wal", data_dir);
if (wo_wal_replay(wal_path, &DB) < 0) {
fprintf(stderr, "wovm: %s: replay found corruption beyond a torn tail\n", wal_path);
wo_db_destroy(&DB);
wo_vm_destroy(&VM);
wo_module_free(&mod);
return 2;
}
if (wo_wal_open(&WAL, wal_path, 1u << 20) != 0) {
fprintf(stderr, "wovm: cannot open %s\n", wal_path);
wo_db_destroy(&DB);
wo_vm_destroy(&VM);
wo_module_free(&mod);
return 2;
}
VM.rt.wal = &WAL;
}
/* Program mode: an entry that declares one parameter gets the program's /* Program mode: an entry that declares one parameter gets the program's
* OWN arguments as a `multi Text` — not the program name, and not the * OWN arguments as a `multi Text` — not the program name, and not the
* image path a plain `wovm image.wob args...` invocation carries. So * image path a plain `wovm image.wob args...` invocation carries. So
@ -207,6 +244,8 @@ int main(int argc, char **argv) {
* heap is torn down, and after a trap too: the container outlives the * heap is torn down, and after a trap too: the container outlives the
* unwind. */ * unwind. */
if (argv_val) wo_drop_kind(&VM.rt, WO_K_MULTI, argv_val); if (argv_val) wo_drop_kind(&VM.rt, WO_K_MULTI, argv_val);
if (VM.rt.wal) wo_wal_close(&WAL);
wo_db_destroy(&DB);
gc_pump(&VM); gc_pump(&VM);
wo_vm_destroy(&VM); wo_vm_destroy(&VM);
wo_module_free(&mod); wo_module_free(&mod);

View file

@ -38,6 +38,12 @@ typedef struct wo_rt {
size_t len, cap; size_t len, cap;
} cycbuf; } cycbuf;
void *out; /* FILE*; kept void* so obj.h needn't pull in stdio */ void *out; /* FILE*; kept void* so obj.h needn't pull in stdio */
/* the database engine's handles (database/src), opaque here so the VM
core needn't include engine headers: db = wo_db*, wal = wo_wal*.
NULL = engine absent (test binaries) / durability off (no WO_DATA).
Set by main.c at boot; db.c casts. */
void *db;
void *wal;
} wo_rt; } wo_rt;
int wo_rt_init(wo_rt *rt, size_t heap_cap, const wo_classdesc *classes, int wo_rt_init(wo_rt *rt, size_t heap_cap, const wo_classdesc *classes,

View file

@ -12,7 +12,7 @@
/* ---- file header (44 bytes, absolute offsets) ---- */ /* ---- file header (44 bytes, absolute offsets) ---- */
#define WOB_MAGIC 0x31424F57u /* "WOB1" read as LE u32 */ #define WOB_MAGIC 0x31424F57u /* "WOB1" read as LE u32 */
#define WOB_VERSION 2u /* v2 adds per-field names/types to the class table */ #define WOB_VERSION 3u /* v3: v2's field metadata + per-class index metadata */
#define WOB_HDR_SIZE 44u #define WOB_HDR_SIZE 44u
#define WOB_OFF_MAGIC 0u #define WOB_OFF_MAGIC 0u
#define WOB_OFF_VERSION 4u #define WOB_OFF_VERSION 4u
@ -121,6 +121,13 @@ enum {
* message rides along in the error record, and `try ... catch` is how * message rides along in the error record, and `try ... catch` is how
* a program that expects the failure handles it. */ * a program that expects the failure handles it. */
WO_T_IO = 9, WO_T_IO = 9,
/* iteration 9 Task 4: a unique-index violation on insert/update —
raised by the engine at the row choke point, catchable like any
trap (the employee sample's SEED-DUP line) */
WO_T_UNIQUE = 10,
/* iteration 9b: deleting a row still referenced by a `ref` traps here
(restrict) — the employee sample's DROP-of-a-department-with-staff */
WO_T_FK = 11,
}; };
/* ---- opcodes (spec section 5; semantics in the format doc) ---- */ /* ---- opcodes (spec section 5; semantics in the format doc) ---- */
@ -285,8 +292,38 @@ enum {
* out of a function (`return` of a borrowed place, which is what this id * out of a function (`return` of a borrowed place, which is what this id
* exists for: the callee's borrow must not become the caller's owner). */ * exists for: the callee's borrow must not become the caller's owner). */
WO_B_TEXT_COPY = 60, WO_B_TEXT_COPY = 60,
/* ---- database engine (iteration 9; database/src/db.c) ----
* DB_INSERT window: R[B] = class id, R[B+1..] = one slot per declared
* field in declaration order. Result R[A] = the new row's id. Engine
* failures trap WO_T_DB; a failed WAL commit traps WO_T_IO (the write
* was applied to RAM but never acknowledged). */
WO_B_DB_INSERT = 61,
/* DB_UPDATE_FIELD: R[B] = class id, R[B+1] = row id, R[B+2] = field
* index, R[B+3] = the value. R[A] = 0. Unique violation traps
* WO_T_UNIQUE with the row untouched. */
WO_B_DB_UPDATE_FIELD = 62,
/* DB_DELETE: R[B] = class id, R[B+1] = row id. R[A] = 0. A missing row
* traps WO_T_DB (deleting what is not there is a fault, not a no-op). */
WO_B_DB_DELETE = 63,
/* the query surface's reads (iteration 9b). A table-class value IS its
* row id at runtime (the "objects are rows" model), so these are how the
* compiled query loop touches storage:
* DB_SCAN (64): R[B] = class -> R[A] = multi<Int> of every id
* DB_GET_FIELD(65): R[B]=class, R[B+1]=id, R[B+2]=field
* -> R[A] = that field, decoded to a VM value (a
* Text field decodes to a fresh Text; a ref field
* decodes to the target id). Missing row traps
* WO_T_DB.
* DB_PROBE (66): R[B]=class, R[B+1]=index, R[B+2]=key
* -> R[A] = multi<Int> of ids whose first indexed
* column equals key (backlink + indexed where). */
WO_B_DB_SCAN = 64,
WO_B_DB_GET_FIELD = 65,
WO_B_DB_PROBE = 66,
WO_B_STR_LT = 67, /* (a, b) text -> 1 if a < b by content, else 0 (query
* order-by on a Text key; scalars use the LT opcode) */
}; };
#define WO_B_MAX 60u #define WO_B_MAX 67u
/* ids at or above this one live in sysio.c, not builtin.c */ /* ids at or above this one live in sysio.c, not builtin.c */
#define WO_B_SYS_FIRST WO_B_FS_EXISTS #define WO_B_SYS_FIRST WO_B_FS_EXISTS
@ -324,6 +361,11 @@ typedef struct wo_classdesc {
const uint32_t *field_names; /* constant index of each field's name */ const uint32_t *field_names; /* constant index of each field's name */
const uint32_t *field_class; /* referenced class id / JSON_RAW / NONE */ const uint32_t *field_class; /* referenced class id / JSON_RAW / NONE */
const uint32_t *field_elem; /* container element kinds */ const uint32_t *field_elem; /* container element kinds */
/* v3 (iteration 9 Task 4): the class's secondary indexes, flat-encoded
[flags, col_cnt, col...]* — flags bit0 = unique. idx_cnt entries.
Columns are field indices, scalar/Text kinds only (loader-checked). */
uint32_t idx_cnt;
const uint32_t *idx_meta;
} wo_classdesc; } wo_classdesc;
#define WO_CLASSF_GC 0x01u #define WO_CLASSF_GC 0x01u

243
runtime/test/test_table.c Normal file
View file

@ -0,0 +1,243 @@
/* test_table — iteration 9 Task 1: class-shaped row storage.
* Round-trips across kinds, nil encodings, id interleave across shards,
* slab growth past one slab, slot reuse after removal, and the out-gate
* invariant (a read hands back FRESH VM values, never slab pointers). */
#include <string.h>
#include "cont.h"
#include "gc.h"
#include "obj.h"
#include "t.h"
#include "table.h"
/* class 0: Addr { city: Text }
* class 1: Emp { name: Text, salary: Int(scalar), addr: OWNED Addr,
* tags: multi Text, meta: map<Text, scalar> }
* class 2: Tiny { n: scalar } (slab-growth workhorse) */
static const uint8_t addr_kinds[] = {WO_K_TEXT};
static const uint8_t emp_kinds[] = {WO_K_TEXT, WO_K_SCALAR, WO_K_OWNED, WO_K_MULTI,
WO_K_MAP};
static const uint8_t tiny_kinds[] = {WO_K_SCALAR};
static const wo_classdesc CLASSES[] = {
{.name = 0, .flags = 0, .field_cnt = 1, .kinds = addr_kinds},
{.name = 0, .flags = 0, .field_cnt = 5, .kinds = emp_kinds},
{.name = 0, .flags = 0, .field_cnt = 1, .kinds = tiny_kinds},
};
static void test_roundtrip_all_kinds(void) {
wo_rt rt;
T_EQ(wo_rt_init(&rt, 1 << 20, CLASSES, 3), 0);
wo_db db;
T_EQ(wo_db_init(&db, CLASSES, 3, 0, 1), 0);
const char *msg = "";
/* build the VM-side value: Emp{"Asha", 9200000, Addr{"Pune"}, ["a","b"], {"k": 7}} */
wo_str *name = wo_str_new(&rt, "Asha", 4);
wo_hdr *addr = wo_obj_new(&rt, 0);
wo_fields(addr)[0] = (uint64_t)(uintptr_t)wo_str_new(&rt, "Pune", 4);
wo_multi *tags = wo_multi_new(&rt, WO_K_TEXT);
wo_multi_push(tags, (uint64_t)(uintptr_t)wo_str_new(&rt, "a", 1));
wo_multi_push(tags, (uint64_t)(uintptr_t)wo_str_new(&rt, "b", 1));
wo_map *meta = wo_map_new(&rt, WO_K_TEXT, WO_K_SCALAR);
uint64_t old;
wo_map_set(meta, (uint64_t)(uintptr_t)wo_str_new(&rt, "k", 1), 7, &old);
uint64_t vals[5] = {(uint64_t)(uintptr_t)name, 9200000,
(uint64_t)(uintptr_t)addr, (uint64_t)(uintptr_t)tags,
(uint64_t)(uintptr_t)meta};
uint64_t id = wo_row_insert(&db, 1, vals, &msg, NULL);
T_EQ(id, 1); /* shard 0 of 1: first id is 1 */
/* the row stored COPIES: mutate the VM originals, then read back */
name->data[0] = 'X';
((wo_str *)(uintptr_t)wo_fields(addr)[0])->data[0] = 'X';
uint64_t out[5] = {0};
T_EQ(wo_row_read(&db, &rt, 1, id, out, &msg), 0);
wo_str *rname = (wo_str *)(uintptr_t)out[0];
T_EQ(rname->len, 4);
T_CHECK(memcmp(rname->data, "Asha", 4) == 0); /* not "Xsha" */
T_CHECK(rname != name); /* fresh allocation */
T_EQ(out[1], 9200000);
wo_hdr *raddr = (wo_hdr *)(uintptr_t)out[2];
T_CHECK(raddr != addr);
wo_str *rcity = (wo_str *)(uintptr_t)wo_fields(raddr)[0];
T_CHECK(memcmp(rcity->data, "Pune", 4) == 0); /* not "Xune" */
wo_multi *rtags = (wo_multi *)(uintptr_t)out[3];
T_EQ(rtags->len, 2);
T_CHECK(memcmp(((wo_str *)(uintptr_t)rtags->items[1])->data, "b", 1) == 0);
wo_map *rmeta = (wo_map *)(uintptr_t)out[4];
uint64_t got = 0;
wo_str *k = wo_str_new(&rt, "k", 1);
T_EQ(wo_map_get(rmeta, (uint64_t)(uintptr_t)k, &got), 0);
T_EQ(got, 7);
/* nil TEXT / nil OWNED / WO_NIL_SCALAR round-trip */
uint64_t nilvals[5] = {0, WO_NIL_SCALAR, 0, 0, 0};
uint64_t id2 = wo_row_insert(&db, 1, nilvals, &msg, NULL);
T_EQ(id2, 2);
uint64_t out2[5] = {(uint64_t)-1, 0, (uint64_t)-1, (uint64_t)-1, (uint64_t)-1};
T_EQ(wo_row_read(&db, &rt, 1, id2, out2, &msg), 0);
T_EQ(out2[0], 0);
T_EQ(out2[1], WO_NIL_SCALAR);
T_EQ(out2[2], 0);
T_EQ(out2[3], 0);
/* the VM-side values are containers with malloc'd backing arrays:
real drops, not arena teardown, are what frees them */
wo_drop_obj(&rt, (wo_hdr *)name);
wo_drop_obj(&rt, addr);
wo_drop_obj(&rt, (wo_hdr *)tags);
wo_drop_obj(&rt, (wo_hdr *)meta);
wo_drop_obj(&rt, (wo_hdr *)k);
for (int i = 0; i < 5; i++)
if (i != 1 && out[i]) wo_drop_obj(&rt, (wo_hdr *)(uintptr_t)out[i]);
wo_db_destroy(&db);
wo_rt_destroy(&rt);
}
static void test_id_interleave_across_shards(void) {
const char *msg = "";
wo_db a, b, c;
T_EQ(wo_db_init(&a, CLASSES, 3, 0, 3), 0);
T_EQ(wo_db_init(&b, CLASSES, 3, 1, 3), 0);
T_EQ(wo_db_init(&c, CLASSES, 3, 2, 3), 0);
uint64_t v[1] = {42};
T_EQ(wo_row_insert(&a, 2, v, &msg, NULL), 1); /* shard 0: 1, 4, 7 */
T_EQ(wo_row_insert(&a, 2, v, &msg, NULL), 4);
T_EQ(wo_row_insert(&b, 2, v, &msg, NULL), 2); /* shard 1: 2, 5 */
T_EQ(wo_row_insert(&b, 2, v, &msg, NULL), 5);
T_EQ(wo_row_insert(&c, 2, v, &msg, NULL), 3); /* shard 2: 3, 6 */
T_EQ(wo_row_insert(&c, 2, v, &msg, NULL), 6);
/* owner-shard discipline: (id-1) % N names the shard */
T_EQ((4 - 1) % 3, 0);
T_EQ((5 - 1) % 3, 1);
T_EQ((6 - 1) % 3, 2);
/* shard/nshards misuse refused */
wo_db bad;
T_EQ(wo_db_init(&bad, CLASSES, 3, 3, 3), -1);
T_EQ(wo_db_init(&bad, CLASSES, 3, 0, 0), -1);
wo_db_destroy(&a);
wo_db_destroy(&b);
wo_db_destroy(&c);
}
static void test_slab_growth_and_reuse(void) {
wo_rt rt;
T_EQ(wo_rt_init(&rt, 1 << 20, CLASSES, 3), 0);
const char *msg = "";
wo_db db;
T_EQ(wo_db_init(&db, CLASSES, 3, 0, 1), 0);
/* three slabs' worth of Tiny rows */
enum { N = 3 * DB_SLAB_ROWS + 5 };
uint64_t ids[N];
for (uint32_t i = 0; i < N; i++) {
uint64_t v[1] = {i};
ids[i] = wo_row_insert(&db, 2, v, &msg, NULL);
T_CHECK(ids[i] == i + 1);
}
T_EQ(db.tables[2].slab_cnt, 4);
T_EQ(db.tables[2].count, N);
/* every row readable after growth (addresses were never moved) */
uint64_t out[1];
T_EQ(wo_row_read(&db, &rt, 2, ids[0], out, &msg), 0);
T_EQ(out[0], 0);
T_EQ(wo_row_read(&db, &rt, 2, ids[N - 1], out, &msg), 0);
T_EQ(out[0], N - 1);
/* remove a middle row: its slot is reused BEFORE any new slab grows */
db_row *victim = wo_row_ptr(&db, 2, ids[100]);
T_CHECK(victim != NULL);
T_EQ(wo_row_remove(&db, 2, ids[100]), 0);
T_EQ(wo_row_read(&db, &rt, 2, ids[100], out, &msg), -1); /* gone */
T_EQ(wo_row_remove(&db, 2, ids[100]), -1); /* twice = miss */
uint64_t v[1] = {777};
uint64_t fresh = wo_row_insert(&db, 2, v, &msg, NULL);
T_CHECK(fresh > (uint64_t)N); /* ids never reused ... */
db_row *fresh_row = wo_row_ptr(&db, 2, fresh);
T_CHECK(fresh_row == victim); /* ... but the SLOT is */
T_EQ(db.tables[2].slab_cnt, 4);
wo_db_destroy(&db);
wo_rt_destroy(&rt);
}
/* Task 5: single-field update through the choke point — value swapped,
* indexes moved, unique violations leave the row untouched. Class 3 in a
* local table: User { email: Text @unique-ish } via idx metadata. */
static const uint8_t user_kinds[] = {WO_K_TEXT, WO_K_SCALAR};
static const uint32_t user_idx_meta[] = {1 /*unique*/, 1, 0 /*col: email*/,
0 /*non-unique*/, 1, 1 /*col: n*/};
static const wo_classdesc UCLASSES[] = {
{.name = 0, .flags = 0, .field_cnt = 2, .kinds = user_kinds, .idx_cnt = 2,
.idx_meta = user_idx_meta},
};
static void test_update_field(void) {
wo_rt rt;
T_EQ(wo_rt_init(&rt, 1 << 20, UCLASSES, 1), 0);
wo_db db;
T_EQ(wo_db_init(&db, UCLASSES, 1, 0, 1), 0);
const char *msg = "";
int ek = 0;
wo_str *e1 = wo_str_new(&rt, "a@x", 3);
wo_str *e2 = wo_str_new(&rt, "b@x", 3);
uint64_t v1[2] = {(uint64_t)(uintptr_t)e1, 10};
uint64_t v2[2] = {(uint64_t)(uintptr_t)e2, 20};
uint64_t a = wo_row_insert(&db, 0, v1, &msg, NULL);
uint64_t b = wo_row_insert(&db, 0, v2, &msg, NULL);
T_CHECK(a && b);
/* scalar update: value moves, non-unique index follows */
T_EQ(wo_row_update_field(&db, 0, a, 1, 99, &msg, &ek), 0);
uint64_t out[2];
T_EQ(wo_row_read(&db, &rt, 0, a, out, &msg), 0);
T_EQ(out[1], 99);
wo_str_free(&rt, (wo_str *)(uintptr_t)out[0]);
/* unique violation: updating a's email to b's must refuse, row untouched */
wo_str *dupe = wo_str_new(&rt, "b@x", 3);
T_EQ(wo_row_update_field(&db, 0, a, 0, (uint64_t)(uintptr_t)dupe, &msg, &ek), -1);
T_EQ(ek, DB_ERR_UNIQUE);
T_EQ(wo_row_read(&db, &rt, 0, a, out, &msg), 0);
wo_str *still = (wo_str *)(uintptr_t)out[0];
T_CHECK(still->len == 3 && memcmp(still->data, "a@x", 3) == 0);
wo_str_free(&rt, still);
/* legal text update: old engine value freed (ASan), index moved — the
old email becomes free for someone else */
wo_str *fresh = wo_str_new(&rt, "c@x", 3);
T_EQ(wo_row_update_field(&db, 0, a, 0, (uint64_t)(uintptr_t)fresh, &msg, &ek), 0);
wo_str *take_a = wo_str_new(&rt, "a@x", 3);
uint64_t v3[2] = {(uint64_t)(uintptr_t)take_a, 30};
uint64_t cid = wo_row_insert(&db, 0, v3, &msg, &ek);
T_CHECK(cid != 0); /* "a@x" released by the update */
wo_drop_obj(&rt, (wo_hdr *)e1);
wo_drop_obj(&rt, (wo_hdr *)e2);
wo_drop_obj(&rt, (wo_hdr *)dupe);
wo_drop_obj(&rt, (wo_hdr *)fresh);
wo_drop_obj(&rt, (wo_hdr *)take_a);
wo_db_destroy(&db);
wo_rt_destroy(&rt);
}
static void test_misuse(void) {
const char *msg = "";
wo_db db;
T_EQ(wo_db_init(&db, CLASSES, 3, 0, 1), 0);
uint64_t v[1] = {1};
T_EQ(wo_row_insert(&db, 99, v, &msg, NULL), 0); /* unknown class */
T_CHECK(wo_row_ptr(&db, 99, 1) == NULL);
T_CHECK(wo_row_ptr(&db, 2, 1) == NULL); /* table never touched */
T_EQ(wo_row_remove(&db, 2, 1), -1);
wo_db_destroy(&db);
}
int main(void) {
test_roundtrip_all_kinds();
test_id_interleave_across_shards();
test_slab_growth_and_reuse();
test_update_field();
test_misuse();
return t_report("test_table");
}

282
runtime/test/test_wal.c Normal file
View file

@ -0,0 +1,282 @@
/* test_wal — iteration 9 Task 2: typed WAL + boot replay.
* Round-trip through a replay, torn-tail drop, reopen-overwrites-tear,
* and the commit-then-kill crash battery: a forked child inserts rows and
* acks each COMMITTED id over a pipe; SIGKILL lands mid-stream; the parent
* verifies with the offline oracle and a replay that every acked id is
* present with the right contents. */
#define _POSIX_C_SOURCE 200809L
#include <fcntl.h>
#include <signal.h>
#include <stdlib.h>
#include <string.h>
#include <sys/wait.h>
#include <time.h>
#include <unistd.h>
#include "gc.h"
#include "obj.h"
#include "t.h"
#include "table.h"
#include "wal.h"
/* class 0: Row { n: scalar, label: Text } */
static const uint8_t row_kinds[] = {WO_K_SCALAR, WO_K_TEXT};
static const wo_classdesc CLASSES[] = {
{.name = 0, .flags = 0, .field_cnt = 2, .kinds = row_kinds},
};
static char g_dir[64];
static void test_roundtrip_replay(void) {
char path[128];
snprintf(path, sizeof path, "%s/basic.wal", g_dir);
wo_rt rt;
T_EQ(wo_rt_init(&rt, 1 << 20, CLASSES, 1), 0);
wo_db db;
T_EQ(wo_db_init(&db, CLASSES, 1, 0, 1), 0);
wo_wal w;
T_EQ(wo_wal_open(&w, path, 1 << 16), 0);
const char *msg = "";
/* three inserts and one remove, RAM first, WAL second, one commit */
uint64_t ids[3];
for (int i = 0; i < 3; i++) {
wo_str *s = wo_str_new(&rt, "abcXYZ" + i, 3); /* "abc","bcX","cXY" */
uint64_t vals[2] = {(uint64_t)(i * 10), (uint64_t)(uintptr_t)s};
ids[i] = wo_row_insert(&db, 0, vals, &msg, NULL);
T_CHECK(ids[i] != 0);
T_EQ(wo_wal_append_insert(&w, &db, 0, ids[i]), 0);
wo_str_free(&rt, s);
}
T_EQ(wo_row_remove(&db, 0, ids[1]), 0);
T_EQ(wo_wal_append_remove(&w, 0, ids[1]), 0);
T_EQ(wo_wal_commit(&w), 0);
wo_wal_close(&w);
wo_db_destroy(&db);
/* boot: fresh engine, replay, deep-compare */
wo_db db2;
T_EQ(wo_db_init(&db2, CLASSES, 1, 0, 1), 0);
T_EQ(wo_wal_replay(path, &db2), 4);
uint64_t out[2];
T_EQ(wo_row_read(&db2, &rt, 0, ids[0], out, &msg), 0);
T_EQ(out[0], 0);
wo_str *s0 = (wo_str *)(uintptr_t)out[1];
T_CHECK(s0->len == 3 && memcmp(s0->data, "abc", 3) == 0);
wo_str_free(&rt, s0);
T_EQ(wo_row_read(&db2, &rt, 0, ids[1], out, &msg), -1); /* removed */
T_EQ(wo_row_read(&db2, &rt, 0, ids[2], out, &msg), 0);
T_EQ(out[0], 20);
wo_str_free(&rt, (wo_str *)(uintptr_t)out[1]);
/* next_id advanced past the replayed ids: a fresh insert never collides */
uint64_t vals[2] = {99, 0};
uint64_t fresh = wo_row_insert(&db2, 0, vals, &msg, NULL);
T_CHECK(fresh > ids[2]);
wo_db_destroy(&db2);
/* update record: re-log, replay replaces */
{
char upath[128];
snprintf(upath, sizeof upath, "%s/upd.wal", g_dir);
wo_db du;
T_EQ(wo_db_init(&du, CLASSES, 1, 0, 1), 0);
wo_wal wu;
T_EQ(wo_wal_open(&wu, upath, 0), 0);
wo_str *s1 = wo_str_new(&rt, "old", 3);
uint64_t uv[2] = {7, (uint64_t)(uintptr_t)s1};
uint64_t uid = wo_row_insert(&du, 0, uv, &msg, NULL);
T_EQ(wo_wal_append_insert(&wu, &du, 0, uid), 0);
int ek = 0;
wo_str *s2 = wo_str_new(&rt, "new!", 4);
T_EQ(wo_row_update_field(&du, 0, uid, 1, (uint64_t)(uintptr_t)s2, &msg, &ek), 0);
T_EQ(wo_row_update_field(&du, 0, uid, 0, 8, &msg, &ek), 0);
T_EQ(wo_wal_append_update(&wu, &du, 0, uid), 0);
T_EQ(wo_wal_commit(&wu), 0);
wo_wal_close(&wu);
wo_db_destroy(&du);
wo_db db4;
T_EQ(wo_db_init(&db4, CLASSES, 1, 0, 1), 0);
T_EQ(wo_wal_replay(upath, &db4), 2);
uint64_t uo[2];
T_EQ(wo_row_read(&db4, &rt, 0, uid, uo, &msg), 0);
T_EQ(uo[0], 8);
wo_str *us = (wo_str *)(uintptr_t)uo[1];
T_CHECK(us->len == 4 && memcmp(us->data, "new!", 4) == 0);
wo_str_free(&rt, us);
wo_drop_obj(&rt, (wo_hdr *)s1);
wo_drop_obj(&rt, (wo_hdr *)s2);
wo_db_destroy(&db4);
}
/* replay of a missing file is a fresh boot, not an error */
wo_db db3;
T_EQ(wo_db_init(&db3, CLASSES, 1, 0, 1), 0);
T_EQ(wo_wal_replay("/nonexistent/nope.wal", &db3), 0);
wo_db_destroy(&db3);
wo_rt_destroy(&rt);
}
static void test_torn_tail(void) {
char path[128];
snprintf(path, sizeof path, "%s/torn.wal", g_dir);
wo_db db;
T_EQ(wo_db_init(&db, CLASSES, 1, 0, 1), 0);
wo_wal w;
T_EQ(wo_wal_open(&w, path, 0), 0);
const char *msg = "";
for (int i = 0; i < 5; i++) {
uint64_t vals[2] = {(uint64_t)i, 0};
uint64_t id = wo_row_insert(&db, 0, vals, &msg, NULL);
T_EQ(wo_wal_append_insert(&w, &db, 0, id), 0);
T_EQ(wo_wal_commit(&w), 0);
}
uint64_t intact_end = w.off;
/* tear: append half a record's worth of a valid-looking header + junk */
uint32_t fake_len = 40, fake_crc = 0xDEAD;
uint8_t junk[20] = {7, 7, 7};
T_CHECK(pwrite(w.fd, &fake_len, 4, (off_t)intact_end) == 4);
T_CHECK(pwrite(w.fd, &fake_crc, 4, (off_t)(intact_end + 4)) == 4);
T_CHECK(pwrite(w.fd, junk, sizeof junk, (off_t)(intact_end + 8)) == (ssize_t)sizeof junk);
wo_wal_close(&w);
wo_db_destroy(&db);
/* the oracle sees exactly the intact prefix */
uint64_t at = 0;
T_EQ(wo_wal_check(path, &at), 5);
T_EQ(at, intact_end);
/* replay drops the tear whole */
wo_db db2;
T_EQ(wo_db_init(&db2, CLASSES, 1, 0, 1), 0);
T_EQ(wo_wal_replay(path, &db2), 5);
T_EQ(db2.tables[0].count, 5);
wo_db_destroy(&db2);
/* reopen positions AT the tear: the next commit overwrites it */
wo_db db3;
T_EQ(wo_db_init(&db3, CLASSES, 1, 0, 1), 0);
T_EQ(wo_wal_replay(path, &db3), 5);
wo_wal w2;
T_EQ(wo_wal_open(&w2, path, 0), 0);
T_EQ(w2.off, intact_end);
uint64_t vals[2] = {100, 0};
uint64_t id = wo_row_insert(&db3, 0, vals, &msg, NULL);
T_EQ(wo_wal_append_insert(&w2, &db3, 0, id), 0);
T_EQ(wo_wal_commit(&w2), 0);
wo_wal_close(&w2);
T_EQ(wo_wal_check(path, NULL), 6); /* tear gone, record in its place */
wo_db_destroy(&db3);
}
/* ---- the crash battery -------------------------------------------------- */
/* Child: insert forever — RAM, WAL, COMMIT, and only then ack the id down
* the pipe. Killed mid-stream by the parent. */
static void battery_child(const char *path, int ack_fd) {
wo_rt rt;
wo_db db;
wo_wal w;
if (wo_rt_init(&rt, 1 << 20, CLASSES, 1) != 0) _exit(9);
if (wo_db_init(&db, CLASSES, 1, 0, 1) != 0) _exit(9);
if (wo_wal_open(&w, path, 1 << 20) != 0) _exit(9);
const char *msg = "";
for (uint64_t i = 0;; i++) {
char label[32];
int n = snprintf(label, sizeof label, "row-%llu", (unsigned long long)i);
wo_str *s = wo_str_new(&rt, label, (uint32_t)n);
uint64_t vals[2] = {i * 3 + 1, (uint64_t)(uintptr_t)s};
uint64_t id = wo_row_insert(&db, 0, vals, &msg, NULL);
wo_str_free(&rt, s);
if (!id) _exit(9);
if (wo_wal_append_insert(&w, &db, 0, id) != 0) _exit(9);
if (wo_wal_commit(&w) != 0) _exit(9); /* durable BEFORE the ack */
ssize_t wr = write(ack_fd, &id, 8);
if (wr != 8) _exit(0); /* parent went away */
}
}
static void test_crash_battery(void) {
for (int round = 0; round < 5; round++) {
char path[128];
snprintf(path, sizeof path, "%s/crash-%d.wal", g_dir, round);
int pipefd[2];
T_EQ(pipe(pipefd), 0);
pid_t pid = fork();
T_CHECK(pid >= 0);
if (pid == 0) {
close(pipefd[0]);
battery_child(path, pipefd[1]);
_exit(0);
}
close(pipefd[1]);
/* collect acks for a few ms, then kill mid-stream — no sync with
the child's commit loop, which is the point */
struct timespec ts = {0, (20 + round * 13) * 1000000L};
while (nanosleep(&ts, &ts) != 0) {}
kill(pid, SIGKILL);
int status;
waitpid(pid, &status, 0);
/* drain every ack that made it into the pipe */
uint64_t acked[65536];
size_t n_acked = 0;
for (;;) {
uint64_t id;
ssize_t n = read(pipefd[0], &id, 8);
if (n != 8) break;
if (n_acked < 65536) acked[n_acked++] = id;
}
close(pipefd[0]);
T_CHECK(n_acked > 0); /* the child got at least one commit out */
/* offline oracle: the file's intact prefix covers every ack */
int64_t intact = wo_wal_check(path, NULL);
T_CHECK(intact >= (int64_t)n_acked);
/* replay and verify: every acked id present, contents exact */
wo_rt rt;
T_EQ(wo_rt_init(&rt, 1 << 22, CLASSES, 1), 0);
wo_db db;
T_EQ(wo_db_init(&db, CLASSES, 1, 0, 1), 0);
int64_t applied = wo_wal_replay(path, &db);
T_CHECK(applied >= (int64_t)n_acked);
const char *msg = "";
int bad = 0;
for (size_t i = 0; i < n_acked; i++) {
uint64_t out[2];
if (wo_row_read(&db, &rt, 0, acked[i], out, &msg) != 0) {
bad++;
continue;
}
/* id = i+1 (shard 0 of 1), field 0 = i*3+1, label = "row-i" */
char want[32];
int wl = snprintf(want, sizeof want, "row-%llu",
(unsigned long long)(acked[i] - 1));
wo_str *s = (wo_str *)(uintptr_t)out[1];
if (out[0] != (acked[i] - 1) * 3 + 1 || s->len != (uint32_t)wl ||
memcmp(s->data, want, (size_t)wl) != 0)
bad++;
wo_str_free(&rt, s);
}
T_EQ(bad, 0); /* zero acked-but-missing, zero acked-but-wrong */
wo_db_destroy(&db);
wo_rt_destroy(&rt);
}
}
int main(void) {
snprintf(g_dir, sizeof g_dir, "/tmp/wo-wal-test-XXXXXX");
if (!mkdtemp(g_dir)) return 1;
test_roundtrip_replay();
test_torn_tail();
test_crash_battery();
/* leave the dir for a failed run's forensics only */
if (!t_fail) {
char cmd[128];
snprintf(cmd, sizeof cmd, "rm -rf %s", g_dir);
if (system(cmd) != 0) fprintf(stderr, "cleanup failed, kept %s\n", g_dir);
} else {
fprintf(stderr, "kept %s\n", g_dir);
}
return t_report("test_wal");
}

View file

@ -65,6 +65,7 @@ uint32_t wb_class(wb_t *b, uint32_t name_const, uint32_t flags,
for (uint32_t j = 0; j < field_cnt; j++) put_u32(&b->classes, WOB_NONE); for (uint32_t j = 0; j < field_cnt; j++) put_u32(&b->classes, WOB_NONE);
for (uint32_t j = 0; j < field_cnt; j++) put_u32(&b->classes, WOB_NONE); for (uint32_t j = 0; j < field_cnt; j++) put_u32(&b->classes, WOB_NONE);
for (uint32_t j = 0; j < field_cnt; j++) put_u32(&b->classes, 0); for (uint32_t j = 0; j < field_cnt; j++) put_u32(&b->classes, 0);
put_u32(&b->classes, 0); /* v3: no secondary indexes in hand-built images */
return b->class_cnt++; return b->class_cnt++;
} }

106
scripts/employee-accept.sh Executable file
View file

@ -0,0 +1,106 @@
#!/usr/bin/env bash
# scripts/employee-accept.sh — the database track's acceptance workload.
#
# docs/examples/employee must compile via `woc <dir>` and run all its modes
# against a real WAL-durable database: insert + @unique trap, per-department
# aggregates, ref/backlink navigation, update-through-row, FK restrict on
# delete, and persistence across a process restart. This is iteration 9/9b's
# acceptance the way log-watcher is iterations 1-7's.
#
# Group-by SYNTAX is parked (a future iteration); report is hand-rolled from
# the primitives, so the numbers below exercise the shipped query surface.
set -uo pipefail
ROOT="$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)"
WOC="$ROOT/compiler/_build/default/bin/woc"
WOVM="$ROOT/runtime/wovm"
SAMPLE="$ROOT/docs/examples/employee"
pass=0
fail=0
ok() { echo "ok $1"; pass=$((pass + 1)); }
bad() { echo "FAIL $1 -- $2"; fail=$((fail + 1)); }
if [[ ! -x "$WOC" || ! -x "$WOVM" ]]; then
echo "employee-accept: build woc and wovm first (just woc-build; just wovm-build)" >&2
exit 1
fi
WORK="$(mktemp -d "${TMPDIR:-/tmp}/emp-accept.XXXXXX")"
DATA="$WORK/data"
mkdir -p "$DATA"
IMG="$WORK/employee.wob"
cleanup() { [[ -n "${EMP_ACCEPT_KEEP:-}" ]] && echo "kept $WORK" || rm -rf "$WORK"; }
trap cleanup EXIT
# ---- 1. compile ------------------------------------------------------
if "$WOC" --emit "$SAMPLE" -o "$IMG" >"$WORK/compile.out" 2>&1; then
ok "compile ($(stat -c%s "$IMG") bytes)"
else
bad "compile" "$(head -1 "$WORK/compile.out")"
echo; printf 'employee-accept: %d checks, %d failures\n' "$((pass + fail))" "$fail"; exit 1
fi
run() { WO_DATA="$DATA" "$WOVM" "$IMG" "$@"; }
# ---- 2. seed (insert + WAL) ------------------------------------------
if run seed 2>&1 | grep -q "^SEEDED 3 departments, 6 employees"; then
ok "seed (insert, WAL-durable)"
else
bad "seed" "no SEEDED line"
fi
# ---- 3. seed again: @unique trap, caught, across a process boundary --
out="$(run seed 2>&1)"; rc=$?
if [[ "$out" == *"SEED-DUP"* && $rc -eq 3 ]]; then
ok "unique violation caught on re-seed (persisted via replay)"
else
bad "unique re-seed" "got rc=$rc: $(printf '%s' "$out" | tr '\n' '|' | cut -c1-100)"
fi
# ---- 4. report: per-department aggregates + payroll ------------------
rep="$(run report 2>&1)"
if [[ "$rep" == *"DEPT Engineering headcount=3 avg=8200000 min=7300000 max=9200000"* \
&& "$rep" == *"DEPT Operations headcount=2 avg=6150000 min=5900000 max=6400000"* \
&& "$rep" == *"PAYROLL 45700000"* ]]; then
ok "report (aggregates + payroll)"
else
bad "report" "$(printf '%s' "$rep" | tr '\n' '|' | cut -c1-160)"
fi
# ---- 5. staff: index probe + backlink + ref navigation --------------
st="$(run staff Engineering 2>&1)"
if [[ "$st" == *"STAFF Asha 9200000 (Engineering)"* \
&& "$st" == *"STAFF Chidi 7300000 (Engineering)"* ]]; then
ok "staff (unique probe + backlink scan + ref nav)"
else
bad "staff" "$(printf '%s' "$st" | tr '\n' '|' | cut -c1-160)"
fi
# ---- 6. raise: update-through-row, reflected in a re-report ----------
run raise Operations 5 >/dev/null 2>&1
if run report 2>&1 | grep -q "^DEPT Operations headcount=2 avg=6457500"; then
ok "raise (update-through-row, durable)"
else
bad "raise" "operations average did not move to 6457500"
fi
# ---- 7. drop: FK restrict (Engineering still has staff) --------------
out="$(run drop Engineering 2>&1)"; rc=$?
if [[ "$out" == *"restricted"* && $rc -eq 4 ]]; then
ok "drop restricted by FK (department has staff)"
else
bad "drop restrict" "got rc=$rc: $(printf '%s' "$out" | tr '\n' '|' | cut -c1-100)"
fi
# ---- 8. persistence: a fresh process still sees every acked write ----
if run report 2>&1 | grep -q "^PAYROLL 46315000"; then
ok "persistence (replay: raised payroll survives restart)"
else
bad "persistence" "payroll after restart not 46315000"
fi
echo
printf 'employee-accept: %d checks, %d failures\n' "$((pass + fail))" "$fail"
[[ $fail -eq 0 ]]

View file

@ -0,0 +1,2 @@
eng restricted
ops deleted

View file

@ -0,0 +1,23 @@
-- the restrict trap is catchable; a free (unreferenced) row deletes.
@table(name: "dept", index: [name])
class Dept {
name: Text
staff: backlink Emp.dept
}
@table(name: "emp", index: [dept])
class Emp {
name: Text
dept: ref Dept
}
fn main() {
let eng = insert Dept { name: "eng" }
let ops = insert Dept { name: "ops" }
insert Emp { name: "asha", dept: eng }
let de = from x in Dept where x.name == "eng" take 1 select x
let r1 = try delete de[0] catch (e) nil
if r1 == nil { print("eng restricted") }
let dop = from x in Dept where x.name == "ops" take 1 select x
let r2 = try delete dop[0] catch (e) nil
if r2 == nil { print("ops restricted (WRONG)") }
print("ops deleted")
}

Some files were not shown because too many files have changed in this diff Show more