diff --git a/.gitignore b/.gitignore index ee9ec15..e5975fb 100644 --- a/.gitignore +++ b/.gitignore @@ -3,6 +3,7 @@ # `woc .` manifest builds (wo.toml [build] target) /docs/examples/log-watcher/target +/docs/examples/employee/target # Rust runtime (crates/rt/): compiled binary + build artifacts /crates/rt/target diff --git a/README.md b/README.md index 113ddcc..8954eb0 100644 --- a/README.md +++ b/README.md @@ -1,77 +1,323 @@ # writeonce -A declarative full-stack programming language. You write `.wo` files; the runtime compiles them into a binary that owns the database, serves REST, and pushes live subscriptions — no external database, no external web server, no frontend framework. +**A small compiled language with a database built in.** You write `.wo` +files; one command turns them into a single native binary that carries its +own storage engine — a typed, WAL-durable, crash-recoverable database — with +no server to install, no ORM, and no query strings. Tables are just classes, +queries are written in the language and checked by the compiler, and the whole +program ships as one file that depends only on the system C library. -Think **Go + Postgres + `net/http` + Phoenix LiveView, folded into one language and one binary.** +> **Status: early, honest.** Everything documented on this page compiles and +> runs today and is exercised by the acceptance tests in this repository. +> Features that are planned but **not yet available** are listed separately +> under [Roadmap](#roadmap) — they are not described as if they work. Nothing +> here is API-stable yet. -# persistant database +--- -- reads and writes database to RAM, persist data to postgres SQL. -- The entire database lives in RAM; every committed write is mirrored to PostgreSQL **as a backup** — asynchronously, behind the runtime's own WAL, never in the read or ack path. Set `WO_PG=postgres://user@host:5432/db` and every type's rows appear as a Postgres table (named by its `@table(name: ...)` annotation) that you can query with plain `psql`. Plan and phases: [`docs/plan/16-postgres-mirror.md`](docs/plan/16-postgres-mirror.md); try it: `just pricing-pg-demo`. +## Why writeonce -## Quickstart +- **The database is part of the language.** A `class` marked `@table` *is* a + table. Its rows persist through a write-ahead log, survive a restart, and are + reached by navigating typed relations — not by assembling SQL text. +- **Queries are compiled, not interpreted.** `from e in Employee where + e.salary > 90000 select e` lowers to bytecode loops over the engine. A + mistyped field name is a **compile error**, not a runtime surprise. There is + no SQL string anywhere in the shipped binary. +- **One binary, no runtime dependencies.** `woc .` produces a self-contained + executable (~100 KB for the sample programs) that links only libc. Copy it to + a server and run it. +- **Small on purpose.** No FFI, no package manager, no framework. The standard + library is a handful of OS modules. The language is designed to be read. + +writeonce is **not** a web framework and does not (yet) serve HTTP, WebSockets, +or a UI. It is a systems language whose distinguishing feature is the embedded +database. If you have seen an older "writeonce" that served REST from `cargo +run`, that was a separate, earlier runtime; this page documents the current +`woc`/`wovm` toolchain. + +--- + +## System requirements + +**To run a compiled writeonce program:** + +- Linux on x86-64. The produced binary is a native executable that links only + the system C library (`libc`); nothing else is required at runtime. + +**To build programs from source (the toolchain), you need:** + +| Tool | Version tested | Purpose | +| --- | --- | --- | +| OCaml | 4.14+ | builds `woc`, the compiler front end | +| dune | 3.14+ | OCaml build driver | +| A C11 compiler | gcc 13 / clang | builds `wovm`, the runtime VM | +| just | 1.x | task runner for the build/test recipes | +| make | any | drives the runtime build | + +Other POSIX platforms (macOS, BSD) are untested. The toolchain itself has no +network or package-download step — it builds entirely from the checked-in +source. + +--- + +## Getting the toolchain + +Two artifacts make up the toolchain: + +- **`woc`** — the compiler (OCaml). Reads `.wo` source, type-checks it, runs + the ownership pass, and emits a `.wob` image or a standalone binary. +- **`wovm`** — the runtime (C11). Loads a `.wob` image and executes it. When + `woc` builds a standalone binary, it embeds the image into a copy of `wovm`. + +Build both from the repository root: ```bash -git clone https://github.com/shoneyJ/writeonce -cd writeonce -cargo run --bin wo -- run docs/examples/blog # serve the sample blog on :8080 -curl http://127.0.0.1:8080/api/articles # it's a real REST API now +just woc-build # builds compiler/_build/default/bin/woc +just wovm-build # builds runtime/wovm + +# gate them (optional but recommended) +just woc-test # compiler unit + golden suites +just wovm-test # runtime unit suites, both dispatch flavors, ASan-clean ``` -See [`.dev/reference/rest/blog.rest`](.dev/reference/rest/blog.rest) for a preconfigured HTTP-request file that drives the whole sample — open it in VS Code (with the REST Client extension) or JetBrains and click "Send Request" on each block. +--- -## What this repository contains +## Your first program -| Path | What it is | -| ------------------------------------------------------------------------------------------ | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | -| [`crates/rt/`](crates/rt/) | The new `.wo` language runtime — lexer, type-DSL parser, in-memory engine, axum REST server. Produces the `wo` binary. | -| [`crates/{ql,value,engine,txn,db,wal,sub,http,gen,policy,logic,service,ui,app}/`](crates/) | 14 empty placeholder crates scaffolded for Phases 2–6. Real code extracts from `rt/` as each phase activates. | -| [`docs/runtime/wo-language.md`](docs/runtime/wo-language.md) | **Start here.** The language overview: toolchain, hello-world, stdlib, client model. | -| [`docs/runtime/database.md`](docs/runtime/database.md) | The 7-phase engineering series that drives the runtime's design. | -| [`docs/examples/blog/`](docs/examples/blog/) | Sample `.wo` project: blog with articles, authors, tags, comments. ~200 lines. | -| [`docs/examples/ecommerce/`](docs/examples/ecommerce/) | Sample `.wo` project: storefront + live order-ops table + cross-paradigm checkout. ~300 lines. | -| [`prototypes/wo-db/`](prototypes/wo-db/) | C++ prototype of the query-layer engine (SQL + Cypher + document paths, `RETURNING` aliases, `LIVE` stub). ~2k lines, smoke tests pass. Reference implementation the Rust port follows. | -| [`.dev/reference/rest/`](.dev/reference/rest/) | `.rest` files (VS Code REST Client / JetBrains HTTP format) for manually testing the running prototype. | -| [`.dev/reference/crates/`](.dev/reference/crates/) | The v1 writeonce blog — 13 Rust crates implementing the original `.seg` + sidecar-index storage engine and `.htmlx` templating. Preserved as a nested workspace; see [`.dev/reference/README.md`](.dev/reference/README.md). | +A writeonce project is a directory with a `wo.toml` manifest and one or more +`.wo` files. Every program has an entry point: -## Current stage +``` +-- hello/main.wo +fn main(args: multi Text) -> Int { + print("hello, writeonce"); + return 0; +} +``` -The runtime is under active development. Each stage lands as an independently shippable cut: +```toml +# hello/wo.toml +name = "hello" +version = "0.1.0" -| Stage | What works | Status | -| ------ | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------------------------------------------------------- | -| **1** | `wo run ` discovers every `.wo` file under a directory | ✅ shipped | -| **2** | Type-DSL parser, in-memory engine, REST CRUD (`list` / `get` / `create` / `update` / `delete`) generated from `service rest` blocks, JSON bodies with auto-id, default-value seeding, partial-update PATCH | ✅ shipped — `cargo run -- run docs/examples/blog` | -| **3** | LIVE subscriptions over WebSocket, delta frames on commit, `me` / session layer | pending | -| **4+** | Transactional fns (`fn checkout in txn snapshot`), row-level policies, type-attached triggers, `##ui` SSR, WAL durability, codegen | see [docs/runtime/database.md](docs/runtime/database.md) | +[runtime] +wo = ">= 0.1" +``` -`cargo test --lib` at the root runs 14 unit tests covering the lexer, parser, compiler, and engine. Stage-3 endpoints respond `501 Not Implemented` until they land. - -## Build & test +Compile the directory into a single binary and run it: ```bash -cargo build # builds all 15 crates (only `rt` has real code) -cargo test --lib # 14 unit tests - -cargo run --bin wo -- run docs/examples/blog # serve the blog sample -cargo run --bin wo -- run docs/examples/ecommerce # serve the ecommerce sample - -# Override the listen address -WO_LISTEN=127.0.0.1:9000 cargo run --bin wo -- run docs/examples/blog +woc hello/ # produces hello/target/hello +./hello/target/hello +# hello, writeonce ``` -## The v1 codebase (reference) +`main` returns an `Int` — that value is the process **exit code**. `args` is +the command-line arguments (the program name is not included). -The original writeonce blog engine — 13 crates, flat-file `.seg` storage, sidecar indexes, `.htmlx` templates, hand-rolled `epoll` event loop — moved to [`.dev/reference/crates/`](.dev/reference/crates/) when the new runtime was scaffolded. It's a nested Cargo workspace: +### The two build paths ```bash -cd .dev/reference/crates -cargo build # all 13 v1 crates still compile -cargo test # 12 unit tests, 1 ignored integration test +# 1. standalone binary (what you ship): woc reads wo.toml, emits target/ +woc myproject/ + +# 2. image + VM (handy while developing): emit a .wob, run it with wovm +woc --emit myproject/ -o app.wob +wovm app.wob arg1 arg2 ``` -V1 crates keep the `wo-` prefix (`wo-seg`, `wo-store`, …). The new runtime crates dropped it (`ql`, `value`, `engine`, …). [`docs/runtime/database/07-wo-seg-migration.md`](docs/runtime/database/07-wo-seg-migration.md) is the phased coexistence plan for replacing v1 with the new runtime — abstract behind a trait, dual-write, cut over, decommission. +Both paths run the same program. The standalone binary is the release artifact; +the image path lets you inspect or move the image around. -## License & status +--- -Work in progress. Nothing here is stable. Read the language overview in [`docs/runtime/wo-language.md`](docs/runtime/wo-language.md) if you want to know the shape; read the phase docs if you want to see the engineering plan; look in [`docs/examples/`](docs/examples/) if you want to see what the end product feels like. +## Language at a glance + +writeonce is statically typed with a compile-time ownership model — every value +has a known owner, memory is freed deterministically, and values that form +cycles are collected by an inferred garbage collector (you never annotate GC- +ness; the compiler infers it). The surface will look familiar: + +- **Types:** `Int`, `Text`, `Bool`, and user `class` types. `?T` marks an + optional (nullable) value; `nil` is the empty case. +- **Containers:** `multi T` (a growable list) and `map`. Literals: + `[]`, `[a, b]`, `{}`. +- **Classes & records:** classes with fields and methods, `static const` / + `static fn` members, module-scoped across files. +- **Control flow:** `if`/`else`, `for x in xs`, `for k, v in m`, `switch` + expressions, and `try { … } catch (e) { … }` (also an expression form). +- **Strings:** interpolation with `${expr}` inside a `"…"` literal. +- **Functions:** free functions and methods; arguments and returns are typed. + +``` +fn classify(n: Int) -> Text { + if n < 0 { return "negative"; } + return switch n { + case 0: "zero"; + default: "positive"; + }; +} +``` + +### Standard library + +A compact set of OS modules, reached by their reserved names — no imports: + +| Module | What it does | +| --- | --- | +| `fs` | `exists`, `list`, `stat`, `read_all`, `read_at`, `append` | +| `time` | `sleep`, `now`, `local`, `iso` | +| `env` | `get`, `stopping` (a cooperative shutdown flag) | +| `net` | TCP `listen` / `accept` / `read` / `write` / `close` (host + port) | +| `proc` | `run` a child process, capture stdout/stderr/exit | +| `json` | `encode` / `decode` (`json.decode(t) as T` yields `?T`) | + +These are deliberately minimal — the surface a real program needs, and no more. + +--- + +## The database + +This is the point of the language. Declaring storage is declaring a class: + +``` +@table(name: "departments", index: [name]) +class Department { + name: Text @unique + staff: backlink Employee.dept -- reverse relation, not a stored column +} + +@table(name: "employees", index: [dept], index: [dept, salary]) +class Employee { + name: Text + salary: Int + hired: Int + dept: ref Department -- foreign key: stored as the row id +} +``` + +- **`@table`** makes a class persistent — named storage plus declared secondary + indexes. Every instance you `insert` is written to a write-ahead log **before** + it is acknowledged, so an acked write survives a crash; on the next start the + log is replayed. +- **`ref T`** is a typed foreign key (a forward relation). **`backlink T.f`** is + its inverse — a virtual field, no stored column, resolved by an index scan. +- **`@unique`** enforces uniqueness at insert/update; a violation is a + **catchable** trap. +- **Foreign keys restrict deletes**: deleting a row that another row still + references traps rather than orphaning it. + +### Writing and reading data + +Mutation is direct; queries are a comprehension the compiler lowers to engine +operations: + +``` +-- insert (WAL-durable); @unique makes a re-insert trap, and try/catch it: +let eng = try insert Department { name: "Engineering" } catch (e) nil; +insert Employee { name: "Asha", salary: 9200000, hired: 1704067200000, dept: eng }; + +-- query: filter, order, limit, project — checked at compile time +for e in from s in Employee where s.salary > 8000000 order by s.salary desc select s { + print("${e.name} ${e.salary} (${e.dept.name})"); -- ref navigation +} + +-- navigate a backlink (the department's staff), update through the result +for e in from s in dept.staff select s { + e.salary = e.salary + e.salary * 5 / 100; -- update-through-row +} + +-- delete (restricted if still referenced) +let ok = try delete row catch (e) nil; +``` + +The query surface available today is **`from v in where … +[order by k [desc]] [take n] select v | v.field`**, plus `insert`, delete, and +update-through-a-row. It is proven end to end by the `employee` sample, whose +data survives a process restart via log replay. + +--- + +## Project layout & the manifest + +``` +myproject/ +├── wo.toml # manifest: name, version, [runtime], [build] +├── main.wo # entry point (fn main) +├── types.wo # your @table classes, other types +└── target/ # build output (the standalone binary lands here) +``` + +```toml +name = "myproject" +version = "0.1.0" + +[runtime] +wo = ">= 0.1" + +[build] +runtime = "../../../runtime/wovm" # path to the wovm the binary is built from +``` + +`woc myproject/` compiles every `.wo` file under the directory as one program. + +Programs that create tables read their data directory from the `WO_DATA` +environment variable at run time: + +```bash +WO_DATA=./data ./target/myproject seed +WO_DATA=./data ./target/myproject report # a fresh process still sees the data +``` + +--- + +## Worked examples + +Two complete sample programs live in the repository and double as the language's +acceptance tests: + +- **`docs/examples/employee/`** — departments and employees related by + `ref`/`backlink`, `@unique`, foreign-key restrict on delete, per-department + reports, and persistence across a restart. Run it: + + ```bash + just employee # compile + run every mode against a durable database + ``` + +- **`docs/examples/log-watcher/`** — a long-running daemon that watches log + files for silent death, using the `fs`/`time`/`net`/`proc` stdlib. Run it: + + ```bash + just log-watcher + ``` + +Read either program's `main.wo` for idiomatic, working writeonce. + +--- + +## Roadmap + +Planned, **not yet available** — listed so the shipped surface above stays +honest. These exist as design iterations and/or work-in-progress branches, not +as features you can use today: + +- **Query aggregates** — `group … by … into g` with `count`/`avg`/`min`/`max` + and projection records. (Today the same result is written by hand from the + shipped primitives.) +- **HTTP service layer** — `service` blocks that route requests to methods. +- **Concurrency** — a shard-actor runtime and green-threaded fibers. +- **Cross-program database access** — one program attaching to another's + database over a local channel, with keypair authentication and per-client + rights. +- **Blue-green deployment** — in-process recompile and atomic version switch. +- **Compile-time metaprogramming** — `@derive(Json/Csv/Eq/…)` generated from a + class's own metadata, no reflection. + +Known current limits worth naming: `net` is TCP host+port only; `proc.run` has +no timeout or signal control; there is no stdin/stdout byte I/O and no FFI. + +--- + +*writeonce is a work in progress. Interfaces will change. If you build +something with it, pin to a commit.* diff --git a/compiler/bin/main.ml b/compiler/bin/main.ml index a8f1276..a3240bb 100644 --- a/compiler/bin/main.ml +++ b/compiler/bin/main.ml @@ -600,10 +600,23 @@ let manifest_parse (path : string) : (string * string) list = if line = "" || (String.length line >= 1 && line.[0] = '#') then () else if line.[0] = '[' then begin if line.[String.length line - 1] <> ']' then fail !lineno "malformed section header"; - section := String.sub line 1 (String.length line - 2); - if !section <> "runtime" && !section <> "build" then - fail !lineno (Printf.sprintf "unknown section [%s] (runtime and build exist)" !section) + (* accept `[[table.array]]` headers too (iteration 9c's + [[share.clients]]) by trimming the doubled brackets *) + let inner = String.sub line 1 (String.length line - 2) in + let inner = + if String.length inner >= 2 && inner.[0] = '[' && inner.[String.length inner - 1] = ']' + then String.sub inner 1 (String.length inner - 2) + else inner + in + section := inner; + if !section <> "runtime" && !section <> "build" && !section <> "share" + && !section <> "share.clients" + then + fail !lineno + (Printf.sprintf "unknown section [%s] (runtime and build exist)" !section) end + else if !section = "share" || !section = "share.clients" then + () (* iteration 9c manifest keys — parsed by the attach feature, ignored here *) else match String.index_opt line '=' with | None -> fail !lineno "expected `key = \"value\"`" diff --git a/compiler/src/ast.ml b/compiler/src/ast.ml index d02879b..840f2f0 100644 --- a/compiler/src/ast.ml +++ b/compiler/src/ast.ml @@ -69,6 +69,9 @@ type field_ty = | Ref of string | Multi of string | Map of string * string (* key type, value type: map *) + | Backlink of string * string (* backlink C.f: the computed inverse of a + `ref` — NOT a stored column; reading it + scans C's index on f. Types as multi C. *) | Nullable of field_ty (* ?T wrapper *) (* Parameter passing convention (spec section 3, rule 2): default is an @@ -211,6 +214,15 @@ and expr_kind = | Binary of binop * expr * expr | Ctor of string * (string * expr) list | DbStub of Token.t list + (* `insert Class { field: expr, ... }` — the FIRST DB statement to leave + the stub behind (iteration 9, Task 3). Typed like a constructor + literal, returns the new row's id (Int), legal in statement and + expression position both. `select` stays a DbStub until Task 5. *) + | Insert of string * (string * expr) list + (* `delete ` (iteration 9b): removes the row a table-class value + names; an expression yielding the deleted id (restrict/trap surfaces + through the engine like any DB fault, catchable). *) + | Delete of expr (* haxe-parity Task 2: one `${expr}` interpolation site, produced only by the string-interpolation desugar (parser.ml) — never written directly by a parse rule the way every other expr_kind is. Its @@ -274,6 +286,31 @@ and expr_kind = ename : string; handler : stmt list; } + (* iteration 9b: a language-integrated query. `from in + where * [group by into ] [order by [desc]] [take ] + select ` — lowered to a bytecode loop over engine cursor builtins, + never SQL text. A table-class value is its row id at runtime, so field + access on a range variable reads through the engine. Slice scope today: + from/where/order/take/select and group-by aggregation; join is later. *) + | Query of query + +and query_source = + | QTable of string (* a table class by name: `from e in Employee` *) + | QNav of expr (* a backlink/multi navigation: `from s in d.staff` *) + +and query = { + q_var : string; + q_src : query_source; + q_wheres : expr list; + (* group by ... into : present iff this is an aggregating + query. q_group_key is the whole grouped element (`e`), q_group_by the + key, q_gvar the group binding whose `.f` columns feed aggregates. *) + q_group : (string * expr) option; (* (gvar, key_expr) *) + q_order : (expr * bool) option; (* (key, desc?) *) + q_take : expr option; + q_select : expr; + q_pos : pos; +} (* ---- statements (Task 5) --------------------------------------------- diff --git a/compiler/src/disasm.ml b/compiler/src/disasm.ml index a30ca0c..ecd0f86 100644 --- a/compiler/src/disasm.ml +++ b/compiler/src/disasm.ml @@ -164,7 +164,7 @@ let dump (img : string) : string = let line fmt = Buffer.add_string out (fmt ^ "\n") in if u32 img 0 <> magic then raise (Bad "bad magic"); let ver = u32 img 4 in - if ver <> 2 then raise (Bad (Printf.sprintf "unsupported version %d" ver)); + if ver <> 3 then raise (Bad (Printf.sprintf "unsupported version %d" ver)); let coff = u32 img 8 and ccnt = u32 img 12 in let koff = u32 img 16 and kcnt = u32 img 20 in let ioff = u32 img 24 and icnt = u32 img 28 in @@ -213,6 +213,14 @@ let dump (img : string) : string = renders as keys, so a wrong one is worth seeing. *) let names = List.init fcnt (fun j -> u32 img (!o + (j * 4))) in o := !o + (fcnt * 12); + (* v3 index tail: walk past (the disassembly prints class shape, not + indexes — dump goldens stay byte-stable across the version bump) *) + let icnt = u32 img !o in + o := !o + 4; + for _ = 1 to icnt do + let ccnt = u32 img (!o + 4) in + o := !o + 8 + (ccnt * 4) + done; let fields = List.map2 (fun nmk k -> if nmk = 0xFFFFFFFF then k else Printf.sprintf "%s:%s" (kname nmk) k) diff --git a/compiler/src/dump.ml b/compiler/src/dump.ml index b2530f2..3cfe237 100644 --- a/compiler/src/dump.ml +++ b/compiler/src/dump.ml @@ -150,6 +150,7 @@ let rec field_ty_str : Ast.field_ty -> string = function | Ast.Ref s -> Printf.sprintf "ref %s" s | Ast.Multi s -> Printf.sprintf "multi %s" s | Ast.Map (k, v) -> Printf.sprintf "map<%s, %s>" k v + | Ast.Backlink (c, f) -> Printf.sprintf "backlink %s.%s" c f | Ast.Nullable t -> "?" ^ field_ty_str t let param_str (p : Ast.param) : string = Printf.sprintf "%s%s: %s" (conv_str p.conv) p.name (field_ty_str p.ty) @@ -222,6 +223,19 @@ let rec expr_str (e : Ast.expr) : string = Printf.sprintf "%s { %s }" name (String.concat ", " (List.map (fun (fname, fval) -> Printf.sprintf "%s: %s" fname (expr_str fval)) fields)) + | Ast.Insert (name, fields) -> + Printf.sprintf "INSERT %s { %s }" name + (String.concat ", " + (List.map (fun (fname, fval) -> Printf.sprintf "%s: %s" fname (expr_str fval)) fields)) + | Ast.Query q -> + let src = match q.Ast.q_src with Ast.QTable cn -> cn | Ast.QNav e -> expr_str e in + Printf.sprintf "QUERY from %s in %s%s%s select %s" q.Ast.q_var src + (String.concat "" (List.map (fun w -> " where " ^ expr_str w) q.Ast.q_wheres)) + (match q.Ast.q_group with + | Some (g, k) -> Printf.sprintf " group by %s into %s" (expr_str k) g + | None -> "") + (expr_str q.Ast.q_select) + | Ast.Delete t -> Printf.sprintf "DELETE %s" (expr_str t) | Ast.DbStub toks -> Printf.sprintf "DB_STUB(%s)" (dbstub_tokens_str toks) | Ast.Interp inner -> Printf.sprintf "INTERP(%s)" (expr_str inner) | Ast.ListLit items -> Printf.sprintf "[%s]" (String.concat ", " (List.map expr_str items)) diff --git a/compiler/src/emit.ml b/compiler/src/emit.ml index 70255cc..5dcf173 100644 --- a/compiler/src/emit.ml +++ b/compiler/src/emit.ml @@ -152,7 +152,7 @@ let stdlib_not_linked_code = Diag.emitter_prefix ^ "06" ============================================================ *) let wob_magic = 0x31424F57 (* "WOB1" read as an LE u32 *) -let wob_version = 2 (* v2: per-field class-table metadata *) +let wob_version = 3 (* v3: v2 + per-class secondary-index metadata *) let wob_hdr_size = 44 let wob_none = 0xFFFFFFFF let k_int = 0 @@ -257,6 +257,7 @@ let b_map_val_at = 38 let b_multi_set = 39 let b_map_get_opt = 59 let b_text_copy = 60 +let b_db_insert = 61 (* json (runtime/src/json.c): encode takes the value's static kind as its second argument, decode the class id to build as its second. *) @@ -348,6 +349,16 @@ type clsrec = { cr_gc : bool; cr_fields : (string * Ast.field_ty) array; cr_methods : string list; (* method names, declaration order *) + (* iteration 9 Task 4: (unique, column indices) per secondary index — + `@table(index: [a, b])` entries (non-unique, composite) plus one + unique single-column entry per `@unique` field. Serialized as the v3 + class-record tail; the engine builds its runtime indexes from this. *) + cr_indexes : (bool * int array) list; + cr_is_table : bool; (* has @table — its instances are row ids (iteration 9b) *) + (* backlink fields (iteration 9b): name -> (source class, source field). + Virtual — not in cr_fields, no stored column; `d.staff` reads them by + probing the source class's index on the source field. *) + cr_backlinks : (string * (string * string)) list; } type ifacerec = { @@ -772,6 +783,44 @@ let field_kind (p : pctx) (ft : Ast.field_ty) : int = let class_of_name (p : pctx) (n : string) : int option = SM.find_opt n p.p_class_id +(* iteration 9b: a @table class's instances are row ids, so field access on + one reads through the engine (DB_GET_FIELD) rather than GETF. *) +let is_table_class (p : pctx) (cid : int) : bool = + cid >= 0 && cid < Array.length p.p_classes && p.p_classes.(cid).cr_is_table + +let b_str_lt = 67 +let b_db_update_field = 62 +let b_db_delete = 63 +let b_db_scan = 64 +let b_db_get_field = 65 +let b_db_probe = 66 + +(* iteration 9b: `d.staff` where staff is `backlink Employee.dept` reads by + probing Employee's index on its `dept` column. Resolve to (source cid, + index number) — None if the source field is not a declared index (a + backlink without a backing index has no efficient read and is rejected). *) +let backlink_target (p : pctx) (base_cid : int) (fname : string) : (int * int) option = + match List.assoc_opt fname p.p_classes.(base_cid).cr_backlinks with + | None -> None + | Some (src_class, src_field) -> ( + match class_of_name p src_class with + | None -> None + | Some scid -> + let sc = p.p_classes.(scid) in + (* stored column index of the source field *) + let col = ref (-1) in + Array.iteri (fun i (n, _) -> if n = src_field then col := i) sc.cr_fields; + if !col < 0 then None + else + (* the index whose single column is that field *) + let rec find n = function + | [] -> None + | (_, cols) :: tl -> + if Array.length cols = 1 && cols.(0) = !col then Some (scid, n) + else find (n + 1) tl + in + find 0 sc.cr_indexes) + let field_of (p : pctx) (cid : int) (fname : string) : (int * Ast.field_ty) option = let fs = p.p_classes.(cid).cr_fields in let rec go i = if i >= Array.length fs then None else @@ -941,6 +990,22 @@ let variant_tag_value (p : pctx) (u : Types.union_info) (vi : Types.variant_info | None -> 0 (* unreachable: pass 1 registers every payload-union variant *) else vi.Types.vi_tag +(* iteration 9b: a query's element type, as the name a `Multi` carries. + `select x` yields the source class (a row id typed as the class); + `select x.field` yields that field's type; anything else falls back to + Int (the slice's shapes are these two). *) +let query_elem_scalar (p : pctx) (q : Ast.query) ~(src : string) : string = + match q.Ast.q_select.Ast.kind with + | Ast.Ident v when v = q.Ast.q_var -> src (* select the whole row: element = source class *) + | Ast.Field ({ Ast.kind = Ast.Ident v; _ }, fname) when v = q.Ast.q_var -> ( + match class_of_name p src with + | Some cid -> ( + match field_of p cid fname with + | Some (_, ty) -> ( match unwrap ty with Scalar n -> n | _ -> "Int") + | None -> "Int") + | None -> "Int") + | _ -> "Int" + let rec ty_of_expr (p : pctx) (f : fstate) (e : Ast.expr) : Ast.field_ty option = match e.kind with | IntLit _ -> Some (Scalar "Int") @@ -970,10 +1035,17 @@ let rec ty_of_expr (p : pctx) (f : fstate) (e : Ast.expr) : Ast.field_ty option | Field (base, fname) -> ( match ty_of_expr p f base with | Some bt -> ( - match unwrap bt with + (* a `ref C` navigates into C: the target is a table row id *) + match (match unwrap bt with Ref c -> Scalar c | other -> other) with | Scalar cn -> ( match class_of_name p cn with - | Some cid -> ( match field_of p cid fname with Some (_, t) -> Some t | None -> None) + | Some cid -> ( + match field_of p cid fname with + | Some (_, t) -> Some t + | None -> ( + match List.assoc_opt fname p.p_classes.(cid).cr_backlinks with + | Some (sc, _) -> Some (Multi sc) + | None -> None)) | None -> None) | _ -> None) | None -> None) @@ -1054,6 +1126,16 @@ let rec ty_of_expr (p : pctx) (f : fstate) (e : Ast.expr) : Ast.field_ty option | Eq | Ne | Lt | Le | Gt | Ge | And | Or -> Some (Scalar "Bool") | Add | Sub | Mul | Div | Mod -> ( match ty_of_expr p f l with Some t -> Some t | None -> Some (Scalar "Int"))) | Ctor (cn, _) -> Some (Scalar cn) + | Insert _ -> Some (Scalar "Int") + | Delete _ -> Some (Scalar "Int") + | Query q -> + let src = + match q.Ast.q_src with + | Ast.QTable cn -> cn + | Ast.QNav nav -> ( + match ty_of_expr p f nav with Some t -> (match unwrap t with Multi c -> c | Scalar c -> c | _ -> "") | None -> "") + in + Some (Multi (query_elem_scalar p q ~src)) | Interp _ -> Some (Scalar "Text") | DbStub _ -> None | Switch (subject, arms) -> ( @@ -1344,7 +1426,8 @@ let field_class_meta (p : pctx) (ty : Ast.field_ty) : int = match name_of (Ast.Scalar e) with | Some n -> ( match class_of_name p n with Some cid -> cid | None -> wob_none) | None -> wob_none) - | Ast.Ref _ | Ast.Nullable _ -> wob_none + | Ast.Ref n -> ( match class_of_name p n with Some cid -> cid | None -> wob_none) + | Ast.Backlink _ | Ast.Nullable _ -> wob_none let field_elem_meta (p : pctx) (ty : Ast.field_ty) : int = match unwrap ty with @@ -1583,9 +1666,22 @@ let rec emit_expr (p : pctx) (f : fstate) (v : views) ~(dst : int) ?expected (e | Field (base, fname) -> ( match ty_of_expr p f base with | Some bt -> ( - match unwrap bt with + match (match unwrap bt with Ref c -> Scalar c | other -> other) with | Scalar cn -> ( match class_of_name p cn with + | Some cid when is_table_class p cid && backlink_target p cid fname <> None -> ( + (* `d.staff`: probe the source class's index for rows referencing + this row's id. Window: [class, index, key(=base id)]. *) + match backlink_target p cid fname with + | Some (scid, ino) -> + let b = emit_operand p f v base in + let w = alloc_temps p f e.pos 3 in + put f (ins_abx op_loadk w (check_bx p f e.pos "constant" (const_int p scid))); + put f (ins_abx op_loadk (w + 1) (check_bx p f e.pos "constant" (const_int p ino))); + put f (ins_abc op_move (w + 2) b 0); + sync_mask p f v e.id; + put f (ins_abc op_builtin dst w b_db_probe) + | None -> ()) | Some cid -> ( match field_of p cid fname with | Some (idx, _) -> @@ -1603,7 +1699,17 @@ let rec emit_expr (p : pctx) (f : fstate) (v : views) ~(dst : int) ?expected (e f.f_stmt_drops <- g :: f.f_stmt_drops; f.f_esc_drops <- g :: f.f_esc_drops end; - put f (ins_abc op_getf dst b (check_field_idx p f e.pos idx)) + if is_table_class p cid then begin + (* a table-class value is its row id; read the column from the + engine. Window: [class-id, id, field-idx]. *) + let w = alloc_temps p f e.pos 3 in + put f (ins_abx op_loadk w (check_bx p f e.pos "constant" (const_int p cid))); + put f (ins_abc op_move (w + 1) b 0); + put f (ins_abx op_loadk (w + 2) (check_bx p f e.pos "constant" (const_int p idx))); + sync_mask p f v e.id; + put f (ins_abc op_builtin dst w b_db_get_field) + end + else put f (ins_abc op_getf dst b (check_field_idx p f e.pos idx)) | None -> err p ~code:cannot_lower_code ~file:f.f_file ~pos:e.pos ~message:(Printf.sprintf "`%s` has no field `%s`" cn fname); @@ -1654,6 +1760,43 @@ let rec emit_expr (p : pctx) (f : fstate) (v : views) ~(dst : int) ?expected (e put f (ins_abc op_neg dst b 0) | Binary (op, l, r) -> emit_binary p f v ~dst op l r | Ctor (cn, fields) -> emit_ctor p f v ~dst e cn fields + | Insert (cn, fields) -> emit_insert p f v ~dst e cn fields + | Delete target -> ( + match ty_of_expr p f target with + | Some bt -> ( + match (match unwrap bt with Ref c -> Scalar c | o -> o) with + | Scalar cn -> ( + match class_of_name p cn with + | Some cid when is_table_class p cid -> + (* reserve dst past the window: in tail position dst == the first + window reg, and moving the id into dst would clobber the class + id — the disassembly-caught bug *) + let outer = f.f_temp in + if f.f_temp <= dst then f.f_temp <- dst + 1; + let w = alloc_temps p f e.pos 2 in + put f (ins_abx op_loadk w (check_bx p f e.pos "constant" (const_int p cid))); + let save = f.f_temp in + emit_expr p f v ~dst:(w + 1) target; + f.f_temp <- save; + (* keep the id so `delete x` can be used as an expression *) + put f (ins_abc op_move dst (w + 1) 0); + sync_mask p f v e.id; + f.f_cur_line <- e.pos.line; + put f (ins_abc op_builtin w w b_db_delete); + f.f_temp <- outer + | _ -> + err p ~code:cannot_lower_code ~file:f.f_file ~pos:e.pos + ~message:"`delete` target is not a table row"; + put f (ins_abx op_loadk dst (const_int p 0))) + | _ -> + err p ~code:cannot_lower_code ~file:f.f_file ~pos:e.pos + ~message:"`delete` target is not a table row"; + put f (ins_abx op_loadk dst (const_int p 0))) + | None -> + err p ~code:cannot_lower_code ~file:f.f_file ~pos:e.pos + ~message:"cannot resolve the `delete` target's type"; + put f (ins_abx op_loadk dst (const_int p 0))) + | Query q -> emit_query p f v ~dst e q | Interp inner -> ( (* haxe-parity Task 2: the type-directed half of the interpolation desugar (parser.ml's own doc comment on Ast.Interp) — a Text @@ -2347,6 +2490,324 @@ and emit_ctor (p : pctx) (f : fstate) (v : views) ~(dst : int) (e : Ast.expr) (c ci.Types.fields); f.f_temp <- outer +(* iteration 9 Task 3: `insert Class { ... }` lowers to one DB_INSERT + builtin whose window is [class-id const, then one slot per DECLARED + field in declaration order] — the executor walks the class table's + kinds, so slot order must be the table's, not the literal's. A field + the literal omits gets its default (same emit_default_value the ctor + uses) or, for a `?` field, its kind's own nil (WO_NIL_SCALAR for a + nullable scalar, the zero word otherwise). The engine COPIES every + value at the row API, so after the builtin every freshly built + argument is still this frame's to drop — same reap as push/set. *) +and emit_query (p : pctx) (f : fstate) (v : views) ~(dst : int) (e : Ast.expr) + (q : Ast.query) : unit = + (* iteration 9b slice: from/where/select over a table scan. group/order/ + take/navigation are diagnosed in types.ml, so a written image never + reaches this with them set. Lowered to an ordinary bytecode loop over + DB_SCAN's materialized id list — no plan tree, no text. *) + let cn = + match q.Ast.q_src with + | Ast.QTable cn -> cn + | Ast.QNav nav -> ( + (* the source's element type is the range var's class *) + match ty_of_expr p f nav with + | Some t -> ( match unwrap t with Multi c -> c | Scalar c -> c | _ -> "") + | None -> "") + in + match class_of_name p cn with + | None -> + err p ~code:cannot_lower_code ~file:f.f_file ~pos:e.pos + ~message:(Printf.sprintf "query over `%s`, which is not a declared table class" cn); + put f (ins_abx op_loadk dst (const_int p 0)) + | Some cid -> + let elem_name = query_elem_scalar p q ~src:cn in + let elem = Scalar elem_name in + (* a table-class element is a row ID (a scalar), not a heap pointer — so + the result container is SCALAR-kinded even though the element TYPES as + the class; getting this wrong drops an id as a pointer (ASan SEGV) *) + let elem_kind = + match class_of_name p elem_name with + | Some ecid when is_table_class p ecid -> 0 (* WO_K_SCALAR *) + | _ -> field_kind p elem + in + (* reserve dst past the loop's working registers (same guard emit_ctor + uses): dst holds the result multi every push writes into *) + let outer = f.f_temp in + if f.f_temp <= dst then f.f_temp <- dst + 1; + (* loop-carried registers, allocated once above dst, never reset *) + let scan = alloc_temp p f e.pos in + let idx = alloc_temp p f e.pos in + let len = alloc_temp p f e.pos in + let idreg = alloc_temp p f e.pos in + let body_base = f.f_temp in + (* scan -> multi of ids; result multi -> dst *) + sync_mask p f v e.id; + f.f_cur_line <- e.pos.line; + (match q.Ast.q_src with + | Ast.QTable _ -> + put f (ins_abx op_loadk scan (check_bx p f e.pos "constant" (const_int p cid))); + put f (ins_abc op_builtin scan scan b_db_scan) + | Ast.QNav nav -> + (* the navigation (a backlink) already yields a multi of source ids *) + let save = f.f_temp in + f.f_temp <- scan + 1; + emit_expr p f v ~dst:scan nav; + f.f_temp <- save); + put f (ins_abc op_builtin dst elem_kind b_multi_new); + put f (ins_abc op_builtin len scan b_len); + put f (ins_abx op_loadk idx (check_bx p f e.pos "constant" (const_int p 0))); + (* bind the range var to the current id (typed as the class), so field + access inside where/select routes through DB_GET_FIELD *) + let saved_env = f.f_env in + f.f_env <- (q.Ast.q_var, (idreg, Scalar cn)) :: f.f_env; + ignore body_base; + let top = here f in + f.f_temp <- body_base; + let tc = alloc_temp p f e.pos in + put f (ins_abc op_lt tc idx len); + let jz_exit = here f in + put f (ins_asbx op_jz tc 0); + (* id = multi_get(scan, idx) *) + let w = alloc_temps p f e.pos 2 in + put f (ins_abc op_move w scan 0); + put f (ins_abc op_move (w + 1) idx 0); + put f (ins_abc op_builtin idreg w b_multi_get); + (* where guards: any false skips the push *) + let skips = ref [] in + List.iter + (fun w_expr -> + let save = f.f_temp in + let wr = emit_operand p f v w_expr in + skips := here f :: !skips; + put f (ins_asbx op_jz wr 0); + f.f_temp <- save) + q.Ast.q_wheres; + (* select -> push into dst (copying a Text element the container owns) *) + let save = f.f_temp in + let sel = alloc_temp p f e.pos in + emit_expr p f v ~dst:sel q.Ast.q_select; + if elem_kind = 3 then put f (ins_abc op_builtin sel sel b_text_copy); + let pw = alloc_temps p f e.pos 2 in + put f (ins_abc op_move pw dst 0); + put f (ins_abc op_move (pw + 1) sel 0); + put f (ins_abc op_builtin pw pw b_multi_push); + f.f_temp <- save; + (* skip target: increment and loop *) + let cont = here f in + List.iter (fun pc -> patch_jump p f ~file:f.f_file ~pos:e.pos pc cont) !skips; + f.f_temp <- body_base; + let one = alloc_temp p f e.pos in + put f (ins_abx op_loadk one (check_bx p f e.pos "constant" (const_int p 1))); + put f (ins_abc op_add idx idx one); + let back = here f in + put f (ins_asbx op_jmp 0 0); + patch_jump p f ~file:f.f_file ~pos:e.pos back top; + let exit_pc = here f in + patch_jump p f ~file:f.f_file ~pos:e.pos jz_exit exit_pc; + f.f_env <- saved_env; + (* the scan's id list was this query's own, dropped now *) + put f (ins_abc op_drop scan 0 0); + (* ---- order by (whole-row selection sort) ---------------------------- + Elements of dst are row ids; the key re-reads a field through the + range var. Selection sort is O(n^2) but the result sets here are + small and this is KISS by design (no cost planner). Only the + whole-row + field-key shape is supported; grouped/projection ordering + lands with group-by. *) + (match q.Ast.q_order with + | Some (key, desc) -> + f.f_temp <- body_base; + let n = alloc_temp p f e.pos in + put f (ins_abc op_builtin n dst b_count); + let i = alloc_temp p f e.pos in + let j = alloc_temp p f e.pos in + let best = alloc_temp p f e.pos in + let elem_j = alloc_temp p f e.pos in + let elem_b = alloc_temp p f e.pos in + let sort_scratch = f.f_temp in + put f (ins_abx op_loadk i (check_bx p f e.pos "constant" (const_int p 0))); + let oi = here f in (* outer: while i < n *) + let oc = alloc_temp p f e.pos in + put f (ins_abc op_lt oc i n); + let ojz = here f in + put f (ins_asbx op_jz oc 0); + put f (ins_abc op_move best i 0); + let oneA = alloc_temp p f e.pos in + put f (ins_abx op_loadk oneA (check_bx p f e.pos "constant" (const_int p 1))); + put f (ins_abc op_add j i oneA); + let ij = here f in (* inner: while j < n *) + let ic = alloc_temp p f e.pos in + put f (ins_abc op_lt ic j n); + let ijz = here f in + put f (ins_asbx op_jz ic 0); + (* elem_j = multi_get(dst,j); elem_b = multi_get(dst,best) *) + let gw = alloc_temps p f e.pos 2 in + put f (ins_abc op_move gw dst 0); + put f (ins_abc op_move (gw + 1) j 0); + put f (ins_abc op_builtin elem_j gw b_multi_get); + put f (ins_abc op_move (gw + 1) best 0); + put f (ins_abc op_builtin elem_b gw b_multi_get); + (* keys: bind range var to elem_j / elem_b, eval key expr *) + let saved_env2 = f.f_env in + f.f_temp <- sort_scratch; + f.f_env <- (q.Ast.q_var, (elem_j, Scalar cn)) :: saved_env2; + (* key kind must be read with the range var BOUND — else ty_of_expr of + `x.name` sees x unbound, returns None, and a Text key silently falls + to the pointer-comparing op_lt (the wrong-order bug) *) + let key_is_text = + match ty_of_expr p f key with Some t -> field_kind p t = 3 | None -> false + in + let kj = alloc_temp p f e.pos in + emit_expr p f v ~dst:kj key; + f.f_env <- (q.Ast.q_var, (elem_b, Scalar cn)) :: saved_env2; + let kb = alloc_temp p f e.pos in + emit_expr p f v ~dst:kb key; + f.f_env <- saved_env2; + (* cmp: for asc, kj < kb -> best=j; for desc, kj > kb (== kb < kj). *) + let cmp = alloc_temp p f e.pos in + let lt a b = + if key_is_text then begin + let save = f.f_temp in + let w = alloc_temps p f e.pos 2 in + put f (ins_abc op_move w a 0); + put f (ins_abc op_move (w + 1) b 0); + put f (ins_abc op_builtin cmp w b_str_lt); + f.f_temp <- save + end + else put f (ins_abc op_lt cmp a b) + in + if desc then lt kb kj else lt kj kb; + let cjz = here f in + put f (ins_asbx op_jz cmp 0); + put f (ins_abc op_move best j 0); + let after = here f in + patch_jump p f ~file:f.f_file ~pos:e.pos cjz after; + f.f_temp <- sort_scratch; + let oneB = alloc_temp p f e.pos in + put f (ins_abx op_loadk oneB (check_bx p f e.pos "constant" (const_int p 1))); + put f (ins_abc op_add j j oneB); + let iback = here f in + put f (ins_asbx op_jmp 0 0); + patch_jump p f ~file:f.f_file ~pos:e.pos iback ij; + let iexit = here f in + patch_jump p f ~file:f.f_file ~pos:e.pos ijz iexit; + (* swap dst[i], dst[best]: read both, multi_set both *) + f.f_temp <- sort_scratch; + let vi = alloc_temp p f e.pos in + let vb = alloc_temp p f e.pos in + let sw = alloc_temps p f e.pos 3 in + put f (ins_abc op_move sw dst 0); + put f (ins_abc op_move (sw + 1) i 0); + put f (ins_abc op_builtin vi sw b_multi_get); + put f (ins_abc op_move (sw + 1) best 0); + put f (ins_abc op_builtin vb sw b_multi_get); + (* dst[i] = vb *) + put f (ins_abc op_move sw dst 0); + put f (ins_abc op_move (sw + 1) i 0); + put f (ins_abc op_move (sw + 2) vb 0); + put f (ins_abc op_builtin sw sw b_multi_set); + (* dst[best] = vi *) + put f (ins_abc op_move sw dst 0); + put f (ins_abc op_move (sw + 1) best 0); + put f (ins_abc op_move (sw + 2) vi 0); + put f (ins_abc op_builtin sw sw b_multi_set); + f.f_temp <- sort_scratch; + let oneC = alloc_temp p f e.pos in + put f (ins_abx op_loadk oneC (check_bx p f e.pos "constant" (const_int p 1))); + put f (ins_abc op_add i i oneC); + let oback = here f in + put f (ins_asbx op_jmp 0 0); + patch_jump p f ~file:f.f_file ~pos:e.pos oback oi; + let oexit = here f in + patch_jump p f ~file:f.f_file ~pos:e.pos ojz oexit + | None -> ()); + (* ---- take N: slice dst to [0, N) --------------------------------- *) + (match q.Ast.q_take with + | Some tk -> + f.f_temp <- body_base; + let nreg = alloc_temp p f e.pos in + emit_expr p f v ~dst:nreg tk; + (* clamp N to count(dst) so slice never runs past the end *) + let cnt = alloc_temp p f e.pos in + put f (ins_abc op_builtin cnt dst b_count); + let over = alloc_temp p f e.pos in + put f (ins_abc op_lt over cnt nreg); (* count < N ? use count *) + let jz2 = here f in + put f (ins_asbx op_jz over 0); + put f (ins_abc op_move nreg cnt 0); + let aft = here f in + patch_jump p f ~file:f.f_file ~pos:e.pos jz2 aft; + let sw = alloc_temps p f e.pos 3 in + let zero = alloc_temp p f e.pos in + put f (ins_abx op_loadk zero (check_bx p f e.pos "constant" (const_int p 0))); + put f (ins_abc op_move sw dst 0); + put f (ins_abc op_move (sw + 1) zero 0); + put f (ins_abc op_move (sw + 2) nreg 0); + let sliced = alloc_temp p f e.pos in + put f (ins_abc op_builtin sliced sw b_slice); + put f (ins_abc op_drop dst 0 0); (* the pre-slice multi is discarded *) + put f (ins_abc op_move dst sliced 0) + | None -> ()); + f.f_temp <- outer + +and emit_insert (p : pctx) (f : fstate) (v : views) ~(dst : int) (e : Ast.expr) (cn : string) + (fields : (string * Ast.expr) list) : unit = + match class_of_name p cn with + | None -> + err p ~code:cannot_lower_code ~file:f.f_file ~pos:e.pos + ~message:(Printf.sprintf "insert into `%s`, which is not a declared class" cn); + put f (ins_abx op_loadk dst (const_int p 0)) + | Some cid -> + let fcnt = Array.length p.p_classes.(cid).cr_fields in + let base = alloc_temps p f e.pos (fcnt + 1) in + put f (ins_abx op_loadk base (check_bx p f e.pos "constant" (const_int p cid))); + (* every field the literal names lands in ITS declared slot *) + List.iter + (fun ((fname : string), (fe : Ast.expr)) -> + match field_of p cid fname with + | None -> + err p ~code:cannot_lower_code ~file:f.f_file ~pos:e.pos + ~message:(Printf.sprintf "`%s` has no field `%s`" cn fname) + | Some (idx, fty) -> + let save = f.f_temp in + emit_expr p f v ~dst:(base + 1 + idx) ~expected:fty fe; + f.f_temp <- save) + fields; + (* omitted fields: declared default, else the kind's own nil *) + let provided = List.map fst fields in + (match Types.StringMap.find_opt cn p.p_syms.Types.classes with + | None -> () + | Some (ci : Types.class_info) -> + List.iter + (fun (fname, fty, fdefault, _) -> + if not (List.mem fname provided) then + match field_of p cid fname with + | None -> () + | Some (idx, dfty) -> ( + match fdefault with + | Some d -> + let save = f.f_temp in + emit_default_value p f ~dst:(base + 1 + idx) ~fty:dfty ~pos:e.pos d; + f.f_temp <- save + | None -> + let nil_word = + if is_nullable_scalar p fty then const_int p nil_scalar_word + else const_int p 0 + in + put f (ins_abx op_loadk (base + 1 + idx) (check_bx p f e.pos "constant" nil_word)))) + ci.Types.fields); + sync_mask p f v e.id; + f.f_cur_line <- e.pos.line; + put f (ins_abc op_builtin dst base b_db_insert); + (* the engine copied: fresh argument values die here *) + List.iter + (fun ((fname : string), (fe : Ast.expr)) -> + match field_of p cid fname with + | None -> () + | Some (idx, _) -> + drop_fresh_owned ~keep:dst p f (base + 1 + idx) fe; + drop_fresh_text ~keep:dst p f (base + 1 + idx) fe) + fields + (* The default expressions the emitter can lower (haxe-parity Task 4): the literal shapes the sample's own typedefs use — Int (optionally negated), Text, Bool, `now()` (parse_default_expr's own recognized @@ -3224,6 +3685,27 @@ and emit_assign (p : pctx) (f : fstate) (v : views) (s : Ast.stmt) (target : Ast | None -> err p ~code:cannot_lower_code ~file:f.f_file ~pos:target.pos ~message:(Printf.sprintf "assignment into `%s`, which is not a declared class" cn) + | Some cid when is_table_class p cid -> ( + (* iteration 9b: `e.salary = v` where e is a table row updates the + engine (DB_UPDATE_FIELD: class, id, field, value) — the row's + own indexes are maintained at the choke point *) + match field_of p cid fname with + | None -> + err p ~code:cannot_lower_code ~file:f.f_file ~pos:target.pos + ~message:(Printf.sprintf "`%s` has no field `%s`" cn fname) + | Some (idx, fty) -> + let w = alloc_temps p f target.pos 4 in + put f (ins_abx op_loadk w (check_bx p f target.pos "constant" (const_int p cid))); + let save = f.f_temp in + emit_expr p f v ~dst:(w + 1) base; + f.f_temp <- save; + put f (ins_abx op_loadk (w + 2) (check_bx p f target.pos "constant" (const_int p idx))); + let save = f.f_temp in + emit_expr p f v ~dst:(w + 3) ~expected:fty value; + f.f_temp <- save; + sync_mask p f v s.s_id; + f.f_cur_line <- s.s_pos.line; + put f (ins_abc op_builtin w w b_db_update_field)) | Some cid -> ( match field_of p cid fname with | None -> @@ -3852,6 +4334,18 @@ let emit ~(syms : Types.symbols) ~(module_of : string -> string) ~(module_syms : (string, Types.symbols) Hashtbl.t) (coll : Diag.Collector.t) (units : input list) : string = let colliding = compute_colliding_fn_names ~module_of units in + (* iteration 9 Task 4: index-declaration problems found while building + clsrecs — reported once a file/pos-bearing context exists below *) + let index_col_err : (Ast.pos * string) option ref = ref None in + let index_err_file = ref "" in + let ref_index_of_name (fnames : string list) (n : string) : int = + let rec go i = function + | [] -> 0 (* unknown column: the caller records the diagnostic *) + | x :: tl -> if x = n then i else go (i + 1) tl + in + go 0 fnames + in + let p_syms_for_indexes = syms in (* ---- pass 1: declarations, in discovery then declaration order ---- *) let classes = ref [] and class_id = ref SM.empty and nclasses = ref 0 in let ifaces = ref [] and iface_id = ref SM.empty and nifaces = ref 0 and nslots = ref 0 in @@ -3890,6 +4384,7 @@ let emit ~(syms : Types.symbols) ~(module_of : string -> string) (function | Ast.Class (c : Ast.class_decl) -> if not (SM.mem c.name !class_id) then begin + (if !index_col_err = None then index_err_file := u.file); let shape = if c.is_record then Some (record_shape_key c) else None in let alias_of = match shape with Some key -> Hashtbl.find_opt record_shape key | None -> None @@ -3907,10 +4402,77 @@ let emit ~(syms : Types.symbols) ~(module_of : string -> string) | Some key -> Hashtbl.replace record_shape key cid | None -> ()); classes := - { cr_name = c.name; cr_gc = c.is_gc; - cr_fields = - Array.of_list (List.map (fun (fl : Ast.field) -> (fl.name, fl.ty)) c.fields); - cr_methods = List.map (fun (m : Ast.method_decl) -> m.name) c.methods } + (let fnames = + List.filter_map + (fun (fl : Ast.field) -> + match fl.Ast.ty with Ast.Backlink _ -> None | _ -> Some fl.Ast.name) + c.fields + in + let col_of n = ref_index_of_name fnames n in + let is_indexable (fl : Ast.field) = + match Types.wob_kind_of_typ p_syms_for_indexes (Types.typ_of_field_ty (unwrap fl.Ast.ty)) with + | Types.WO_K_SCALAR | Types.WO_K_TEXT -> true + | _ -> false + in + let table_indexes = + match c.Ast.table with + | None -> [] + | Some cfg -> + List.map + (fun cols -> (false, Array.of_list (List.map col_of cols))) + cfg.Ast.indexes + in + let unique_indexes = + List.concat_map + (fun (fl : Ast.field) -> + if List.mem "unique" fl.Ast.annotations then begin + if not (is_indexable fl) then + index_col_err := Some (c.Ast.pos, Printf.sprintf + "`@unique` on `%s.%s`: only scalar and Text fields can be indexed" + c.Ast.name fl.Ast.name); + [ (true, [| col_of fl.Ast.name |]) ] + end + else []) + c.fields + in + (match c.Ast.table with + | Some cfg -> + List.iter + (fun cols -> + List.iter + (fun cn -> + match List.find_opt (fun (fl : Ast.field) -> fl.Ast.name = cn) c.fields with + | None -> + index_col_err := Some (c.Ast.pos, Printf.sprintf + "`@table(index: ...)` on `%s` names `%s`, which is not a field" + c.Ast.name cn) + | Some fl -> + if not (is_indexable fl) then + index_col_err := Some (c.Ast.pos, Printf.sprintf + "`@table(index: ...)` on `%s`: `%s` is not a scalar or Text field" + c.Ast.name cn)) + cols) + cfg.Ast.indexes + | None -> ()); + { cr_name = c.name; cr_gc = c.is_gc; + cr_fields = + Array.of_list + (List.filter_map + (fun (fl : Ast.field) -> + match fl.Ast.ty with + | Ast.Backlink _ -> None (* virtual: no stored column *) + | _ -> Some (fl.Ast.name, fl.Ast.ty)) + c.fields); + cr_methods = List.map (fun (m : Ast.method_decl) -> m.name) c.methods; + cr_indexes = table_indexes @ unique_indexes; + cr_is_table = (c.Ast.table <> None); + cr_backlinks = + List.filter_map + (fun (fl : Ast.field) -> + match fl.Ast.ty with + | Ast.Backlink (sc, sf) -> Some (fl.Ast.name, (sc, sf)) + | _ -> None) + c.fields }) :: !classes end | Ast.Union (ud : Ast.union_decl) -> @@ -3930,7 +4492,8 @@ let emit ~(syms : Types.symbols) ~(module_of : string -> string) class_id := SM.add key cid !class_id; incr nclasses; classes := - { cr_name = key; cr_gc = false; + { cr_name = key; cr_gc = false; cr_indexes = []; cr_is_table = false; + cr_backlinks = []; cr_fields = Array.of_list vd.Ast.v_fields; cr_methods = [] } :: !classes @@ -3973,11 +4536,18 @@ let emit ~(syms : Types.symbols) ~(module_of : string -> string) class_id := SM.add name cid !class_id; incr nclasses; classes := - { cr_name = name; cr_gc = false; cr_fields = Array.of_list fields; cr_methods = [] } + { cr_name = name; cr_gc = false; cr_fields = Array.of_list fields; cr_methods = []; + cr_indexes = []; cr_is_table = false; cr_backlinks = [] } :: !classes end) Types.predeclared_records; let class_id = !class_id in + (match !index_col_err with + | Some (pos, msg) -> + Diag.Collector.add coll + (Diag.error ~code:cannot_lower_code ~file:!index_err_file ~line:pos.Ast.line + ~col:pos.Ast.col ~message:msg ()) + | None -> ()); let p_classes = Array.of_list (List.rev !classes) in let p_ifaces = Array.of_list (List.rev !ifaces) in List.iter @@ -4149,7 +4719,16 @@ let emit ~(syms : Types.symbols) ~(module_of : string -> string) needs when decode creates one. *) Array.iter (fun kidx -> Buf.u32 cls kidx) class_field_names.(cid); Array.iter (fun (_, ty) -> Buf.u32 cls (field_class_meta p ty)) c.cr_fields; - Array.iter (fun (_, ty) -> Buf.u32 cls (field_elem_meta p ty)) c.cr_fields) + Array.iter (fun (_, ty) -> Buf.u32 cls (field_elem_meta p ty)) c.cr_fields; + (* v3 tail (iteration 9 Task 4): the class's secondary indexes — + index_cnt, then per index: flags (bit0 unique), col_cnt, cols *) + Buf.u32 cls (List.length c.cr_indexes); + List.iter + (fun (uniq, cols) -> + Buf.u32 cls (if uniq then 1 else 0); + Buf.u32 cls (Array.length cols); + Array.iter (fun ci -> Buf.u32 cls ci) cols) + c.cr_indexes) p_classes; let ifs = Buf.create () in Array.iteri diff --git a/compiler/src/owner.ml b/compiler/src/owner.ml index c268bf3..c61e1db 100644 --- a/compiler/src/owner.ml +++ b/compiler/src/owner.ml @@ -439,6 +439,7 @@ let oclass_of (ctx : ctx) (ft : Ast.field_ty) : oclass = | Some u -> if u.Types.u_has_payload then Owned else Copy | None -> Copy (* unknown type: WO-E225 already reported by types.ml *)) | Ref _ -> Copy + | Backlink _ -> Copy (* a virtual collection of row ids read on demand *) | Multi _ | Map _ -> Owned | Nullable _ -> Copy (* unreachable: unwrapped above *) @@ -560,6 +561,9 @@ let rec expr_ty (ctx : ctx) (e : Ast.expr) : Ast.field_ty option = | Binary (Concat, _, _) -> Some (Scalar "Text") | Binary _ -> None (* arithmetic/comparison: Copy either way *) | Ctor (cn, _) -> Some (Scalar cn) + | Insert _ -> Some (Scalar "Int") (* the new row's id — Copy, nothing to drop *) + | Query _ -> Some (Multi "Int") (* a query yields a fresh multi of ids — owned *) + | Delete _ -> Some (Scalar "Int") (* the deleted id — Copy *) | Interp _ -> Some (Scalar "Text") (* an interpolation always produces Text *) | DbStub _ -> None | Switch (subject, arms) -> @@ -1148,6 +1152,14 @@ let rec read_expr (ctx : ctx) (e : Ast.expr) : unit = read_place_parts ctx e | Call (callee, args) -> analyze_call ctx e callee args | Ctor (cn, fields) -> analyze_ctor ctx cn fields + | Insert (_, fields) -> + (* iteration 9 Task 3: the engine copies every field value at the row + API (the two-worlds bulkhead), so an insert BORROWS its values — + no transfer, no E304, the source keeps what it had. Trap-capable + (unique violations arrive with Task 4), so the drop map is + recorded exactly like DbStub's. *) + List.iter (fun (_, fe) -> read_expr ctx fe) fields; + record_drop ctx ~node:e.id ~pos:e.pos ~kind:DLiveMask ~items:(mask_items (live_holders ctx)) | Unary (_, o) -> read_expr ctx o | Binary (_, a, b) -> read_expr ctx a; @@ -1170,6 +1182,20 @@ let rec read_expr (ctx : ctx) (e : Ast.expr) : unit = | DbStub _ -> (* trap-capable: the frame needs its drop map here *) record_drop ctx ~node:e.id ~pos:e.pos ~kind:DLiveMask ~items:(mask_items (live_holders ctx)) + | Delete t -> + read_expr ctx t; + record_drop ctx ~node:e.id ~pos:e.pos ~kind:DLiveMask ~items:(mask_items (live_holders ctx)) + | Query q -> + (* iteration 9b: the sub-expressions only READ (engine field-reads copy + out at the boundary); the query is trap-capable (engine faults), so + the frame needs its drop map here, exactly like DbStub. *) + (match q.q_src with QNav e2 -> read_expr ctx e2 | QTable _ -> ()); + List.iter (read_expr ctx) q.q_wheres; + (match q.q_group with Some (_, k) -> read_expr ctx k | None -> ()); + (match q.q_order with Some (k, _) -> read_expr ctx k | None -> ()); + (match q.q_take with Some t -> read_expr ctx t | None -> ()); + read_expr ctx q.q_select; + record_drop ctx ~node:e.id ~pos:e.pos ~kind:DLiveMask ~items:(mask_items (live_holders ctx)) | Switch (subject, arms) -> analyze_switch ctx e.id subject arms (* The root of a place expression is already accounted for by use_place; diff --git a/compiler/src/parser.ml b/compiler/src/parser.ml index c4b1695..ed64b95 100644 --- a/compiler/src/parser.ml +++ b/compiler/src/parser.ml @@ -356,6 +356,12 @@ let parse_field_ty (st : state) : Ast.field_ty = | Token.Ident "multi" -> ignore (advance st); Ast.Multi (expect_ident st "multi target type") + | Token.Ident "backlink" -> + ignore (advance st); + let cls = expect_ident st "backlink source class" in + expect st Token.Dot "'.'"; + let fld = expect_ident st "backlink source field" in + Ast.Backlink (cls, fld) | Token.Ident "map" -> ignore (advance st); expect st Token.Lt "'<'"; @@ -1009,9 +1015,121 @@ and parse_switch_expr (st : state) : Ast.expr = done; { Ast.id; pos; kind = Ast.Switch (subject, List.rev !arms) } +and parse_insert_expr (st : state) : Ast.expr = + (* `insert` + a constructor literal, sharing parse_ctor_literal so the + field-list grammar (trailing commas, newlines) can never drift from the + ctor's. The literal's node is unwrapped into Insert — its id is reused, + which is safe because the Ctor node itself is discarded whole. *) + let pos = peek_pos st in + ignore (advance st) (* the `insert` trigger token *); + skip_newlines st; + let lit = parse_ctor_literal st in + (match lit.Ast.kind with + | Ast.Ctor (cn, fields) -> { lit with Ast.pos; kind = Ast.Insert (cn, fields) } + | _ -> lit (* unreachable: parse_ctor_literal only builds Ctor *)) + +and is_query_trigger (st : state) : bool = + (* `from in` — positional, so `from` stays a usable identifier + everywhere else (same discipline as insert/select) *) + (match peek st with Token.Ident "from" -> true | _ -> false) + && (match (tok_at st (st.pos + 1)).kind with Token.Ident _ -> true | _ -> false) + && (tok_at st (st.pos + 2)).kind = Token.KwIn + +and parse_query_expr (st : state) : Ast.expr = + let pos = peek_pos st in + let id = fresh_id st in + ignore (advance st) (* from *); + let var = expect_ident st "query range variable" in + expect st Token.KwIn "`in`"; + (* source: a bare class name is a table scan; any other expression is a + navigation (`d.staff`). One token of lookahead: Ident not followed by a + `.`/`(`/`[` and sitting where a clause keyword follows is a table name. *) + let src = + match peek st with + | Token.Ident cn + when (match (tok_at st (st.pos + 1)).kind with + | Token.Dot | Token.LParen | Token.LBracket -> false + | _ -> true) -> + ignore (advance st); + Ast.QTable cn + | _ -> Ast.QNav (parse_expr_no_brace st) + in + (* clauses may sit on their own lines; skip the separating newlines when + looking for the next clause keyword (the query is one expression) *) + let clause name = + skip_newlines st; + match peek st with Token.Ident n when n = name -> true | _ -> false + in + let wheres = ref [] in + while clause "where" do + ignore (advance st); + wheres := parse_expr_no_brace st :: !wheres + done; + let group = + if clause "group" then begin + ignore (advance st); + let key_elem = parse_expr_no_brace st in + ignore key_elem (* the grouped element is the range var; `group e by k` *); + if not (clause "by") then fail st (peek_pos st) syntax_code "expected `by` in a group clause"; + ignore (advance st); + let key = parse_expr_no_brace st in + if not (clause "into") then fail st (peek_pos st) syntax_code "expected `into` in a group clause"; + ignore (advance st); + let gvar = expect_ident st "group variable" in + Some (gvar, key) + end + else None + in + let order = + if clause "order" then begin + ignore (advance st); + if not (clause "by") then fail st (peek_pos st) syntax_code "expected `by` after `order`"; + ignore (advance st); + let key = parse_expr_no_brace st in + let desc = clause "desc" in + if desc then ignore (advance st); + Some (key, desc) + end + else None + in + (* `take` is a reserved keyword (KwTake, the param convention), not an + Ident — so match the token, not the name *) + skip_newlines st; + let take = + if peek st = Token.KwTake then (ignore (advance st); Some (parse_expr_no_brace st)) else None + in + if not (clause "select") then fail st (peek_pos st) syntax_code "a query must end in `select`"; + ignore (advance st); + let sel = parse_expr st in + { + Ast.id; + pos; + kind = + Ast.Query + { + Ast.q_var = var; + q_src = src; + q_wheres = List.rev !wheres; + q_group = group; + q_order = order; + q_take = take; + q_select = sel; + q_pos = pos; + }; + } + and parse_primary (st : state) : Ast.expr = match peek st with + | _ when is_query_trigger st -> parse_query_expr st + | Token.Ident "delete" when (match (tok_at st (st.pos + 1)).kind with + | Token.Newline | Token.Semicolon | Token.Eof -> false | _ -> true) -> + let pos = peek_pos st in + let id = fresh_id st in + ignore (advance st); + let target = parse_expr st in + { Ast.id; pos; kind = Ast.Delete target } | k when is_select_trigger k -> parse_dbstub_expr st + | k when is_insert_trigger k -> parse_insert_expr st | Token.KwSwitch -> parse_switch_expr st | Token.Int n -> let pos = peek_pos st in @@ -1286,7 +1404,7 @@ and parse_stmt (st : state) : Ast.stmt = | k when is_insert_trigger k -> let pos = peek_pos st in let id = fresh_id st in - let e = parse_dbstub_expr st in + let e = parse_insert_expr st in end_of_stmt st; { Ast.s_id = id; s_pos = pos; s_kind = Ast.ExprStmt e } | Token.KwLet -> parse_let_stmt st @@ -1785,6 +1903,23 @@ let rec subst_expr (consts : Ast.expr StringMap.t) (bound : StringSet.t) (e : As { e with Ast.kind = Ast.Binary (op, subst_expr consts bound l, subst_expr consts bound r) } | Ast.Ctor (cn, fields) -> { e with Ast.kind = Ast.Ctor (cn, List.map (fun (n, v) -> (n, subst_expr consts bound v)) fields) } + | Ast.Insert (cn, fields) -> + { e with Ast.kind = Ast.Insert (cn, List.map (fun (n, v) -> (n, subst_expr consts bound v)) fields) } + | Ast.Delete t -> { e with Ast.kind = Ast.Delete (subst_expr consts bound t) } + | Ast.Query q -> + (* the range/group vars shadow consts inside the query body *) + let bound' = StringSet.add q.Ast.q_var bound in + let bound' = match q.Ast.q_group with Some (g, _) -> StringSet.add g bound' | None -> bound' in + let sub = subst_expr consts bound' in + { e with Ast.kind = Ast.Query { + q with Ast.q_src = (match q.Ast.q_src with + | Ast.QTable cn -> Ast.QTable cn + | Ast.QNav e2 -> Ast.QNav (subst_expr consts bound e2)); + q_wheres = List.map sub q.Ast.q_wheres; + q_group = (match q.Ast.q_group with Some (g, k) -> Some (g, sub k) | None -> None); + q_order = (match q.Ast.q_order with Some (k, d) -> Some (sub k, d) | None -> None); + q_take = (match q.Ast.q_take with Some t -> Some (sub t) | None -> None); + q_select = sub q.Ast.q_select } } | Ast.Interp inner -> { e with Ast.kind = Ast.Interp (subst_expr consts bound inner) } | Ast.ListLit items -> { e with Ast.kind = Ast.ListLit (List.map (subst_expr consts bound) items) } | Ast.MapLit | Ast.NilLit -> e diff --git a/compiler/src/types.ml b/compiler/src/types.ml index 3a5395f..4b904ec 100644 --- a/compiler/src/types.ml +++ b/compiler/src/types.ml @@ -306,6 +306,7 @@ let rec has_recursive_structure (cls : class_info) : bool = | Ast.Ref name -> name = cls.name | Ast.Multi name -> name = cls.name (* multi Self *) | Ast.Map (k, v) -> k = cls.name || v = cls.name (* map<_, Self> / map *) + | Ast.Backlink _ -> false | Ast.Nullable inner -> has_recursive_structure_type inner cls.name ) cls.fields @@ -315,6 +316,7 @@ and has_recursive_structure_type (ty : Ast.field_ty) (cls_name : string) : bool | Ast.Ref name -> name = cls_name | Ast.Multi name -> name = cls_name | Ast.Map (k, v) -> k = cls_name || v = cls_name + | Ast.Backlink _ -> false (* a computed inverse holds no owned structure *) | Ast.Nullable inner -> has_recursive_structure_type inner cls_name (* @unique field -> persistent identity (plan's "When NOT to emit": a @@ -363,6 +365,7 @@ let rec typ_of_field_ty (ft : field_ty) : typ = | Ref name -> TRef name | Multi inner_name -> TMulti (TScalar inner_name) | Map (k_name, v_name) -> TMap (TScalar k_name, TScalar v_name) + | Backlink (c, _) -> TMulti (TScalar c) (* reads as a collection of C *) | Nullable inner -> TNullable (typ_of_field_ty inner) (* wob_kind_of_typ: maps internal typ to .wob field kind *) @@ -412,6 +415,7 @@ let unknown_fn_code = Diag.types_prefix ^ "04" let unsatisfied_interface_code = Diag.types_prefix ^ "05" let incomplete_ctor_code = Diag.types_prefix ^ "06" let unknown_type_code = Diag.types_prefix ^ "07" +let query_code = Diag.types_prefix ^ "50" (* WO-E250: query surface (iteration 9b) *) let non_exhaustive_switch_code = Diag.types_prefix ^ "08" let invalid_builtin_code = Diag.types_prefix ^ "09" let module_not_imported_code = Diag.types_prefix ^ "10" @@ -632,7 +636,7 @@ let rec scalar_name_of (ft : field_ty) : string option = match ft with | Scalar name -> Some name | Nullable inner -> scalar_name_of inner - | Ref _ | Multi _ | Map _ -> None + | Ref _ | Multi _ | Map _ | Backlink _ -> None (* Checked once per field declaration (not at every access/use site), so the diagnostic lands at the field's own declaration position and @@ -1165,6 +1169,11 @@ let typecheck_program ~file ~(module_of : string -> string) needs `int_to_text` first) -- unlike the placeholders below, this is a fact, not a guess. *) Some (TScalar "Text") + | Insert _ -> + (* the new row's id — the one thing an insert produces *) + Some (TScalar "Int") + | Query _ -> None (* a query's type is chased only by typecheck_expr *) + | Delete _ -> Some (TScalar "Int") | Unary _ | Binary _ | DbStub _ -> (* Not chased: the arithmetic-ladder `Binary` ops have no reliable per-node type in this pass at all (see above); `Unary`/`DbStub` @@ -1198,7 +1207,7 @@ let typecheck_program ~file ~(module_of : string -> string) with Not_found -> { typ = TScalar "Int"; is_nil = false }) | Field (base, field_name) -> let base_res = typecheck_expr env cenv base in - (match base_res.typ with + (match (match base_res.typ with TRef c -> TScalar c | other -> other) with | TScalar class_name -> (* Only a *declared* class can be checked for a missing field. typecheck_expr falls back to `TScalar "Int"` for everything @@ -1222,9 +1231,14 @@ let typecheck_program ~file ~(module_of : string -> string) { typ = TScalar "Int"; is_nil = false })) | _ -> { typ = TScalar "Int"; is_nil = false }) | Index (base, idx) -> - let _ = typecheck_expr env cenv base in + let base_res = typecheck_expr env cenv base in let _ = typecheck_expr env cenv idx in - { typ = TScalar "Int"; is_nil = false } + (* `xs[i]` yields the container's element type — a `multi C` indexed + is a C (iteration 9b: query results are indexed to pick a row) *) + (match base_res.typ with + | TMulti et -> { typ = et; is_nil = false } + | TMap (_, vt) -> { typ = vt; is_nil = false } + | _ -> { typ = TScalar "Int"; is_nil = false }) | Call (callee, args) -> List.iter (fun arg -> ignore (typecheck_expr env cenv arg)) args; (match callee.kind with @@ -1391,7 +1405,8 @@ let typecheck_program ~file ~(module_of : string -> string) the zero word NEW already leaves there). Everything else stays WO-E206, classes and records alike. *) let omittable (default : default_expr option) (fty : field_ty) : bool = - Option.is_some default || (match fty with Nullable _ -> true | _ -> false) + Option.is_some default + || (match fty with Nullable _ | Backlink _ -> true | _ -> false) in List.iter (fun (fname, fty, fdefault, _) -> if not (List.mem fname provided) && not (omittable fdefault fty) then @@ -1405,6 +1420,100 @@ let typecheck_program ~file ~(module_of : string -> string) (Diag.error ~code:unknown_type_code ~file ~line:e.pos.line ~col:e.pos.col ~message:(Printf.sprintf "unknown type `%s` in constructor" class_name) ()); { typ = TScalar "Int"; is_nil = false }) + | Insert (class_name, fields) -> + (* iteration 9 Task 3: typed exactly like a constructor literal — + same missing-field rule (defaults and `?` fields omittable), + same unknown-class diagnostic — but the VALUE is the new row's + id. The engine copies every field at the choke point, so field + values keep their owners (owner.ml's stores_by_copy). *) + (try + let cls = StringMap.find class_name syms.classes in + let provided = List.map (fun (n, _) -> n) fields in + let omittable (default : default_expr option) (fty : field_ty) : bool = + Option.is_some default + || (match fty with Nullable _ | Backlink _ -> true | _ -> false) + in + List.iter (fun (fname, fty, fdefault, _) -> + if not (List.mem fname provided) && not (omittable fdefault fty) then + Diag.Collector.add collector + (Diag.error ~code:incomplete_ctor_code ~file ~line:e.pos.line ~col:e.pos.col + ~message:(Printf.sprintf "missing field `%s` in insert of `%s`" fname class_name) ()) + ) cls.fields; + { typ = TScalar "Int"; is_nil = false } + with Not_found -> + Diag.Collector.add collector + (Diag.error ~code:unknown_type_code ~file ~line:e.pos.line ~col:e.pos.col + ~message:(Printf.sprintf "unknown type `%s` in insert" class_name) ()); + { typ = TScalar "Int"; is_nil = false }) + | Delete target -> + let tr = typecheck_expr env cenv target in + (match (match tr.typ with TRef c -> TScalar c | o -> o) with + | TScalar cn when StringMap.mem cn syms.classes -> () + | _ -> + Diag.Collector.add collector + (Diag.error ~code:query_code ~file ~line:e.pos.line ~col:e.pos.col + ~message:"`delete` takes a table-row value" ())); + { typ = TScalar "Int"; is_nil = false } + | Query q -> + (* iteration 9b slice: from/where/select over a table class. The + range variable is bound to the class type; a table-class value is + its row id at runtime but types AS the class, so `e.field` checks + against the class's fields exactly like a heap instance. group / + order / take / navigation sources are diagnosed as not-yet so the + surface is honest about its edge. *) + let elem_err () = + { typ = TMulti (TScalar "Int"); is_nil = false } + in + (match q.q_src with + | Ast.QNav nav -> + (* `from s in d.staff`: the navigation yields `multi C`, so the + range var is a C. Reuse the QTable body by resolving C. *) + let nav_res = typecheck_expr env cenv nav in + let cn = + match nav_res.typ with + | TMulti (TScalar c) -> c + | _ -> "" + in + if not (StringMap.mem cn syms.classes) then begin + Diag.Collector.add collector + (Diag.error ~code:query_code ~file ~line:q.q_pos.line ~col:q.q_pos.col + ~message:"query navigation source must be a `backlink`/`multi` of a table class" ()); + elem_err () + end + else begin + (if q.q_group <> None then + Diag.Collector.add collector + (Diag.error ~code:query_code ~file ~line:q.q_pos.line ~col:q.q_pos.col + ~message:"group-by on a navigation query is not supported yet" ())); + let env' = StringMap.add q.q_var (TScalar cn) env in + let cenv' = StringMap.add q.q_var (TScalar cn) cenv in + List.iter (fun w -> ignore (typecheck_expr env' cenv' w)) q.q_wheres; + (match q.q_order with Some (k, _) -> ignore (typecheck_expr env' cenv' k) | None -> ()); + (match q.q_take with Some t -> ignore (typecheck_expr env cenv t) | None -> ()); + let sel = typecheck_expr env' cenv' q.q_select in + { typ = TMulti sel.typ; is_nil = false } + end + | Ast.QTable cn -> + if not (StringMap.mem cn syms.classes) then begin + Diag.Collector.add collector + (Diag.error ~code:query_code ~file ~line:q.q_pos.line ~col:q.q_pos.col + ~message:(Printf.sprintf "`from %s in %s`: `%s` is not a declared table class" + q.q_var cn cn) ()); + elem_err () + end + else begin + (if q.q_group <> None then + Diag.Collector.add collector + (Diag.error ~code:query_code ~file ~line:q.q_pos.line ~col:q.q_pos.col + ~message:"group-by aggregation is not supported yet" ())); + let env' = StringMap.add q.q_var (TScalar cn) env in + let cenv' = StringMap.add q.q_var (TScalar cn) cenv in + List.iter (fun w -> ignore (typecheck_expr env' cenv' w)) q.q_wheres; + (match q.q_order with Some (k, _) -> ignore (typecheck_expr env' cenv' k) | None -> ()); + (match q.q_take with Some t -> ignore (typecheck_expr env cenv t) | None -> ()); + let sel = typecheck_expr env' cenv' q.q_select in + { typ = TMulti sel.typ; is_nil = false } + end) | DbStub _ -> { typ = TVoid; is_nil = false } | Switch (subject, arms) -> typecheck_switch ~want_value:true env cenv subject arms | ListLit items -> @@ -2102,7 +2211,8 @@ and walk_expr (bound : StringSet.t) (visit : StringSet.t -> expr -> unit) (e : e | Binary (_, l, r) -> walk_expr bound visit l; walk_expr bound visit r - | Ctor (_, fields) -> List.iter (fun (_, v) -> walk_expr bound visit v) fields + | Ctor (_, fields) | Insert (_, fields) -> + List.iter (fun (_, v) -> walk_expr bound visit v) fields | Interp inner -> walk_expr bound visit inner | ListLit items -> List.iter (walk_expr bound visit) items | MapLit | NilLit -> () @@ -2111,6 +2221,16 @@ and walk_expr (bound : StringSet.t) (visit : StringSet.t -> expr -> unit) (e : e walk_expr bound visit body; walk_block (StringSet.add ename bound) visit handler | DbStub _ -> () + | Delete t -> walk_expr bound visit t + | Query q -> + (match q.q_src with QNav e -> walk_expr bound visit e | QTable _ -> ()); + let b = StringSet.add q.q_var bound in + let b = match q.q_group with Some (g, _) -> StringSet.add g b | None -> b in + List.iter (walk_expr b visit) q.q_wheres; + (match q.q_group with Some (_, k) -> walk_expr b visit k | None -> ()); + (match q.q_order with Some (k, _) -> walk_expr b visit k | None -> ()); + (match q.q_take with Some t -> walk_expr b visit t | None -> ()); + walk_expr b visit q.q_select | Switch (subject, arms) -> walk_expr bound visit subject; List.iter @@ -2370,6 +2490,7 @@ let rec field_ty_str (ft : field_ty) : string = | Ref s -> "ref " ^ s | Multi s -> "multi " ^ s | Map (k, v) -> "map<" ^ k ^ ", " ^ v ^ ">" + | Backlink (c, f) -> "backlink " ^ c ^ "." ^ f | Nullable t -> "?" ^ field_ty_str t let dump_symbols (syms : symbols) : string = diff --git a/compiler/test/golden/ast/db-stub.expected b/compiler/test/golden/ast/db-stub.expected index 17db105..3494df8 100644 --- a/compiler/test/golden/ast/db-stub.expected +++ b/compiler/test/golden/ast/db-stub.expected @@ -1,6 +1,6 @@ 1:1 METHOD sync() - 2:3 DB_STUB IDENT(insert) IDENT(Product) LBRACE IDENT(sku) COLON STR(A1) COMMA IDENT(price) COLON INT(10) RBRACE + 2:3 EXPR INSERT Product { sku: "A1", price: 10 } 3:3 DB_STUB IDENT(select) IDENT(Product) LBRACE IDENT(sku) EQEQ STR(A1) RBRACE 4:3 LET rows = DB_STUB(IDENT(select) IDENT(Product) LBRACE IDENT(price) GT INT(5) RBRACE) - 5:3 DB_STUB KW_INSERT IDENT(Product) LBRACE IDENT(sku) COLON STR(A2) RBRACE + 5:3 EXPR INSERT Product { sku: "A2" } 6:3 DB_STUB KW_SELECT IDENT(Product) LBRACE IDENT(sku) EQEQ STR(A2) RBRACE diff --git a/compiler/test/golden/owner/drops.wo b/compiler/test/golden/owner/drops.wo index f50111f..e487997 100644 --- a/compiler/test/golden/owner/drops.wo +++ b/compiler/test/golden/owner/drops.wo @@ -38,7 +38,7 @@ fn pick(take a: Item, take b: Item, flag: Bool) -> Int { } fn store(take r: Item) -> Int { - insert into rows values (1) + insert Row { n: 1 } return 0 } @@ -63,3 +63,7 @@ fn reinit_after_move(take a: Item) -> Int { a = Item { n: 7 } return 0 } + +class Row { + n: Int +} diff --git a/compiler/test/runner.ml b/compiler/test/runner.ml index 28635c7..8d644a5 100644 --- a/compiler/test/runner.ml +++ b/compiler/test/runner.ml @@ -531,35 +531,35 @@ let () = | _ -> check "ctor literal: exactly one free fn" false let () = - (* The brief's stated asymmetry: `insert` is a statement-only trigger - (parser.ml's is_insert_trigger, checked only in parse_stmt) — a - bare `insert` reached from parse_primary is just an ordinary - identifier reference, exactly like self/me/on/service/policy's own - "recognized positionally, not a reserved word" rule (this task's - own keyword-discipline note). `select` (is_select_trigger) is - checked unconditionally *inside* parse_primary, so the same - position always builds a DbStub instead. Neither is an error on - its own — the difference shows up in which Ast.expr_kind comes - back. *) + (* Iteration 9 Task 3 retired the old asymmetry: `insert` is grammar-owned + in BOTH positions now — a typed Insert node validated like a ctor, + returning the id — while `select` stays the opaque DbStub until + Task 5. The old contract ("bare insert is a plain Ident") is gone + with the stub that motivated it. *) let prog, collector = - parse_str ~file:"insert-vs-select.wo" "fn f() {\n let a = insert\n let b = select\n}\n" + parse_str ~file:"insert-vs-select.wo" + "fn f() {\n let a = insert Product { sku: \"A1\" }\n let b = select\n}\n" in - check_eq "insert vs. select as bare expressions: no diagnostics" ~expected:0 + check_eq "typed insert + stub select: no diagnostics" ~expected:0 ~actual:(List.length (Diag.Collector.diagnostics collector)) string_of_int; - match prog.Ast.decls with + (match prog.Ast.decls with | [ Ast.Fn m ] -> ( match m.body with | [ { Ast.s_kind = Ast.Let { name = "a"; value = a_val; _ }; _ }; { Ast.s_kind = Ast.Let { name = "b"; value = b_val; _ }; _ }; ] -> - check "bare `insert` in expression position is a plain Ident" - (match a_val.Ast.kind with Ast.Ident "insert" -> true | _ -> false); + check "`insert` in expression position is a typed Insert node" + (match a_val.Ast.kind with Ast.Insert ("Product", [ ("sku", _) ]) -> true | _ -> false); check "bare `select` in expression position always becomes a DbStub" (match b_val.Ast.kind with Ast.DbStub _ -> true | _ -> false) | _ -> check "insert vs. select: exactly two `let` statements" false) - | _ -> check "insert vs. select: exactly one free fn" false + | _ -> check "insert vs. select: exactly one free fn" false); + (* and a bare `insert` with no literal is a parse error now, not an Ident *) + let _, c2 = parse_str ~file:"bare-insert.wo" "fn f() {\n let a = insert\n}\n" in + check "bare `insert` with no constructor literal is a diagnostic" + (List.length (Diag.Collector.diagnostics c2) > 0) let () = (* The no_brace guard (parser.ml's state.no_brace / looks_like_ctor): @@ -2382,7 +2382,7 @@ let validate_image (img : string) : string list = let u64 o = if ok 8 o then String.get_int64_le img o else 0L in let none = 0xFFFFFFFF in if u32 0 <> 0x31424F57 then fail "bad magic"; - if u32 4 <> 2 then fail "unsupported version"; + if u32 4 <> 3 then fail "unsupported version"; let coff = u32 8 and ccnt = u32 12 in let koff = u32 16 and kcnt = u32 20 in let ioff = u32 24 and icnt = u32 28 in @@ -2419,6 +2419,7 @@ let validate_image (img : string) : string list = if flags land lnot 0x01 <> 0 then fail (Printf.sprintf "class %d: unknown flags" i); if fcnt > 65535 then fail (Printf.sprintf "class %d: too many fields" i); class_fields.(i) <- fcnt; + let kco = !o in (* the kind bytes' offset: the v3 index walk re-reads them *) for j = 0 to fcnt - 1 do if u8 (!o + j) > 5 then fail (Printf.sprintf "class %d field %d: bad kind" i j) done; @@ -2435,6 +2436,25 @@ let validate_image (img : string) : string list = fail (Printf.sprintf "class %d field %d: field class out of range" i j) done; o := !o + (fcnt * 12); + (* v3: the index tail — flags (bit0 only), col_cnt 1..8, columns in + range and scalar/Text-kinded. Mirrors loader.c's checks. *) + let icnt_x = u32 !o in + o := !o + 4; + if icnt_x > 64 then fail (Printf.sprintf "class %d: too many indexes" i); + for x = 0 to icnt_x - 1 do + let ifl = u32 !o and ccnt = u32 (!o + 4) in + o := !o + 8; + if ifl land lnot 1 <> 0 then fail (Printf.sprintf "class %d index %d: unknown flags" i x); + if ccnt = 0 || ccnt > 8 then fail (Printf.sprintf "class %d index %d: bad column count" i x); + for c = 0 to ccnt - 1 do + let col = u32 !o in + o := !o + 4; + if col >= fcnt then fail (Printf.sprintf "class %d index %d: column out of range" i x); + let kind = u8 (kco + col) in + if kind <> 0 && kind <> 3 then + fail (Printf.sprintf "class %d index %d: column %d is not scalar or Text" i x c) + done + done; if !o > len then fail (Printf.sprintf "class %d: truncated" i) done; (* interfaces + vtable rows *) @@ -2580,7 +2600,12 @@ let validate_image (img : string) : string list = | 22 | 23 | 24 | 25 | 26 | 27 | 28 -> rchk pc a | 29 -> rchk pc a; - if c > 12 then fail (Printf.sprintf "method %d pc %d: builtin out of range" i pc) + (* the mirror's ceiling tracks wob.h's WO_B_MAX only for ids the + golden lowering suite actually emits; 61 = DB_INSERT (arity 1: + the class-id slot — field slots are runtime-validated, same as + the C loader) *) + if c > 12 && (c < 61 || c > 67) then + fail (Printf.sprintf "method %d pc %d: builtin out of range" i pc) else if c = 4 then begin if b > 5 then fail (Printf.sprintf "method %d pc %d: bad element kind" i pc) end @@ -2595,6 +2620,13 @@ let validate_image (img : string) : string list = | 1 | 2 | 3 | 7 | 8 -> 1 | 5 | 6 | 11 | 12 -> 2 | 10 -> 3 + | 61 -> 1 + | 62 -> 4 + | 63 -> 2 + | 64 -> 1 + | 65 -> 3 + | 66 -> 3 + | 67 -> 2 | _ -> 0 in if arity > 0 then begin @@ -2951,7 +2983,8 @@ let () = ( "text: concat, equality, words", "fn f(a: Text, b: Text) -> Int {\n let joined = a .. b\n\ \ if joined == a {\n return 1\n }\n return words(joined)\n}\n" ); - ("db stub statement", "fn f() -> Int {\n insert into rows values (1)\n return 0\n}\n"); + ( "db insert statement", + "class Row {\n n: Int\n}\n\nfn f() -> Int {\n insert Row { n: 1 }\n return 0\n}\n" ); ( "nested calls in arguments", "fn one() -> Int {\n return 1\n}\n\nfn add(a: Int, b: Int) -> Int {\n\ \ return a + b\n}\n\nfn f() -> Int {\n return add(add(one(), one()), one())\n}\n" ); diff --git a/database/src/CODE-LOGIC.md b/database/src/CODE-LOGIC.md new file mode 100644 index 0000000..ad5987e --- /dev/null +++ b/database/src/CODE-LOGIC.md @@ -0,0 +1,63 @@ +# database/src — how the engine hangs together + +The database engine is its own top-level directory, statically linked into +every `wovm` and every runtime test binary (`runtime/Makefile`'s `DBSRC`). +One binary, unchanged. Format doc: `docs/plan/oop-vm/04-db-binding.md`. +Memory-safety doctrine: the 9b design's section 6. + +## table.c — rows (iteration 9, Task 1) + +``` +VM values ──copy──▶ row slots (engine-owned malloc) ──copy──▶ fresh VM values + wo_row_insert wo_row_read +``` + +- **No VM pointer ever enters a slab; no slab pointer ever leaves.** Encode + copies per kind (Texts to `db_text`, owned objects flattened recursively to + `db_rec`, containers element-wise); decode allocates fresh VM values from + the caller's `wo_rt`. The GCREF kind is refused at encode — the compiler + should have made that impossible (the GC bulkhead), the engine refuses it + anyway. +- **Rows never move.** Slabs of 256 are malloc'd and kept for the table's + life; the free-slot list recycles removed slots before any slab grows; + the id hash maps id → slot. Ids are never reused (per-table counter, + shard-interleaved `S+1, S+1+N, …`), which is also what makes the hash's + tombstone sentinel safe. +- **Choke points**: `wo_row_insert` / `wo_row_remove` carry the `INDEX HOOK` + comments where Task 4's secondary indexes attach and Task 2's WAL stages + its record. Nothing else may mutate storage. +- One deliberate file-static: `g_classes` for recursive frees (`db_val_free` + has no context parameter). One process, one class table; revisit at + iteration 8 (shards share the same immutable table). + +## wal.c — durability (iteration 9, Task 2) + +The commit order IS the module: RAM apply → stage → one pwrite + one +fdatasync → ack. `wo_wal_commit` returning 0 is the only thing "durable" +means. Replay never touches the VM heap — payloads decode straight into +engine-owned values and re-enter through the row API, so whatever hooks the +choke points (indexes, Task 4) applies to replayed rows identically. Torn +tails end the intact prefix and get overwritten by the next commit; +CRC-valid-but-undecodable records fail replay loudly (corruption is not a +tear). The crash battery in `runtime/test/test_wal.c` is the module's +meaning proven: acked-over-a-pipe after commit, SIGKILL mid-stream, replay, +zero acked-but-missing. + +## db.c — statement executors (iteration 9, Task 3) + +One dispatcher, the builtin contract (0 ok, else WO_T_* + msg). The engine +handles ride `wo_rt.db` / `wo_rt.wal` as opaque pointers set by main.c — +NULL db traps WO_T_DB, NULL wal means RAM-only (the corpus's mode; WO_DATA +opts into durability). Insert's contract: RAM apply through the row API, +then stage + commit BEFORE returning — the builtin's return is the +acknowledgment, so a failed commit un-applies the row and traps WO_T_IO +rather than acknowledging what disk never got. + +## Verifying a change + +- `make -C runtime test` — `test_table` is this directory's suite (round + trips across kinds, nil encodings, shard interleave, slab growth, slot + reuse, misuse), ASan+UBSan like every runtime test. +- `just oop-e2e`, `just log-watcher` — regression that linking the engine + into wovm changed nothing observable (it is dead code until Task 3 wires + the first builtin). diff --git a/database/src/db.c b/database/src/db.c new file mode 100644 index 0000000..3f40d88 --- /dev/null +++ b/database/src/db.c @@ -0,0 +1,163 @@ +#include "db.h" + +#include + +#include "cont.h" +#include "table.h" +#include "wal.h" + +int wo_builtin_db(wo_vm *vm, uint64_t *R, uint32_t ins, const char **msg) { + uint32_t A = wo_ins_a(ins), B = wo_ins_b(ins), C = wo_ins_c(ins); + wo_db *db = (wo_db *)vm->rt.db; + if (!db) { + *msg = "database engine not initialized"; + return WO_T_DB; + } + switch (C) { + case WO_B_DB_INSERT: { + uint32_t cid = (uint32_t)R[B]; + int ek = 0; + uint64_t id = wo_row_insert(db, cid, &R[B + 1], msg, &ek); + if (!id) + return ek == DB_ERR_UNIQUE ? WO_T_UNIQUE + : ek == DB_ERR_OOM ? WO_T_OOM + : WO_T_DB; + wo_wal *w = (wo_wal *)vm->rt.wal; + if (w) { + /* RAM applied, record staged, ONE commit before the ack (the + * builtin's return). A failed commit is a failed write: the + * row is removed again so RAM never claims what disk never + * acknowledged, and the statement traps. */ + if (wo_wal_append_insert(w, db, cid, id) != 0 || wo_wal_commit(w) != 0) { + wo_row_remove(db, cid, id); + *msg = "wal commit failed"; + return WO_T_IO; + } + } + R[A] = id; + return 0; + } + case WO_B_DB_UPDATE_FIELD: { + uint32_t cid = (uint32_t)R[B]; + uint64_t id = R[B + 1]; + uint32_t field = (uint32_t)R[B + 2]; + int ek = 0; + if (wo_row_update_field(db, cid, id, field, R[B + 3], msg, &ek) != 0) + return ek == DB_ERR_UNIQUE ? WO_T_UNIQUE : ek == DB_ERR_OOM ? WO_T_OOM : WO_T_DB; + wo_wal *w = (wo_wal *)vm->rt.wal; + if (w) { + if (wo_wal_append_update(w, db, cid, id) != 0 || wo_wal_commit(w) != 0) { + *msg = "wal commit failed"; /* RAM ahead of disk: trap, do not ack */ + return WO_T_IO; + } + } + R[A] = 0; + return 0; + } + case WO_B_DB_DELETE: { + uint32_t cid = (uint32_t)R[B]; + uint64_t id = R[B + 1]; + /* FK restrict: refuse if another row still references this one + (iteration 9b) — nothing is removed, the statement traps */ + if (wo_row_has_referrers(db, cid, id)) { + *msg = "row is still referenced (restrict)"; + return WO_T_FK; + } + if (wo_row_remove(db, cid, id) != 0) { + *msg = "no such row"; + return WO_T_DB; + } + wo_wal *w = (wo_wal *)vm->rt.wal; + if (w) { + if (wo_wal_append_remove(w, cid, id) != 0 || wo_wal_commit(w) != 0) { + *msg = "wal commit failed"; + return WO_T_IO; + } + } + R[A] = 0; + return 0; + } + case WO_B_DB_SCAN: { + uint32_t cid = (uint32_t)R[B]; + if (cid >= db->class_cnt) { + *msg = "no such class"; + return WO_T_DB; + } + wo_multi *ids = wo_multi_new(&vm->rt, WO_K_SCALAR); + if (!ids) return WO_T_OOM; + /* materialize the id list up front — the 9b cursor-stability rule: + * the loop body then point-reads each id, so a row updated mid-loop + * (even an indexed column) cannot disturb the iteration */ + db_table *t = &db->tables[cid]; + if (t->row_size) { + uint32_t total = t->slab_cnt * DB_SLAB_ROWS; + for (uint32_t g = 0; g < total; g++) { + if (!(t->bitmap[g >> 6] & (1ull << (g & 63)))) continue; + db_row *row = + (db_row *)(t->slabs[g / DB_SLAB_ROWS] + (size_t)(g % DB_SLAB_ROWS) * t->row_size); + if (wo_multi_push(ids, row->id) != 0) return WO_T_OOM; + } + } + R[A] = (uint64_t)(uintptr_t)ids; + return 0; + } + case WO_B_DB_GET_FIELD: { + uint32_t cid = (uint32_t)R[B]; + uint64_t id = R[B + 1]; + uint32_t field = (uint32_t)R[B + 2]; + if (cid >= db->class_cnt || field >= db->classes[cid].field_cnt) { + *msg = "no such field"; + return WO_T_DB; + } + db_row *row = wo_row_ptr(db, cid, id); + if (!row) { + *msg = "no such row"; + return WO_T_DB; + } + int ok = 1; + uint64_t v = wo_val_decode_vm(db, &vm->rt, db->classes[cid].kinds[field], + row->slots[field], &ok, msg); + if (!ok) return WO_T_OOM; + R[A] = v; + return 0; + } + case WO_B_DB_PROBE: { + uint32_t cid = (uint32_t)R[B]; + uint32_t index = (uint32_t)R[B + 1]; + if (cid >= db->class_cnt) { + *msg = "no such class"; + return WO_T_DB; + } + wo_multi *ids = wo_multi_new(&vm->rt, WO_K_SCALAR); + if (!ids) return WO_T_OOM; + db_table *t = &db->tables[cid]; + if (t->row_size && index < t->index_cnt) { + db_index *ix = &t->indexes[index]; + uint32_t col = ix->cols[0]; + uint8_t kind = db->classes[cid].kinds[col]; + uint64_t key = R[B + 2]; + uint32_t total = t->slab_cnt * DB_SLAB_ROWS; + for (uint32_t g = 0; g < total; g++) { + if (!(t->bitmap[g >> 6] & (1ull << (g & 63)))) continue; + db_row *row = + (db_row *)(t->slabs[g / DB_SLAB_ROWS] + (size_t)(g % DB_SLAB_ROWS) * t->row_size); + int eq; + if (kind == WO_K_TEXT) { + const wo_str *want = (const wo_str *)(uintptr_t)key; + const db_text *have = (const db_text *)(uintptr_t)row->slots[col]; + eq = (!want && !have) || + (want && have && want->len == have->len && + memcmp(want->data, have->bytes, have->len) == 0); + } else + eq = row->slots[col] == key; + if (eq && wo_multi_push(ids, row->id) != 0) return WO_T_OOM; + } + } + R[A] = (uint64_t)(uintptr_t)ids; + return 0; + } + default: + *msg = "unknown db builtin"; + return WO_T_DB; + } +} diff --git a/database/src/db.h b/database/src/db.h new file mode 100644 index 0000000..95eb97d --- /dev/null +++ b/database/src/db.h @@ -0,0 +1,24 @@ +/* db.h — DB statement executors (iteration 9, Task 3+). + * + * The VM reaches the engine through one dispatcher with the same contract + * as every builtin family: 0 = ok, else a WO_T_* code with *msg set. The + * engine and WAL handles ride the runtime context as opaque pointers + * (obj.h's rt.db / rt.wal) — set by main.c at boot, NULL in test binaries + * that never touch DB statements (a DB builtin with rt.db == NULL traps + * WO_T_DB "engine not initialized"). + * + * Commit contract per statement (until iteration 8 brings ticks): the + * insert applies to RAM, stages its WAL record, and COMMITS before the + * builtin returns — the builtin returning IS the acknowledgment, so the + * ack-after-fsync doctrine holds at statement granularity. No WAL + * (rt.wal == NULL, no WO_DATA) means RAM-only: every test and every + * corpus fixture runs that way; durability is opt-in by pointing WO_DATA + * at a directory. */ +#ifndef WO_DB_H +#define WO_DB_H + +#include "vm.h" + +int wo_builtin_db(wo_vm *vm, uint64_t *R, uint32_t ins, const char **msg); + +#endif /* WO_DB_H */ diff --git a/database/src/table.c b/database/src/table.c new file mode 100644 index 0000000..a8dc324 --- /dev/null +++ b/database/src/table.c @@ -0,0 +1,753 @@ +#include "table.h" + +#include +#include + +#include "cont.h" + +/* ---- engine-owned value encode / free / decode ------------------------- */ + +/* Free one encoded slot value of [kind]. Recursion mirrors encoding. */ +static void db_val_free(uint8_t kind, uint64_t v); + +static void db_rec_free(db_rec *r, const wo_classdesc *classes) { + const wo_classdesc *c = &classes[r->class_id]; + for (uint32_t i = 0; i < c->field_cnt; i++) db_val_free(c->kinds[i], r->slots[i]); + free(r); +} + +/* db_val_free needs the class table for nested records; a file-static is + * the honest signature here — one engine per process today (N=1), and the + * pointer is set once at init. Revisit when iteration 8 brings N>1 shards + * (each shard's wo_db shares the same immutable class table anyway). */ +static const wo_classdesc *g_classes; + +static void db_val_free(uint8_t kind, uint64_t v) { + if (!v) return; + switch (kind) { + case WO_K_SCALAR: return; + case WO_K_TEXT: free((db_text *)(uintptr_t)v); return; + case WO_K_OWNED: db_rec_free((db_rec *)(uintptr_t)v, g_classes); return; + case WO_K_MULTI: { + db_multi *m = (db_multi *)(uintptr_t)v; + for (uint32_t i = 0; i < m->len; i++) db_val_free(m->elem_kind, m->items[i]); + free(m); + return; + } + case WO_K_MAP: { + db_map *m = (db_map *)(uintptr_t)v; + for (uint32_t i = 0; i < m->len; i++) { + db_val_free(m->key_kind, m->kv[2 * i]); + db_val_free(m->val_kind, m->kv[2 * i + 1]); + } + free(m); + return; + } + default: return; /* GCREF never stored */ + } +} + +/* Encode one VM value into an engine-owned slot value. 0-with-*ok=0 means + * failure (OOM or a GCREF); a genuine nil encodes as 0 with *ok=1. */ +static uint64_t db_val_encode(const wo_classdesc *classes, uint8_t kind, uint64_t v, + int *ok, const char **msg) { + *ok = 1; + switch (kind) { + case WO_K_SCALAR: return v; + case WO_K_TEXT: { + if (!v) return 0; + const wo_str *s = (const wo_str *)(uintptr_t)v; + db_text *t = malloc(sizeof(db_text) + s->len); + if (!t) goto oom; + t->len = s->len; + memcpy(t->bytes, s->data, s->len); + return (uint64_t)(uintptr_t)t; + } + case WO_K_OWNED: { + if (!v) return 0; + const wo_hdr *o = (const wo_hdr *)(uintptr_t)v; + const wo_classdesc *c = &classes[o->class_id]; + db_rec *r = malloc(sizeof(db_rec) + (size_t)c->field_cnt * 8u); + if (!r) goto oom; + r->class_id = o->class_id; + r->_pad = 0; + const uint64_t *f = (const uint64_t *)(const void *)(o + 1); + for (uint32_t i = 0; i < c->field_cnt; i++) { + r->slots[i] = db_val_encode(classes, c->kinds[i], f[i], ok, msg); + if (!*ok) { /* free what we built so far, then fail upward */ + for (uint32_t j = 0; j < i; j++) db_val_free(c->kinds[j], r->slots[j]); + free(r); + return 0; + } + } + return (uint64_t)(uintptr_t)r; + } + case WO_K_MULTI: { + if (!v) return 0; + const wo_multi *m = (const wo_multi *)(uintptr_t)v; + db_multi *d = malloc(sizeof(db_multi) + (size_t)m->len * 8u); + if (!d) goto oom; + d->elem_kind = m->elem_kind; + d->len = m->len; + for (uint32_t i = 0; i < m->len; i++) { + d->items[i] = db_val_encode(classes, m->elem_kind, m->items[i], ok, msg); + if (!*ok) { + for (uint32_t j = 0; j < i; j++) db_val_free(d->elem_kind, d->items[j]); + free(d); + return 0; + } + } + return (uint64_t)(uintptr_t)d; + } + case WO_K_MAP: { + if (!v) return 0; + const wo_map *m = (const wo_map *)(uintptr_t)v; + db_map *d = malloc(sizeof(db_map) + (size_t)m->len * 16u); + if (!d) goto oom; + d->key_kind = m->key_kind; + d->val_kind = m->val_kind; + d->len = m->len; + for (uint32_t i = 0; i < m->len; i++) { + d->kv[2 * i] = db_val_encode(classes, m->key_kind, m->keys[i], ok, msg); + uint64_t dv = 0; + if (*ok) dv = db_val_encode(classes, m->val_kind, m->vals[i], ok, msg); + d->kv[2 * i + 1] = dv; + if (!*ok) { + for (uint32_t j = 0; j <= i; j++) { + db_val_free(d->key_kind, d->kv[2 * j]); + db_val_free(d->val_kind, d->kv[2 * j + 1]); + } + free(d); + return 0; + } + } + return (uint64_t)(uintptr_t)d; + } + default: + *ok = 0; + *msg = "a garbage-collected value cannot be stored in a table field"; + return 0; + } +oom: + *ok = 0; + *msg = "out of memory encoding a row"; + return 0; +} + +/* Decode one engine slot back into a fresh VM value (the out-gate: always + * a copy). 0-with-*ok=0 = OOM; nil decodes as 0 with *ok=1. */ +static uint64_t db_val_decode(wo_rt *rt, uint8_t kind, uint64_t v, int *ok, + const char **msg) { + *ok = 1; + switch (kind) { + case WO_K_SCALAR: return v; + case WO_K_TEXT: { + if (!v) return 0; + const db_text *t = (const db_text *)(uintptr_t)v; + wo_str *s = wo_str_new(rt, t->bytes, t->len); + if (!s) goto oom; + return (uint64_t)(uintptr_t)s; + } + case WO_K_OWNED: { + if (!v) return 0; + const db_rec *r = (const db_rec *)(uintptr_t)v; + wo_hdr *o = wo_obj_new(rt, r->class_id); + if (!o) goto oom; + const wo_classdesc *c = &rt->classes[r->class_id]; + uint64_t *f = wo_fields(o); + for (uint32_t i = 0; i < c->field_cnt; i++) { + f[i] = db_val_decode(rt, c->kinds[i], r->slots[i], ok, msg); + if (!*ok) return 0; /* partial object: rt teardown reclaims (test scope) */ + } + return (uint64_t)(uintptr_t)o; + } + case WO_K_MULTI: { + if (!v) return 0; + const db_multi *d = (const db_multi *)(uintptr_t)v; + wo_multi *m = wo_multi_new(rt, d->elem_kind); + if (!m) goto oom; + for (uint32_t i = 0; i < d->len; i++) { + uint64_t ev = db_val_decode(rt, d->elem_kind, d->items[i], ok, msg); + if (!*ok || wo_multi_push(m, ev) != 0) goto oom; + } + return (uint64_t)(uintptr_t)m; + } + case WO_K_MAP: { + if (!v) return 0; + const db_map *d = (const db_map *)(uintptr_t)v; + wo_map *m = wo_map_new(rt, d->key_kind, d->val_kind); + if (!m) goto oom; + for (uint32_t i = 0; i < d->len; i++) { + uint64_t kv = db_val_decode(rt, d->key_kind, d->kv[2 * i], ok, msg); + uint64_t vv = 0; + if (*ok) vv = db_val_decode(rt, d->val_kind, d->kv[2 * i + 1], ok, msg); + uint64_t old; + if (!*ok || wo_map_set(m, kv, vv, &old) < 0) goto oom; + } + return (uint64_t)(uintptr_t)m; + } + default: return 0; /* GCREF never stored, so never decoded */ + } +oom: + *ok = 0; + *msg = "out of memory decoding a row"; + return 0; +} + +/* ---- id hash (open addressing, pow2, id -> global slot + 1) ----------- */ + +static uint64_t hmix(uint64_t x) { /* splitmix64 finalizer */ + x += 0x9e3779b97f4a7c15ull; + x = (x ^ (x >> 30)) * 0xbf58476d1ce4e5b9ull; + x = (x ^ (x >> 27)) * 0x94d049bb133111ebull; + return x ^ (x >> 31); +} + +/* Ids are never 0 (0 spells "empty bucket") and never reused, so all-ones + * can never collide with a live id — it marks a deleted bucket that probes + * walk straight past. */ +#define H_DELETED ((uint64_t)-1) + +static int hgrow(db_table *t) { + size_t ncap = t->hcap ? t->hcap * 2 : 64; + uint64_t *nk = calloc(ncap, 8), *nv = calloc(ncap, 8); + if (!nk || !nv) { + free(nk); + free(nv); + return -1; + } + for (size_t i = 0; i < t->hcap; i++) { + if (!t->hkeys[i] || t->hkeys[i] == H_DELETED) continue; + size_t j = hmix(t->hkeys[i]) & (ncap - 1); + while (nk[j]) j = (j + 1) & (ncap - 1); + nk[j] = t->hkeys[i]; + nv[j] = t->hvals[i]; + } + free(t->hkeys); + free(t->hvals); + t->hkeys = nk; + t->hvals = nv; + t->hcap = ncap; + return 0; +} + +static int hput(db_table *t, uint64_t id, uint64_t slot1) { + if (t->hlen * 10 >= t->hcap * 7 && hgrow(t) != 0) return -1; + size_t j = hmix(id) & (t->hcap - 1); + while (t->hkeys[j] && t->hkeys[j] != id) j = (j + 1) & (t->hcap - 1); + if (!t->hkeys[j]) t->hlen++; + t->hkeys[j] = id; + t->hvals[j] = slot1; + return 0; +} + +static uint64_t hget(const db_table *t, uint64_t id) { + if (!t->hcap) return 0; + size_t j = hmix(id) & (t->hcap - 1); + while (t->hkeys[j]) { + if (t->hkeys[j] == id) return t->hvals[j]; + j = (j + 1) & (t->hcap - 1); + } + return 0; +} + +static void hdel(db_table *t, uint64_t id) { + if (!t->hcap) return; + size_t j = hmix(id) & (t->hcap - 1); + while (t->hkeys[j]) { + if (t->hkeys[j] == id) { + t->hkeys[j] = H_DELETED; + t->hvals[j] = 0; + return; + } + j = (j + 1) & (t->hcap - 1); + } +} + +/* ---- tables and rows ---------------------------------------------------- */ + +/* ---- secondary indexes (Task 4) ---------------------------------------- */ + +/* hash of one row's index columns: kind-driven, never trusted for equality */ +static uint64_t idx_hash(const wo_classdesc *c, const db_index *ix, const db_row *r) { + uint64_t h = 0x9e3779b97f4a7c15ull; + for (uint32_t i = 0; i < ix->col_cnt; i++) { + uint32_t col = ix->cols[i]; + uint64_t v = r->slots[col]; + if (c->kinds[col] == WO_K_TEXT) { + const db_text *t = (const db_text *)(uintptr_t)v; + uint64_t th = 1469598103934665603ull; /* FNV-1a over bytes; nil = 0 */ + if (t) + for (uint32_t b = 0; b < t->len; b++) th = (th ^ (uint8_t)t->bytes[b]) * 1099511628211ull; + else th = 0; + v = th; + } + h ^= hmix(v + i); + } + return h ? h : 1; /* 0 marks an empty bucket */ +} + +static int idx_cols_equal(const wo_classdesc *c, const db_index *ix, const db_row *a, + const db_row *b) { + for (uint32_t i = 0; i < ix->col_cnt; i++) { + uint32_t col = ix->cols[i]; + if (c->kinds[col] == WO_K_TEXT) { + const db_text *x = (const db_text *)(uintptr_t)a->slots[col]; + const db_text *y = (const db_text *)(uintptr_t)b->slots[col]; + if (!x || !y) { + if (x != y) return 0; + } else if (x->len != y->len || memcmp(x->bytes, y->bytes, x->len) != 0) + return 0; + } else if (a->slots[col] != b->slots[col]) + return 0; + } + return 1; +} + +static db_ibucket *idx_bucket(db_index *ix, uint64_t h, int create) { + if (ix->bcap == 0) { + if (!create) return NULL; + ix->buckets = calloc(64, sizeof(db_ibucket)); + if (!ix->buckets) return NULL; + ix->bcap = 64; + } + if (create && ix->blen * 10 >= ix->bcap * 7) { + size_t ncap = ix->bcap * 2; + db_ibucket *nb = calloc(ncap, sizeof(db_ibucket)); + if (!nb) return NULL; + for (size_t i = 0; i < ix->bcap; i++) { + if (!ix->buckets[i].hash) continue; + size_t j = ix->buckets[i].hash & (ncap - 1); + while (nb[j].hash) j = (j + 1) & (ncap - 1); + nb[j] = ix->buckets[i]; + } + free(ix->buckets); + ix->buckets = nb; + ix->bcap = ncap; + } + size_t j = h & (ix->bcap - 1); + while (ix->buckets[j].hash) { + if (ix->buckets[j].hash == h) return &ix->buckets[j]; + j = (j + 1) & (ix->bcap - 1); + } + if (!create) return NULL; + ix->buckets[j].hash = h; + ix->blen++; + return &ix->buckets[j]; +} + +/* Add [r] to every index; unique violation reports which without mutating + * anything (checks run before any add). 0 ok, DB_ERR_* otherwise. */ +static int idx_add_row(wo_db *db, db_table *t, db_row *r) { + const wo_classdesc *c = &db->classes[t->class_id]; + for (uint32_t x = 0; x < t->index_cnt; x++) { + db_index *ix = &t->indexes[x]; + if (!(ix->flags & 1u)) continue; + db_ibucket *b = idx_bucket(ix, idx_hash(c, ix, r), 0); + if (!b) continue; + for (uint32_t i = 0; i < b->len; i++) { + db_row *other = wo_row_ptr(db, t->class_id, b->ids[i]); + if (other && idx_cols_equal(c, ix, r, other)) return DB_ERR_UNIQUE; + } + } + for (uint32_t x = 0; x < t->index_cnt; x++) { + db_index *ix = &t->indexes[x]; + db_ibucket *b = idx_bucket(ix, idx_hash(c, ix, r), 1); + if (!b) return DB_ERR_OOM; + if (b->len == b->cap) { + uint32_t ncap = b->cap ? b->cap * 2 : 4; + uint64_t *ni = realloc(b->ids, (size_t)ncap * 8u); + if (!ni) return DB_ERR_OOM; + b->ids = ni; + b->cap = ncap; + } + b->ids[b->len++] = r->id; + } + return 0; +} + +static void idx_remove_row(wo_db *db, db_table *t, db_row *r) { + const wo_classdesc *c = &db->classes[t->class_id]; + for (uint32_t x = 0; x < t->index_cnt; x++) { + db_index *ix = &t->indexes[x]; + db_ibucket *b = idx_bucket(ix, idx_hash(c, ix, r), 0); + if (!b) continue; + for (uint32_t i = 0; i < b->len; i++) + if (b->ids[i] == r->id) { + b->ids[i] = b->ids[--b->len]; + break; + } + } +} + +int wo_db_init(wo_db *db, const wo_classdesc *classes, uint32_t class_cnt, + uint32_t shard, uint32_t nshards) { + if (!nshards || shard >= nshards) return -1; + memset(db, 0, sizeof(*db)); + db->classes = classes; + db->class_cnt = class_cnt; + db->shard = shard; + db->nshards = nshards; + db->tables = calloc(class_cnt ? class_cnt : 1, sizeof(db_table)); + if (!db->tables) return -1; + g_classes = classes; + return 0; +} + +static void table_destroy(wo_db *db, db_table *t) { + /* free every live row's engine-owned values, then the slabs */ + const wo_classdesc *c = &db->classes[t->class_id]; + for (uint32_t s = 0; s < t->slab_cnt; s++) { + for (uint32_t i = 0; i < DB_SLAB_ROWS; i++) { + uint32_t g = s * DB_SLAB_ROWS + i; + if (!(t->bitmap[g >> 6] & (1ull << (g & 63)))) continue; + db_row *r = (db_row *)(t->slabs[s] + (size_t)i * t->row_size); + for (uint32_t f = 0; f < c->field_cnt; f++) + db_val_free(c->kinds[f], r->slots[f]); + } + free(t->slabs[s]); + } + free(t->slabs); + free(t->bitmap); + free(t->free_slots); + free(t->hkeys); + free(t->hvals); + for (uint32_t x = 0; x < t->index_cnt; x++) { + for (size_t b = 0; b < t->indexes[x].bcap; b++) free(t->indexes[x].buckets[b].ids); + free(t->indexes[x].buckets); + } + free(t->indexes); +} + +void wo_db_destroy(wo_db *db) { + if (!db->tables) return; + for (uint32_t i = 0; i < db->class_cnt; i++) + if (db->tables[i].slab_cnt || db->tables[i].hkeys) table_destroy(db, &db->tables[i]); + free(db->tables); + db->tables = NULL; +} + +static db_table *table_of(wo_db *db, uint32_t class_id) { + if (class_id >= db->class_cnt) return NULL; + db_table *t = &db->tables[class_id]; + if (!t->row_size) { /* lazy init on first touch */ + const wo_classdesc *c = &db->classes[class_id]; + t->class_id = class_id; + t->row_size = sizeof(db_row) + (size_t)c->field_cnt * 8u; + t->next_id = db->shard + 1; /* S+1, then += N: interleaved, local-only */ + if (c->idx_cnt) { + t->indexes = calloc(c->idx_cnt, sizeof(db_index)); + if (!t->indexes) return NULL; + const uint32_t *im = c->idx_meta; + for (uint32_t x = 0; x < c->idx_cnt; x++) { + t->indexes[x].flags = im[0]; + t->indexes[x].col_cnt = im[1]; + t->indexes[x].cols = im + 2; + im += 2 + im[1]; + } + t->index_cnt = c->idx_cnt; + } + } + return t; +} + +static db_row *slot_row(db_table *t, uint32_t g) { + return (db_row *)(t->slabs[g / DB_SLAB_ROWS] + (size_t)(g % DB_SLAB_ROWS) * t->row_size); +} + +/* Pick the slot a new row lands in: recycled first, else the next free bit, + * else grow a slab. Returns the global slot or UINT32_MAX on OOM. */ +static uint32_t slot_alloc(db_table *t) { + if (t->free_cnt) return t->free_slots[--t->free_cnt]; + uint32_t total = t->slab_cnt * DB_SLAB_ROWS; + for (uint32_t g = 0; g < total; g++) /* cheap at slab granularity: only + reached when free list is empty, and the bitmap scan is bounded by + one word test per 64 slots */ + if (!(t->bitmap[g >> 6] & (1ull << (g & 63)))) return g; + /* grow */ + if (t->slab_cnt == t->slab_cap) { + uint32_t ncap = t->slab_cap ? t->slab_cap * 2 : 4; + uint8_t **ns = realloc(t->slabs, (size_t)ncap * sizeof(uint8_t *)); + if (!ns) return UINT32_MAX; + t->slabs = ns; + t->slab_cap = ncap; + } + uint8_t *slab = malloc((size_t)DB_SLAB_ROWS * t->row_size); + if (!slab) return UINT32_MAX; + size_t nwords = ((size_t)(t->slab_cnt + 1) * DB_SLAB_ROWS + 63) / 64; + uint64_t *nb = realloc(t->bitmap, nwords * 8); + if (!nb) { + free(slab); + return UINT32_MAX; + } + memset(nb + ((size_t)t->slab_cnt * DB_SLAB_ROWS) / 64, 0, + (nwords - ((size_t)t->slab_cnt * DB_SLAB_ROWS) / 64) * 8); + t->bitmap = nb; + t->slabs[t->slab_cnt] = slab; + return t->slab_cnt++ * DB_SLAB_ROWS; +} + +uint64_t wo_row_insert(wo_db *db, uint32_t class_id, const uint64_t *vals, + const char **msg, int *err_kind) { + if (err_kind) *err_kind = DB_ERR_MISC; + db_table *t = table_of(db, class_id); + if (!t) { + *msg = "no such class"; + return 0; + } + const wo_classdesc *c = &db->classes[class_id]; + uint32_t g = slot_alloc(t); + if (g == UINT32_MAX) { + *msg = "out of memory growing a table"; + return 0; + } + db_row *r = slot_row(t, g); + r->class_id = class_id; + r->flags = 0; + int ok = 1; + uint32_t i = 0; + for (; i < c->field_cnt; i++) { + r->slots[i] = db_val_encode(db->classes, c->kinds[i], vals[i], &ok, msg); + if (!ok) { + if (err_kind) *err_kind = DB_ERR_BADKIND; + break; + } + } + if (!ok) { + for (uint32_t j = 0; j < i; j++) db_val_free(c->kinds[j], r->slots[j]); + /* slot never became live: recycle it (bitmap bit was never set) */ + if (t->free_cnt == t->free_cap) { + uint32_t ncap = t->free_cap ? t->free_cap * 2 : 16; + uint32_t *nf = realloc(t->free_slots, (size_t)ncap * 4); + if (nf) { + t->free_slots = nf; + t->free_cap = ncap; + } + } + if (t->free_cnt < t->free_cap) t->free_slots[t->free_cnt++] = g; + return 0; + } + r->id = t->next_id; + t->next_id += db->nshards; + if (hput(t, r->id, (uint64_t)g + 1) != 0) { + for (uint32_t j = 0; j < c->field_cnt; j++) db_val_free(c->kinds[j], r->slots[j]); + if (err_kind) *err_kind = DB_ERR_OOM; + *msg = "out of memory indexing a row"; + return 0; + } + t->bitmap[g >> 6] |= 1ull << (g & 63); + t->count++; + /* THE index hook (Task 4): inside the choke point, never anywhere else. + A unique violation un-applies the row entirely — id never handed out + twice matters less than the row never having existed. */ + int irc = idx_add_row(db, t, r); + if (irc != 0) { + t->bitmap[g >> 6] &= ~(1ull << (g & 63)); + hdel(t, r->id); + t->count--; + t->next_id -= db->nshards; /* the id was never observable: reclaim it */ + for (uint32_t j = 0; j < c->field_cnt; j++) db_val_free(c->kinds[j], r->slots[j]); + if (t->free_cnt < t->free_cap) t->free_slots[t->free_cnt++] = g; + if (err_kind) *err_kind = irc; + *msg = irc == DB_ERR_UNIQUE ? "unique index violation" : "out of memory indexing a row"; + return 0; + } + if (err_kind) *err_kind = DB_ERR_NONE; + return r->id; +} + +db_row *wo_row_ptr(wo_db *db, uint32_t class_id, uint64_t id) { + if (class_id >= db->class_cnt) return NULL; + db_table *t = &db->tables[class_id]; + if (!t->row_size) return NULL; + uint64_t s1 = hget(t, id); + if (!s1) return NULL; + return slot_row(t, (uint32_t)(s1 - 1)); +} + +int wo_row_read(wo_db *db, wo_rt *rt, uint32_t class_id, uint64_t id, + uint64_t *out_vals, const char **msg) { + db_row *r = wo_row_ptr(db, class_id, id); + if (!r) return -1; + const wo_classdesc *c = &db->classes[class_id]; + int ok = 1; + for (uint32_t i = 0; i < c->field_cnt; i++) { + out_vals[i] = db_val_decode(rt, c->kinds[i], r->slots[i], &ok, msg); + if (!ok) return -2; + } + return 0; +} + +db_row *wo_row_create_raw(wo_db *db, uint32_t class_id, uint64_t id) { + db_table *t = table_of(db, class_id); + if (!t || !id) return NULL; + if (hget(t, id)) return NULL; /* duplicate id: corruption, not a tear */ + uint32_t g = slot_alloc(t); + if (g == UINT32_MAX) return NULL; + db_row *r = slot_row(t, g); + r->id = id; + r->class_id = class_id; + r->flags = 0; + memset(r->slots, 0, t->row_size - sizeof(db_row)); + if (hput(t, id, (uint64_t)g + 1) != 0) return NULL; + t->bitmap[g >> 6] |= 1ull << (g & 63); + t->count++; + /* keep the interleave: only ids this shard owns move its counter */ + if ((id - 1) % db->nshards == db->shard && id >= t->next_id) + t->next_id = id + db->nshards; + /* indexes: NOT here — the slots are still zero. wal.c fills them and + then calls wo_row_raw_commit, which is where replayed rows re-index. */ + return r; +} + +int wo_row_raw_commit(wo_db *db, uint32_t class_id, db_row *r) { + db_table *t = &db->tables[class_id]; + return idx_add_row(db, t, r) == 0 ? 0 : -1; +} + +void wo_db_val_free(wo_db *db, uint8_t kind, uint64_t v) { + (void)db; + db_val_free(kind, v); +} + +uint64_t wo_val_decode_vm(wo_db *db, wo_rt *rt, uint8_t kind, uint64_t engine_val, + int *ok, const char **msg) { + (void)db; + return db_val_decode(rt, kind, engine_val, ok, msg); +} + +int wo_row_update_field(wo_db *db, uint32_t class_id, uint64_t id, uint32_t field, + uint64_t vm_val, const char **msg, int *err_kind) { + if (err_kind) *err_kind = DB_ERR_MISC; + db_row *r = wo_row_ptr(db, class_id, id); + if (!r) { + *msg = "no such row"; + return -1; + } + const wo_classdesc *c = &db->classes[class_id]; + if (field >= c->field_cnt) { + *msg = "no such field"; + return -1; + } + db_table *t = &db->tables[class_id]; + int ok = 1; + uint64_t nv = db_val_encode(db->classes, c->kinds[field], vm_val, &ok, msg); + if (!ok) { + if (err_kind) *err_kind = DB_ERR_BADKIND; + return -1; + } + /* indexes containing this column: unique checks against the NEW value + run first, against a shadow of the row, before anything mutates */ + uint64_t old = r->slots[field]; + r->slots[field] = nv; + for (uint32_t x = 0; x < t->index_cnt; x++) { + db_index *ix = &t->indexes[x]; + if (!(ix->flags & 1u)) continue; + int touches = 0; + for (uint32_t i = 0; i < ix->col_cnt; i++) + if (ix->cols[i] == field) touches = 1; + if (!touches) continue; + db_ibucket *b = idx_bucket(ix, idx_hash(c, ix, r), 0); + if (!b) continue; + for (uint32_t i = 0; i < b->len; i++) { + if (b->ids[i] == id) continue; + db_row *other = wo_row_ptr(db, class_id, b->ids[i]); + if (other && idx_cols_equal(c, ix, r, other)) { + r->slots[field] = old; /* untouched, promised */ + db_val_free(c->kinds[field], nv); + if (err_kind) *err_kind = DB_ERR_UNIQUE; + *msg = "unique index violation"; + return -1; + } + } + } + /* commit: fix every index containing the column (old entry out under + the OLD value's hash, new entry in), then free the old value */ + r->slots[field] = old; + for (uint32_t x = 0; x < t->index_cnt; x++) { + db_index *ix = &t->indexes[x]; + int touches = 0; + for (uint32_t i = 0; i < ix->col_cnt; i++) + if (ix->cols[i] == field) touches = 1; + if (!touches) continue; + db_ibucket *b = idx_bucket(ix, idx_hash(c, ix, r), 0); + if (b) + for (uint32_t i = 0; i < b->len; i++) + if (b->ids[i] == id) { + b->ids[i] = b->ids[--b->len]; + break; + } + } + r->slots[field] = nv; + for (uint32_t x = 0; x < t->index_cnt; x++) { + db_index *ix = &t->indexes[x]; + int touches = 0; + for (uint32_t i = 0; i < ix->col_cnt; i++) + if (ix->cols[i] == field) touches = 1; + if (!touches) continue; + db_ibucket *b = idx_bucket(ix, idx_hash(c, ix, r), 1); + if (b) { + if (b->len == b->cap) { + uint32_t ncap = b->cap ? b->cap * 2 : 4; + uint64_t *ni = realloc(b->ids, (size_t)ncap * 8u); + if (ni) { + b->ids = ni; + b->cap = ncap; + } + } + if (b->len < b->cap) b->ids[b->len++] = id; + } + } + db_val_free(c->kinds[field], old); + if (err_kind) *err_kind = DB_ERR_NONE; + return 0; +} + +int wo_row_has_referrers(wo_db *db, uint32_t class_id, uint64_t id) { + if (!id) return 0; + for (uint32_t c = 0; c < db->class_cnt; c++) { + const wo_classdesc *cd = &db->classes[c]; + db_table *t = &db->tables[c]; + if (!t->row_size || !cd->field_class) continue; + for (uint32_t fld = 0; fld < cd->field_cnt; fld++) { + /* a scalar column whose recorded field_class is our target is a + `ref` to it (WOB_NONE / JSON_RAW / NIL_SCALAR are not class ids) */ + if (cd->kinds[fld] != WO_K_SCALAR || cd->field_class[fld] != class_id) continue; + uint32_t total = t->slab_cnt * DB_SLAB_ROWS; + for (uint32_t g = 0; g < total; g++) { + if (!(t->bitmap[g >> 6] & (1ull << (g & 63)))) continue; + db_row *r = (db_row *)(t->slabs[g / DB_SLAB_ROWS] + + (size_t)(g % DB_SLAB_ROWS) * t->row_size); + if (r->slots[fld] == id) return 1; + } + } + } + return 0; +} + +int wo_row_remove(wo_db *db, uint32_t class_id, uint64_t id) { + if (class_id >= db->class_cnt) return -1; + db_table *t = &db->tables[class_id]; + if (!t->row_size) return -1; + uint64_t s1 = hget(t, id); + if (!s1) return -1; + uint32_t g = (uint32_t)(s1 - 1); + db_row *r = slot_row(t, g); + /* the index hook's remove side: before the row's values die, while the + columns are still comparable */ + idx_remove_row(db, t, r); + const wo_classdesc *c = &db->classes[class_id]; + for (uint32_t i = 0; i < c->field_cnt; i++) db_val_free(c->kinds[i], r->slots[i]); + t->bitmap[g >> 6] &= ~(1ull << (g & 63)); + hdel(t, id); + t->count--; + if (t->free_cnt == t->free_cap) { + uint32_t ncap = t->free_cap ? t->free_cap * 2 : 16; + uint32_t *nf = realloc(t->free_slots, (size_t)ncap * 4); + if (!nf) return 0; /* slot simply not recycled; bitmap still frees it */ + t->free_slots = nf; + t->free_cap = ncap; + } + t->free_slots[t->free_cnt++] = g; + return 0; +} diff --git a/database/src/table.h b/database/src/table.h new file mode 100644 index 0000000..cf0dd62 --- /dev/null +++ b/database/src/table.h @@ -0,0 +1,196 @@ +/* table.h — class-shaped row storage (iteration 9, Task 1). + * + * The engine and the VM heap are two memory worlds crossed only by copy + * (the 9b design's section 6): a row stores NO VM pointer. Every field + * lands in one 8-byte slot, kind-driven: + * + * SCALAR the 8 bytes themselves (WO_NIL_SCALAR spells a ?scalar's nil) + * TEXT engine-owned db_text* (0 = nil) + * OWNED engine-owned db_rec* — the object flattened by value, + * recursively, through these same rules (0 = nil) + * MULTI engine-owned db_multi* — elements encoded element-wise + * MAP engine-owned db_map* — keys and values encoded pair-wise + * GCREF never stored: the compiler rejects it (the GC bulkhead); + * the engine refuses it defensively as an encode error + * + * `ref T` is a SCALAR at this layer — the target row's id, an ordinary + * number the compiler produced; the engine learns nothing about it until + * the FK checks (9b plan, Task 3). + * + * Row layout: a 16-byte header (id, class, flags) then field_cnt 8-byte + * slots — deliberately the VM object layout's shape, so encode/decode walk + * the same class-table kinds the VM walks. Rows live in per-class SLABS + * (fixed-count, malloc'd, never moved: a row's address is stable for its + * lifetime, which is what lets 9b hand out loop-scoped row views). A + * per-table bitmap tracks occupancy; removed slots go on a free list and + * are reused before any slab grows. The id->row map is an open-addressing + * hash owned by the table. + * + * Id discipline (the c-runtime plan's shipped behavior): per table, per + * shard, ids interleave — shard S of N allocates S+1, S+1+N, S+1+2N, … — + * so creation is coordination-free and a row's owner shard is (id-1) % N. + * Milestone runs at N=1 (iteration 8 not yet landed); everything here is + * N-parametric and degenerates cleanly. + * + * CHOKE POINT DOCTRINE: wo_row_insert / wo_row_remove are the only paths + * that touch storage. Task 4's secondary indexes hook exactly these two + * functions; anything else mutating a slab is a defect by definition. + */ +#ifndef WO_TABLE_H +#define WO_TABLE_H + +#include "obj.h" /* wo_rt, wo_classdesc, kinds, wo_str, containers */ + +/* ---- engine-owned value shapes (all malloc'd, all reachable only from + * row slots, all freed through db_val_free) ---- */ + +typedef struct db_text { + uint32_t len; + char bytes[]; /* len bytes, no NUL */ +} db_text; + +typedef struct db_rec { /* an owned object flattened by value */ + uint32_t class_id; /* index into the SAME class table the VM uses */ + uint32_t _pad; + uint64_t slots[]; /* field_cnt slots, encoded by these rules */ +} db_rec; + +typedef struct db_multi { + uint8_t elem_kind; + uint32_t len; + uint64_t items[]; +} db_multi; + +typedef struct db_map { + uint8_t key_kind, val_kind; + uint32_t len; + uint64_t kv[]; /* len pairs: k0 v0 k1 v1 … */ +} db_map; + +/* ---- rows and tables ---- */ + +typedef struct db_row { + uint64_t id; + uint32_t class_id; + uint32_t flags; /* reserved (0) */ + uint64_t slots[]; +} db_row; + +#define DB_SLAB_ROWS 256u + +/* Secondary index (iteration 9, Task 4): built from the class table's v3 + * metadata at first touch, maintained ONLY inside the row choke points. + * Hash multimap: bucket per column-value hash, ids within; equality is + * re-checked against the actual rows on the unique path (a hash is a hint, + * never an answer). */ +typedef struct db_ibucket { + uint64_t hash; + uint64_t *ids; + uint32_t len, cap; +} db_ibucket; + +typedef struct db_index { + uint32_t flags; /* bit0 = unique */ + uint32_t col_cnt; + const uint32_t *cols; /* into the loader's idx pool */ + db_ibucket *buckets; /* open addressing by hash; hash==0 stored as 1 */ + size_t bcap, blen; +} db_index; + +/* wo_row_insert failure classes — *msg carries the sentence, this carries + * the machine-readable kind so db.c maps to the right trap. */ +enum { DB_ERR_NONE = 0, DB_ERR_OOM = 1, DB_ERR_BADKIND = 2, DB_ERR_UNIQUE = 3, DB_ERR_MISC = 4 }; + +typedef struct db_table { + uint32_t class_id; + size_t row_size; /* 16 + field_cnt * 8 */ + /* slabs of DB_SLAB_ROWS rows each; addresses stable forever */ + uint8_t **slabs; + uint32_t slab_cnt, slab_cap; + uint64_t *bitmap; /* one bit per slot, slab-major */ + /* removed slots, reused LIFO before any slab grows */ + uint32_t *free_slots; + uint32_t free_cnt, free_cap; + uint64_t next_id; /* next id THIS shard hands out for this table */ + uint64_t count; /* live rows */ + /* id -> (global slot + 1); 0 = empty. Open addressing, pow2. */ + uint64_t *hkeys; + uint64_t *hvals; + size_t hcap, hlen; + /* secondary indexes, from the class table's v3 metadata */ + db_index *indexes; + uint32_t index_cnt; +} db_table; + +typedef struct wo_db { + const wo_classdesc *classes; + uint32_t class_cnt; + uint32_t shard, nshards; /* S of N; ids interleave S+1, S+1+N, … */ + db_table *tables; /* class_cnt entries, created lazily on first insert */ +} wo_db; + +/* 0 ok, -1 alloc failure. nshards >= 1, shard < nshards. */ +int wo_db_init(wo_db *db, const wo_classdesc *classes, uint32_t class_cnt, + uint32_t shard, uint32_t nshards); +void wo_db_destroy(wo_db *db); + +/* Insert: encode field_cnt VM values (register words, kinds from the class + * table) into a fresh row. Returns the new id, or 0 with *msg set (OOM, or + * a GCREF field — which the compiler should have refused upstream). */ +uint64_t wo_row_insert(wo_db *db, uint32_t class_id, const uint64_t *vals, + const char **msg, int *err_kind); + +/* Read: decode the row's fields into VM values freshly allocated from + * [rt] — always copies, never a pointer into the slab (the out-gate). + * 0 ok, -1 no such row, -2 OOM (*msg set). */ +int wo_row_read(wo_db *db, wo_rt *rt, uint32_t class_id, uint64_t id, + uint64_t *out_vals, const char **msg); + +/* Remove: free the row's engine-owned field values, clear the slot, recycle + * it. 0 ok, -1 no such row. */ +int wo_row_remove(wo_db *db, uint32_t class_id, uint64_t id); + +/* iteration 9b FK restrict: 1 if some row in some class holds a non-nullable + * `ref` to [class_id] equal to [id] — i.e. deleting this row would dangle a + * reference. The compiler records a ref field's target class in the class + * table's field_class metadata; this scans those columns. Correctness-first + * (a full scan of referencing tables); the backlink index is the later + * optimization the spec records. */ +int wo_row_has_referrers(wo_db *db, uint32_t class_id, uint64_t id); + +/* Update one field in place (iteration 9 Task 5): encode the VM value, + * swap it into the slot, keep every index containing that column honest — + * remove-old/add-new with the unique re-check running BEFORE anything + * mutates, so a violating update leaves the row untouched. 0 ok, -1 no + * such row / bad field, DB_ERR_* codes via *err_kind like insert. */ +int wo_row_update_field(wo_db *db, uint32_t class_id, uint64_t id, uint32_t field, + uint64_t vm_val, const char **msg, int *err_kind); + +/* Borrowed row pointer for engine-internal callers (the WAL writes a row's + * encoded bytes; indexes read key slots). NULL = no such row. NEVER handed + * to the VM. */ +db_row *wo_row_ptr(wo_db *db, uint32_t class_id, uint64_t id); + +/* Engine-internal, for WAL replay only: create a row with a FIXED id, + * slots zeroed — the caller (wal.c) fills them with engine-encoded values + * it built while decoding. Advances the table's next_id past [id] when the + * id belongs to this shard, so post-replay inserts never collide. NULL = + * OOM or duplicate id (corruption beyond a torn tail). */ +db_row *wo_row_create_raw(wo_db *db, uint32_t class_id, uint64_t id); + +/* Engine-internal: free one engine-encoded slot value of [kind] (wal.c's + * decode error paths). */ +void wo_db_val_free(wo_db *db, uint8_t kind, uint64_t v); + +/* Decode one engine slot value to a FRESH VM value in [rt] (the out-gate: + * always a copy). The query builtins' field reads go through this. */ +uint64_t wo_val_decode_vm(wo_db *db, wo_rt *rt, uint8_t kind, uint64_t engine_val, + int *ok, const char **msg); + +/* Engine-internal, replay only: after wal.c fills a raw row's slots, this + * runs the index maintenance the normal insert runs inline — including the + * unique check, whose violation during replay is corruption, not data + * (0 ok, -1). */ +int wo_row_raw_commit(wo_db *db, uint32_t class_id, db_row *r); + +#endif /* WO_TABLE_H */ diff --git a/database/src/wal.c b/database/src/wal.c new file mode 100644 index 0000000..4f58161 --- /dev/null +++ b/database/src/wal.c @@ -0,0 +1,474 @@ +/* pread/pwrite/fdatasync/posix_fallocate under -std=c11 */ +#define _POSIX_C_SOURCE 200809L + +#include "wal.h" + +#include +#include +#include +#include +#include + +/* ---- crc32 (poly 0xEDB88320) — ported from runtime/wo-rt.c ------------- */ + +static uint32_t crc_table[256]; +static int crc_ready; + +static void crc32_init(void) { + for (uint32_t i = 0; i < 256; i++) { + uint32_t c = i; + for (int k = 0; k < 8; k++) c = (c & 1) ? 0xEDB88320u ^ (c >> 1) : c >> 1; + crc_table[i] = c; + } + crc_ready = 1; +} + +static uint32_t crc32(const void *buf, size_t len) { + if (!crc_ready) crc32_init(); + const uint8_t *p = buf; + uint32_t c = 0xFFFFFFFFu; + while (len--) c = crc_table[(c ^ *p++) & 0xFF] ^ (c >> 8); + return c ^ 0xFFFFFFFFu; +} + +/* ---- byte buffer -------------------------------------------------------- */ + +typedef struct { + uint8_t *b; + size_t len, cap; + int oom; +} wbuf; + +static void wput(wbuf *w, const void *p, size_t n) { + if (w->oom) return; + if (w->len + n > w->cap) { + size_t nc = w->cap ? w->cap * 2 : 256; + while (nc < w->len + n) nc *= 2; + uint8_t *nb = realloc(w->b, nc); + if (!nb) { + w->oom = 1; + return; + } + w->b = nb; + w->cap = nc; + } + memcpy(w->b + w->len, p, n); + w->len += n; +} + +static void wput_u8(wbuf *w, uint8_t v) { wput(w, &v, 1); } +static void wput_u32(wbuf *w, uint32_t v) { wput(w, &v, 4); } +static void wput_u64(wbuf *w, uint64_t v) { wput(w, &v, 8); } + +/* bounds-checked reader */ +typedef struct { + const uint8_t *p, *end; + int bad; +} rbuf; + +static int rtake(rbuf *r, void *out, size_t n) { + if (r->bad || (size_t)(r->end - r->p) < n) { + r->bad = 1; + return -1; + } + memcpy(out, r->p, n); + r->p += n; + return 0; +} + +static uint8_t rd_u8(rbuf *r) { + uint8_t v = 0; + rtake(r, &v, 1); + return v; +} +static uint32_t rd_u32(rbuf *r) { + uint32_t v = 0; + rtake(r, &v, 4); + return v; +} +static uint64_t rd_u64(rbuf *r) { + uint64_t v = 0; + rtake(r, &v, 8); + return v; +} + +#define WAL_NIL_TEXT 0xFFFFFFFFu + +/* ---- engine-value <-> bytes (kind-driven, mirrors table.c's encoding) --- */ + +static void enc_val(wbuf *w, const wo_classdesc *classes, uint8_t kind, uint64_t v) { + switch (kind) { + case WO_K_SCALAR: wput_u64(w, v); return; + case WO_K_TEXT: { + if (!v) { + wput_u32(w, WAL_NIL_TEXT); + return; + } + const db_text *t = (const db_text *)(uintptr_t)v; + wput_u32(w, t->len); + wput(w, t->bytes, t->len); + return; + } + case WO_K_OWNED: { + if (!v) { + wput_u8(w, 0); + return; + } + const db_rec *r = (const db_rec *)(uintptr_t)v; + wput_u8(w, 1); + wput_u32(w, r->class_id); + const wo_classdesc *c = &classes[r->class_id]; + for (uint32_t i = 0; i < c->field_cnt; i++) + enc_val(w, classes, c->kinds[i], r->slots[i]); + return; + } + case WO_K_MULTI: { + if (!v) { + wput_u8(w, 0); + return; + } + const db_multi *m = (const db_multi *)(uintptr_t)v; + wput_u8(w, 1); + wput_u8(w, m->elem_kind); + wput_u32(w, m->len); + for (uint32_t i = 0; i < m->len; i++) enc_val(w, classes, m->elem_kind, m->items[i]); + return; + } + case WO_K_MAP: { + if (!v) { + wput_u8(w, 0); + return; + } + const db_map *m = (const db_map *)(uintptr_t)v; + wput_u8(w, 1); + wput_u8(w, m->key_kind); + wput_u8(w, m->val_kind); + wput_u32(w, m->len); + for (uint32_t i = 0; i < m->len; i++) { + enc_val(w, classes, m->key_kind, m->kv[2 * i]); + enc_val(w, classes, m->val_kind, m->kv[2 * i + 1]); + } + return; + } + default: return; /* GCREF never stored, so never logged */ + } +} + +/* Decode one value into an engine-owned allocation. Returns 0 on success + * with *out set (0 = genuine nil); -1 on truncation/corruption/OOM — the + * caller frees what it already built. */ +static int dec_val(rbuf *r, wo_db *db, uint8_t kind, uint64_t *out) { + *out = 0; + switch (kind) { + case WO_K_SCALAR: { + uint64_t v = rd_u64(r); + if (r->bad) return -1; + *out = v; + return 0; + } + case WO_K_TEXT: { + uint32_t len = rd_u32(r); + if (r->bad) return -1; + if (len == WAL_NIL_TEXT) return 0; + if ((size_t)(r->end - r->p) < len) return -1; + db_text *t = malloc(sizeof(db_text) + len); + if (!t) return -1; + t->len = len; + memcpy(t->bytes, r->p, len); + r->p += len; + *out = (uint64_t)(uintptr_t)t; + return 0; + } + case WO_K_OWNED: { + uint8_t tag = rd_u8(r); + if (r->bad) return -1; + if (!tag) return 0; + uint32_t cid = rd_u32(r); + if (r->bad || cid >= db->class_cnt) return -1; + const wo_classdesc *c = &db->classes[cid]; + db_rec *rec = malloc(sizeof(db_rec) + (size_t)c->field_cnt * 8u); + if (!rec) return -1; + rec->class_id = cid; + rec->_pad = 0; + for (uint32_t i = 0; i < c->field_cnt; i++) { + if (dec_val(r, db, c->kinds[i], &rec->slots[i]) != 0) { + for (uint32_t j = 0; j < i; j++) wo_db_val_free(db, c->kinds[j], rec->slots[j]); + free(rec); + return -1; + } + } + *out = (uint64_t)(uintptr_t)rec; + return 0; + } + case WO_K_MULTI: { + uint8_t tag = rd_u8(r); + if (r->bad) return -1; + if (!tag) return 0; + uint8_t ek = rd_u8(r); + uint32_t len = rd_u32(r); + if (r->bad || ek > WO_K_MAX) return -1; + if (len > (size_t)(r->end - r->p)) return -1; /* each elem >= 1 byte */ + db_multi *m = malloc(sizeof(db_multi) + (size_t)len * 8u); + if (!m) return -1; + m->elem_kind = ek; + m->len = len; + for (uint32_t i = 0; i < len; i++) { + if (dec_val(r, db, ek, &m->items[i]) != 0) { + for (uint32_t j = 0; j < i; j++) wo_db_val_free(db, ek, m->items[j]); + free(m); + return -1; + } + } + *out = (uint64_t)(uintptr_t)m; + return 0; + } + case WO_K_MAP: { + uint8_t tag = rd_u8(r); + if (r->bad) return -1; + if (!tag) return 0; + uint8_t kk = rd_u8(r), vk = rd_u8(r); + uint32_t len = rd_u32(r); + if (r->bad || kk > WO_K_MAX || vk > WO_K_MAX) return -1; + if (len > (size_t)(r->end - r->p)) return -1; + db_map *m = malloc(sizeof(db_map) + (size_t)len * 16u); + if (!m) return -1; + m->key_kind = kk; + m->val_kind = vk; + m->len = len; + for (uint32_t i = 0; i < len; i++) { + if (dec_val(r, db, kk, &m->kv[2 * i]) != 0 || + dec_val(r, db, vk, &m->kv[2 * i + 1]) != 0) { + m->len = i; /* free only the fully-built pairs plus a possible key */ + for (uint32_t j = 0; j < i; j++) { + wo_db_val_free(db, kk, m->kv[2 * j]); + wo_db_val_free(db, vk, m->kv[2 * j + 1]); + } + wo_db_val_free(db, kk, m->kv[2 * i]); /* 0 if the key failed */ + free(m); + return -1; + } + } + *out = (uint64_t)(uintptr_t)m; + return 0; + } + default: return -1; + } +} + +/* ---- record scan (shared by open, replay, check) ------------------------ */ + +/* Read the record at [off]. 0 = intact (*len_out = payload length, payload + * malloc'd into *payload_out if non-NULL); 1 = end of intact prefix (zero + * length, short read, bad crc, missing mark). */ +static int scan_record(int fd, uint64_t off, uint32_t *len_out, uint8_t **payload_out) { + uint8_t hdr[8]; + ssize_t n = pread(fd, hdr, 8, (off_t)off); + if (n != 8) return 1; + uint32_t len, crc; + memcpy(&len, hdr, 4); + memcpy(&crc, hdr + 4, 4); + if (len == 0 || len > (64u << 20)) return 1; /* preallocated tail or garbage */ + uint8_t *payload = malloc(len + 4); + if (!payload) return 1; + n = pread(fd, payload, len + 4, (off_t)(off + 8)); + if (n != (ssize_t)(len + 4)) { + free(payload); + return 1; + } + uint32_t mark; + memcpy(&mark, payload + len, 4); + if (mark != WO_WAL_MARK || crc32(payload, len) != crc) { + free(payload); + return 1; + } + *len_out = len; + if (payload_out) *payload_out = payload; + else free(payload); + return 0; +} + +/* ---- public API ---------------------------------------------------------- */ + +int wo_wal_open(wo_wal *w, const char *path, uint64_t prealloc) { + memset(w, 0, sizeof(*w)); + w->fd = open(path, O_RDWR | O_CREAT, 0644); + if (w->fd < 0) return -1; + if (prealloc) { + /* best-effort: a filesystem without fallocate still works */ + (void)posix_fallocate(w->fd, 0, (off_t)prealloc); + } + /* position after the intact prefix: a torn tail is OVERWRITTEN by the + * next append, never appended after */ + uint64_t off = 0; + uint32_t len; + while (scan_record(w->fd, off, &len, NULL) == 0) off += 8u + len + 4u; + w->off = off; + return 0; +} + +void wo_wal_close(wo_wal *w) { + if (w->fd >= 0) close(w->fd); + free(w->buf); + memset(w, 0, sizeof(*w)); + w->fd = -1; +} + +/* frame one payload into the staged batch */ +static int stage(wo_wal *w, const wbuf *payload) { + if (payload->oom) return -1; + wbuf rec = {0}; + wput_u32(&rec, (uint32_t)payload->len); + wput_u32(&rec, crc32(payload->b, payload->len)); + wput(&rec, payload->b, payload->len); + wput_u32(&rec, WO_WAL_MARK); + if (rec.oom) { + free(rec.b); + return -1; + } + if (w->len + rec.len > w->cap) { + size_t nc = w->cap ? w->cap * 2 : 4096; + while (nc < w->len + rec.len) nc *= 2; + uint8_t *nb = realloc(w->buf, nc); + if (!nb) { + free(rec.b); + return -1; + } + w->buf = nb; + w->cap = nc; + } + memcpy(w->buf + w->len, rec.b, rec.len); + w->len += rec.len; + free(rec.b); + return 0; +} + +int wo_wal_append_insert(wo_wal *w, wo_db *db, uint32_t class_id, uint64_t id) { + db_row *r = wo_row_ptr(db, class_id, id); + if (!r) return -1; /* commit order: RAM apply comes FIRST */ + wbuf p = {0}; + wput_u8(&p, WO_WAL_INSERT); + wput_u32(&p, class_id); + wput_u64(&p, id); + const wo_classdesc *c = &db->classes[class_id]; + for (uint32_t i = 0; i < c->field_cnt; i++) enc_val(&p, db->classes, c->kinds[i], r->slots[i]); + int rc = stage(w, &p); + free(p.b); + return rc; +} + +int wo_wal_append_update(wo_wal *w, wo_db *db, uint32_t class_id, uint64_t id) { + db_row *r = wo_row_ptr(db, class_id, id); + if (!r) return -1; + wbuf p = {0}; + wput_u8(&p, WO_WAL_UPDATE); + wput_u32(&p, class_id); + wput_u64(&p, id); + const wo_classdesc *c = &db->classes[class_id]; + for (uint32_t i = 0; i < c->field_cnt; i++) enc_val(&p, db->classes, c->kinds[i], r->slots[i]); + int rc = stage(w, &p); + free(p.b); + return rc; +} + +int wo_wal_append_remove(wo_wal *w, uint32_t class_id, uint64_t id) { + wbuf p = {0}; + wput_u8(&p, WO_WAL_REMOVE); + wput_u32(&p, class_id); + wput_u64(&p, id); + int rc = stage(w, &p); + free(p.b); + return rc; +} + +int wo_wal_commit(wo_wal *w) { + if (!w->len) return 0; + size_t at = 0; + while (at < w->len) { + ssize_t n = pwrite(w->fd, w->buf + at, w->len - at, (off_t)(w->off + at)); + if (n < 0) { + if (errno == EINTR) continue; + return -1; + } + at += (size_t)n; + } + if (fdatasync(w->fd) != 0) return -1; + w->off += w->len; + w->len = 0; /* acked: the batch is durable */ + return 0; +} + +static int apply_record(wo_db *db, const uint8_t *payload, uint32_t len) { + rbuf r = {payload, payload + len, 0}; + uint8_t kind = rd_u8(&r); + uint32_t cid = rd_u32(&r); + uint64_t id = rd_u64(&r); + if (r.bad || cid >= db->class_cnt) return -1; + if (kind == WO_WAL_REMOVE) return wo_row_remove(db, cid, id); + if (kind != WO_WAL_INSERT && kind != WO_WAL_UPDATE) return -1; + if (kind == WO_WAL_UPDATE) { + /* replace: the row must exist (its insert precedes its update in a + correct log); anything else is corruption */ + if (wo_row_remove(db, cid, id) != 0) return -1; + } + db_row *row = wo_row_create_raw(db, cid, id); + if (!row) return -1; + const wo_classdesc *c = &db->classes[cid]; + for (uint32_t i = 0; i < c->field_cnt; i++) { + if (dec_val(&r, db, c->kinds[i], &row->slots[i]) != 0) { + /* a record that CRC-passed but does not decode is corruption, + * not a tear: fail loudly (the row's built slots are freed by + * wo_row_remove, which also unregisters the id) */ + wo_row_remove(db, cid, id); + return -1; + } + } + if ((size_t)(r.end - r.p) != 0) { /* trailing bytes = corrupt */ + wo_row_remove(db, cid, id); + return -1; + } + /* slots are real now: re-index (Task 4). A unique violation during + * replay is corruption — the live insert would have refused it. */ + if (wo_row_raw_commit(db, cid, row) != 0) { + wo_row_remove(db, cid, id); + return -1; + } + return 0; +} + +int64_t wo_wal_replay(const char *path, wo_db *db) { + int fd = open(path, O_RDONLY); + if (fd < 0) return errno == ENOENT ? 0 : -1; /* no WAL yet = fresh boot */ + uint64_t off = 0; + int64_t applied = 0; + for (;;) { + uint32_t len; + uint8_t *payload; + if (scan_record(fd, off, &len, &payload) != 0) break; /* intact prefix ends */ + int rc = apply_record(db, payload, len); + free(payload); + if (rc != 0) { + close(fd); + return -1; + } + off += 8u + len + 4u; + applied++; + } + close(fd); + return applied; +} + +int64_t wo_wal_check(const char *path, uint64_t *intact_bytes) { + int fd = open(path, O_RDONLY); + if (fd < 0) return -1; + uint64_t off = 0; + int64_t records = 0; + for (;;) { + uint32_t len; + if (scan_record(fd, off, &len, NULL) != 0) break; + off += 8u + len + 4u; + records++; + } + if (intact_bytes) *intact_bytes = off; + close(fd); + return records; +} diff --git a/database/src/wal.h b/database/src/wal.h new file mode 100644 index 0000000..4acbaf2 --- /dev/null +++ b/database/src/wal.h @@ -0,0 +1,91 @@ +/* wal.h — typed-row write-ahead log + boot replay (iteration 9, Task 2). + * + * The c-runtime plan's shipped pattern (phases D/E), generalized to typed + * rows. The commit order is doctrine, verbatim: + * + * RAM apply → wal_append (staged) → wal_commit (write + fdatasync) + * → only then is the write ACKNOWLEDGED + * + * Record framing — replay-whole-or-not-at-all: + * + * record := len u32 | crc u32 | payload | mark u32 + * len = payload byte count (never 0; 0 = preallocated tail, stop) + * crc = CRC32 of payload + * mark = 0x574F4C31 "WOL1" — written LAST, so a record without its + * mark is torn by definition + * payload := kind u8 | class_id u32 | row_id u64 | body + * kind : 1 insert (body = the row's fields, engine encoding below) + * 2 remove (no body) + * 3 update (reserved for Task 5) + * + * Field encoding in a body walks the class table's kinds: + * SCALAR 8 bytes + * TEXT u32 len | bytes (0xFFFFFFFF = nil) + * OWNED u8 0 = nil, or u8 1 | u32 class_id | fields recursively + * MULTI u8 0 = nil, or u8 1 | u8 elem_kind | u32 len | elements + * MAP u8 0 = nil, or u8 1 | u8 kk | u8 vk | u32 len | k v pairs + * + * Replay decodes payloads STRAIGHT into engine-owned values — the VM heap + * is never involved (boot must not depend on a VM existing yet), and rows + * re-enter through the same choke-point row API, so Task 4's indexes are + * rebuilt for free. A torn tail (short record, bad CRC, missing mark) drops + * everything from the tear onward — never a partial record, never a record + * after a tear. Little-endian on-disk, matching the .wob loader's platform + * note. + * + * wo_wal_check is the offline oracle the crash battery verifies with: it + * walks a WAL file with no engine at all and reports how many records are + * intact and where the intact prefix ends. */ +#ifndef WO_WAL_H +#define WO_WAL_H + +#include "table.h" + +#define WO_WAL_MARK 0x574F4C31u /* "WOL1" LE */ + +enum { WO_WAL_INSERT = 1, WO_WAL_REMOVE = 2, WO_WAL_UPDATE = 3 }; + +typedef struct wo_wal { + int fd; + uint64_t off; /* next write offset (the intact tail) */ + /* staged batch: appended by wal_append_*, flushed by wal_commit */ + uint8_t *buf; + size_t len, cap; +} wo_wal; + +/* Open (create if missing) and preallocate [prealloc] bytes (best-effort; + * a filesystem without fallocate still works). Positions the write offset + * at the end of the INTACT record prefix — an existing file is scanned the + * same way replay scans it, so a torn tail is overwritten, not appended + * after. 0 ok, -1 errno-style failure. */ +int wo_wal_open(wo_wal *w, const char *path, uint64_t prealloc); +void wo_wal_close(wo_wal *w); + +/* Stage a record for the row that MUST already be applied to RAM (the + * commit-order doctrine). Insert/update read the row via wo_row_ptr. + * 0 ok, -1 OOM / no such row. */ +int wo_wal_append_insert(wo_wal *w, wo_db *db, uint32_t class_id, uint64_t id); +int wo_wal_append_remove(wo_wal *w, uint32_t class_id, uint64_t id); +/* UPDATE re-logs the whole row (KISS: replay replaces — remove + re-create + * with the same id; the prefix/suffix delta trick from the survey is a + * later optimization, recorded). Call AFTER the RAM update. */ +int wo_wal_append_update(wo_wal *w, wo_db *db, uint32_t class_id, uint64_t id); + +/* Write the staged batch and fdatasync — the ack line. Empty batch = ok, + * no syscall. 0 ok, -1 write/sync failure (the batch stays staged). */ +int wo_wal_commit(wo_wal *w); + +/* Boot replay: apply every intact record to [db] in order. Ids re-enter + * exactly as logged; each table's next_id advances past the replayed ids + * that belong to this shard. Returns the number of records applied, or -1 + * on open failure / a record naming an unknown class (corruption beyond + * what a torn tail explains). A torn tail is NOT an error: replay applies + * the intact prefix and reports it. */ +int64_t wo_wal_replay(const char *path, wo_db *db); + +/* Offline verification (no engine): scan [path], count intact records. + * *intact_bytes (optional) = where the intact prefix ends. -1 = open + * failure. */ +int64_t wo_wal_check(const char *path, uint64_t *intact_bytes); + +#endif /* WO_WAL_H */ diff --git a/docs/00-status.md b/docs/00-status.md index 49c7d2d..5e41ca6 100644 --- a/docs/00-status.md +++ b/docs/00-status.md @@ -115,11 +115,18 @@ that sequences its tasks. Read one, approve, then the next starts. | 7 | [log-watcher proof](stories/language-runtime-database/07-logwatcher-proof.md) | 🔄 **runs; executable in progress** | | 7b | [Inferred GC + mark-sweep](stories/language-runtime-database/07b-inferred-gc-mark-sweep.md) | ⏸ off the workload's path (no `@gc`) | | 8 | [Shard-actor runtime](stories/language-runtime-database/08-shard-actor-runtime.md) | ⬜ | -| 9 | [Database engine](stories/language-runtime-database/09-database-engine.md) | ⬜ | -| 9b | [`@table`, relations, query](stories/language-runtime-database/09b-table-relations-query.md) | ⬜ needs a spec first | +| 9 | [Database engine](stories/language-runtime-database/09-database-engine.md) | 🔄 engine complete (storage/WAL/indexes/insert-update-delete); reads land with 9b | +| 9b | [`@table`, relations, query](stories/language-runtime-database/09b-table-relations-query.md) | 🔄 query surface + relations + FK done (branch query-surface); group-by parked | +| 9c | [Cross-program tables](stories/language-runtime-database/09c-cross-program-tables.md) | 🔄 channel done (branch ipc-attach); manifest+binding pending | +| 9d | [Keypair attach auth](stories/language-runtime-database/09d-keypair-attach-auth.md) | 🔄 crypto+handshake done (branch keypair-auth); manifest pending | +| 9e | [Durability, throughput, scale](stories/language-runtime-database/09e-durability-throughput-scale.md) | ⬜ needs a spec first | +| 9f | [io_uring group-commit](stories/language-runtime-database/09f-io-uring-commit.md) | ⬜ after 8 + 9e | +| 9g | [Query grammar corpus](stories/language-runtime-database/09g-query-grammar-corpus.md) | ⬜ needs a spec first | | 10 | [HTTP service layer](stories/language-runtime-database/10-http-service.md) | ⬜ | Hold | | 11 | [Fibers](stories/language-runtime-database/11-fibers.md) | ⬜ | Hold | | 12 | [Blue-green deploy](stories/language-runtime-database/12-blue-green-deploy.md) | ⬜ | Hold | +| 13 | [Compile-time metaprogramming](stories/language-runtime-database/13-compile-time-metaprogramming.md) | ⬜ needs a spec first | +| 14 | [skillhost host workload](stories/language-runtime-database/14-skillhost-host-workload.md) | ⬜ gaps recorded (branch query-grammar found skillhost needs no new query grammar); each gap a candidate iteration | --- @@ -322,7 +329,13 @@ Ecommerce sample (verified 2026-06-13): `api.rest` 17/17 expected statuses pass. | 7b | Inferred GC + incremental mark-sweep — `@gc` removed, GC-ness inferred, RC retired | [spec](superpowers/specs/2026-08-11-inferred-gc-mark-sweep-design.md) — plan to be written | | 8 | Shard-actor runtime | [plan 4](superpowers/plans/2026-08-01-shard-actor-vm-runtime.md) | | 9 | Database engine binding | [plan 5](superpowers/plans/2026-08-01-db-engine-binding.md) | -| 9b | `@table` + relations + language-integrated query | **no spec yet** — three open forks recorded in the iteration; brainstorm before planning | +| 9b | `@table` + relations + language-integrated query — comprehension queries, `ref`/`backlink` navigation, GroupBy aggregates; acceptance: new `docs/examples/employee` sample | [spec](superpowers/specs/2026-08-15-table-relations-query-design.md) · [plan](plan/compiler/2026-08-15-employee-relations-query.md) | +| 9c | Cross-program tables — attach to a running program's database (IPC string in wo.toml, manifest-granted rights, owner stays the single writer) | **no spec yet** — four open forks recorded in the iteration; brainstorm before planning | +| 9d | Keypair attach auth — mutual challenge–response, grants name public keys, uid superseded | **no spec yet** — four forks recorded; plan folds into 9c's | +| 9e | Durability + throughput + scale — restart-persistence, read/write benchmark, ~1M rows; the gate every later optimization re-runs | **no spec yet** — four forks recorded; the measurement backbone | +| 9f | io_uring group-commit write path — batched durability overlapped on shard threads, fsync fallback | **no spec yet** — brainstorm after iterations 8 + 9e | +| 9g | Query grammar from real embedded-DB corpora — whole-query count + correlated exists, driven by the skillhost SQL catalogue; add only what a corpus uses | **no spec yet** — three forks; may collapse to "confirm len(query) + add exists" | +| 14 | skillhost host workload — port skillhost (MCP host + confined script runner) to writeonce; drives the missing host capabilities into the open (bounded subprocess, stdin/stdout transport, fs metadata, FFI-vs-out-of-process) | **no spec yet** — gaps recorded in the iteration; each gap brainstormed on demand, bounded-subprocess first | | 10 | HTTP service layer | [plan 6](superpowers/plans/2026-08-01-http-service-layer.md) | | 11 | Fibers | vision §3, [blue-green exploration](plan/exploration/blue-green-vm/00-vision.md) | | 12 | Blue-green deploy | [spec](superpowers/specs/2026-08-03-blue-green-vm-design.md) — plan authored after iterations 9–10 | @@ -361,13 +374,14 @@ log-watcher proof. | ⬜ | 15a–15e MCP over streamable HTTP | [15](plan/15-mcp-streamable-http.md) | 15e needs 13c + 09d | | ⬜ | 16c–16f typed columns, lossless resync, restore, SCRAM | [16](plan/16-postgres-mirror.md) | | -### Frontend — parked +### Frontend — removed as stale (2026-08-17) -| Status | Phase | Doc | -| ------ | -------------------------------- | ------------------------------------------------------------------- | -| ⏸ | 13d pricing UI | [13](plan/13-class-model-live-pricing.md) | -| ⏸ | 14 MVC UI implementation (14a–f) | [14](plan/14-mvc-ui-implementation.md) | -| ⏸ | UI exploration track | [exploration/ui/00-overview.md](plan/exploration/ui/00-overview.md) | +The `##ui` / `.htmlx` LiveView frontend track — 13d pricing UI, the 14-MVC-UI +implementation plan, the 7-of-7 `ui-htmlx-live` plan, and the 9-doc +`plan/exploration/ui/` design set — was **removed**. It was built entirely on +the non-advancing Rust runtime (`.dev/reference/crates/wo-htmlx`, `cargo run`, +WebSocket live-patches) and contradicts the current woc/wovm direction. Recorded +in [`discarded.md`](plan/discarded.md). --- diff --git a/docs/01-problem.md b/docs/01-problem.md index 584e2a8..afee068 100644 --- a/docs/01-problem.md +++ b/docs/01-problem.md @@ -1,6 +1,6 @@ # Problem Statement -The current writeonce architecture works, but it carries weight that the project doesn't need. This document identifies the structural problems that motivate the redesign described in [02-recovery.md](./02-recovery.md). +The current writeonce architecture works, but it carries weight that the project doesn't need. This document identifies the structural problems that motivate the redesign described in 02-recovery.md. ## Too Many Moving Parts diff --git a/docs/02-recovery.md b/docs/02-recovery.md deleted file mode 100644 index fd7b770..0000000 --- a/docs/02-recovery.md +++ /dev/null @@ -1,172 +0,0 @@ -# Recovery — The Target Architecture - -This document describes where writeonce is going: a single, self-contained binary that owns its own storage, serves its own content, and pushes updates to connected clients in real-time — with no external database, no cloud pipeline, and no separate API server. - -## Guiding Principle - -**Everything in one process.** The database, the server logic, and the client-facing interface all live in a single codebase and ship as a single executable. If you can run the binary, you have the full platform. - -## Own Database - -The current PostgreSQL instance is a derived cache — it stores JSONB copies of files that already exist as the source of truth. The recovery architecture eliminates this indirection entirely. - -### What Changes - -- **No external database.** No PostgreSQL, no Diesel ORM, no connection pooling, no migrations. -- **Local file storage.** Markdown files and JSON metadata files are stored in a local directory, just as they are today in `writeonce-articles-s3/`. The file system *is* the database. -- **Custom storage segments (.seg files).** Research area: segment files that provide efficient read access, indexing, and potentially append-only writes for content. Think of these as a lightweight, purpose-built storage layer — not a general-purpose database engine, but enough to support indexed lookups by `blog-title` and ordered listing by date. -- **Indexed by blog-title.** The `sys_title` / blog-title field remains the primary key for content retrieval. The embedded storage must support O(1) or O(log n) lookups by this field. - -### What Stays the Same - -- Articles are still structured as JSON metadata + Markdown content pairs. -- The `sys_title`, `published`, `tags`, `author`, and section structure remain the content model. -- Content is still the source of truth — but now it's read directly from local storage instead of being derived through a sync pipeline. - -## No AWS Infrastructure - -The current architecture uses S3 as a file host and Lambda as a sync trigger. In the target architecture, there is nothing to sync *to* — the files are already where they need to be. - -### What Gets Removed - -| Current Component | Why It Existed | Why It's No Longer Needed | -|---|---|---| -| S3 bucket | Remote file storage | Files live locally alongside the binary | -| Lambda function (Go) | Watch S3 for changes, call API | No remote store to watch — file changes are local | -| aws-infra service (Rust) | Bridge to AWS S3/EC2 APIs | No AWS dependency | -| Pulumi IaC | Manage Lambda + S3 resources | No cloud resources to manage | - -### What Replaces It - -The binary watches its own content directory. When a file changes (new article, updated metadata), the embedded database re-indexes and notifies subscribers. The deployment model becomes: - -``` -1. Place the binary on a server -2. Point it at a content directory -3. It serves -``` - -No credentials, no IAM roles, no SDK configuration. - -## No Separate API - -Today, `writeonce-api` is a standalone Actix-web server that mediates between the frontend and the database. In the target architecture, the server logic is embedded in the same process as the database and the content renderer. - -### What This Means - -- **No HTTP hop between database and server.** Queries go directly from the request handler to the storage engine in-process. No network serialization, no connection pool, no ORM layer. -- **Single codebase.** No multi-repo coordination. A new article field is added once — in the content model — and it flows through storage, indexing, and rendering in the same compilation unit. -- **Single deployment.** One binary, one container, one process. No docker-compose orchestrating API + database + infra services. - -The binary still exposes HTTP endpoints — it's still a web server. But it's a web server with an embedded database, not a web server that talks to an external one. - -## Real-Time Subscriptions Without WebSocket - -The current architecture has no mechanism for pushing content updates to connected clients. The target architecture adds real-time subscriptions, but explicitly without WebSocket. - -### Why Not WebSocket - -WebSocket adds connection state management, heartbeat logic, reconnection handling, and protocol upgrade complexity. For a content platform where updates are infrequent (articles are published, not streamed), the overhead isn't justified. - -### Subscription Model - -The target is a subscription mechanism where: - -- A client subscribes to a content query (e.g., "all published articles" or "article with sys_title X") -- When the underlying data changes, the server pushes the relevant diff to the subscriber -- No polling from the client side - -Candidate approaches to research: - -- **Server-Sent Events (SSE)** — unidirectional push over HTTP. Simple, well-supported, no protocol upgrade. Natural fit for infrequent content updates. -- **SpacetimeDB-style subscriptions** — clients register queries, the engine tracks which rows match, and only sends diffs when the result set changes. This is the aspirational model. -- **Long polling** — fallback option. Simple but less efficient than SSE for multiple subscribers. - -The key constraint: the subscription mechanism must work without requiring clients to maintain persistent bidirectional connections. - -## Target Architecture - -``` - content directory - (JSON + MD files, .seg index) - | - | file watch + re-index - v - +---------------------------+ - | writeonce binary | - | | - | +-------------------+ | - | | embedded storage | | .seg files, blog-title index - | | (read/write/index)| | - | +-------------------+ | - | | | - | +-------------------+ | - | | server logic | | route handlers, content queries - | | (HTTP endpoints) | | - | +-------------------+ | - | | | - | +-------------------+ | - | | subscription mgr | | SSE / query-based push - | | (real-time push) | | - | +-------------------+ | - | | - +---------------------------+ - | - HTTP / SSE - | - v - +-------------------+ - | frontend app | Angular or successor - | (browser client) | - +-------------------+ -``` - -## Single Repository - -The five current repos collapse into one: - -``` -writeonce/ - content/ # articles (JSON + MD), images, assets - storage/ # embedded database engine (.seg files, indexing) - server/ # HTTP handlers, subscription manager - frontend/ # client application - writeonce.toml # configuration (port, content dir, index settings) -``` - -One repo. One build. One deploy artifact. - -## What Needs Research - -| Area | Question | Notes | -|------|----------|-------| -| **.seg file format** | What storage format gives efficient indexed reads over JSON+MD content? | Look at LSM trees, append-only logs, SQLite's page format for inspiration | -| **File watching** | How to efficiently detect content changes on Linux/macOS? | `inotify` on Linux, `kqueue` on macOS, or cross-platform via `notify` crate | -| **SSE vs alternatives** | Is SSE sufficient for the subscription model, or is something custom needed? | SSE handles the "push diffs to subscribers" case well for low-frequency updates | -| **Index structure** | What index structure supports `blog-title` lookup + date-ordered listing? | B-tree or hash index for title, sorted set for date ordering | -| **Language choice** | Continue with Rust for the unified binary? | Rust fits: single binary output, no runtime, strong typing, existing team knowledge | -| **Frontend coupling** | Should the frontend be embedded in the binary (serve static assets) or remain separate? | Embedding simplifies deployment; separate allows independent frontend iteration | - -## Migration Path - -The transition from current to target doesn't have to be all-or-nothing: - -1. **Phase 1** — Build the embedded storage engine. Read JSON+MD files from a local directory, index by `blog-title`, serve via HTTP. No AWS, no PostgreSQL. This alone replaces `writeonce-api` + `aws-infra` + `lambda-function` + PostgreSQL. -2. **Phase 2** — Add real-time subscriptions (SSE). Clients subscribe to content queries and receive push updates when files change. -3. **Phase 3** — Collapse repositories. Move frontend into the unified codebase. Ship as a single binary that serves both API and static assets. - -Each phase produces a working system. The current architecture can run in parallel until the new one is ready. - -## Implementation phases - -The "embedded storage engine" of Phase 1 above lands in three numbered plan docs under [`docs/plan/`](./plan/): - -| Phase | Doc | What it ships | -| --- | --- | --- | -| 10 | [`plan/10-storage-foundations.md`](./plan/10-storage-foundations.md) | On-disk row codec (length-prefix + flags + LSN + CRC32C); per-type segment files (`data/.seg`); `posix_fallocate` preallocation; `pwrite`-only append path. Reads still in-memory. | -| 11 | [`plan/11-wal-and-recovery.md`](./plan/11-wal-and-recovery.md) | WAL log with `fdatasync` at commit; group commit per loop tick; control file with `last_durable_lsn` (rename-on-write); replay loop on startup. `kill -9` mid-write loses nothing acknowledged. | -| 12 | [`plan/12-engine-disk-cutover.md`](./plan/12-engine-disk-cutover.md) | `Engine`'s row payload moves to disk; in-memory map becomes `BTreeMap`. Periodic checkpoint flushes segments + advances the control file. RAM bounded by id-count, not row size. | - -Postgres' storage subsystem is the design reference — see [`docs/plan/exploration/postgresql/`](./plan/exploration/postgresql/) for which Postgres modules informed which decision and what writeonce skips (multi-process IPC, latches, separate writer processes). - -The durability syscalls themselves live in [`docs/plan/exploration/linux/12-pwrite-fsync.md`](./plan/exploration/linux/12-pwrite-fsync.md). diff --git a/docs/03-data.md b/docs/03-data.md deleted file mode 100644 index d444392..0000000 --- a/docs/03-data.md +++ /dev/null @@ -1,184 +0,0 @@ -# Data Layer — Local Storage with Subscriptions - -This document describes the embedded data layer that replaces PostgreSQL: local `.seg` files with indexing, and a subscription model where clients register queries and receive diffs on route visit — no polling required. - -## .seg File Storage - -The `.seg` (segment) format is the on-disk representation of article data. Each segment file holds serialized article content with positional indexing for fast lookups. - -### Design Goals - -- **No external database process.** The binary reads and writes `.seg` files directly. No socket connections, no protocol negotiation, no separate daemon. -- **Indexed by blog-title.** The primary access pattern is `GET /blog/:sys_title`. The storage layer must resolve a `sys_title` to its article content without scanning all files. -- **Append-friendly.** New articles and updates append to the segment. Deletes are tombstoned and compacted later. -- **Human-readable source.** The JSON + Markdown files remain the authoring format. `.seg` files are a derived index — if they're deleted, they can be rebuilt from the content directory. - -### Proposed Structure - -``` -content/ - linux-misc/ - linux-misc.json # authored metadata (source of truth) - linux-misc.md # authored content (source of truth) - aws-lambda-pulumi/ - aws-lambda-pulumi.json - aws-lambda-pulumi.md - -data/ - articles.seg # serialized article records - index/ - title.idx # blog-title -> offset mapping - date.idx # publish date -> offset (sorted) - tags.idx # tag -> [offsets] (inverted index) -``` - -The `content/` directory is what the author edits. The `data/` directory is what the engine builds and queries. Losing `data/` is a cold start, not data loss. - -### Segment File Internals - -``` -+------------------+ -| segment header | magic bytes, version, record count -+------------------+ -| record 0 | length-prefixed serialized article -+------------------+ -| record 1 | -+------------------+ -| ... | -+------------------+ -| record N | -+------------------+ -``` - -Each record is a length-prefixed byte sequence containing the full article (metadata + content merged). Records are addressed by byte offset from the start of the file. - -### Index Files - -**title.idx** — Hash map serialized to disk. Maps `sys_title` (string) to byte offset in `articles.seg`. Loaded into memory at startup for O(1) lookups. - -**date.idx** — Sorted array of `(timestamp, offset)` pairs. Supports range queries for "articles published between X and Y" and ordered listing for the homepage. - -**tags.idx** — Inverted index. Maps each tag string to a list of offsets. Supports "all articles tagged with X" queries. - -On startup, index files are memory-mapped or loaded into heap. On content change, affected indexes are rebuilt incrementally. - -## Subscription Model - -The subscription model is inspired by SpacetimeDB: clients register queries, and the engine tracks which results match. When underlying data changes, only the relevant diffs are pushed to subscribers. - -### How It Works - -``` - Client A Server Content Dir - | | | - |--- GET /blog/linux-misc -| | - | |-- read from .seg index ---->| - |<-- article + SSE stream -| | - | | | - | (subscribed to | | - | sys_title=linux-misc) | | - | | | - | |<-- file change detected ----| - | | | - | |-- re-index article -------->| - | |-- diff against last push -->| - | | | - |<-- SSE: updated content -| | - | | | -``` - -### Route-Based Subscription - -When a user visits a route, the response includes both the current content and an SSE stream. The client is automatically subscribed to changes for that query — no explicit subscription handshake needed. - -``` -GET /blog/linux-misc -``` - -Response: -``` -HTTP/1.1 200 OK -Content-Type: text/html - - - - - -``` - -The subscription lives as long as the browser tab is open. When the user navigates away, the EventSource closes and the server drops the subscription. No heartbeat management, no reconnection logic beyond what SSE provides natively (automatic reconnect is built into the EventSource API). - -### Query Registration - -Subscriptions are not limited to single-article lookups. The engine supports registering arbitrary content queries: - -| Query Type | Example | Subscription Behavior | -|---|---|---| -| Single article | `sys_title = "linux-misc"` | Push when this specific article changes | -| All published | `published = true` | Push when any article is published or unpublished | -| By tag | `tags contains "rust"` | Push when a rust-tagged article is added, removed, or updated | -| Homepage list | `published = true ORDER BY date DESC LIMIT 10` | Push when the top-10 list changes | - -The server maintains a registry of active subscriptions. On each content change, it evaluates which subscriptions are affected and pushes diffs only to those clients. - -### Diff Format - -When content changes, the server doesn't resend the full article. It sends a minimal diff: - -```json -{ - "type": "update", - "sys_title": "linux-misc", - "changes": { - "content.sections[2].paragraphs[0]": "Updated paragraph text...", - "content.tags": ["linux", "kernel", "new-tag"] - }, - "version": 42 -} -``` - -The `version` field enables clients to detect missed updates and request a full resync if needed. - -## Sample Dataset - -To validate the storage engine and subscription model, a sample dataset should exercise the core access patterns: - -### Articles - -| sys_title | tags | published | purpose | -|---|---|---|---| -| `sample-getting-started` | `[tutorial, beginner]` | true | Basic article, tests single-article subscription | -| `sample-rust-patterns` | `[rust, patterns]` | true | Tests tag-based queries | -| `sample-draft-wip` | `[draft]` | false | Tests published filter — should not appear in public queries | -| `sample-long-form` | `[deep-dive, rust]` | true | Multiple sections, images, code snippets — tests complex content rendering | -| `sample-frequently-updated` | `[changelog]` | true | Updated often — tests subscription diff delivery | - -### Test Scenarios - -1. **Cold start** — Delete `data/`, start the binary. It should rebuild `.seg` and index files from `content/` and serve all articles. -2. **Single article query** — `GET /blog/sample-getting-started` returns the article and opens an SSE subscription. -3. **Live update** — Edit `sample-frequently-updated.json` while a client is subscribed. The client should receive an SSE event with the diff. -4. **Tag query** — Subscribe to `tags contains "rust"`. Both `sample-rust-patterns` and `sample-long-form` should be in the result set. Adding a new article tagged `rust` should trigger a push. -5. **Publish toggle** — Change `sample-draft-wip` from `published: false` to `true`. Clients subscribed to the homepage list should receive a push with the new article added. - -## SpacetimeDB Reference - -SpacetimeDB is the primary architectural inspiration for the subscription model. Key concepts to study: - -- **Modules** — server logic that runs inside the database, not beside it -- **Subscription queries** — clients register SQL-like queries; the engine evaluates them incrementally on each transaction -- **Incremental view maintenance** — only recompute the parts of a query result that changed -- **Client SDK generation** — type-safe client code generated from the server schema - -Add SpacetimeDB as a reference submodule for quick access to their implementation patterns: - -```bash -git submodule add https://github.com/clockworklabs/SpacetimeDB.git references/spacetimedb -``` - -The goal is not to replicate SpacetimeDB — it's to take its subscription semantics and apply them to a much narrower domain (blog content), where the simplicity of the problem allows a simpler implementation. diff --git a/docs/04-ui.md b/docs/04-ui.md deleted file mode 100644 index 74dfd6f..0000000 --- a/docs/04-ui.md +++ /dev/null @@ -1,216 +0,0 @@ -# User Interface — Server-Rendered HTMLX - -No Angular. No React. No frontend framework. The UI is a set of `.htmlx` template files that the server parses, populates with content from the embedded database, and serves as plain HTML. Real-time updates arrive via SSE and are applied with minimal client-side scripting. - -## Why Not Angular - -The current `writeonce-app` is an Angular 18 SPA with Tailwind, PrismJS, ngx-markdown, and FontAwesome. It works, but it's a heavy delivery mechanism for what is fundamentally a read-heavy content site: - -- **~200MB of `node_modules`** for a site that renders markdown articles -- **Client-side routing** for content that doesn't need it — every article is a distinct URL, not an interactive application -- **JavaScript-dependent rendering** — content doesn't exist until Angular boots, hydrates, and fetches from the API -- **Separate build pipeline** — `npm run build` produces static assets that must be deployed to nginx independently of the API - -The content is static between updates. The interactivity is limited to navigation and code highlighting. A server-rendered approach matches the actual requirements. - -## HTMLX Templates - -The author defines the site layout using `.htmlx` files — HTML with embedded data bindings that the server resolves at render time. - -### Template Structure - -``` -templates/ - layout.htmlx # outer shell: , , - header.htmlx # site header, navigation - footer.htmlx # site footer - home.htmlx # homepage: article list - article.htmlx # single article view - about.htmlx # static page - contact.htmlx # static page - components/ - article-card.htmlx # summary card for article listings - code-snippet.htmlx # code block with language + title - img-caption.htmlx # image with caption - section.htmlx # article section (heading + paragraphs) -``` - -### Template Syntax - -Templates use a binding syntax that references content from the database. The server parses these bindings, resolves them against the current content, and outputs plain HTML. - -```html - -
- -
-``` - -```html - -
-

{{article.title}}

-

by {{article.author}} · {{article.tags}}

- - {{#each article.sections}} -
-

{{heading}}

- {{#each paragraphs}} -

{{this}}

- {{/each}} -
- {{/each}} - - {{#each article.codes}} - {{> code-snippet snippet=this}} - {{/each}} - - {{#each article.images}} - {{> img-caption image=this}} - {{/each}} -
-``` - -```html - -
-

articles

- {{#each articles}} - {{> article-card article=this}} - {{/each}} -
-``` - -The `{{> partial}}` syntax includes another `.htmlx` file as a component. The server resolves these at render time — no client-side component tree. - -### Content Subscription in Templates - -Templates declare what data they need. The server resolves these declarations against the embedded database and subscribes the client to changes: - -```html - - - -
-

{{article.title}}

- ... -
-``` - -```html - - - -
- {{#each articles}} - {{> article-card article=this}} - {{/each}} -
-``` - -The `` comment is a directive to the server. It declares the query that populates the template's data context. The same query is used to register an SSE subscription for live updates (as described in [03-data.md](./03-data.md)). - -## Rendering Pipeline - -``` - Browser request - | - v - Route match (/blog/linux-misc) - | - v - Load template (article.htmlx) - | - v - Parse subscribe directive - (article WHERE sys_title = "linux-misc") - | - v - Query embedded database (.seg index) - | - v - Resolve template bindings ({{article.title}}, etc.) - | - v - Compose with layout.htmlx + header.htmlx + footer.htmlx - | - v - Inject SSE subscription script - | - v - Send complete HTML response -``` - -The browser receives a fully rendered page on first load. No JavaScript framework boots. No API call fires. The content is already in the HTML. - -## Live Updates via SSE - -After the initial HTML is delivered, a small inline script opens an SSE connection for the page's subscription query: - -```html - -``` - -The `applyDiff` function is a lightweight client-side updater — it targets DOM elements by data attribute and patches their content. No virtual DOM, no reconciliation, no framework. For a content site where updates are infrequent and localized (a paragraph changed, a tag was added), direct DOM manipulation is sufficient. - -```html -

Linux Misc

-

First paragraph...

-``` - -When a diff arrives for `article.title`, the script finds the element with `data-bind="article.title"` and replaces its text content. This is the minimal client-side code the architecture requires. - -## Code Highlighting - -The current frontend uses PrismJS for syntax highlighting. In the server-rendered model, highlighting can happen at either layer: - -**Server-side (preferred):** The server parses code blocks during template rendering and emits pre-highlighted HTML with CSS classes. The browser only needs the PrismJS CSS theme, not the JavaScript library. This eliminates client-side parsing entirely. - -**Client-side (fallback):** Include PrismJS as a small script that runs on page load and on SSE update. Simpler to implement initially but adds a JavaScript dependency. - -## Markdown Rendering - -The current frontend uses `ngx-markdown` and `marked` to parse markdown in the browser. In the target architecture, markdown is rendered to HTML on the server during template composition. The browser never sees raw markdown. - -This aligns with the content model: the JSON metadata already defines the article structure (sections, paragraphs, code snippets, images). The markdown file provides prose content. The server combines both into final HTML — the template just places the pre-rendered blocks. - -## What Gets Removed - -| Current (Angular) | Target (HTMLX) | -|---|---| -| `writeonce-app/` (full Angular project) | `templates/` (handful of .htmlx files) | -| `node_modules/` (~200MB) | None | -| `angular.json`, `tsconfig.json`, `karma.conf.js` | None | -| npm build pipeline | Template parsed at request time | -| Nginx static file serving | Binary serves its own HTML | -| Client-side routing | Server-side route matching | -| Client-side markdown parsing | Server-side rendering | -| Client-side code highlighting | Server-side or minimal JS | - -## Styling - -Templates use plain CSS. Tailwind can optionally be used as a build-time utility (generating a static CSS file), but there is no runtime CSS framework. The author writes styles in a `styles.css` file that the server serves as a static asset. - -``` -templates/ - styles/ - main.css # site-wide styles - article.css # article-specific styles - code-theme.css # syntax highlighting theme (PrismJS compatible) -``` - -## Template Authoring Experience - -The `.htmlx` files are editable by the same author who writes articles. The template syntax is intentionally close to HTML — there's no JSX, no TypeScript, no build step. An author who knows HTML can modify the site layout. - -This closes the loop on the writeonce philosophy: the author writes content (markdown + JSON) and layout (`.htmlx` + CSS) as files, and the binary turns them into a live site. diff --git a/docs/05-datalayer.md b/docs/05-datalayer.md deleted file mode 100644 index 97150e3..0000000 --- a/docs/05-datalayer.md +++ /dev/null @@ -1,139 +0,0 @@ -# Data Layer — Implementation Status - -The embedded data layer described in [02-recovery.md](./02-recovery.md) and [03-data.md](./03-data.md) has been implemented as a Cargo workspace with 8 crates. All 44 tests pass. No external database, no AWS, no tokio — direct Linux syscalls on a custom event loop. - -## Workspace Structure - -``` -writeonce-all/ - Cargo.toml # workspace root - docs/ # architecture documentation - sample-content/ # 5 test articles for validation - crates/ - wo-model/ # content model - wo-seg/ # .seg file format - wo-index/ # index files - wo-store/ # unified storage engine - wo-watch/ # inotify file watcher - wo-event/ # epoll event loop - wo-sub/ # subscription system - wo-rt/ # custom runtime -``` - -## Crate Summary - -| Crate | Purpose | Tests | Key Types | -|-------|---------|-------|-----------| -| **wo-model** | Article structs matching existing JSON schema, `ContentLoader` for directory walking | 8 | `Article`, `ArticleContent`, `ArticleBody`, `Section`, `CodeSnippet`, `ContentLoader` | -| **wo-seg** | Binary `.seg` file format — length-prefixed records, tombstoning, positional I/O | 6 | `SegWriter`, `SegReader`, `SegHeader` | -| **wo-index** | Three index types for O(1) and O(log n) access patterns | 8 | `TitleIndex`, `DateIndex`, `TagIndex` | -| **wo-store** | Unified storage engine composing seg + indexes, cold-start rebuild | 3 | `Store` | -| **wo-watch** | Content directory watcher using inotify | 4 | `ContentWatcher`, `ContentChange` | -| **wo-event** | Custom event loop on epoll with eventfd, timerfd, signalfd | 5 | `EventLoop`, `EventFd`, `TimerFd`, `SignalFd` | -| **wo-sub** | Subscription manager with fd-based notifications, `register!` macro | 6 | `SubscriptionManager`, `Subscription`, `Notification` | -| **wo-rt** | Runtime tying all crates together — single process, single event loop | 4 | `Runtime`, `RuntimeHandle`, `Config` | - -## Linux Kernel Syscalls Used - -| Syscall | Crate | Purpose | -|---------|-------|---------| -| `pread` / `pwrite` | wo-seg | Positional read/write for .seg records without seeking | -| `fallocate` | wo-seg | Pre-allocate .seg file space to reduce fragmentation | -| `epoll_create1` / `epoll_ctl` / `epoll_wait` | wo-event | Event-driven I/O multiplexing for the main loop | -| `eventfd` | wo-event, wo-sub | Lightweight signaling between watcher and subscription manager | -| `timerfd_create` / `timerfd_settime` | wo-event | Periodic tasks (compaction, keepalive) as file descriptors | -| `signalfd` | wo-event | SIGINT/SIGTERM delivered as fd events for graceful shutdown | -| `inotify_init1` / `inotify_add_watch` | wo-watch | File system change detection on the content directory | -| `pipe2` | wo-sub (tests) | Mock subscriber fds for testing notification delivery | - -## .seg File Format - -``` -Offset Size Field -0 4 Magic: b"WOSF" -4 2 Version: u16 LE (1) -6 2 Flags: u16 LE (reserved) -8 8 Record count: u64 LE -16 8 Data start offset: u64 LE -24 8 Reserved -32+ variable Records: [u32 length][u8 flags][bincode payload]... -``` - -- Records are addressed by byte offset from file start -- Flags: `0x00` = active, `0x01` = tombstoned -- Payload: bincode-serialized `Article` struct - -## Index Files - -| File | Format | Access Pattern | -|------|--------|----------------| -| `title.idx` | On-disk hash table (Robin Hood, load factor 0.5), 138 bytes/slot | O(1) lookup by `sys_title` | -| `date.idx` | Sorted `(i64 timestamp, u64 offset)` array, 16 bytes/entry | Binary search for date ranges, latest N | -| `tags.idx` | Bincode-serialized `HashMap>` | Tag-to-offsets inverted index | - -All indexes are derived from `.seg` and rebuildable from `content/` on cold start. - -## Subscription Model - -No SSE. No WebSocket. Notifications are written directly to subscriber file descriptors. - -- **Subscribe**: `SubscriptionManager::subscribe(fd, Subscription::ByTitle("linux-misc"))` -- **Notify**: on content change, length-prefixed `Notification` written to matching fds -- **Cleanup**: `EPOLLHUP` on epoll triggers automatic `unsubscribe(fd)` -- **Dedup**: if a fd matches multiple patterns (title + tag), it receives only one notification - -Subscription patterns: -- `Subscription::ByTitle(sys_title)` — single article -- `Subscription::ByTag(tag)` — all articles with tag -- `Subscription::All` — all content changes - -## Store Query API - -```rust -store.get_by_title("linux-misc") -> Option
-store.list_published(skip, limit) -> Vec
-store.list_by_tag("rust") -> Vec
-store.list_by_date_range(start, end) -> Vec
-store.count_published() -> usize -store.article_version("linux-misc") -> Option -store.rebuild() // full rebuild from content/ -``` - -## Runtime Event Loop - -Single `epoll` instance multiplexing all file descriptors: - -| Token | Fd | Handler | -|-------|----|---------| -| `WATCHER` | inotify fd | Process file changes → update store → notify subscribers | -| `SIGNAL` | signalfd | SIGINT/SIGTERM → graceful shutdown | -| `TIMER` | timerfd | Periodic tasks (compaction, stats) | -| `NOTIFY` | eventfd | Subscription notification signal | -| `1000+` | subscriber fds | Hangup detection → unsubscribe + cleanup | - -## External Dependencies - -| Crate | Version | Purpose | -|-------|---------|---------| -| `serde` | 1.x | Serialization derives | -| `serde_json` | 1.x | JSON parsing for article files | -| `bincode` | 1.x | Compact binary serialization for .seg records and notifications | -| `libc` | 0.2.x | Raw Linux syscall bindings | - -No tokio. No async-std. No database driver. No HTTP framework (yet). - -## What Comes Next - -The data layer delivers everything the HTTP server and UI layers need: - -1. **`Store` with zero-copy query access** — all article queries resolve in-process -2. **Subscription system accepting raw fds** — HTTP layer hands socket fds to `subscribe()` -3. **Shared event loop** — HTTP listener socket registers on the same epoll -4. **Automatic cold-start** — if `data/` is missing, rebuilds from `content/` on startup -5. **Graceful shutdown** — SIGTERM triggers clean fd cleanup - -Next phases per [02-recovery.md](./02-recovery.md): -- **HTTP server** — route handlers using the `Store` query API, embedded in the same binary -- **HTMLX templates** — server-rendered HTML with `{{bindings}}` per [04-ui.md](./04-ui.md) -- **Frontend collapse** — serve static assets from the binary, eliminate the Angular app -3 \ No newline at end of file diff --git a/docs/06-markdown-render.md b/docs/06-markdown-render.md deleted file mode 100644 index 224b84f..0000000 --- a/docs/06-markdown-render.md +++ /dev/null @@ -1,246 +0,0 @@ -# Markdown File Rendering - -## Current State (writeonce-articles-s3) - -Each article is a directory containing a JSON metadata file and one or more `.md` files: - -``` -auto-scale-gitlab-runner-using-aws-spot-instance/ - docker-machine-test-with-t2.md - gitlab-runner-config.md - stop-test-gitlab-docker-machine.md - -gitlab-runner-with-kubernetes-executor/ - gitlab-runner-with-kubernetes-executor.json - deploy.md - permission.md - role-binding.md - role-defination.md - gitlab-runnergitlab-runner-deploy.md -``` - -The JSON metadata currently defines the full article structure — sections, headings, paragraphs, and code snippet references. Markdown files are limited to code blocks referenced via the `codes[].snippet` field. - -## Problem - -The JSON metadata carries too much content. Headings, paragraphs, prose — all of this is duplicated as JSON strings inside `content.content.sections`. The markdown files only hold code snippets, referenced by `sectionIndex` and `paragraphIndex`. - -This is backwards. The markdown file should be the content. The JSON should be minimal metadata. - -## Target: Markdown-First Content Model - -**The markdown file is the article.** All prose, headings, code blocks, and inline formatting live in the `.md` file. The JSON metadata file holds only what markdown cannot express: system fields, tags, publication state, and author. - -### Minimal JSON Metadata - -```json -{ - "sys_title": "gitlab-runner-with-kubernetes-executor", - "title": "Gitlab Runner with Kubernetes Executor", - "published": true, - "author": "Shoney Arickathil", - "tags": ["kubernetes", "gitlab", "ci-cd"], - "published_on": 1740950884 -} -``` - -No `content.content.sections`. No `content.content.codes`. No `paragraphs[]` arrays. No `sectionIndex`/`paragraphIndex` mapping. - -### Markdown File = Full Article Content - -````markdown -# Introduction - -Deploying a Gitlab runner using kubernetes is a great option to overcome -the limitations of other gitlab runner executor such as docker and docker machine. - -## Running Gitlab Runner in gitlab namespace - -Create the namespace and apply the deployment: - -```yaml -apiVersion: apps/v1 -kind: Deployment -metadata: - name: gitlab-runner - namespace: gitlab -``` -```` - -## Permissions - -The runner needs RBAC permissions to create pods: - -```yaml -apiVersion: rbac.authorization.k8s.io/v1 -kind: Role -metadata: - name: gitlab-runner -``` - -Everything is in the markdown — headings, paragraphs, code blocks with language hints, links, images. The rendering pipeline parses the markdown directly. - -### Directory Structure - -``` -content/ - gitlab-runner-with-kubernetes-executor/ - gitlab-runner-with-kubernetes-executor.json # minimal metadata - gitlab-runner-with-kubernetes-executor.md # full article content - linux-misc/ - linux-misc.json - linux-misc.md -``` - -One JSON for metadata. One markdown for content. No scattered `.md` files per code snippet. - -## What Changes - -| Before | After | -| -------------------------------------------------------------- | ----------------------------------------------------------------------- | -| JSON holds sections, headings, paragraphs as structured arrays | JSON holds only sys_title, title, published, author, tags, published_on | -| Markdown files hold only code snippets | Markdown file holds the entire article | -| `codes[].snippet` maps filename to sectionIndex/paragraphIndex | No mapping needed — headings and code blocks are inline in markdown | -| Renderer reads JSON structure, injects code from .md files | Renderer parses markdown directly into HTML | -| Multiple .md files per article (one per code snippet) | One .md file per article | - -## Impact on the Data Layer - -### wo-model - -The `Article` struct simplifies: - -```rust -pub struct Article { - pub sys_title: String, - pub title: String, - pub published: bool, - pub author: String, - pub tags: Vec, - pub published_on: Option, -} -``` - -The nested `ArticleContent` / `ArticleBody` / `Section` / `CodeSnippet` hierarchy is no longer needed. Article content comes from parsing the `.md` file at render time, not from the JSON. - -### wo-md - -Currently handles only inline markdown (`**bold**`, `` `code` ``, links). Needs to become a full markdown-to-HTML renderer: - -- Block elements: headings (`#`, `##`), paragraphs, code fences (` `lang ```), lists, blockquotes -- Inline elements: bold, italic, code, links, images -- Code fence language extraction for `wo-md::highlight()` -- The renderer reads `{sys_title}/{sys_title}.md`, parses it, and returns HTML - -### wo-htmlx - -The `article.htmlx` template simplifies. Instead of iterating `{{#each article.content.content.sections}}`, it renders the pre-parsed markdown HTML: - -```html -
-

{{article.title}}

-

by {{article.author}} · {{article.tags}}

- {{article.content_html}} -
-``` - -Where `content_html` is the full HTML output from the markdown renderer. - -### wo-store - -`ContentLoader` reads the `.json` for metadata and the `.md` for content. The `.seg` file stores both. At query time, the markdown is either: - -- Pre-rendered to HTML during ingestion (stored in .seg alongside metadata) -- Rendered on-demand at request time (read .md from disk) - -Pre-rendering is preferred — it avoids parsing markdown on every HTTP request. - -## Migration Path - -1. Update `wo-model` with the simplified `Article` struct -2. Extend `wo-md` to handle full markdown (block-level parsing, code fences) -3. Update `ContentLoader` to read `.json` + `.md` pairs -4. Update `wo-store` to store pre-rendered HTML in the .seg file -5. Simplify `article.htmlx` template -6. Migrate existing articles: extract prose from JSON into `.md` files - -Existing articles with the old JSON format can coexist during migration — `ContentLoader` checks for a `.md` file and falls back to the JSON structure if none exists. - -## Blog Subscription — Live Content Reload - -When a user visits `http://localhost:3000/blog/sample-rust-patterns`, the content should stay live. Any edit to `sample-content/sample-rust-patterns/sample-rust-patterns.md` must auto-reflect in the browser without a page refresh. - -### How It Works - -``` -Browser visits /blog/sample-rust-patterns - │ - ▼ -1. Server renders article HTML from .seg (pre-rendered from .md) -2. Server writes HTML response to socket fd -3. Server registers socket fd in subscription table: - register!(sub_manager, socket_fd, ByTitle("sample-rust-patterns")) -4. Connection transitions to Subscribed state (stays open) - │ - │ (user edits sample-rust-patterns.md) - │ - ▼ -5. inotify fires IN_MODIFY on sample-rust-patterns.md -6. ContentWatcher maps file → sys_title "sample-rust-patterns" -7. Store rebuilds: re-reads .json + .md, re-renders markdown to HTML, updates .seg + indexes -8. SubscriptionManager::notify("sample-rust-patterns", ...) fires -9. For each subscribed fd: write(fd, diff_payload) - │ - ▼ -10. Browser receives payload on the open connection -11. Client-side script applies the update to the DOM -``` - -### What Needs to Work - -| Component | Requirement | -|-----------|-------------| -| **inotify** (wo-watch) | Already watches `content/` directory. `.md` file changes must trigger `ContentChange::Modified(sys_title)` | -| **Store rebuild** (wo-store) | On `.md` change: re-read file, re-render markdown to HTML, update `.seg` and indexes | -| **Subscription table** (wo-sub) | Route handler registers the browser's socket fd via `register!` after sending initial HTML | -| **Notification** (wo-sub) | On content change, write updated `content_html` to all subscribed fds as JSON payload | -| **Event loop** (wo-rt) | After writing initial response, transition connection to `Subscribed` state. Keep fd on epoll for hangup detection. | -| **Client script** | Injected in the HTML. Reads payloads from the open connection. Replaces article content in the DOM. | - -### Client-Side Script - -Injected by the template renderer into every article page: - -```html - -``` - -### inotify and .md Files - -The current `ContentWatcher` watches for `.json` file changes. It must also trigger on `.md` file changes: - -- `IN_MODIFY` on `*.md` → `ContentChange::Modified(sys_title)` -- The sys_title is derived from the parent directory name (same as for JSON) -- Both `.json` and `.md` changes trigger a store rebuild and subscriber notification diff --git a/docs/07-ssl.md b/docs/07-ssl.md deleted file mode 100644 index 9c47f56..0000000 --- a/docs/07-ssl.md +++ /dev/null @@ -1,252 +0,0 @@ -# SSL and Deployment - -## Problem - -In [01-problem.md](./01-problem.md), the infrastructure overhead was identified — multiple repos, AWS dependencies, separate deployment pipelines. But one problem went unaddressed: the server-side infrastructure that sits in front of the application — nginx reverse proxy, SSL certificates, systemd service management, and deployment to the production host. - -Currently this requires manual SSH, manual nginx config, manual certbot runs. For a single-binary platform, the deployment should be as simple as the architecture. - -## Target - -Given: -- SSH access to `writeonce.de` is configured -- nginx exists at the default path `/etc/nginx/` -- The writeonce binary listens on a local port (e.g., `127.0.0.1:3000`) - -The deployment pipeline should: -1. Build the binary -2. Copy it to the server -3. Create/update the systemd service -4. Restart the service -5. Configure nginx as a reverse proxy -6. Obtain and auto-renew SSL certificates via Let's Encrypt - -## Systemd Service - -The writeonce binary runs as a systemd service for automatic restart, logging, and boot-start. - -### Service File - -```ini -# /etc/systemd/system/writeonce.service -[Unit] -Description=writeonce content platform -After=network.target - -[Service] -Type=simple -User=writeonce -Group=writeonce -WorkingDirectory=/opt/writeonce -ExecStart=/opt/writeonce/writeonce -Restart=on-failure -RestartSec=5 -StandardOutput=journal -StandardError=journal - -# Security hardening -NoNewPrivileges=true -ProtectSystem=strict -ProtectHome=true -ReadWritePaths=/opt/writeonce/data -PrivateTmp=true - -[Install] -WantedBy=multi-user.target -``` - -### Directory Layout on Server - -``` -/opt/writeonce/ - writeonce # the binary - content/ # article .json + .md files - data/ # derived .seg + .idx (rebuilt on start) - templates/ # .htmlx templates - static/ # CSS, images -``` - -### Service Management - -```bash -# Install / update -sudo systemctl daemon-reload -sudo systemctl enable writeonce -sudo systemctl restart writeonce - -# Check status -sudo systemctl status writeonce -journalctl -u writeonce -f -``` - -## Nginx Reverse Proxy - -Nginx sits in front of the writeonce binary, handling SSL termination and proxying requests to `127.0.0.1:3000`. - -### Nginx Config - -```nginx -# /etc/nginx/sites-available/writeonce.de -server { - listen 80; - server_name writeonce.de www.writeonce.de; - return 301 https://$server_name$request_uri; -} - -server { - listen 443 ssl http2; - server_name writeonce.de www.writeonce.de; - - ssl_certificate /etc/letsencrypt/live/writeonce.de/fullchain.pem; - ssl_certificate_key /etc/letsencrypt/live/writeonce.de/privkey.pem; - ssl_protocols TLSv1.2 TLSv1.3; - ssl_ciphers HIGH:!aNULL:!MD5; - ssl_prefer_server_ciphers on; - - # HSTS - add_header Strict-Transport-Security "max-age=31536000; includeSubDomains" always; - - location / { - proxy_pass http://127.0.0.1:3000; - proxy_set_header Host $host; - proxy_set_header X-Real-IP $remote_addr; - proxy_set_header X-Forwarded-For $proxy_add_x_forwarded_for; - proxy_set_header X-Forwarded-Proto $scheme; - - # Keep connections open for database subscriptions - proxy_http_version 1.1; - proxy_set_header Connection ""; - proxy_read_timeout 86400s; - proxy_send_timeout 86400s; - } - - # Static assets — let nginx serve directly for better caching - location /static/ { - alias /opt/writeonce/static/; - expires 1y; - add_header Cache-Control "public, immutable"; - } -} -``` - -### Enable Site - -```bash -sudo ln -sf /etc/nginx/sites-available/writeonce.de /etc/nginx/sites-enabled/ -sudo nginx -t -sudo systemctl reload nginx -``` - -## SSL with Let's Encrypt - -### Initial Certificate - -```bash -sudo apt install certbot python3-certbot-nginx -sudo certbot --nginx -d writeonce.de -d www.writeonce.de -``` - -Certbot modifies the nginx config to add SSL directives and obtains the certificate. - -### Auto-Renewal - -Certbot installs a systemd timer that runs twice daily: - -```bash -# Check timer -systemctl list-timers | grep certbot - -# Manual test -sudo certbot renew --dry-run -``` - -Certificates auto-renew before expiry. Nginx reloads automatically via certbot's deploy hook. - -### Deploy Hook for Nginx Reload - -```bash -# /etc/letsencrypt/renewal-hooks/deploy/reload-nginx.sh -#!/bin/bash -systemctl reload nginx -``` - -## Deployment Script - -A single script that builds, copies, and restarts: - -```bash -#!/bin/bash -# deploy.sh — run from the development machine -set -e - -SERVER="writeonce.de" -REMOTE_DIR="/opt/writeonce" - -echo "Building release binary..." -cargo build --release -p wo-rt --bin writeonce - -echo "Copying binary to server..." -scp target/release/writeonce $SERVER:$REMOTE_DIR/writeonce.new - -echo "Syncing content and templates..." -rsync -az --delete content/ $SERVER:$REMOTE_DIR/content/ -rsync -az --delete templates/ $SERVER:$REMOTE_DIR/templates/ -rsync -az --delete static/ $SERVER:$REMOTE_DIR/static/ - -echo "Swapping binary and restarting..." -ssh $SERVER " - sudo mv $REMOTE_DIR/writeonce.new $REMOTE_DIR/writeonce - sudo systemctl restart writeonce -" - -echo "Deployed. Checking status..." -ssh $SERVER "sudo systemctl status writeonce --no-pager" -``` - -### First-Time Setup - -Run once on the server to create the user, directory, and service: - -```bash -#!/bin/bash -# setup.sh — run on the server -set -e - -# Create user -sudo useradd -r -s /bin/false writeonce - -# Create directory -sudo mkdir -p /opt/writeonce/{content,data,templates,static} -sudo chown -R writeonce:writeonce /opt/writeonce - -# Install service -sudo cp writeonce.service /etc/systemd/system/ -sudo systemctl daemon-reload -sudo systemctl enable writeonce - -# Configure nginx -sudo cp writeonce.de.nginx /etc/nginx/sites-available/writeonce.de -sudo ln -sf /etc/nginx/sites-available/writeonce.de /etc/nginx/sites-enabled/ -sudo nginx -t -sudo systemctl reload nginx - -# SSL -sudo certbot --nginx -d writeonce.de -d www.writeonce.de -``` - -## What This Replaces - -| Before | After | -|--------|-------| -| Pulumi IaC managing Lambda + S3 | `deploy.sh` with scp + rsync | -| AWS Lambda deployment pipeline | `systemctl restart writeonce` | -| S3 bucket for content hosting | `rsync content/` to server | -| Docker Compose for API + DB | Single binary, one systemd service | -| Multiple nginx configs for API + frontend | One nginx config, one proxy_pass | -| Manual SSL setup | `certbot --nginx` with auto-renewal | - -## Connection Keepalive for Subscriptions - -The nginx config sets `proxy_read_timeout 86400s` (24 hours) to keep persistent connections open for the database subscription model. When a browser visits an article page and the connection transitions to `Subscribed` state, nginx must not timeout and close the upstream connection. - -If nginx is removed in the future (the binary handles TLS directly via `rustls`), this concern disappears — the binary owns the socket end-to-end. diff --git a/docs/08-project-structure.md b/docs/08-project-structure.md index f9e110d..777f958 100644 --- a/docs/08-project-structure.md +++ b/docs/08-project-structure.md @@ -90,19 +90,18 @@ Unchanged from CLAUDE.md's description: `rt` is the monolithic Stage-2 runtime p ``` docs/ -├── 01…08-*.md numbered design docs (this file is 08) -├── writeonce-pl.md language positioning -├── runtime/ user-facing language overview + the 7-phase database series -├── examples/ blog/, ecommerce/, pricing/ samples; ⏳ log-watcher/ (plan 10) +├── 00-*,01,08-*.md status / principles / problem / structure docs +├── runtime/ the 7-phase database design series + runtime concept refs +├── examples/ log-watcher/, employee/, employee-list/ samples ├── plan/ numbered engineering plans 00–16, linux/ cards, assembly/, -│ ├── exploration/ c-runtime/ (A–F, done), ui/ (htmlx track), colibri/ +│ ├── exploration/ c-runtime/ (A–F, done), linux/, postgresql/, assembly/ │ └── oop-vm/ ⏳ the OOP-track contracts: 00-wob-format, 01-error-catalog, │ 02-corpus, 03-shard-actor, 04-db-binding, 05-http-service, │ 06-ui-live, 07-systems-stdlib ├── superpowers/ │ ├── specs/ the two approved track specs (2026-08-01) │ └── plans/ implementation plans 1–10 (2026-08-01, prose-only) -└── future-scope/, cm.md legacy notes +└── cm.md legacy notes ``` Repo rule restated: documentation belongs here; code directories keep one orientation README each. diff --git a/docs/examples/employee-list/README.md b/docs/examples/employee-list/README.md new file mode 100644 index 0000000..a755455 --- /dev/null +++ b/docs/examples/employee-list/README.md @@ -0,0 +1,42 @@ +# employee-list — program B: attach, authenticate, read + +> **Status: target workload — does not compile on today's toolchain.** +> Written ahead of iterations +> [9c (cross-program tables)](../../stories/language-runtime-database/09c-cross-program-tables.md) +> and [9d (keypair attach auth)](../../stories/language-runtime-database/09d-keypair-attach-auth.md), +> the way every acceptance sample here precedes its features. It also leans +> on 9/9b (the [employee sample](../employee/) it attaches to must run +> first). + +Two programs, one database, one writer: + +``` +employee (A) employee-list (B) + owns WO_DATA + WAL no database of its own + [share] listen = unix:...sock [connect.employee] ipc = unix:...sock + [[share.clients]] public_key = + public_key = project = ../employee (shapes) + rights = "read" + ▲ │ + └── every statement executes here ◄───┘ (typed, over the wire) +``` + +The manifests are the design: **A grants, B pins.** A's `[share]` names B's +public-key fingerprint with rights (`read` here); B's `[connect.employee]` +names A's IPC string AND A's fingerprint, so neither side talks to an +impostor. Fingerprints are printed by each program's `--identity` after +first boot (keys are generated into `WO_DATA`, never written into a toml) +and pasted — the `PASTE-…-HERE` placeholders mark exactly where. The +connect-section name is the code's namespace: `[connect.employee]` is why +the source says `employee.Employee`. + +| Mode | What it proves | +| --- | --- | +| `employee-list list` | typed reads over the wire, `e.dept.name` ref navigation executing inside A | +| `employee-list report` | **byte-identical output to A's own `report`** — attach + GroupBy compose, the wire changes nothing | +| `employee-list staff ` | unique-name index probe + `staff` backlink scan, both in A | +| `employee-list probe-write` | the rights matrix: registered read-only, so the insert traps with access-denied (caught, `DENIED …`, exit 4) and A's row count is unchanged | + +The 9d acceptance drives the rest from the outside: wrong key, no key, +same-uid-wrong-key, impostor socket, handshake replay, key rotation — see +the iteration's criteria; this sample is the workload they run against. diff --git a/docs/examples/employee-list/main.wo b/docs/examples/employee-list/main.wo new file mode 100644 index 0000000..b08993c --- /dev/null +++ b/docs/examples/employee-list/main.wo @@ -0,0 +1,73 @@ +-- employee-list — program B of the cross-program story (iterations 9c/9d). +-- Attaches to the RUNNING employee program (A) named by [connect.employee] +-- in wo.toml: keypair handshake first (9d), then typed statements over the +-- wire (9c). A registered this program read-only, so every mode here reads — +-- except probe-write, which exists to prove the rights matrix refuses. +-- +-- The `employee.` prefix is the manifest's connect-section name: these are +-- A's tables, checked against A's shapes at compile time and re-verified by +-- the schema handshake at attach. A stays the single writer; every statement +-- below executes inside A. + +fn main(args: multi Text) -> Int { + if len(args) >= 1 and args[0] == "list" { return list(); } + if len(args) >= 1 and args[0] == "report" { return report(); } + if len(args) >= 2 and args[0] == "staff" { return staff(args[1]); } + if len(args) >= 1 and args[0] == "probe-write" { return probe_write(); } + print_err("usage:"); + print_err(" employee-list list every employee, with department"); + print_err(" employee-list report aggregates by department (A's own report, over the wire)"); + print_err(" employee-list staff one department's staff"); + print_err(" employee-list probe-write prove read-only: the insert must be DENIED"); + return 1; +} + +-- Every employee with forward ref navigation — each `e.dept.name` is a +-- point read executing inside A. +fn list() -> Int { + for e in from x in employee.Employee order by x.name select x { + print("EMP ${e.name} ${e.salary} ${e.dept.name}"); + } + return 0; +} + +-- Byte-identical output to A's own `employee report` — the 9c acceptance +-- line: attach + query compose, and the wire changes nothing. +fn report() -> Int { + let rows = from e in employee.Employee + group e by e.dept into g + order by avg(g.salary) desc + select { dept: g.key.name, headcount: count(g), + avg_salary: avg(g.salary), min_salary: min(g.salary), + max_salary: max(g.salary) }; + for r in rows { + print("DEPT ${r.dept} headcount=${r.headcount} avg=${r.avg_salary} min=${r.min_salary} max=${r.max_salary}"); + } + let payroll = sum(from e in employee.Employee select e.salary); + print("PAYROLL ${payroll}"); + return 0; +} + +-- Unique-name index probe + backlink scan, both executing in A. +fn staff(name: Text) -> Int { + let ds = from d in employee.Department where d.name == name take 1 select d; + if len(ds) == 0 { print_err("no such department: ${name}"); return 1; } + for e in from s in ds[0].staff order by s.salary desc select s { + print("STAFF ${e.name} ${e.salary}"); + } + return 0; +} + +-- The rights matrix, exercised: this program is registered READ-ONLY, so +-- the insert must trap with the access-denied code — caught here, printed, +-- and nothing was applied or WAL-logged in A (the acceptance asserts A's +-- row count is unchanged). +fn probe_write() -> Int { + let id = try insert employee.Department { name: "Intruders" } catch (e) nil; + if id == nil { + print("DENIED write to employee.Department (registered read-only)"); + return 4; + } + print_err("UNEXPECTED: write succeeded with id ${id} — rights not enforced"); + return 1; +} diff --git a/docs/examples/employee-list/wo.toml b/docs/examples/employee-list/wo.toml new file mode 100644 index 0000000..646ed57 --- /dev/null +++ b/docs/examples/employee-list/wo.toml @@ -0,0 +1,29 @@ +name = "employee-list" +version = "0.1.0" +description = "Program B of the cross-program story: attaches read-only to the running employee program (A) over its IPC channel after keypair authentication, and lists/aggregates A's tables over the wire" + +[runtime] +wo = ">= 0.1" + +[build] +runtime = "../../../runtime/wovm" + +# Iteration 9c/9d surface (target — compiles once those land). +# The section NAME is the namespace this program's code uses: [connect.employee] +# makes A's tables reachable as `employee.Department` / `employee.Employee`. +[connect.employee] +# A's IPC string (9c fork 1: unix socket, listening beside A's WO_DATA; +# A declares the same path in its [share] section). +ipc = "unix:../employee/target/data/employee.sock" + +# A's PINNED public-key fingerprint (9d): B refuses to speak to anything on +# that socket path that cannot sign A's challenge with this key — the +# impostor-socket acceptance line. Printed by `employee --identity` after +# A's first boot; paste it here. Never a private key, never a secret. +public_key = "ed25519:PASTE-EMPLOYEE-FINGERPRINT-HERE" + +# Where B's compiler reads A's table shapes at compile time (9c fork 2's +# milestone lean: project-directory reference). The runtime handshake +# re-verifies the shapes against A's live class table at attach — a stale +# checkout refuses the attachment with both shapes named. +project = "../employee" diff --git a/docs/examples/employee/README.md b/docs/examples/employee/README.md new file mode 100644 index 0000000..49ac282 --- /dev/null +++ b/docs/examples/employee/README.md @@ -0,0 +1,32 @@ +# employee — the database track's acceptance workload + +> **Status: target workload — does not compile on today's toolchain.** +> This sample is written *ahead of* the features it exercises, exactly as +> log-watcher was written ahead of iterations 5–7: the sample is the test, +> and the plans compile toward it. It becomes buildable when iteration 9 +> (engine: [`2026-08-01-db-engine-binding.md`](../../superpowers/plans/2026-08-01-db-engine-binding.md)) +> and iteration 9b (query surface: +> [`2026-08-15-employee-relations-query.md`](../../plan/compiler/2026-08-15-employee-relations-query.md)) +> land. Normative semantics: +> [the 9b spec](../../superpowers/specs/2026-08-15-table-relations-query-design.md). + +Two `@table` classes and every 9b feature load-bearing: + +| Mode | What it proves | +| --- | --- | +| `employee seed` | `insert` + WAL-before-ack; a second run catches the `departments.name` `@unique` trap (`SEED-DUP`, exit 3) | +| `employee report` | `group … by … into g` lowered as one hash pass; `count`/`avg`/`min`/`max` per department; `order by avg(g.salary) desc`; whole-query `sum` for payroll | +| `employee staff ` | unique-name **index probe** (asserted via the engine's probe counter, not assumed), `backlink` scan one way, `e.dept.name` ref navigation the other | +| `employee raise ` | update through a query result; a missing department takes the empty-query path | +| `employee drop ` | delete **restrict**: a department with staff traps (exit 4); the trap code is the acceptance's assertion | + +The engine lives in `database/` (statically linked into `wovm` — still one +binary). Rows are RAM-authoritative, WAL-durable; a kill -9 between `seed` +and `report` followed by an identical `report` is part of the acceptance +script (iteration 9's replay, proven on this workload). + +Data shape: `Department { name @unique, staff: backlink Employee.dept }`, +`Employee { name, salary (cents), hired (epoch ms), dept: ref Department }`, +indexes `[name]`, `[dept]`, `[dept, salary]`. Salaries wrap like all language +arithmetic; `avg`/`min`/`max` are `?Int` because an empty group is data, not +a fault. diff --git a/docs/examples/employee/justfile b/docs/examples/employee/justfile new file mode 100644 index 0000000..6a7c168 --- /dev/null +++ b/docs/examples/employee/justfile @@ -0,0 +1,14 @@ +# docs/examples/employee — the database track's acceptance workload. +# `just employee::` from the repo root, or plain `just ` here. +ROOT := source_directory() / "../../.." + +default: accept + +# build the standalone binary from wo.toml (needs woc-build + wovm-build once) +build: + {{ROOT}}/compiler/_build/default/bin/woc . + @ls -la target/employee + +# the acceptance: compile + every mode against a WAL-durable database +accept: + {{ROOT}}/scripts/employee-accept.sh diff --git a/docs/examples/employee/main.wo b/docs/examples/employee/main.wo new file mode 100644 index 0000000..cedd19d --- /dev/null +++ b/docs/examples/employee/main.wo @@ -0,0 +1,124 @@ +-- Employee management — iteration 9/9b's acceptance workload. +-- Four modes, each existing to make one feature load-bearing: +-- seed inserts (WAL, @unique trap on the second run) +-- report the GroupBy showcase (headcount/avg/min/max, payroll) +-- staff relation navigation both directions + index probe +-- raise

update through a query result; empty query => nil path +-- drop delete restrict: a department with staff must trap + +fn main(args: multi Text) -> Int { + if len(args) >= 1 and args[0] == "seed" { return seed(); } + if len(args) >= 1 and args[0] == "report" { return report(); } + if len(args) >= 2 and args[0] == "staff" { return staff(args[1]); } + if len(args) >= 3 and args[0] == "raise" { + let pct = parse_int(args[2]); + if pct == nil { print_err("raise: must be a number"); return 2; } + return raise(args[1], pct); + } + if len(args) >= 2 and args[0] == "drop" { return drop_dept(args[1]); } + print_err("usage:"); + print_err(" employee seed seed departments and employees"); + print_err(" employee report aggregates by department"); + print_err(" employee staff list a department's staff"); + print_err(" employee raise

raise a department's salaries p%"); + print_err(" employee drop delete a department (restrict demo)"); + return 1; +} + +-- Inserts are WAL-logged before acknowledgment (iteration 9). A second run +-- hits the departments.name @unique index and traps; catching it here is the +-- sample's unique-violation acceptance line. +fn seed() -> Int { + let eng = try insert Department { name: "Engineering" } catch (e) nil; + if eng == nil { + print("SEED-DUP departments already seeded (unique violation caught)"); + return 3; + } + let ops = insert Department { name: "Operations" }; + let sales = insert Department { name: "Sales" }; + + insert Employee { name: "Asha", salary: 9200000, hired: 1704067200000, dept: eng }; + insert Employee { name: "Bram", salary: 8100000, hired: 1706745600000, dept: eng }; + insert Employee { name: "Chidi", salary: 7300000, hired: 1709251200000, dept: eng }; + insert Employee { name: "Dora", salary: 6400000, hired: 1711929600000, dept: ops }; + insert Employee { name: "Emil", salary: 5900000, hired: 1714521600000, dept: ops }; + insert Employee { name: "Farah", salary: 8800000, hired: 1717200000000, dept: sales }; + + print("SEEDED 3 departments, 6 employees"); + return 0; +} + +-- Per-department aggregates. The group-by SYNTAX +-- from e in Employee group e by e.dept into g order by avg(g.salary) desc +-- select { dept: g.key.name, headcount: count(g), avg_salary: avg(g.salary), ... } +-- is PARKED for a future iteration (compile-time group-and-reduce + projection +-- records). Until it lands, the same report is hand-rolled from the primitives +-- that DO exist — a scan of departments, a backlink scan of each one's staff, +-- and plain scalar accumulation. Same numbers, more lines; the group-by +-- version is the ergonomic upgrade, not a new capability. +fn report() -> Int { + let payroll = 0; + for d in from x in Department order by x.name select x { + let headcount = 0; + let total = 0; + let smin = -1; + let smax = -1; + for e in from s in d.staff select s { + headcount = headcount + 1; + total = total + e.salary; + payroll = payroll + e.salary; + if smin == -1 or e.salary < smin { smin = e.salary; } + if smax == -1 or e.salary > smax { smax = e.salary; } + } + if headcount == 0 { + print("DEPT ${d.name} headcount=0 avg=nil min=nil max=nil"); + } else { + print("DEPT ${d.name} headcount=${headcount} avg=${total / headcount} min=${smin} max=${smax}"); + } + } + print("PAYROLL ${payroll}"); + return 0; +} + +-- Both navigation directions on one screen: the department found by its +-- unique-name index (a probe, not a scan — the acceptance asserts the probe +-- counter), its `staff` backlink scanned, and each employee's forward +-- `e.dept.name` printed to prove ref navigation. +fn staff(name: Text) -> Int { + let ds = from d in Department where d.name == name take 1 select d; + if len(ds) == 0 { print_err("no such department: ${name}"); return 1; } + let d = ds[0]; + for e in from s in d.staff order by s.salary desc select s { + print("STAFF ${e.name} ${e.salary} (${e.dept.name})"); + } + return 0; +} + +-- Update through a query result; a missing department exercises the +-- empty-query path (the loop body never runs, nothing prints but the DONE). +fn raise(name: Text, pct: Int) -> Int { + let ds = from d in Department where d.name == name take 1 select d; + if len(ds) == 0 { print("RAISE ${name}: no such department (0 rows)"); return 0; } + let d = ds[0]; + let n = 0; + for e in from s in d.staff select s { + e.salary = e.salary + e.salary * pct / 100; + n = n + 1; + } + print("RAISE ${name} ${pct}% applied to ${n} employees"); + return 0; +} + +-- Restrict is the only FK action: deleting a department that employees still +-- reference traps, and the trap code is the acceptance's assertion. +fn drop_dept(name: Text) -> Int { + let ds = from d in Department where d.name == name take 1 select d; + if len(ds) == 0 { print_err("no such department: ${name}"); return 1; } + let ok = try delete ds[0] catch (e) nil; + if ok == nil { + print("DROP ${name}: restricted (staff still reference it)"); + return 4; + } + print("DROP ${name}: deleted"); + return 0; +} diff --git a/docs/examples/employee/types.wo b/docs/examples/employee/types.wo new file mode 100644 index 0000000..324074d --- /dev/null +++ b/docs/examples/employee/types.wo @@ -0,0 +1,24 @@ +-- The two tables the whole sample exists to relate. Every class IS a table; +-- @table only configures storage (name, indexes) — the language's oldest +-- doctrine. The [dept] index serves the `staff` backlink and the restrict +-- check; [dept, salary] serves the per-department salary queries and the +-- report's ordering inside a department. + +@table(name: "departments", index: [name]) +class Department { + name: Text @unique + + -- Not a stored column: the declared inverse of Employee.dept. Reading + -- `d.staff` is a secondary-index scan of employees.dept and yields + -- `multi Employee`. + staff: backlink Employee.dept +} + +@table(name: "employees", index: [dept], index: [dept, salary]) +class Employee { + name: Text + salary: Int -- cents; sum wraps like all language arithmetic + hired: Int -- epoch ms + dept: ref Department -- FK: stored as the department's row id, checked + -- by a primary-index probe on insert/update +} diff --git a/docs/examples/employee/wo.toml b/docs/examples/employee/wo.toml new file mode 100644 index 0000000..6c2a151 --- /dev/null +++ b/docs/examples/employee/wo.toml @@ -0,0 +1,28 @@ +name = "employee" +version = "0.1.0" +description = "Employee management — iteration 9/9b acceptance workload: @table storage, ref/backlink relations, compiler-checked GroupBy queries" + +[runtime] +wo = ">= 0.1" + +# `woc .` builds target/employee once iterations 9 (engine) and 9b (query +# surface) land; until then this sample is the target the plans compile +# toward, sample-first like log-watcher was. +[build] +runtime = "../../../runtime/wovm" + +# Iteration 9c/9d surface (target): this program OWNS its database and +# shares it. The runtime listens on `listen` beside WO_DATA; every client +# below is a grant — no registration, no attach, same uid included. +[share] +listen = "unix:target/data/employee.sock" + +# Program B (docs/examples/employee-list), granted read-only. Identity is +# B's public-key fingerprint (9d): printed by `employee-list --identity` +# after B's first boot; paste it here. Rights: "read" or "rw" — B's +# probe-write mode exists to prove "read" refuses. Rotation and revocation +# are edits to this table plus a restart, never an API. +[[share.clients]] +name = "employee-list" +public_key = "ed25519:PASTE-EMPLOYEE-LIST-FINGERPRINT-HERE" +rights = "read" diff --git a/docs/future-scope/ai-agents-content-management.md b/docs/future-scope/ai-agents-content-management.md deleted file mode 100644 index 59119ed..0000000 --- a/docs/future-scope/ai-agents-content-management.md +++ /dev/null @@ -1,182 +0,0 @@ -# AI Agents and Content Management - -## Context - -AI agents (Claude Code, Copilot, Cursor, custom agents) work within a specific project or working directory. Their sessions, context, and understanding are scoped to that directory. This works well when projects are completely different domains. - -But writeonce content is not isolated — articles reference each other, share tags, build on concepts from other articles. An agent editing `gitlab-runner-with-kubernetes-executor.md` would benefit from knowing that `auto-scale-gitlab-runner-using-aws-spot-instance.md` exists and covers related infrastructure. Without explicit mappings, the agent treats each article as an island. - -## Problem - -1. **Agents lack cross-article awareness.** When asked to write or update an article about Kubernetes, the agent doesn't know that related articles about Docker, CI/CD, or AWS already exist in the content directory — unless it manually searches. - -2. **No semantic grouping.** Tags provide flat categorization (`kubernetes`, `ci-cd`), but they don't express relationships: "this article is a prerequisite for that one", "these three articles form a series", "this article supersedes that one." - -3. **Context window waste.** Without mappings, the agent must scan all articles to find related content. With explicit mappings, it can load exactly the relevant files. - -## Solution: Metadata-Driven Content Mappings - -Users define relationships between articles in the JSON metadata. These mappings serve two purposes: - -1. **Human navigation** — rendered as "related articles" links on the site -2. **Agent context** — when an agent works on an article, it loads the mapped articles into its context for cross-referencing - -### Mapping Fields in JSON Metadata - -Per [06-markdown-render.md](../06-markdown-render.md), the JSON metadata is minimal. Add a `mappings` field: - -```json -{ - "sys_title": "gitlab-runner-with-kubernetes-executor", - "title": "Gitlab Runner with Kubernetes Executor", - "published": true, - "author": "Shoney Arickathil", - "tags": ["kubernetes", "gitlab", "ci-cd"], - "published_on": 1740950884, - "mappings": { - "related": ["auto-scale-gitlab-runner-using-aws-spot-instance"], - "prerequisite": ["linux-misc"], - "series": { - "name": "gitlab-runner", - "order": 2 - } - } -} -``` - -### Mapping Types - -| Type | Meaning | Agent Use | -| -------------- | ----------------------------------------------- | --------------------------------------------------------------------------------------- | -| `related` | Topically related articles | Agent loads these for cross-reference when editing | -| `prerequisite` | Articles the reader should read first | Agent ensures no concept duplication, references prerequisites instead of re-explaining | -| `series` | Articles that form an ordered sequence | Agent maintains narrative continuity across the series | -| `supersedes` | This article replaces an older one | Agent can mark the old article as outdated or unpublished | -| `references` | External articles or URLs the content builds on | Agent checks links are still valid, cites them properly | - -### Directory Structure with Mappings - -``` -content/ - gitlab-runner-with-kubernetes-executor/ - gitlab-runner-with-kubernetes-executor.json # metadata + mappings - gitlab-runner-with-kubernetes-executor.md # full article - auto-scale-gitlab-runner-using-aws-spot-instance/ - auto-scale-gitlab-runner-using-aws-spot-instance.json - auto-scale-gitlab-runner-using-aws-spot-instance.md - linux-misc/ - linux-misc.json - linux-misc.md -``` - -## Agent Workflows - -### 1. Writing a New Article - -The author asks an agent: "Write an article about deploying GitLab Runner on ECS." - -The agent: - -1. Scans the content directory for existing articles with tags `gitlab`, `ci-cd`, `aws` -2. Finds `gitlab-runner-with-kubernetes-executor` and `auto-scale-gitlab-runner-using-aws-spot-instance` -3. Reads their `.md` files to understand what's already covered -4. Writes the new article, referencing existing articles rather than re-explaining shared concepts -5. Suggests `mappings.related` entries for the new article's JSON - -### 2. Updating an Existing Article - -The author asks: "Update the Kubernetes executor article with the new runner token format." - -The agent: - -1. Reads the article's JSON metadata and `.md` content -2. Reads the `mappings.related` articles to check for consistency -3. Makes the update in the `.md` file -4. Checks if the change affects any prerequisite or series articles -5. inotify detects the `.md` change → store rebuilds → subscribers notified - -### 3. Content Audit - -The author asks: "Which articles reference outdated AWS configurations?" - -The agent: - -1. Loads all article metadata (the Store already indexes everything) -2. Follows `mappings` to build a dependency graph -3. Reads the `.md` files of articles tagged with `aws` -4. Identifies outdated patterns (old SDK versions, deprecated services) -5. Reports findings with links to specific articles and line numbers - -### 4. Series Management - -The author asks: "Add a new part to the gitlab-runner series." - -The agent: - -1. Finds all articles with `mappings.series.name == "gitlab-runner"` -2. Reads them in order to understand the narrative arc -3. Writes the new article continuing from where the series left off -4. Sets `mappings.series.order` to the next number -5. Updates the previous article's mappings to reference the new one - -## Integration with writeonce Architecture - -### Store Index - -Add a mappings index alongside the existing title, date, and tag indexes: - -``` -data/ - articles.seg - index/ - title.idx - date.idx - tags.idx - mappings.idx # sys_title → related sys_titles -``` - -The mappings index allows efficient traversal: "give me all articles related to X" without scanning every article's JSON. - -### Template Rendering - -The `article.htmlx` template can render related articles: - -```html -

-

{{article.title}}

- {{article.content_html}} {{#each article.related}} - - {{/each}} -
-``` - -### Subscription - -When a mapped article changes, subscribers to related articles can optionally be notified. If article A lists article B in `mappings.related`, and article B is updated, subscribers to article A can receive a notification that related content changed. - -## Agent Configuration - -For agents to use the mappings effectively, the project can include an agent instruction file (e.g., `CLAUDE.md` or `.agent/instructions.md`): - -```markdown -## Content Management - -- Articles are in `content/{sys_title}/{sys_title}.md` -- Metadata is in `content/{sys_title}/{sys_title}.json` -- Before writing or editing an article, read its `mappings` field and load related articles for context -- When creating a new article, suggest appropriate `mappings` based on tags and content overlap -- Maintain narrative continuity within `series` mappings -- Do not duplicate explanations that exist in `prerequisite` articles — reference them instead -``` - -This turns the content directory into an agent-navigable knowledge graph where the metadata provides the edges and the markdown files provide the nodes. - -## queryable graph database - -- traversable knowledge graphs available on RAM. -- which linux kernels, develop in C ++. User wants to learn it. diff --git a/docs/plan/07-inotify-content-watcher.md b/docs/plan/07-inotify-content-watcher.md index a2a2f26..0380339 100644 --- a/docs/plan/07-inotify-content-watcher.md +++ b/docs/plan/07-inotify-content-watcher.md @@ -2,7 +2,7 @@ > **Status: ⬜ not started** — Track 1 (runtime foundations). Board: [00-status.md](../00-status.md) -**Context sources:** [`./02-event-loop-epoll.md`](./done/02-event-loop-epoll.md), [`./linux/00-linux.md`](./exploration/linux/00-linux.md) § File Watching, [`../02-recovery.md`](../02-recovery.md) § No AWS Infrastructure. +**Context sources:** [`./02-event-loop-epoll.md`](./done/02-event-loop-epoll.md), [`./linux/00-linux.md`](./exploration/linux/00-linux.md) § File Watching, `../02-recovery.md` § No AWS Infrastructure. ## Goal diff --git a/docs/plan/08-sendfile-static-assets.md b/docs/plan/08-sendfile-static-assets.md index 92cf90f..9002f2f 100644 --- a/docs/plan/08-sendfile-static-assets.md +++ b/docs/plan/08-sendfile-static-assets.md @@ -2,7 +2,7 @@ > **Status: ⬜ not started** — Track 1 (runtime foundations); also a prerequisite of the parked UI track. Board: [00-status.md](../00-status.md) -**Context sources:** [`./03-hand-rolled-http.md`](./done/03-hand-rolled-http.md), [`./linux/00-linux.md`](./exploration/linux/00-linux.md) § Efficient File Serving, [`../02-recovery.md`](../02-recovery.md). +**Context sources:** [`./03-hand-rolled-http.md`](./done/03-hand-rolled-http.md), [`./linux/00-linux.md`](./exploration/linux/00-linux.md) § Efficient File Serving, `../02-recovery.md`. ## Goal diff --git a/docs/plan/11-wal-and-recovery.md b/docs/plan/11-wal-and-recovery.md index 3675e15..cb03269 100644 --- a/docs/plan/11-wal-and-recovery.md +++ b/docs/plan/11-wal-and-recovery.md @@ -2,7 +2,7 @@ > **Status: ⬜ not started (scope reduced)** — replay + ack-after-fsync + group commit landed via 09c and its follow-ups; remaining here: snapshots (`.data`), compaction, WAL rotation. Board: [00-status.md](../00-status.md) -**Context sources:** [`./10-storage-foundations.md`](./10-storage-foundations.md), [`../runtime/database/02-wo-language.md#concurrency-model`](../runtime/database/02-wo-language.md#concurrency-model), [`../runtime/database/03-inmemory-engine.md`](../runtime/database/03-inmemory-engine.md), [`./exploration/postgresql/wal.md`](./exploration/postgresql/wal.md), [`./exploration/postgresql/buffer-and-checkpoint.md`](./exploration/postgresql/buffer-and-checkpoint.md), [`./exploration/linux/12-pwrite-fsync.md`](./exploration/linux/12-pwrite-fsync.md), [`../02-recovery.md`](../02-recovery.md). +**Context sources:** [`./10-storage-foundations.md`](./10-storage-foundations.md), [`../runtime/database/02-wo-language.md#concurrency-model`](../runtime/database/02-wo-language.md#concurrency-model), [`../runtime/database/03-inmemory-engine.md`](../runtime/database/03-inmemory-engine.md), [`./exploration/postgresql/wal.md`](./exploration/postgresql/wal.md), [`./exploration/postgresql/buffer-and-checkpoint.md`](./exploration/postgresql/buffer-and-checkpoint.md), [`./exploration/linux/12-pwrite-fsync.md`](./exploration/linux/12-pwrite-fsync.md), `../02-recovery.md`. ## Goal diff --git a/docs/plan/13-class-model-live-pricing.md b/docs/plan/13-class-model-live-pricing.md index 43ef678..3bcb7dc 100644 --- a/docs/plan/13-class-model-live-pricing.md +++ b/docs/plan/13-class-model-live-pricing.md @@ -2,7 +2,7 @@ > **Status: 🔄 in progress** — 13a ✅ shipped; 13b ✅ shipped (methods execute over RPC); 13c (LIVE push) is next; 13d ⏸ parked (frontend); 13e ⬜. Board: [00-status.md](../00-status.md) -**Context sources:** [`../runtime/database/02-wo-language.md`](../runtime/database/02-wo-language.md) (schema layer, § Schema-Layer DML brace disambiguation, § Cross-Paradigm Transaction Coordinator), [`../runtime/database/04-client-api.md`](../runtime/database/04-client-api.md) (subscription engine), [`./09-concurrency-scaleout.md`](./09-concurrency-scaleout.md) (thread-per-core scale-out), [`./exploration/ui/00-overview.md`](./exploration/ui/00-overview.md) + [`./exploration/ui/01-htmlx-format-spec.md`](./exploration/ui/01-htmlx-format-spec.md) (live UI), [`../examples/pricing/`](../examples/pricing/) (the demo this phase makes real), [`../examples/ecommerce/shared/logic/checkout.wo`](../examples/ecommerce/shared/logic/checkout.wo) (the existing `fn … in txn snapshot` signature style methods reuse). +**Context sources:** [`../runtime/database/02-wo-language.md`](../runtime/database/02-wo-language.md) (schema layer, § Schema-Layer DML brace disambiguation, § Cross-Paradigm Transaction Coordinator), [`../runtime/database/04-client-api.md`](../runtime/database/04-client-api.md) (subscription engine), [`./09-concurrency-scaleout.md`](./09-concurrency-scaleout.md) (thread-per-core scale-out), `./exploration/ui/00-overview.md` + `./exploration/ui/01-htmlx-format-spec.md` (live UI), [`../examples/pricing/`](../examples/pricing/) (the demo this phase makes real), [`../examples/ecommerce/shared/logic/checkout.wo`](../examples/ecommerce/shared/logic/checkout.wo) (the existing `fn … in txn snapshot` signature style methods reuse). ## Context @@ -73,7 +73,7 @@ Each lands as its own numbered plan doc (`13a-…`, `13b-…`) when ready. The b ### `13a-class-surface.md` — lexer, parser, AST, spec amendments — ✅ shipped -`class` joins the keyword map (`crates/rt/src/lexer.rs` keyword match, ~line 202 — note `self` stays an ident per decision 4). `parse_type` (`crates/rt/src/parser.rs:120`) takes the leading keyword as a parameter and serves both constructs; `fn` members inside the body parse-and-discard through the existing brace-depth skip — the same mechanism that already swallows `on update … do { … }` triggers. `ast::TypeDecl` gains `is_class: bool`; `Catalog::from_schemas` ignores it (decision 5), so REST CRUD works the moment parsing does. Docs amended in the same change: a "Class Model" subsection in [`02-wo-language.md`](../runtime/database/02-wo-language.md) next to § Schema-Layer DML, the "Isn't OO" paragraph in [`wo-language.md`](../runtime/wo-language.md), the class line in [`writeonce-pl.md`](../writeonce-pl.md), and `just pricing` / `just pricing-demo` recipes. +`class` joins the keyword map (`crates/rt/src/lexer.rs` keyword match, ~line 202 — note `self` stays an ident per decision 4). `parse_type` (`crates/rt/src/parser.rs:120`) takes the leading keyword as a parameter and serves both constructs; `fn` members inside the body parse-and-discard through the existing brace-depth skip — the same mechanism that already swallows `on update … do { … }` triggers. `ast::TypeDecl` gains `is_class: bool`; `Catalog::from_schemas` ignores it (decision 5), so REST CRUD works the moment parsing does. Docs amended in the same change: a "Class Model" subsection in [`02-wo-language.md`](../runtime/database/02-wo-language.md) next to § Schema-Layer DML, the "Isn't OO" paragraph in `wo-language.md`, the class line in `writeonce-pl.md`, and `just pricing` / `just pricing-demo` recipes. **Exit (met):** `wo run docs/examples/pricing` parses 2 classes, serves `/api/products` CRUD; parser unit tests (`parses_class_with_methods`, `class_method_braces_do_not_truncate_body`) green; blog/ecommerce/hello unchanged. ### `13b-method-execution.md` — methods over RPC — ✅ shipped @@ -88,7 +88,7 @@ The subscription registry from [`04-client-api.md`](../runtime/database/04-clien ### `13d-pricing-ui.md` — the `/pricing` screen, MVC -The screen ships as an **MVC triplet** per [`exploration/ui/08-mvc-structure.md`](./exploration/ui/08-mvc-structure.md), built in the sub-phase sequence of [`14-mvc-ui-implementation.md`](./14-mvc-ui-implementation.md): model = the classes themselves, view = [`pricing.htmlx`](../examples/pricing/ui/pricing/pricing.htmlx) (plain htmlx, logic-free) + external [`pricing.scss`](../examples/pricing/ui/pricing/pricing.scss) (strict SCSS subset compiled at `wo build`, no external deps), controller = [`pricing.wo`](../examples/pricing/ui/pricing/pricing.wo) (`route:`/`view:`/`styles:`, `model:` bindings, `actions:` calling the 13b class methods). SSR per [`exploration/ui/01-htmlx-format-spec.md`](./exploration/ui/01-htmlx-format-spec.md), compiler glue per [`02-ui-compiler.md`](./exploration/ui/02-ui-compiler.md), and the vanilla-JS client runtime ([`03-client-runtime.md`](./exploration/ui/03-client-runtime.md)) patches the price cell when the 13c delta lands. The controller's `model:` block is the M→V binding; the watchlist narrows the subscription predicate server-side. +The screen ships as an **MVC triplet** per `exploration/ui/08-mvc-structure.md`, built in the sub-phase sequence of `14-mvc-ui-implementation.md`: model = the classes themselves, view = [`pricing.htmlx`](../examples/pricing/ui/pricing/pricing.htmlx) (plain htmlx, logic-free) + external [`pricing.scss`](../examples/pricing/ui/pricing/pricing.scss) (strict SCSS subset compiled at `wo build`, no external deps), controller = [`pricing.wo`](../examples/pricing/ui/pricing/pricing.wo) (`route:`/`view:`/`styles:`, `model:` bindings, `actions:` calling the 13b class methods). SSR per `exploration/ui/01-htmlx-format-spec.md`, compiler glue per `02-ui-compiler.md`, and the vanilla-JS client runtime (`03-client-runtime.md`) patches the price cell when the 13c delta lands. The controller's `model:` block is the M→V binding; the watchlist narrows the subscription predicate server-side. **Exit:** browser at `/pricing` shows selected products; a `set_price` commit from curl changes the price cell in every open browser without reload. ### `13e-pricing-at-scale.md` — millions of readers, millions of live updates @@ -122,5 +122,4 @@ No new architecture — this sub-phase wires the demo to [`09-concurrency-scaleo - [`../runtime/database/02-wo-language.md`](../runtime/database/02-wo-language.md) — schema layer the class grammar extends; transaction coordinator methods reuse. - [`../runtime/database/04-client-api.md`](../runtime/database/04-client-api.md) — subscription engine 13c scopes down. - [`./09-concurrency-scaleout.md`](./09-concurrency-scaleout.md) — the scale architecture 13e instantiates. -- [`./exploration/ui/00-overview.md`](./exploration/ui/00-overview.md) — UI track 13d draws on. - [`../examples/hello/main.wo`](../examples/hello/main.wo) — the minimal example whose `Revision`-trigger pattern is the declarative ancestor of methods. diff --git a/docs/plan/14-mvc-ui-implementation.md b/docs/plan/14-mvc-ui-implementation.md deleted file mode 100644 index ecef360..0000000 --- a/docs/plan/14-mvc-ui-implementation.md +++ /dev/null @@ -1,90 +0,0 @@ -# 14 — MVC UI implementation: model = class, view = htmlx + scss, controller = .wo - -> **Status: ⏸ parked (frontend)** — backend focus first; design stays current. Board: [00-status.md](../00-status.md) - -**Context sources:** [`./exploration/ui/08-mvc-structure.md`](./exploration/ui/08-mvc-structure.md) (the design this plan implements), [`./exploration/ui/01-htmlx-format-spec.md`](./exploration/ui/01-htmlx-format-spec.md) / [`02-ui-compiler.md`](./exploration/ui/02-ui-compiler.md) / [`03-client-runtime.md`](./exploration/ui/03-client-runtime.md) (the three UI-track pieces this plan sequences, each with port sources and LOC budgets), [`./13-class-model-live-pricing.md`](./13-class-model-live-pricing.md) (the class methods controllers call: 13a/13b; the LIVE deltas views consume: 13c), [`../examples/pricing/ui/pricing/`](../examples/pricing/ui/pricing/) (the reference MVC triplet), [`reference/crates/wo-htmlx/`](../../.dev/reference/crates/wo-htmlx/) (the v1 template engine, primary port source). - -## Context - -[`exploration/ui/08-mvc-structure.md`](./exploration/ui/08-mvc-structure.md) locks the screen anatomy: **model** = the `class`/`type` itself, **view** = plain `.htmlx` + external `.scss`, **controller** = a `.wo` file (`route:`/`view:`/`styles:`/`model:`/`actions:`) that binds the model into the view and is the only place UI may call class methods. The UI exploration docs 01–03 already specify the htmlx engine, the `##ui` compiler, and the client runtime in implementable detail. What's missing is the build order, the two genuinely new pieces (the controller format and the SCSS subset compiler), and the wiring into the `13` class-model track. This doc is that sequence. - -Everything lands in **`crates/ui`** (currently a placeholder) and small deltas to `crates/rt` — consistent with the crate inventory in [`crates/README.md`](../../crates/README.md). The deployment shape never changes: one binary serving SSR + database + API on the kernel-primitive runtime. - -## Goal - -`cargo run --bin wo -- run docs/examples/pricing` (after plan 13a–13c land) serves `GET /pricing` as styled SSR HTML; clicking ☆ dispatches a controller action; an Ops `set-price` action calls `Product.set_price`, the commit pushes a delta, and the price cell patches in every open browser without reload — the [plan 13d exit criterion](./13-class-model-live-pricing.md), implemented MVC-shaped. - -## Dependency graph - -``` -14a htmlx engine ──────┬─→ 14c controller format ─→ 14d SSR routes ─→ 14e actions ─→ 14f live patch -14b scss compiler ─────┘ │ │ │ - (14a ∥ 14b — no shared code) needs engine needs 13a+13b needs 13c - (exists today) -``` - -Phases 05/06 (hand-rolled JSON / bespoke error) are orthogonal: `crates/ui` adopts `serde`/`serde_json` per the ui/01 decision and migrates when 05 lands. Phase 08 (`sendfile`) upgrades static-asset serving in 14d when it arrives; 14d ships with plain buffered writes first. - -## Sub-phase sequence - -### `14a-htmlx-engine.md` — port the view engine into `crates/ui` - -Execute [`exploration/ui/01-htmlx-format-spec.md`](./exploration/ui/01-htmlx-format-spec.md) as written: port `reference/crates/wo-htmlx` (585 LOC — `parser.rs`, `ast.rs`, `value.rs`, `registry.rs`, `render.rs` carried over per its table) into `crates/ui/src/htmlx/`, extend with `` structured nodes, `wo:bind` capture, and the `data-wo-manifest` JSON emitter (~250 LOC new). One addition beyond the 01 spec, from the MVC design: `` records whether `source` is a bare name (controller model binding, resolved in 14d) or an inline query — a one-field change to `LiveSubscription`. -**Exit:** the 01 spec's criteria — `cargo build -p ui` green, golden parse+render for every `.htmlx` under `docs/examples/{blog,ecommerce}` **plus** [`pricing/ui/pricing/pricing.htmlx`](../examples/pricing/ui/pricing/pricing.htmlx), manifest matches the 01 schema. - -### `14b-scss-subset.md` — the stylesheet compiler - -New, no port source (~400 LOC at `crates/ui/src/scss/`): scanner → rule tree → flattener. Exactly the subset locked in [08-mvc decision 3](./exploration/ui/08-mvc-structure.md): `$variables`, nesting (including `&`-less descendant flattening), `@use "partials"` (`ui/styles/_*.scss`), comments. **No mixins, functions, `@extend`, or color math** — `rgba($accent, 0.06)` in the reference file compiles by literal substitution of `$accent` and is the only function-form supported. Output is one flat `.css` per screen, written to `target/wo//static/`. -**Exit:** [`pricing.scss`](../examples/pricing/ui/pricing/pricing.scss) → golden-file CSS; unknown construct = compile error naming file:line (never silent passthrough); runs standalone (`wo build` integration is 14d). - -### `14c-controller-format.md` — parse the controller, keep the shorthand - -The `Kind::HashHash("ui")` skip arm in `crates/rt/src/parser.rs` (L80–88) parses for real, into one of two IRs by key-shape dispatch: - -- **Controller form** (`route:`/`view:`/`styles:`/`model:`/`actions:` — the MVC triplet's `pricing.wo`) → new `Controller` IR: route pattern, view/styles paths resolved relative to the screen directory, `model:` entries as named query strings (`LIVE` flag captured, execution deferred), `actions:` entries as `(name, params, target-method-or-fn, role-set)`. -- **Shorthand form** (`source:`/`columns:`/… — the existing ecommerce/blog screens) → the `Screen` IR of [`exploration/ui/02-ui-compiler.md`](./exploration/ui/02-ui-compiler.md), whose codegen emits a generated view + a synthesized `Controller` — [08-mvc decision 5](./exploration/ui/08-mvc-structure.md): the shorthand is sugar over the triplet, one downstream path. - -`##app` and `##component` keep their current skip behaviour (owned by ui/05 and ui/02 respectively). -**Exit:** `pricing.wo` parses to a `Controller` with 2 model bindings + 3 actions; every existing `##ui` screen in blog/ecommerce parses to `Screen` and compiles to an `.htmlx` that 14a round-trips; a controller naming a missing view file is a compile error. - -### `14d-ssr-routes.md` — the single binary serves the screen - -Wire controllers into `crates/rt`'s router (`crates/rt/src/server.rs`): each `Controller.route` becomes a GET route; the handler resolves `model:` bindings against the in-process engine (snapshot `select` now — the `LIVE` flag additionally registers a 13c subscription when that phase is live), renders the view via 14a with the model names as root scope, and emits HTML + manifest + ``. `wo build`/`wo run` gain the asset step: compile SCSS (14b), bake `/_wo/runtime.js` (`include_bytes!`, per ui/03 decision 6), serve `target/wo/.../static/` with buffered writes (upgraded to `sendfile` when [phase 08](./08-sendfile-static-assets.md) lands). -**Exit:** `GET /pricing` returns styled SSR HTML with a valid manifest and resolvable CSS/JS links; `GET /` on the blog sample is unaffected; route table printed at boot includes UI routes alongside REST. - -### `14e-action-dispatch.md` — controller actions call class methods - -`wo:action` buttons POST to `/_wo/action//` with `wo:args` + form payload. The dispatcher looks up the controller's action table, enforces the `role:` set server-side (per-app policy model of [`exploration/ui/07-per-app-policies.md`](./exploration/ui/07-per-app-policies.md)), and invokes the target: a class method via the 13b row-scoped RPC path (`Product{ id == id }.set_price(amount)`) or a free `fn`. Response is 204 — the UI never re-renders from the action response; the visible change arrives as a 13c delta, keeping one update path. -**Requires:** 13a + 13b. **Exit:** the ☆/★ watch toggle round-trips; `set-price` with an Ops session commits a Price; without the role it's 403 and no transaction starts. - -### `14f-live-patching.md` — the browser follows commits - -Execute [`exploration/ui/03-client-runtime.md`](./exploration/ui/03-client-runtime.md) as written (~500 LOC vanilla JS at `crates/ui/assets/wo-runtime.js`, JSON frames over one WebSocket, targeted DOM patching by `data-key` + `wo:bind`, coalescing backpressure, snapshot resync on reconnect), pointed at the 13c subscription endpoint. -**Requires:** 13c. **Exit:** the plan-13d criterion — `set_price` via curl in one terminal, the price cell changes in every open `/pricing` browser without reload; kill the server, restart, the page resyncs on reconnect. - -## Verification targets (after 14f) - -| Check | Target | How | -| --- | --- | --- | -| Golden corpus | every `.htmlx` in blog/ecommerce/pricing parses + renders byte-stable | `cargo test -p ui` golden files | -| SCSS | `pricing.scss` → golden CSS; errors carry file:line | `cargo test -p ui scss` | -| SSR | `GET /pricing` < 5 ms p99 on the dev box (RAM engine, no I/O on read path) | scripted curl loop | -| End-to-end | 13d criterion green | two-terminal demo, scripted in `just pricing-demo` | -| Single binary | UI + DB + API + WS in one `wo build` output, no Node anywhere | `ldd` shows libc only; no build-step JS | -| Dep budget | `crates/ui`: `serde`/`serde_json` only (dropped when phase 05 lands) | `Cargo.toml` review | - -## Non-scope - -- **No SPA router, no client-side templates.** Navigation is full-page SSR; only `wo:bind` cells and `` subtrees mutate in place. (ui/00 decision; unchanged.) -- **No SCSS mixins/functions/`@extend`/color math** beyond literal variable substitution — the subset is a floor, widened only by demonstrated need in the sample corpus. -- **No theme system / design tokens.** Shared partials under `ui/styles/_*.scss` are the only sharing mechanism for now. -- **No component framework.** `##component` partials render server-side via the 14a engine; they have no client behaviour beyond inherited `wo:bind` sites. -- **No changes to the REST API surface.** UI routes live beside `/api/*`; nothing under `/api` changes shape in this plan. - -## Cross-references - -- [`./exploration/ui/08-mvc-structure.md`](./exploration/ui/08-mvc-structure.md) — the design; its exit criteria are satisfied by 14c/14b/14f respectively. -- [`./13-class-model-live-pricing.md`](./13-class-model-live-pricing.md) — 13a/13b gate 14e; 13c gates 14f; 13d's exit criterion is this plan's end-to-end target. -- [`./exploration/ui/00-overview.md`](./exploration/ui/00-overview.md) — the UI track's master frame (per-app binaries, shared DB daemon) that 14d's asset/serving choices stay compatible with. -- [`reference/crates/wo-htmlx/`](../../.dev/reference/crates/wo-htmlx/) — primary port source (585 LOC), per ui/01. -- [`../examples/pricing/ui/pricing/`](../examples/pricing/ui/pricing/) — the reference triplet every sub-phase tests against. diff --git a/docs/plan/compiler/2026-08-15-employee-relations-query.md b/docs/plan/compiler/2026-08-15-employee-relations-query.md new file mode 100644 index 0000000..37de242 --- /dev/null +++ b/docs/plan/compiler/2026-08-15-employee-relations-query.md @@ -0,0 +1,226 @@ +# `@table`, Relations, Query — Implementation Plan (employee sample) + +> **Status: ⬜ pending** (story iteration 9b) — blocked on iteration 9's engine +> plan ([`2026-08-01-db-engine-binding.md`](../../superpowers/plans/2026-08-01-db-engine-binding.md)): +> Tasks 3–6 below consume its row storage, WAL, indexes and select subset. +> Board: [00-status.md](../../00-status.md) + +> **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (`- [ ]`) syntax for tracking. +> +> **Style rule (user convention):** concept, reason, and required behavior in words only; the executor writes the code. + +**Spec:** [`docs/superpowers/specs/2026-08-15-table-relations-query-design.md`](../../superpowers/specs/2026-08-15-table-relations-query-design.md) (normative: the settled forks, the clause grammar, the aggregate semantics table, the lowering model), amended by the systems-track spec's program-mode contract for the sample's CLI. + +**Goal:** `@table` classes queried **in the language**: comprehension queries +with `where`/`join`/`group`/`order`/`take`/`select`, typed `ref`/`backlink` +navigation with restrict integrity, and the LINQ aggregate vocabulary +(`count`/`sum`/`avg`/`min`/`max`, `GroupBy` as group-and-reduce) — all lowered +to bytecode loops over engine cursor builtins, proven by a new +`docs/examples/employee` sample whose acceptance script is the gate. + +**Architecture:** compiler front (`compiler/src/{lexer,parser,ast,types,owner,emit}.ml`) +for the query surface; the database engine (`database/src/` — its own +top-level directory per the 2026-08-15 decision, statically linked into wovm; +`table.c`/`db.c` from iteration 9, plus a new `query.c` for cursors and group +hashing). The compiler +is the planner: index selection happens at lowering, the VM never sees a plan +tree. Reference semantics: System.Linq for operator meaning, PostgreSQL's +nodeAgg/ri_triggers for execution and integrity vocabulary (both surveyed in +the spec, fork 3). + +**Tech Stack:** OCaml stdlib (compiler), C11 libc (runtime). No new opcodes; +new builtin ids appended to the format doc. + +## Global Constraints + +- **The sample is the test.** `docs/examples/employee` plus its acceptance + script is 9b's gate; the only corpus additions are the db-corpus fixtures + iteration 9 already plans. `just oop-e2e`, `just woc-test`, `just wovm-test` + stay green as regression after every task. +- **Index doctrine** (iteration 9, verbatim): secondary indexes are maintained + only through the engine's row choke points; FK checks and backlink reads are + index probes, never storage walks. +- **No SQL text anywhere** — the 9b acceptance's disassembly criterion. +- **No function values.** Every predicate/projection is expression syntax; a + query value is never deferred or passed around. +- Plans/specs in words; commits local only, never push; docs under `docs/`; + CODE-LOGIC.md updated beside changed code. + +## File Structure + +``` +compiler/src/lexer.ml parser.ml ast.ml query-expression grammar, clause AST (Task 1) +compiler/src/types.ml range/group scopes, navigation, projection synthesis (Task 2) +compiler/src/owner.ml emit.ml query ownership + lowering to cursor builtins (Task 5) +database/src/query.c query.h cursors, group hash, transition/finalize aggregates (Tasks 3, 4) +database/src/db.c table.c FK restrict checks at the row choke points (Task 3) +docs/examples/employee/ the acceptance workload (Task 6) +scripts/employee-accept.sh the gate (Task 6) +docs/plan/oop-vm/00-wob-format.md appended builtin ids (Tasks 3-5) +``` + +--- + +### Task 1: Query expressions parse + +**Concept & reason:** the comprehension grammar from spec section 3 — +`from`/`where`/`join`/`group…into`/`order by`/`take`/`skip`/`select`, clause +order fixed, aggregates as clause functions — becomes lexer keywords (contextual, +so `from`/`group`/`order` stay legal identifiers outside a query), a clause AST, +and parse diagnostics in a new WO-E5xx range (clause out of order, missing +`select`, aggregate named outside a query). Fixed clause order is a parse rule, +not a type rule, so the error lands on the exact token. + +- [ ] Grammar and AST for every clause; contextual keywords verified against + the existing samples (no `.wo` file in the tree breaks). +- [ ] WO-E5xx diagnostics with golden coverage in the existing `woc-test` + suite (dump-ast goldens for well-formed queries, diagnostic goldens for + each malformed shape). +- [ ] Gates green; commit locally. + +### Task 2: Queries typecheck + +**Concept & reason:** the surface's whole promise is compile-time checking. +`from e in Employee` opens a scope where `e`'s fields are the class's; `ref` +navigation substitutes the target class's field set (`e.dept.name`); +`backlink` reads type as `multi` of the source class; `group … into g` closes +the range scope and opens the group scope, where `g.key` has the key's type +and `g.f` is legal **only** inside an aggregate call; `select { … }` +synthesizes an anonymous record type so later use of a dropped column is a +compile error. Aggregate result types follow the spec's table exactly — +`count` is Int, `sum` Int, `avg`/`min`/`max` are `?T` because an empty group +is data. Unknown table, unknown column (naming the class), navigation through +a non-`ref`, bare group-member column, join sides not separable: each is its +own WO-E5xx with a golden. + +- [ ] Scope machinery for range/group variables; navigation typing both + directions; projection record synthesis interned like other class + shapes. +- [ ] Aggregate typing per the spec table; `?T` results force the existing + nil-handling style at use sites. +- [ ] Diagnostic goldens for every error class named above; gates green; + commit locally. + +### Task 3: Cursors and integrity in the engine + +**Concept & reason:** the runtime side queries need, built on iteration 9's +storage: a table-scan cursor (open by class id, advance, borrowed row view), +a primary-index point read (`ref` navigation), a secondary-index cursor with +longest-prefix probe (backlink reads, indexed `where`), all as builtins +appended to the format doc. Integrity lands at the row choke points where the +indexes already live: inserting/updating a non-nil `ref` probes the referenced +primary index and traps on a miss (foreign-key violation, sibling of the +unique trap); deleting a row still referenced traps (restrict) via the same +secondary index a `backlink` requires — one probe, no new structure. Nil +`ref` never probes (the MATCH SIMPLE rule); an update that leaves the key +unchanged skips the check (both from the PostgreSQL RI survey, mechanism +discarded — a check is an index probe, never query text). + +- [ ] Cursor builtins + borrowed-row-view lifetime rules written into the + binding doc (a row view never escapes the loop that opened the cursor — + the ownership pass enforces it, mirror of the container-read borrow). + **Rows have no borrow word** (they share the VM's field encoding, not + its header), so unlike every VM-heap borrow there is no runtime trap + behind this rule — the compile-time check is load-bearing alone, which + is why it gets its own diagnostic and goldens rather than riding on + E30x. +- [ ] Cursor stability per spec section 6: scans **materialize their id list** + before the body runs and point-read per iteration; updates through the + row view stay legal (exclusive row borrow, index maintenance at the row + API); `insert`/`delete` targeting a table with an open cursor is a new + WO-E5xx (the ownership pass carries the open-cursor table set through + the loop body); read-only nested queries over the same table stay legal. + The `raise` mode — updating an indexed column mid-scan — is the fixture + that proves the materialized-id semantics. +- [ ] FK trap + restrict trap wired through the row API; a debug-build probe + counter exposed for Task 5's index-selection proof. +- [ ] The `delete` statement (point delete of a row value) lands here too: + iteration 9's subset is insert/select/update-point, and restrict has + nothing to restrict without it — same typed-AST-plus-builtin shape as + insert, WAL remove record already specified by iteration 9's Task 2. +- [ ] db-corpus fixtures from iteration 9's plan extended with one FK-violation + and one restrict fixture (trap-code exact); ASan green; commit locally. + +### Task 4: Group hash + aggregate execution + +**Concept & reason:** the `AggregateBy` shape from the spec — group-and-reduce +in one pass, no group objects. A group hash table builtin set: create (keyed +by the group key's kind, nil legal), upsert-advance (locate-or-create the +group, advance each aggregate's transition slot), drain (iterate groups, +finalize, hand key + finals to the compiled projection loop). Transition/ +finalize is PostgreSQL's split: `avg` carries sum+count in its transition +state and divides at finalize; the "no row seen yet" state is distinct from +"transition value is nil" so nil-skipping aggregates need no first-row special +case. Aggregate semantics are the spec's table: empty ⇒ nil (or 0 for +count/sum), nil elements skipped, sum wraps, avg truncates. + +- [ ] Builtins implemented; the input projected to only the columns the + aggregates read before hashing (the nodeAgg memory lesson). +- [ ] Fixture-level verification through iteration 9's db corpus (one grouped + query, exact output) plus the sample in Task 6; ASan green; commit + locally. + +### Task 5: Lowering — the compiler is the planner + +**Concept & reason:** a query desugars to the bytecode loops the language +already has. Scan or probe chosen at compile time: leading `where` equality +conjuncts matched against declared indexes longest-prefix-first, probe emitted +on a hit, scan otherwise — and the choice is **demonstrated** via Task 3's +probe counter in the acceptance, not assumed. `join` builds a transient hash +on the inner side and probes with the outer (nil never inserted, never +probes). `group` emits the two-phase hash-aggregation loop from Task 4. +`order by` materializes and stable-sorts (original-index tiebreak — the LINQ +stability guarantee); `take`/`skip` slice the result. Ownership: row views +stay loop-bound borrows; everything `select` emits is copied/built at the +boundary under the established Text-copy and fresh-value rules — queries add +no new ownership classes, and the ownership pass's existing drop machinery +covers the query's temporaries because the lowering IS ordinary loops. + +- [ ] Desugar + lowering for every clause; disassembly of the sample's report + mode shows loops and builtins, no plan tree, no text. The select + boundary is the ownership bulkhead (spec section 6): everything a query + returns is copied or freshly built, so no value anywhere points into a + row slab after the query ends — asserted under ASan by mutating rows + after a query and re-reading the query's results. +- [ ] Index selection proven: the acceptance asserts probe-counter deltas for + the indexed `staff ` path versus a full-scan query. +- [ ] `oop-e2e`, `woc-test`, ASan-corpus gates green; commit locally. + +### Task 6: The employee sample is the acceptance + +**Concept & reason:** spec section 7 verbatim — `docs/examples/employee` with +`Department`/`Employee` (`@table`, `@unique` name, composite `[dept, salary]` +index, `ref`/`backlink` pair), wo.toml manifest so `woc .` builds it, a `just` +module beside it (the log-watcher convention), and `scripts/employee-accept.sh` +as the gate: `seed` (insert + WAL + duplicate-department trap on the second +run), `report` (headcount/avg/min/max by department, total payroll, ordered by +average salary — byte-exact lines), `staff` (both navigation directions + +index-probe proof), `raise` (update through a query; missing department ⇒ the +empty-query nil path), the restrict-trap demo, and a kill -9 between `seed` +and `report` proving replay on this workload. + +- [ ] Sample source pre-authored 2026-08-15 (`docs/examples/employee/` — + the target workload, sample-first like log-watcher was); this task makes + `woc docs/examples/employee` compile it with zero diagnostics and the + binary's modes run. The sample is authoritative: divergence between it + and the spec is resolved in the spec's favor and committed. +- [ ] `scripts/employee-accept.sh` + `just employee` module: every mode + checked with exact expectations, trap codes asserted, crash step + included; soak-style RSS/fd sampling reused from the log-watcher + script's pattern for the `report` loop. +- [ ] ASan run of the full acceptance: zero leaks (the log-watcher bar). +- [ ] Docs: README beside the sample, CODE-LOGIC.md updates beside changed + compiler/runtime code, format-doc builtin table final, board + story + rows updated with measured numbers; commit locally. + +## Out of scope — deferred by name + +- Everything spec section 8 lists: set operators, outer joins, subqueries, + composite group keys, groups as values, FK cascade/set-nil, deferred + checks, SQL text in any role, cross-shard queries, `LIVE`, migrations, + cost-based planning. +- Sorted-grouping and partial-sort optimizations (recorded LINQ/postgres + precedents; hash + full stable sort are this plan's only strategies). +- The ecommerce sample's query rewrite (9b's fifth acceptance criterion) — + it lands as its own follow-up once the employee gate is green, so this + plan's blast radius stays one new sample. diff --git a/docs/plan/discarded.md b/docs/plan/discarded.md index dc7ef25..a0257c2 100644 --- a/docs/plan/discarded.md +++ b/docs/plan/discarded.md @@ -52,3 +52,5 @@ Status board: [`00-status.md`](../00-status.md) · Doctrine: [`../00-principles. | **Per-example `principle.md` files** | 2026-08-08: one canonical repo-level [`docs/00-principles.md`](../00-principles.md) instead; examples link to it. | | **Minimal 3-file log-watcher sample** | Breaks the file-for-file `.hx` → `.wo` mapping and leaves the "could not express" column unproven — which is the sample's entire acceptance criterion. | | **Raw code in plan documents** | Plans carry concept, reason, and required behavior in words; the executor writes the code. | +| **`##ui` / `.htmlx` LiveView frontend track** | 2026-08-17: removed the 9-doc `exploration/ui/` design set, the `14-mvc-ui-implementation` plan, and the `ui-htmlx-live` plan. All were built on the non-advancing Rust runtime (`.dev/reference/crates/wo-htmlx`, `cargo run`, WebSocket live-patches) and contradict the current woc/wovm direction. The 13d pricing-UI row went with them. Revisit only if a UI story is re-opened on the woc/wovm stack. | +| **Old-runtime "front door" + v1 design docs** | 2026-08-17: removed `writeonce-pl.md`, `runtime/wo-language.md`, `future-scope/ai-agents-content-management.md`, the numbered v1 set `02-recovery`/`03-data`/`04-ui`/`05-datalayer`/`06-markdown-render`/`07-ssl`, and `runtime/database/05-go-sdk.md`. They pitched the old Rust `wo` runtime (REST + LiveView + SQL/Cypher) as the current language and contradicted the shipped woc/wovm toolchain. The `runtime/database/` design series is kept as cited design history; the Rust-track plans/`done` are kept per the status board. | diff --git a/docs/plan/done/02-event-loop-epoll.md b/docs/plan/done/02-event-loop-epoll.md index 9ca2a92..f0106f4 100644 --- a/docs/plan/done/02-event-loop-epoll.md +++ b/docs/plan/done/02-event-loop-epoll.md @@ -2,7 +2,7 @@ > **Status: ✅ done** (Rust Stage 2 — shipped, maintained, not advancing) — `runtime/netpoll_epoll.rs`: the hand-rolled `epoll` loop that replaced the async runtime. Board: [00-status.md](../../00-status.md) -**Context sources:** [`../01-problem.md`](../../01-problem.md), [`../02-recovery.md`](../../02-recovery.md), [`./linux/00-linux.md`](../exploration/linux/00-linux.md), [`./done/01-scafolding-crates.md`](01-scafolding-crates.md). +**Context sources:** [`../01-problem.md`](../../01-problem.md), `../02-recovery.md`, [`./linux/00-linux.md`](../exploration/linux/00-linux.md), [`./done/01-scafolding-crates.md`](01-scafolding-crates.md). ## Goal diff --git a/docs/plan/done/03-hand-rolled-http.md b/docs/plan/done/03-hand-rolled-http.md index c66f6b1..0265dc2 100644 --- a/docs/plan/done/03-hand-rolled-http.md +++ b/docs/plan/done/03-hand-rolled-http.md @@ -2,7 +2,7 @@ > **Status: ✅ done** (Rust Stage 2 — shipped, maintained, not advancing) — hand-rolled HTTP/1.1, plus keep-alive and pipelining. Board: [00-status.md](../../00-status.md) -**Context sources:** [`./02-event-loop-epoll.md`](./02-event-loop-epoll.md), [`./linux/00-linux.md`](../exploration/linux/00-linux.md), [`../02-recovery.md`](../../02-recovery.md). +**Context sources:** [`./02-event-loop-epoll.md`](./02-event-loop-epoll.md), [`./linux/00-linux.md`](../exploration/linux/00-linux.md), `../02-recovery.md`. ## Goal diff --git a/docs/plan/exploration/c-runtime/01-architecture.md b/docs/plan/exploration/c-runtime/01-architecture.md index b5206e3..ae7d820 100644 --- a/docs/plan/exploration/c-runtime/01-architecture.md +++ b/docs/plan/exploration/c-runtime/01-architecture.md @@ -146,4 +146,4 @@ Work stealing (breaks single-writer ACID), shared-heap locking (the doctrine exi - [`../../docs/plan/09-concurrency-scaleout.md`](../../09-concurrency-scaleout.md) — the thread-per-core doctrine. - [`../../docs/plan/exploration/linux/07-io_uring.md`](../linux/07-io_uring.md), [`08-mmap.md`](../linux/08-mmap.md) — the two shared-page mechanisms. - [`../../docs/plan/13-class-model-live-pricing.md`](../../13-class-model-live-pricing.md) — 13e's read-replica alternative, contrasted in improvement 1. -- [`../../docs/writeonce-pl.md`](../../../writeonce-pl.md) — the C/assembly "one address" pedagogy this doc extends to a full runtime. +- [`README.md`](../../../../README.md) — the C/assembly "one address" pedagogy the single-binary story extends to a full runtime. diff --git a/docs/plan/exploration/c-runtime/02-single-binary.md b/docs/plan/exploration/c-runtime/02-single-binary.md index a6a2be7..65e3bb2 100644 --- a/docs/plan/exploration/c-runtime/02-single-binary.md +++ b/docs/plan/exploration/c-runtime/02-single-binary.md @@ -1,6 +1,6 @@ # 02 — The end goal: the writeonce single binary on this runtime environment -**Context sources:** [`00-plan.md`](./00-plan.md) (the runtime-environment phases, A–B ✅), [`01-architecture.md`](./01-architecture.md) (the one-address trace), [`../../../runtime/wo-language.md`](../../../runtime/wo-language.md) ("one binary per project; no runtime to install on the target host"; `.wo` has "its own lexer, parser, analyzer, and bytecode"), [`../../09-concurrency-scaleout.md`](../../09-concurrency-scaleout.md)–[`12`](../../12-engine-disk-cutover.md) (the Rust product track this proves out), [`../../../runtime/database/02-wo-language.md`](../../../runtime/database/02-wo-language.md) (catalog + transaction semantics the payload carries). +**Context sources:** [`00-plan.md`](./00-plan.md) (the runtime-environment phases, A–B ✅), [`01-architecture.md`](./01-architecture.md) (the one-address trace), `../../../runtime/wo-language.md` ("one binary per project; no runtime to install on the target host"; `.wo` has "its own lexer, parser, analyzer, and bytecode"), [`../../09-concurrency-scaleout.md`](../../09-concurrency-scaleout.md)–[`12`](../../12-engine-disk-cutover.md) (the Rust product track this proves out), [`../../../runtime/database/02-wo-language.md`](../../../runtime/database/02-wo-language.md) (catalog + transaction semantics the payload carries). ## The end goal, stated once @@ -85,6 +85,6 @@ Steps 1–2 and 4–6 exist in `wo-rt-c` today with the notes store standing in ## Cross-references - [`00-plan.md`](./00-plan.md) — the kernel phases; [`01-architecture.md`](./01-architecture.md) — the one-address trace through the same stack. -- [`../../../runtime/wo-language.md`](../../../runtime/wo-language.md) — the user-facing single-binary promise this document implements. -- [`../../13-class-model-live-pricing.md`](../../13-class-model-live-pricing.md) (13b methods), [`../../14-mvc-ui-implementation.md`](../../14-mvc-ui-implementation.md) (UI assets) — the payload-side tracks. +- [`README.md`](../../../../README.md) — the user-facing single-binary promise this document implements. +- [`../../13-class-model-live-pricing.md`](../../13-class-model-live-pricing.md) (13b methods) — the payload-side track. - [`../../09-concurrency-scaleout.md`](../../09-concurrency-scaleout.md) — shard-key routing and 2PC the contract defers to. diff --git a/docs/plan/exploration/linux/00-linux.md b/docs/plan/exploration/linux/00-linux.md index d9a8a4d..2a00420 100644 --- a/docs/plan/exploration/linux/00-linux.md +++ b/docs/plan/exploration/linux/00-linux.md @@ -1,6 +1,6 @@ ## Linux Kernel Features -Kernel primitives that the writeonce binary can leverage, mapped to the architectural needs identified in [docs/01-problem.md](../../../01-problem.md) and [docs/02-recovery.md](../../../02-recovery.md). +Kernel primitives that the writeonce binary can leverage, mapped to the architectural needs identified in [docs/01-problem.md](../../../01-problem.md) and docs/02-recovery.md. ### Per-primitive reference cards diff --git a/docs/plan/exploration/ui/00-overview.md b/docs/plan/exploration/ui/00-overview.md deleted file mode 100644 index 45d5d29..0000000 --- a/docs/plan/exploration/ui/00-overview.md +++ /dev/null @@ -1,213 +0,0 @@ -# UI track — `.htmlx` live templates + Angular-style monorepo - -**Context sources:** [`docs/examples/ecommerce/ui/`](../../examples/ecommerce/ui/) (current `##ui` screens — storefront, order_tracker, admin_orders), [`docs/examples/ecommerce/types/`](../../examples/ecommerce/types/) + [`docs/examples/ecommerce/logic/`](../../examples/ecommerce/logic/) (the shared-schema + shared-fn anchor), [`reference/crates/wo-htmlx/`](../../../../.dev/reference/crates/wo-htmlx/) (v1 template engine — `{{path}}`, `{{#each}}`, `{{> partial}}`, `data-bind` attributes), [`templates/`](../../../../templates/) (v1 blog's concrete `.htmlx` usage), [`docs/runtime/database/06-lowcode-fullstack.md`](../../../runtime/database/06-lowcode-fullstack.md) (Phase 6's `##ui` + `##app` block spec). - -## Context - -Three threads converge into one plan: - -1. **`##ui` needs a concrete output format.** Phase 6's spec says screens "compile to a render tree" served as SSR HTML with a thin client runtime, but the actual template format isn't named. The v1 `.htmlx` engine at [`reference/crates/wo-htmlx/`](../../../../.dev/reference/crates/wo-htmlx/) already speaks `{{bindings}}`, `{{#each}}`, `{{> partials}}`, and `data-bind` attributes — it's 90% of what the new runtime needs and already has a working parser + renderer. Adopting it (and extending it with live-subscription semantics) is cheaper than inventing a new format. - -2. **The samples want a home that matches how real frontends are organised.** The ecommerce sample today is one flat directory with `types/`, `logic/`, and `ui/` beside each other. A real deployment has *multiple apps* against the same data: a customer storefront, an admin dashboard, a fulfillment console, maybe a read-only analytics viewer. Each has its own routes, its own policies, its own ideal binary shape. Angular (via Nx / Angular CLI workspaces) solved this with `apps/*` + `libs/*` on top of a shared root config — writeonce adopts the same shape. - -3. **Each app wants to be its own binary but share a database.** Running storefront and admin as one monolith conflates concerns: a CPU spike in admin stalls customer checkout; an admin auth bug opens customer data paths. Splitting into per-app binaries that share a single database backend (via the Phase-4 native wire protocol) gives blast-radius isolation without duplicating data. - -Intended outcome: `docs/examples/ecommerce/` refactors into a workspace with `shared/` + `apps/storefront/` + `apps/admin/`. `wo build apps/storefront` produces a `storefront` binary. `wo db serve` runs the shared database. The apps connect over `wo://…` and serve `.htmlx` SSR pages that subscribe to LIVE queries without a page reload. - -## Goal - -After this track's sub-phases land: - -- A **monorepo workspace** at `docs/examples/ecommerce/` with `shared/{types,logic,components}` + `apps/{storefront,admin}` structure. -- A **per-app binary** for each app under `apps/`: `wo build apps/storefront` → `./target/wo/storefront`, `wo build apps/admin` → `./target/wo/admin`. Each binary includes only its own `##ui` / `##app` / app-local types and logic; shared code compiles in by path reference. -- A **shared DB daemon** (`wo db serve`) — one process, no UI, just the engine and wire-protocol server. Each app binary connects as a client via the Phase-4 native protocol. -- **`.htmlx` as the compiled UI output.** Every `##ui` screen compiles into an `.htmlx` template file that the app binary serves; hand-written `.htmlx` files under `apps/X/ui/*.htmlx` are accepted as a first-class authoring alternative. -- **`.htmlx` subscribes.** A `` subtree in the template registers a LIVE query at page load; a ~20 KB client JS runtime patches DOM nodes on delta frames without reloading the page. -- **Per-app users + policies.** Each app declares its role set in `apps/X/app.wo`; row-level policies on shared types stay global, app-scoped policies layer on top per route. - -## Design decisions (locked) - -1. **`.htmlx` is the template format; `##ui` is the DSL that emits it.** Authors choose per-screen: declare `##ui #home { source: Product, columns: [...] }` in `.wo` and let the compiler produce `home.htmlx`; OR hand-write `home.htmlx` for a custom page. Both flow through the same `wo-htmlx` renderer. -2. **Extend v1 `.htmlx` with two new constructs** — `...` (subscription subtree) and `wo:bind="field"` (field-level live binding). The rest of the v1 syntax (`{{path}}`, `{{#each}}`, `{{> partial}}`) carries through unchanged. -3. **One binary per app, shared database process.** Not a shared library, not a monolith. Apps connect via the Phase-4 wire protocol (`wo://host:port`) — the same connection a Go/TS client would use. No in-process shared state between apps; their isolation is enforced by the OS process boundary. -4. **Angular-parallel workspace layout.** `apps/*` for deployable binaries, `shared/*` for libs shared across apps (types, logic, UI components), `wo.toml` at the workspace root. Each app also has its own `wo.toml` that names which `shared/` directories it depends on. -5. **File structure mirrors Angular component organisation.** Each UI screen lives in its own directory: `apps/storefront/ui/home/{home.wo, home.htmlx, home.css}`. Tests go in `home.test.wo`. Matches the Angular component pattern (`home.component.ts`, `home.component.html`, `home.component.scss`). -6. **Per-app routes, not per-screen routes.** `apps/storefront/app.wo` declares route table; each route maps to a `ui.` declared under `apps/storefront/ui/*/`. Cross-app navigation is an external redirect, not an internal route. -7. **Policy composition.** Global policies live in `shared/types/.wo` next to the `type` declaration (today). App-scoped policies live in `apps/X/app.wo` and AND with the global set — an admin app might relax a storefront policy for ops roles but can never relax beyond what the type's own policy permits. - -## Angular parallels — what writeonce copies, what it doesn't - -| Angular feature | Writeonce translation | Notes | -| --- | --- | --- | -| `nx workspace` / `angular.json` | Root `wo.toml` with `[workspace] apps = [...], shared = [...]` | Path references, not package registry | -| `apps//` | `apps//` with `app.wo` + `ui/` + local `types/` + local `logic/` | 1:1 naming | -| `libs//` | `shared//` | Used `shared/` instead of `libs/` — matches the more common monorepo idiom (Nx defaults to `libs`, but `shared` is clearer for this audience) | -| `.ts` + `.html` + `.scss` | `.wo` + `.htmlx` + `.css` under `ui//` | One-directory-per-screen | -| `ng build ` | `wo build apps/` | Per-app binary output | -| Dependency injection | Service-block resolution — `service rest` blocks in shared types are callable from any app by import | No runtime DI container; bindings are resolved at compile time | -| RxJS observables | LIVE subscription frames on a WebSocket | Declarative `live` attribute instead of imperative `.subscribe(...)` | -| Zone.js change detection | Per-row delta dispatch + field-level `wo:bind` | No full-tree change detection — only the rows/fields the delta names get repainted | -| `HttpClient` | Built-in wire-protocol client inside the app binary | No separate library to import; always present | - -### What we don't copy - -- **No TypeScript.** Authoring is `.wo` (for logic) + `.htmlx` (for templates) + `.css`. If a page needs bespoke JS interactivity beyond what `wo:bind` covers, a `` and consumed only by phase 03's client runtime. Schema below; `version: 1` is a constant for this phase. - -## Scope - -### New files inside `crates/ui/src/htmlx/` - -| File | Responsibility | Port source | -| --- | --- | --- | -| `mod.rs` | Re-exports `Template`, `Manifest`, `LiveSubscription`, `BindSite`, `ParseError`, `RenderError` | [`reference/crates/wo-htmlx/src/lib.rs`](../../../../.dev/reference/crates/wo-htmlx/src/lib.rs) (11 LOC) | -| `ast.rs` | Adds `Node::Live { attrs, body }` and `wo_bind: Option` on element nodes | [`reference/crates/wo-htmlx/src/ast.rs`](../../../../.dev/reference/crates/wo-htmlx/src/ast.rs) (18 LOC) — extend by ~50 LOC | -| `parser.rs` | Adds `` body in `
` for the runtime | [`reference/crates/wo-htmlx/src/render.rs`](../../../../.dev/reference/crates/wo-htmlx/src/render.rs) (140 LOC) — extend by ~70 LOC | -| `manifest.rs` | Walks the AST, collects subscriptions + bind sites, serialises JSON | new (~150 LOC) | - -Total: ~835 LOC (585 ported + ~250 new). - -### `Cargo.toml` change - -```toml -[dependencies] -serde = { version = "1", features = ["derive"] } -serde_json = "1" -``` - -(Both already in `crates/rt`; `crates/ui` adopts them rather than hand-rolling JSON for the manifest at this phase.) - -### Manifest schema - -```json -{ - "version": 1, - "subscriptions": [ - { - "id": "orders-live-0", - "source": "Order{ status != Cancelled }", - "key": "id", - "sort": "placed_at desc", - "filter": null, - "root_selector": "[data-wo-subscription=\"orders-live-0\"]" - } - ], - "bind_sites": [ - { "subscription_id": "orders-live-0", "key": "id", "field": "status" }, - { "subscription_id": "orders-live-0", "key": "id", "field": "total" } - ] -} -``` - -## API shape (target) - -```rust -use ui::htmlx::{Template, Manifest, HelperRegistry}; - -let tmpl = Template::parse(src)?; // Result -let html = tmpl.render(&ctx, ®istry)?; // Result -let mani = tmpl.manifest(); // Manifest - -// SSR pattern: page = head + html + "" -``` - -## Exit criteria - -1. `cargo build -p ui` green; no new top-level dependencies beyond `serde`/`serde_json`. -2. **Golden parse + render** for every `.htmlx` file under [`docs/examples/blog/ui/components/`](../../examples/blog/ui/components/) and [`docs/examples/ecommerce/shared/components/`](../../examples/ecommerce/shared/components/) — output is HTML and parses back into an isomorphic AST. -3. **Manifest emission** for `…` produces a `LiveSubscription` with the source string preserved verbatim and the body wrapped under `data-wo-subscription="orders-live-0"`. -4. **Bind-site collection** for `
` inside `` records `(subscription_id, key="id", field="status")` once and only once. -5. **v1 regression**: `cargo run --bin wo -- run docs/examples/blog` continues to start without parser errors. The blog sample has no `` or `wo:bind` today; nothing should regress. -6. All 14 existing `crates/rt` tests pass. - -## Non-scope - -- **No SSR.** Phase 02 emits templates; phase 03 ships the runtime; serving them is a downstream concern that the per-app binary in phase 05 wires together. -- **No nested ``.** Parse error in this phase. Author can compose live subtrees by partial inclusion (`{{> child}}`). -- **No author-extensible helpers.** The closed enum is the contract for this phase. -- **No streaming render.** Templates are rendered to a single `String`. -- **No manifest version negotiation.** Wire format is `version: 1` always. - -## Verification - -```bash -cargo build -p ui -cargo test -p ui --test golden_v1 # blog/ecommerce templates byte-identical -cargo test -p ui --test golden_extensions # + wo:bind cases -cargo test -p ui --test manifest # manifest emission - -# v1 regression -cargo run --bin wo -- run docs/examples/blog & -PID=$!; sleep 1; curl -fsS http://127.0.0.1:8080/ >/dev/null; kill $PID - -cd reference/crates && cargo build && cargo test -``` - -## After this phase - -Phase 02 (`02-ui-compiler.md`) consumes the `Template` + `Manifest` types defined here as its emission target — every `##ui` block compiles down to an `.htmlx` file that parses cleanly under this phase's parser. Phase 03 (`03-client-runtime.md`) consumes the manifest JSON schema as its wire input. diff --git a/docs/plan/exploration/ui/02-ui-compiler.md b/docs/plan/exploration/ui/02-ui-compiler.md deleted file mode 100644 index 292cdd4..0000000 --- a/docs/plan/exploration/ui/02-ui-compiler.md +++ /dev/null @@ -1,95 +0,0 @@ -# 02 — `##ui` → `.htmlx` compiler - -**Context sources:** [`./00-overview.md`](./00-overview.md) §§ "Sub-phase sequence" (L172–180), "Design decisions" 1–6, [`./01-htmlx-format-spec.md`](./01-htmlx-format-spec.md) (the emission target), [`../../runtime/database/06-lowcode-fullstack.md`](../../../runtime/database/06-lowcode-fullstack.md) (the `##ui` block spec), [`docs/examples/blog/ui/article_list.wo`](../../../examples/blog/ui/article_list.wo), [`docs/examples/blog/ui/article_detail.wo`](../../../examples/blog/ui/article_detail.wo), [`docs/examples/ecommerce/apps/admin/ui/orders/orders.wo`](../../examples/ecommerce/apps/admin/ui/orders/orders.wo) (the test corpus), [`crates/rt/src/parser.rs:80–116`](../../../../crates/rt/src/parser.rs) (the parse-and-discard call site to replace). - -## Goal - -Walk a parsed `##ui` block — every key listed in 00-overview's locked grammar (`title`, `source`, `live`, `role`, `filter`, `quick-filters`, `columns`, `sort`, `actions`, `pagination`, `refresh`, `highlight-new`, `key`, `sections`, `inputs`, `use`, `with`, `template`, `styles`) — and emit a complete `.htmlx` template that parses under phase 01's grammar. When a hand-written `.htmlx` sits beside the `.wo`, the compiler honours it and only validates that the manifest still aligns. - -## Design decisions (locked) - -1. **File-presence dispatch.** If `apps//ui//.htmlx` exists alongside `.wo`, the hand-written file wins. The compiler still emits a manifest; it errors if the manifest's `bind_sites` reference fields the hand-written template doesn't expose. Anchored in [`./00-overview.md`](./00-overview.md) L24, L30, L46. -2. **`live: true` ⇒ `` wrapper.** The compiler emits `` around the auto-generated `
{{status}}
`. `key` defaults to `id` when not declared. -3. **`renderer: ` ⇒ helper invocation by rule table.** A static rule table maps each renderer name to its emission form: `markdown` → `{{markdown }}`, `money` → `{{> money amount=}}`, `relative-date` → `{{relative }}`, `code` / `tag-chips` / `pill` / `image` / `stock-badge` / `list` similarly. The set is closed and matches phase 01's helper registry. -4. **`actions: row-* / bulk-*` ⇒ `data-action` + `data-role` attributes.** Click → POST `/api/fn/`. Role gating is a `data-role=""` attribute the runtime hides on; phase 07 wires the server-side check. -5. **Generated templates land in `target/wo//ui/.htmlx`.** Same path the per-app binary in phase 05 reads from at startup. Build artefact, not committed. -6. **`crates/rt/src/parser.rs:80–88` no longer skips `##ui`.** The `Kind::HashHash` arm parses into a `Screen` IR (this phase's new type). All other `##` blocks (`##app`, `##component`) keep their current `skip_top_level_chunk` behaviour for now — `##app` is owned by phase 05 and `##component` parses inline as a sibling of `##ui` but emits no template (it's a partial that other screens reference). - -## Scope - -### New files inside `crates/ui/src/compiler/` - -| File | Responsibility | Port source | -| --- | --- | --- | -| `mod.rs` | Re-exports `compile_screen`, `Screen`, `Column`, `Action`, `CompileError` | new (~30 LOC) | -| `screen.rs` | `Screen` IR — every key listed in the goal section above | new (~150 LOC) | -| `codegen.rs` | Walk `Screen` → emit `.htmlx` source string | new (~280 LOC) | -| `renderers.rs` | Closed `renderer:` → helper-emission rule table | new (~120 LOC) | -| `fallback.rs` | File-presence dispatch + manifest cross-check | new (~80 LOC) | - -### Modified file - -| File | Change | Notes | -| --- | --- | --- | -| `crates/rt/src/parser.rs` | Replace the `Kind::HashHash(_)` skip arm at L80–88 with a real parse into `Screen` when the tag is `ui` | +60 LOC delta | - -Total: ~720 LOC (all new — the v1 codebase has no `##ui` precedent to port). - -### `Cargo.toml` change - -`crates/ui` already depends on `serde`/`serde_json` from phase 01. No new deps. - -## API shape (target) - -```rust -use ui::compiler::{compile_screen, compile_app, Screen}; - -let screen: Screen = ql::parse_ui_block(src)?; -let template: String = compile_screen(&screen, &ctx)?; // an .htmlx string -let mani = ui::htmlx::Template::parse(&template)?.manifest(); - -// Whole-app pipeline used by phase 05's `wo build`: -let outputs: Vec<(PathBuf, String)> = compile_app(&app_dir)?; -for (path, src) in outputs { fs::write(path, src)?; } -``` - -## Exit criteria - -1. `cargo build -p ui` and `cargo build -p rt` green. -2. **Compile every sample `##ui`** in the test corpus: blog `article_list.wo`, blog `article_detail.wo`, ecommerce `apps/storefront/ui/home/home.wo`, `apps/storefront/ui/orders/orders.wo`, `apps/admin/ui/orders/orders.wo`. Output template parses cleanly under phase 01. -3. **Hand-written fallback honoured.** With a hand-written `apps/admin/ui/orders/orders.htmlx` present, the compiler returns its source unchanged but still emits the manifest. -4. **Manifest cross-check fires.** Renaming `body` to `text` in a hand-written template that the `##ui` block expects under `wo:bind="body"` produces a `CompileError::HandWrittenMissingField` diagnostic. -5. **Parser change is non-breaking.** `crates/rt`'s 14 unit tests still pass; `cargo run --bin wo -- run docs/examples/blog` boots and serves REST as before. -6. `cd reference/crates && cargo build && cargo test`. - -## Non-scope - -- **No SSR.** This phase emits files only — the runtime in phase 05 reads them at startup. -- **No author-defined renderers.** The closed table from phase 01's helper registry is the contract. -- **No build-output caching.** The compiler runs on every `wo build`. Caching is a future concern. -- **No partial recovery.** First parse or codegen error aborts compilation for the screen; whole-app compilation reports per-screen status. -- **No `##component` codegen in this phase.** Components remain their existing partial-include shape (`{{> name args}}`) — only `##ui` screens drive new emission. - -## Verification - -```bash -cargo build -p ui -p rt -cargo test -p ui --test compile_blog -cargo test -p ui --test compile_ecommerce -cargo test -p ui --test fallback_handwritten - -# end-to-end: emit templates for the storefront and inspect them -cargo run --bin wo -- build docs/examples/ecommerce/apps/storefront --emit-templates-only -ls target/wo/storefront/ui/ # home.htmlx, orders.htmlx -head -1 target/wo/storefront/ui/orders.htmlx # starts with - -# v1 regression -cargo run --bin wo -- run docs/examples/blog & -PID=$!; sleep 1; curl -fsS http://127.0.0.1:8080/ >/dev/null; kill $PID - -cd reference/crates && cargo build && cargo test -``` - -## After this phase - -Phase 03 (`03-client-runtime.md`) is now unblocked: every screen has a manifest the client runtime can read. Phase 04 (`04-workspace-layout.md`) places the compiled outputs under `target/wo//ui/`, which phase 05 then bakes into the per-app binary. diff --git a/docs/plan/exploration/ui/03-client-runtime.md b/docs/plan/exploration/ui/03-client-runtime.md deleted file mode 100644 index fe90e3d..0000000 --- a/docs/plan/exploration/ui/03-client-runtime.md +++ /dev/null @@ -1,109 +0,0 @@ -# 03 — Client runtime - -**Context sources:** [`./00-overview.md`](./00-overview.md) §§ "`.htmlx` with live subscriptions — target format" (L127–166) and decisions 1–2, [`./01-htmlx-format-spec.md`](./01-htmlx-format-spec.md) (the manifest schema this runtime consumes), [`reference/crates/wo-sub/src/lib.rs`](../../../../.dev/reference/crates/wo-sub/src/lib.rs) (the v1 frame model the wire format mirrors), [`docs/examples/ecommerce/shared/components/order-row.htmlx`](../../../examples/ecommerce/shared/components/order-row.htmlx) (the live workload the runtime must update without reload). - -## Goal - -Ship a ~500-line vanilla-JS client at `crates/ui/assets/wo-runtime.js` that, on page load, reads the `