From 41422d316bec2a49f2a1068da707b5926c22ca50 Mon Sep 17 00:00:00 2001 From: "shoney.arickathil" Date: Tue, 18 Aug 2026 01:50:20 +0200 Subject: [PATCH] chore: remove the old Rust runtime track from master MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Master now reflects only the current woc/wovm project. The Rust `wo` runtime was the prior, abandoned architecture; it is fully independent of the woc/wovm stack (dune + make, no Cargo dependency), so it lifts out cleanly. Recoverable via git history. Removed: - crates/ (29 files) + Cargo.toml + Cargo.lock — the Rust runtime workspace - prototypes/ — wo-rt-c (a stale duplicate of runtime/) + wo-db C++ ref - justfile: the rt-c-demo / rt-c-bench recipes (drove prototypes/wo-rt-c) - docs/plan/05..16 (11 Rust engineering plans) + docs/plan/done/ (4 Rust Stage-2 done plans) - docs/runtime/ (11): the old runtime overview + 7-phase DB design series + async/fibers/gc/surreal concept essays - docs/cm.md — legacy scratch note Kept: compiler/, runtime/ (C VM), database/, the current-track docs (stories, superpowers, plan/{oop-vm,compiler,exploration}, examples), the discarded/learnings registers, and the syscall/postgres/assembly/c-runtime studies. Follow-up commits fix the status board, project-structure doc, and any dangling links to the removed docs. Verified: woc + wovm still build; woc/wovm --version green. Co-Authored-By: Claude Opus 5 (1M context) --- Cargo.lock | 447 ------ Cargo.toml | 41 - crates/README.md | 54 - crates/rt/Cargo.toml | 22 - crates/rt/src/ast.rs | 204 --- crates/rt/src/bin/wo.rs | 209 --- crates/rt/src/compile.rs | 149 -- crates/rt/src/engine.rs | 748 ---------- crates/rt/src/http/connection.rs | 329 ----- crates/rt/src/http/listener.rs | 187 --- crates/rt/src/http/mod.rs | 21 - crates/rt/src/http/request.rs | 192 --- crates/rt/src/http/response.rs | 118 -- crates/rt/src/http/route.rs | 185 --- crates/rt/src/lexer.rs | 343 ----- crates/rt/src/lib.rs | 103 -- crates/rt/src/method.rs | 592 -------- crates/rt/src/mirror.rs | 270 ---- crates/rt/src/parser.rs | 1240 ----------------- crates/rt/src/pg.rs | 460 ------ crates/rt/src/runtime/eventfd.rs | 66 - crates/rt/src/runtime/mod.rs | 25 - crates/rt/src/runtime/netpoll_epoll.rs | 170 --- crates/rt/src/runtime/netpoll_io_uring.rs | 237 ---- crates/rt/src/runtime/scheduler.rs | 276 ---- crates/rt/src/runtime/signalfd.rs | 52 - crates/rt/src/runtime/timerfd.rs | 102 -- crates/rt/src/server.rs | 601 -------- crates/rt/src/shard.rs | 259 ---- crates/rt/src/token.rs | 139 -- crates/rt/src/wal.rs | 337 ----- docs/cm.md | 17 - docs/plan/05-hand-rolled-json.md | 116 -- docs/plan/06-bespoke-error.md | 122 -- docs/plan/07-inotify-content-watcher.md | 126 -- docs/plan/08-sendfile-static-assets.md | 110 -- docs/plan/09-concurrency-scaleout.md | 138 -- docs/plan/10-storage-foundations.md | 148 -- docs/plan/11-wal-and-recovery.md | 211 --- docs/plan/12-engine-disk-cutover.md | 175 --- docs/plan/13-class-model-live-pricing.md | 125 -- docs/plan/15-mcp-streamable-http.md | 124 -- docs/plan/16-postgres-mirror.md | 92 -- docs/plan/done/01-scafolding-crates.md | 70 - docs/plan/done/02-event-loop-epoll.md | 101 -- docs/plan/done/03-hand-rolled-http.md | 102 -- .../plan/done/04-cutover-remove-tokio-axum.md | 121 -- docs/runtime/async.md | 338 ----- docs/runtime/database.md | 60 - docs/runtime/database/01-evaluation.md | 270 ---- docs/runtime/database/02-wo-language.md | 508 ------- docs/runtime/database/03-inmemory-engine.md | 205 --- docs/runtime/database/04-client-api.md | 279 ---- docs/runtime/database/06-lowcode-fullstack.md | 389 ------ docs/runtime/database/07-wo-seg-migration.md | 207 --- docs/runtime/fibers.md | 704 ---------- docs/runtime/garbage-collection.md | 248 ---- docs/runtime/surreal-case-study.md | 157 --- justfile | 40 - prototypes/wo-db/README.md | 140 -- 60 files changed, 13621 deletions(-) delete mode 100644 Cargo.lock delete mode 100644 Cargo.toml delete mode 100644 crates/README.md delete mode 100644 crates/rt/Cargo.toml delete mode 100644 crates/rt/src/ast.rs delete mode 100644 crates/rt/src/bin/wo.rs delete mode 100644 crates/rt/src/compile.rs delete mode 100644 crates/rt/src/engine.rs delete mode 100644 crates/rt/src/http/connection.rs delete mode 100644 crates/rt/src/http/listener.rs delete mode 100644 crates/rt/src/http/mod.rs delete mode 100644 crates/rt/src/http/request.rs delete mode 100644 crates/rt/src/http/response.rs delete mode 100644 crates/rt/src/http/route.rs delete mode 100644 crates/rt/src/lexer.rs delete mode 100644 crates/rt/src/lib.rs delete mode 100644 crates/rt/src/method.rs delete mode 100644 crates/rt/src/mirror.rs delete mode 100644 crates/rt/src/parser.rs delete mode 100644 crates/rt/src/pg.rs delete mode 100644 crates/rt/src/runtime/eventfd.rs delete mode 100644 crates/rt/src/runtime/mod.rs delete mode 100644 crates/rt/src/runtime/netpoll_epoll.rs delete mode 100644 crates/rt/src/runtime/netpoll_io_uring.rs delete mode 100644 crates/rt/src/runtime/scheduler.rs delete mode 100644 crates/rt/src/runtime/signalfd.rs delete mode 100644 crates/rt/src/runtime/timerfd.rs delete mode 100644 crates/rt/src/server.rs delete mode 100644 crates/rt/src/shard.rs delete mode 100644 crates/rt/src/token.rs delete mode 100644 crates/rt/src/wal.rs delete mode 100644 docs/cm.md delete mode 100644 docs/plan/05-hand-rolled-json.md delete mode 100644 docs/plan/06-bespoke-error.md delete mode 100644 docs/plan/07-inotify-content-watcher.md delete mode 100644 docs/plan/08-sendfile-static-assets.md delete mode 100644 docs/plan/09-concurrency-scaleout.md delete mode 100644 docs/plan/10-storage-foundations.md delete mode 100644 docs/plan/11-wal-and-recovery.md delete mode 100644 docs/plan/12-engine-disk-cutover.md delete mode 100644 docs/plan/13-class-model-live-pricing.md delete mode 100644 docs/plan/15-mcp-streamable-http.md delete mode 100644 docs/plan/16-postgres-mirror.md delete mode 100644 docs/plan/done/01-scafolding-crates.md delete mode 100644 docs/plan/done/02-event-loop-epoll.md delete mode 100644 docs/plan/done/03-hand-rolled-http.md delete mode 100644 docs/plan/done/04-cutover-remove-tokio-axum.md delete mode 100644 docs/runtime/async.md delete mode 100644 docs/runtime/database.md delete mode 100644 docs/runtime/database/01-evaluation.md delete mode 100644 docs/runtime/database/02-wo-language.md delete mode 100644 docs/runtime/database/03-inmemory-engine.md delete mode 100644 docs/runtime/database/04-client-api.md delete mode 100644 docs/runtime/database/06-lowcode-fullstack.md delete mode 100644 docs/runtime/database/07-wo-seg-migration.md delete mode 100644 docs/runtime/fibers.md delete mode 100644 docs/runtime/garbage-collection.md delete mode 100644 docs/runtime/surreal-case-study.md delete mode 100644 prototypes/wo-db/README.md diff --git a/Cargo.lock b/Cargo.lock deleted file mode 100644 index 9a56072..0000000 --- a/Cargo.lock +++ /dev/null @@ -1,447 +0,0 @@ -# This file is automatically @generated by Cargo. -# It is not intended for manual editing. -version = 4 - -[[package]] -name = "anyhow" -version = "1.0.102" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "7f202df86484c868dbad7eaa557ef785d5c66295e41b460ef922eca0723b842c" - -[[package]] -name = "bitflags" -version = "2.11.0" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "843867be96c8daad0d758b57df9392b6d8d271134fce549de6ce169ff98a92af" - -[[package]] -name = "cfg-if" -version = "1.0.4" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "9330f8b2ff13f34540b44e946ef35111825727b38d33286ef986142615121801" - -[[package]] -name = "equivalent" -version = "1.0.2" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "877a4ace8713b0bcf2a4e7eec82529c029f1d0619886d18145fea96c3ffe5c0f" - -[[package]] -name = "errno" -version = "0.3.14" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "39cab71617ae0d63f51a36d69f866391735b51691dbda63cf6f96d042b63efeb" -dependencies = [ - "libc", - "windows-sys", -] - -[[package]] -name = "fastrand" -version = "2.3.0" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "37909eebbb50d72f9059c3b6d82c0463f2ff062c9e95845c43a6c9c0355411be" - -[[package]] -name = "foldhash" -version = "0.1.5" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "d9c4f5dac5e15c24eb999c26181a6ca40b39fe946cbe4c263c7209467bc83af2" - -[[package]] -name = "getrandom" -version = "0.4.2" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "0de51e6874e94e7bf76d726fc5d13ba782deca734ff60d5bb2fb2607c7406555" -dependencies = [ - "cfg-if", - "libc", - "r-efi", - "wasip2", - "wasip3", -] - -[[package]] -name = "hashbrown" -version = "0.15.5" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "9229cfe53dfd69f0609a49f65461bd93001ea1ef889cd5529dd176593f5338a1" -dependencies = [ - "foldhash", -] - -[[package]] -name = "hashbrown" -version = "0.16.1" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "841d1cc9bed7f9236f321df977030373f4a4163ae1a7dbfe1a51a2c1a51d9100" - -[[package]] -name = "heck" -version = "0.5.0" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "2304e00983f87ffb38b55b444b5e3b60a884b5d30c0fca7d82fe33449bbe55ea" - -[[package]] -name = "id-arena" -version = "2.3.0" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "3d3067d79b975e8844ca9eb072e16b31c3c1c36928edf9c6789548c524d0d954" - -[[package]] -name = "indexmap" -version = "2.13.0" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "7714e70437a7dc3ac8eb7e6f8df75fd8eb422675fc7678aff7364301092b1017" -dependencies = [ - "equivalent", - "hashbrown 0.16.1", - "serde", - "serde_core", -] - -[[package]] -name = "itoa" -version = "1.0.18" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "8f42a60cbdf9a97f5d2305f08a87dc4e09308d1276d28c869c684d7777685682" - -[[package]] -name = "leb128fmt" -version = "0.1.0" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "09edd9e8b54e49e587e4f6295a7d29c3ea94d469cb40ab8ca70b288248a81db2" - -[[package]] -name = "libc" -version = "0.2.183" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "b5b646652bf6661599e1da8901b3b9522896f01e736bad5f723fe7a3a27f899d" - -[[package]] -name = "linux-raw-sys" -version = "0.12.1" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "32a66949e030da00e8c7d4434b251670a91556f4144941d37452769c25d58a53" - -[[package]] -name = "log" -version = "0.4.29" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "5e5032e24019045c762d3c0f28f5b6b8bbf38563a65908389bf7978758920897" - -[[package]] -name = "memchr" -version = "2.8.0" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "f8ca58f447f06ed17d5fc4043ce1b10dd205e060fb3ce5b979b8ed8e59ff3f79" - -[[package]] -name = "once_cell" -version = "1.21.4" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "9f7c3e4beb33f85d45ae3e3a1792185706c8e16d043238c593331cc7cd313b50" - -[[package]] -name = "prettyplease" -version = "0.2.37" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "479ca8adacdd7ce8f1fb39ce9ecccbfe93a3f1344b3d0d97f20bc0196208f62b" -dependencies = [ - "proc-macro2", - "syn", -] - -[[package]] -name = "proc-macro2" -version = "1.0.106" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "8fd00f0bb2e90d81d1044c2b32617f68fcb9fa3bb7640c23e9c748e53fb30934" -dependencies = [ - "unicode-ident", -] - -[[package]] -name = "quote" -version = "1.0.45" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "41f2619966050689382d2b44f664f4bc593e129785a36d6ee376ddf37259b924" -dependencies = [ - "proc-macro2", -] - -[[package]] -name = "r-efi" -version = "6.0.0" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "f8dcc9c7d52a811697d2151c701e0d08956f92b0e24136cf4cf27b57a6a0d9bf" - -[[package]] -name = "rt" -version = "0.1.0" -dependencies = [ - "anyhow", - "libc", - "serde", - "serde_json", - "tempfile", -] - -[[package]] -name = "rustix" -version = "1.1.4" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "b6fe4565b9518b83ef4f91bb47ce29620ca828bd32cb7e408f0062e9930ba190" -dependencies = [ - "bitflags", - "errno", - "libc", - "linux-raw-sys", - "windows-sys", -] - -[[package]] -name = "semver" -version = "1.0.27" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "d767eb0aabc880b29956c35734170f26ed551a859dbd361d140cdbeca61ab1e2" - -[[package]] -name = "serde" -version = "1.0.228" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "9a8e94ea7f378bd32cbbd37198a4a91436180c5bb472411e48b5ec2e2124ae9e" -dependencies = [ - "serde_core", - "serde_derive", -] - -[[package]] -name = "serde_core" -version = "1.0.228" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "41d385c7d4ca58e59fc732af25c3983b67ac852c1a25000afe1175de458b67ad" -dependencies = [ - "serde_derive", -] - -[[package]] -name = "serde_derive" -version = "1.0.228" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "d540f220d3187173da220f885ab66608367b6574e925011a9353e4badda91d79" -dependencies = [ - "proc-macro2", - "quote", - "syn", -] - -[[package]] -name = "serde_json" -version = "1.0.149" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "83fc039473c5595ace860d8c4fafa220ff474b3fc6bfdb4293327f1a37e94d86" -dependencies = [ - "itoa", - "memchr", - "serde", - "serde_core", - "zmij", -] - -[[package]] -name = "syn" -version = "2.0.117" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "e665b8803e7b1d2a727f4023456bbbbe74da67099c585258af0ad9c5013b9b99" -dependencies = [ - "proc-macro2", - "quote", - "unicode-ident", -] - -[[package]] -name = "tempfile" -version = "3.27.0" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "32497e9a4c7b38532efcdebeef879707aa9f794296a4f0244f6f69e9bc8574bd" -dependencies = [ - "fastrand", - "getrandom", - "once_cell", - "rustix", - "windows-sys", -] - -[[package]] -name = "unicode-ident" -version = "1.0.24" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "e6e4313cd5fcd3dad5cafa179702e2b244f760991f45397d14d4ebf38247da75" - -[[package]] -name = "unicode-xid" -version = "0.2.6" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "ebc1c04c71510c7f702b52b7c350734c9ff1295c464a03335b00bb84fc54f853" - -[[package]] -name = "wasip2" -version = "1.0.2+wasi-0.2.9" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "9517f9239f02c069db75e65f174b3da828fe5f5b945c4dd26bd25d89c03ebcf5" -dependencies = [ - "wit-bindgen", -] - -[[package]] -name = "wasip3" -version = "0.4.0+wasi-0.3.0-rc-2026-01-06" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "5428f8bf88ea5ddc08faddef2ac4a67e390b88186c703ce6dbd955e1c145aca5" -dependencies = [ - "wit-bindgen", -] - -[[package]] -name = "wasm-encoder" -version = "0.244.0" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "990065f2fe63003fe337b932cfb5e3b80e0b4d0f5ff650e6985b1048f62c8319" -dependencies = [ - "leb128fmt", - "wasmparser", -] - -[[package]] -name = "wasm-metadata" -version = "0.244.0" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "bb0e353e6a2fbdc176932bbaab493762eb1255a7900fe0fea1a2f96c296cc909" -dependencies = [ - "anyhow", - "indexmap", - "wasm-encoder", - "wasmparser", -] - -[[package]] -name = "wasmparser" -version = "0.244.0" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "47b807c72e1bac69382b3a6fb3dbe8ea4c0ed87ff5629b8685ae6b9a611028fe" -dependencies = [ - "bitflags", - "hashbrown 0.15.5", - "indexmap", - "semver", -] - -[[package]] -name = "windows-link" -version = "0.2.1" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "f0805222e57f7521d6a62e36fa9163bc891acd422f971defe97d64e70d0a4fe5" - -[[package]] -name = "windows-sys" -version = "0.61.2" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "ae137229bcbd6cdf0f7b80a31df61766145077ddf49416a728b02cb3921ff3fc" -dependencies = [ - "windows-link", -] - -[[package]] -name = "wit-bindgen" -version = "0.51.0" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "d7249219f66ced02969388cf2bb044a09756a083d0fab1e566056b04d9fbcaa5" -dependencies = [ - "wit-bindgen-rust-macro", -] - -[[package]] -name = "wit-bindgen-core" -version = "0.51.0" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "ea61de684c3ea68cb082b7a88508a8b27fcc8b797d738bfc99a82facf1d752dc" -dependencies = [ - "anyhow", - "heck", - "wit-parser", -] - -[[package]] -name = "wit-bindgen-rust" -version = "0.51.0" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "b7c566e0f4b284dd6561c786d9cb0142da491f46a9fbed79ea69cdad5db17f21" -dependencies = [ - "anyhow", - "heck", - "indexmap", - "prettyplease", - "syn", - "wasm-metadata", - "wit-bindgen-core", - "wit-component", -] - -[[package]] -name = "wit-bindgen-rust-macro" -version = "0.51.0" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "0c0f9bfd77e6a48eccf51359e3ae77140a7f50b1e2ebfe62422d8afdaffab17a" -dependencies = [ - "anyhow", - "prettyplease", - "proc-macro2", - "quote", - "syn", - "wit-bindgen-core", - "wit-bindgen-rust", -] - -[[package]] -name = "wit-component" -version = "0.244.0" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "9d66ea20e9553b30172b5e831994e35fbde2d165325bec84fc43dbf6f4eb9cb2" -dependencies = [ - "anyhow", - "bitflags", - "indexmap", - "log", - "serde", - "serde_derive", - "serde_json", - "wasm-encoder", - "wasm-metadata", - "wasmparser", - "wit-parser", -] - -[[package]] -name = "wit-parser" -version = "0.244.0" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "ecc8ac4bc1dc3381b7f59c34f00b67e18f910c2c0f50015669dde7def656a736" -dependencies = [ - "anyhow", - "id-arena", - "indexmap", - "log", - "semver", - "serde", - "serde_derive", - "serde_json", - "unicode-xid", - "wasmparser", -] - -[[package]] -name = "zmij" -version = "1.0.21" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "b8848ee67ecc8aedbaf3e4122217aff892639231befc6a1b58d29fff4c2cabaa" diff --git a/Cargo.toml b/Cargo.toml deleted file mode 100644 index 618937d..0000000 --- a/Cargo.toml +++ /dev/null @@ -1,41 +0,0 @@ -# Root workspace — the new `.wo` runtime. -# -# Only `crates/rt` carries real code today (Stage 2 of the runtime); the -# fourteen sibling crates are empty skeletons populated phase-by-phase per -# docs/plan/done/01-scafolding-crates.md. They are commented out of the -# workspace until their phase activates — uncomment each one as code lands. -# -# The v1 writeonce blog crates at `reference/crates/` are a separate nested -# workspace, excluded here so the root build stays focused on the new runtime. - -[workspace] -resolver = "2" -members = [ - "crates/rt", - - # Uncomment as each phase extracts code from `rt/` into its target crate. - # See docs/plan/02..08 for the sequence. - # - # "crates/ql", # phase 02 target — lexer / parser / AST - # "crates/value", # phase 02 target — tagged Value + path helpers - # "crates/engine", # phase 02 target — rel/doc/graph executor - # "crates/txn", # phase 02 target — MVCC + RETURNING alias table - # "crates/db", # phase 02 target — top-level facade - # "crates/wal", # phase 03 target — io_uring + fsync WAL - # "crates/sub", # phase 04 target — LIVE subscriptions - # "crates/http", # phase 04 target — wire protocol + router - # "crates/gen", # phase 05 target — client SDK codegen - # "crates/policy", # phase 06 target — RBAC planner rewrites - # "crates/logic", # phase 06 target — triggers + fn interpreter - # "crates/service", # phase 06 target — endpoint dispatch - # "crates/ui", # phase 06 target — ##ui SSR + client runtime - # "crates/app", # phase 06 target — ##app manifest -] -exclude = [ - "reference/crates", -] - -[workspace.dependencies] -serde = { version = "1", features = ["derive"] } -serde_json = "1" -anyhow = "1" diff --git a/crates/README.md b/crates/README.md deleted file mode 100644 index 652a362..0000000 --- a/crates/README.md +++ /dev/null @@ -1,54 +0,0 @@ -# `crates/` — the `.wo` runtime - -Fifteen crates make up the new runtime. Only `rt/` carries real code today (Stage 2); the other fourteen are **empty placeholders** scaffolded to match the 7-phase design so each phase's extraction work becomes a mechanical code move into an existing home. - -> The crate-name prefix `wo-` was dropped when the active project namespaced itself under `wo` (the binary, the file extension, the language). Internal imports read cleanly: `use ql::Parser`, `use db::Tx`, `use http::router`. The v1 codebase keeps its `wo-*` prefix in [`reference/crates/`](../reference/crates/) to distinguish the generations. - -## Map - -| Phase | Crate | Purpose | Status | -| --- | --- | --- | --- | -| 2 | [`ql`](./ql/) | `.wo` grammar — lexer, parser, AST | placeholder | -| 2 | [`value`](./value/) | tagged `Value` + dotted-path helpers | placeholder | -| 2 | [`engine`](./engine/) | in-memory executor (rel / doc / graph) + schema catalog | placeholder | -| 2 | [`txn`](./txn/) | transaction coordinator — MVCC, `RETURNING` alias table | placeholder | -| 2 | [`db`](./db/) | top-level facade — `open()`, `Tx`, `Query`, `Subscribe` | placeholder | -| 3 | [`wal`](./wal/) | write-ahead log — io_uring + fsync + recovery | placeholder | -| 4 | [`sub`](./sub/) | live subscriptions — delta frames on commit | placeholder | -| 4 | [`http`](./http/) | wire protocol — REST / GraphQL-over-WS / native codec | placeholder | -| 5 | [`gen`](./gen/) | codegen — `.wo type` → Go / TS / Rust / Python clients | placeholder | -| 6 | [`policy`](./policy/) | RBAC + row-level rules compiled into planner rewrites | placeholder | -| 6 | [`logic`](./logic/) | `on ` triggers + `fn ... in txn` interpreter | placeholder | -| 6 | [`service`](./service/) | `service rest/graphql/native` endpoint dispatch | placeholder | -| 6 | [`ui`](./ui/) | `##ui` screens → SSR HTML + client runtime | placeholder | -| 6 | [`app`](./app/) | `##app` route manifest + startup hooks | placeholder | -| — | [`rt`](./rt/) | **active** — Stage-2 monolith + the `wo` binary | **shipped** | - -## Why `rt/` is monolithic right now - -`rt/` currently holds every module the runtime needs — lexer, parser, AST, in-memory engine, axum REST server — because **shipping working Stage 2 was more important than hitting the final crate layout on day one**. Each module inside `rt/src/` is written with a target home in mind: - -| `rt` module | Moves to | Phase | -| --- | --- | --- | -| `token.rs` + `lexer.rs` + `ast.rs` + `parser.rs` | `ql/` | 2 | -| `engine.rs` (Value + Row helpers) | `value/` | 2 | -| `engine.rs` (Engine + Catalog) | `engine/` | 2 | -| `compile.rs` | `engine/` | 2 | -| `server.rs` | `http/` + `service/` | 4 / 6 | -| `bin/wo.rs` | stays in `rt/` (the binary) | — | - -Extractions happen phase-by-phase — first one lands when a second caller appears (likely when Stage 3 needs the parser for raw-`.wo` HTTP requests). - -## Build & test - -```bash -cargo build # compiles all 15 crates -cargo test --lib # 14 unit tests (all in rt today) -cargo run --bin wo -- run docs/examples/blog # serve the blog sample -``` - -## What's outside this directory - -- [`../reference/crates/`](../reference/crates/) — the v1 writeonce blog (13 crates, nested workspace). Preserved for reference per [docs/runtime/database/07-wo-seg-migration.md](../docs/runtime/database/07-wo-seg-migration.md). Keeps its `wo-*` prefix. -- [`../prototypes/wo-db/`](../prototypes/wo-db/) — C++ prototype of the query-layer engine (~2k lines). The reference implementation this Rust port follows at the language level. -- [`../docs/plan/`](../docs/plan/) — planning documents for in-flight work (the `.md` files directly under `plan/` are upcoming phases; `plan/done/` holds completed ones). [`plan/done/01-scafolding-crates.md`](../docs/plan/done/01-scafolding-crates.md) is the authoritative scope doc for the 14 new placeholders. diff --git a/crates/rt/Cargo.toml b/crates/rt/Cargo.toml deleted file mode 100644 index 5b5694f..0000000 --- a/crates/rt/Cargo.toml +++ /dev/null @@ -1,22 +0,0 @@ -[package] -name = "rt" -version = "0.1.0" -edition = "2021" -description = "writeonce runtime — the `.wo` language engine (v2, Phase 2+ design)" - -[lib] -name = "rt" -path = "src/lib.rs" - -[[bin]] -name = "wo" -path = "src/bin/wo.rs" - -[dependencies] -anyhow = "1" -serde = { version = "1", features = ["derive"] } -serde_json = "1" -libc = "0.2" # phase 02 — direct epoll/eventfd/timerfd/signalfd syscalls - -[dev-dependencies] -tempfile = "3" diff --git a/crates/rt/src/ast.rs b/crates/rt/src/ast.rs deleted file mode 100644 index f1475cf..0000000 --- a/crates/rt/src/ast.rs +++ /dev/null @@ -1,204 +0,0 @@ -//! AST for `.wo` schema-layer declarations. -//! -//! Stage 2 scope: `type` declarations with fields and `service rest` blocks. -//! Policies, triggers, computed fields, link types, `##ui`, `##app`, `fn` bodies, -//! and `##sql/##doc/##graph` blocks are parsed-and-discarded for now — the AST -//! carries just enough to stand up a REST server that serves the declared types. - -#[derive(Debug, Clone, Default)] -pub struct Schema { - pub types: Vec, -} - -#[derive(Debug, Clone)] -pub struct TypeDecl { - pub name: String, - pub fields: Vec, - pub services: Vec, - /// Declared with `class` instead of `type`. Storage and REST are - /// class-blind (plan 13 decision 5); the flag gates method parsing — - /// `fn` members of a `class` compile into [`MethodDecl`]s (13b), while - /// `fn` inside a plain `type` keeps the 13a parse-and-discard behaviour. - pub is_class: bool, - /// Row-scoped methods (plan 13b). Only populated for classes. - pub methods: Vec, - /// Storage configuration from a type-level `@table(...)` annotation. - /// Every type/class IS a table regardless (plan 13 decisions 3/5) — - /// `@table` configures storage, it never toggles it. - pub table: TableCfg, -} - -/// `@table(name: "prices", index: [product, at], index: [sku])` — optional -/// storage configuration. `shard_key:`/`retention:` are reserved for later -/// phases and rejected by the parser until they land. -#[derive(Debug, Clone, Default)] -pub struct TableCfg { - /// Storage/table name override. Defaults to the type name. Consumed by - /// the SQL layer and plans 10–12; WAL records keep the type name as the - /// stable identifier. - pub name: Option, - /// Composite secondary indexes — one `index: [a, b]` entry each. - pub indexes: Vec>, -} - -/// `fn name(args) -> Ret [in txn [snapshot]] { body }` — a row-scoped -/// transactional function with an implicit `self` receiver (plan 13 -/// decision 2). Served over RPC as `POST /:id/`. -#[derive(Debug, Clone)] -pub struct MethodDecl { - pub name: String, - /// `(name, declared type)` — the type is diagnostic-only in Stage 2. - pub params: Vec<(String, String)>, - pub ret: Option, - pub txn: TxnMode, - pub body: Vec, -} - -/// Transaction annotation on a method. The Stage-2 engine is single-threaded -/// per shard, so every method already executes atomically and in isolation; -/// the mode is recorded for diagnostics and for the future MVCC coordinator. -#[derive(Debug, Clone, Copy, PartialEq, Eq)] -pub enum TxnMode { None, Txn, Snapshot, Serializable } - -/// Method-body statement — the schema-layer DML subset of -/// `02-wo-language.md § Schema-Layer DML` that 13b executes. -#[derive(Debug, Clone)] -pub enum Stmt { - /// `let name = expr` — binding is a snapshot (spec rule). - Let { name: String, expr: Expr }, - /// `insert Type { field: expr, ... }` — construction form. - Insert { ty: String, fields: Vec<(String, Expr)> }, - /// `return [expr]` - Return { expr: Option }, - /// `assert expr [otherwise abort ["msg"]]` — false aborts the txn. - Assert { cond: Expr, msg: Option }, - /// `if cond { ... } [else { ... }]` (else-if chains nest in `otherwise`). - If { cond: Expr, then: Vec, otherwise: Vec }, -} - -#[derive(Debug, Clone)] -pub enum Expr { - Int(i64), - Str(String), - Bool(bool), - Null, - /// Argument, `let` binding, or `self` (bound positionally — `self` stays - /// a plain identifier in the lexer, plan 13 decision 4). - Ident(String), - /// `base.field` — plain object access, or relation resolution when the - /// base is `self` and the field is a `multi`/`backlink` relation. - Field(Box, String), - /// `name(args)` — builtins: `latest`, `count`, `now`. - Call(String, Vec), - /// `select Type{ field == expr, other_field }` — schema-layer select per - /// § Brace Disambiguation: operator entries are predicates, bare idents - /// are projections. Evaluates to a set (array) of row objects, shaped by - /// the projection when one is given. Equality predicates route through - /// the engine's secondary indexes when the type declares a matching - /// `@table(index: ...)`. - Select { - ty: String, - predicates: Vec<(String, BinOp, Box)>, - projection: Vec, - }, - Unary(UnOp, Box), - Binary(BinOp, Box, Box), -} - -#[derive(Debug, Clone, Copy, PartialEq, Eq)] -pub enum UnOp { Neg, Not } - -#[derive(Debug, Clone, Copy, PartialEq, Eq)] -pub enum BinOp { - Add, Sub, Mul, Div, Mod, - Eq, Ne, Lt, Le, Gt, Ge, - And, Or, -} - -#[derive(Debug, Clone)] -pub struct Field { - pub name: String, - pub ty: FieldTy, - pub nullable:bool, - pub unique: bool, - pub default: Option, - /// True if this is a `ref T` / `multi T ...` / `backlink ...` field that doesn't - /// correspond to a storage column in this type's table. The server ignores it - /// for create/update/list scalar projection but exposes it via sub-endpoints. - pub is_relation: bool, -} - -#[derive(Debug, Clone)] -pub enum FieldTy { - /// Plain scalar — one of the well-known names (`Id`, `Text`, `Int`, etc.) or - /// an unrecognised identifier that's compiled as opaque text for Stage 2. - Scalar(String), - /// `[T]` — array of the inner type. - Array(Box), - /// `{ k: T, ... }` — inline embedded-document struct. - Struct(Vec), - /// Tagged union: `A | B | C`. Stage 2 stores variants as strings. - Union(Vec), - /// `ref TypeName` — scalar FK to another type. - Ref(String), - /// `multi TypeName @edge(:TAG)` — zero-prop graph edge. - MultiEdge { target: String, tag: Option }, - /// `multi TypeName via LinkType` — graph edge with properties. - MultiVia { target: String, link: String }, - /// `backlink TypeName.field` — inverse relation. - Backlink { target: String, field: String }, -} - -#[derive(Debug, Clone)] -pub enum DefaultExpr { - Str(String), - Int(i64), - Bool(bool), - Null, - Now, - Enum(String), // a bare identifier — e.g. `Customer` in a union default - Opaque(String), // anything we didn't bother to evaluate (computed, etc.) -} - -#[derive(Debug, Clone)] -pub struct ServiceDecl { - pub kind: ServiceKind, - pub path: String, - pub expose: Vec, -} - -#[derive(Debug, Clone, Copy, PartialEq, Eq)] -pub enum ServiceKind { Rest, Graphql, Native } - -#[derive(Debug, Clone, Copy, PartialEq, Eq, Hash)] -pub enum Operation { - List, - Get, - Create, - Update, - Delete, - Subscribe, - Me, - Custom, // anything we don't special-case — reserved -} - -impl Operation { - pub fn from_ident(s: &str) -> Operation { - match s { - "list" => Operation::List, - "get" => Operation::Get, - "create" => Operation::Create, - "update" => Operation::Update, - "delete" => Operation::Delete, - "subscribe" => Operation::Subscribe, - "me" => Operation::Me, - _ => Operation::Custom, - } - } -} - -impl Schema { - pub fn merge(&mut self, other: Schema) { - self.types.extend(other.types); - } -} diff --git a/crates/rt/src/bin/wo.rs b/crates/rt/src/bin/wo.rs deleted file mode 100644 index f0f06e5..0000000 --- a/crates/rt/src/bin/wo.rs +++ /dev/null @@ -1,209 +0,0 @@ -//! `wo` — the writeonce toolchain binary. -//! -//! Stage 2 scope: -//! wo run discover `.wo` files under , parse the type DSL, -//! compile a catalog, and serve REST CRUD on :8080. -//! wo --help print usage. -//! -//! After the phase-04 cutover this binary owned one event loop on one -//! thread; plan 09a upgrades that to thread-per-core — `WO_THREADS` pinned -//! workers, each with its own event loop and `SO_REUSEPORT` listener -//! (`rt::runtime::scheduler`). Engine state is still globally shared until -//! 09b. No tokio. See `docs/plan/09-concurrency-scaleout.md`. - -use std::path::PathBuf; -use std::process::ExitCode; - -fn usage() { - eprintln!( - "\ -wo — writeonce toolchain - -USAGE: - wo run parse .wo files under , serve REST CRUD on :8080 - wo --help print this message - -ENV: - WO_LISTEN override the listen address (default: 127.0.0.1:8080) - WO_PG postgres://user[:pass]@host[:port]/db — mirror every - committed write to Postgres as a backup (reads stay in - RAM; see docs/plan/16-postgres-mirror.md) -" - ); -} - -fn main() -> ExitCode { - let args: Vec = std::env::args().skip(1).collect(); - let slice: Vec<&str> = args.iter().map(|s| s.as_str()).collect(); - match slice.as_slice() { - [] | ["--help"] | ["-h"] => { usage(); ExitCode::from(0) } - ["run"] => { - eprintln!("wo run: directory argument required"); - usage(); - ExitCode::from(2) - } - ["run", dir] => match run(PathBuf::from(dir)) { - Ok(c) => c, - Err(e) => { eprintln!("error: {e}"); ExitCode::from(1) } - }, - _ => { usage(); ExitCode::from(2) } - } -} - -fn run(dir: PathBuf) -> anyhow::Result { - // 1. Discover .wo files. - let files = rt::discover(&dir)?; - if files.is_empty() { - anyhow::bail!("no .wo files found under {}", dir.display()); - } - println!("[wo] discovered {} .wo file{} under {}", - files.len(), - if files.len() == 1 { "" } else { "s" }, - dir.display(), - ); - - // 2. Parse each into a Schema. - let mut schemas = Vec::new(); - for f in &files { - match rt::parser::parse(&f.src) { - Ok(s) => { - let n = s.types.len(); - println!(" parsed {} — {} type{}", f.rel.display(), n, if n == 1 { "" } else { "s" }); - schemas.push(s); - } - Err(e) => { - eprintln!(" parse error in {}: {}", f.rel.display(), e); - return Ok(ExitCode::from(1)); - } - } - } - - // 3. Compile the catalog. - let catalog = rt::compile::Catalog::from_schemas(schemas)?; - println!("[wo] compiled catalog — {} type{}", catalog.order.len(), - if catalog.order.len() == 1 { "" } else { "s" }); - - // 4. Print the route banner. - println!(); - println!("[wo] routes:"); - print!("{}", rt::server::describe_routes(&catalog)); - - // 5. Serve — thread-per-core with a SHARDED engine (plan 09b): each - // worker owns its own Engine (interleaved ids) and its own router; - // cross-shard operations travel the shard bus (mailbox + eventfd). - // No Arc> anywhere. - let addr = std::env::var("WO_LISTEN").unwrap_or_else(|_| "127.0.0.1:8080".to_string()); - let n = rt::runtime::scheduler::thread_count(); - println!(); - println!("[wo] listening on http://{addr} — {n} shard{} (thread-per-core, SO_REUSEPORT, sharded engine)", - if n == 1 { "" } else { "s" }); - println!("[wo] ctrl-C to stop"); - - // Durability (plan 09c): per-shard WAL under WO_DATA (default ./wo-data; - // WO_DATA=off disables). A `meta` file pins the shard count — replaying - // a 4-shard data dir with WO_THREADS=8 would strand logs and break the - // id interleave, so a mismatch refuses to boot (resharding is 09f). - let data_dir = std::env::var("WO_DATA").unwrap_or_else(|_| "./wo-data".to_string()); - let durable = data_dir != "off"; - if durable { - std::fs::create_dir_all(&data_dir)?; - let meta = std::path::Path::new(&data_dir).join("meta"); - match std::fs::read_to_string(&meta) { - Ok(prev) => { - let prev: usize = prev.trim().parse().unwrap_or(0); - if prev != 0 && prev != n { - anyhow::bail!("{data_dir} was written with WO_THREADS={prev} — restart with that, or wipe the dir"); - } - } - Err(_) => std::fs::write(&meta, format!("{n}\n"))?, - } - } - - // Postgres backup mirror (plan 16b): RAM stays authoritative — the - // `wo-pg` thread receives every committed mutation on a bounded channel - // and upserts it as JSONB. Never in the ack path; off unless WO_PG set. - let mirror_tx = match std::env::var("WO_PG") { - Ok(url) => { - let cfg = rt::pg::PgConfig::from_url(&url) - .map_err(|e| anyhow::anyhow!("WO_PG: {e}"))?; - let tables: Vec<(String, String)> = catalog.order.iter() - .map(|name| (name.clone(), catalog.get(name).unwrap().storage_name.clone())) - .collect(); - let (tx, rx) = std::sync::mpsc::sync_channel(rt::mirror::QUEUE_CAP); - rt::mirror::spawn(cfg.clone(), rx, tables); - println!("[wo] postgres mirror: {}:{}/{} (backup only — reads stay in RAM)", - cfg.host, cfg.port, cfg.database); - Some(tx) - } - Err(_) => None, - }; - - let bus = rt::shard::ShardBus::new(n)?; - let catalog_for_workers = catalog.clone(); - rt::runtime::scheduler::serve(&addr, move |id| { - use std::os::unix::io::AsRawFd; - let mut engine = rt::engine::Engine::for_shard(catalog_for_workers.clone(), id, n); - let mut has_group_wal = false; - if durable { - let path = std::path::Path::new(&data_dir).join(format!("shard-{id}.rwal")); - let t0 = std::time::Instant::now(); - match rt::wal::Wal::open_and_replay(&path, &mut engine) { - Ok((wal, recs)) => { - if recs > 0 { - println!("[wo] shard {id}: replayed {recs} wal records in {:?}", t0.elapsed()); - } - // Group commit (io_uring): batch the tick's frames into - // one WRITE→FSYNC pair; acks ride the fsync CQE. Falls - // back to per-commit fsync if the ring is unavailable - // (or WO_GROUP_COMMIT=off, kept for A/B measurement). - let group_enabled = std::env::var("WO_GROUP_COMMIT").map(|v| v != "off").unwrap_or(true); - match (if group_enabled { rt::runtime::Uring::new(256) } else { Err(std::io::Error::other("disabled")) }) - .and_then(|ring| rt::wal::WalGroup::new(wal, ring)) - { - Ok(group) => { engine.attach_wal_group(group); has_group_wal = true; } - Err(e) => { - eprintln!("[wo] shard {id}: io_uring unavailable ({e}) — per-commit fsync"); - // wal moved; reopen in per-commit mode - let mut scratch = rt::engine::Engine::for_shard(catalog_for_workers.clone(), id, n); - if let Ok((w2, _)) = rt::wal::Wal::open_and_replay(&path, &mut scratch) { - engine.attach_wal(w2); - } - } - } - } - Err(e) => eprintln!("[wo] shard {id}: WAL unavailable ({e}) — running non-durable"), - } - } - // Mirror attaches AFTER replay: the replayed state goes to Postgres - // once, as a boot-time bulk sync, then live mutations stream. - if let Some(tx) = &mirror_tx { - engine.attach_mirror(tx.clone()); - engine.mirror_sync_all(); - } - let ctx = rt::shard::ShardCtx::new(id, n, engine, bus.clone()); - let router = rt::server::router(ctx.clone(), &catalog_for_workers); - let mail_fd = bus.mail_fd(id).as_raw_fd(); - let wal_hooks = if has_group_wal { - ctx.engine.borrow().wal_ring_fd().map(|rfd| { - let pump_ctx = ctx.clone(); - let unpark_ctx = ctx.clone(); - let park_ctx = ctx.clone(); - rt::runtime::scheduler::WalHooks { - ring_fd: rfd, - pump: Box::new(move || pump_ctx.wal_pump()), - unparks: Box::new(move || unpark_ctx.take_unparks()), - park_conn: Box::new(move |fd, gen| park_ctx.engine.borrow_mut().park_conn(fd, gen)), - } - }) - } else { None }; - let mail_ctx = ctx.clone(); - rt::runtime::scheduler::Worker { - router, - mail: Some((mail_fd, Box::new(move || mail_ctx.drain_inbox()))), - wal: wal_hooks, - } - })?; - - println!("[wo] all {n} shards joined — bye"); - Ok(ExitCode::from(0)) -} diff --git a/crates/rt/src/compile.rs b/crates/rt/src/compile.rs deleted file mode 100644 index 3b8bb11..0000000 --- a/crates/rt/src/compile.rs +++ /dev/null @@ -1,149 +0,0 @@ -//! Compile parsed [`Schema`] objects into a runtime [`Catalog`] — the engine's -//! view of the declared types, their storage columns, and the REST routes we -//! need to expose. - -use crate::ast::*; -use anyhow::{bail, Result}; -use std::collections::HashMap; - -/// A catalog of all compiled types, ready to feed to the engine and the server. -#[derive(Debug, Clone, Default)] -pub struct Catalog { - pub types: HashMap, - /// Preserves declaration order so REST route registration is deterministic. - pub order: Vec, -} - -#[derive(Debug, Clone)] -pub struct CompiledType { - pub name: String, - /// Columns in source order. Includes relation fields; the server treats - /// those specially but keeps them in the type's row object for echo. - pub fields: Vec, - pub services: Vec, - /// Row-scoped methods (plan 13b) — populated for `class` declarations, - /// empty for plain `type`s. Storage stays class-blind; methods only add - /// RPC routes on top. - pub methods: Vec, - /// True iff the type declared an `id: Id` column. Auto-populated on insert. - pub has_id: bool, - /// Storage/table name — `@table(name: "...")` override or the type name. - /// Metadata for the SQL layer and plans 10–12; the engine and WAL key - /// everything by type name (the stable identifier). - pub storage_name: String, - /// Composite secondary indexes from `@table(index: [...])`, validated - /// against the fields. The engine maintains these on every mutation. - pub indexes: Vec>, -} - -impl Catalog { - pub fn from_schemas(schemas: Vec) -> Result { - let mut cat = Catalog::default(); - let mut storage_names = std::collections::HashSet::new(); - for s in schemas { - for t in s.types { - if cat.types.contains_key(&t.name) { - bail!("duplicate type declaration: {}", t.name); - } - let has_id = t.fields.iter().any(|f| - f.name == "id" && - matches!(f.ty, FieldTy::Scalar(ref n) if n == "Id") - ); - - // @table validation (plan 13 follow-up): the name must be - // catalog-unique; index columns must be stored scalar - // columns (scalars, unions, and `ref` FKs — not relations - // without a column, not arrays/structs). - let storage_name = t.table.name.clone().unwrap_or_else(|| t.name.clone()); - if !storage_names.insert(storage_name.clone()) { - bail!("{}: @table name \"{storage_name}\" is already used by another type", - t.name); - } - for cols in &t.table.indexes { - for col in cols { - let Some(f) = t.fields.iter().find(|f| &f.name == col) else { - bail!("{}: @table index names unknown field `{col}`", t.name); - }; - match &f.ty { - FieldTy::Scalar(_) | FieldTy::Union(_) | FieldTy::Ref(_) => {} - FieldTy::MultiEdge { .. } | FieldTy::MultiVia { .. } - | FieldTy::Backlink { .. } => bail!( - "{}: @table index field `{col}` is a relation without a \ - stored column — index the `ref` side instead", t.name), - FieldTy::Array(_) | FieldTy::Struct(_) => bail!( - "{}: @table index field `{col}` is not a scalar column", t.name), - } - } - } - - cat.order.push(t.name.clone()); - cat.types.insert(t.name.clone(), CompiledType { - name: t.name.clone(), - fields: t.fields, - services: t.services, - methods: t.methods, - has_id, - storage_name, - indexes: t.table.indexes, - }); - } - } - Ok(cat) - } - - pub fn get(&self, name: &str) -> Option<&CompiledType> { - self.types.get(name) - } -} - -#[cfg(test)] -mod tests { - use super::*; - use crate::parser::parse; - - #[test] - fn table_annotation_validation() { - // Duplicate storage names collide across types. - let sch = parse("@table(name: \"t\")\ntype A { id: Id }\n@table(name: \"t\")\ntype B { id: Id }").unwrap(); - let err = Catalog::from_schemas(vec![sch]).unwrap_err().to_string(); - assert!(err.contains("already used"), "{err}"); - - // A name override colliding with another type's default name. - let sch = parse("@table(name: \"B\")\ntype A { id: Id }\ntype B { id: Id }").unwrap(); - assert!(Catalog::from_schemas(vec![sch]).is_err()); - - // Index on a missing field. - let sch = parse("@table(index: [nope])\ntype C { id: Id }").unwrap(); - let err = Catalog::from_schemas(vec![sch]).unwrap_err().to_string(); - assert!(err.contains("unknown field `nope`"), "{err}"); - - // Index on a relation without a stored column. - let sch = parse("@table(index: [prices])\nclass P { id: Id\n prices: multi Price }").unwrap(); - let err = Catalog::from_schemas(vec![sch]).unwrap_err().to_string(); - assert!(err.contains("relation without a stored column"), "{err}"); - - // Valid: ref FK + scalar composite; storage_name defaults to type name. - let sch = parse( - "@table(name: \"prices\", index: [product, at])\nclass Price { id: Id\n product: ref Product\n at: Text }" - ).unwrap(); - let cat = Catalog::from_schemas(vec![sch]).unwrap(); - let t = cat.get("Price").unwrap(); - assert_eq!(t.storage_name, "prices"); - assert_eq!(t.indexes, vec![vec!["product".to_string(), "at".to_string()]]); - } - - #[test] - fn catalog_collects_types_and_detects_id() { - let sch = parse(r#" -type Article { id: Id - title: Text - service rest "/api/articles" expose list, get } -type Tag { slug: Slug - service rest "/api/tags" expose list } -"#).unwrap(); - let cat = Catalog::from_schemas(vec![sch]).unwrap(); - assert_eq!(cat.order, vec!["Article", "Tag"]); - assert!(cat.types["Article"].has_id); - assert!(!cat.types["Tag"].has_id); - } -} diff --git a/crates/rt/src/engine.rs b/crates/rt/src/engine.rs deleted file mode 100644 index 7bb58ff..0000000 --- a/crates/rt/src/engine.rs +++ /dev/null @@ -1,748 +0,0 @@ -//! In-memory CRUD engine for the compiled schema. -//! -//! Storage model: `HashMap>`. Rows are -//! `serde_json::Value::Object`. `Id` columns are auto-populated on insert. -//! All queries are plain iteration — fine for Stage 2. - -use crate::ast::{DefaultExpr, FieldTy}; -use crate::compile::{Catalog, CompiledType}; - -use anyhow::Result; -use serde_json::{json, Map, Value}; -use std::collections::{BTreeMap, BTreeSet}; -use std::time::{SystemTime, UNIX_EPOCH}; - -pub type Row = Map; - -/// Comparable encoding of an indexed column value — the key space of the -/// secondary indexes (`@table(index: [...])`). Ordering: Null < Bool < Int -/// < Str, then natural order within each. Residual filters always re-check -/// with real JSON equality, so encoding collisions cannot produce wrong -/// results — only wasted candidates. -#[derive(Debug, Clone, PartialEq, Eq, PartialOrd, Ord)] -enum IndexKey { - Null, - Bool(bool), - Int(i64), - Str(String), -} - -impl IndexKey { - fn from_value(v: Option<&Value>) -> IndexKey { - match v { - None | Some(Value::Null) => IndexKey::Null, - Some(Value::Bool(b)) => IndexKey::Bool(*b), - Some(other) => match other.as_i64() { - Some(n) => IndexKey::Int(n), - None => match other { - Value::String(s) => IndexKey::Str(s.clone()), - v => IndexKey::Str(v.to_string()), - }, - }, - } - } -} - -/// One composite secondary index: ordered key tuples → row ids. Per-shard, -/// in RAM, maintained incrementally by [`Engine::row_insert`]/[`row_remove`]. -#[derive(Debug)] -struct Index { - cols: Vec, - map: BTreeMap, BTreeSet>, -} - -impl Index { - fn key_for(&self, row: &Row) -> Vec { - self.cols.iter().map(|c| IndexKey::from_value(row.get(c))).collect() - } -} - -#[derive(Debug, Default)] -pub struct Engine { - catalog: Catalog, - /// type_name → { id → row } - tables: std::collections::HashMap>, - /// type_name → its secondary indexes (`@table(index: [...])`). Only - /// mutated by `row_insert`/`row_remove` — every table mutation path - /// (CRUD, replay, txn undo) goes through those two helpers. - indexes: std::collections::HashMap>, - /// per-type id allocator - next_id: std::collections::HashMap, - /// id stride — 1 for a standalone engine, `n_shards` for a 09b shard so - /// ids interleave (shard t mints t+1, t+1+n, …) and the owner of any id - /// is recoverable as `(id-1) % n` with zero coordination. - id_step: i64, - /// Per-shard write-ahead log (plan 09c). `PerCommit` fsyncs inside the - /// mutating call (simple, used by tests); `Group` stages frames and - /// parks acks for the worker's per-tick io_uring flush — the C - /// prototype's phase-D group commit. - wal: Option, - /// Set when the last mutation staged a group-commit frame — the caller - /// (handler or shard job) must park its ack. Cleared by `take_staged`. - staged: bool, - /// Active method transaction (plan 13b): mutations defer their WAL - /// records here and journal undo entries; `commit_txn` emits one - /// [`WalRec::Txn`] frame, `abort_txn` reverts RAM in reverse order. - txn: Option, - /// Postgres backup mirror (plan 16b): committed mutations are cloned - /// onto this channel AFTER the WAL accepted them — the mirror never - /// gates an ack. `None` when `WO_PG` is unset. - mirror: Option, - /// Records dropped because the mirror channel was full/closed — - /// counted per shard, logged loudly but never blocking. - mirror_dropped: u64, -} - -#[derive(Debug, Default)] -struct TxnState { - wal: Vec, - undo: Vec, - /// Mirror records for this transaction — sent as ONE - /// [`crate::mirror::MirrorRec::Txn`] on commit, dropped on abort. - mirror: Vec, -} - -/// Inverse of one applied mutation — enough to restore the pre-txn RAM state. -#[derive(Debug)] -enum Undo { - Created { ty: String, id: i64 }, - Updated { ty: String, id: i64, prev: Row }, - Deleted { ty: String, id: i64, row: Row }, -} - -#[derive(Debug)] -enum WalBackend { - PerCommit(crate::wal::Wal), - Group(crate::wal::WalGroup), -} - -impl Engine { - pub fn new(catalog: Catalog) -> Self { - Self::for_shard(catalog, 0, 1) - } - - /// One shard of a thread-per-core deployment (plan 09b): same engine, - /// interleaved id minting. - pub fn for_shard(catalog: Catalog, shard: usize, n_shards: usize) -> Self { - let mut tables = std::collections::HashMap::new(); - let mut next_id = std::collections::HashMap::new(); - let mut indexes = std::collections::HashMap::new(); - for name in catalog.order.iter() { - tables.insert(name.clone(), BTreeMap::new()); - next_id.insert(name.clone(), shard as i64 + 1); - let t = catalog.get(name).expect("type present"); - if !t.indexes.is_empty() { - indexes.insert(name.clone(), t.indexes.iter().map(|cols| Index { - cols: cols.clone(), - map: BTreeMap::new(), - }).collect()); - } - } - Self { catalog, tables, indexes, next_id, id_step: n_shards.max(1) as i64, - wal: None, staged: false, txn: None, mirror: None, mirror_dropped: 0 } - } - - /// Attach the Postgres backup mirror (plan 16b). Like `attach_wal`, - /// this happens AFTER boot replay — replayed rows are pushed once via - /// [`mirror_sync_all`], not re-mirrored record by record. - pub fn attach_mirror(&mut self, tx: crate::mirror::MirrorSender) { - self.mirror = Some(tx); - } - - /// Enqueue this shard's ENTIRE current state as upserts — boot-time - /// initial sync so a fresh Postgres catches up with a replayed WAL. - pub fn mirror_sync_all(&mut self) { - if self.mirror.is_none() { return; } - let snapshot: Vec<(String, i64, Row)> = self.tables.iter() - .flat_map(|(ty, table)| table.iter() - .map(|(id, row)| (ty.clone(), *id, row.clone()))) - .collect(); - for (ty, id, row) in snapshot { - self.mirror_dispatch(crate::mirror::MirrorRec::Upsert { ty, id, row }); - } - } - - /// Route one committed mutation to the mirror: buffered while a method - /// transaction is open (sent atomically on commit), dispatched - /// immediately otherwise. No-op when no mirror is attached. - fn mirror_send(&mut self, rec: crate::mirror::MirrorRec) { - if self.mirror.is_none() { return; } - match self.txn.as_mut() { - Some(t) => t.mirror.push(rec), - None => self.mirror_dispatch(rec), - } - } - - /// Non-blocking send; a full or closed channel drops the record and - /// counts it — Postgres lags, clients never do (16b policy; the - /// dirty-flag resync is plan 16d). - fn mirror_dispatch(&mut self, rec: crate::mirror::MirrorRec) { - let Some(tx) = self.mirror.as_ref() else { return }; - if tx.try_send(rec).is_err() { - self.mirror_dropped += 1; - if self.mirror_dropped.is_power_of_two() { - eprintln!("[wo] pg mirror: queue full/closed — {} records dropped on this \ - shard (Postgres is behind RAM until resync, plan 16d)", - self.mirror_dropped); - } - } - } - - /// Attach a per-commit WAL (fsync inside each mutation). Must happen - /// AFTER replay — replayed mutations must not be re-logged. - pub fn attach_wal(&mut self, wal: crate::wal::Wal) { - self.wal = Some(WalBackend::PerCommit(wal)); - } - - /// Attach a group-commit WAL (io_uring): mutations stage frames; the - /// worker flushes once per tick and releases parked acks on the CQE. - pub fn attach_wal_group(&mut self, wal: crate::wal::WalGroup) { - self.wal = Some(WalBackend::Group(wal)); - } - - /// Did the last mutation stage a group-commit frame? (Cleared on read.) - /// The caller must park its ack on the batch when this is true. - pub fn take_staged(&mut self) -> bool { - std::mem::take(&mut self.staged) - } - - /// Park a cross-shard reply on the active batch — sent on fsync. - pub fn park_reply(&mut self, cb: Box) { - match self.wal.as_mut() { - Some(WalBackend::Group(g)) => g.park(crate::wal::Parked::Reply(cb)), - _ => cb(), // no group WAL: durability already settled (or off) - } - } - - /// Park a local connection's response on the active batch. - pub fn park_conn(&mut self, fd: std::os::unix::io::RawFd, gen: u64) { - if let Some(WalBackend::Group(g)) = self.wal.as_mut() { - g.park(crate::wal::Parked::Conn { fd, gen }); - } - } - - /// Worker hooks — flush at tick end; reap on ring-fd readable. - pub fn wal_flush(&mut self) { - if let Some(WalBackend::Group(g)) = self.wal.as_mut() { - if let Err(e) = g.flush() { eprintln!("[wo] wal flush: {e}"); } - } - } - - pub fn wal_complete(&mut self) -> Option<(bool, Vec)> { - match self.wal.as_mut() { - Some(WalBackend::Group(g)) => g.complete(), - _ => None, - } - } - - pub fn wal_ring_fd(&self) -> Option { - match self.wal.as_ref() { - Some(WalBackend::Group(g)) => Some(g.ring_fd()), - _ => None, - } - } - - /// Apply one replayed WAL record. Bypasses default-seeding and id - /// minting — the log carries exact state — but advances the id - /// high-water mark so post-recovery mints never collide. - pub fn replay(&mut self, rec: &crate::wal::WalRec) { - use crate::wal::WalRec; - match rec { - WalRec::Create { ty, row } => { - let Some(id) = row.get("id").and_then(|v| v.as_i64()) else { return }; - if self.tables.contains_key(ty) { - self.row_insert(ty, id, row.clone()); - let step = self.id_step; - let counter = self.next_id.entry(ty.clone()).or_insert(1); - while *counter <= id { *counter += step; } - } - } - WalRec::Update { ty, id, body } => { - // Remove-then-insert keeps the secondary indexes in step. - let Some(mut row) = self.row_remove(ty, *id) else { return }; - if let Value::Object(input) = body { - for (k, v) in input { - if k != "id" { row.insert(k.clone(), v.clone()); } - } - } - self.row_insert(ty, *id, row); - } - WalRec::Delete { ty, id } => { - self.row_remove(ty, *id); - } - // A method's mutations — the frame validated whole, apply all. - WalRec::Txn { recs } => { - for r in recs { self.replay(r); } - } - } - } - - // --- method transactions (plan 13b) --- - - /// Enter method-transaction mode: subsequent mutations journal undo - /// entries and defer their WAL records until [`commit_txn`]. - pub fn begin_txn(&mut self) -> Result<()> { - if self.txn.is_some() { - anyhow::bail!("nested method transactions are not supported"); - } - self.txn = Some(TxnState::default()); - Ok(()) - } - - /// Commit the active transaction: every deferred record leaves as ONE - /// `WalRec::Txn` frame (atomic on replay). `Err` means the WAL rejected - /// the frame — RAM has been rolled back and the caller must not ack. - pub fn commit_txn(&mut self) -> Result<()> { - let Some(t) = self.txn.take() else { - anyhow::bail!("commit_txn without begin_txn"); - }; - if t.wal.is_empty() { return Ok(()); } // read-only method — nothing to log - if let Err(e) = self.wal_log(crate::wal::WalRec::Txn { recs: t.wal }) { - self.apply_undo(t.undo); // never ack non-durable - return Err(e); - } - // Mirror the whole method as one atomic Postgres transaction — - // committed only; an aborted method never reaches this point. - if !t.mirror.is_empty() { - self.mirror_dispatch(crate::mirror::MirrorRec::Txn(t.mirror)); - } - Ok(()) - } - - /// Abort the active transaction: revert RAM in reverse order, log nothing. - pub fn abort_txn(&mut self) { - if let Some(t) = self.txn.take() { - self.apply_undo(t.undo); - } - } - - fn apply_undo(&mut self, undo: Vec) { - for u in undo.into_iter().rev() { - match u { - Undo::Created { ty, id } => { - self.row_remove(&ty, id); - } - Undo::Updated { ty, id, prev } | Undo::Deleted { ty, id, row: prev } => { - // Clear the current version's index keys (if any row is - // present) before restoring the previous one. - self.row_remove(&ty, id); - self.row_insert(&ty, id, prev); - } - } - } - } - - /// Make a mutation durable (per-commit) or stage it (group). `Err` - /// means the caller must undo the RAM apply. Inside a method - /// transaction the record is deferred instead — durability happens - /// once, at [`commit_txn`], as a single atomic frame. - fn wal_log(&mut self, rec: crate::wal::WalRec) -> Result<()> { - if let Some(t) = self.txn.as_mut() { - t.wal.push(rec); - return Ok(()); - } - match self.wal.as_mut() { - Some(WalBackend::PerCommit(w)) => { - w.append(&rec).map_err(|e| anyhow::anyhow!("wal append: {e}"))?; - } - Some(WalBackend::Group(g)) => { - g.stage(&rec).map_err(|e| anyhow::anyhow!("wal stage: {e}"))?; - self.staged = true; - } - None => {} - } - Ok(()) - } - - pub fn catalog(&self) -> &Catalog { &self.catalog } - - /// List every row of `ty` in insertion (id) order. - pub fn list(&self, ty: &str) -> Result> { - self.table(ty).map(|t| t.values().cloned().collect()) - } - - pub fn get(&self, ty: &str, id: i64) -> Result> { - Ok(self.table(ty)?.get(&id).cloned()) - } - - /// Create a row. `body` is the JSON object from the request; missing columns - /// fill in from defaults. Returns the finalized row (including the auto-id). - pub fn create(&mut self, ty: &str, body: Value) -> Result { - let t = self.compiled(ty)?.clone(); - let mut row = self.seed_defaults(&t); - - if let Value::Object(input) = body { - for (k, v) in input { - row.insert(k, v); - } - } - - // Assign id if the type declares one and the caller didn't provide. - if t.has_id && !row.contains_key("id") { - let id = self.mint_id(ty); - row.insert("id".into(), json!(id)); - } - - let id = row.get("id") - .and_then(|v| v.as_i64()) - .unwrap_or_else(|| self.mint_id(ty)); - row.insert("id".into(), json!(id)); - - self.row_insert(ty, id, row.clone()); - if let Some(t) = self.txn.as_mut() { - t.undo.push(Undo::Created { ty: ty.into(), id }); - } - // Dual-write order: RAM applied above, durable now, ack after return. - if let Err(e) = self.wal_log(crate::wal::WalRec::Create { ty: ty.into(), row: row.clone() }) { - self.row_remove(ty, id); // never ack non-durable - return Err(e); - } - if self.mirror.is_some() { - self.mirror_send(crate::mirror::MirrorRec::Upsert { - ty: ty.into(), id, row: row.clone(), - }); - } - Ok(row) - } - - /// Merge-update a row. Remove-then-insert so the secondary indexes see - /// both the old and the new key tuples. - pub fn update(&mut self, ty: &str, id: i64, body: Value) -> Result> { - let Some(prev) = self.table(ty)?.get(&id).cloned() else { return Ok(None); }; - let mut row = prev.clone(); - if let Value::Object(input) = &body { - for (k, v) in input { - if k == "id" { continue; } // don't let the client mutate the primary key - row.insert(k.clone(), v.clone()); - } - } - let updated = row.clone(); - self.row_remove(ty, id); - self.row_insert(ty, id, row); - if let Some(t) = self.txn.as_mut() { - t.undo.push(Undo::Updated { ty: ty.into(), id, prev: prev.clone() }); - } - if let Err(e) = self.wal_log(crate::wal::WalRec::Update { ty: ty.into(), id, body }) { - self.row_remove(ty, id); // undo: never ack non-durable - self.row_insert(ty, id, prev); - return Err(e); - } - if self.mirror.is_some() { - // The mirror needs the FULL post-merge row, not the merge body. - self.mirror_send(crate::mirror::MirrorRec::Upsert { - ty: ty.into(), id, row: updated.clone(), - }); - } - Ok(Some(updated)) - } - - pub fn delete(&mut self, ty: &str, id: i64) -> Result { - self.table(ty)?; // surface unknown-type as an error, not a silent false - let Some(removed) = self.row_remove(ty, id) else { return Ok(false) }; - if let Some(t) = self.txn.as_mut() { - t.undo.push(Undo::Deleted { ty: ty.into(), id, row: removed.clone() }); - } - if let Err(e) = self.wal_log(crate::wal::WalRec::Delete { ty: ty.into(), id }) { - self.row_insert(ty, id, removed); // undo - return Err(e); - } - self.mirror_send(crate::mirror::MirrorRec::Delete { ty: ty.into(), id }); - Ok(true) - } - - /// Equality lookup, index-accelerated. Picks the index whose leading - /// columns form the longest prefix of the queried fields (prefix range - /// scan on its BTreeMap); remaining predicates filter the candidates; - /// no matching index → full scan. Results in id order. `eq` empty = - /// plain `list`. - pub fn find_by(&self, ty: &str, eq: &[(String, Value)]) -> Result> { - let table = self.table(ty)?; - if eq.is_empty() { - return Ok(table.values().cloned().collect()); - } - // Real-equality re-check over ALL queried fields — the index only - // narrows candidates, it never decides membership. - let matches = |row: &Row| eq.iter().all(|(f, v)| { - match row.get(f) { - Some(rv) => rv == v, - None => v.is_null(), - } - }); - - let mut best: Option<(&Index, usize)> = None; - if let Some(idxs) = self.indexes.get(ty) { - for idx in idxs { - let mut k = 0; - for col in &idx.cols { - if eq.iter().any(|(f, _)| f == col) { k += 1; } else { break; } - } - if k > 0 && best.map_or(true, |(_, bk)| k > bk) { - best = Some((idx, k)); - } - } - } - - let Some((idx, k)) = best else { - return Ok(table.values().filter(|r| matches(r)).cloned().collect()); - }; - let prefix: Vec = idx.cols[..k].iter() - .map(|c| IndexKey::from_value(eq.iter().find(|(f, _)| f == c).map(|(_, v)| v))) - .collect(); - let mut ids: Vec = Vec::new(); - // A shorter Vec sorts before any longer Vec sharing its prefix, so - // range(prefix..) starts exactly at the first candidate key. - for (key, set) in idx.map.range(prefix.clone()..) { - if key.len() < k || key[..k] != prefix[..] { break; } - ids.extend(set.iter().copied()); - } - ids.sort_unstable(); - Ok(ids.into_iter() - .filter_map(|id| table.get(&id)) - .filter(|r| matches(r)) - .cloned() - .collect()) - } - - // --- helpers --- - - fn table(&self, ty: &str) -> Result<&BTreeMap> { - self.tables.get(ty).ok_or_else(|| anyhow::anyhow!("no such type: {ty}")) - } - - /// THE two table-mutation primitives — every path that changes a row - /// (CRUD, WAL replay, txn undo) goes through these so the secondary - /// indexes can never drift from the tables. - fn row_insert(&mut self, ty: &str, id: i64, row: Row) { - if let Some(idxs) = self.indexes.get_mut(ty) { - for idx in idxs { - let key = idx.key_for(&row); - idx.map.entry(key).or_default().insert(id); - } - } - if let Some(t) = self.tables.get_mut(ty) { - t.insert(id, row); - } - } - - fn row_remove(&mut self, ty: &str, id: i64) -> Option { - let row = self.tables.get_mut(ty)?.remove(&id)?; - if let Some(idxs) = self.indexes.get_mut(ty) { - for idx in idxs { - let key = idx.key_for(&row); - if let Some(set) = idx.map.get_mut(&key) { - set.remove(&id); - if set.is_empty() { idx.map.remove(&key); } - } - } - } - Some(row) - } - - fn compiled(&self, ty: &str) -> Result<&CompiledType> { - self.catalog.get(ty).ok_or_else(|| anyhow::anyhow!("no such type: {ty}")) - } - - fn mint_id(&mut self, ty: &str) -> i64 { - let counter = self.next_id.entry(ty.to_string()).or_insert(1); - let id = *counter; - *counter += self.id_step; - id - } - - /// Produce the initial row object — default values for every non-relation - /// field that declared one. - fn seed_defaults(&self, t: &CompiledType) -> Row { - let mut row = Map::new(); - for f in &t.fields { - if f.is_relation { continue; } - if let Some(def) = &f.default { - if let Some(v) = Self::eval_default(def, &f.ty) { - row.insert(f.name.clone(), v); - } - } else if matches!(f.ty, FieldTy::Array(_)) && !f.nullable { - row.insert(f.name.clone(), Value::Array(Vec::new())); - } else if let FieldTy::Struct(inner) = &f.ty { - let mut nested = Map::new(); - for g in inner { - if let Some(d) = &g.default { - if let Some(v) = Self::eval_default(d, &g.ty) { - nested.insert(g.name.clone(), v); - } - } - } - if !nested.is_empty() { - row.insert(f.name.clone(), Value::Object(nested)); - } - } - } - row - } - - fn eval_default(def: &DefaultExpr, _ty: &FieldTy) -> Option { - Some(match def { - DefaultExpr::Str(s) => Value::String(s.clone()), - DefaultExpr::Int(n) => json!(n), - DefaultExpr::Bool(b) => json!(b), - DefaultExpr::Null => Value::Null, - DefaultExpr::Now => Value::String(now_iso8601()), - DefaultExpr::Enum(s) => Value::String(s.clone()), - // Opaque expressions (computed fields like `total = sum(...)`) are - // Stage 2-unevaluated. Omit rather than echo parser-debug tokens. - DefaultExpr::Opaque(_) => return None, - }) - } -} - -pub(crate) fn now_iso8601() -> String { - let d = SystemTime::now() - .duration_since(UNIX_EPOCH) - .unwrap_or_default(); - let secs = d.as_secs() as i64; - let nanos = d.subsec_nanos(); - // Minimal ISO-8601 UTC formatter — good enough for Stage 2. - // (`chrono` would be nicer but we're keeping deps small.) - let (yr, mo, da, hr, mi, se) = ymdhms(secs); - format!("{yr:04}-{mo:02}-{da:02}T{hr:02}:{mi:02}:{se:02}.{:03}Z", nanos / 1_000_000) -} - -/// Pure-Rust epoch → (year, month, day, hour, minute, second) for UTC. -fn ymdhms(mut secs: i64) -> (i32, u32, u32, u32, u32, u32) { - const SECS_PER_DAY: i64 = 86_400; - let se = (secs.rem_euclid(60)) as u32; - secs = secs.div_euclid(60); - let mi = (secs.rem_euclid(60)) as u32; - secs = secs.div_euclid(60); - let hr = (secs.rem_euclid(24)) as u32; - let mut days = secs.div_euclid(24); - - // Days → calendar. Epoch 1970-01-01 is a Thursday but we don't need weekday. - let mut year: i32 = 1970; - loop { - let dy = if is_leap(year) { 366 } else { 365 }; - if days >= dy { - days -= dy; - year += 1; - } else { - break; - } - } - - let months = if is_leap(year) { - [31, 29, 31, 30, 31, 30, 31, 31, 30, 31, 30, 31] - } else { - [31, 28, 31, 30, 31, 30, 31, 31, 30, 31, 30, 31] - }; - let mut mo = 1u32; - for (i, d) in months.iter().enumerate() { - if days < *d as i64 { - mo = i as u32 + 1; - break; - } - days -= *d as i64; - } - let _ = SECS_PER_DAY; - let da = days as u32 + 1; - (year, mo, da, hr, mi, se) -} - -fn is_leap(y: i32) -> bool { - (y % 4 == 0 && y % 100 != 0) || (y % 400 == 0) -} - -#[cfg(test)] -mod tests { - use super::*; - use crate::parser::parse; - - fn engine_from(src: &str) -> Engine { - let cat = Catalog::from_schemas(vec![parse(src).unwrap()]).unwrap(); - Engine::new(cat) - } - - const INDEXED: &str = r#" -@table(index: [owner, at]) -type Item { id: Id - owner: Int - at: Text - service rest "/api/items" expose list } -"#; - - #[test] - fn secondary_index_tracks_create_update_delete() { - let mut eng = engine_from(INDEXED); - for i in 0..3 { - eng.create("Item", json!({"owner": 1, "at": format!("t{i}")})).unwrap(); - } - eng.create("Item", json!({"owner": 2, "at": "t9"})).unwrap(); - - // prefix match (owner) and full composite (owner, at) - let one = eng.find_by("Item", &[("owner".into(), json!(1))]).unwrap(); - assert_eq!(one.len(), 3); - let exact = eng.find_by("Item", - &[("owner".into(), json!(1)), ("at".into(), json!("t1"))]).unwrap(); - assert_eq!(exact.len(), 1); - - // update moves the row between index keys - let id = exact[0]["id"].as_i64().unwrap(); - eng.update("Item", id, json!({"owner": 2})).unwrap(); - assert_eq!(eng.find_by("Item", &[("owner".into(), json!(1))]).unwrap().len(), 2); - assert_eq!(eng.find_by("Item", &[("owner".into(), json!(2))]).unwrap().len(), 2); - - // delete clears its entries - eng.delete("Item", id).unwrap(); - assert_eq!(eng.find_by("Item", &[("owner".into(), json!(2))]).unwrap().len(), 1); - - // non-indexed field → scan fallback, same semantics - assert_eq!(eng.find_by("Item", &[("at".into(), json!("t0"))]).unwrap().len(), 1); - - // index answers equal scan answers (ground truth) - let scan: Vec<_> = eng.list("Item").unwrap().into_iter() - .filter(|r| r["owner"] == json!(1)).collect(); - assert_eq!(eng.find_by("Item", &[("owner".into(), json!(1))]).unwrap(), scan); - } - - #[test] - fn txn_abort_restores_index_state() { - let mut eng = engine_from(INDEXED); - eng.create("Item", json!({"owner": 1, "at": "a"})).unwrap(); // id 1 - - eng.begin_txn().unwrap(); - eng.create("Item", json!({"owner": 1, "at": "b"})).unwrap(); - eng.update("Item", 1, json!({"owner": 5})).unwrap(); - eng.abort_txn(); - - assert_eq!(eng.find_by("Item", &[("owner".into(), json!(1))]).unwrap().len(), 1); - assert!(eng.find_by("Item", &[("owner".into(), json!(5))]).unwrap().is_empty()); - assert_eq!(eng.list("Item").unwrap().len(), 1); - } - - #[test] - fn crud_roundtrip_auto_id() { - let mut eng = engine_from(r#" -type Article { id: Id - title: Text - published: Bool = false - service rest "/api/articles" expose list, get, create, update, delete } -"#); - let a = eng.create("Article", json!({"title": "Hello"})).unwrap(); - assert_eq!(a.get("id").unwrap().as_i64().unwrap(), 1); - assert_eq!(a.get("title").unwrap().as_str().unwrap(), "Hello"); - assert_eq!(a.get("published").unwrap(), &json!(false)); - - let b = eng.create("Article", json!({"title": "World", "published": true})).unwrap(); - assert_eq!(b.get("id").unwrap().as_i64().unwrap(), 2); - - let list = eng.list("Article").unwrap(); - assert_eq!(list.len(), 2); - - let got = eng.get("Article", 1).unwrap().unwrap(); - assert_eq!(got.get("title").unwrap().as_str().unwrap(), "Hello"); - - let upd = eng.update("Article", 1, json!({"title": "Hi"})).unwrap().unwrap(); - assert_eq!(upd.get("title").unwrap().as_str().unwrap(), "Hi"); - - assert!(eng.delete("Article", 1).unwrap()); - assert!(!eng.delete("Article", 1).unwrap()); - assert_eq!(eng.list("Article").unwrap().len(), 1); - } -} diff --git a/crates/rt/src/http/connection.rs b/crates/rt/src/http/connection.rs deleted file mode 100644 index 246245b..0000000 --- a/crates/rt/src/http/connection.rs +++ /dev/null @@ -1,329 +0,0 @@ -//! Per-connection state machine driven by the phase-02 [`EventLoop`]. -//! -//! Lifecycle (HTTP/1.1 keep-alive — the C prototype's phase-C sequence): -//! Reading → drain `read(2)` to `EAGAIN`, parse one request, dispatch -//! through the `Router`, queue the response. Consumed bytes are trimmed -//! so a pipelined follow-up request carries over. -//! Writing → drain `write(2)` to `EAGAIN`. If a write was partial, the -//! loop re-arms the fd as `WRITABLE` and we continue on the next event. -//! Once flushed: keep-alive resets to Reading (and immediately serves -//! any buffered pipelined request); `Connection: close` goes to Done. -//! Done → loop closes the fd. -//! -//! Adapted from `reference/crates/wo-http/src/connection.rs`. The owning -//! [`EventLoop`] supplies `read`/`write` readiness via edge-triggered -//! `epoll`; this struct is the per-fd part of the state. -//! -//! [`EventLoop`]: crate::runtime::EventLoop - -use std::io; -use std::os::unix::io::{AsRawFd, RawFd}; - -use super::request::{self, ParseResult}; -use super::response::Response; -use super::route::Router; - -#[derive(Debug, Clone, Copy, PartialEq, Eq)] -pub enum ConnState { - Reading, - Writing, - /// Response built but gated on the WAL batch's fsync (group commit). - Parked, - Done, -} - -pub struct Connection { - fd: RawFd, - state: ConnState, - read_buf: Vec, - write_buf: Vec, - write_offset: usize, - keep_alive: bool, - /// Incarnation stamp — parked acks are released only when the stamp - /// matches, so a reused fd can never receive another commit's ack. - gen: u64, -} - -impl Connection { - pub fn new(fd: RawFd) -> Self { - Self { - fd, - state: ConnState::Reading, - read_buf: Vec::with_capacity(4096), - write_buf: Vec::new(), - write_offset: 0, - keep_alive: true, - gen: 0, - } - } - - pub fn with_gen(fd: RawFd, gen: u64) -> Self { - let mut c = Self::new(fd); - c.gen = gen; - c - } - - pub fn gen(&self) -> u64 { self.gen } - pub fn is_parked(&self) -> bool { self.state == ConnState::Parked } - - /// The batch fsync landed — the gated response may leave now. - pub fn unpark(&mut self) { - if self.state == ConnState::Parked { - self.state = ConnState::Writing; - } - } - - pub fn state(&self) -> ConnState { self.state } - pub fn is_done(&self) -> bool { self.state == ConnState::Done } - - /// Drain the socket into `read_buf` until `EAGAIN` or EOF. - /// Returns `false` when the peer closed (connection should be torn down). - fn drain_read(&mut self) -> io::Result { - let mut tmp = [0u8; 4096]; - loop { - let n = unsafe { - libc::read(self.fd, tmp.as_mut_ptr() as *mut libc::c_void, tmp.len()) - }; - if n < 0 { - let err = io::Error::last_os_error(); - if err.raw_os_error() == Some(libc::EAGAIN) { - return Ok(true); - } - return Err(err); - } - if n == 0 { - return Ok(false); - } - self.read_buf.extend_from_slice(&tmp[..n as usize]); - // Cap at MAX_BODY_BYTES + headers — refuse pathological requests. - if self.read_buf.len() > 32 * 1024 * 1024 { - return Err(io::Error::new(io::ErrorKind::InvalidData, "request too large")); - } - } - } - - /// Drain `write_buf[write_offset..]` to the socket until `EAGAIN`. - /// Returns `true` once everything has been flushed. - fn drain_write(&mut self) -> io::Result { - loop { - let remaining = &self.write_buf[self.write_offset..]; - if remaining.is_empty() { return Ok(true); } - let n = unsafe { - libc::write( - self.fd, - remaining.as_ptr() as *const libc::c_void, - remaining.len(), - ) - }; - if n < 0 { - let err = io::Error::last_os_error(); - if err.raw_os_error() == Some(libc::EAGAIN) { - return Ok(false); - } - return Err(err); - } - if n == 0 { return Ok(false); } - self.write_offset += n as usize; - } - } - - fn try_parse(&self) -> ParseResult { - request::parse(&self.read_buf) - } - - fn queue_response(&mut self, response: &Response) { - self.write_buf = response.to_bytes(self.keep_alive); - self.write_offset = 0; - self.state = if response.gate { ConnState::Parked } else { ConnState::Writing }; - } - - /// One step of the state machine, given a readiness event from the - /// loop. Serves as many buffered requests as it can (keep-alive + - /// pipelining). Returns `true` if the connection now wants `WRITABLE` - /// (the caller should switch interest from `READABLE`). - pub fn drive( - &mut self, - readable: bool, - _writable: bool, - hangup: bool, - error: bool, - router: &Router, - ) -> io::Result { - if error { - self.state = ConnState::Done; - return Ok(false); - } - - let mut peer_open = true; - if readable && self.state == ConnState::Reading { - peer_open = self.drain_read()?; - } - - loop { - if self.state == ConnState::Reading { - match self.try_parse() { - ParseResult::Complete(req, consumed) => { - // The response's Connection header — and what we do - // after flushing it — follow the request's wish. - self.keep_alive = req.keep_alive; - self.read_buf.drain(..consumed); - let resp = router.dispatch(&req); - self.queue_response(&resp); - } - ParseResult::Incomplete => { - if !peer_open || hangup { - self.state = ConnState::Done; // peer gone mid-request / idle EOF - } - return Ok(false); - } - ParseResult::Error(msg) => { - self.keep_alive = false; // protocol state is suspect - let resp = Response::status(super::Status::BAD_REQUEST).text(msg); - self.queue_response(&resp); - } - } - } - - if self.state == ConnState::Writing { - let flushed = self.drain_write()?; - if !flushed { - return Ok(true); // wait for WRITABLE - } - if self.keep_alive { - self.write_buf.clear(); - self.write_offset = 0; - self.state = ConnState::Reading; - continue; // pipelined request may be buffered - } - self.state = ConnState::Done; - } - - return Ok(false); - } - } -} - -impl AsRawFd for Connection { - fn as_raw_fd(&self) -> RawFd { self.fd } -} - -impl Drop for Connection { - fn drop(&mut self) { - unsafe { libc::close(self.fd); } - } -} - -#[cfg(test)] -mod tests { - use super::*; - use crate::http::{Method, Response, Router}; - - fn socketpair_nonblock() -> (RawFd, RawFd) { - let mut fds = [0i32; 2]; - let r = unsafe { - libc::socketpair( - libc::AF_UNIX, - libc::SOCK_STREAM | libc::SOCK_NONBLOCK | libc::SOCK_CLOEXEC, - 0, - fds.as_mut_ptr(), - ) - }; - assert_eq!(r, 0); - (fds[0], fds[1]) - } - - #[test] - fn keep_alive_serves_many_requests_on_one_connection() { - let (server_fd, client_fd) = socketpair_nonblock(); - - let router = Router::new() - .route(Method::Get, "/healthz", |_, _| Response::ok().text("ok")); - let mut conn = Connection::new(server_fd); - - for i in 0..3 { - let req = b"GET /healthz HTTP/1.1\r\nHost: localhost\r\n\r\n"; - let n = unsafe { libc::write(client_fd, req.as_ptr() as *const _, req.len()) }; - assert_eq!(n, req.len() as isize); - - let want_writable = conn.drive(true, false, false, false, &router).unwrap(); - assert!(!want_writable, "small response fits in one write"); - assert!(!conn.is_done(), "keep-alive must survive request {i}"); - - let mut buf = [0u8; 4096]; - let n = unsafe { libc::read(client_fd, buf.as_mut_ptr() as *mut _, buf.len()) }; - assert!(n > 0); - let s = std::str::from_utf8(&buf[..n as usize]).unwrap(); - assert!(s.starts_with("HTTP/1.1 200 OK\r\n"), "got: {s}"); - assert!(s.contains("Connection: keep-alive\r\n"), "got: {s}"); - assert!(s.ends_with("\r\n\r\nok")); - } - - unsafe { libc::close(client_fd); } - } - - #[test] - fn pipelined_requests_are_served_in_order() { - let (server_fd, client_fd) = socketpair_nonblock(); - - let router = Router::new() - .route(Method::Get, "/healthz", |_, _| Response::ok().text("ok")); - let mut conn = Connection::new(server_fd); - - // Two requests in ONE write — the second must be served from the - // carried-over buffer without another readable event. - let req = b"GET /healthz HTTP/1.1\r\n\r\nGET /healthz HTTP/1.1\r\n\r\n"; - unsafe { libc::write(client_fd, req.as_ptr() as *const _, req.len()) }; - - conn.drive(true, false, false, false, &router).unwrap(); - assert!(!conn.is_done()); - - let mut buf = [0u8; 4096]; - let n = unsafe { libc::read(client_fd, buf.as_mut_ptr() as *mut _, buf.len()) }; - let s = std::str::from_utf8(&buf[..n as usize]).unwrap(); - assert_eq!(s.matches("HTTP/1.1 200 OK").count(), 2, "got: {s}"); - - unsafe { libc::close(client_fd); } - } - - #[test] - fn connection_close_header_is_honored() { - let (server_fd, client_fd) = socketpair_nonblock(); - - let router = Router::new() - .route(Method::Get, "/healthz", |_, _| Response::ok().text("ok")); - let mut conn = Connection::new(server_fd); - - let req = b"GET /healthz HTTP/1.1\r\nConnection: close\r\n\r\n"; - unsafe { libc::write(client_fd, req.as_ptr() as *const _, req.len()) }; - - conn.drive(true, false, false, false, &router).unwrap(); - assert!(conn.is_done(), "Connection: close must end the connection"); - - let mut buf = [0u8; 4096]; - let n = unsafe { libc::read(client_fd, buf.as_mut_ptr() as *mut _, buf.len()) }; - let s = std::str::from_utf8(&buf[..n as usize]).unwrap(); - assert!(s.contains("Connection: close\r\n"), "got: {s}"); - - unsafe { libc::close(client_fd); } - } - - #[test] - fn returns_404_for_unknown_path() { - let (server_fd, client_fd) = socketpair_nonblock(); - - let req = b"GET /missing HTTP/1.1\r\n\r\n"; - unsafe { libc::write(client_fd, req.as_ptr() as *const _, req.len()); } - - let router = Router::new() - .route(Method::Get, "/healthz", |_, _| Response::ok().text("ok")); - let mut conn = Connection::new(server_fd); - conn.drive(true, false, false, false, &router).unwrap(); - - let mut buf = [0u8; 4096]; - let n = unsafe { libc::read(client_fd, buf.as_mut_ptr() as *mut _, buf.len()) }; - let s = std::str::from_utf8(&buf[..n as usize]).unwrap(); - assert!(s.starts_with("HTTP/1.1 404 Not Found\r\n"), "got: {s}"); - - unsafe { libc::close(client_fd); } - } -} diff --git a/crates/rt/src/http/listener.rs b/crates/rt/src/http/listener.rs deleted file mode 100644 index 1ce6359..0000000 --- a/crates/rt/src/http/listener.rs +++ /dev/null @@ -1,187 +0,0 @@ -//! Non-blocking TCP listener — `socket(2)` + `bind(2)` + `listen(2)` + `accept4(2)`. -//! -//! Adapted from `reference/crates/wo-http/src/listener.rs`. The v1 hand-rolled -//! IPv4 parser had a byte-order bug for non-localhost addresses; here we -//! defer to `std::net::SocketAddr` (stdlib, no extra crate) and convert the -//! resulting octets to a `sockaddr_in` correctly. - -use std::io; -use std::net::SocketAddr; -use std::os::unix::io::{AsRawFd, RawFd}; - -pub struct Listener { - fd: RawFd, - addr: SocketAddr, -} - -impl Listener { - /// Bind to `addr` (IPv4 only for now) and start listening with backlog 128. - /// Socket is created `SOCK_NONBLOCK | SOCK_CLOEXEC`. - pub fn bind(addr: &str) -> io::Result { - Self::bind_inner(addr, false) - } - - /// Like [`bind`](Self::bind), but with `SO_REUSEPORT`: every shard thread - /// binds its own listener to the same port and the kernel load-balances - /// incoming connections across them by 4-tuple hash (plan 09 decision 3). - pub fn bind_reuseport(addr: &str) -> io::Result { - Self::bind_inner(addr, true) - } - - fn bind_inner(addr: &str, reuseport: bool) -> io::Result { - let parsed: SocketAddr = addr.parse().map_err(|e| { - io::Error::new(io::ErrorKind::InvalidInput, format!("bad addr {addr:?}: {e}")) - })?; - let SocketAddr::V4(v4) = parsed else { - return Err(io::Error::new(io::ErrorKind::InvalidInput, "IPv4 only for now")); - }; - - let fd = unsafe { - libc::socket( - libc::AF_INET, - libc::SOCK_STREAM | libc::SOCK_NONBLOCK | libc::SOCK_CLOEXEC, - 0, - ) - }; - if fd < 0 { return Err(io::Error::last_os_error()); } - - // SO_REUSEADDR so the same port restarts cleanly between `wo run`s. - let one: libc::c_int = 1; - let ret = unsafe { - libc::setsockopt( - fd, libc::SOL_SOCKET, libc::SO_REUSEADDR, - &one as *const _ as *const libc::c_void, - std::mem::size_of::() as libc::socklen_t, - ) - }; - if ret < 0 { - let err = io::Error::last_os_error(); - unsafe { libc::close(fd); } - return Err(err); - } - - if reuseport { - let ret = unsafe { - libc::setsockopt( - fd, libc::SOL_SOCKET, libc::SO_REUSEPORT, - &one as *const _ as *const libc::c_void, - std::mem::size_of::() as libc::socklen_t, - ) - }; - if ret < 0 { - let err = io::Error::last_os_error(); - unsafe { libc::close(fd); } - return Err(err); - } - } - - let s_addr = u32::from_be_bytes(v4.ip().octets()).to_be(); - let sock = libc::sockaddr_in { - sin_family: libc::AF_INET as libc::sa_family_t, - sin_port: v4.port().to_be(), - sin_addr: libc::in_addr { s_addr }, - sin_zero: [0; 8], - }; - let ret = unsafe { - libc::bind( - fd, - &sock as *const _ as *const libc::sockaddr, - std::mem::size_of::() as libc::socklen_t, - ) - }; - if ret < 0 { - let err = io::Error::last_os_error(); - unsafe { libc::close(fd); } - return Err(err); - } - - if unsafe { libc::listen(fd, 128) } < 0 { - let err = io::Error::last_os_error(); - unsafe { libc::close(fd); } - return Err(err); - } - - // Resolve the actual bound address — caller may have asked for port 0. - let local = read_local_addr(fd)?; - Ok(Self { fd, addr: local }) - } - - /// Accept the next pending connection. Returns `None` on `EAGAIN`. - pub fn accept(&self) -> io::Result> { - let cfd = unsafe { - libc::accept4( - self.fd, - std::ptr::null_mut(), std::ptr::null_mut(), - libc::SOCK_NONBLOCK | libc::SOCK_CLOEXEC, - ) - }; - if cfd < 0 { - let err = io::Error::last_os_error(); - // On Linux EAGAIN == EWOULDBLOCK; one branch is enough. - return match err.raw_os_error() { - Some(libc::EAGAIN) => Ok(None), - _ => Err(err), - }; - } - Ok(Some(cfd)) - } - - pub fn local_addr(&self) -> SocketAddr { self.addr } -} - -impl AsRawFd for Listener { - fn as_raw_fd(&self) -> RawFd { self.fd } -} - -impl Drop for Listener { - fn drop(&mut self) { - unsafe { libc::close(self.fd); } - } -} - -fn read_local_addr(fd: RawFd) -> io::Result { - let mut sock: libc::sockaddr_in = unsafe { std::mem::zeroed() }; - let mut len = std::mem::size_of::() as libc::socklen_t; - let ret = unsafe { - libc::getsockname(fd, &mut sock as *mut _ as *mut libc::sockaddr, &mut len) - }; - if ret < 0 { return Err(io::Error::last_os_error()); } - let ip = u32::from_be(sock.sin_addr.s_addr).to_be_bytes(); - let port = u16::from_be(sock.sin_port); - Ok(SocketAddr::from(([ip[0], ip[1], ip[2], ip[3]], port))) -} - -#[cfg(test)] -mod tests { - use super::*; - - #[test] - fn binds_and_accepts_a_client() { - let listener = Listener::bind("127.0.0.1:0").unwrap(); - let port = listener.local_addr().port(); - assert!(port > 0); - assert!(listener.accept().unwrap().is_none(), "no clients pending yet"); - - let stream = std::net::TcpStream::connect(("127.0.0.1", port)).unwrap(); - // Block briefly until the server-side accept sees it. - let mut accepted = None; - for _ in 0..50 { - if let Some(fd) = listener.accept().unwrap() { accepted = Some(fd); break; } - std::thread::sleep(std::time::Duration::from_millis(10)); - } - let cfd = accepted.expect("accept produced a client fd"); - unsafe { libc::close(cfd); } - drop(stream); - } - - #[test] - fn reuseport_allows_two_listeners_on_one_port() { - let a = Listener::bind_reuseport("127.0.0.1:0").unwrap(); - let port = a.local_addr().port(); - let b = Listener::bind_reuseport(&format!("127.0.0.1:{port}")) - .expect("second SO_REUSEPORT bind on the same port must succeed"); - assert_eq!(b.local_addr().port(), port); - // Plain bind on the same port must still fail (no REUSEPORT). - assert!(Listener::bind(&format!("127.0.0.1:{port}")).is_err()); - } -} diff --git a/crates/rt/src/http/mod.rs b/crates/rt/src/http/mod.rs deleted file mode 100644 index 5a412c3..0000000 --- a/crates/rt/src/http/mod.rs +++ /dev/null @@ -1,21 +0,0 @@ -//! Hand-rolled HTTP/1.1 — phase 03 of the runtime plan. -//! -//! The transport layer for the `wo` binary after the phase-04 cutover. -//! Drives non-blocking accept + per-connection state machines off the -//! phase-02 [`EventLoop`](super::runtime::EventLoop). No `tokio`, no -//! `axum`, no `hyper`. Synchronous handlers; close-after-response (the -//! v1 model — keep-alive lands when a sample needs it). -//! -//! See `docs/plan/03-hand-rolled-http.md` and `docs/plan/04-cutover-remove-tokio-axum.md`. - -mod connection; -mod listener; -mod request; -mod response; -mod route; - -pub use connection::{Connection, ConnState}; -pub use listener::Listener; -pub use request::{Method, Request}; -pub use response::{Response, Status}; -pub use route::{RouteParams, Router}; diff --git a/crates/rt/src/http/request.rs b/crates/rt/src/http/request.rs deleted file mode 100644 index 1e595f7..0000000 --- a/crates/rt/src/http/request.rs +++ /dev/null @@ -1,192 +0,0 @@ -//! Incremental HTTP/1.1 request parser. -//! -//! Adapted from `reference/crates/wo-http/src/request.rs`. v1 only parsed -//! request headers (the v1 blog is read-only HTML). The phase-04 cutover -//! needs JSON request bodies, so this parser also drains a -//! `Content-Length`-delimited body. Chunked transfer encoding is not -//! supported (no sample sends one — see `docs/plan/03-hand-rolled-http.md`). - -use std::collections::HashMap; - -#[derive(Debug, Clone, Copy, PartialEq, Eq, Hash)] -pub enum Method { - Get, - Post, - Patch, - Put, - Delete, - Head, - Options, - Other, -} - -impl Method { - fn parse(token: &str) -> Method { - match token { - "GET" => Method::Get, - "POST" => Method::Post, - "PATCH" => Method::Patch, - "PUT" => Method::Put, - "DELETE" => Method::Delete, - "HEAD" => Method::Head, - "OPTIONS" => Method::Options, - _ => Method::Other, - } - } -} - -#[derive(Debug, Clone)] -pub struct Request { - pub method: Method, - pub path: String, - pub query: Option, - pub headers: HashMap, - pub body: Vec, - /// HTTP/1.1 defaults to keep-alive unless `Connection: close`; - /// HTTP/1.0 defaults to close unless `Connection: keep-alive`. - pub keep_alive: bool, -} - -pub enum ParseResult { - /// A full request plus the number of bytes it consumed — the caller - /// trims its buffer so a pipelined follow-up request survives. - Complete(Request, usize), - Incomplete, - Error(String), -} - -const MAX_HEADER_BYTES: usize = 8 * 1024; -const MAX_BODY_BYTES: usize = 16 * 1024 * 1024; - -/// Try to parse a complete HTTP/1.1 request out of `buf`. -/// Returns `Complete(req, bytes_consumed)` once headers + body are present. -pub fn parse(buf: &[u8]) -> ParseResult { - let header_end = match find_header_end(buf) { - Some(p) => p, - None => { - if buf.len() > MAX_HEADER_BYTES { - return ParseResult::Error("request headers too large".into()); - } - return ParseResult::Incomplete; - } - }; - - let header_str = match std::str::from_utf8(&buf[..header_end]) { - Ok(s) => s, - Err(_) => return ParseResult::Error("non-UTF-8 in headers".into()), - }; - - let mut lines = header_str.lines(); - - let request_line = match lines.next() { - Some(l) => l, - None => return ParseResult::Error("empty request".into()), - }; - let mut parts = request_line.split_whitespace(); - let method = match parts.next() { - Some(t) => Method::parse(t), - None => return ParseResult::Error("missing method".into()), - }; - let raw_path = match parts.next() { - Some(p) => p, - None => return ParseResult::Error("missing path".into()), - }; - let http10 = parts.next() == Some("HTTP/1.0"); - let (path, query) = match raw_path.split_once('?') { - Some((p, q)) => (p.to_string(), Some(q.to_string())), - None => (raw_path.to_string(), None), - }; - - let mut headers = HashMap::new(); - for line in lines { - if line.is_empty() { break; } - if let Some((k, v)) = line.split_once(':') { - headers.insert(k.trim().to_ascii_lowercase(), v.trim().to_string()); - } - } - - let body_len: usize = headers - .get("content-length") - .and_then(|v| v.parse().ok()) - .unwrap_or(0); - if body_len > MAX_BODY_BYTES { - return ParseResult::Error("Content-Length exceeds limit".into()); - } - - let header_bytes = header_end + 4; // include the trailing \r\n\r\n - let total = header_bytes + body_len; - if buf.len() < total { - return ParseResult::Incomplete; - } - - let body = buf[header_bytes..total].to_vec(); - let conn_hdr = headers.get("connection").map(|v| v.to_ascii_lowercase()); - let keep_alive = if http10 { - conn_hdr.as_deref() == Some("keep-alive") - } else { - conn_hdr.as_deref() != Some("close") - }; - ParseResult::Complete(Request { method, path, query, headers, body, keep_alive }, total) -} - -fn find_header_end(buf: &[u8]) -> Option { - buf.windows(4).position(|w| w == b"\r\n\r\n") -} - -#[cfg(test)] -mod tests { - use super::*; - - #[test] - fn parses_simple_get() { - let raw = b"GET /api/articles HTTP/1.1\r\nHost: localhost\r\n\r\n"; - let ParseResult::Complete(req, _) = parse(raw) else { panic!("expected Complete") }; - assert_eq!(req.method, Method::Get); - assert_eq!(req.path, "/api/articles"); - assert!(req.query.is_none()); - assert!(req.body.is_empty()); - } - - #[test] - fn parses_query_string() { - let raw = b"GET /tag/rust?page=2 HTTP/1.1\r\n\r\n"; - let ParseResult::Complete(req, _) = parse(raw) else { panic!() }; - assert_eq!(req.path, "/tag/rust"); - assert_eq!(req.query.as_deref(), Some("page=2")); - } - - #[test] - fn parses_post_with_body() { - let body = b"{\"title\":\"hi\"}"; - let mut raw = Vec::new(); - raw.extend_from_slice(b"POST /api/articles HTTP/1.1\r\n"); - raw.extend_from_slice(b"Host: localhost\r\n"); - raw.extend_from_slice(b"Content-Type: application/json\r\n"); - raw.extend_from_slice(format!("Content-Length: {}\r\n", body.len()).as_bytes()); - raw.extend_from_slice(b"\r\n"); - raw.extend_from_slice(body); - let ParseResult::Complete(req, consumed) = parse(&raw) else { panic!() }; - assert_eq!(consumed, raw.len()); - assert_eq!(req.method, Method::Post); - assert_eq!(req.body, body); - } - - #[test] - fn incomplete_when_body_truncated() { - let raw = b"POST / HTTP/1.1\r\nContent-Length: 10\r\n\r\nshort"; - assert!(matches!(parse(raw), ParseResult::Incomplete)); - } - - #[test] - fn incomplete_when_headers_truncated() { - let raw = b"GET / HTTP/1.1\r\nHost: local"; - assert!(matches!(parse(raw), ParseResult::Incomplete)); - } - - #[test] - fn parses_patch_method() { - let raw = b"PATCH /api/articles/1 HTTP/1.1\r\nContent-Length: 0\r\n\r\n"; - let ParseResult::Complete(req, _) = parse(raw) else { panic!() }; - assert_eq!(req.method, Method::Patch); - } -} diff --git a/crates/rt/src/http/response.rs b/crates/rt/src/http/response.rs deleted file mode 100644 index 21eed05..0000000 --- a/crates/rt/src/http/response.rs +++ /dev/null @@ -1,118 +0,0 @@ -//! HTTP/1.1 response builder + serializer. -//! -//! Adapted from `reference/crates/wo-http/src/response.rs`. Adds: -//! * `Status` constants for the codes the REST samples assert on -//! (200/201/204/400/404/405/500/501). -//! * `Response::json(&serde_json::Value)` matching the cutover-handler -//! shape in `docs/plan/04-cutover-remove-tokio-axum.md`. - -use serde_json::Value; - -#[derive(Debug, Clone, Copy)] -pub struct Status(pub u16, pub &'static str); - -impl Status { - pub const OK: Status = Status(200, "OK"); - pub const CREATED: Status = Status(201, "Created"); - pub const NO_CONTENT: Status = Status(204, "No Content"); - pub const BAD_REQUEST: Status = Status(400, "Bad Request"); - pub const NOT_FOUND: Status = Status(404, "Not Found"); - pub const METHOD_NOT_ALLOWED: Status = Status(405, "Method Not Allowed"); - pub const CONFLICT: Status = Status(409, "Conflict"); - pub const INTERNAL_SERVER_ERROR: Status = Status(500, "Internal Server Error"); - pub const NOT_IMPLEMENTED: Status = Status(501, "Not Implemented"); -} - -#[derive(Debug, Clone)] -pub struct Response { - pub status: Status, - pub headers: Vec<(String, String)>, - pub body: Vec, - /// Group-commit gate: this response acknowledges a staged WAL frame and - /// must not leave until the batch's fsync CQE (the connection parks). - pub gate: bool, -} - -impl Response { - pub fn status(s: Status) -> Self { - Self { status: s, headers: Vec::new(), body: Vec::new(), gate: false } - } - - pub fn ok() -> Self { Self::status(Status::OK) } - pub fn created() -> Self { Self::status(Status::CREATED) } - pub fn no_content() -> Self { Self::status(Status::NO_CONTENT) } - - pub fn header(mut self, k: &str, v: &str) -> Self { - self.headers.push((k.to_string(), v.to_string())); - self - } - - pub fn body(mut self, body: impl Into>) -> Self { - self.body = body.into(); - self - } - - /// Plain-text body with `Content-Type: text/plain; charset=utf-8`. - pub fn text(self, body: impl Into) -> Self { - let body: String = body.into(); - self.header("Content-Type", "text/plain; charset=utf-8") - .body(body.into_bytes()) - } - - /// JSON body with `Content-Type: application/json`. - pub fn json(self, value: &Value) -> Self { - let buf = serde_json::to_vec(value).unwrap_or_else(|_| b"null".to_vec()); - self.header("Content-Type", "application/json").body(buf) - } - - /// Serialize to the wire format. Auto-injects `Content-Length` and - /// Serialize with the connection disposition the state machine decided. - pub fn to_bytes(&self, keep_alive: bool) -> Vec { - let mut buf = Vec::with_capacity(256 + self.body.len()); - buf.extend_from_slice( - format!("HTTP/1.1 {} {}\r\n", self.status.0, self.status.1).as_bytes(), - ); - for (k, v) in &self.headers { - buf.extend_from_slice(format!("{k}: {v}\r\n").as_bytes()); - } - buf.extend_from_slice(format!("Content-Length: {}\r\n", self.body.len()).as_bytes()); - buf.extend_from_slice(if keep_alive { b"Connection: keep-alive\r\n".as_slice() } - else { b"Connection: close\r\n".as_slice() }); - buf.extend_from_slice(b"\r\n"); - buf.extend_from_slice(&self.body); - buf - } -} - -#[cfg(test)] -mod tests { - use super::*; - use serde_json::json; - - #[test] - fn ok_text_body() { - let r = Response::ok().text("ok"); - let s = String::from_utf8(r.to_bytes(false)).unwrap(); - assert!(s.starts_with("HTTP/1.1 200 OK\r\n")); - assert!(s.contains("Content-Type: text/plain")); - assert!(s.contains("Content-Length: 2\r\n")); - assert!(s.ends_with("\r\n\r\nok")); - } - - #[test] - fn json_body() { - let r = Response::created().json(&json!({"id": 1, "title": "Hi"})); - let s = String::from_utf8(r.to_bytes(false)).unwrap(); - assert!(s.starts_with("HTTP/1.1 201 Created\r\n")); - assert!(s.contains("Content-Type: application/json")); - assert!(s.contains(r#"{"id":1,"title":"Hi"}"#)); - } - - #[test] - fn no_content_status() { - let r = Response::no_content(); - let s = String::from_utf8(r.to_bytes(false)).unwrap(); - assert!(s.starts_with("HTTP/1.1 204 No Content\r\n")); - assert!(s.ends_with("\r\n\r\n")); - } -} diff --git a/crates/rt/src/http/route.rs b/crates/rt/src/http/route.rs deleted file mode 100644 index 0799630..0000000 --- a/crates/rt/src/http/route.rs +++ /dev/null @@ -1,185 +0,0 @@ -//! Method + URL pattern → handler dispatch. -//! -//! Combined adaptation of `reference/crates/wo-route/src/{router,pattern}.rs`. -//! Handler shape is `Fn(&Request, &RouteParams) -> Response`, captured as a -//! boxed closure so each route closes over its own state (typically an -//! `Arc>` — see `crates/rt/src/server.rs`). -//! -//! Dispatch distinguishes 404 (no path matches any registered route) from -//! 405 (path matches at least one route but not for the request's method), -//! which the REST sample asserts (`expose list, get, ...` → POST → 405). - -use std::collections::HashMap; - -use super::{Method, Request, Response, Status}; - -// NOT `Send`/`Sync`: since plan 09b each worker thread builds and owns its -// own Router over its own shard engine (`Rc` captures) — routers -// never cross threads. -pub type HandlerFn = dyn Fn(&Request, &RouteParams) -> Response + 'static; - -#[derive(Debug, Clone, PartialEq)] -enum Segment { - Literal(String), - Param(String), - Wildcard(String), -} - -#[derive(Debug, Clone)] -struct Pattern { - segments: Vec, -} - -impl Pattern { - fn compile(s: &str) -> Self { - let segments = s.trim_start_matches('/') - .split('/') - .filter(|s| !s.is_empty()) - .map(|seg| { - if let Some(name) = seg.strip_prefix(':') { - Segment::Param(name.to_string()) - } else if let Some(name) = seg.strip_prefix('*') { - Segment::Wildcard(name.to_string()) - } else { - Segment::Literal(seg.to_string()) - } - }) - .collect(); - Self { segments } - } - - fn matches(&self, path: &str) -> Option> { - let parts: Vec<&str> = path.trim_start_matches('/') - .split('/') - .filter(|s| !s.is_empty()) - .collect(); - let mut params = Vec::new(); - let mut pi = 0usize; - for seg in &self.segments { - match seg { - Segment::Literal(lit) => { - if pi >= parts.len() || parts[pi] != lit { return None; } - pi += 1; - } - Segment::Param(name) => { - if pi >= parts.len() { return None; } - params.push((name.clone(), parts[pi].to_string())); - pi += 1; - } - Segment::Wildcard(name) => { - if pi >= parts.len() { return None; } - params.push((name.clone(), parts[pi..].join("/"))); - return Some(params); - } - } - } - if pi == parts.len() { Some(params) } else { None } - } -} - -#[derive(Debug, Clone, Default)] -pub struct RouteParams { - params: HashMap, -} - -impl RouteParams { - pub fn get(&self, k: &str) -> Option<&str> { - self.params.get(k).map(String::as_str) - } - - fn from_pairs(pairs: Vec<(String, String)>) -> Self { - Self { params: pairs.into_iter().collect() } - } -} - -struct Route { - method: Method, - pattern: Pattern, - handler: Box, -} - -#[derive(Default)] -pub struct Router { - routes: Vec, -} - -impl Router { - pub fn new() -> Self { Self::default() } - - pub fn route(mut self, method: Method, pattern: &str, handler: F) -> Self - where - F: Fn(&Request, &RouteParams) -> Response + 'static, - { - self.routes.push(Route { - method, - pattern: Pattern::compile(pattern), - handler: Box::new(handler), - }); - self - } - - /// Resolve a request to a response. 404 if no path matches; 405 if the - /// path matches a registered route under a different method. - pub fn dispatch(&self, req: &Request) -> Response { - let mut path_matched_any = false; - for r in &self.routes { - if let Some(pairs) = r.pattern.matches(&req.path) { - if r.method == req.method { - let params = RouteParams::from_pairs(pairs); - return (r.handler)(req, ¶ms); - } - path_matched_any = true; - } - } - if path_matched_any { - Response::status(Status::METHOD_NOT_ALLOWED).text("method not allowed") - } else { - Response::status(Status::NOT_FOUND).text("no route") - } - } -} - -#[cfg(test)] -mod tests { - use super::*; - use std::collections::HashMap; - - fn req(method: Method, path: &str) -> Request { - Request { method, path: path.into(), query: None, headers: HashMap::new(), body: vec![], keep_alive: true } - } - - #[test] - fn dispatches_static_and_param_routes() { - let r = Router::new() - .route(Method::Get, "/healthz", |_, _| Response::ok().text("ok")) - .route(Method::Get, "/api/articles", |_, _| Response::ok().text("list")) - .route(Method::Get, "/api/articles/:id", |_, p| Response::ok().text(format!("get {}", p.get("id").unwrap()))) - .route(Method::Post, "/api/articles", |_, _| Response::created().text("create")); - - let resp = r.dispatch(&req(Method::Get, "/healthz")); - assert_eq!(resp.status.0, 200); - assert_eq!(resp.body, b"ok"); - - let resp = r.dispatch(&req(Method::Get, "/api/articles/42")); - assert_eq!(resp.body, b"get 42"); - - let resp = r.dispatch(&req(Method::Post, "/api/articles")); - assert_eq!(resp.status.0, 201); - } - - #[test] - fn returns_405_when_path_matches_but_method_does_not() { - let r = Router::new() - .route(Method::Get, "/api/articles", |_, _| Response::ok().text("list")); - let resp = r.dispatch(&req(Method::Post, "/api/articles")); - assert_eq!(resp.status.0, 405); - } - - #[test] - fn returns_404_when_no_path_matches() { - let r = Router::new() - .route(Method::Get, "/api/articles", |_, _| Response::ok().text("list")); - let resp = r.dispatch(&req(Method::Get, "/nope")); - assert_eq!(resp.status.0, 404); - } -} diff --git a/crates/rt/src/lexer.rs b/crates/rt/src/lexer.rs deleted file mode 100644 index 73ff0ac..0000000 --- a/crates/rt/src/lexer.rs +++ /dev/null @@ -1,343 +0,0 @@ -//! Tokenizer for `.wo` source. Emits a sequential stream of [`Token`]s -//! suitable for the recursive-descent parser in [`crate::parser`]. - -use crate::token::{Kind, Token}; -use anyhow::{bail, Result}; - -pub fn tokenize(src: &str) -> Result> { - Lexer::new(src).lex() -} - -struct Lexer<'a> { - bytes: &'a [u8], - pos: usize, - line: u32, - col: u32, -} - -impl<'a> Lexer<'a> { - fn new(src: &'a str) -> Self { - Self { bytes: src.as_bytes(), pos: 0, line: 1, col: 1 } - } - - fn peek(&self) -> Option { self.bytes.get(self.pos).copied() } - fn peek_at(&self, n: usize) -> Option { self.bytes.get(self.pos + n).copied() } - - fn advance(&mut self) -> Option { - let c = self.peek()?; - self.pos += 1; - if c == b'\n' { self.line += 1; self.col = 1; } else { self.col += 1; } - Some(c) - } - - fn lex(mut self) -> Result> { - let mut out = Vec::new(); - while let Some(c) = self.peek() { - let line = self.line; - let col = self.col; - - // line comment: -- ... EOL - if c == b'-' && self.peek_at(1) == Some(b'-') { - while let Some(c) = self.peek() { - if c == b'\n' { break; } - self.advance(); - } - continue; - } - - // newline → significant (ends trigger/policy lines, etc.) - if c == b'\n' { - self.advance(); - if out.last().map(|t: &Token| matches!(t.kind, Kind::Newline)) != Some(true) { - out.push(Token { kind: Kind::Newline, line, col }); - } - continue; - } - - // plain whitespace - if c == b' ' || c == b'\t' || c == b'\r' { - self.advance(); - continue; - } - - // ## block marker - if c == b'#' && self.peek_at(1) == Some(b'#') { - self.advance(); self.advance(); - let name = self.read_ident_chars(); - out.push(Token { kind: Kind::HashHash(name), line, col }); - continue; - } - - // # name - if c == b'#' { - self.advance(); - let name = self.read_ident_chars(); - out.push(Token { kind: Kind::Hash(name), line, col }); - continue; - } - - // $name - if c == b'$' { - self.advance(); - let name = self.read_ident_chars(); - if name.is_empty() { - bail!("line {line}: expected parameter name after '$'"); - } - out.push(Token { kind: Kind::Param(name), line, col }); - continue; - } - - // string literal (single or double quote) - if c == b'"' || c == b'\'' { - let quote = c; - self.advance(); - let mut s = String::new(); - while let Some(c) = self.peek() { - if c == quote { self.advance(); break; } - if c == b'\\' { - self.advance(); - match self.advance() { - Some(b'n') => s.push('\n'), - Some(b't') => s.push('\t'), - Some(b'\\') => s.push('\\'), - Some(b'"') => s.push('"'), - Some(b'\'') => s.push('\''), - Some(other) => s.push(other as char), - None => bail!("line {line}: unterminated string escape"), - } - continue; - } - s.push(self.advance().unwrap() as char); - } - out.push(Token { kind: Kind::Str(s), line, col }); - continue; - } - - // integer literal - if c.is_ascii_digit() { - let mut n: i64 = 0; - while let Some(d) = self.peek() { - if !d.is_ascii_digit() { break; } - n = n.saturating_mul(10) + (d - b'0') as i64; - self.advance(); - } - out.push(Token { kind: Kind::Int(n), line, col }); - continue; - } - - // identifier / keyword - if c.is_ascii_alphabetic() || c == b'_' { - let name = self.read_ident_chars(); - let kind = match name.as_str() { - "type" => Kind::KwType, - "class" => Kind::KwClass, - "ref" => Kind::KwRef, - "multi" => Kind::KwMulti, - "via" => Kind::KwVia, - "backlink" => Kind::KwBacklink, - "link" => Kind::KwLink, - "service" => Kind::KwService, - "rest" => Kind::KwRest, - "graphql" => Kind::KwGraphql, - "native" => Kind::KwNative, - "expose" => Kind::KwExpose, - "policy" => Kind::KwPolicy, - "for" => Kind::KwFor, - "role" => Kind::KwRole, - "when" => Kind::KwWhen, - "anyone" => Kind::KwAnyone, - "on" => Kind::KwOn, - "do" => Kind::KwDo, - "set" => Kind::KwSet, - "call" => Kind::KwCall, - "emit" => Kind::KwEmit, - "enqueue" => Kind::KwEnqueue, - "assert" => Kind::KwAssert, - "otherwise" => Kind::KwOtherwise, - "abort" => Kind::KwAbort, - "return" => Kind::KwReturn, - "returning" => Kind::KwReturning, - "RETURNING" => Kind::KwReturning, - "as" => Kind::KwAs, - "AS" => Kind::KwAs, - "fn" => Kind::KwFn, - "in" => Kind::KwIn, - "txn" => Kind::KwTxn, - "snapshot" => Kind::KwSnapshot, - "serializable" => Kind::KwSerializable, - "BEGIN" => Kind::KwBegin, - "COMMIT" => Kind::KwCommit, - "ROLLBACK" => Kind::KwRollback, - "SAVEPOINT" => Kind::KwSavepoint, - "TO" => Kind::KwTo, - "LIVE" => Kind::KwLive, - "live" => Kind::KwLive, // lowercase `live` used in UI blocks - // `subscribe`, `receive`, `expect_abort` stay as plain idents so - // `expose ... subscribe` works in `service rest` blocks. - "INSERT" => Kind::KwInsert, - "INTO" => Kind::KwInto, - "VALUES" => Kind::KwValues, - "UPDATE" => Kind::KwUpdate, - "DELETE" => Kind::KwDelete, - "SELECT" => Kind::KwSelect, - "FROM" => Kind::KwFrom, - "WHERE" => Kind::KwWhere, - "MATCH" => Kind::KwMatch, - "CREATE" => Kind::KwCreate, - "SET" => Kind::KwSet, - "let" => Kind::KwLet, - "if" => Kind::KwIf, - "else" => Kind::KwElse, - "each" => Kind::KwEach, - "contains" => Kind::KwContains, - "and" => Kind::KwAnd, - "AND" => Kind::KwAnd, - "or" => Kind::KwOr, - "OR" => Kind::KwOr, - "not" => Kind::KwNot, - "NOT" => Kind::KwNot, - "true" => Kind::KwTrue, - "false" => Kind::KwFalse, - "null" => Kind::KwNull, - "test" => Kind::KwTest, - "main" => Kind::KwMain, - "startup" => Kind::KwStartup, - _ => Kind::Ident(name), - }; - out.push(Token { kind, line, col }); - continue; - } - - // punctuation & operators - let kind = match c { - b'{' => { self.advance(); Kind::LBrace } - b'}' => { self.advance(); Kind::RBrace } - b'(' => { self.advance(); Kind::LParen } - b')' => { self.advance(); Kind::RParen } - b'[' => { self.advance(); Kind::LBracket } - b']' => { self.advance(); Kind::RBracket } - b',' => { self.advance(); Kind::Comma } - b';' => { self.advance(); Kind::Semicolon } - b':' => { self.advance(); Kind::Colon } - b'.' => { - self.advance(); - match self.peek() { - Some(b'.') => { self.advance(); Kind::DotDot } - Some(b'*') => { self.advance(); Kind::DotStar } - _ => Kind::Dot, - } - } - b'?' => { self.advance(); Kind::Question } - b'@' => { self.advance(); Kind::At } - b'|' => { self.advance(); Kind::Pipe } - b'-' => { - self.advance(); - match self.peek() { - Some(b'>') => { self.advance(); Kind::Arrow } - Some(b'=') => { self.advance(); Kind::MinusEq } - _ => Kind::Dash, - } - } - b'+' => { - self.advance(); - match self.peek() { - Some(b'=') => { self.advance(); Kind::PlusEq } - _ => Kind::Plus, - } - } - b'*' => { self.advance(); Kind::Star } - b'/' => { self.advance(); Kind::Slash } - b'%' => { self.advance(); Kind::Percent } - b'=' => { - self.advance(); - match self.peek() { - Some(b'=') => { self.advance(); Kind::EqEq } - Some(b'>') => { self.advance(); Kind::FatArrow } - _ => Kind::Eq, - } - } - b'!' => { - self.advance(); - match self.peek() { - Some(b'=') => { self.advance(); Kind::NotEq } - _ => bail!("line {line}: expected '!=' got '!'"), - } - } - b'<' => { - self.advance(); - match self.peek() { - Some(b'=') => { self.advance(); Kind::LtEq } - _ => Kind::Lt, - } - } - b'>' => { - self.advance(); - match self.peek() { - Some(b'=') => { self.advance(); Kind::GtEq } - _ => Kind::Gt, - } - } - other => bail!("line {line}, col {col}: unexpected character {:?}", other as char), - }; - out.push(Token { kind, line, col }); - } - - out.push(Token { kind: Kind::End, line: self.line, col: self.col }); - Ok(out) - } - - fn read_ident_chars(&mut self) -> String { - let start = self.pos; - while let Some(c) = self.peek() { - if c.is_ascii_alphanumeric() || c == b'_' || c == b'-' { - self.advance(); - } else { - break; - } - } - String::from_utf8_lossy(&self.bytes[start..self.pos]).into_owned() - } -} - -#[cfg(test)] -mod tests { - use super::*; - - #[test] - fn lexes_type_header() { - let toks = tokenize("type Article {").unwrap(); - let kinds: Vec<_> = toks.iter().map(|t| format!("{:?}", t.kind)).collect(); - assert_eq!(kinds, vec![ - "KwType".to_string(), - "Ident(\"Article\")".to_string(), - "LBrace".to_string(), - "End".to_string(), - ]); - } - - #[test] - fn lexes_string_and_int_and_param() { - let toks = tokenize(r#"VALUES ($uid, 'hello', 42)"#).unwrap(); - let kinds: Vec<_> = toks.iter().map(|t| &t.kind).cloned().collect(); - assert!(kinds.contains(&Kind::KwValues)); - assert!(kinds.contains(&Kind::Param("uid".into()))); - assert!(kinds.contains(&Kind::Str("hello".into()))); - assert!(kinds.contains(&Kind::Int(42))); - } - - #[test] - fn skips_line_comments() { - let toks = tokenize("-- comment\ntype X {}").unwrap(); - // First non-newline non-comment token should be `type`. - let first_meaningful = toks.iter().find(|t| !matches!(t.kind, Kind::Newline)).unwrap(); - assert_eq!(first_meaningful.kind, Kind::KwType); - } - - #[test] - fn lexes_hash_markers() { - let toks = tokenize("##ui\n#article-list").unwrap(); - assert!(matches!(toks[0].kind, Kind::HashHash(ref s) if s == "ui")); - assert!(matches!(toks.iter().find(|t| matches!(t.kind, Kind::Hash(_))).unwrap().kind, - Kind::Hash(ref s) if s == "article-list")); - } -} diff --git a/crates/rt/src/lib.rs b/crates/rt/src/lib.rs deleted file mode 100644 index 7e4b907..0000000 --- a/crates/rt/src/lib.rs +++ /dev/null @@ -1,103 +0,0 @@ -//! writeonce runtime — `.wo` language engine. -//! -//! Stage 1: file discovery. ← src/lib.rs::discover() -//! Stage 2: parser + engine + server. -//! Stage 3: LIVE subscriptions over WebSocket. - -pub mod ast; -pub mod compile; -pub mod engine; -pub mod http; -pub mod lexer; -pub mod method; -pub mod mirror; -pub mod parser; -pub mod pg; -pub mod runtime; -pub mod server; -pub mod shard; -pub mod token; -pub mod wal; - -use std::fs; -use std::path::{Path, PathBuf}; - -pub use ast::Schema; -pub use compile::Catalog; -pub use engine::Engine; - -/// A discovered `.wo` source file, resolved to an absolute path with its -/// contents slurped into memory. -#[derive(Debug, Clone)] -pub struct WoFile { - pub path: PathBuf, - pub rel: PathBuf, - pub src: String, -} - -/// Discover every `.wo` file rooted at `dir`, returning them in stable -/// (sorted-by-relative-path) order. Recursive. -pub fn discover(dir: &Path) -> anyhow::Result> { - let root = dir.canonicalize().map_err(|e| { - anyhow::anyhow!("cannot resolve {}: {}", dir.display(), e) - })?; - let mut out = Vec::new(); - walk(&root, &root, &mut out)?; - out.sort_by(|a, b| a.rel.cmp(&b.rel)); - Ok(out) -} - -fn walk(root: &Path, dir: &Path, out: &mut Vec) -> anyhow::Result<()> { - for entry in fs::read_dir(dir)? { - let entry = entry?; - let path = entry.path(); - let ty = entry.file_type()?; - - let name = entry.file_name(); - let name = name.to_string_lossy(); - if name.starts_with('.') || name == "target" || name == "data" || name == "node_modules" { - continue; - } - - if ty.is_dir() { - walk(root, &path, out)?; - } else if ty.is_file() - && path.extension().and_then(|s| s.to_str()) == Some("wo") - { - let src = fs::read_to_string(&path)?; - let rel = path.strip_prefix(root).unwrap_or(&path).to_path_buf(); - out.push(WoFile { path: path.clone(), rel, src }); - } - } - Ok(()) -} - -#[cfg(test)] -mod tests { - use super::*; - use std::fs; - - #[test] - fn discovers_wo_files_recursively() { - let tmp = tempfile::tempdir().unwrap(); - let root = tmp.path(); - - fs::write(root.join("app.wo"), "-- app").unwrap(); - fs::create_dir(root.join("types")).unwrap(); - fs::write(root.join("types/article.wo"), "-- article").unwrap(); - fs::write(root.join("types/README.md"), "not a wo file").unwrap(); - - fs::create_dir(root.join(".hidden")).unwrap(); - fs::write(root.join(".hidden/x.wo"), "").unwrap(); - fs::create_dir(root.join("target")).unwrap(); - fs::write(root.join("target/built.wo"), "").unwrap(); - fs::create_dir(root.join("data")).unwrap(); - fs::write(root.join("data/runtime.wo"), "").unwrap(); - - let files = discover(root).unwrap(); - assert_eq!(files.len(), 2, "expected 2 .wo files, got {:?}", - files.iter().map(|f| &f.rel).collect::>()); - assert!(files.iter().any(|f| f.rel == PathBuf::from("app.wo"))); - assert!(files.iter().any(|f| f.rel == PathBuf::from("types/article.wo"))); - } -} diff --git a/crates/rt/src/method.rs b/crates/rt/src/method.rs deleted file mode 100644 index b508e59..0000000 --- a/crates/rt/src/method.rs +++ /dev/null @@ -1,592 +0,0 @@ -//! Row-scoped method execution — plan 13b (`docs/plan/13-class-model-live-pricing.md`). -//! -//! A class method is a free `fn` with a hidden first parameter: `self` binds -//! to the receiving row (fetched by the RPC route's `:id` on the owning -//! shard), arguments arrive as a JSON object, and the body — the -//! schema-layer DML of `02-wo-language.md` — runs inside an engine method -//! transaction. Commit emits ONE `WalRec::Txn` frame (all-or-nothing on -//! replay); any abort or execution error rolls RAM back completely. -//! -//! The Stage-2 engine is single-threaded per shard, so `in txn` / -//! `in txn snapshot` are already serializable by construction — the mode is -//! accepted and recorded, and every method (pure ones included) runs under -//! begin/commit so the semantics stay uniform when MVCC arrives. - -use crate::ast::{BinOp, Expr, FieldTy, MethodDecl, Stmt, UnOp}; -use crate::engine::Engine; - -use serde_json::{json, Map, Value}; - -/// Why a method call failed — shaped for the RPC layer's status mapping. -#[derive(Debug)] -pub enum MethodError { - /// No row of the receiving type with the requested id (→ 404). - NoSuchRow, - /// Missing / malformed arguments (→ 400). - BadArgs(String), - /// `assert … otherwise abort` fired — transaction rolled back (→ 409). - Abort(String), - /// Execution error (bad field, arithmetic on non-numbers, WAL failure…) - /// — transaction rolled back (→ 500). - Exec(String), -} - -impl std::fmt::Display for MethodError { - fn fmt(&self, f: &mut std::fmt::Formatter<'_>) -> std::fmt::Result { - match self { - MethodError::NoSuchRow => write!(f, "no such row"), - MethodError::BadArgs(m) => write!(f, "bad arguments: {m}"), - MethodError::Abort(m) => write!(f, "aborted: {m}"), - MethodError::Exec(m) => write!(f, "{m}"), - } - } -} - -/// Call `m` on row `id` of type `ty`. Runs entirely on the engine it's -/// handed — the RPC route routes to the owning shard before calling this. -/// Returns the method's `return` value (`Null` if it falls off the end). -pub fn call( - e: &mut Engine, - ty: &str, - id: i64, - m: &MethodDecl, - args: &Map, -) -> Result { - // `self` is a snapshot of the receiving row at entry (spec: bindings - // are snapshots; the row-scoped txn reads one consistent state). - let row = e.get(ty, id) - .map_err(|err| MethodError::Exec(err.to_string()))? - .ok_or(MethodError::NoSuchRow)?; - - // Bind declared parameters. Missing → 400; extras are ignored. - let mut scope: Map = Map::new(); - for (pname, pty) in &m.params { - let Some(v) = args.get(pname) else { - return Err(MethodError::BadArgs(format!("missing argument `{pname}` ({pty})"))); - }; - scope.insert(pname.clone(), v.clone()); - } - - e.begin_txn().map_err(|err| MethodError::Exec(err.to_string()))?; - let mut cx = Cx { e, self_ty: ty, self_row: &row, scope }; - match exec_block(&mut cx, &m.body) { - Ok(flow) => { - cx.e.commit_txn().map_err(|err| MethodError::Exec(err.to_string()))?; - Ok(match flow { Flow::Return(v) => v, Flow::Continue => Value::Null }) - } - Err(err) => { - cx.e.abort_txn(); - Err(err) - } - } -} - -/// Statement-level control flow. -enum Flow { - Continue, - Return(Value), -} - -struct Cx<'a> { - e: &'a mut Engine, - self_ty: &'a str, - self_row: &'a Map, - scope: Map, -} - -fn exec_block(cx: &mut Cx, stmts: &[Stmt]) -> Result { - for s in stmts { - match exec_stmt(cx, s)? { - Flow::Continue => {} - r @ Flow::Return(_) => return Ok(r), - } - } - Ok(Flow::Continue) -} - -fn exec_stmt(cx: &mut Cx, s: &Stmt) -> Result { - match s { - Stmt::Let { name, expr } => { - let v = eval(cx, expr)?; - cx.scope.insert(name.clone(), v); - Ok(Flow::Continue) - } - Stmt::Insert { ty, fields } => { - let mut body = Map::new(); - for (fname, fexpr) in fields { - body.insert(fname.clone(), eval(cx, fexpr)?); - } - cx.e.create(ty, Value::Object(body)) - .map_err(|e| MethodError::Exec(format!("insert {ty}: {e}")))?; - Ok(Flow::Continue) - } - Stmt::Return { expr } => { - let v = match expr { - Some(e) => eval(cx, e)?, - None => Value::Null, - }; - Ok(Flow::Return(v)) - } - Stmt::Assert { cond, msg } => { - if truthy(&eval(cx, cond)?) { - Ok(Flow::Continue) - } else { - Err(MethodError::Abort( - msg.clone().unwrap_or_else(|| "assertion failed".into()))) - } - } - Stmt::If { cond, then, otherwise } => { - if truthy(&eval(cx, cond)?) { - exec_block(cx, then) - } else { - exec_block(cx, otherwise) - } - } - } -} - -fn truthy(v: &Value) -> bool { - match v { - Value::Bool(b) => *b, - Value::Null => false, - _ => true, - } -} - -fn eval(cx: &mut Cx, e: &Expr) -> Result { - match e { - Expr::Int(n) => Ok(json!(n)), - Expr::Str(s) => Ok(Value::String(s.clone())), - Expr::Bool(b) => Ok(json!(b)), - Expr::Null => Ok(Value::Null), - Expr::Ident(name) => { - if name == "self" { - return Ok(Value::Object(cx.self_row.clone())); - } - cx.scope.get(name).cloned() - .ok_or_else(|| MethodError::Exec(format!("unknown name `{name}`"))) - } - Expr::Field(base, field) => { - // `self.` resolves the relation; anything else is - // plain object access on the evaluated base. - if matches!(&**base, Expr::Ident(n) if n == "self") { - if let Some(rel) = relation_rows(cx, field)? { - return Ok(rel); - } - } - let b = eval(cx, base)?; - match b { - Value::Object(m) => m.get(field).cloned().ok_or_else(|| - MethodError::Exec(format!("no field `{field}`"))), - // Dotted access distributes over a set — the spec's - // cardinality rule: `select Price{...}.amount` is the set - // of amounts. - Value::Array(items) => { - let mut out = Vec::with_capacity(items.len()); - for it in items { - match it { - Value::Object(m) => out.push(m.get(field).cloned() - .ok_or_else(|| MethodError::Exec( - format!("no field `{field}` in set element")))?), - other => return Err(MethodError::Exec(format!( - "`.{field}` on a non-object set element ({other})"))), - } - } - Ok(Value::Array(out)) - } - other => Err(MethodError::Exec(format!( - "`.{field}` on a non-object value ({other})"))), - } - } - Expr::Select { ty, predicates, projection } => { - // Equality predicates route through the engine's secondary - // indexes (`@table(index: ...)`) via find_by; other comparison - // operators filter the candidates. - let mut eq = Vec::new(); - let mut rest = Vec::new(); - for (field, op, rhs) in predicates { - let v = eval(cx, rhs)?; - if *op == BinOp::Eq { eq.push((field.clone(), v)); } - else { rest.push((field.as_str(), *op, v)); } - } - let rows = cx.e.find_by(ty, &eq) - .map_err(|e| MethodError::Exec(format!("select {ty}: {e}")))?; - let mut out = Vec::new(); - for row in rows { - if !rest.iter().all(|(f, op, v)| pred_holds(row.get(*f), *op, v)) { - continue; - } - out.push(if projection.is_empty() { - Value::Object(row) - } else { - let mut shaped = Map::new(); - for p in projection { - if let Some(v) = row.get(p) { - shaped.insert(p.clone(), v.clone()); - } - } - Value::Object(shaped) - }); - } - Ok(Value::Array(out)) - } - Expr::Call(name, args) => { - let mut vals = Vec::with_capacity(args.len()); - for a in args { vals.push(eval(cx, a)?); } - builtin(name, vals) - } - Expr::Unary(op, inner) => { - let v = eval(cx, inner)?; - match op { - UnOp::Neg => { - let n = as_i64(&v)?; - Ok(json!(-n)) - } - UnOp::Not => Ok(json!(!truthy(&v))), - } - } - Expr::Binary(op, l, r) => { - // Short-circuit boolean operators. - match op { - BinOp::And => { - let lv = eval(cx, l)?; - if !truthy(&lv) { return Ok(json!(false)); } - return Ok(json!(truthy(&eval(cx, r)?))); - } - BinOp::Or => { - let lv = eval(cx, l)?; - if truthy(&lv) { return Ok(json!(true)); } - return Ok(json!(truthy(&eval(cx, r)?))); - } - _ => {} - } - let lv = eval(cx, l)?; - let rv = eval(cx, r)?; - match op { - BinOp::Eq => Ok(json!(lv == rv)), - BinOp::Ne => Ok(json!(lv != rv)), - BinOp::Lt | BinOp::Le | BinOp::Gt | BinOp::Ge => { - let (a, b) = (as_i64(&lv)?, as_i64(&rv)?); - Ok(json!(match op { - BinOp::Lt => a < b, - BinOp::Le => a <= b, - BinOp::Gt => a > b, - _ => a >= b, - })) - } - BinOp::Add | BinOp::Sub | BinOp::Mul | BinOp::Div | BinOp::Mod => { - let (a, b) = (as_i64(&lv)?, as_i64(&rv)?); - if b == 0 && matches!(op, BinOp::Div | BinOp::Mod) { - return Err(MethodError::Exec("division by zero".into())); - } - Ok(json!(match op { - BinOp::Add => a + b, - BinOp::Sub => a - b, - BinOp::Mul => a * b, - BinOp::Div => a / b, - _ => a % b, - })) - } - BinOp::And | BinOp::Or => unreachable!("handled above"), - } - } - } -} - -fn as_i64(v: &Value) -> Result { - v.as_i64().ok_or_else(|| MethodError::Exec(format!("expected a number, got {v}"))) -} - -/// Non-equality select predicate over a row column. Numbers compare -/// numerically, strings lexicographically (covers `at > "2026-…"`); -/// mismatched or missing values fail the predicate rather than erroring — -/// a filter, not an expression. -fn pred_holds(actual: Option<&Value>, op: BinOp, wanted: &Value) -> bool { - let Some(a) = actual else { return op == BinOp::Ne && !wanted.is_null() }; - match op { - BinOp::Ne => a != wanted, - BinOp::Lt | BinOp::Le | BinOp::Gt | BinOp::Ge => { - let ord = match (a.as_i64(), wanted.as_i64()) { - (Some(x), Some(y)) => x.cmp(&y), - _ => match (a.as_str(), wanted.as_str()) { - (Some(x), Some(y)) => x.cmp(y), - _ => return false, - }, - }; - match op { - BinOp::Lt => ord.is_lt(), - BinOp::Le => ord.is_le(), - BinOp::Gt => ord.is_gt(), - _ => ord.is_ge(), - } - } - _ => unreachable!("equality predicates go through find_by"), - } -} - -/// If `field` names a relation on the receiving type, materialize it: -/// `multi T` / `backlink T.f` → array of the related rows, in id order. -/// Returns `Ok(None)` when `field` is not a relation (plain column access). -/// -/// Shard note: the scan runs on the executing (owner) shard only. Related -/// rows created BY methods land here too (creates mint locally on the shard -/// that runs the method — `set_price`'s Price is by construction visible to -/// the same product's `current_price`). Cross-shard relation reads are 09d/ -/// 09e territory. -fn relation_rows(cx: &mut Cx, field: &str) -> Result, MethodError> { - let Some(t) = cx.e.catalog().get(cx.self_ty) else { return Ok(None) }; - let Some(f) = t.fields.iter().find(|f| f.name == field && f.is_relation) else { - return Ok(None); - }; - let (target, link_field) = match &f.ty { - FieldTy::MultiEdge { target, .. } | FieldTy::MultiVia { target, .. } => { - // Find the ref column in the target type that points back at us. - let tt = cx.e.catalog().get(target).ok_or_else(|| - MethodError::Exec(format!("relation `{field}`: unknown type {target}")))?; - let mut backrefs = tt.fields.iter().filter(|g| - matches!(&g.ty, FieldTy::Ref(r) if r == cx.self_ty)); - let Some(link) = backrefs.next() else { - return Err(MethodError::Exec(format!( - "relation `{field}`: {target} has no `ref {}` field", cx.self_ty))); - }; - if backrefs.next().is_some() { - return Err(MethodError::Exec(format!( - "relation `{field}`: {target} has multiple refs to {} — ambiguous", cx.self_ty))); - } - (target.clone(), link.name.clone()) - } - FieldTy::Backlink { target, field: link } => (target.clone(), link.clone()), - _ => return Ok(None), // `ref T` is a stored scalar column — plain access - }; - - let self_id = cx.self_row.get("id").cloned().unwrap_or(Value::Null); - // Index-accelerated when the target declares @table(index: [, …]); - // find_by falls back to a scan otherwise. - let rows = cx.e.find_by(&target, &[(link_field, self_id)]) - .map_err(|e| MethodError::Exec(e.to_string()))? - .into_iter() - .map(Value::Object) - .collect::>(); - Ok(Some(Value::Array(rows))) -} - -/// Built-in functions available in method bodies. -fn builtin(name: &str, mut args: Vec) -> Result { - match (name, args.len()) { - // `latest(set)` — the row with the highest id. Prices are - // append-only, so highest id = most recently inserted (per shard). - ("latest", 1) => { - let Value::Array(items) = args.remove(0) else { - return Err(MethodError::Exec("latest(): expected a set".into())); - }; - items.into_iter() - .max_by_key(|v| v.get("id").and_then(|i| i.as_i64()).unwrap_or(i64::MIN)) - .ok_or_else(|| MethodError::Exec("latest(): empty set".into())) - } - ("count", 1) => { - match &args[0] { - Value::Array(items) => Ok(json!(items.len() as i64)), - other => Err(MethodError::Exec(format!("count(): expected a set, got {other}"))), - } - } - ("now", 0) => Ok(Value::String(crate::engine::now_iso8601())), - _ => Err(MethodError::Exec(format!( - "unknown function `{name}`/{} (built-ins: latest, count, now)", args.len()))), - } -} - -#[cfg(test)] -mod tests { - use super::*; - use crate::compile::Catalog; - use crate::parser::parse; - - const PRICING: &str = r#" -@table(name: "prices", index: [product, at]) -class Price { - id: Id - product: ref Product - amount: Money - currency: Text = "EUR" - at: Text = "t0" - - fn discounted(pct: Int) -> Money { - return self.amount * (100 - pct) / 100; - } -} - -class Product { - id: Id - sku: SKU @unique - name: Text - prices: multi Price - - fn current_price() -> Money in txn { - return latest(self.prices).amount; - } - - fn set_price(amount: Money) in txn { - assert amount > 0 otherwise abort "price must be positive" - insert Price { product: self.id, amount: amount }; - } - - fn history() -> [Money] in txn { - return select Price{ product == self.id, amount }; - } - - service rest "/api/products" expose list, get, create -} -"#; - - fn engine() -> Engine { - let cat = Catalog::from_schemas(vec![parse(PRICING).unwrap()]).unwrap(); - Engine::new(cat) - } - - fn method<'a>(e: &'a Engine, ty: &str, name: &str) -> MethodDecl { - e.catalog().get(ty).unwrap().methods.iter() - .find(|m| m.name == name).unwrap().clone() - } - - fn args(v: Value) -> Map { - match v { Value::Object(m) => m, _ => Map::new() } - } - - #[test] - fn set_price_inserts_and_current_price_reads_it_back() { - let mut e = engine(); - let p = e.create("Product", json!({"sku": "SKU-1", "name": "Gadget"})).unwrap(); - let id = p["id"].as_i64().unwrap(); - - let set = method(&e, "Product", "set_price"); - call(&mut e, "Product", id, &set, &args(json!({"amount": 4999}))).unwrap(); - call(&mut e, "Product", id, &set, &args(json!({"amount": 5999}))).unwrap(); - - let prices = e.list("Price").unwrap(); - assert_eq!(prices.len(), 2); - assert_eq!(prices[0]["product"].as_i64().unwrap(), id); - assert_eq!(prices[0]["currency"], "EUR"); // default seeded by create - - let cur = method(&e, "Product", "current_price"); - let v = call(&mut e, "Product", id, &cur, &Map::new()).unwrap(); - assert_eq!(v, json!(5999)); - } - - #[test] - fn select_expression_filters_and_projects() { - let mut e = engine(); - let p1 = e.create("Product", json!({"sku": "A", "name": "A"})).unwrap(); - let p2 = e.create("Product", json!({"sku": "B", "name": "B"})).unwrap(); - let (id1, id2) = (p1["id"].as_i64().unwrap(), p2["id"].as_i64().unwrap()); - - let set = method(&e, "Product", "set_price"); - call(&mut e, "Product", id1, &set, &args(json!({"amount": 100}))).unwrap(); - call(&mut e, "Product", id1, &set, &args(json!({"amount": 200}))).unwrap(); - call(&mut e, "Product", id2, &set, &args(json!({"amount": 999}))).unwrap(); - - // `history` = select Price{ product == self.id, amount } — the - // equality predicate rides the (product, at) index; the projection - // shapes each row down to { amount }. - let hist = method(&e, "Product", "history"); - let v = call(&mut e, "Product", id1, &hist, &Map::new()).unwrap(); - assert_eq!(v, json!([{"amount": 100}, {"amount": 200}])); - - // Indexed relation read (self.prices) equals the select's row set. - let cur = method(&e, "Product", "current_price"); - assert_eq!(call(&mut e, "Product", id1, &cur, &Map::new()).unwrap(), json!(200)); - assert_eq!(call(&mut e, "Product", id2, &cur, &Map::new()).unwrap(), json!(999)); - } - - #[test] - fn pure_method_computes_from_self() { - let mut e = engine(); - let p = e.create("Product", json!({"sku": "S", "name": "N"})).unwrap(); - let pid = p["id"].as_i64().unwrap(); - let set = method(&e, "Product", "set_price"); - call(&mut e, "Product", pid, &set, &args(json!({"amount": 1000}))).unwrap(); - - let price_id = e.list("Price").unwrap()[0]["id"].as_i64().unwrap(); - let disc = method(&e, "Price", "discounted"); - let v = call(&mut e, "Price", price_id, &disc, &args(json!({"pct": 25}))).unwrap(); - assert_eq!(v, json!(750)); - } - - #[test] - fn abort_rolls_back_completely() { - let mut e = engine(); - let p = e.create("Product", json!({"sku": "S", "name": "N"})).unwrap(); - let id = p["id"].as_i64().unwrap(); - let set = method(&e, "Product", "set_price"); - - // amount <= 0 trips the assert AFTER nothing, but build a stronger - // case: a method that inserts and THEN aborts must leave no row. - let src = r#" -class Product { - id: Id - fn bad(amount: Money) in txn { - insert Price { product: self.id, amount: amount }; - assert false otherwise abort "always" - } -} -"#; - let sch = parse(src).unwrap(); - let bad = sch.types[0].methods[0].clone(); - - let err = call(&mut e, "Product", id, &bad, &args(json!({"amount": 1}))).unwrap_err(); - assert!(matches!(err, MethodError::Abort(ref m) if m == "always")); - assert_eq!(e.list("Price").unwrap().len(), 0, "aborted insert must roll back"); - - // The plain assert path also rejects without side effects. - let err = call(&mut e, "Product", id, &set, &args(json!({"amount": 0}))).unwrap_err(); - assert!(matches!(err, MethodError::Abort(_))); - assert_eq!(e.list("Price").unwrap().len(), 0); - } - - #[test] - fn missing_row_and_missing_arg_are_typed_errors() { - let mut e = engine(); - let set = method(&e, "Product", "set_price"); - assert!(matches!( - call(&mut e, "Product", 999, &set, &args(json!({"amount": 1}))), - Err(MethodError::NoSuchRow))); - - let p = e.create("Product", json!({"sku": "S", "name": "N"})).unwrap(); - let id = p["id"].as_i64().unwrap(); - assert!(matches!( - call(&mut e, "Product", id, &set, &Map::new()), - Err(MethodError::BadArgs(_)))); - } - - #[test] - fn method_commit_is_one_atomic_wal_frame() { - use crate::wal::Wal; - let path = std::env::temp_dir() - .join(format!("wo-method-wal-{}", std::process::id())); - let _ = std::fs::remove_file(&path); - - let cat = Catalog::from_schemas(vec![parse(PRICING).unwrap()]).unwrap(); - let id; - { - let mut e = Engine::new(cat.clone()); - let (wal, n) = Wal::open_and_replay(&path, &mut e).unwrap(); - assert_eq!(n, 0); - e.attach_wal(wal); - let p = e.create("Product", json!({"sku": "S", "name": "N"})).unwrap(); - id = p["id"].as_i64().unwrap(); - let set = method(&e, "Product", "set_price"); - call(&mut e, "Product", id, &set, &args(json!({"amount": 4999}))).unwrap(); - } - // Recovery: the create frame + ONE txn frame replay into a fresh engine. - let mut e = Engine::new(cat); - let (_, n) = Wal::open_and_replay(&path, &mut e).unwrap(); - assert_eq!(n, 2, "one create frame + one method-txn frame"); - let prices = e.list("Price").unwrap(); - assert_eq!(prices.len(), 1); - assert_eq!(prices[0]["amount"], json!(4999)); - - let cur = method(&e, "Product", "current_price"); - let v = call(&mut e, "Product", id, &cur, &Map::new()).unwrap(); - assert_eq!(v, json!(4999)); - let _ = std::fs::remove_file(&path); - } -} diff --git a/crates/rt/src/mirror.rs b/crates/rt/src/mirror.rs deleted file mode 100644 index d341b50..0000000 --- a/crates/rt/src/mirror.rs +++ /dev/null @@ -1,270 +0,0 @@ -//! PostgreSQL backup mirror — plan 16b (`docs/plan/16-postgres-mirror.md`). -//! -//! RAM is authoritative; this module is the **backup mechanism**: every -//! committed mutation is cloned onto an mpsc channel by the shard engines -//! (after the WAL made it durable — never before, never gating the ack) and -//! a single dedicated `wo-pg` thread drains the channel into Postgres as -//! JSONB upserts. Reads never touch Postgres. -//! -//! Failure doctrine (16b): Postgres being down costs clients nothing — the -//! thread reconnects with capped backoff while the bounded channel absorbs -//! the burst; if the channel fills, records are dropped **loudly** (counted -//! and logged). Lossless catch-up (dirty-flag full resync) is plan 16d; -//! restore-from-Postgres at boot is 16e. -//! -//! Schema (16b): one table per type — `"" (id BIGINT PRIMARY -//! KEY, row JSONB NOT NULL)` — the `@table(name: "prices")` annotation -//! names the table. Typed-column projection is 16c. - -use std::sync::mpsc::{Receiver, RecvTimeoutError, SyncSender}; -use std::time::Duration; - -use serde_json::Value; - -use crate::engine::Row; -use crate::pg::{escape_ident, escape_literal, Conn, PgConfig}; - -/// Mirror channel capacity. At ~200 bytes/record this bounds the buffered -/// backlog around a few tens of MB — enough to ride out a Postgres restart -/// under load without threatening the RAM budget. -pub const QUEUE_CAP: usize = 65_536; - -/// One committed mutation, as the mirror needs it. Unlike `WalRec::Update` -/// (which carries the merge body), `Upsert` always carries the FULL -/// post-merge row — the mirror's `ON CONFLICT ... DO UPDATE` replaces the -/// whole JSONB value. -#[derive(Debug, Clone)] -pub enum MirrorRec { - Upsert { ty: String, id: i64, row: Row }, - Delete { ty: String, id: i64 }, - /// One method transaction (plan 13b) — applied inside one Postgres - /// transaction, mirroring the WAL's atomic `WalRec::Txn` frame. - Txn(Vec), -} - -pub type MirrorSender = SyncSender; - -/// Spawn the `wo-pg` mirror thread. `tables` maps type name → storage -/// (table) name for every catalog type; DDL is bootstrapped on every -/// (re)connect so a fresh database works out of the box. -pub fn spawn( - cfg: PgConfig, - rx: Receiver, - tables: Vec<(String, String)>, -) -> std::thread::JoinHandle<()> { - std::thread::Builder::new() - .name("wo-pg".into()) - .spawn(move || run(cfg, rx, tables)) - .expect("spawn wo-pg mirror thread") -} - -fn run(cfg: PgConfig, rx: Receiver, tables: Vec<(String, String)>) { - let mut dropped: u64 = 0; - loop { - // (Re)connect with capped backoff, bootstrapping DDL each time. - let Some(mut c) = connect_with_backoff(&cfg, &tables, &rx, &mut dropped) else { - return; // channel closed — shutdown - }; - - eprintln!("[wo] pg mirror: connected to {}:{}/{} ({} tables)", - cfg.host, cfg.port, cfg.database, tables.len()); - if dropped > 0 { - eprintln!("[wo] pg mirror: WARNING — {dropped} records were dropped while \ - disconnected; Postgres is behind RAM until a resync (plan 16d)"); - } - - // Drain loop: batch what's queued, one round-trip per batch. - loop { - let batch = match next_batch(&rx) { - Some(b) => b, - None => return, // senders gone — shutdown - }; - if let Err(e) = apply_batch(&mut c, &tables, &batch) { - eprintln!("[wo] pg mirror: connection lost ({e}) — reconnecting"); - break; // outer loop reconnects - } - } - } -} - -/// Block for the next record, then opportunistically drain up to a batch. -/// `None` = all senders dropped (process shutting down). -fn next_batch(rx: &Receiver) -> Option> { - const BATCH: usize = 512; - let first = rx.recv().ok()?; - let mut batch = vec![first]; - while batch.len() < BATCH { - match rx.try_recv() { - Ok(rec) => batch.push(rec), - Err(_) => break, - } - } - Some(batch) -} - -/// `None` = every sender is gone (process shutting down). -fn connect_with_backoff( - cfg: &PgConfig, - tables: &[(String, String)], - rx: &Receiver, - dropped: &mut u64, -) -> Option { - let mut delay = Duration::from_millis(200); - loop { - match Conn::connect(cfg) { - Ok(mut c) => match bootstrap_ddl(&mut c, tables) { - Ok(()) => return Some(c), - Err(e) => eprintln!("[wo] pg mirror: DDL bootstrap failed ({e}) — retrying"), - }, - Err(e) => eprintln!("[wo] pg mirror: connect failed ({e}) — retrying in {delay:?}"), - } - // While waiting, keep the channel from silently backing up forever: - // absorb what we can into the void, counting the loss (16b policy — - // 16d replaces this with dirty-flag resync). - let wait_until = std::time::Instant::now() + delay; - loop { - let left = wait_until.saturating_duration_since(std::time::Instant::now()); - if left.is_zero() { break; } - match rx.recv_timeout(left.min(Duration::from_millis(100))) { - Ok(_) => { *dropped += 1; } - Err(RecvTimeoutError::Timeout) => {} - Err(RecvTimeoutError::Disconnected) => return None, - } - } - delay = (delay * 2).min(Duration::from_secs(5)); - } -} - -fn bootstrap_ddl(c: &mut Conn, tables: &[(String, String)]) -> Result<(), crate::pg::PgError> { - for (_, storage) in tables { - c.simple_query(&format!( - "CREATE TABLE IF NOT EXISTS {} (id BIGINT PRIMARY KEY, row JSONB NOT NULL)", - escape_ident(storage)))?; - } - Ok(()) -} - -/// Apply one batch. Statement/data errors are isolated per record and -/// logged (the batch continues); only I/O errors propagate (→ reconnect). -fn apply_batch( - c: &mut Conn, - tables: &[(String, String)], - batch: &[MirrorRec], -) -> Result<(), crate::pg::PgError> { - for rec in batch { - let sql = rec_sql(tables, rec); - match c.simple_query(&sql) { - Ok(_) => {} - Err(e) if e.severity == "CLIENT" => return Err(e), // socket-level: reconnect - Err(e) => eprintln!("[wo] pg mirror: statement rejected ({e}) — record skipped"), - } - } - Ok(()) -} - -/// Render one record as SQL. A `Txn` becomes BEGIN; …; COMMIT in a single -/// simple-query message — atomic on the Postgres side like its WAL frame. -fn rec_sql(tables: &[(String, String)], rec: &MirrorRec) -> String { - match rec { - MirrorRec::Upsert { ty, id, row } => { - let table = storage_for(tables, ty); - let json = serde_json::to_string(&Value::Object(row.clone())) - .unwrap_or_else(|_| "{}".into()); - format!( - "INSERT INTO {} (id, row) VALUES ({}, {}::jsonb) \ - ON CONFLICT (id) DO UPDATE SET row = EXCLUDED.row", - escape_ident(table), id, escape_literal(&json)) - } - MirrorRec::Delete { ty, id } => { - format!("DELETE FROM {} WHERE id = {}", escape_ident(storage_for(tables, ty)), id) - } - MirrorRec::Txn(recs) => { - let mut sql = String::from("BEGIN"); - for r in recs { - sql.push_str("; "); - sql.push_str(&rec_sql(tables, r)); - } - sql.push_str("; COMMIT"); - sql - } - } -} - -fn storage_for<'a>(tables: &'a [(String, String)], ty: &'a str) -> &'a str { - tables.iter() - .find(|(t, _)| t == ty) - .map(|(_, s)| s.as_str()) - .unwrap_or(ty) -} - -#[cfg(test)] -mod tests { - use super::*; - use serde_json::json; - - fn row(v: Value) -> Row { - match v { Value::Object(m) => m, _ => panic!() } - } - - #[test] - fn sql_rendering_upsert_delete_txn() { - let tables = vec![("Price".to_string(), "prices".to_string())]; - let up = MirrorRec::Upsert { - ty: "Price".into(), id: 2, - row: row(json!({"amount": 4999, "note": "it's"})), - }; - let sql = rec_sql(&tables, &up); - assert!(sql.starts_with(r#"INSERT INTO "prices" (id, row) VALUES (2, '{"#), "{sql}"); - assert!(sql.contains("''s"), "quote must be doubled: {sql}"); - assert!(sql.ends_with("ON CONFLICT (id) DO UPDATE SET row = EXCLUDED.row")); - - let del = MirrorRec::Delete { ty: "Price".into(), id: 7 }; - assert_eq!(rec_sql(&tables, &del), r#"DELETE FROM "prices" WHERE id = 7"#); - - // Unmapped type falls back to the type name. - let other = MirrorRec::Delete { ty: "Ghost".into(), id: 1 }; - assert_eq!(rec_sql(&tables, &other), r#"DELETE FROM "Ghost" WHERE id = 1"#); - - let txn = MirrorRec::Txn(vec![up, del]); - let sql = rec_sql(&tables, &txn); - assert!(sql.starts_with("BEGIN; ")); - assert!(sql.ends_with("; COMMIT")); - } - - /// Integration: full pipeline against a live server (WO_PG_TEST gated). - #[test] - fn mirror_pipeline_against_live_server() { - let Ok(url) = std::env::var("WO_PG_TEST") else { - eprintln!("mirror_pipeline: skipped (set WO_PG_TEST=postgres://... to run)"); - return; - }; - let cfg = PgConfig::from_url(&url).unwrap(); - { - let mut c = Conn::connect(&cfg).unwrap(); - c.simple_query(r#"DROP TABLE IF EXISTS "mirror_prices""#).unwrap(); - } - - let tables = vec![("Price".to_string(), "mirror_prices".to_string())]; - let (tx, rx) = std::sync::mpsc::sync_channel(QUEUE_CAP); - let handle = spawn(cfg.clone(), rx, tables); - - tx.send(MirrorRec::Upsert { - ty: "Price".into(), id: 1, row: row(json!({"amount": 100})), - }).unwrap(); - tx.send(MirrorRec::Txn(vec![ - MirrorRec::Upsert { ty: "Price".into(), id: 3, row: row(json!({"amount": 300})) }, - MirrorRec::Upsert { ty: "Price".into(), id: 1, row: row(json!({"amount": 150})) }, - ])).unwrap(); - tx.send(MirrorRec::Delete { ty: "Price".into(), id: 3 }).unwrap(); - drop(tx); // close channel → thread drains and exits - handle.join().unwrap(); - - let mut c = Conn::connect(&cfg).unwrap(); - let r = c.simple_query( - r#"SELECT id, row->>'amount' FROM "mirror_prices" ORDER BY id"#).unwrap(); - assert_eq!(r.rows.len(), 1, "id 3 deleted, id 1 remains: {:?}", r.rows); - assert_eq!(r.rows[0][0].as_deref(), Some("1")); - assert_eq!(r.rows[0][1].as_deref(), Some("150"), "txn upsert must have applied"); - c.simple_query(r#"DROP TABLE "mirror_prices""#).unwrap(); - } -} diff --git a/crates/rt/src/parser.rs b/crates/rt/src/parser.rs deleted file mode 100644 index f54d741..0000000 --- a/crates/rt/src/parser.rs +++ /dev/null @@ -1,1240 +0,0 @@ -//! Recursive-descent parser for `.wo` source. -//! -//! Stage 2 scope: -//! * `type Name { ... }` declarations with fields and `service rest` blocks -//! * graceful skip of constructs we don't yet execute: `policy`, `on `, -//! computed field defaults, `fn`, `main`, `##ui`/`##app`/`##sql`/`##doc`/ -//! `##graph`/`##policy`/`##service`/`##logic`/`##logic` blocks. -//! -//! "Skip" means: consume until the matching close brace / next top-level start, -//! so the parser survives and later phases can do nothing. - -use crate::ast::*; -use crate::lexer::tokenize; -use crate::token::{Kind, Token}; - -use anyhow::{bail, Context, Result}; - -pub fn parse(src: &str) -> Result { - let toks = tokenize(src).context("tokenize")?; - Parser::new(toks).parse_schema() -} - -struct Parser { - toks: Vec, - pos: usize, -} - -impl Parser { - fn new(toks: Vec) -> Self { Self { toks, pos: 0 } } - - // --- primitives --- - - fn peek(&self) -> &Kind { &self.toks[self.pos.min(self.toks.len() - 1)].kind } - fn peek_line(&self) -> u32 { self.toks[self.pos.min(self.toks.len() - 1)].line } - - fn advance(&mut self) -> &Token { - let t = &self.toks[self.pos]; - if !matches!(t.kind, Kind::End) { self.pos += 1; } - &self.toks[self.pos.saturating_sub(1)] - } - - fn skip_newlines(&mut self) { - while matches!(self.peek(), Kind::Newline) { self.advance(); } - } - - fn accept(&mut self, want: &Kind) -> bool { - if std::mem::discriminant(self.peek()) == std::mem::discriminant(want) { - self.advance(); - true - } else { false } - } - - fn expect(&mut self, want: &Kind, what: &str) -> Result<&Token> { - if std::mem::discriminant(self.peek()) == std::mem::discriminant(want) { - Ok(self.advance()) - } else { - bail!("line {}: expected {what}, got {}", self.peek_line(), self.peek()) - } - } - - fn expect_ident(&mut self, what: &str) -> Result { - match self.peek().clone() { - Kind::Ident(s) => { self.advance(); Ok(s) } - k => bail!("line {}: expected {what}, got {k}", self.peek_line()), - } - } - - fn at_end(&self) -> bool { matches!(self.peek(), Kind::End) } - - // --- top-level --- - - fn parse_schema(&mut self) -> Result { - let mut sch = Schema::default(); - loop { - self.skip_newlines(); - if self.at_end() { break; } - match self.peek() { - Kind::KwType | Kind::KwClass => - sch.types.push(self.parse_type(TableCfg::default())?), - // Type-level annotation: `@table(...)` configures the - // declaration that follows; unknown names skip silently - // (the field-annotation precedent). - Kind::At => { - if let Some(table) = self.parse_type_annotations()? { - self.skip_newlines(); - if !matches!(self.peek(), Kind::KwType | Kind::KwClass) { - bail!("line {}: expected `type` or `class` after @table, got {}", - self.peek_line(), self.peek()); - } - sch.types.push(self.parse_type(table)?); - } - } - // Skip constructs we don't execute yet. - Kind::HashHash(_) - | Kind::KwFn - | Kind::KwMain - | Kind::KwOn - | Kind::KwPolicy - | Kind::KwTest - | Kind::KwLet // stray `let` at top level (in `main { ... }` probably) - | Kind::Hash(_) // ##sql #table etc. - => self.skip_top_level_chunk()?, - _ => self.skip_top_level_chunk()?, - } - } - Ok(sch) - } - - /// Parse the type-level annotation list ahead of a `type`/`class`. - /// Returns `Some(cfg)` when a recognised `@table` was consumed, `None` - /// when the annotation was unknown and skipped (caller resumes the loop). - fn parse_type_annotations(&mut self) -> Result> { - self.expect(&Kind::At, "'@'")?; - let name = self.expect_ident("annotation name")?; - if name != "table" { - // Unknown type-level annotation: consume an optional (...) block - // and let the schema loop decide what the next token means. - if matches!(self.peek(), Kind::LParen) { - let mut depth = 1i32; - self.advance(); - while depth > 0 && !self.at_end() { - match self.peek() { - Kind::LParen => { depth += 1; self.advance(); } - Kind::RParen => { depth -= 1; self.advance(); } - _ => { self.advance(); } - } - } - } - return Ok(None); - } - - let mut cfg = TableCfg::default(); - if !self.accept(&Kind::LParen) { - return Ok(Some(cfg)); // bare `@table` — legal no-op - } - loop { - self.skip_newlines(); - if self.accept(&Kind::RParen) { break; } - let key = self.expect_ident("@table argument")?; - self.expect(&Kind::Colon, "':'")?; - match key.as_str() { - "name" => { - if cfg.name.is_some() { - bail!("line {}: @table(name: ...) given twice", self.peek_line()); - } - match self.peek().clone() { - Kind::Str(s) => { self.advance(); cfg.name = Some(s); } - other => bail!("line {}: @table name must be a string, got {other}", - self.peek_line()), - } - } - "index" => { - self.expect(&Kind::LBracket, "'['")?; - let mut cols = Vec::new(); - loop { - cols.push(self.expect_ident("index column")?); - if !self.accept(&Kind::Comma) { break; } - } - self.expect(&Kind::RBracket, "']'")?; - if cols.is_empty() { - bail!("line {}: @table index needs at least one column", self.peek_line()); - } - cfg.indexes.push(cols); - } - // `shard_key`/`retention` are reserved for later phases — - // reject loudly rather than silently ignoring (no silent - // passthrough on surface we own). - other => bail!( - "line {}: unknown @table argument `{other}` \ - (supported: name, index)", self.peek_line()), - } - self.skip_newlines(); - if !self.accept(&Kind::Comma) { - self.skip_newlines(); - self.expect(&Kind::RParen, "')' or ','")?; - break; - } - } - Ok(Some(cfg)) - } - - /// Walk forward until we reach the start of the next top-level construct - /// (another `type`, `##…`, or EOF), balancing braces in between. - fn skip_top_level_chunk(&mut self) -> Result<()> { - // Always consume at least one token so we don't loop forever. - let mut depth = 0i32; - let start = self.pos; - loop { - match self.peek() { - Kind::End => break, - Kind::LBrace => { depth += 1; self.advance(); } - Kind::RBrace => { depth -= 1; self.advance(); if depth <= 0 { break; } } - Kind::LBracket => { depth += 1; self.advance(); } - Kind::RBracket => { depth -= 1; self.advance(); } - Kind::LParen => { depth += 1; self.advance(); } - Kind::RParen => { depth -= 1; self.advance(); } - Kind::KwType if depth == 0 && self.pos > start => break, - Kind::KwClass if depth == 0 && self.pos > start => break, - Kind::HashHash(_) if depth == 0 && self.pos > start => break, - _ => { self.advance(); } - } - } - Ok(()) - } - - // --- type declaration --- - - fn parse_type(&mut self, table: TableCfg) -> Result { - // `class` is the behavior-bearing sibling of `type` — identical field - // grammar plus `fn` methods (plan 13a). Storage/REST are class-blind. - let is_class = matches!(self.peek(), Kind::KwClass); - if is_class { - self.advance(); - } else { - self.expect(&Kind::KwType, "`type` or `class`")?; - } - let name = self.expect_ident("type name")?; - - // Link types have a different header: `type Purchase link Customer -> Product { ... }`. - // For Stage 2 we don't bind link-type behaviour, so parse-and-discard the body. - let is_link = matches!(self.peek(), Kind::KwLink); - if is_link { - // consume link A -> B - while !matches!(self.peek(), Kind::LBrace | Kind::End) { - self.advance(); - } - } - - self.expect(&Kind::LBrace, "'{'")?; - - let mut decl = TypeDecl { - name, fields: Vec::new(), services: Vec::new(), is_class, - methods: Vec::new(), table, - }; - loop { - self.skip_newlines(); - match self.peek() { - Kind::RBrace => { self.advance(); break; } - Kind::End => bail!("unexpected end of input inside type body"), - Kind::KwPolicy => self.skip_block_line()?, // policy ... - Kind::KwOn => self.skip_on_block()?, // on update when ... do ... - Kind::KwFn if is_class => { - // 13b: class methods parse for real — signature + DML body. - decl.methods.push(self.parse_method()?); - } - Kind::KwFn => self.skip_block_line()?, // fn inside a plain `type`: - // parse-and-discard (13a); - // brace depth keeps the body's - // `}` from closing the type - Kind::KwService => decl.services.push(self.parse_service()?), - Kind::Ident(_) => { - if is_link { - // Stage 2: absorb link-type bodies without interpreting them. - self.skip_block_line()?; - } else { - // Distinguish a field from a trailing junk line. Fields look like - // `ident : type ...`. Anything else → skip. - if self.looks_like_field() { - decl.fields.push(self.parse_field()?); - } else { - self.skip_block_line()?; - } - } - } - _ => self.skip_block_line()?, - } - } - Ok(decl) - } - - fn looks_like_field(&self) -> bool { - // Lookahead: Ident followed by Colon (possibly after a hyphenated ident). - let mut i = self.pos; - match self.toks.get(i).map(|t| &t.kind) { - Some(Kind::Ident(_)) => {} - _ => return false, - } - i += 1; - matches!(self.toks.get(i).map(|t| &t.kind), Some(Kind::Colon)) - } - - /// Skip tokens until the end of the current logical line (up to Newline or - /// the outer RBrace). Used for policies, triggers, etc., whose full grammar - /// is out of Stage 2 scope. - fn skip_block_line(&mut self) -> Result<()> { - let mut depth = 0i32; - loop { - match self.peek() { - Kind::End => break, - Kind::Newline if depth == 0 => { self.advance(); break; } - Kind::RBrace if depth == 0 => break, // stop before the outer `}` — caller handles it - Kind::LBrace | Kind::LBracket | Kind::LParen => { depth += 1; self.advance(); } - Kind::RBrace | Kind::RBracket | Kind::RParen => { depth -= 1; self.advance(); } - _ => { self.advance(); } - } - } - Ok(()) - } - - /// `on [when ...] do ` possibly spanning many lines. Skip - /// until the next top-level keyword inside the type body. Trigger bodies - /// commonly contain `{ k: v, ... }` object literals and `( ... )` calls, - /// so we track brace depth — RBrace only terminates when we're at the - /// outermost level of the `on` block. - fn skip_on_block(&mut self) -> Result<()> { - self.advance(); // consume `on` - let mut depth = 0i32; - loop { - match self.peek() { - Kind::End => break, - Kind::RBrace if depth == 0 => break, // closes the surrounding type body - Kind::LBrace | Kind::LBracket | Kind::LParen => { - depth += 1; - self.advance(); - } - Kind::RBrace | Kind::RBracket | Kind::RParen => { - depth -= 1; - self.advance(); - } - Kind::Newline if depth == 0 => { - // At depth 0, a newline may end the `on` block if the next - // meaningful token starts a new type-body item. - while matches!(self.peek(), Kind::Newline) { self.advance(); } - if matches!(self.peek(), - Kind::RBrace | Kind::KwPolicy | Kind::KwService - | Kind::KwOn | Kind::KwFn | Kind::End - ) { return Ok(()); } - if matches!(self.peek(), Kind::Ident(_)) && self.looks_like_field() { - return Ok(()); - } - } - _ => { self.advance(); } - } - } - Ok(()) - } - - // --- field parsing --- - - fn parse_field(&mut self) -> Result { - let name = self.expect_ident("field name")?; - self.expect(&Kind::Colon, "':'")?; - let (ty, is_relation) = self.parse_field_ty()?; - - let mut nullable = false; - if matches!(self.peek(), Kind::Question) { - self.advance(); - nullable = true; - } - - let mut unique = false; - let mut default = None; - loop { - match self.peek() { - Kind::At => { - self.advance(); - let name = self.expect_ident("annotation name")?; - // Consume optional (...) argument block without interpreting it. - if matches!(self.peek(), Kind::LParen) { - let mut depth = 1i32; - self.advance(); - while depth > 0 { - match self.peek() { - Kind::End => break, - Kind::LParen => { depth += 1; self.advance(); } - Kind::RParen => { depth -= 1; self.advance(); } - _ => { self.advance(); } - } - } - } - if name == "unique" { unique = true; } - } - Kind::Eq => { - self.advance(); - default = Some(self.parse_default_expr()?); - } - Kind::Newline | Kind::RBrace | Kind::End => break, - _ => { - // Skip any stray tokens until end-of-line — resilient to unhandled - // annotation forms like `@check(between 1 and 5)`. - self.advance(); - } - } - } - - Ok(Field { name, ty, nullable, unique, default, is_relation }) - } - - /// Returns (type, is_relation) — `ref`/`multi`/`backlink` are relations - /// and don't carry a stored scalar column in this type. - fn parse_field_ty(&mut self) -> Result<(FieldTy, bool)> { - match self.peek() { - Kind::KwRef => { - self.advance(); - let target = self.expect_ident("ref target type")?; - Ok((FieldTy::Ref(target), true)) - } - Kind::KwMulti => { - self.advance(); - let target = self.expect_ident("multi target type")?; - // Either `@edge(:TAG)` or `via LinkType` or nothing (= `@edge(:target_upper)`). - match self.peek() { - Kind::At => { - self.advance(); - let ann = self.expect_ident("annotation name")?; - if ann != "edge" { - bail!("line {}: expected @edge, got @{ann}", self.peek_line()); - } - self.expect(&Kind::LParen, "'('")?; - self.expect(&Kind::Colon, "':'")?; - let tag = self.expect_ident("edge tag")?; - self.expect(&Kind::RParen, "')'")?; - Ok((FieldTy::MultiEdge { target, tag: Some(tag) }, true)) - } - Kind::KwVia => { - self.advance(); - let link = self.expect_ident("link type name")?; - Ok((FieldTy::MultiVia { target, link }, true)) - } - _ => Ok((FieldTy::MultiEdge { target, tag: None }, true)) - } - } - Kind::KwBacklink => { - self.advance(); - let target = self.expect_ident("backlink target type")?; - self.expect(&Kind::Dot, "'.'")?; - let field = self.expect_ident("backlink field")?; - Ok((FieldTy::Backlink { target, field }, true)) - } - Kind::LBracket => { - self.advance(); - let (inner, _) = self.parse_field_ty()?; - self.expect(&Kind::RBracket, "']'")?; - Ok((FieldTy::Array(Box::new(inner)), false)) - } - Kind::LBrace => { - self.advance(); - let mut fields = Vec::new(); - loop { - self.skip_newlines(); - if matches!(self.peek(), Kind::RBrace) { self.advance(); break; } - if matches!(self.peek(), Kind::End) { bail!("unexpected end in struct type"); } - fields.push(self.parse_field()?); - // Field end can be comma or just newline - self.accept(&Kind::Comma); - } - Ok((FieldTy::Struct(fields), false)) - } - Kind::Ident(_) => { - let first = self.expect_ident("type name")?; - // Try tagged union: IDENT | IDENT | IDENT - if matches!(self.peek(), Kind::Pipe) { - let mut variants = vec![first]; - while matches!(self.peek(), Kind::Pipe) { - self.advance(); - let next = self.expect_ident("union variant")?; - variants.push(next); - } - Ok((FieldTy::Union(variants), false)) - } else { - Ok((FieldTy::Scalar(first), false)) - } - } - other => bail!("line {}: expected type, got {other}", self.peek_line()), - } - } - - fn parse_default_expr(&mut self) -> Result { - // Fast path: a single literal followed by newline / `}` / `,`. - if let Some(lit) = self.try_standalone_literal() { - return Ok(lit); - } - - // `now` or `now()` — recognise before falling into opaque slurp. - if matches!(self.peek(), Kind::Ident(ref s) if s == "now") { - self.advance(); - if matches!(self.peek(), Kind::LParen) { - self.advance(); - if matches!(self.peek(), Kind::RParen) { self.advance(); } - } - // Only honour if the expression ends here; otherwise fall through - // to the opaque path — `now + 5` etc. becomes opaque. - if matches!(self.peek(), Kind::Newline | Kind::Comma | Kind::RBrace | Kind::End) { - return Ok(DefaultExpr::Now); - } - } - - // Otherwise slurp a balanced expression until newline at depth 0. - // This handles `now()`, `count(...)`, `words(meta.body_md)`, `self.xxx`, - // and anything else we haven't modelled explicitly. Stage 2 treats the - // result as opaque for execution, but the parser survives intact. - let mut buf = String::new(); - let mut depth = 0i32; - loop { - match self.peek() { - Kind::End => break, - Kind::Newline | Kind::Comma if depth == 0 => break, - Kind::RBrace if depth == 0 => break, - Kind::LBrace | Kind::LBracket | Kind::LParen => { - depth += 1; - let t = self.advance(); - buf.push_str(&format!("{} ", t.kind)); - } - Kind::RBrace | Kind::RBracket | Kind::RParen => { - depth -= 1; - let t = self.advance(); - buf.push_str(&format!("{} ", t.kind)); - } - _ => { - let t = self.advance(); - buf.push_str(&format!("{} ", t.kind)); - } - } - } - let trimmed = buf.trim().to_string(); - // Friendly recognition of common forms. - if trimmed.starts_with("now ") || trimmed == "now" { - return Ok(DefaultExpr::Now); - } - Ok(DefaultExpr::Opaque(trimmed)) - } - - /// If the next token is a self-contained literal (not followed by more - /// expression tokens), consume and return it. Otherwise leave the cursor - /// alone and return None so the opaque path takes over. - fn try_standalone_literal(&mut self) -> Option { - // Peek one ahead to decide if the literal is the whole expression. - let follower_ends_expr = match self.toks.get(self.pos + 1).map(|t| &t.kind) { - Some(Kind::Newline) | Some(Kind::End) | Some(Kind::RBrace) | Some(Kind::Comma) => true, - _ => false, - }; - if !follower_ends_expr { return None; } - match self.peek().clone() { - Kind::Str(s) => { self.advance(); Some(DefaultExpr::Str(s)) } - Kind::Int(n) => { self.advance(); Some(DefaultExpr::Int(n)) } - Kind::KwTrue => { self.advance(); Some(DefaultExpr::Bool(true)) } - Kind::KwFalse => { self.advance(); Some(DefaultExpr::Bool(false)) } - Kind::KwNull => { self.advance(); Some(DefaultExpr::Null) } - Kind::Ident(s) => { self.advance(); Some(DefaultExpr::Enum(s)) } - _ => None, - } - } - - // --- service --- - - fn parse_service(&mut self) -> Result { - self.expect(&Kind::KwService, "`service`")?; - let kind = match self.peek() { - Kind::KwRest => { self.advance(); ServiceKind::Rest } - Kind::KwGraphql => { self.advance(); ServiceKind::Graphql } - Kind::KwNative => { self.advance(); ServiceKind::Native } - Kind::Ident(s) if s == "rest" => { self.advance(); ServiceKind::Rest } - Kind::Ident(s) if s == "graphql" => { self.advance(); ServiceKind::Graphql } - Kind::Ident(s) if s == "native" => { self.advance(); ServiceKind::Native } - other => bail!("line {}: expected `rest`/`graphql`/`native`, got {other}", self.peek_line()), - }; - - let path = match self.peek().clone() { - Kind::Str(s) => { self.advance(); s } - other => bail!("line {}: expected service path string, got {other}", self.peek_line()), - }; - - self.skip_newlines(); - self.expect(&Kind::KwExpose, "`expose`")?; - let mut expose = Vec::new(); - loop { - let name = self.expect_ident("operation name")?; - expose.push(Operation::from_ident(&name)); - if !self.accept(&Kind::Comma) { break; } - } - Ok(ServiceDecl { kind, path, expose }) - } - - // --- class methods (plan 13b) --- - - /// `fn name([p: T, ...]) [-> Ret] [in txn [snapshot|serializable]] { body }` - fn parse_method(&mut self) -> Result { - self.expect(&Kind::KwFn, "`fn`")?; - let name = self.expect_ident("method name")?; - - self.expect(&Kind::LParen, "'('")?; - let mut params = Vec::new(); - self.skip_newlines(); - while !matches!(self.peek(), Kind::RParen) { - let pname = self.expect_ident("parameter name")?; - self.expect(&Kind::Colon, "':'")?; - let pty = self.expect_ident("parameter type")?; - params.push((pname, pty)); - self.skip_newlines(); - if !self.accept(&Kind::Comma) { break; } - self.skip_newlines(); - } - self.expect(&Kind::RParen, "')'")?; - - let mut ret = None; - if self.accept(&Kind::Arrow) { - // `-> Money` or `-> [Price]` — diagnostic-only in Stage 2. - if self.accept(&Kind::LBracket) { - let inner = self.expect_ident("return type")?; - self.expect(&Kind::RBracket, "']'")?; - ret = Some(format!("[{inner}]")); - } else { - ret = Some(self.expect_ident("return type")?); - } - } - - let mut txn = TxnMode::None; - if self.accept(&Kind::KwIn) { - self.expect(&Kind::KwTxn, "`txn`")?; - txn = match self.peek() { - Kind::KwSnapshot => { self.advance(); TxnMode::Snapshot } - Kind::KwSerializable => { self.advance(); TxnMode::Serializable } - _ => TxnMode::Txn, - }; - } - - self.skip_newlines(); - self.expect(&Kind::LBrace, "'{' to open method body")?; - let body = self.parse_stmt_block()?; - Ok(MethodDecl { name, params, ret, txn, body }) - } - - /// Statements until the matching `}` (consumed). - fn parse_stmt_block(&mut self) -> Result> { - let mut stmts = Vec::new(); - loop { - self.skip_newlines(); - while self.accept(&Kind::Semicolon) { self.skip_newlines(); } - match self.peek() { - Kind::RBrace => { self.advance(); return Ok(stmts); } - Kind::End => bail!("unexpected end of input inside method body"), - _ => stmts.push(self.parse_stmt()?), - } - } - } - - fn parse_stmt(&mut self) -> Result { - let stmt = match self.peek() { - Kind::KwLet => { - self.advance(); - let name = self.expect_ident("binding name")?; - self.expect(&Kind::Eq, "'='")?; - let expr = self.parse_expr()?; - Stmt::Let { name, expr } - } - // Schema-layer DML `insert` is lowercase and deliberately NOT a - // lexer keyword (only SQL-layer `INSERT` is) — same rule that - // keeps `subscribe`/`me`/`self` usable as plain identifiers. - Kind::Ident(s) if s == "insert" => { - self.advance(); - let ty = self.expect_ident("type name after `insert`")?; - self.expect(&Kind::LBrace, "'{'")?; - let mut fields = Vec::new(); - loop { - self.skip_newlines(); - if self.accept(&Kind::RBrace) { break; } - let fname = self.expect_ident("field name")?; - self.expect(&Kind::Colon, "':'")?; - let expr = self.parse_expr()?; - fields.push((fname, expr)); - self.skip_newlines(); - self.accept(&Kind::Comma); - } - Stmt::Insert { ty, fields } - } - Kind::KwReturn => { - self.advance(); - let expr = if matches!(self.peek(), - Kind::Semicolon | Kind::Newline | Kind::RBrace | Kind::End) - { None } else { Some(self.parse_expr()?) }; - Stmt::Return { expr } - } - Kind::KwAssert => { - self.advance(); - let cond = self.parse_expr()?; - let mut msg = None; - if self.accept(&Kind::KwOtherwise) { - self.expect(&Kind::KwAbort, "`abort`")?; - if let Kind::Str(s) = self.peek().clone() { - self.advance(); - msg = Some(s); - } - } - Stmt::Assert { cond, msg } - } - Kind::KwIf => { - self.advance(); - let cond = self.parse_expr()?; - self.skip_newlines(); - self.expect(&Kind::LBrace, "'{' after if condition")?; - let then = self.parse_stmt_block()?; - let mut otherwise = Vec::new(); - // `else` may sit on the next line. - let mark = self.pos; - self.skip_newlines(); - if self.accept(&Kind::KwElse) { - self.skip_newlines(); - if matches!(self.peek(), Kind::KwIf) { - otherwise.push(self.parse_stmt()?); // else-if chain - } else { - self.expect(&Kind::LBrace, "'{' after else")?; - otherwise = self.parse_stmt_block()?; - } - } else { - self.pos = mark; // no else — restore consumed newlines - } - return Ok(Stmt::If { cond, then, otherwise }); - } - other => bail!( - "line {}: unsupported statement in method body: {other} \ - (13b executes `let`/`insert`/`return`/`assert`/`if`)", - self.peek_line() - ), - }; - // Statement terminator: `;`, newline, or the closing `}`. - match self.peek() { - Kind::Semicolon | Kind::Newline => { self.advance(); } - Kind::RBrace | Kind::End => {} - other => bail!("line {}: expected end of statement, got {other}", self.peek_line()), - } - Ok(stmt) - } - - // --- expressions (precedence climbing: or < and < cmp < add < mul < unary < postfix) --- - - fn parse_expr(&mut self) -> Result { - self.parse_or() - } - - fn parse_or(&mut self) -> Result { - let mut lhs = self.parse_and()?; - while self.accept(&Kind::KwOr) { - let rhs = self.parse_and()?; - lhs = Expr::Binary(BinOp::Or, Box::new(lhs), Box::new(rhs)); - } - Ok(lhs) - } - - fn parse_and(&mut self) -> Result { - let mut lhs = self.parse_cmp()?; - while self.accept(&Kind::KwAnd) { - let rhs = self.parse_cmp()?; - lhs = Expr::Binary(BinOp::And, Box::new(lhs), Box::new(rhs)); - } - Ok(lhs) - } - - fn parse_cmp(&mut self) -> Result { - let lhs = self.parse_add()?; - let op = match self.peek() { - Kind::EqEq => BinOp::Eq, - Kind::NotEq => BinOp::Ne, - Kind::Lt => BinOp::Lt, - Kind::LtEq => BinOp::Le, - Kind::Gt => BinOp::Gt, - Kind::GtEq => BinOp::Ge, - _ => return Ok(lhs), - }; - self.advance(); - let rhs = self.parse_add()?; - Ok(Expr::Binary(op, Box::new(lhs), Box::new(rhs))) - } - - fn parse_add(&mut self) -> Result { - let mut lhs = self.parse_mul()?; - loop { - let op = match self.peek() { - Kind::Plus => BinOp::Add, - Kind::Dash => BinOp::Sub, - _ => return Ok(lhs), - }; - self.advance(); - let rhs = self.parse_mul()?; - lhs = Expr::Binary(op, Box::new(lhs), Box::new(rhs)); - } - } - - fn parse_mul(&mut self) -> Result { - let mut lhs = self.parse_unary()?; - loop { - let op = match self.peek() { - Kind::Star => BinOp::Mul, - Kind::Slash => BinOp::Div, - Kind::Percent => BinOp::Mod, - _ => return Ok(lhs), - }; - self.advance(); - let rhs = self.parse_unary()?; - lhs = Expr::Binary(op, Box::new(lhs), Box::new(rhs)); - } - } - - fn parse_unary(&mut self) -> Result { - match self.peek() { - Kind::Dash => { self.advance(); Ok(Expr::Unary(UnOp::Neg, Box::new(self.parse_unary()?))) } - Kind::KwNot => { self.advance(); Ok(Expr::Unary(UnOp::Not, Box::new(self.parse_unary()?))) } - _ => self.parse_postfix(), - } - } - - /// Primary followed by `.field` chains. - fn parse_postfix(&mut self) -> Result { - let mut e = self.parse_primary()?; - while self.accept(&Kind::Dot) { - let field = self.expect_ident("field name after '.'")?; - e = Expr::Field(Box::new(e), field); - } - Ok(e) - } - - fn parse_primary(&mut self) -> Result { - match self.peek().clone() { - Kind::Int(n) => { self.advance(); Ok(Expr::Int(n)) } - Kind::Str(s) => { self.advance(); Ok(Expr::Str(s)) } - Kind::KwTrue => { self.advance(); Ok(Expr::Bool(true)) } - Kind::KwFalse => { self.advance(); Ok(Expr::Bool(false)) } - Kind::KwNull => { self.advance(); Ok(Expr::Null) } - Kind::LParen => { - self.advance(); - let e = self.parse_expr()?; - self.expect(&Kind::RParen, "')'")?; - Ok(e) - } - Kind::Ident(name) => { - self.advance(); - // `select Type{ ... }` — schema-layer select expression. - // Lowercase `select` is an ident (only SQL `SELECT` is a - // keyword); the two-token shape Ident + LBrace disambiguates - // it from a variable named `select`. - if name == "select" - && matches!(self.peek(), Kind::Ident(_)) - && matches!(self.toks.get(self.pos + 1).map(|t| &t.kind), Some(Kind::LBrace)) - { - return self.parse_select_expr(); - } - // `name(args)` — call form. - if self.accept(&Kind::LParen) { - let mut args = Vec::new(); - if !matches!(self.peek(), Kind::RParen) { - loop { - args.push(self.parse_expr()?); - if !self.accept(&Kind::Comma) { break; } - } - } - self.expect(&Kind::RParen, "')'")?; - Ok(Expr::Call(name, args)) - } else { - Ok(Expr::Ident(name)) - } - } - other => bail!("line {}: expected expression, got {other}", self.peek_line()), - } - } - - /// `select Type{ entries }` — per § Brace Disambiguation: an entry with a - /// comparison operator is a predicate; a bare identifier is a projection. - /// (`select` and the type name are already consumed up to the ident.) - fn parse_select_expr(&mut self) -> Result { - let ty = self.expect_ident("type name after `select`")?; - self.expect(&Kind::LBrace, "'{'")?; - let mut predicates = Vec::new(); - let mut projection = Vec::new(); - loop { - self.skip_newlines(); - if self.accept(&Kind::RBrace) { break; } - let field = self.expect_ident("field name")?; - let op = match self.peek() { - Kind::EqEq => Some(BinOp::Eq), - Kind::NotEq => Some(BinOp::Ne), - Kind::Lt => Some(BinOp::Lt), - Kind::LtEq => Some(BinOp::Le), - Kind::Gt => Some(BinOp::Gt), - Kind::GtEq => Some(BinOp::Ge), - _ => None, - }; - match op { - Some(op) => { - self.advance(); - // RHS is an additive expression — comparisons don't chain. - let rhs = self.parse_add()?; - predicates.push((field, op, Box::new(rhs))); - } - None => projection.push(field), - } - self.skip_newlines(); - self.accept(&Kind::Comma); - } - Ok(Expr::Select { ty, predicates, projection }) - } -} - -#[cfg(test)] -mod tests { - use super::*; - - #[test] - fn parses_simple_type() { - let src = r#" -type Article { - id: Id - title: Text - service rest "/api/articles" expose list, get, create -} -"#; - let sch = parse(src).unwrap(); - assert_eq!(sch.types.len(), 1); - let t = &sch.types[0]; - assert_eq!(t.name, "Article"); - assert_eq!(t.fields.len(), 2); - assert_eq!(t.fields[0].name, "id"); - assert_eq!(t.fields[1].name, "title"); - assert_eq!(t.services.len(), 1); - assert_eq!(t.services[0].path, "/api/articles"); - assert_eq!(t.services[0].expose, vec![ - Operation::List, Operation::Get, Operation::Create, - ]); - } - - #[test] - fn parses_nullable_and_unique_and_default() { - let src = r#" -type User { - email: Email @unique - name: Text - avatar: Url? - joined: Timestamp = now() - admin: Bool = false -} -"#; - let sch = parse(src).unwrap(); - let t = &sch.types[0]; - assert!(t.fields[0].unique); - assert!(t.fields[2].nullable); - assert!(matches!(t.fields[3].default, Some(DefaultExpr::Now))); - assert!(matches!(t.fields[4].default, Some(DefaultExpr::Bool(false)))); - } - - #[test] - fn parses_relations() { - let src = r#" -type Article { - author: ref User - tags: multi Tag @edge(:TAGGED_AS) - related: multi Article @edge(:RELATED_TO) -} -"#; - let sch = parse(src).unwrap(); - let t = &sch.types[0]; - assert!(matches!(t.fields[0].ty, FieldTy::Ref(ref n) if n == "User")); - assert!(matches!(t.fields[1].ty, FieldTy::MultiEdge { ref target, .. } if target == "Tag")); - } - - #[test] - fn skips_policy_and_on_trigger() { - let src = r#" -type Article { - id: Id - title: Text - - policy read anyone - policy write for role Admin - - on update when old.published == false and new.published == true - do set self.published_at = now() - - service rest "/api/articles" expose list, get -} -"#; - let sch = parse(src).unwrap(); - let t = &sch.types[0]; - assert_eq!(t.name, "Article"); - assert_eq!(t.fields.len(), 2); - assert_eq!(t.services.len(), 1); - } - - #[test] - fn parses_inline_struct() { - let src = r#" -type Product { - meta: { - title: Text - tags: [Text] - } -} -"#; - let sch = parse(src).unwrap(); - let t = &sch.types[0]; - if let FieldTy::Struct(fs) = &t.fields[0].ty { - assert_eq!(fs.len(), 2); - assert!(matches!(fs[1].ty, FieldTy::Array(_))); - } else { panic!("expected struct") } - } - - #[test] - fn parses_union() { - let src = r#" -type Order { status: Pending | Paid | Shipped } -"#; - let sch = parse(src).unwrap(); - if let FieldTy::Union(v) = &sch.types[0].fields[0].ty { - assert_eq!(v, &vec!["Pending".to_string(), "Paid".into(), "Shipped".into()]); - } else { panic!("expected union") } - } - - #[test] - fn article_debug() { - // formerly read docs/examples/blog/types/article.wo; the blog sample - // moved to the cleanup branch, so the fixture lives inline now — - // same shape: scalars, embedded doc, edges, backlink, computed, - // policies, triggers, service - let src = r#" -type Article { - id: Id - slug: Slug @unique - title: Text - author: ref Author - published: Bool = false - published_at: Timestamp? - created_at: Timestamp = now() - updated_at: Timestamp = now() - - meta: { - excerpt: Text - body_md: Markdown - hero_image: Url? - reading_min: Int? - } - - tags: multi Tag @edge(:TAGGED_AS) - related: multi Article @edge(:RELATED_TO) - prerequisites: multi Article @edge(:PREREQUISITE) - comments: backlink Comment.article - word_count: Int = words(meta.body_md) - - policy read anyone when published == true - policy read for role Admin - policy read for role Author when author == $session.user - policy write for role Admin - policy write for role Author when author == $session.user - policy delete for role Admin - - on update - when old.published == false and new.published == true - do set self.published_at = now() - do emit "article.published"(self) - do enqueue "send-subscriber-emails" with { article_id: self.id } - - on update - do set self.updated_at = now() - - service rest "/api/articles" - expose list, get, create, update, delete, subscribe -} -"#; - let sch = parse(src).unwrap(); - eprintln!("types: {}", sch.types.len()); - for t in &sch.types { - eprintln!(" type {} — {} fields, {} services", t.name, t.fields.len(), t.services.len()); - for f in &t.fields { eprintln!(" field: {}", f.name); } - for svc in &t.services { eprintln!(" service {:?} {} expose {:?}", svc.kind, svc.path, svc.expose); } - } - assert!(sch.types.iter().any(|t| t.name == "Article" && !t.services.is_empty()), - "Article should have a service"); - } - - #[test] - fn parses_class_with_methods() { - let src = r#" -class Product { - id: Id - sku: SKU @unique - name: Text - prices: multi Price - - fn current_price() -> Money in txn { - return latest(self.prices).amount; - } - - fn set_price(amount: Money) in txn { - insert Price { product: self.id, amount: amount }; - } - - service rest "/api/products" - expose list, get, create, update, delete, subscribe -} -"#; - let sch = parse(src).unwrap(); - assert_eq!(sch.types.len(), 1); - let t = &sch.types[0]; - assert!(t.is_class); - assert_eq!(t.name, "Product"); - assert_eq!(t.fields.len(), 4); - assert!(matches!(t.fields[3].ty, FieldTy::MultiEdge { ref target, .. } if target == "Price")); - assert_eq!(t.services.len(), 1); - assert_eq!(t.services[0].path, "/api/products"); - assert_eq!(t.services[0].expose.len(), 6); - - // 13b: methods parse into real AST. - assert_eq!(t.methods.len(), 2); - let cp = &t.methods[0]; - assert_eq!(cp.name, "current_price"); - assert!(cp.params.is_empty()); - assert_eq!(cp.ret.as_deref(), Some("Money")); - assert_eq!(cp.txn, TxnMode::Txn); - assert_eq!(cp.body.len(), 1); - assert!(matches!(cp.body[0], Stmt::Return { expr: Some(_) })); - - let sp = &t.methods[1]; - assert_eq!(sp.name, "set_price"); - assert_eq!(sp.params, vec![("amount".to_string(), "Money".to_string())]); - assert_eq!(sp.txn, TxnMode::Txn); - assert!(matches!(&sp.body[0], - Stmt::Insert { ty, fields } if ty == "Price" && fields.len() == 2)); - } - - #[test] - fn parses_table_annotation() { - let src = r#" -@table(name: "prices", index: [product, at], index: [sku]) -class Price { - id: Id - product: ref Product - amount: Money - at: Timestamp = now() - sku: SKU -} - -@table -type Note { id: Id } - -type Plain { id: Id } -"#; - let sch = parse(src).unwrap(); - assert_eq!(sch.types.len(), 3); - let p = &sch.types[0]; - assert_eq!(p.table.name.as_deref(), Some("prices")); - assert_eq!(p.table.indexes, vec![ - vec!["product".to_string(), "at".to_string()], - vec!["sku".to_string()], - ]); - // bare @table = legal no-op config - let n = &sch.types[1]; - assert!(n.table.name.is_none() && n.table.indexes.is_empty()); - assert!(sch.types[2].table.indexes.is_empty()); - } - - #[test] - fn table_annotation_error_cases() { - // Reserved-for-later keys error loudly — no silent passthrough. - let err = parse("@table(shard_key: sku)\ntype T { id: Id }").unwrap_err().to_string(); - assert!(err.contains("unknown @table argument `shard_key`"), "{err}"); - - // @table must be followed by a type/class. - assert!(parse("@table(name: \"x\")\nfn stray() {}").is_err()); - - // Unknown annotation NAMES skip silently (field-annotation precedent). - let sch = parse("@experimental(anything, at: all)\ntype T { id: Id }").unwrap(); - assert_eq!(sch.types.len(), 1); - assert_eq!(sch.types[0].name, "T"); - } - - #[test] - fn parses_select_expression() { - let src = r#" -class Product { - id: Id - fn history() -> [Price] in txn { - return select Price{ product == self.id, amount, at }; - } -} -"#; - let m = &parse(src).unwrap().types[0].methods[0]; - let Stmt::Return { expr: Some(Expr::Select { ty, predicates, projection }) } = &m.body[0] - else { panic!("expected return select, got {:?}", m.body[0]) }; - assert_eq!(ty, "Price"); - assert_eq!(predicates.len(), 1); - assert_eq!(predicates[0].0, "product"); - assert!(matches!(predicates[0].1, BinOp::Eq)); - assert!(matches!(&*predicates[0].2, Expr::Field(b, f) if f == "id" - && matches!(&**b, Expr::Ident(s) if s == "self"))); - assert_eq!(projection, &vec!["amount".to_string(), "at".to_string()]); - } - - #[test] - fn method_body_expression_ast() { - let src = r#" -class Price { - id: Id - amount: Money - - fn discounted(pct: Int) -> Money { - return self.amount * (100 - pct) / 100; - } -} -"#; - let sch = parse(src).unwrap(); - let m = &sch.types[0].methods[0]; - assert_eq!(m.txn, TxnMode::None); - let Stmt::Return { expr: Some(e) } = &m.body[0] else { panic!("expected return") }; - // ((self.amount * (100 - pct)) / 100) — mul level is left-associative. - let Expr::Binary(BinOp::Div, lhs, rhs) = e else { panic!("expected /: {e:?}") }; - assert!(matches!(**rhs, Expr::Int(100))); - let Expr::Binary(BinOp::Mul, base, paren) = &**lhs else { panic!("expected *") }; - assert!(matches!(&**base, Expr::Field(b, f) if f == "amount" - && matches!(&**b, Expr::Ident(s) if s == "self"))); - assert!(matches!(&**paren, Expr::Binary(BinOp::Sub, _, _))); - } - - #[test] - fn class_method_braces_do_not_truncate_body() { - // The method body's `}` and nested `{ ... }` literals must not be - // mistaken for the class's closing brace — fields AFTER the methods - // must still parse, and a following declaration must be seen. - let src = r#" -class Price { - id: Id - - fn discounted(pct: Int) -> Money { - if pct > 0 { - return self.amount * (100 - pct) / 100; - } - return self.amount; - } - - amount: Money - currency: Text = "EUR" -} - -type Audit { id: Id } -"#; - let sch = parse(src).unwrap(); - assert_eq!(sch.types.len(), 2); - let p = &sch.types[0]; - assert!(p.is_class); - assert_eq!(p.fields.len(), 3, "fields after the method must parse"); - assert_eq!(p.fields[1].name, "amount"); - assert!(!sch.types[1].is_class); - assert_eq!(sch.types[1].name, "Audit"); - } -} diff --git a/crates/rt/src/pg.rs b/crates/rt/src/pg.rs deleted file mode 100644 index 4b094ea..0000000 --- a/crates/rt/src/pg.rs +++ /dev/null @@ -1,460 +0,0 @@ -//! Hand-rolled PostgreSQL wire-protocol client — plan 16a -//! (`docs/plan/16-postgres-mirror.md`). -//! -//! The mirror's outbound half: protocol v3 over a blocking -//! `std::net::TcpStream`, zero external crates — the same doctrine as the -//! hand-rolled HTTP layer and the CRC32 in `wal.rs`. Scope is exactly what -//! the backup mirror needs: -//! -//! * startup + auth: `trust`, `password` (cleartext), `md5` -//! (SCRAM-SHA-256 is plan 16f) -//! * the **simple query protocol** only (`Query` → `RowDescription` / -//! `DataRow` / `CommandComplete` / `ErrorResponse` / `ReadyForQuery`) — -//! no extended protocol, no prepared statements, no TLS -//! * literal/identifier escaping for SQL the mirror generates -//! -//! Protocol reference: PostgreSQL docs “Frontend/Backend Protocol” and -//! `reference/postgresql/src/include/libpq/` (research symlink). -//! -//! Blocking I/O is deliberate: the only caller is the dedicated `wo-pg` -//! mirror thread (plan 16b) — never a shard worker. - -use std::fmt; -use std::io::{self, Read, Write}; -use std::net::TcpStream; -use std::time::Duration; - -/// Parsed `postgres://user[:password]@host[:port]/database` URL. -/// (No percent-decoding — keep credentials URL-safe.) -#[derive(Debug, Clone)] -pub struct PgConfig { - pub user: String, - pub password: Option, - pub host: String, - pub port: u16, - pub database: String, -} - -impl PgConfig { - pub fn from_url(url: &str) -> Result { - let rest = url.strip_prefix("postgres://") - .or_else(|| url.strip_prefix("postgresql://")) - .ok_or_else(|| format!("WO_PG url must start with postgres:// — got {url}"))?; - let (userinfo, hostpart) = rest.split_once('@') - .ok_or_else(|| "WO_PG url needs user@host".to_string())?; - let (user, password) = match userinfo.split_once(':') { - Some((u, p)) => (u.to_string(), Some(p.to_string())), - None => (userinfo.to_string(), None), - }; - let (hostport, database) = hostpart.split_once('/') - .ok_or_else(|| "WO_PG url needs /database".to_string())?; - let (host, port) = match hostport.split_once(':') { - Some((h, p)) => (h.to_string(), - p.parse::().map_err(|_| format!("bad port `{p}`"))?), - None => (hostport.to_string(), 5432), - }; - if user.is_empty() || host.is_empty() || database.is_empty() { - return Err(format!("incomplete WO_PG url: {url}")); - } - Ok(PgConfig { user, password, host, port, database: database.to_string() }) - } -} - -/// A backend `ErrorResponse` (or client-side failure talking to it). -#[derive(Debug)] -pub struct PgError { - pub severity: String, - pub code: String, - pub message: String, -} - -impl fmt::Display for PgError { - fn fmt(&self, f: &mut fmt::Formatter<'_>) -> fmt::Result { - write!(f, "{} {}: {}", self.severity, self.code, self.message) - } -} - -impl PgError { - fn client(msg: impl Into) -> PgError { - PgError { severity: "CLIENT".into(), code: "XX000".into(), message: msg.into() } - } -} - -impl From for PgError { - fn from(e: io::Error) -> PgError { PgError::client(format!("io: {e}")) } -} - -/// Result of one simple query (possibly multi-statement). -#[derive(Debug, Default)] -pub struct QueryResult { - pub columns: Vec, - /// Text-format values, `None` = SQL NULL. Rows of the LAST result set. - pub rows: Vec>>, - /// One CommandComplete tag per statement, e.g. `INSERT 0 1`. - pub tags: Vec, -} - -pub struct Conn { - stream: TcpStream, -} - -impl Conn { - /// Connect and authenticate. Blocking, with a connect timeout. - pub fn connect(cfg: &PgConfig) -> Result { - let addr = format!("{}:{}", cfg.host, cfg.port); - let sockaddr = addr.parse() - .map_err(|_| { - // Not a literal ip:port — resolve via ToSocketAddrs. - PgError::client("resolve") - }); - let stream = match sockaddr { - Ok(sa) => TcpStream::connect_timeout(&sa, Duration::from_secs(5))?, - Err(_) => TcpStream::connect(&addr)?, // DNS path - }; - stream.set_nodelay(true).ok(); - stream.set_read_timeout(Some(Duration::from_secs(30)))?; - stream.set_write_timeout(Some(Duration::from_secs(30)))?; - let mut conn = Conn { stream }; - conn.startup(cfg)?; - Ok(conn) - } - - fn startup(&mut self, cfg: &PgConfig) -> Result<(), PgError> { - // StartupMessage: no type byte — i32 len | i32 196608 | k\0v\0 ... \0 - let mut body = Vec::new(); - body.extend_from_slice(&196_608i32.to_be_bytes()); // protocol 3.0 - for (k, v) in [("user", cfg.user.as_str()), - ("database", cfg.database.as_str()), - ("client_encoding", "UTF8"), - ("application_name", "wo-pg-mirror")] { - body.extend_from_slice(k.as_bytes()); body.push(0); - body.extend_from_slice(v.as_bytes()); body.push(0); - } - body.push(0); - let mut msg = Vec::with_capacity(body.len() + 4); - msg.extend_from_slice(&((body.len() as i32 + 4).to_be_bytes())); - msg.extend_from_slice(&body); - self.stream.write_all(&msg)?; - - // Authentication exchange, then drain to ReadyForQuery. - loop { - let (kind, payload) = self.read_message()?; - match kind { - b'R' => { - let auth = be_i32(&payload, 0)?; - match auth { - 0 => {} // AuthenticationOk - 3 => { // CleartextPassword - let pw = cfg.password.clone().ok_or_else(|| - PgError::client("server wants a password; none in WO_PG url"))?; - self.send_password(&pw)?; - } - 5 => { // MD5Password + 4B salt - let pw = cfg.password.clone().ok_or_else(|| - PgError::client("server wants md5 auth; no password in WO_PG url"))?; - let salt = payload.get(4..8).ok_or_else(|| - PgError::client("short md5 salt"))?; - // "md5" + md5hex(md5hex(password + user) + salt) - let inner = md5_hex(format!("{pw}{}", cfg.user).as_bytes()); - let mut outer_in = inner.into_bytes(); - outer_in.extend_from_slice(salt); - let digest = format!("md5{}", md5_hex(&outer_in)); - self.send_password(&digest)?; - } - 10 => return Err(PgError::client( - "server requires SCRAM-SHA-256 — not supported until plan 16f; \ - configure md5/password/trust auth for the mirror role")), - n => return Err(PgError::client(format!("unsupported auth type {n}"))), - } - } - b'S' | b'K' | b'N' => {} // ParameterStatus / BackendKeyData / Notice - b'Z' => return Ok(()), // ReadyForQuery - b'E' => return Err(parse_error(&payload)), - other => return Err(PgError::client(format!( - "unexpected message '{}' during startup", other as char))), - } - } - } - - fn send_password(&mut self, pw: &str) -> Result<(), PgError> { - let mut msg = Vec::with_capacity(pw.len() + 6); - msg.push(b'p'); - msg.extend_from_slice(&((pw.len() as i32 + 5).to_be_bytes())); - msg.extend_from_slice(pw.as_bytes()); - msg.push(0); - self.stream.write_all(&msg)?; - Ok(()) - } - - /// Run one simple query (may contain multiple `;`-separated statements — - /// the backend wraps them in an implicit transaction). Returns the last - /// result set + all command tags; a backend error is returned AFTER the - /// stream is drained to ReadyForQuery, so the connection stays usable. - pub fn simple_query(&mut self, sql: &str) -> Result { - let mut msg = Vec::with_capacity(sql.len() + 6); - msg.push(b'Q'); - msg.extend_from_slice(&((sql.len() as i32 + 5).to_be_bytes())); - msg.extend_from_slice(sql.as_bytes()); - msg.push(0); - self.stream.write_all(&msg)?; - - let mut out = QueryResult::default(); - let mut err: Option = None; - loop { - let (kind, payload) = self.read_message()?; - match kind { - b'T' => { // RowDescription - out.columns.clear(); - let n = be_i16(&payload, 0)? as usize; - let mut off = 2; - for _ in 0..n { - let name = read_cstr(&payload, off)?; - off += name.len() + 1 + 18; // 4+2+4+2+4+2 fixed fields - out.columns.push(name); - } - out.rows.clear(); // keep the last result set - } - b'D' => { // DataRow - let n = be_i16(&payload, 0)? as usize; - let mut off = 2; - let mut row = Vec::with_capacity(n); - for _ in 0..n { - let len = be_i32(&payload, off)?; - off += 4; - if len < 0 { row.push(None); continue; } - let len = len as usize; - let bytes = payload.get(off..off + len) - .ok_or_else(|| PgError::client("short DataRow"))?; - row.push(Some(String::from_utf8_lossy(bytes).into_owned())); - off += len; - } - out.rows.push(row); - } - b'C' => out.tags.push(read_cstr(&payload, 0)?), // CommandComplete - b'E' => { if err.is_none() { err = Some(parse_error(&payload)); } } - b'Z' => break, // ReadyForQuery - b'N' | b'S' | b'I' | b'G' | b'H' | b'W' => {} // notices etc. - other => return Err(PgError::client(format!( - "unexpected message '{}' in query response", other as char))), - } - } - match err { - Some(e) => Err(e), - None => Ok(out), - } - } - - /// Read one backend message: 1-byte type + i32 length (incl. itself). - fn read_message(&mut self) -> Result<(u8, Vec), PgError> { - let mut head = [0u8; 5]; - self.stream.read_exact(&mut head)?; - let len = i32::from_be_bytes([head[1], head[2], head[3], head[4]]); - if !(4..=64 * 1024 * 1024).contains(&len) { - return Err(PgError::client(format!("bad message length {len}"))); - } - let mut payload = vec![0u8; len as usize - 4]; - self.stream.read_exact(&mut payload)?; - Ok((head[0], payload)) - } -} - -// --- wire helpers --- - -fn be_i32(b: &[u8], off: usize) -> Result { - b.get(off..off + 4) - .map(|s| i32::from_be_bytes(s.try_into().unwrap())) - .ok_or_else(|| PgError::client("short message")) -} - -fn be_i16(b: &[u8], off: usize) -> Result { - b.get(off..off + 2) - .map(|s| i16::from_be_bytes(s.try_into().unwrap())) - .ok_or_else(|| PgError::client("short message")) -} - -fn read_cstr(b: &[u8], off: usize) -> Result { - let end = b[off..].iter().position(|&c| c == 0) - .ok_or_else(|| PgError::client("unterminated string"))?; - Ok(String::from_utf8_lossy(&b[off..off + end]).into_owned()) -} - -/// ErrorResponse / NoticeResponse: (field-code byte, cstring) pairs. -fn parse_error(payload: &[u8]) -> PgError { - let mut e = PgError { severity: "ERROR".into(), code: String::new(), message: String::new() }; - let mut off = 0; - while off < payload.len() && payload[off] != 0 { - let code = payload[off]; - let Ok(val) = read_cstr(payload, off + 1) else { break }; - off += 1 + val.len() + 1; - match code { - b'S' => e.severity = val, - b'C' => e.code = val, - b'M' => e.message = val, - _ => {} - } - } - e -} - -// --- SQL text helpers (the mirror builds statements as text) --- - -/// `'…'` literal with single quotes doubled. Standard-conforming strings -/// (the server default) treat backslashes literally, so quotes are the only -/// metacharacter. -pub fn escape_literal(s: &str) -> String { - let mut out = String::with_capacity(s.len() + 2); - out.push('\''); - for c in s.chars() { - if c == '\'' { out.push('\''); } - out.push(c); - } - out.push('\''); - out -} - -/// `"…"` identifier with double quotes doubled. -pub fn escape_ident(s: &str) -> String { - let mut out = String::with_capacity(s.len() + 2); - out.push('"'); - for c in s.chars() { - if c == '"' { out.push('"'); } - out.push(c); - } - out.push('"'); - out -} - -// --- hand-rolled MD5 (RFC 1321) — for the `md5` auth exchange only, the -// --- same no-crates spirit as the CRC32 in wal.rs. Not for new designs. - -pub fn md5_hex(data: &[u8]) -> String { - const S: [u32; 64] = [ - 7, 12, 17, 22, 7, 12, 17, 22, 7, 12, 17, 22, 7, 12, 17, 22, - 5, 9, 14, 20, 5, 9, 14, 20, 5, 9, 14, 20, 5, 9, 14, 20, - 4, 11, 16, 23, 4, 11, 16, 23, 4, 11, 16, 23, 4, 11, 16, 23, - 6, 10, 15, 21, 6, 10, 15, 21, 6, 10, 15, 21, 6, 10, 15, 21, - ]; - const K: [u32; 64] = [ - 0xd76aa478, 0xe8c7b756, 0x242070db, 0xc1bdceee, 0xf57c0faf, 0x4787c62a, - 0xa8304613, 0xfd469501, 0x698098d8, 0x8b44f7af, 0xffff5bb1, 0x895cd7be, - 0x6b901122, 0xfd987193, 0xa679438e, 0x49b40821, 0xf61e2562, 0xc040b340, - 0x265e5a51, 0xe9b6c7aa, 0xd62f105d, 0x02441453, 0xd8a1e681, 0xe7d3fbc8, - 0x21e1cde6, 0xc33707d6, 0xf4d50d87, 0x455a14ed, 0xa9e3e905, 0xfcefa3f8, - 0x676f02d9, 0x8d2a4c8a, 0xfffa3942, 0x8771f681, 0x6d9d6122, 0xfde5380c, - 0xa4beea44, 0x4bdecfa9, 0xf6bb4b60, 0xbebfbc70, 0x289b7ec6, 0xeaa127fa, - 0xd4ef3085, 0x04881d05, 0xd9d4d039, 0xe6db99e5, 0x1fa27cf8, 0xc4ac5665, - 0xf4292244, 0x432aff97, 0xab9423a7, 0xfc93a039, 0x655b59c3, 0x8f0ccc92, - 0xffeff47d, 0x85845dd1, 0x6fa87e4f, 0xfe2ce6e0, 0xa3014314, 0x4e0811a1, - 0xf7537e82, 0xbd3af235, 0x2ad7d2bb, 0xeb86d391, - ]; - - let mut msg = data.to_vec(); - let bit_len = (data.len() as u64).wrapping_mul(8); - msg.push(0x80); - while msg.len() % 64 != 56 { msg.push(0); } - msg.extend_from_slice(&bit_len.to_le_bytes()); - - let (mut a0, mut b0, mut c0, mut d0) = - (0x6745_2301u32, 0xefcd_ab89u32, 0x98ba_dcfeu32, 0x1032_5476u32); - - for chunk in msg.chunks_exact(64) { - let m: Vec = chunk.chunks_exact(4) - .map(|w| u32::from_le_bytes(w.try_into().unwrap())) - .collect(); - let (mut a, mut b, mut c, mut d) = (a0, b0, c0, d0); - for i in 0..64 { - let (f, g) = match i { - 0..=15 => ((b & c) | (!b & d), i), - 16..=31 => ((d & b) | (!d & c), (5 * i + 1) % 16), - 32..=47 => (b ^ c ^ d, (3 * i + 5) % 16), - _ => (c ^ (b | !d), (7 * i) % 16), - }; - let f2 = f.wrapping_add(a).wrapping_add(K[i]).wrapping_add(m[g]); - a = d; d = c; c = b; - b = b.wrapping_add(f2.rotate_left(S[i])); - } - a0 = a0.wrapping_add(a); - b0 = b0.wrapping_add(b); - c0 = c0.wrapping_add(c); - d0 = d0.wrapping_add(d); - } - - let mut out = String::with_capacity(32); - for word in [a0, b0, c0, d0] { - for byte in word.to_le_bytes() { - out.push_str(&format!("{byte:02x}")); - } - } - out -} - -#[cfg(test)] -mod tests { - use super::*; - - #[test] - fn md5_matches_rfc_vectors() { - assert_eq!(md5_hex(b""), "d41d8cd98f00b204e9800998ecf8427e"); - assert_eq!(md5_hex(b"abc"), "900150983cd24fb0d6963f7d28e17f72"); - assert_eq!(md5_hex(b"message digest"), "f96b697d7cb7938d525a2f31aaf161d0"); - // > one block - assert_eq!( - md5_hex(b"12345678901234567890123456789012345678901234567890123456789012345678901234567890"), - "57edf4a22be3c955ac49da2e2107b67a"); - } - - #[test] - fn url_parse_covers_the_forms() { - let c = PgConfig::from_url("postgres://wo:secret@db.example:6432/prod").unwrap(); - assert_eq!((c.user.as_str(), c.password.as_deref(), c.host.as_str(), c.port, c.database.as_str()), - ("wo", Some("secret"), "db.example", 6432, "prod")); - let c = PgConfig::from_url("postgres://postgres@127.0.0.1/wo").unwrap(); - assert_eq!(c.port, 5432); - assert!(c.password.is_none()); - assert!(PgConfig::from_url("mysql://nope@x/y").is_err()); - assert!(PgConfig::from_url("postgres://user-only-no-host").is_err()); - } - - #[test] - fn escaping_doubles_quotes() { - assert_eq!(escape_literal("it's"), "'it''s'"); - assert_eq!(escape_literal(r#"back\slash"#), r#"'back\slash'"#); - assert_eq!(escape_ident(r#"we"ird"#), r#""we""ird""#); - } - - /// Integration: needs a reachable server — set WO_PG_TEST to run, e.g. - /// WO_PG_TEST=postgres://postgres@127.0.0.1:54329/wo cargo test pg_ - #[test] - fn pg_roundtrip_against_live_server() { - let Ok(url) = std::env::var("WO_PG_TEST") else { - eprintln!("pg_roundtrip: skipped (set WO_PG_TEST=postgres://... to run)"); - return; - }; - let cfg = PgConfig::from_url(&url).unwrap(); - let mut c = Conn::connect(&cfg).unwrap(); - - c.simple_query("DROP TABLE IF EXISTS wo_pg_smoke").unwrap(); - c.simple_query("CREATE TABLE wo_pg_smoke (id BIGINT PRIMARY KEY, row JSONB NOT NULL)").unwrap(); - c.simple_query(&format!( - "INSERT INTO wo_pg_smoke (id, row) VALUES (1, {}::jsonb) \ - ON CONFLICT (id) DO UPDATE SET row = EXCLUDED.row", - escape_literal(r#"{"amount":4999,"note":"it's fine"}"#))).unwrap(); - - let r = c.simple_query("SELECT row->>'amount', row->>'note' FROM wo_pg_smoke").unwrap(); - assert_eq!(r.rows.len(), 1); - assert_eq!(r.rows[0][0].as_deref(), Some("4999")); - assert_eq!(r.rows[0][1].as_deref(), Some("it's fine")); - - // A backend error must leave the connection usable. - assert!(c.simple_query("SELECT * FROM does_not_exist_xyz").is_err()); - let r = c.simple_query("SELECT count(*) FROM wo_pg_smoke").unwrap(); - assert_eq!(r.rows[0][0].as_deref(), Some("1")); - - // Multi-statement query = implicit transaction; both tags come back. - let r = c.simple_query( - "INSERT INTO wo_pg_smoke VALUES (2, '{}'::jsonb); DELETE FROM wo_pg_smoke WHERE id = 2" - ).unwrap(); - assert_eq!(r.tags.len(), 2); - c.simple_query("DROP TABLE wo_pg_smoke").unwrap(); - } -} diff --git a/crates/rt/src/runtime/eventfd.rs b/crates/rt/src/runtime/eventfd.rs deleted file mode 100644 index 642a0b5..0000000 --- a/crates/rt/src/runtime/eventfd.rs +++ /dev/null @@ -1,66 +0,0 @@ -//! `eventfd(2)` wrapper — counter semaphore exposed as a file descriptor. -//! -//! Used as the cross-flow wake-up primitive: any code path that needs the -//! event loop to come back and run something writes a `1` to the eventfd, -//! which becomes readable on the loop's next `wait_once`. -//! -//! Ported from `reference/crates/wo-event/src/eventfd.rs` with an added -//! `AsRawFd` impl so callers can drop the fd straight into `EventLoop`. - -use std::io; -use std::os::unix::io::{AsRawFd, RawFd}; - -pub struct EventFd { - fd: RawFd, -} - -impl EventFd { - /// Create a non-blocking, close-on-exec eventfd with initial counter 0. - pub fn new() -> io::Result { - let fd = unsafe { libc::eventfd(0, libc::EFD_NONBLOCK | libc::EFD_CLOEXEC) }; - if fd < 0 { return Err(io::Error::last_os_error()); } - Ok(Self { fd }) - } - - /// Add `val` to the counter. Wakes any waiter when the counter goes 0→non-zero. - pub fn write(&self, val: u64) -> io::Result<()> { - let buf = val.to_ne_bytes(); - let ret = unsafe { - libc::write(self.fd, buf.as_ptr() as *const libc::c_void, 8) - }; - if ret < 0 { Err(io::Error::last_os_error()) } else { Ok(()) } - } - - /// Atomically read and zero the counter. Returns `EAGAIN` if the - /// counter is already 0 (since we set `EFD_NONBLOCK`). - pub fn read(&self) -> io::Result { - let mut buf = [0u8; 8]; - let ret = unsafe { - libc::read(self.fd, buf.as_mut_ptr() as *mut libc::c_void, 8) - }; - if ret < 0 { Err(io::Error::last_os_error()) } else { Ok(u64::from_ne_bytes(buf)) } - } -} - -impl AsRawFd for EventFd { - fn as_raw_fd(&self) -> RawFd { self.fd } -} - -impl Drop for EventFd { - fn drop(&mut self) { - unsafe { libc::close(self.fd) }; - } -} - -#[cfg(test)] -mod tests { - use super::*; - - #[test] - fn writes_accumulate() { - let efd = EventFd::new().unwrap(); - efd.write(5).unwrap(); - efd.write(3).unwrap(); - assert_eq!(efd.read().unwrap(), 8); - } -} diff --git a/crates/rt/src/runtime/mod.rs b/crates/rt/src/runtime/mod.rs deleted file mode 100644 index 8a1453c..0000000 --- a/crates/rt/src/runtime/mod.rs +++ /dev/null @@ -1,25 +0,0 @@ -//! Single-threaded event loop on Linux kernel primitives. -//! -//! Phase 02 of the runtime plan — see `docs/plan/02-event-loop-epoll.md`. -//! Wraps `epoll`, `eventfd`, `timerfd`, and `signalfd` directly via `libc`, -//! with no `tokio` / `mio` / `nix` involvement. Every fd is registered -//! edge-triggered (`EPOLLET`); the loop reads to `EAGAIN`. -//! -//! Lives alongside the tokio-backed `server` module until phase 04 cuts -//! the binary over. -//! -//! Filename convention mirrors Go's `src/runtime/netpoll_.go`; -//! when phase 03 adds io_uring it lands as `netpoll_io_uring.rs` next door. - -mod eventfd; -mod netpoll_epoll; -mod netpoll_io_uring; -pub mod scheduler; -mod signalfd; -mod timerfd; - -pub use eventfd::EventFd; -pub use netpoll_epoll::{Event, EventLoop, Interest, Token}; -pub use netpoll_io_uring::Uring; -pub use signalfd::SignalFd; -pub use timerfd::TimerFd; diff --git a/crates/rt/src/runtime/netpoll_epoll.rs b/crates/rt/src/runtime/netpoll_epoll.rs deleted file mode 100644 index e251130..0000000 --- a/crates/rt/src/runtime/netpoll_epoll.rs +++ /dev/null @@ -1,170 +0,0 @@ -//! `epoll`-backed event loop. Single-threaded, edge-triggered. -//! -//! Ported from `reference/crates/wo-event/src/epoll.rs`. Differences: -//! * `Token` is a newtype rather than a `u64` alias. -//! * `Interest` is a struct exposing `READABLE`, `WRITABLE`, `READ_WRITE` -//! constants, matching the API in `docs/plan/02-event-loop-epoll.md`. -//! * Every registration sets `EPOLLET`; callers must read to `EAGAIN`. -//! * `wait_once` reuses an internal event buffer instead of allocating -//! a fresh `[epoll_event; 64]` per call. - -use std::io; -use std::os::unix::io::RawFd; -use std::time::Duration; - -/// Caller-assigned identifier for a registered file descriptor. -#[derive(Debug, Clone, Copy, PartialEq, Eq, Hash)] -pub struct Token(pub u64); - -/// Interest flags for epoll registration. -#[derive(Debug, Clone, Copy, PartialEq, Eq)] -pub struct Interest(u32); - -impl Interest { - pub const READABLE: Interest = Interest(libc::EPOLLIN as u32); - pub const WRITABLE: Interest = Interest(libc::EPOLLOUT as u32); - pub const READ_WRITE: Interest = Interest((libc::EPOLLIN | libc::EPOLLOUT) as u32); - - fn bits(self) -> u32 { self.0 } -} - -/// One readiness event from the loop. -#[derive(Debug, Clone)] -pub struct Event { - token: Token, - pub readable: bool, - pub writable: bool, - pub error: bool, - pub hangup: bool, -} - -impl Event { - pub fn token(&self) -> Token { self.token } -} - -const EVENT_BUF: usize = 64; - -/// Single-threaded event loop built on `epoll_create1` + `epoll_wait`. -pub struct EventLoop { - epoll_fd: RawFd, - events: Vec, -} - -impl EventLoop { - pub fn new() -> io::Result { - let fd = unsafe { libc::epoll_create1(libc::EPOLL_CLOEXEC) }; - if fd < 0 { return Err(io::Error::last_os_error()); } - Ok(Self { - epoll_fd: fd, - events: vec![libc::epoll_event { events: 0, u64: 0 }; EVENT_BUF], - }) - } - - pub fn register(&self, fd: RawFd, interest: Interest, token: Token) -> io::Result<()> { - self.ctl(libc::EPOLL_CTL_ADD, fd, interest, token) - } - - pub fn modify(&self, fd: RawFd, interest: Interest, token: Token) -> io::Result<()> { - self.ctl(libc::EPOLL_CTL_MOD, fd, interest, token) - } - - pub fn deregister(&self, fd: RawFd) -> io::Result<()> { - let ret = unsafe { - libc::epoll_ctl(self.epoll_fd, libc::EPOLL_CTL_DEL, fd, std::ptr::null_mut()) - }; - if ret < 0 { Err(io::Error::last_os_error()) } else { Ok(()) } - } - - fn ctl(&self, op: libc::c_int, fd: RawFd, interest: Interest, token: Token) -> io::Result<()> { - let mut ev = libc::epoll_event { - events: interest.bits() | libc::EPOLLET as u32 | libc::EPOLLRDHUP as u32, - u64: token.0, - }; - let ret = unsafe { libc::epoll_ctl(self.epoll_fd, op, fd, &mut ev) }; - if ret < 0 { Err(io::Error::last_os_error()) } else { Ok(()) } - } - - /// Block until at least one fd is ready or `timeout` expires. - /// `None` blocks indefinitely. Returns up to 64 events per call. - pub fn wait_once(&mut self, timeout: Option) -> io::Result> { - let timeout_ms: i32 = match timeout { - None => -1, - Some(d) => d.as_millis().min(i32::MAX as u128) as i32, - }; - let n = unsafe { - libc::epoll_wait( - self.epoll_fd, - self.events.as_mut_ptr(), - self.events.len() as i32, - timeout_ms, - ) - }; - if n < 0 { - let err = io::Error::last_os_error(); - if err.raw_os_error() == Some(libc::EINTR) { return Ok(vec![]); } - return Err(err); - } - Ok((0..n as usize) - .map(|i| { - let bits = self.events[i].events; - Event { - token: Token(self.events[i].u64), - readable: bits & libc::EPOLLIN as u32 != 0, - writable: bits & libc::EPOLLOUT as u32 != 0, - error: bits & libc::EPOLLERR as u32 != 0, - hangup: bits & (libc::EPOLLHUP | libc::EPOLLRDHUP) as u32 != 0, - } - }) - .collect()) - } - - pub fn fd(&self) -> RawFd { self.epoll_fd } -} - -impl Drop for EventLoop { - fn drop(&mut self) { - unsafe { libc::close(self.epoll_fd) }; - } -} - -#[cfg(test)] -mod tests { - use super::*; - use crate::runtime::EventFd; - use std::os::unix::io::AsRawFd; - - #[test] - fn wait_once_returns_eventfd_token() { - let mut eloop = EventLoop::new().unwrap(); - let efd = EventFd::new().unwrap(); - - eloop.register(efd.as_raw_fd(), Interest::READABLE, Token(42)).unwrap(); - efd.write(1).unwrap(); - - let events = eloop.wait_once(Some(Duration::from_millis(100))).unwrap(); - assert_eq!(events.len(), 1); - assert_eq!(events[0].token(), Token(42)); - assert!(events[0].readable); - assert_eq!(efd.read().unwrap(), 1); - } - - #[test] - fn wait_once_times_out_with_no_events() { - let mut eloop = EventLoop::new().unwrap(); - let events = eloop.wait_once(Some(Duration::from_millis(10))).unwrap(); - assert!(events.is_empty()); - } - - #[test] - fn deregister_silences_fd() { - let mut eloop = EventLoop::new().unwrap(); - let efd = EventFd::new().unwrap(); - - eloop.register(efd.as_raw_fd(), Interest::READABLE, Token(1)).unwrap(); - eloop.deregister(efd.as_raw_fd()).unwrap(); - - efd.write(1).unwrap(); - let events = eloop.wait_once(Some(Duration::from_millis(10))).unwrap(); - assert!(events.is_empty()); - } -} diff --git a/crates/rt/src/runtime/netpoll_io_uring.rs b/crates/rt/src/runtime/netpoll_io_uring.rs deleted file mode 100644 index b891017..0000000 --- a/crates/rt/src/runtime/netpoll_io_uring.rs +++ /dev/null @@ -1,237 +0,0 @@ -//! Raw io_uring — no liburing, kernel ABI structs defined by hand, exactly -//! the sequence proven in C (`prototypes/wo-rt-c/wo-rt.c` ring_init/enter; -//! card: `docs/plan/exploration/linux/07-io_uring.md`). -//! -//! Scope (this phase): the **storage ring** for per-shard group commit — -//! batched WAL `WRITE` + hard-linked `FSYNC` SQEs, one `io_uring_enter` per -//! flush. The ring fd is pollable, so it registers in the existing epoll -//! loop and completions arrive as just another readable event; the full -//! network port (accept/recv/send SQEs) is a later phase. -//! -//! Requires `IORING_FEAT_SINGLE_MMAP` (kernel ≥ 5.4). - -use std::io; -use std::os::unix::io::RawFd; -use std::sync::atomic::{AtomicU32, Ordering}; - -const SYS_SETUP: libc::c_long = 425; -const SYS_ENTER: libc::c_long = 426; - -const IORING_OFF_SQ_RING: i64 = 0; -const IORING_OFF_SQES: i64 = 0x1000_0000; -const IORING_ENTER_GETEVENTS: u32 = 1; -const IORING_FEAT_SINGLE_MMAP: u32 = 1; - -pub const OP_FSYNC: u8 = 3; -pub const OP_WRITE: u8 = 23; -pub const IOSQE_IO_LINK: u8 = 1 << 2; -pub const FSYNC_DATASYNC: u32 = 1; - -#[repr(C)] -#[derive(Default)] -struct SqOffsets { head: u32, tail: u32, ring_mask: u32, ring_entries: u32, flags: u32, dropped: u32, array: u32, resv1: u32, user_addr: u64 } - -#[repr(C)] -#[derive(Default)] -struct CqOffsets { head: u32, tail: u32, ring_mask: u32, ring_entries: u32, overflow: u32, cqes: u32, flags: u32, resv1: u32, user_addr: u64 } - -#[repr(C)] -#[derive(Default)] -struct Params { - sq_entries: u32, cq_entries: u32, flags: u32, - sq_thread_cpu: u32, sq_thread_idle: u32, - features: u32, wq_fd: u32, resv: [u32; 3], - sq_off: SqOffsets, cq_off: CqOffsets, -} - -/// One submission-queue entry — the 64-byte kernel layout. -#[repr(C)] -#[derive(Clone, Copy, Default)] -struct Sqe { - opcode: u8, flags: u8, ioprio: u16, fd: i32, - off: u64, addr: u64, len: u32, op_flags: u32, - user_data: u64, - buf_index: u16, personality: u16, splice_fd_in: i32, - _pad: [u64; 2], -} - -#[repr(C)] -#[derive(Clone, Copy)] -struct Cqe { user_data: u64, res: i32, flags: u32 } - -pub struct Uring { - fd: RawFd, - // SQ ring pointers (into the shared mmap) - sq_tail: *const AtomicU32, - sq_mask: u32, - sq_array: *mut u32, - // CQ ring pointers - cq_head: *const AtomicU32, - cq_tail: *const AtomicU32, - cq_mask: u32, - cqes: *const Cqe, - sqes: *mut Sqe, - local_tail: u32, - to_submit: u32, -} - -// The ring is owned and driven by exactly one worker thread (plan 09 -// decision 4); raw pointers into its own mmaps don't change that. -unsafe impl Send for Uring {} - -impl Uring { - pub fn new(entries: u32) -> io::Result { - let mut p = Params::default(); - let fd = unsafe { libc::syscall(SYS_SETUP, entries, &mut p as *mut Params) } as RawFd; - if fd < 0 { return Err(io::Error::last_os_error()); } - if p.features & IORING_FEAT_SINGLE_MMAP == 0 { - unsafe { libc::close(fd) }; - return Err(io::Error::other("kernel lacks IORING_FEAT_SINGLE_MMAP (need >= 5.4)")); - } - - let sq_sz = p.sq_off.array as usize + p.sq_entries as usize * 4; - let cq_sz = p.cq_off.cqes as usize + p.cq_entries as usize * std::mem::size_of::(); - let ring_sz = sq_sz.max(cq_sz); - let ring = unsafe { - libc::mmap(std::ptr::null_mut(), ring_sz, libc::PROT_READ | libc::PROT_WRITE, - libc::MAP_SHARED | libc::MAP_POPULATE, fd, IORING_OFF_SQ_RING) - }; - if ring == libc::MAP_FAILED { let e = io::Error::last_os_error(); unsafe { libc::close(fd) }; return Err(e); } - - let sqes_sz = p.sq_entries as usize * std::mem::size_of::(); - let sqes = unsafe { - libc::mmap(std::ptr::null_mut(), sqes_sz, libc::PROT_READ | libc::PROT_WRITE, - libc::MAP_SHARED | libc::MAP_POPULATE, fd, IORING_OFF_SQES) - }; - if sqes == libc::MAP_FAILED { let e = io::Error::last_os_error(); unsafe { libc::close(fd) }; return Err(e); } - - let at = |off: u32| unsafe { (ring as *mut u8).add(off as usize) }; - let sq_mask = unsafe { *(at(p.sq_off.ring_mask) as *const u32) }; - let cq_mask = unsafe { *(at(p.cq_off.ring_mask) as *const u32) }; - let sq_tail = at(p.sq_off.tail) as *const AtomicU32; - let local_tail = unsafe { (*sq_tail).load(Ordering::Relaxed) }; - - Ok(Self { - fd, - sq_tail, - sq_mask, - sq_array: at(p.sq_off.array) as *mut u32, - cq_head: at(p.cq_off.head) as *const AtomicU32, - cq_tail: at(p.cq_off.tail) as *const AtomicU32, - cq_mask, - cqes: at(p.cq_off.cqes) as *const Cqe, - sqes: sqes as *mut Sqe, - local_tail, - to_submit: 0, - }) - } - - /// The ring fd — readable when completions are pending, so it registers - /// straight into the epoll loop. - pub fn as_raw_fd(&self) -> RawFd { self.fd } - - fn sqe(&mut self) -> &mut Sqe { - let idx = self.local_tail & self.sq_mask; - unsafe { *self.sq_array.add(idx as usize) = idx; } - self.local_tail = self.local_tail.wrapping_add(1); - self.to_submit += 1; - let s = unsafe { &mut *self.sqes.add(idx as usize) }; - *s = Sqe::default(); - s - } - - /// Queue a positional write. SAFETY contract: `buf` must stay alive and - /// unmoved until this op's CQE is reaped — the caller double-buffers. - pub fn push_write(&mut self, fd: RawFd, buf: &[u8], offset: u64, link: bool, user_data: u64) { - let s = self.sqe(); - s.opcode = OP_WRITE; - s.fd = fd; - s.addr = buf.as_ptr() as u64; - s.len = buf.len() as u32; - s.off = offset; - s.flags = if link { IOSQE_IO_LINK } else { 0 }; - s.user_data = user_data; - } - - pub fn push_fsync(&mut self, fd: RawFd, user_data: u64) { - let s = self.sqe(); - s.opcode = OP_FSYNC; - s.fd = fd; - s.op_flags = FSYNC_DATASYNC; - s.user_data = user_data; - } - - /// Publish queued SQEs with one syscall. Non-blocking — completions are - /// observed via epoll on the ring fd. - pub fn submit(&mut self) -> io::Result<()> { - if self.to_submit == 0 { return Ok(()); } - unsafe { (*self.sq_tail).store(self.local_tail, Ordering::Release); } - let n = self.to_submit; - self.to_submit = 0; - loop { - let r = unsafe { libc::syscall(SYS_ENTER, self.fd, n, 0u32, IORING_ENTER_GETEVENTS, 0usize, 0usize) }; - if r >= 0 { return Ok(()); } - let e = io::Error::last_os_error(); - if e.raw_os_error() != Some(libc::EINTR) { return Err(e); } - } - } - - /// Reap every pending completion as `(user_data, res)`. - pub fn pop_cqes(&mut self) -> Vec<(u64, i32)> { - let mut out = Vec::new(); - unsafe { - let mut head = (*self.cq_head).load(Ordering::Relaxed); - let tail = (*self.cq_tail).load(Ordering::Acquire); - while head != tail { - let cqe = &*self.cqes.add((head & self.cq_mask) as usize); - out.push((cqe.user_data, cqe.res)); - head = head.wrapping_add(1); - } - (*self.cq_head).store(head, Ordering::Release); - } - out - } -} - -impl Drop for Uring { - fn drop(&mut self) { - unsafe { libc::close(self.fd) }; - } -} - -#[cfg(test)] -mod tests { - use super::*; - use std::io::Read; - - #[test] - fn write_then_linked_fsync_round_trips() { - let path = std::env::temp_dir().join(format!("wo-uring-test-{}", std::process::id())); - let _ = std::fs::remove_file(&path); - let file = std::fs::OpenOptions::new().read(true).write(true).create(true).open(&path).unwrap(); - use std::os::unix::io::AsRawFd; - - let mut ring = Uring::new(8).expect("io_uring available"); - let buf = b"hello-from-the-ring".to_vec(); - ring.push_write(file.as_raw_fd(), &buf, 0, true, 1); - ring.push_fsync(file.as_raw_fd(), 2); - ring.submit().unwrap(); - - // Poll the ring fd until both CQEs arrive. - let mut got = Vec::new(); - for _ in 0..200 { - got.extend(ring.pop_cqes()); - if got.len() >= 2 { break; } - std::thread::sleep(std::time::Duration::from_millis(1)); - } - assert_eq!(got.len(), 2, "write + fsync completions"); - assert_eq!(got[0], (1, buf.len() as i32), "write res = full length"); - assert_eq!(got[1].0, 2); - assert!(got[1].1 >= 0, "fsync ok"); - - let mut s = String::new(); - std::fs::File::open(&path).unwrap().read_to_string(&mut s).unwrap(); - assert_eq!(s, "hello-from-the-ring"); - let _ = std::fs::remove_file(&path); - } -} diff --git a/crates/rt/src/runtime/scheduler.rs b/crates/rt/src/runtime/scheduler.rs deleted file mode 100644 index 89992bf..0000000 --- a/crates/rt/src/runtime/scheduler.rs +++ /dev/null @@ -1,276 +0,0 @@ -//! Thread-per-core scheduler — plan 09a (`docs/plan/09-concurrency-scaleout.md`). -//! -//! Go parallel: `src/runtime/proc.go`, drastically simplified — there are no -//! goroutines to schedule. `WO_THREADS` OS threads (default: online cores) -//! spawn at boot; each pins itself to a core with `sched_setaffinity`, binds -//! its own `SO_REUSEPORT` listener on the shared port, and runs its own -//! [`EventLoop`] over its own connections. A connection accepted on thread K -//! is driven and closed on thread K — no migration, no work stealing. -//! -//! Per plan 09a, **engine state stays globally shared** (`Arc>` -//! inside the per-thread `Router`s) — one thing at a time; the sharded engine -//! is 09b. The C proving ground for this exact sequence is -//! `prototypes/wo-rt-c` phase A (see `docs/plan/exploration/c-runtime/`). -//! -//! Shutdown: signals are blocked in `main` before any worker spawns (the -//! mask is inherited), so only worker 0 — which owns the `signalfd` — ever -//! sees SIGINT/SIGTERM. It broadcasts by writing every worker's `eventfd`; -//! each loop wakes, drains its connections, and joins. - -use std::collections::HashMap; -use std::io; -use std::os::unix::io::{AsRawFd, RawFd}; -use std::sync::Arc; -use std::time::Duration; - -use crate::http::{Connection, Listener, Router}; - -use super::{EventFd, EventLoop, Interest, SignalFd, Token}; - -const MAX_THREADS: usize = 64; - -/// Resolve the worker count: `WO_THREADS` env override, else online cores. -pub fn thread_count() -> usize { - std::env::var("WO_THREADS") - .ok() - .and_then(|s| s.parse::().ok()) - .filter(|&n| n >= 1) - .unwrap_or_else(|| { - std::thread::available_parallelism().map(|n| n.get()).unwrap_or(1) - }) - .min(MAX_THREADS) -} - -fn pin_to_core(core: usize) { - let cores = std::thread::available_parallelism().map(|n| n.get()).unwrap_or(1); - unsafe { - let mut set: libc::cpu_set_t = std::mem::zeroed(); - libc::CPU_ZERO(&mut set); - libc::CPU_SET(core % cores, &mut set); - // 0 = the calling thread. Best-effort: a denied affinity (cgroup - // restrictions) must not stop the worker from serving. - libc::sched_setaffinity(0, std::mem::size_of::(), &set); - } -} - -/// Everything a worker needs beyond its listener: the router over its own -/// shard, and (since 09b) an optional auxiliary fd + callback — the shard -/// bus's mail eventfd, drained into the local engine when it fires. -pub struct Worker { - pub router: Router, - pub mail: Option<(RawFd, Box)>, - /// Group-commit hooks (io_uring WAL): the pollable ring fd, a pump - /// (flush staged batch / reap completions), a drain of pending - /// connection unparks `(fd, gen, durable_ok)`, and a parker invoked - /// when a drive leaves a connection gated on the next fsync. - pub wal: Option, -} - -pub struct WalHooks { - pub ring_fd: RawFd, - pub pump: Box, - pub unparks: Box Vec<(RawFd, u64, bool)>>, - pub park_conn: Box, -} - -/// Spawn `thread_count()` pinned workers, each serving `addr` behind -/// `SO_REUSEPORT` with the [`Worker`] built by `worker_fn(id)`. Blocks -/// until a SIGINT/SIGTERM shuts every worker down. -pub fn serve(addr: &str, worker_fn: F) -> anyhow::Result<()> -where - F: Fn(usize) -> Worker + Send + Sync + 'static, -{ - let n = thread_count(); - - // Block SIGINT/SIGTERM NOW — every worker inherits the mask, so the - // signalfd (owned by worker 0) is the only delivery path. - let signals = SignalFd::new()?; - let sig_raw = signals.as_raw_fd(); - - let wake: Arc> = - Arc::new((0..n).map(|_| EventFd::new()).collect::>>()?); - let worker_fn = Arc::new(worker_fn); - - let mut handles = Vec::with_capacity(n); - for t in 0..n { - let wake = Arc::clone(&wake); - let worker_fn = Arc::clone(&worker_fn); - let addr = addr.to_string(); - let sigfd = (t == 0).then_some(sig_raw); - handles.push( - std::thread::Builder::new() - .name(format!("wo-shard-{t}")) - .spawn(move || worker(t, &addr, sigfd, &wake, &*worker_fn))?, - ); - } - - for h in handles { - match h.join() { - Ok(Ok(())) => {} - Ok(Err(e)) => eprintln!("[wo] worker error: {e}"), - Err(_) => eprintln!("[wo] worker panicked"), - } - } - drop(signals); - Ok(()) -} - -fn worker( - id: usize, - addr: &str, - sigfd: Option, - wake: &[EventFd], - worker_fn: &(dyn Fn(usize) -> Worker + Send + Sync), -) -> anyhow::Result<()> { - pin_to_core(id); - - let listener = match Listener::bind_reuseport(addr) { - Ok(l) => l, - Err(e) => { - // Without a listener this worker is useless — take the whole - // process down cleanly rather than serving with a hole. - for w in wake { let _ = w.write(1); } - anyhow::bail!("shard {id}: bind {addr}: {e}"); - } - }; - let Worker { router, mut mail, mut wal } = worker_fn(id); - - let mut eloop = EventLoop::new()?; - let listen_fd = listener.as_raw_fd(); - let wake_fd = wake[id].as_raw_fd(); - - eloop.register(listen_fd, Interest::READABLE, Token(listen_fd as u64))?; - eloop.register(wake_fd, Interest::READABLE, Token(wake_fd as u64))?; - if let Some(sfd) = sigfd { - eloop.register(sfd, Interest::READABLE, Token(sfd as u64))?; - } - let mail_fd = mail.as_ref().map(|(fd, _)| *fd); - if let Some(mfd) = mail_fd { - eloop.register(mfd, Interest::READABLE, Token(mfd as u64))?; - } - let ring_fd = wal.as_ref().map(|w| w.ring_fd); - if let Some(rfd) = ring_fd { - eloop.register(rfd, Interest::READABLE, Token(rfd as u64))?; - } - let mut next_gen: u64 = 1; - - let mut conns: HashMap = HashMap::new(); - - 'outer: loop { - let events = match eloop.wait_once(Some(Duration::from_secs(60))) { - Ok(evs) => evs, - Err(e) => { - eprintln!("[wo] shard {id}: event loop error: {e}"); - continue; - } - }; - - for ev in events { - let fd = ev.token().0 as RawFd; - - if fd == wake_fd { - let _ = wake[id].read(); - break 'outer; // shutdown broadcast - } - - if Some(fd) == ring_fd { - if let Some(w) = wal.as_mut() { (w.pump)(); } - continue; - } - - if Some(fd) == mail_fd { - // Reset the edge, then drain the shard-bus inbox into the - // local engine. (The eventfd counter is read-and-zeroed; - // jobs arriving mid-drain re-arm the edge.) - let mut buf = [0u8; 8]; - unsafe { libc::read(fd, buf.as_mut_ptr() as *mut libc::c_void, 8) }; - if let Some((_, drain)) = mail.as_mut() { drain(); } - continue; - } - - if Some(fd) == sigfd { - // Drain the siginfo and broadcast shutdown to every shard - // (including ourselves — we exit through the wake path). - let mut buf = [0u8; 128]; // sizeof(signalfd_siginfo) - let r = unsafe { - libc::read(fd, buf.as_mut_ptr() as *mut libc::c_void, buf.len()) - }; - let signo = if r >= 4 { - u32::from_ne_bytes([buf[0], buf[1], buf[2], buf[3]]) - } else { 0 }; - println!(); - println!("[wo] received signal {signo} — broadcasting shutdown to {} shards", wake.len()); - for w in wake { let _ = w.write(1); } - continue; - } - - if fd == listen_fd { - // Drain the accept queue (edge-triggered). - while let Some(cfd) = listener.accept()? { - eloop.register(cfd, Interest::READABLE, Token(cfd as u64))?; - conns.insert(cfd, Connection::with_gen(cfd, next_gen)); - next_gen += 1; - } - continue; - } - - // Connection event — same state machine as the single-threaded - // loop; the only difference is whose loop it runs on. - let Some(conn) = conns.get_mut(&fd) else { continue }; - let want_writable = match conn.drive(ev.readable, ev.writable, ev.hangup, ev.error, &router) { - Ok(w) => w, - Err(_) => { conns.remove(&fd); continue; } - }; - - if conn.is_done() { - eloop.deregister(fd).ok(); - conns.remove(&fd); // Drop closes the fd. - } else if conn.is_parked() { - let g = conn.gen(); - if let Some(w) = wal.as_mut() { (w.park_conn)(fd, g); } - } else if want_writable { - let _ = eloop.modify(fd, Interest::READ_WRITE, Token(fd as u64)); - } - } - - // End of tick: flush the group-commit batch (one WRITE→FSYNC pair, - // one syscall) and release any unparks the pumps produced. Released - // connections may serve pipelined requests that commit again — loop - // until quiescent so nothing sleeps on an unflushed batch. - if let Some(w) = wal.as_mut() { - for _ in 0..64 { - (w.pump)(); - let pending = (w.unparks)(); - if pending.is_empty() { break; } - for (fd, gen, ok) in pending { - let Some(conn) = conns.get_mut(&fd) else { continue }; - if conn.gen() != gen || !conn.is_parked() { continue; } - if !ok { - eloop.deregister(fd).ok(); - conns.remove(&fd); // never ack non-durable - continue; - } - conn.unpark(); - let want_writable = match conn.drive(false, true, false, false, &router) { - Ok(wb) => wb, - Err(_) => { conns.remove(&fd); continue; } - }; - if conn.is_done() { - eloop.deregister(fd).ok(); - conns.remove(&fd); - } else if conn.is_parked() { - let g = conn.gen(); - (w.park_conn)(fd, g); - } else if want_writable { - let _ = eloop.modify(fd, Interest::READ_WRITE, Token(fd as u64)); - } - } - } - } - } - - for (fd, _) in conns.drain() { - let _ = eloop.deregister(fd); - } - Ok(()) -} diff --git a/crates/rt/src/runtime/signalfd.rs b/crates/rt/src/runtime/signalfd.rs deleted file mode 100644 index a9ef929..0000000 --- a/crates/rt/src/runtime/signalfd.rs +++ /dev/null @@ -1,52 +0,0 @@ -//! `signalfd(2)` wrapper — POSIX signals delivered as fd reads. -//! -//! Replaces `tokio::signal::unix::signal` for graceful shutdown: the loop -//! gets `SIGINT` / `SIGTERM` as just another readable fd it can poll. -//! -//! Blocks the captured signals in the calling thread's mask, so the -//! kernel routes them to the signalfd instead of running default handlers. -//! Ported from `reference/crates/wo-event/src/signalfd.rs`. - -use std::io; -use std::os::unix::io::{AsRawFd, RawFd}; - -pub struct SignalFd { - fd: RawFd, -} - -impl SignalFd { - /// Capture `SIGINT` and `SIGTERM`. Both are blocked process-wide. - pub fn new() -> io::Result { - let mut mask: libc::sigset_t = unsafe { std::mem::zeroed() }; - unsafe { - libc::sigemptyset(&mut mask); - libc::sigaddset(&mut mask, libc::SIGINT); - libc::sigaddset(&mut mask, libc::SIGTERM); - let ret = libc::pthread_sigmask(libc::SIG_BLOCK, &mask, std::ptr::null_mut()); - if ret != 0 { return Err(io::Error::from_raw_os_error(ret)); } - } - let fd = unsafe { libc::signalfd(-1, &mask, libc::SFD_NONBLOCK | libc::SFD_CLOEXEC) }; - if fd < 0 { return Err(io::Error::last_os_error()); } - Ok(Self { fd }) - } - - /// Read one pending signal. Returns the signal number (e.g. `SIGINT = 2`). - pub fn read(&self) -> io::Result { - let mut info: libc::signalfd_siginfo = unsafe { std::mem::zeroed() }; - let size = std::mem::size_of::(); - let ret = unsafe { - libc::read(self.fd, &mut info as *mut _ as *mut libc::c_void, size) - }; - if ret < 0 { Err(io::Error::last_os_error()) } else { Ok(info.ssi_signo as i32) } - } -} - -impl AsRawFd for SignalFd { - fn as_raw_fd(&self) -> RawFd { self.fd } -} - -impl Drop for SignalFd { - fn drop(&mut self) { - unsafe { libc::close(self.fd) }; - } -} diff --git a/crates/rt/src/runtime/timerfd.rs b/crates/rt/src/runtime/timerfd.rs deleted file mode 100644 index bf8f8f0..0000000 --- a/crates/rt/src/runtime/timerfd.rs +++ /dev/null @@ -1,102 +0,0 @@ -//! `timerfd_create(2)` wrapper — timers as file descriptors. -//! -//! Both one-shot and periodic timers are armed with `timerfd_settime`. The -//! fd becomes readable when the timer expires; reading drains the -//! expiration count. -//! -//! Ported from `reference/crates/wo-event/src/timerfd.rs`. Adds `oneshot` -//! and `periodic` constructors that match the API in the phase-02 plan. - -use std::io; -use std::os::unix::io::{AsRawFd, RawFd}; -use std::time::Duration; - -pub struct TimerFd { - fd: RawFd, -} - -impl TimerFd { - /// Create a disarmed monotonic timer fd (non-blocking, close-on-exec). - pub fn new() -> io::Result { - let fd = unsafe { - libc::timerfd_create( - libc::CLOCK_MONOTONIC, - libc::TFD_NONBLOCK | libc::TFD_CLOEXEC, - ) - }; - if fd < 0 { return Err(io::Error::last_os_error()); } - Ok(Self { fd }) - } - - /// Convenience: a fresh fd armed to fire once after `after`. - pub fn oneshot(after: Duration) -> io::Result { - let t = Self::new()?; - t.set(after, Duration::ZERO)?; - Ok(t) - } - - /// Convenience: a fresh fd that fires every `every` (first tick at +`every`). - pub fn periodic(every: Duration) -> io::Result { - let t = Self::new()?; - t.set(every, every)?; - Ok(t) - } - - /// Arm: fire once after `initial`, then repeat every `interval`. - /// Pass `Duration::ZERO` for `interval` to make it one-shot. - pub fn set(&self, initial: Duration, interval: Duration) -> io::Result<()> { - let spec = libc::itimerspec { - it_interval: timespec(interval), - it_value: timespec(initial), - }; - let ret = unsafe { - libc::timerfd_settime(self.fd, 0, &spec, std::ptr::null_mut()) - }; - if ret < 0 { Err(io::Error::last_os_error()) } else { Ok(()) } - } - - /// Read the number of expirations since the previous read. - pub fn read(&self) -> io::Result { - let mut buf = [0u8; 8]; - let ret = unsafe { - libc::read(self.fd, buf.as_mut_ptr() as *mut libc::c_void, 8) - }; - if ret < 0 { Err(io::Error::last_os_error()) } else { Ok(u64::from_ne_bytes(buf)) } - } -} - -impl AsRawFd for TimerFd { - fn as_raw_fd(&self) -> RawFd { self.fd } -} - -impl Drop for TimerFd { - fn drop(&mut self) { - unsafe { libc::close(self.fd) }; - } -} - -fn timespec(d: Duration) -> libc::timespec { - libc::timespec { - tv_sec: d.as_secs() as libc::time_t, - tv_nsec: d.subsec_nanos() as libc::c_long, - } -} - -#[cfg(test)] -mod tests { - use super::*; - use crate::runtime::{EventLoop, Interest, Token}; - - #[test] - fn oneshot_fires_within_window() { - let mut eloop = EventLoop::new().unwrap(); - let timer = TimerFd::oneshot(Duration::from_millis(100)).unwrap(); - - eloop.register(timer.as_raw_fd(), Interest::READABLE, Token(99)).unwrap(); - - let events = eloop.wait_once(Some(Duration::from_millis(500))).unwrap(); - assert!(!events.is_empty(), "expected timer event within 500ms"); - assert_eq!(events[0].token(), Token(99)); - assert!(timer.read().unwrap() >= 1); - } -} diff --git a/crates/rt/src/server.rs b/crates/rt/src/server.rs deleted file mode 100644 index 430212e..0000000 --- a/crates/rt/src/server.rs +++ /dev/null @@ -1,601 +0,0 @@ -//! REST routing built from `service rest` blocks in the compiled catalog. -//! -//! For each type that declares `service rest "/path" expose ...`, we bind the -//! exposed operations at the given path: -//! -//! ```text -//! list GET /path -//! get GET /path/:id -//! create POST /path -//! update PATCH /path/:id -//! delete DELETE /path/:id -//! subscribe stubbed in Stage 2; wires up in Stage 3 (501) -//! me GET /path/me (501) -//! ``` -//! -//! Since plan 09b the engine is **sharded**: each worker thread owns its own -//! [`Engine`] behind a [`ShardCtx`](crate::shard::ShardCtx) — no mutex, no -//! shared heap. Handlers resolve the owning shard from the row id -//! (`owner = (id-1) % n`, the interleaved-mint rule), run locally when it's -//! ours, ship a job over the shard bus when it isn't. Creates always mint -//! locally; lists fan out to every shard and merge by id. - -use std::rc::Rc; - -use crate::ast::{Operation, ServiceKind}; -use crate::compile::Catalog; -use crate::engine::Row; -use crate::http::{Method, Request, Response, RouteParams, Router, Status}; -use crate::shard::ShardCtx; - -use serde_json::{json, Value}; - -/// Build the fully-wired [`Router`] for one shard's worker thread. -pub fn router(ctx: Rc, catalog: &Catalog) -> Router { - let shard = ctx.id; - let n = ctx.n; - let mut r = Router::new() - .route(Method::Get, "/", move |_, _| { - Response::ok().json(&json!({ - "runtime": "wo", - "stage": 2, - "threads": n, - "shard": shard, - "notes": "REST CRUD for each `service rest` block. /healthz for liveness. LIVE subscribe in Stage 3." - })) - }) - .route(Method::Get, "/healthz", |_, _| Response::ok().text("ok")); - - for name in &catalog.order { - let t = catalog.get(name).expect("type present"); - for svc in &t.services { - if svc.kind != ServiceKind::Rest { continue; } - r = attach_rest(r, ctx.clone(), t.name.clone(), svc.path.clone(), &svc.expose, - &t.methods); - } - } - r -} - -fn attach_rest( - mut r: Router, - ctx: Rc, - ty: String, - path: String, - ops: &[Operation], - methods: &[crate::ast::MethodDecl], -) -> Router { - let id_path = format!("{path}/:id"); - - // Register literal sub-paths (`/live`, `/me`) BEFORE the `/:id` param - // route — the first matching pattern wins, so `:id` would otherwise - // swallow "live" / "me" and produce a 400 invalid-id response. - for op in ops { - match op { - Operation::Subscribe => { - r = r.route(Method::Get, &format!("{path}/live"), - |_, _| Response::status(Status::NOT_IMPLEMENTED).text("LIVE subscriptions arrive in Stage 3")); - } - Operation::Me => { - r = r.route(Method::Get, &format!("{path}/me"), - |_, _| Response::status(Status::NOT_IMPLEMENTED).text("session layer not yet implemented")); - } - _ => {} - } - } - - for op in ops { - match op { - Operation::List => { - let ctx = ctx.clone(); let ty = ty.clone(); - r = r.route(Method::Get, &path, move |req, params| list_h(&ctx, &ty, req, params)); - } - Operation::Create => { - let ctx = ctx.clone(); let ty = ty.clone(); - r = r.route(Method::Post, &path, move |req, params| create_h(&ctx, &ty, req, params)); - } - Operation::Get => { - let ctx = ctx.clone(); let ty = ty.clone(); - r = r.route(Method::Get, &id_path, move |req, params| get_h(&ctx, &ty, req, params)); - } - Operation::Update => { - let ctx = ctx.clone(); let ty = ty.clone(); - r = r.route(Method::Patch, &id_path, move |req, params| update_h(&ctx, &ty, req, params)); - } - Operation::Delete => { - let ctx = ctx.clone(); let ty = ty.clone(); - r = r.route(Method::Delete, &id_path, move |req, params| delete_h(&ctx, &ty, req, params)); - } - Operation::Subscribe | Operation::Me | Operation::Custom => {} - } - } - - // Class methods (plan 13b): every method of a class with a rest service - // is served as a row-scoped RPC. Registered after the CRUD `/:id` routes - // — the extra path segment makes patterns disjoint either way. - for m in methods { - let ctx = ctx.clone(); - let ty = ty.clone(); - let m = m.clone(); - r = r.route(Method::Post, &format!("{path}/:id/{}", m.name), - move |req, params| method_h(&ctx, &ty, &m, req, params)); - } - r -} - -// --- handlers --- -// -// Cross-shard results travel as `Result<_, String>` (anyhow::Error isn't -// guaranteed Send-friendly to reconstruct losslessly; the string is what we -// put in the HTTP body anyway). `run_on` returning `None` means the owning -// shard is gone — only during shutdown — and maps to 503. - -fn shard_gone() -> Response { - Response::status(Status::INTERNAL_SERVER_ERROR).text("owning shard unavailable") -} - -fn list_h(ctx: &Rc, ty: &str, req: &Request, _params: &RouteParams) -> Response { - // `?field=value` filters run through the engine's find_by — secondary - // indexes (`@table(index: ...)`) accelerate, scan is the fallback. - let filters = match query_filters(ctx, ty, req.query.as_deref()) { - Ok(f) => f, - Err(r) => return r, - }; - let ty_owned = ty.to_string(); - // Fan out to every shard, merge by id — the cross-shard read per 09b. - let per_shard: Vec, String>> = - ctx.fanout(move |e| e.find_by(&ty_owned, &filters).map_err(|e| e.to_string())); - let mut rows = Vec::new(); - for r in per_shard { - match r { - Ok(mut v) => rows.append(&mut v), - Err(e) => return Response::status(Status::INTERNAL_SERVER_ERROR).text(e), - } - } - rows.sort_by_key(|row| row.get("id").and_then(|v| v.as_i64()).unwrap_or(0)); - Response::ok().json(&json!(rows)) -} - -/// Parse `k=v&k2=v2` into equality filters. Fields must be stored columns -/// of the type (unknown → 400); values coerce int → bool → string. -fn query_filters( - ctx: &Rc, - ty: &str, - query: Option<&str>, -) -> Result, Response> { - let Some(q) = query.filter(|q| !q.is_empty()) else { return Ok(Vec::new()) }; - let mut out = Vec::new(); - let engine = ctx.engine.borrow(); - let t = engine.catalog().get(ty) - .ok_or_else(|| Response::status(Status::INTERNAL_SERVER_ERROR).text("type missing"))?; - for pair in q.split('&').filter(|s| !s.is_empty()) { - let (k, v) = pair.split_once('=').unwrap_or((pair, "")); - let (k, v) = (url_decode(k), url_decode(v)); - let stored = k == "id" || t.fields.iter().any(|f| f.name == k && matches!( - f.ty, - crate::ast::FieldTy::Scalar(_) | crate::ast::FieldTy::Union(_) - | crate::ast::FieldTy::Ref(_) - )); - if !stored { - return Err(Response::status(Status::BAD_REQUEST) - .text(format!("unknown filter field `{k}` for {ty}"))); - } - let val = if let Ok(n) = v.parse::() { json!(n) } - else if v == "true" { json!(true) } - else if v == "false" { json!(false) } - else { Value::String(v) }; - out.push((k, val)); - } - Ok(out) -} - -/// Minimal percent-decoding for query values (`%XX` and `+` → space). -fn url_decode(s: &str) -> String { - let b = s.as_bytes(); - let mut out = Vec::with_capacity(b.len()); - let mut i = 0; - while i < b.len() { - match b[i] { - b'%' if i + 2 < b.len() => { - let hex = |c: u8| (c as char).to_digit(16); - match (hex(b[i + 1]), hex(b[i + 2])) { - (Some(h), Some(l)) => { out.push((h * 16 + l) as u8); i += 3; } - _ => { out.push(b[i]); i += 1; } - } - } - b'+' => { out.push(b' '); i += 1; } - c => { out.push(c); i += 1; } - } - } - String::from_utf8_lossy(&out).into_owned() -} - -fn get_h(ctx: &Rc, ty: &str, _req: &Request, params: &RouteParams) -> Response { - let id = match parse_id(params) { - Ok(id) => id, - Err(r) => return r, - }; - let ty_owned = ty.to_string(); - let res = ctx.run_on(ctx.owner_of(id), move |e| { - e.get(&ty_owned, id).map_err(|e| e.to_string()) - }); - match res { - None => shard_gone(), - Some(Ok(Some(row))) => Response::ok().json(&json!(row)), - Some(Ok(None)) => Response::status(Status::NOT_FOUND).text(format!("no {ty} with id {id}")), - Some(Err(e)) => Response::status(Status::INTERNAL_SERVER_ERROR).text(e), - } -} - -fn create_h(ctx: &Rc, ty: &str, req: &Request, _params: &RouteParams) -> Response { - let body = match parse_json_body(req) { - Ok(v) => v, - Err(r) => return r, - }; - // Always local: the receiving shard mints from its own interleaved - // stride, so the row it creates is by construction a row it owns. - let res = ctx.engine.borrow_mut().create(ty, body); - match res { - Ok(row) => gate_if_staged(ctx, Response::status(Status::CREATED).json(&json!(row))), - Err(e) => Response::status(Status::BAD_REQUEST).text(e.to_string()), - } -} - -/// Group commit: a mutation that staged a WAL frame must not be acked until -/// the batch fsync — flag the response so the connection parks it. -fn gate_if_staged(ctx: &Rc, mut resp: Response) -> Response { - if ctx.engine.borrow_mut().take_staged() { - resp.gate = true; - } - resp -} - -fn update_h(ctx: &Rc, ty: &str, req: &Request, params: &RouteParams) -> Response { - let id = match parse_id(params) { - Ok(id) => id, - Err(r) => return r, - }; - let body = match parse_json_body(req) { - Ok(v) => v, - Err(r) => return r, - }; - let ty_owned = ty.to_string(); - let res = ctx.run_on(ctx.owner_of(id), move |e| { - e.update(&ty_owned, id, body).map_err(|e| e.to_string()) - }); - match res { - None => shard_gone(), - Some(Ok(Some(row))) => gate_if_staged(ctx, Response::ok().json(&json!(row))), - Some(Ok(None)) => Response::status(Status::NOT_FOUND).text(format!("no {ty} with id {id}")), - Some(Err(e)) => Response::status(Status::BAD_REQUEST).text(e), - } -} - -fn delete_h(ctx: &Rc, ty: &str, _req: &Request, params: &RouteParams) -> Response { - let id = match parse_id(params) { - Ok(id) => id, - Err(r) => return r, - }; - let ty_owned = ty.to_string(); - let res = ctx.run_on(ctx.owner_of(id), move |e| { - e.delete(&ty_owned, id).map_err(|e| e.to_string()) - }); - match res { - None => shard_gone(), - Some(Ok(true)) => gate_if_staged(ctx, Response::no_content()), - Some(Ok(false)) => Response::status(Status::NOT_FOUND).text(format!("no {ty} with id {id}")), - Some(Err(e)) => Response::status(Status::INTERNAL_SERVER_ERROR).text(e), - } -} - -/// Row-scoped method RPC (plan 13b): `POST /:id/` with a JSON -/// args object. Executes on the shard that owns the receiving row — inserts -/// the body performs mint on that shard, so everything a method writes it -/// also owns. The whole body commits as one WAL frame; failures roll back. -fn method_h( - ctx: &Rc, - ty: &str, - m: &crate::ast::MethodDecl, - req: &Request, - params: &RouteParams, -) -> Response { - let id = match parse_id(params) { - Ok(id) => id, - Err(r) => return r, - }; - let args = match parse_json_body(req) { - Ok(Value::Object(map)) => map, - Ok(Value::Null) => Default::default(), - Ok(_) => return Response::status(Status::BAD_REQUEST) - .text("method arguments must be a JSON object"), - Err(r) => return r, - }; - let ty_owned = ty.to_string(); - let m_owned = m.clone(); - let res = ctx.run_on(ctx.owner_of(id), move |e| { - crate::method::call(e, &ty_owned, id, &m_owned, &args) - .map_err(|err| (status_for(&err), err.to_string())) - }); - match res { - None => shard_gone(), - Some(Ok(v)) => gate_if_staged(ctx, Response::ok().json(&v)), - Some(Err((status, msg))) => Response::status(status).text(msg), - } -} - -fn status_for(err: &crate::method::MethodError) -> Status { - use crate::method::MethodError::*; - match err { - NoSuchRow => Status::NOT_FOUND, - BadArgs(_) => Status::BAD_REQUEST, - Abort(_) => Status::CONFLICT, - Exec(_) => Status::INTERNAL_SERVER_ERROR, - } -} - -fn parse_id(params: &RouteParams) -> Result { - params.get("id") - .and_then(|s| s.parse::().ok()) - .ok_or_else(|| Response::status(Status::BAD_REQUEST).text("invalid id")) -} - -fn parse_json_body(req: &Request) -> Result { - if req.body.is_empty() { - return Ok(Value::Object(Default::default())); - } - serde_json::from_slice::(&req.body) - .map_err(|e| Response::status(Status::BAD_REQUEST).text(format!("invalid JSON: {e}"))) -} - -/// Format the endpoint banner the CLI prints on startup. -pub fn describe_routes(catalog: &Catalog) -> String { - let mut out = String::new(); - out.push_str(" GET / runtime info\n"); - out.push_str(" GET /healthz liveness\n"); - for name in &catalog.order { - let t = catalog.get(name).unwrap(); - for svc in &t.services { - if svc.kind != ServiceKind::Rest { continue; } - for op in &svc.expose { - let (m, p, n) = match op { - Operation::List => ("GET ", svc.path.clone(), format!("list {name}")), - Operation::Get => ("GET ", format!("{}/:id", svc.path), format!("get {name}")), - Operation::Create => ("POST ", svc.path.clone(), format!("create {name}")), - Operation::Update => ("PATCH ", format!("{}/:id", svc.path), format!("update {name}")), - Operation::Delete => ("DELETE", format!("{}/:id", svc.path), format!("delete {name}")), - Operation::Subscribe => ("WS ", format!("{}/live", svc.path), format!("subscribe {name} (Stage 3)")), - Operation::Me => ("GET ", format!("{}/me", svc.path), "(Stage 3)".into()), - Operation::Custom => continue, - }; - out.push_str(&format!(" {m} {p:<30} {n}\n")); - } - for m in &t.methods { - let p = format!("{}/:id/{}", svc.path, m.name); - let mode = match m.txn { - crate::ast::TxnMode::None => "", - _ => " in txn", - }; - out.push_str(&format!(" POST {p:<30} method {}.{}{mode}\n", name, m.name)); - } - // @table storage configuration, when it says anything non-default. - if t.storage_name != *name || !t.indexes.is_empty() { - let mut cfg = format!("table \"{}\"", t.storage_name); - if !t.indexes.is_empty() { - let idx = t.indexes.iter() - .map(|c| c.join("+")) - .collect::>().join(", "); - cfg.push_str(&format!(", index [{idx}]")); - } - out.push_str(&format!(" {:<30} @table {cfg}\n", name)); - } - } - } - out -} - -#[cfg(test)] -mod tests { - use super::*; - use crate::compile::Catalog; - use crate::engine::Engine; - use crate::parser::parse; - use crate::shard::ShardBus; - - /// Single-shard context: `run_on` is always local, `fanout` is just us — - /// exactly the WO_THREADS=1 production shape. - fn build(src: &str) -> (Rc, Router) { - let cat = Catalog::from_schemas(vec![parse(src).unwrap()]).unwrap(); - let bus = ShardBus::new(1).unwrap(); - let ctx = ShardCtx::new(0, 1, Engine::for_shard(cat.clone(), 0, 1), bus); - let r = router(ctx.clone(), &cat); - (ctx, r) - } - - fn req(method: Method, path: &str, body: &[u8]) -> Request { - Request { - method, - path: path.into(), - query: None, - headers: Default::default(), - body: body.to_vec(), - keep_alive: true, - } - } - - #[test] - fn router_serves_crud_for_a_service_rest_block() { - let (_ctx, r) = build(r#" -type Article { id: Id - title: Text - service rest "/api/articles" expose list, get, create, update, delete } -"#); - - // empty list - let resp = r.dispatch(&req(Method::Get, "/api/articles", b"")); - assert_eq!(resp.status.0, 200); - assert_eq!(resp.body, b"[]"); - - // create - let resp = r.dispatch(&req(Method::Post, "/api/articles", br#"{"title":"hello"}"#)); - assert_eq!(resp.status.0, 201); - - // get by id - let resp = r.dispatch(&req(Method::Get, "/api/articles/1", b"")); - assert_eq!(resp.status.0, 200); - - // delete - let resp = r.dispatch(&req(Method::Delete, "/api/articles/1", b"")); - assert_eq!(resp.status.0, 204); - - // 404 after delete - let resp = r.dispatch(&req(Method::Get, "/api/articles/1", b"")); - assert_eq!(resp.status.0, 404); - } - - #[test] - fn unexposed_method_yields_405() { - let (_ctx, r) = build(r#" -type Tag { id: Id - label: Text - service rest "/api/tags" expose list, get } -"#); - let resp = r.dispatch(&req(Method::Post, "/api/tags", br#"{"label":"rust"}"#)); - assert_eq!(resp.status.0, 405); - } - - #[test] - fn live_subscribe_is_501() { - let (_ctx, r) = build(r#" -type Article { id: Id - title: Text - service rest "/api/articles" expose list, subscribe } -"#); - let resp = r.dispatch(&req(Method::Get, "/api/articles/live", b"")); - assert_eq!(resp.status.0, 501); - } - - const PRICING: &str = r#" -@table(name: "prices", index: [product]) -class Price { - id: Id - product: ref Product - amount: Money - service rest "/api/prices" expose list -} - -class Product { - id: Id - sku: SKU @unique - name: Text - prices: multi Price - - fn current_price() -> Money in txn { - return latest(self.prices).amount; - } - - fn set_price(amount: Money) in txn { - assert amount > 0 otherwise abort "price must be positive" - insert Price { product: self.id, amount: amount }; - } - - service rest "/api/products" expose list, get, create -} -"#; - - #[test] - fn class_methods_serve_rpc_routes() { - let (_ctx, r) = build(PRICING); - - let resp = r.dispatch(&req(Method::Post, "/api/products", br#"{"sku":"SKU-1","name":"Gadget"}"#)); - assert_eq!(resp.status.0, 201); - - // The 13b exit criterion: POST :id/set_price inserts a Price atomically… - let resp = r.dispatch(&req(Method::Post, "/api/products/1/set_price", br#"{"amount": 4999}"#)); - assert_eq!(resp.status.0, 200, "{}", String::from_utf8_lossy(&resp.body)); - - // …and current_price returns it. - let resp = r.dispatch(&req(Method::Post, "/api/products/1/current_price", b"")); - assert_eq!(resp.status.0, 200); - assert_eq!(resp.body, b"4999"); - - // The Price row is a real, listable row. - let resp = r.dispatch(&req(Method::Get, "/api/prices", b"")); - let rows: Vec = serde_json::from_slice(&resp.body).unwrap(); - assert_eq!(rows.len(), 1); - assert_eq!(rows[0]["product"], 1); - assert_eq!(rows[0]["amount"], 4999); - } - - #[test] - fn method_errors_map_to_http_statuses() { - let (_ctx, r) = build(PRICING); - r.dispatch(&req(Method::Post, "/api/products", br#"{"sku":"S","name":"N"}"#)); - - // Unknown row → 404. - let resp = r.dispatch(&req(Method::Post, "/api/products/99/set_price", br#"{"amount": 1}"#)); - assert_eq!(resp.status.0, 404); - - // Missing argument → 400. - let resp = r.dispatch(&req(Method::Post, "/api/products/1/set_price", b"{}")); - assert_eq!(resp.status.0, 400); - - // Aborting method → 409, and the transaction rolled back. - let resp = r.dispatch(&req(Method::Post, "/api/products/1/set_price", br#"{"amount": 0}"#)); - assert_eq!(resp.status.0, 409); - let resp = r.dispatch(&req(Method::Get, "/api/prices", b"")); - assert_eq!(resp.body, b"[]", "aborted method must leave no rows"); - - // GET on a method route → 405 (route exists, wrong verb). - let resp = r.dispatch(&req(Method::Get, "/api/products/1/set_price", b"")); - assert_eq!(resp.status.0, 405); - } - - fn req_q(path: &str, query: &str) -> Request { - Request { - method: Method::Get, - path: path.into(), - query: Some(query.to_string()), - headers: Default::default(), - body: vec![], - keep_alive: true, - } - } - - #[test] - fn list_query_filter_is_index_backed() { - let (_ctx, r) = build(PRICING); - r.dispatch(&req(Method::Post, "/api/products", br#"{"sku":"A","name":"A"}"#)); - r.dispatch(&req(Method::Post, "/api/products", br#"{"sku":"B","name":"B"}"#)); - r.dispatch(&req(Method::Post, "/api/products/1/set_price", br#"{"amount": 100}"#)); - r.dispatch(&req(Method::Post, "/api/products/1/set_price", br#"{"amount": 200}"#)); - r.dispatch(&req(Method::Post, "/api/products/2/set_price", br#"{"amount": 999}"#)); - - // Unfiltered list unchanged. - let resp = r.dispatch(&req(Method::Get, "/api/prices", b"")); - let all: Vec = serde_json::from_slice(&resp.body).unwrap(); - assert_eq!(all.len(), 3); - - // ?product=1 → only that product's prices, via the (product) index. - let resp = r.dispatch(&req_q("/api/prices", "product=1")); - assert_eq!(resp.status.0, 200); - let rows: Vec = serde_json::from_slice(&resp.body).unwrap(); - assert_eq!(rows.len(), 2); - assert!(rows.iter().all(|r| r["product"] == 1)); - - // Multiple filters combine (equality AND). - let resp = r.dispatch(&req_q("/api/prices", "product=1&amount=200")); - let rows: Vec = serde_json::from_slice(&resp.body).unwrap(); - assert_eq!(rows.len(), 1); - assert_eq!(rows[0]["amount"], 200); - - // Unknown field → 400. - let resp = r.dispatch(&req_q("/api/prices", "nope=1")); - assert_eq!(resp.status.0, 400); - - // Filtering on a non-indexed stored column falls back to scan. - let resp = r.dispatch(&req_q("/api/prices", "amount=999")); - let rows: Vec = serde_json::from_slice(&resp.body).unwrap(); - assert_eq!(rows.len(), 1); - assert_eq!(rows[0]["product"], 2); - } -} diff --git a/crates/rt/src/shard.rs b/crates/rt/src/shard.rs deleted file mode 100644 index 2edb8be..0000000 --- a/crates/rt/src/shard.rs +++ /dev/null @@ -1,259 +0,0 @@ -//! The shard bus — plan 09b (`docs/plan/09-concurrency-scaleout.md`). -//! -//! With 09b every worker owns its own [`Engine`] — `Arc>` is -//! gone. A request that lands on shard K (kernel `SO_REUSEPORT` hash) but -//! targets a row owned by shard J ships a **job** — a boxed closure — to J's -//! mailbox, wakes J's event loop through its mail `eventfd`, and waits for -//! the reply. Two rules keep this deadlock-free: -//! -//! 1. **Jobs never block.** A job is a pure local engine operation on the -//! owning thread; it cannot itself wait on another shard. -//! 2. **Waiters keep serving.** While shard K waits for J's reply it pumps -//! its own inbox, so J (or anyone) waiting on K is never starved. -//! -//! Row → owner mapping is the interleaved-id rule (`Engine::for_shard`): -//! shard t mints ids t+1, t+1+n, … so `owner(id) = (id-1) % n` with zero -//! coordination. Creates are always local (the receiving shard mints from -//! its own stride); reads/updates/deletes hop at most once; lists fan out -//! to every shard and merge. The C proving ground for the wake mechanism is -//! `prototypes/wo-rt-c` (eventfd broadcast); the mailbox-per-thread design -//! is plan 09 decision 2 and 09d's one-message-per-thread fan-out shape. - -use std::cell::RefCell; -use std::io; -use std::rc::Rc; -use std::sync::mpsc::{channel, Receiver, RecvTimeoutError, Sender}; -use std::sync::{Arc, Mutex}; -use std::time::Duration; - -use crate::engine::Engine; -use crate::runtime::EventFd; - -/// A unit of work shipped to the owning shard. Runs against that shard's -/// engine on that shard's thread; replies through whatever channel it -/// captured. -pub type Job = Box; - -/// Created once in `main`, shared by every worker: each shard's job sender -/// and mail eventfd. Inboxes are taken (once each) by their owning worker. -pub struct ShardBus { - senders: Vec>, - wakes: Vec, - inboxes: Mutex>>>, -} - -impl ShardBus { - pub fn new(n: usize) -> io::Result> { - let mut senders = Vec::with_capacity(n); - let mut inboxes = Vec::with_capacity(n); - let mut wakes = Vec::with_capacity(n); - for _ in 0..n { - let (tx, rx) = channel(); - senders.push(tx); - inboxes.push(Some(rx)); - wakes.push(EventFd::new()?); - } - Ok(Arc::new(Self { senders, wakes, inboxes: Mutex::new(inboxes) })) - } - - /// The owning worker claims its inbox at startup. Panics on double-take — - /// that would be a wiring bug, not a runtime condition. - pub fn take_inbox(&self, shard: usize) -> Receiver { - self.inboxes.lock().unwrap()[shard].take().expect("inbox already taken") - } - - pub fn mail_fd(&self, shard: usize) -> &EventFd { &self.wakes[shard] } -} - -/// Per-worker handle: this shard's engine plus the bus. Deliberately `!Send` -/// (`Rc`/`RefCell`) — it exists on exactly one thread, which is the point. -pub struct ShardCtx { - pub id: usize, - pub n: usize, - pub engine: Rc>, - inbox: Receiver, - bus: Arc, - /// Connection unparks discovered while pumping inside a handler — the - /// worker loop takes and applies them after the handler returns. - unparks: RefCell>, -} - -impl ShardCtx { - pub fn new(id: usize, n: usize, engine: Engine, bus: Arc) -> Rc { - let inbox = bus.take_inbox(id); - Rc::new(Self { id, n, engine: Rc::new(RefCell::new(engine)), inbox, bus, - unparks: RefCell::new(Vec::new()) }) - } - - /// Flush + reap this shard's group-commit WAL. Reply parks release - /// immediately; connection parks queue for the worker loop. Called at - /// tick end, on ring-fd events, AND from every pump-wait — a waiter - /// that didn't flush its own batch would deadlock with a peer waiting - /// on it (cross-shard mutual commit). - pub fn wal_pump(&self) { - let completed = { - let mut e = self.engine.borrow_mut(); - e.wal_flush(); - e.wal_complete() - }; - if let Some((ok, acks)) = completed { - for p in acks { - match p { - crate::wal::Parked::Reply(cb) => { - // A failed batch drops the callback: the requester's - // channel disconnects → 500, never a false ack. - if ok { cb() } - } - crate::wal::Parked::Conn { fd, gen } => { - self.unparks.borrow_mut().push((fd, gen, ok)); - } - } - } - } - } - - /// Worker loop: take any connection unparks the pumps produced. - pub fn take_unparks(&self) -> Vec<(std::os::unix::io::RawFd, u64, bool)> { - std::mem::take(&mut self.unparks.borrow_mut()) - } - - /// Which shard owns a row id, per the interleaved-mint rule. - pub fn owner_of(&self, row_id: i64) -> usize { - ((row_id - 1).rem_euclid(self.n as i64)) as usize - } - - /// Execute every queued job against the local engine. Called from the - /// event loop on a mail-eventfd event, and from the wait loops below. - pub fn drain_inbox(&self) { - while let Ok(job) = self.inbox.try_recv() { - job(&mut self.engine.borrow_mut()); - } - } - - /// Run `f` against the engine that owns `owner` — locally if that's us, - /// else ship it and wait, pumping our own inbox so peers waiting on us - /// make progress. Returns `None` only if the owner is gone (shutdown). - pub fn run_on(&self, owner: usize, f: F) -> Option - where - R: Send + 'static, - F: FnOnce(&mut Engine) -> R + Send + 'static, - { - if owner == self.id { - return Some(f(&mut self.engine.borrow_mut())); - } - let (tx, rx) = channel(); - // Group commit: if the job staged a WAL frame on the owner, its - // reply parks on the owner's batch and is sent on the fsync CQE — - // so our requester-side response leaves only after durability. - let job: Job = Box::new(move |e| { - let r = f(e); - if e.take_staged() { - e.park_reply(Box::new(move || { let _ = tx.send(r); })); - } else { - let _ = tx.send(r); - } - }); - if self.bus.senders[owner].send(job).is_err() { - return None; - } - let _ = self.bus.wakes[owner].write(1); - loop { - match rx.recv_timeout(Duration::from_micros(100)) { - Ok(r) => return Some(r), - Err(RecvTimeoutError::Timeout) => { self.drain_inbox(); self.wal_pump(); } - Err(RecvTimeoutError::Disconnected) => return None, - } - } - } - - /// Run `f` on every shard (self included) and collect the results. - /// Cross-shard reads — `list` — are the fan-out-and-merge case. - pub fn fanout(&self, f: F) -> Vec - where - R: Send + 'static, - F: Fn(&mut Engine) -> R + Send + Sync + Clone + 'static, - { - let (tx, rx) = channel(); - let mut remote = 0usize; - for (t, sender) in self.bus.senders.iter().enumerate() { - if t == self.id { continue; } - let tx = tx.clone(); - let f = f.clone(); - let job: Job = Box::new(move |e| { let _ = tx.send(f(e)); }); - if sender.send(job).is_ok() { - let _ = self.bus.wakes[t].write(1); - remote += 1; - } - } - let mut out = Vec::with_capacity(remote + 1); - out.push(f(&mut self.engine.borrow_mut())); - while out.len() < remote + 1 { - match rx.recv_timeout(Duration::from_micros(100)) { - Ok(r) => out.push(r), - Err(RecvTimeoutError::Timeout) => { self.drain_inbox(); self.wal_pump(); } - Err(RecvTimeoutError::Disconnected) => break, - } - } - out - } -} - -#[cfg(test)] -mod tests { - use super::*; - use crate::compile::Catalog; - use crate::parser::parse; - use serde_json::json; - - fn catalog() -> Catalog { - Catalog::from_schemas(vec![parse( - r#"type Note { id: Id - title: Text - service rest "/api/notes" expose list, get, create }"#, - ).unwrap()]).unwrap() - } - - #[test] - fn interleaved_ids_map_back_to_their_shard() { - let cat = catalog(); - let bus = ShardBus::new(3).unwrap(); - let ctx0 = ShardCtx::new(0, 3, Engine::for_shard(cat.clone(), 0, 3), bus.clone()); - let ctx1 = ShardCtx::new(1, 3, Engine::for_shard(cat.clone(), 1, 3), bus.clone()); - - let a = ctx0.engine.borrow_mut().create("Note", json!({"title":"a"})).unwrap(); - let b = ctx0.engine.borrow_mut().create("Note", json!({"title":"b"})).unwrap(); - let c = ctx1.engine.borrow_mut().create("Note", json!({"title":"c"})).unwrap(); - let (a, b, c) = (a["id"].as_i64().unwrap(), b["id"].as_i64().unwrap(), c["id"].as_i64().unwrap()); - assert_eq!((a, b, c), (1, 4, 2)); - assert_eq!(ctx0.owner_of(a), 0); - assert_eq!(ctx0.owner_of(b), 0); - assert_eq!(ctx0.owner_of(c), 1); - } - - #[test] - fn cross_shard_job_round_trips() { - let cat = catalog(); - let bus = ShardBus::new(2).unwrap(); - let ctx1 = ShardCtx::new(1, 2, Engine::for_shard(cat.clone(), 1, 2), bus.clone()); - - // Shard 1's thread: serve jobs (one drain after the send below). - let bus2 = bus.clone(); - let cat2 = cat.clone(); - let t = std::thread::spawn(move || { - // shard 0 lives on this thread - let ctx0 = ShardCtx::new(0, 2, Engine::for_shard(cat2, 0, 2), bus2); - ctx0.engine.borrow_mut().create("Note", json!({"title":"on-zero"})).unwrap(); - // serve until the job arrives - for _ in 0..1000 { - ctx0.drain_inbox(); - std::thread::sleep(Duration::from_micros(200)); - } - }); - - // From shard 1, read the row owned by shard 0 (id 1). - let row = ctx1.run_on(0, |e| e.get("Note", 1).unwrap()).flatten(); - assert_eq!(row.unwrap()["title"], "on-zero"); - drop(ctx1); - t.join().unwrap(); - } -} diff --git a/crates/rt/src/token.rs b/crates/rt/src/token.rs deleted file mode 100644 index 1d80092..0000000 --- a/crates/rt/src/token.rs +++ /dev/null @@ -1,139 +0,0 @@ -//! Token definitions for the `.wo` lexer. - -use std::fmt; - -#[derive(Debug, Clone, PartialEq, Eq)] -pub enum Kind { - // literals - Ident(String), - Str(String), - Int(i64), - Param(String), // $name - - // keywords (schema + query layer) - KwType, - KwClass, - KwRef, - KwMulti, - KwVia, - KwBacklink, - KwLink, - KwService, - KwRest, - KwGraphql, - KwNative, - KwExpose, - KwPolicy, - KwFor, - KwRole, - KwWhen, - KwAnyone, - KwOn, - KwDo, - KwSet, - KwCall, - KwEmit, - KwEnqueue, - KwAssert, - KwOtherwise, - KwAbort, - KwReturn, - KwReturning, - KwAs, - KwFn, - KwIn, - KwTxn, - KwSnapshot, - KwSerializable, - KwBegin, - KwCommit, - KwRollback, - KwSavepoint, - KwTo, - KwLive, - KwInsert, - KwInto, - KwValues, - KwUpdate, - KwDelete, - KwSelect, - KwFrom, - KwWhere, - KwMatch, - KwCreate, - KwLet, - KwIf, - KwElse, - KwFor1, // the other `for` — for/each loop (disambiguated at parse time) - KwEach, - KwContains, - KwAnd, - KwOr, - KwNot, - KwTrue, - KwFalse, - KwNull, - KwTest, - KwMain, - KwApp, - KwStartup, - - // block markers - HashHash(String), // `##sql`, `##doc`, `##graph`, `##ui`, `##app`, `##policy`, `##service`, `##logic` - Hash(String), // `#table-name` - - // punctuation - LBrace, // { - RBrace, // } - LParen, // ( - RParen, // ) - LBracket, // [ - RBracket, // ] - Comma, - Semicolon, - Colon, - Dot, - DotDot, // .. - DotStar, // .* used in `line_items.*.qty` - Question, // ? - At, // @ - Pipe, // | - Arrow, // -> - FatArrow, // => - Dash, // - - Plus, - Star, - Slash, - Percent, - Eq, // = - EqEq, // == - NotEq, // != - Lt, LtEq, - Gt, GtEq, - PlusEq, MinusEq, - - // meta - Newline, - End, -} - -#[derive(Debug, Clone)] -pub struct Token { - pub kind: Kind, - pub line: u32, - pub col: u32, -} - -impl fmt::Display for Kind { - fn fmt(&self, f: &mut fmt::Formatter<'_>) -> fmt::Result { - match self { - Kind::Ident(s) => write!(f, "ident({s})"), - Kind::Str(s) => write!(f, "\"{s}\""), - Kind::Int(i) => write!(f, "{i}"), - Kind::Param(s) => write!(f, "${s}"), - Kind::HashHash(s) => write!(f, "##{s}"), - Kind::Hash(s) => write!(f, "#{s}"), - k => write!(f, "{k:?}"), - } - } -} diff --git a/crates/rt/src/wal.rs b/crates/rt/src/wal.rs deleted file mode 100644 index e166523..0000000 --- a/crates/rt/src/wal.rs +++ /dev/null @@ -1,337 +0,0 @@ -//! Per-shard write-ahead log — plan 09c (`docs/plan/09-concurrency-scaleout.md`) -//! + the durability core of plan 11, ported from the proven C sequence -//! (`prototypes/wo-rt-c` phases D/E, `docs/plan/exploration/c-runtime/00-plan.md`). -//! -//! One `shard-.rwal` per worker. Frame format (identical shape to the C -//! prototype): `u32 len | u32 crc32(payload) | payload | u32 COMMIT` — a -//! record replays whole or not at all; a torn tail fails CRC/trailer -//! validation and is truncated. Payloads are JSON-serialized [`WalRec`]s. -//! -//! Dual-write order (the C crash-under-load test's hard-won lesson): the -//! engine applies to RAM, appends the frame, `fdatasync`s, and only then -//! returns — so the HTTP ack (written after the handler returns, including -//! for cross-shard jobs whose reply follows the owner's engine call) is -//! always behind the fsync. Group commit (one fsync per loop tick, acks -//! parked on the completion) is deliberately deferred to the io_uring port — -//! doing it on the epoll loop would reopen the exact ack-before-fsync race -//! the C phase-F bench caught. One fsync per commit is slower and correct. -//! -//! Boot: `Wal::open_and_replay` walks the log into the shard's engine — -//! parallel across workers, before any accept is armed. No snapshots yet -//! (the WAL grows unbounded; compaction is phase 11 proper). - -use std::fs::{File, OpenOptions}; -use std::io::{self, Read, Seek, SeekFrom, Write}; -use std::os::unix::io::AsRawFd; -use std::path::Path; - -use serde::{Deserialize, Serialize}; -use serde_json::Value; - -use crate::engine::{Engine, Row}; - -const WAL_COMMIT: u32 = 0xC0FF_EE42; -const WAL_PREALLOC: i64 = 4 * 1024 * 1024; -const MAX_FRAME: u32 = 16 * 1024 * 1024; - -/// One logged mutation. `Create` carries the FULL post-default row (id, -/// timestamps included) so replay is byte-exact; `Update` carries the merge -/// body (merge is deterministic in log order). `Txn` bundles every mutation -/// of one method call (plan 13b) into a single frame: the frame's CRC + -/// trailer make it replay whole-or-not-at-all, so a crash mid-method can -/// never leave a partial method on disk. -#[derive(Debug, Serialize, Deserialize)] -#[serde(tag = "op", rename_all = "snake_case")] -pub enum WalRec { - Create { ty: String, row: Row }, - Update { ty: String, id: i64, body: Value }, - Delete { ty: String, id: i64 }, - Txn { recs: Vec }, -} - -#[derive(Debug)] -pub struct Wal { - file: File, -} - -/// An acknowledgment parked until its batch's fsync CQE (group commit). -/// `Conn` carries the C-proven generation stamp — kernel fds get reused, and -/// releasing by bare fd would ack a NEW connection's commit before ITS batch -/// is durable (the phase-F ABA bug, prevented by construction here). -pub enum Parked { - /// A local connection whose response waits in its write buffer. - Conn { fd: std::os::unix::io::RawFd, gen: u64 }, - /// A cross-shard reply — the owner runs this to release the requester. - Reply(Box), -} - -/// Group-commit WAL (io_uring): mutations STAGE frames + park their acks; -/// once per loop tick the worker flushes the staging buffer as one `WRITE` -/// SQE hard-linked to one `FSYNC` SQE; the fsync completion releases every -/// parked ack in the batch. Double-buffered — while a batch is in flight -/// (its buffer pinned for the kernel), new commits stage into the twin. -impl std::fmt::Debug for WalGroup { - fn fmt(&self, f: &mut std::fmt::Formatter<'_>) -> std::fmt::Result { f.write_str("WalGroup") } -} - -pub struct WalGroup { - file: File, - ring: crate::runtime::Uring, - offset: u64, - staging: Vec, - parked: Vec, - inflight: Option<(Vec, Vec)>, - /// Batch sequence — encoded into user_data so a write CQE and an fsync - /// CQE can never be attributed to the wrong batch. - seq: u64, - got_write: Option, - got_fsync: Option, -} - -impl WalGroup { - /// Wrap a replayed [`Wal`] (same file, offset at the validated tail). - pub fn new(wal: Wal, ring: crate::runtime::Uring) -> io::Result { - let mut file = wal.file; - let offset = file.seek(SeekFrom::Current(0))?; - let _ = file.seek(SeekFrom::Start(offset)); - Ok(Self { file, ring, offset, staging: Vec::new(), parked: Vec::new(), inflight: None, - seq: 0, got_write: None, got_fsync: None }) - } - - pub fn ring_fd(&self) -> std::os::unix::io::RawFd { self.ring.as_raw_fd() } - - /// Stage one record into the active batch. RAM is already applied; the - /// ack must now be parked (see [`Parked`]) until this batch fsyncs. - pub fn stage(&mut self, rec: &WalRec) -> io::Result<()> { - let payload = serde_json::to_vec(rec).map_err(io::Error::other)?; - self.staging.extend_from_slice(&(payload.len() as u32).to_le_bytes()); - self.staging.extend_from_slice(&crc32(&payload).to_le_bytes()); - self.staging.extend_from_slice(&payload); - self.staging.extend_from_slice(&WAL_COMMIT.to_le_bytes()); - Ok(()) - } - - pub fn park(&mut self, p: Parked) { - self.parked.push(p); - } - - /// End-of-tick: if commits are staged and no batch is in flight, submit - /// the whole batch as WRITE→FSYNC linked SQEs — one syscall. - pub fn flush(&mut self) -> io::Result<()> { - if self.inflight.is_some() || self.staging.is_empty() { return Ok(()); } - let buf = std::mem::take(&mut self.staging); - let acks = std::mem::take(&mut self.parked); - let fd = self.file.as_raw_fd(); - // user_data = (batch_seq << 1) | op-bit — CQEs are matched to THIS - // batch only; a stale completion can never release the wrong acks. - let ud_write = self.seq << 1; - let ud_fsync = (self.seq << 1) | 1; - // SAFETY: `buf` moves into `inflight` and stays pinned until the CQE. - self.ring.push_write(fd, &buf, self.offset, true, ud_write); - self.ring.push_fsync(fd, ud_fsync); - self.ring.submit()?; - self.inflight = Some((buf, acks)); - self.got_write = None; - self.got_fsync = None; - Ok(()) - } - - /// Ring-fd readable: reap completions. Returns `Some((ok, acks))` only - /// when BOTH the write and fsync CQEs of the in-flight batch have - /// arrived — they routinely land in different ticks on real disks, and - /// releasing on the write CQE alone would ack before durability. - /// `ok = false` means short write or failed fsync: the caller must DROP - /// the acks (close the connections), never release them. - pub fn complete(&mut self) -> Option<(bool, Vec)> { - for (ud, res) in self.ring.pop_cqes() { - if ud >> 1 != self.seq { continue; } // not this batch (stale/corrupt) - if ud & 1 == 0 { self.got_write = Some(res); } - else { self.got_fsync = Some(res); } - } - if self.inflight.is_none() { return None; } - let (Some(wr), Some(fr)) = (self.got_write, self.got_fsync) else { return None; }; - let (buf, acks) = self.inflight.take()?; - self.got_write = None; - self.got_fsync = None; - self.seq += 1; - let mut ok = true; - if wr != buf.len() as i32 { - eprintln!("[wo] wal: short write {wr} != {}", buf.len()); - ok = false; - } - if fr < 0 { - eprintln!("[wo] wal: fsync failed ({fr})"); - ok = false; - } - if ok { self.offset += buf.len() as u64; } - Some((ok, acks)) - } -} - -impl Wal { - /// Open (creating if absent) the shard's log, replay every valid frame - /// into `engine`, truncate any torn tail, and return a writer positioned - /// at the validated end. Returns `(wal, replayed_records)`. - pub fn open_and_replay(path: &Path, engine: &mut Engine) -> io::Result<(Self, usize)> { - let mut file = OpenOptions::new().read(true).write(true).create(true).open(path)?; - unsafe { libc::fallocate(file.as_raw_fd(), 0, 0, WAL_PREALLOC) }; - - let mut buf = Vec::new(); - file.read_to_end(&mut buf)?; - - let mut off = 0usize; - let mut recs = 0usize; - loop { - let Some(frame) = read_frame(&buf, off) else { break }; - let (payload, next) = frame; - match serde_json::from_slice::(payload) { - Ok(rec) => engine.replay(&rec), - Err(e) => { - eprintln!("[wo] wal {}: undecodable record at byte {off} ({e}) — truncating", path.display()); - break; - } - } - recs += 1; - off = next; - } - - // Resume appends at the validated tail; drop torn bytes. - file.set_len(off as u64)?; - unsafe { libc::fallocate(file.as_raw_fd(), 0, 0, WAL_PREALLOC.max(off as i64)) }; - file.seek(SeekFrom::Start(off as u64))?; - Ok((Self { file }, recs)) - } - - /// Append one record and make it durable. The caller's mutation is only - /// allowed to stand (and its response to leave) after this returns Ok. - pub fn append(&mut self, rec: &WalRec) -> io::Result<()> { - let payload = serde_json::to_vec(rec).map_err(io::Error::other)?; - let mut frame = Vec::with_capacity(payload.len() + 12); - frame.extend_from_slice(&(payload.len() as u32).to_le_bytes()); - frame.extend_from_slice(&crc32(&payload).to_le_bytes()); - frame.extend_from_slice(&payload); - frame.extend_from_slice(&WAL_COMMIT.to_le_bytes()); - self.file.write_all(&frame)?; - self.file.sync_data()?; // the D in ACID — ack ordering lives here - Ok(()) - } -} - -/// Validate and slice one frame at `off`. `None` = clean end or torn tail. -fn read_frame(buf: &[u8], off: usize) -> Option<(&[u8], usize)> { - let u32_at = |o: usize| -> Option { - buf.get(o..o + 4).map(|b| u32::from_le_bytes(b.try_into().unwrap())) - }; - let len = u32_at(off)?; - if len == 0 || len > MAX_FRAME { return None; } // preallocated zeros / garbage - let len = len as usize; - let crc = u32_at(off + 4)?; - let payload = buf.get(off + 8..off + 8 + len)?; - let trailer = u32_at(off + 8 + len)?; - if crc32(payload) != crc || trailer != WAL_COMMIT { return None; } - Some((payload, off + 8 + len + 4)) -} - -/// Hand-rolled CRC32 (poly 0xEDB88320) — same algorithm as the C prototype; -/// no external crate. -fn crc32(data: &[u8]) -> u32 { - let mut c: u32; - let mut table = [0u32; 256]; - for (i, t) in table.iter_mut().enumerate() { - c = i as u32; - for _ in 0..8 { - c = if c & 1 != 0 { 0xEDB8_8320 ^ (c >> 1) } else { c >> 1 }; - } - *t = c; - } - let mut crc = 0xFFFF_FFFFu32; - for &b in data { - crc = table[((crc ^ b as u32) & 0xFF) as usize] ^ (crc >> 8); - } - crc ^ 0xFFFF_FFFF -} - -#[cfg(test)] -mod tests { - use super::*; - use crate::compile::Catalog; - use crate::parser::parse; - use serde_json::json; - - fn catalog() -> Catalog { - Catalog::from_schemas(vec![parse( - r#"@table(index: [title]) - type Note { id: Id - title: Text - service rest "/api/notes" expose list, get, create, update, delete }"#, - ).unwrap()]).unwrap() - } - - fn tmp(name: &str) -> std::path::PathBuf { - let p = std::env::temp_dir().join(format!("wo-wal-test-{}-{name}", std::process::id())); - let _ = std::fs::remove_file(&p); - p - } - - #[test] - fn replay_restores_creates_updates_deletes_and_id_highwater() { - let path = tmp("roundtrip"); - { - let mut e = Engine::for_shard(catalog(), 0, 2); - let (wal, n) = Wal::open_and_replay(&path, &mut e).unwrap(); - assert_eq!(n, 0); - e.attach_wal(wal); - e.create("Note", json!({"title":"a"})).unwrap(); // id 1 - e.create("Note", json!({"title":"b"})).unwrap(); // id 3 - e.update("Note", 1, json!({"title":"a2"})).unwrap(); - e.create("Note", json!({"title":"c"})).unwrap(); // id 5 - e.delete("Note", 3).unwrap(); - } - // Fresh engine, replay from disk — the "first load". - let mut e = Engine::for_shard(catalog(), 0, 2); - let (wal, n) = Wal::open_and_replay(&path, &mut e).unwrap(); - assert_eq!(n, 5); - let rows = e.list("Note").unwrap(); - let ids: Vec = rows.iter().map(|r| r["id"].as_i64().unwrap()).collect(); - assert_eq!(ids, vec![1, 5]); - assert_eq!(rows[0]["title"], "a2"); - // Replay went through row_insert/row_remove — the secondary index is - // rebuilt: the update moved id 1 from "a" to "a2", the delete cleared - // id 3's entry. - assert_eq!(e.find_by("Note", &[("title".into(), json!("a2"))]).unwrap().len(), 1); - assert!(e.find_by("Note", &[("title".into(), json!("a"))]).unwrap().is_empty()); - assert!(e.find_by("Note", &[("title".into(), json!("b"))]).unwrap().is_empty()); - // id high-water restored: the next mint must not collide (and must - // keep the shard-0-of-2 stride: odd ids). - e.attach_wal(wal); - let next = e.create("Note", json!({"title":"d"})).unwrap(); - assert_eq!(next["id"], 7); - let _ = std::fs::remove_file(&path); - } - - #[test] - fn torn_tail_is_dropped_whole() { - let path = tmp("torn"); - { - let mut e = Engine::for_shard(catalog(), 0, 1); - let (wal, _) = Wal::open_and_replay(&path, &mut e).unwrap(); - e.attach_wal(wal); - e.create("Note", json!({"title":"keep"})).unwrap(); - e.create("Note", json!({"title":"casualty"})).unwrap(); - } - // Tear the last record mid-payload. The file is fallocate'd, so the - // data tail is the last non-zero byte (the 0xC0FFEE42 trailer), not - // the file length. - let bytes = std::fs::read(&path).unwrap(); - let tail = bytes.iter().rposition(|&b| b != 0).unwrap() as u64 + 1; - let f = OpenOptions::new().write(true).open(&path).unwrap(); - f.set_len(tail - 7).unwrap(); - drop(f); - - let mut e = Engine::for_shard(catalog(), 0, 1); - let (_, n) = Wal::open_and_replay(&path, &mut e).unwrap(); - assert_eq!(n, 1, "torn record must drop whole"); - assert_eq!(e.list("Note").unwrap().len(), 1); - let _ = std::fs::remove_file(&path); - } -} diff --git a/docs/cm.md b/docs/cm.md deleted file mode 100644 index 775c31f..0000000 --- a/docs/cm.md +++ /dev/null @@ -1,17 +0,0 @@ -scaffold sibling crates, multi-app ecommerce, REST + concurrency docs - -- scaffold 14 placeholder crates (app, db, engine, gen, http, logic, - policy, ql, service, sub, txn, ui, value, wal) — empty Cargo.toml + - src/lib.rs to receive code phase-by-phase from `rt` -- restructure docs/examples/ecommerce into multi-app layout: apps/admin - and apps/storefront, with shared/ types/logic/components, per-app - app.wo + wo.toml, and reusable .htmlx components (layout, money, - order-row) -- add reference/rest/{blog,ecommerce}.rest — VS Code/JetBrains HTTP - request files driving the running prototype, including 501/404/405 - expectations for stubbed endpoints -- add docs/plan/09-concurrency-scaleout.md and docs/plan/ui/00-overview.md; - refine docs/plan/assembly/02-writeonce-stance.md -- refresh templates (about, article, header/footer, home, layout, styles) - and add static favicon/logo -- add infra/sync.sh and tighten .gitignore for reference/ symlinks diff --git a/docs/plan/05-hand-rolled-json.md b/docs/plan/05-hand-rolled-json.md deleted file mode 100644 index 63aabf9..0000000 --- a/docs/plan/05-hand-rolled-json.md +++ /dev/null @@ -1,116 +0,0 @@ -# 05 — Hand-Rolled JSON - -> **Status: ⬜ not started** — Track 1 (runtime foundations), next in the dependency-removal sequence. Board: [00-status.md](../00-status.md) - -**Context sources:** [`./04-cutover-remove-tokio-axum.md`](./done/04-cutover-remove-tokio-axum.md), [`../../prototypes/wo-db/src/value.hpp`](../../prototypes/wo-db/src/value.hpp). - -## Goal - -Replace `serde_json::Value` / `serde_json::Map` with a hand-rolled `Value` type covering exactly the shapes the runtime reads and writes: request bodies, response bodies, and the `Engine`'s in-memory `Row`. Remove `serde` + `serde_json` from `crates/rt/Cargo.toml`. After this phase the dep list is `anyhow` + `libc`. - -## Design decisions (locked) - -1. **Minimal surface.** The runtime's actual JSON needs are small: - - Parse request body bytes → `Value::Object` (single top-level object on every sample endpoint). - - Emit `Value::Object` / `Value::Array` → bytes for the response. - - Pretty-printing is **not** required. Operators reach for `| python3 -m json.tool` if they want it. -2. **No `#[derive(Serialize/Deserialize)]`.** `Value` is the union type; every `Row`, `Product`, `Article` is already a `Value::Object` at the boundary. The only thing that "serializes" is `Value`. The ecosystem of derive-based types doesn't exist in `rt` today — `engine::Row` is `HashMap` via `serde_json` today, becomes `HashMap` via the new module tomorrow. -3. **RFC 8259 compliant, but strict.** No unquoted keys, no trailing commas, no comments. Standard JSON. The runtime isn't serving JSON5. -4. **Parser is recursive descent, zero-copy where possible.** String values borrow from the input buffer unless they contain escapes; objects own their keys. Preserves the "no heavy abstraction" pattern of phase 02 / 03. -5. **Module at `crates/rt/src/json/`.** Same "extract when a second consumer appears" rule. Eventual home is the empty [`crates/value/`](../../crates/value/) sibling — not this phase. - -## Scope - -### New files inside `crates/rt/src/json/` - -| File | Responsibility | Approx LOC | -| --- | --- | --- | -| `mod.rs` | Re-exports `Value`, `Object`, `Array`, `parse`, `emit` | ~10 | -| `value.rs` | `pub enum Value { Null, Bool(bool), Int(i64), Float(f64), Str(String), Array(Vec), Object(BTreeMap) }` + `impl Value` helpers (`as_str`, `as_i64`, `get`, indexing) | ~200 | -| `parse.rs` | `parse(&[u8]) -> Result` — recursive descent: `parse_value` → `parse_object` / `parse_array` / `parse_string` / `parse_number` / `parse_keyword`. Single-pass, no backtracking. | ~300 | -| `emit.rs` | `emit(value: &Value, buf: &mut Vec)` — iterative-ish writer, escapes strings per RFC 8259 §7 | ~150 | - -Total: ~660 LOC. No v1 precedent — no sample parser to port. Reference the target shape against [`prototypes/wo-db/src/value.hpp`](../../prototypes/wo-db/src/value.hpp) for the `Value` variants (same six kinds as the C++ prototype, minus `Float` which that prototype folds into `Int` but we need for HTTP request bodies like `{"qty": 2.5}`). - -### `Cargo.toml` delta - -```diff - [dependencies] - anyhow = "1" --serde = { version = "1", features = ["derive"] } --serde_json = "1" - libc = "0.2" -``` - -### Consumers to update - -Search: `rg 'serde_json|serde::' crates/rt/src | wc -l` — expected ~20 call sites. Each is a mechanical swap: - -| Current | After | -| --- | --- | -| `serde_json::json!({"key": value})` | `json::Value::Object(…)` or a small `json!` macro we ship | -| `serde_json::Value` | `json::Value` | -| `serde_json::Map` | `BTreeMap` (the runtime already uses BTreeMap for stable order) | -| `serde_json::from_slice::(&bytes)?` | `json::parse(&bytes)?` | -| `Json(json!(rows)).into_response()` | `Response::ok().json_body(&rows)` (new helper on phase-03 `Response`) | -| `#[derive(Serialize, Deserialize)]` on any `rt` struct | deleted — no consumer after this phase | - -The biggest consumer is `crates/rt/src/engine.rs` — `Row` is a `serde_json::Map` today. It becomes `BTreeMap`. The `eval_default()` function's `json!(n)` / `json!(b)` calls become `Value::Int(n)` / `Value::Bool(b)`. Minor, all local. - -## A compact `json!` macro (for ergonomics) - -Without `serde_json::json!`, the most-used construction pattern (`json!({"runtime": "wo", "stage": 2})`) gets verbose. Ship a minimal macro: - -```rust -#[macro_export] -macro_rules! json { - (null) => ($crate::json::Value::Null); - (true) => ($crate::json::Value::Bool(true)); - (false) => ($crate::json::Value::Bool(false)); - ([$($e:tt),* $(,)?]) => ( - $crate::json::Value::Array(vec![$($crate::json!($e)),*]) - ); - ({$($k:tt : $v:tt),* $(,)?}) => ({ - let mut m = std::collections::BTreeMap::new(); - $( m.insert(stringify!($k).trim_matches('"').to_string(), $crate::json!($v)); )* - $crate::json::Value::Object(m) - }); - ($e:expr) => ($crate::json::Value::from($e)); -} -``` - -Covers 95% of current `serde_json::json!(...)` uses in the codebase. For the other 5%, build `Value` by hand. - -## Exit criteria - -1. **`cargo build`** — compiles with three deps (`anyhow`, `libc` + `Cargo.toml` itself). -2. **New tests in `crates/rt/src/json/`:** - - `parse_object_simple` — `{"a":1,"b":"x"}` round-trips. - - `parse_nested_and_array` — `{"xs":[1,2,3],"meta":{"k":"v"}}` round-trips. - - `parse_escapes` — `"\\n\\t\\\"\\u0041"` → `"\n\t\"A"`. - - `emit_stable_key_order` — emitting a `BTreeMap`-backed object produces keys in sorted order (matters for `.rest` expected-body stability). - - `parse_errors` — unterminated string, trailing comma, missing comma, unclosed object all return `ParseError` with line/col. -3. **All 14 existing `rt` tests pass** after the swap (the `engine::Engine` and `server::*` tests most affected). -4. **`reference/rest/blog.rest`** — 20 assertions all return the same HTTP status AND the same response body shape (may differ in key ordering if `BTreeMap` ordering differs from `serde_json`'s insertion order — document the shift). -5. **Dep audit.** `cargo tree -p rt --depth 1` shows zero `serde*` lines. - -## Non-scope - -- **No streaming parse.** Request bodies are small (< 1 MB on every sample endpoint). A buffered full-body parse is fine. -- **No `serde_json`-compat feature flag.** Clean break; this is the only consumer that matters, and we control it. -- **No JSON Pointer, no JSON Schema, no JSON Patch.** If needed later, layer on top. -- **No crate extraction to `crates/value/`.** Same rule as phase 02/03: wait for a second consumer. - -## Verification - -```bash -cargo build # three deps -cargo test --lib json # new parser/emitter tests -cargo test --lib # 14 existing tests still green -# full .rest smoke — same script as phase 04 exit criterion 3 -cd reference/crates && cargo build && cargo test -``` - -## After this phase - -Two deps left: `anyhow` and `libc`. Phase 06 removes `anyhow`. After that, `libc` is the only external crate — the stated end goal. diff --git a/docs/plan/06-bespoke-error.md b/docs/plan/06-bespoke-error.md deleted file mode 100644 index d5486cc..0000000 --- a/docs/plan/06-bespoke-error.md +++ /dev/null @@ -1,122 +0,0 @@ -# 06 — Bespoke Error Type - -> **Status: ⬜ not started** — Track 1 (runtime foundations). Board: [00-status.md](../00-status.md) - -**Context sources:** [`./05-hand-rolled-json.md`](./05-hand-rolled-json.md), [`../01-problem.md`](../01-problem.md). - -## Goal - -Replace `anyhow::Error` / `anyhow::Result` with a single `Error` enum rooted in `crates/rt/src/error.rs`. Remove `anyhow` from `crates/rt/Cargo.toml`. At the end of this phase, `[dependencies]` contains only `libc` — the **stated end goal** of this plan sequence. - -## Design decisions (locked) - -1. **One enum per crate.** `rt::Error` covers everything the runtime produces — parse errors, compile errors, engine errors, I/O errors, lex errors, HTTP framing errors, JSON parse errors (phase 05), signal-handler failures. The v1 crates use `std::io::Result` throughout, which is fine for I/O-heavy code but loses context for parse and compile failures. We split the difference with a tagged variant enum. -2. **`From` impls for standard errors.** `io::Error`, `ParseIntError`, `FromUtf8Error`, `Utf8Error` — automatically convert via `?`. Everything else wraps explicitly through constructors like `Error::parse(line, col, msg)`. -3. **`Display` composes a line + context.** No chained backtrace. Error messages stay compact: `parse error at line 42: expected ':', got '}'`. This is what every caller prints today after `anyhow::Error` formats. -4. **`Result` alias.** Shorthand: `pub type Result = core::result::Result`. Replaces `anyhow::Result` every existing call site uses. -5. **No macros.** `anyhow::anyhow!("...")` becomes `Error::msg("...")`. `anyhow::bail!(...)` becomes `return Err(Error::msg(...))`. `anyhow::Context::context(err, "...")` becomes `err.with_context(|| "...")` via a tiny inherent method. - -## Scope - -### New file - -| File | Responsibility | Approx LOC | -| --- | --- | --- | -| `crates/rt/src/error.rs` | `pub enum Error`, `pub type Result`, `From` impls, `Display`, `fn msg`, `fn parse`, `fn with_context` | ~120 | - -### Enum shape (target) - -```rust -#[derive(Debug)] -pub enum Error { - Io(io::Error), - Parse { line: u32, col: u32, message: String }, - Lex { line: u32, col: u32, message: String }, - Compile { message: String }, - Engine { message: String }, - Http { message: String }, - Json { message: String }, - NotFound { ty: String, id: i64 }, - Config { message: String }, - Msg (String), // anyhow::anyhow!-style catch-all -} - -impl Error { - pub fn msg(s: impl Into) -> Self { Error::Msg(s.into()) } - pub fn parse(line: u32, col: u32, m: impl Into) -> Self - { Error::Parse { line, col, message: m.into() } } - // ... one constructor per non-Io variant - - pub fn with_context(self, ctx: F) -> Error - where F: FnOnce() -> String - { - Error::Msg(format!("{}: {}", ctx(), self)) - } -} - -impl From for Error { ... } -impl From for Error { ... } -impl From for Error { ... } -impl From for Error { ... } -impl core::fmt::Display for Error { ... } -impl core::error::Error for Error {} - -pub type Result = core::result::Result; -``` - -### Call-site mechanical sweep - -Find: `rg 'anyhow::|anyhow!|bail!|\.context\(' crates/rt/src | wc -l` — expected ~40 sites. - -| Current | After | -| --- | --- | -| `use anyhow::{Context, Result};` | `use crate::error::{Error, Result};` (from `rt`) | -| `fn foo() -> anyhow::Result` | `fn foo() -> Result` | -| `anyhow::anyhow!("no such table: {}", name)` | `Error::msg(format!("no such table: {}", name))` | -| `anyhow::bail!("...")` | `return Err(Error::msg("..."))` | -| `result.context("while doing X")?` | `result.map_err(\|e\| e.with_context(\|\| "while doing X".into()))?` | -| `Err(anyhow::anyhow!("parse error at line {line}: ..."))` | `Err(Error::parse(line, col, "..."))` | - -`crates/rt/src/parser.rs` and `crates/rt/src/compile.rs` are the biggest consumers. Most edits are verbatim substitutions. - -### `Cargo.toml` delta - -```diff - [dependencies] --anyhow = "1" - libc = "0.2" -``` - -**After this phase:** `[dependencies]` has one line. The stated end goal. - -## Exit criteria - -1. **`cargo build`** — compiles with exactly one external dep. -2. **`cargo test --lib`** — all 14 existing `rt` tests pass. A new test in `error.rs` exercises `From`, `with_context`, and `Display` formatting. -3. **No `anyhow::` references anywhere in the repo.** `rg 'anyhow' crates/ docs/` returns zero hits (docs updated by this phase too). -4. **`reference/rest/blog.rest`** — 20 assertions still pass. Error paths (404, 400) still produce the same response body format (plain-text error message from the handler's `.to_string()`). -5. **`cd reference/crates && cargo build && cargo test`** unchanged. V1 doesn't use `anyhow` — nothing to touch there. -6. **`cargo tree -p rt --depth 1`** lists only `libc` as an external dep (plus transitive ones brought in by libc itself, all of which are kernel-facing). - -## Non-scope - -- **No custom `#[derive(Error)]` macro.** `thiserror` would be nicer but it's a dep. Hand-writing the enum + impls is ~120 lines, done once, maintained rarely. -- **No source-chain traversal.** `Error` stores messages, not source errors (except `Io` which wraps). If we need chained context later, extend the enum then — don't over-engineer now. -- **No `backtrace` crate.** If a panic-style backtrace is ever needed, `RUST_BACKTRACE=1` on `panic!()` gives it. Errors don't carry them. - -## Verification - -```bash -cargo build # one external dep -cargo test --lib # 14 + error.rs test green -rg 'anyhow' crates/ docs/ # zero hits -# full .rest smoke — same script as phase 04 -cd reference/crates && cargo build && cargo test # v1 untouched -cat crates/rt/Cargo.toml | grep -A 20 '\[dependencies\]' # libc is the only line -``` - -## After this phase - -`crates/rt/Cargo.toml` is at its minimum. The runtime drives every I/O operation through direct kernel primitives: `socket`, `bind`, `listen`, `accept4`, `epoll_wait`, `read`, `write`, `signalfd4`, `close`. Nothing between the code and the kernel except `libc`. - -Phases 07 and 08 extend the kernel-primitive surface without adding deps — they're pure feature work on top of the foundation this sequence laid. diff --git a/docs/plan/07-inotify-content-watcher.md b/docs/plan/07-inotify-content-watcher.md deleted file mode 100644 index 0380339..0000000 --- a/docs/plan/07-inotify-content-watcher.md +++ /dev/null @@ -1,126 +0,0 @@ -# 07 — `inotify` Content Watcher - -> **Status: ⬜ not started** — Track 1 (runtime foundations). Board: [00-status.md](../00-status.md) - -**Context sources:** [`./02-event-loop-epoll.md`](./done/02-event-loop-epoll.md), [`./linux/00-linux.md`](./exploration/linux/00-linux.md) § File Watching, `../02-recovery.md` § No AWS Infrastructure. - -## Goal - -First Stage-3 capability. When a `.wo` source file under the active project directory changes, `inotify` fires on the phase-02 event loop and the runtime hot-reloads the affected schema — parser re-run, catalog refreshed, live routes updated in-place. Maps to [`00-linux.md`](./exploration/linux/00-linux.md)'s "watch the content directory for file creates, modifications, and deletes. Triggers re-indexing and subscriber notification when articles change. **Replaces the S3 + Lambda event pipeline entirely.**" - -Also the first real second consumer of the phase-02 `EventLoop` beyond the HTTP listener — validates the abstraction under cross-feature load. - -## Design decisions (locked) - -1. **`inotify_init1(IN_CLOEXEC | IN_NONBLOCK)` + `inotify_add_watch`.** The fd is registered on the phase-02 event loop alongside the HTTP listener. No polling. No cross-platform fallback (`kqueue` on macOS, `ReadDirectoryChangesW` on Windows) — writeonce targets Linux only. -2. **Per-directory watches, not per-file.** `types/`, `ui/`, `logic/`, `tests/`, and `app.wo`'s parent get a watch each; individual files are resolved from the event's `wd` + `name` fields. Prevents fd exhaustion on large projects (the default `fs.inotify.max_user_watches` is 8192 on most distros, but we'd rather spend watches carefully). -3. **Debounce at 150 ms.** Editors issue multiple events per save (create tempfile → write → rename → delete old). A `TimerFd::oneshot(150ms)` per-watch absorbs the burst; only the final "settled" state triggers a recompile. -4. **Full recompile, not incremental.** A file change invalidates the full schema catalog — re-run `rt::discover()` → `rt::parser::parse()` → `rt::compile::Catalog::from_schemas()`. The sample projects are small (8 files for the blog); a full parse is < 50 ms. Incremental type-graph invalidation is a phase 09+ optimization. -5. **Atomic catalog swap.** The `Engine`'s catalog is behind an `Arc>` or equivalent — a new catalog replaces the old under a single pointer write, and in-flight HTTP handlers finish with the old one while new ones see the new. Under single-threaded execution this is essentially free; under sharding it becomes the per-shard atomic. -6. **Module at `crates/rt/src/watch/`.** Same extraction-deferred rule as earlier modules. - -## Scope - -### New files inside `crates/rt/src/watch/` - -| File | Responsibility | Port source | -| --- | --- | --- | -| `mod.rs` | Re-exports `Watcher`, `WatchEvent` | — | -| `inotify.rs` | Raw wrappers: `init()`, `add_watch(path, mask)`, `read_events() -> Vec`. Registers on the `EventLoop`. | [`reference/crates/wo-watch/src/lib.rs`](../../.dev/reference/crates/wo-watch/src/lib.rs) (280 LOC) — v1 already does exactly this | -| `recursive.rs` | Walks the project root, calls `add_watch` for every directory matching `types/\|ui/\|logic/\|tests/` or containing `*.wo` | ~80 new LOC | -| `debounce.rs` | Coalesces bursts per-watch-descriptor, fires a `TimerFd` for the 150 ms settle window | ~100 new LOC | -| `reload.rs` | On debounced fire: re-discover, re-parse, re-compile, `ArcSwap::store(new_catalog)` | ~80 new LOC | - -Total: ~540 LOC (280 ported + ~260 new). - -### `Cargo.toml` change - -None. `libc` already covers `inotify_init1` / `inotify_add_watch` / `inotify_rm_watch`. - -### Routing change in `crates/rt/src/server.rs` - -The router needs to re-resolve the catalog on each request rather than close over a snapshot at boot: - -```rust -// before -let router = Router::new().route("/api/articles", list_h_bound_to_catalog_snapshot); - -// after -let shared = Arc::new(ArcSwap::from_pointee(catalog)); -let router = Router::new().route("/api/articles", move |req, st| { - let cat = shared.load(); - list_h(req, &cat, st) -}); -``` - -One-time rewrite of the 12 handlers (4 types × 3 ops). Mechanical. - -## API shape (target) - -```rust -use rt::event::EventLoop; -use rt::watch::Watcher; - -let mut loop_ = EventLoop::new()?; -let mut watcher = Watcher::recursive(Path::new("docs/examples/blog"), Duration::from_millis(150))?; -watcher.register(&mut loop_)?; - -for event in loop_.wait_once(None)? { - if event.token() == watcher.token() { - for change in watcher.drain() { - eprintln!("[wo] content change: {} ({})", change.path.display(), change.kind); - // reload pipeline fires here - } - } -} -``` - -## Exit criteria - -1. **`cargo build`** green. No new deps. -2. **Unit test:** create a temp dir, write `a.wo`, spin up a `Watcher` on a loop in a test thread, modify `a.wo`, assert the debounced `WatchEvent::Modified(path)` arrives within 250 ms. -3. **End-to-end manual:** - ```bash - cargo run --bin wo -- run docs/examples/blog & - # observe: `curl :8080/api/articles` returns [...] - # edit docs/examples/blog/types/article.wo — add a `nickname: Text?` field - # wait 200 ms - # observe: `curl :8080/api/articles` response shape reflects new field (no restart) - ``` -4. **`[wo]` log lines** match the spec in [00-linux.md](./exploration/linux/00-linux.md) — one line per debounced change, showing the relative path and event kind. -5. **All 14 `rt` unit tests still pass.** `reference/rest/blog.rest` 20-assertion battery still green. -6. **No fd leak** — `ls -la /proc/$PID/fd` before and after ten consecutive edits shows the same count. - -## Non-scope - -- **No cross-platform fallback.** `kqueue` and `ReadDirectoryChangesW` are not on the roadmap. Linux only. -- **No `fanotify`.** [00-linux.md](./exploration/linux/00-linux.md) lists it as "useful if watching needs to span mount points" — writeonce projects live in one directory tree; `inotify` is enough. -- **No incremental reparse.** Full recompile per settled change. If a real project hits the full-recompile wall, phase 09+ can add a dependency-graph-aware rebuilder. -- **No subscription push.** Phase 07 only detects and reloads. Notifying connected clients (the `register! { #{blog-title} => notify(fd) }` model in [00-linux.md](./exploration/linux/00-linux.md)) is phase 09 once `sub` activates. - -## Verification - -```bash -cargo build -cargo test --lib watch -cargo test --lib # 14 existing + watcher tests green - -# manual hot-reload check -cargo run --bin wo -- run docs/examples/blog & -PID=$! -sleep 2 -curl -s http://127.0.0.1:8080/api/articles -echo ' policy read anyone' >> docs/examples/blog/types/article.wo -sleep 0.5 -curl -s http://127.0.0.1:8080/api/articles # server did not restart; catalog refreshed -kill $PID -git checkout docs/examples/blog/types/article.wo # undo the edit - -cd reference/crates && cargo build && cargo test # v1 untouched -``` - -## After this phase - -The runtime now does what `docs/02-recovery.md` originally promised: the binary watches its own content directory with `inotify` and re-indexes on change. The S3 + Lambda + sync-trigger pipeline is fully replaced by one fd on one event loop in one process. - -Phase 08 adds the other half of the v1 kernel-primitive story — `sendfile` for zero-copy static serving. diff --git a/docs/plan/08-sendfile-static-assets.md b/docs/plan/08-sendfile-static-assets.md deleted file mode 100644 index 9002f2f..0000000 --- a/docs/plan/08-sendfile-static-assets.md +++ /dev/null @@ -1,110 +0,0 @@ -# 08 — `sendfile` Zero-Copy Static Serving - -> **Status: ⬜ not started** — Track 1 (runtime foundations); also a prerequisite of the parked UI track. Board: [00-status.md](../00-status.md) - -**Context sources:** [`./03-hand-rolled-http.md`](./done/03-hand-rolled-http.md), [`./linux/00-linux.md`](./exploration/linux/00-linux.md) § Efficient File Serving, `../02-recovery.md`. - -## Goal - -Serve static file bytes — eventually `##ui`-emitted HTML + CSS + JS bundle, today anything put under a project's `static/` directory — via `sendfile(sock_fd, file_fd, NULL, count)`. Zero userspace copy on the payload path: the kernel moves bytes from the page cache directly to the socket's send buffer. Completes the v1 kernel-primitive port started in phase 02. - -## Design decisions (locked) - -1. **`sendfile(2)` only.** Not `splice`, not `vmsplice`. `sendfile` handles "file fd → socket fd" exactly, which is 100% of the use case here. `splice`-through-pipe is ~30% more code for cases we don't have (non-regular-file sources). -2. **`GET /static/...` is the only route mounted.** Hard-coded for Stage 3. When `##ui` arrives in phase 6+ it'll emit bundles into this path; when typed SDK codegen arrives in phase 5 the generated JS client goes here too. -3. **No path traversal.** Canonicalise the requested path; reject anything that escapes the configured static root. Standard directory-traversal defence — `..` segments already stripped by the HTTP request parser from phase 03, but the static resolver double-checks with a realpath comparison. -4. **`open + fstat + sendfile` chain.** No `mmap`. `mmap` wins for repeated reads of the same file (where the page cache warming pays off), but `sendfile` is strictly faster for one-shot delivery since the kernel manages the page cache itself. [00-linux.md](./exploration/linux/00-linux.md) lists both; the runtime's static-asset pattern is one-shot, pick `sendfile`. -5. **`EAGAIN` backoff through the event loop.** If `sendfile` returns partial bytes (send buffer full), re-arm `EPOLLOUT` for the socket and resume when the kernel signals writable. Matches v1 wo-serve's flow. -6. **MIME by extension table.** Compact `match` on `.html`/`.css`/`.js`/`.json`/`.svg`/`.png`/`.jpg`/`.woff2`/`.wasm` covers every asset the SSR layer will emit. Unknown extensions default to `application/octet-stream`. - -## Scope - -### New files inside `crates/rt/src/static_files/` - -| File | Responsibility | Port source | -| --- | --- | --- | -| `mod.rs` | Re-exports `StaticHandler`, `resolve` | — | -| `sendfile.rs` | Raw `sendfile(2)` wrapper + non-blocking `send_all` that co-operates with `EPOLLOUT` | [`reference/crates/wo-serve/src/sendfile.rs`](../../.dev/reference/crates/wo-serve/src/sendfile.rs) (109 LOC) | -| `resolve.rs` | Path canonicalisation + traversal defence + file existence check | [`reference/crates/wo-serve/src/resolve.rs`](../../.dev/reference/crates/wo-serve/src/resolve.rs) (80 LOC) | -| `mime.rs` | Extension → `Content-Type` table | [`reference/crates/wo-serve/src/mime.rs`](../../.dev/reference/crates/wo-serve/src/mime.rs) (44 LOC) | -| `handler.rs` | `StaticHandler` — integrates the three with phase-03's `Response` builder; returns 404 / 403 / 200 as appropriate | ~120 new LOC | - -Total: ~350 LOC (233 ported + ~120 new). - -### `Cargo.toml` change - -None. - -### Router change in `crates/rt/src/server.rs` - -One new route per project: - -```rust -let static_root = project_dir.join("static"); -let handler = StaticHandler::new(static_root); -router.route(Method::GET, "/static/*path", move |req, _| handler.serve(req)); -``` - -`/static/*path` is a new wildcard pattern in the phase-03 router — add it to `route.rs` if not already supported. - -## API shape (target) - -```rust -use rt::static_files::StaticHandler; - -let handler = StaticHandler::new("/app/static"); -handler.serve(&request)?; // returns a Response that streams via sendfile() -``` - -The `Response` returned by `handler.serve()` owns the open `File` fd. The phase-03 connection writer notices it's a `sendfile`-backed response and uses the non-blocking `send_all` path instead of `write`. - -## Exit criteria - -1. **`cargo build`** green. No new deps. -2. **Unit test:** place a 10 MB file under a temp static root, `StaticHandler::serve` on a mock request, assert the `Response` reports 200 / correct `Content-Length` / correct `Content-Type`. A separate integration test validates the actual `sendfile` path using an accepted socket. -3. **`strace` validation:** - ```bash - cargo run --bin wo -- run docs/examples/blog & - PID=$! - # ship a 10 MB file into docs/examples/blog/static/big.bin - strace -p $PID -f -e sendfile,read,write 2>&1 | tee /tmp/strace.log & - curl -s -o /dev/null http://127.0.0.1:8080/static/big.bin - # assert /tmp/strace.log shows sendfile(...) calls and zero read/write - # of the file content - ``` -4. **Path traversal attempts fail closed.** `curl :8080/static/../Cargo.toml` returns 403. `curl :8080/static/nonexistent.png` returns 404. -5. **`EAGAIN` handling.** A test that rate-limits the socket sendbuf to force a partial write exercises the `EPOLLOUT` re-arm path; the full payload still arrives. -6. **All 14 `rt` tests** + phase-02/03/04/05/06/07 additions pass. `reference/rest/blog.rest` 20 assertions still green (no regressions on the JSON endpoints). - -## Non-scope - -- **No `sendfile64`.** Modern glibc aliases `sendfile` to `sendfile64` transparently; explicit 64-bit selection isn't needed. -- **No range requests.** `GET /static/big.bin` with `Range: bytes=...` returns 200 + full payload in Stage 3; proper range-request handling is a follow-on. Nothing in the blog or ecommerce samples uses ranges. -- **No in-memory cache.** The kernel page cache is the only cache. Re-opening the file on every request is cheap; if future profiling says otherwise, add an LRU fd cache — but not pre-emptively. -- **No TLS.** `sendfile` over TLS requires kTLS (`setsockopt(TCP_ULP, "tls")` + kernel 4.13+ and the right cipher suites). Worth doing when TLS lands as its own phase; out of scope here. -- **No compression.** The HTTP response writer in phase 03 doesn't gzip; `sendfile` can't gzip on the fly either. Pre-compress (`.br` / `.gz` sibling files) is a future phase — for now the table lists `.br`/`.gz` extensions with the correct `Content-Encoding` but the caller has to produce the pre-compressed file itself. - -## Verification - -```bash -cargo build -cargo test --lib static_files -cargo test --lib # all existing tests green - -# strace-backed zero-copy proof -cargo run --bin wo -- run docs/examples/blog & -# ... (full script from exit criterion 3) - -# .rest smoke unchanged -# full 20-assertion battery against reference/rest/blog.rest - -cd reference/crates && cargo build && cargo test # v1 untouched -``` - -## After this phase - -The runtime covers every kernel primitive listed in [`00-linux.md`](./exploration/linux/00-linux.md) except `io_uring`, `mmap`, `fallocate`, and `memfd_create` — which all belong to the storage engine (phase 3 of the database series), not the runtime per se. - -Next natural phase: **`09-native-subscriptions.md`** — the `register! { #{blog-title} => notify(fd) }` model from [00-linux.md](./exploration/linux/00-linux.md). Takes the `inotify` watcher from phase 07 and wires it into a subscription table that dispatches delta writes directly to subscriber sockets over the phase-03 HTTP connection. That replaces the Stage-3 `501` stub the `/api//live` endpoint currently returns. - -After that phase, `crates/rt/` is feature-complete for Stages 1–3 of the runtime, with exactly one external dependency. diff --git a/docs/plan/09-concurrency-scaleout.md b/docs/plan/09-concurrency-scaleout.md deleted file mode 100644 index 58a8611..0000000 --- a/docs/plan/09-concurrency-scaleout.md +++ /dev/null @@ -1,138 +0,0 @@ -# 09 — Scale-out: thread-per-core for 10k concurrent users - -> **Status: 🔄 in progress** — 09a/09b/09c ✅ shipped (+ keep-alive and io_uring group-commit follow-ups, measured in the shipped notes below); 09d/09e/09f ⬜ not started. Board: [00-status.md](../00-status.md) - -**Context sources:** [`./08-sendfile-static-assets.md`](./08-sendfile-static-assets.md) (last single-threaded phase), [`./assembly/02-writeonce-stance.md`](./exploration/assembly/02-writeonce-stance.md) (the "single-threaded" policy we're now refining), [`../runtime/database/02-wo-language.md#concurrency-model`](../runtime/database/02-wo-language.md#concurrency-model) (original concurrency stance), [`docs/examples/ecommerce/`](../examples/ecommerce/) (the target workload), [`./linux/`](./exploration/linux/) (kernel primitives), [`reference/go/src/runtime/`](../../.dev/reference/go/src/runtime/) (precedent for a runtime that scales across threads). - -## Context - -Phases 02–08 produce a single-threaded event-loop runtime with zero external Rust dependencies — good enough for the blog sample and the first 500k ops/sec on one core. **The ecommerce workload at `docs/examples/ecommerce/` pushes past that ceiling**: ~10,000 connected websocket subscribers watching `/api/orders/live`, ~1,000 checkouts per second during peak, and every commit fanning out delta frames to a sizeable subset of the connected clients. No single core survives that, regardless of how tight the event loop is. - -This phase is the refinement of the Phase-2 concurrency doctrine "shard to scale past one core" into a concrete architecture. The stance stays the same — **no Go-style goroutines, no work-stealing across threads, no shared mutable heap** — but now we have multiple event loops, each owning its core, its share of connections, and its slice of engine state. - -This doc is a **master plan**. It outlines sub-phases A–F at a high level; each sub-phase lands as its own plan doc (`09a-…`, `09b-…`, …) when implementation starts. No code changes in this pass. The numerical phase slots 10/11/12 are taken by the storage roadmap ([`./10-storage-foundations.md`](./10-storage-foundations.md), [`./11-wal-and-recovery.md`](./11-wal-and-recovery.md), [`./12-engine-disk-cutover.md`](./12-engine-disk-cutover.md)) — single-thread durability has to land before the per-shard WAL of `09c`. - -## Goal - -Serve the ecommerce sample at 10,000 concurrent websocket subscribers + 1,000 checkouts/second on a single 8–16 core box, with per-request tail latencies (`p99`) inside 50 ms for reads and 100 ms for commits. After this phase sequence lands, `cargo run --bin wo -- run docs/examples/ecommerce` can handle production-shaped load with only the process count as the horizontal scale knob. - -## Design decisions (locked) - -1. **Thread-per-core, not M:N.** `N` OS threads pinned to `N` cores via `sched_setaffinity(cpu_set_t)`. Each thread runs its own event loop (the [phase-02 `runtime/` module](./done/02-event-loop-epoll.md)) plus a local shard of engine state. Pinned for the thread's lifetime; a connection accepted on thread K stays on thread K forever. Precedent: Seastar / ScyllaDB / Redis Cluster. -2. **Shared-nothing state.** No cross-thread mutable access to the catalog, engine rows, or subscription registry. Communication is message-passing over single-producer-single-consumer ring buffers (crossbeam-style, built on `std::sync::atomic`, per [`./assembly/02-writeonce-stance.md`](./exploration/assembly/02-writeonce-stance.md) — still no asm). If thread A needs to touch data owned by thread B, it sends a message; B processes it on its own tick. -3. **SO_REUSEPORT for listener-side load balancing.** Every thread binds a socket with `SO_REUSEPORT` on the same `:8080` — the kernel distributes incoming SYNs across the `N` listener sockets with consistent hashing on the connection 4-tuple. No user-space accept-thread bottleneck. Linux ≥ 3.9 is fine; ≥ 4.5 adds `BPF` filters for custom routing if we ever need session-affinity. -4. **Per-thread io_uring ring.** Each thread gets its own `io_uring_setup` ring with `IORING_SETUP_SINGLE_ISSUER` + `IORING_SETUP_SQPOLL` ([per `./linux/07-io_uring.md`](./exploration/linux/07-io_uring.md)). No ring sharing across threads — simpler ordering, no contention. -5. **Shard key: customer id (modulo N).** The ecommerce schema is customer-centric — one customer's orders + purchase edges + cart live on the same shard. Cross-customer queries (admin `list orders`) fan out; same-customer operations (checkout) are local. Blog shard key would be `author.id` for the same reason. -6. **Cross-shard transactions via 2PC.** A checkout that updates inventory on shard A and customer balance on shard B uses two-phase commit between the two engine threads. Phase 4's transaction coordinator (from [`../runtime/database/02-wo-language.md`](../runtime/database/02-wo-language.md) § Cross-Paradigm Transaction Coordinator) already handles this pattern for sql+doc+graph inside one process; it generalises cleanly to cross-thread. -7. **No Go-style goroutines.** Connections are not tasks that migrate. Each connection's state machine runs on its owning thread's event loop, just as it does in the single-threaded model — the difference is there are now `N` event loops running concurrently. - -## What we copy from Go, what we don't - -Read [`reference/go/src/runtime/netpoll_epoll.go`](../../.dev/reference/go/src/runtime/netpoll_epoll.go) and [`reference/go/src/runtime/proc.go`](../../.dev/reference/go/src/runtime/proc.go) for the shape; copy the **ideas** about fd-to-loop mapping and atomic-counter-based wake-up. Do **not** copy: - -| Go feature | Why writeonce skips it | -| --- | --- | -| Goroutines (M:N scheduling, work stealing) | Goroutines pay context-switch + GC-scan costs the thread-per-core model avoids. Scylla benchmarks consistently beat Go-style runtimes at the same hardware. | -| Shared heap + GC | No heap GC — Rust ownership. Data is partitioned across threads, not shared with locks. | -| `gogo` / `mcall` / `systemstack` asm | No scheduler-controlled stack switching. See [`./assembly/02-writeonce-stance.md`](./exploration/assembly/02-writeonce-stance.md). | -| `cgo` boundary | Rust is the only language. `libc` is already ABI-compatible via `extern "C"`. | -| `asyncPreempt` preemption | Handlers run to completion on their owning thread. Back-pressure comes from bounded per-thread queues, not preemption. | - -And what we **do** copy: - -| Go pattern | Writeonce translation | -| --- | --- | -| Per-P netpoller (the `pp.pollDesc` model) | Per-thread `EventLoop` (the phase-02 `runtime::EventLoop`) | -| `netpollBreak` (fd wake-up via sendto) | Per-thread `eventfd` — one fd per thread, write to it to wake a sleeping `epoll_wait`. See [`./linux/02-eventfd.md`](./exploration/linux/02-eventfd.md). | -| `findrunnable` (what to do when idle) | Per-thread idle-state: drain in-process message queues, run compaction, run periodic timers (from [`./linux/03-timerfd.md`](./exploration/linux/03-timerfd.md)). | -| `runtime.GOMAXPROCS` | `WO_THREADS` env var (defaults to `std::thread::available_parallelism()`). | - -## Linux primitives this phase leans on (beyond the phase-02/03/08 set) - -Reference cards already exist for most; this phase adds the ones that are cross-thread-specific: - -| Primitive | Use | Reference | -| --- | --- | --- | -| `SO_REUSEPORT` | N listener sockets on the same port; kernel load-balances accepts | [`reference/linux/net/core/sock_reuseport.c`](../../.dev/reference/linux/net/core/sock_reuseport.c) — worth adding `linux/12-so-reuseport.md` | -| `sched_setaffinity` + `cpu_set_t` | Pin each thread to its core | [`reference/linux/kernel/sched/core.c`](../../.dev/reference/linux/kernel/sched/core.c) | -| `futex(2)` | Fallback cross-thread wait if per-thread eventfd wake-up isn't enough | [`reference/linux/kernel/futex/`](../../.dev/reference/linux/kernel/futex/) — worth `linux/13-futex.md` | -| `membarrier(2)` | Process-wide memory barrier when a rebalance migrates state between threads | [`reference/linux/kernel/sched/membarrier.c`](../../.dev/reference/linux/kernel/sched/membarrier.c) | -| `io_uring` with `IORING_SETUP_SINGLE_ISSUER` | One ring per thread, pinned | [`./linux/07-io_uring.md`](./exploration/linux/07-io_uring.md) | -| `eventfd` per thread | Cross-thread wake-up — thread A writes to thread B's eventfd to deliver a message | [`./linux/02-eventfd.md`](./exploration/linux/02-eventfd.md) | -| `mmap(MAP_HUGETLB)` | Per-thread arena allocator backed by 2 MB pages for cache locality | [`./linux/08-mmap.md`](./exploration/linux/08-mmap.md) | - -## Sub-phase sequence - -Each one lands as its own numbered plan doc when ready for implementation. Smoke test (`cargo run --bin wo -- run docs/examples/ecommerce` serves correctly) stays green after every sub-phase. - -### `09a-thread-per-core.md` — N event loops, `SO_REUSEPORT` — ✅ shipped - -Introduce a thread-pool manager at `crates/rt/src/runtime/scheduler.rs` (Go parallel: `proc.go`). Spawn `WO_THREADS` OS threads at boot; each pins itself and runs an `EventLoop`. Replace the single `Listener` with per-thread listeners bound `SO_REUSEPORT` to the same port. State is still global at first (shared `Arc>`) — one thing at a time. Exit criterion: `wo run` boots N threads visible in `ps -T`, accepts load balanced across them per `ss -tnp`, no regression in the 20-assertion blog smoke. - -**Shipped:** `scheduler.rs` (~190 LOC) ports the proven [`wo-rt-c` phase A](./exploration/c-runtime/00-plan.md) sequence: workers named `wo-shard-`, pinned via `sched_setaffinity` (verified tid→cpu 0,1,2,3); `Listener::bind_reuseport` (+ unit test: two binds on one port succeed, plain bind still fails); signals blocked in `main` before spawn, worker 0 owns the `signalfd` and broadcasts shutdown through per-worker `eventfd`s. Measured: 2,400 concurrent requests spread evenly across 4 shards (49.8–56.6 M ns on-CPU per shard via `schedstat`); blog CRUD + 501-stub smoke green; ecommerce/hello/pricing boot unchanged; `WO_THREADS=1` preserves the old single-threaded behavior; SIGTERM joins all shards. Engine remains `Arc>` per this sub-phase's scope — 09b shards it. - -### `09b-sharded-engine.md` — per-thread engine state — ✅ shipped - -Partition the in-memory engine catalog + row BTreeMaps by shard id (= thread id). Shard key is `customer.id` for ecommerce / `author.id` for blog / per-type default for anything else. Add a shard router in front of every REST/WS handler: resolve the shard from the request's identifying field, send an in-process message to that thread's mailbox, await response. Shared `Arc>` goes away; each thread owns its slice. Cross-shard reads (admin `list orders`) fan out to every thread and merge results. - -**Shipped** (per-type-default shard key; declared-field shard keys await Phase 4's typed wire layer): `crates/rt/src/shard.rs` — `ShardBus` (per-shard mpsc job mailbox + mail `eventfd`) and `ShardCtx` (each worker's own `Engine`). `Engine::for_shard` mints interleaved ids (shard t: t+1, t+1+n, …) so `owner(id) = (id-1) % n` needs zero coordination; creates are always local, point ops hop at most once as boxed-closure jobs, lists fan out and merge by id. Deadlock-free by two rules: jobs never block (pure local engine ops), and waiters pump their own inbox while parked. `Arc>` is deleted; `HandlerFn` dropped its `Send+Sync` bounds (routers are thread-local now). Verified: cross-shard GET/PATCH/404 against rows owned by other shards, merged lists from all shards, all samples green, 42 unit tests (incl. interleave + cross-thread round-trip), clean broadcast shutdown. Measured: durable-free writes 74.9k → **112.9k/s (+51%)**, write p99 4.5 → 3.4 ms vs 09a on the same box — the read path stays connection-setup-bound until keep-alive/io_uring land (09's later phases). - -### `09c-per-shard-wal.md` — one WAL file per shard — ✅ shipped (epoll-stage scope) - -Each thread has its own `foo.wal` + `foo.data` + per-thread `io_uring` ring ([phase 11's durability work](./11-wal-and-recovery.md), repeated per shard). Recovery is parallel across threads. No shared WAL writer thread. Group commit is per-thread. - -**Shipped** (`crates/rt/src/wal.rs` + engine integration): per-shard `shard-.rwal` under `WO_DATA` (default `./wo-data`, `off` disables), C-prototype frame format (`len|crc32|payload|COMMIT`, hand-rolled CRC32, fallocate prealloc), JSON `WalRec` payloads carrying full post-default rows so replay is byte-exact. Durability hooks inside `Engine::{create,update,delete}` — RAM apply → append → `fdatasync` → return, with undo-on-WAL-failure — which puts every ack behind the fsync *including cross-shard jobs* (the reply leaves the owner only after its engine call returns durable). Boot replays per shard in parallel before accepts arm; torn tails drop whole; the id high-water restores per stride; a `meta` file refuses a mismatched `WO_THREADS`. **Deliberately deferred to the io_uring port: group commit** — batching acks on a per-tick fsync over the epoll loop would reopen the ack-before-fsync race the C phase-F crash test caught, so this stage pays one `fdatasync` per commit (~1% on tmpfs; 44 unit tests incl. replay round-trip + torn-tail). Verified e2e: 32-record crash recovery exact (creates/update/delete, ~4 ms/shard), no id collisions post-recovery, durable 178.3k commits/s vs 180.1k non-durable. The `.data` snapshot/compaction half stays with phase 11. - -**Follow-up shipped — io_uring group commit** (`runtime/netpoll_io_uring.rs` — raw ring, kernel ABI structs by hand, no liburing; `wal::WalGroup`): mutations stage frames and **park their acks** instead of fsyncing inline; once per loop tick the worker submits the whole batch as one `WRITE`→`FSYNC` linked SQE pair (one `io_uring_enter`); the fsync CQE releases every parked ack — local responses via gated `Parked` connections with the C-proven generation stamps, cross-shard replies via parked callbacks on the owner's batch. The epoll loop polls the ring fd as an ordinary event source (full network port still pending). **Two bugs found and fixed during real-disk verification:** (1) the write and fsync CQEs of a linked pair routinely land in *different ticks* on ext4 — releasing on the first CQE alone acked before durability (caught because the pre-fix number, 230k/s, was impossibly fast for the disk); (2) reused `user_data` could attribute a stale CQE to the wrong batch — now `(batch_seq << 1) | op-bit`. **A new deadlock class was designed out**: two shards mutually parked on each other's batches would never reach their tick-end flush — `wal_pump` now runs inside every cross-shard wait loop. Measured (8 shards, 64 conns): real ext4/NVMe **27,014 durable commits/s vs 5,765 per-commit (4.7×)**, p50 2.2 ms (one shared fsync per tick); tmpfs 330k/s; reads unaffected (746k/s); crash-under-load on real disk: every acked write recovered; 100 concurrent cross-shard durable PATCHes, zero stalls; `WO_GROUP_COMMIT=off` keeps the per-commit path for A/B. 47 unit tests. - -**Follow-up shipped — HTTP keep-alive** (the C phase-C connection semantics, in `http/{request,response,connection}.rs`): requests report their `keep_alive` wish + consumed byte count; the connection state machine loops over buffered requests (pipelined carry-over included), resets to Reading after each flush, honors `Connection: close` and HTTP/1.0 defaults. Measured on the clean box: reads 227.8k → **770.7k/s (×3.4)**, durable writes 178.3k → **331.5k/s (×1.9)**, p99 993 → 172 µs, zero reconnects at 64 conns — now ahead of Go `net/http` on both axes while fsyncing every write, within ~10% of the C prototype's reads. 46 unit tests (keep-alive ×3, pipelining, close-header). Remaining C-side advantage: io_uring + per-tick group commit (`09g`/io_uring port territory). - -### `09d-cross-shard-subscriptions.md` — LIVE fanout - -A commit on shard K that creates/updates rows of type T needs to wake subscribers on every shard watching T. Via broadcast: K writes the delta to a per-subscriber-thread mailbox — one message per destination thread, not per subscriber. The destination thread then does the fine-grained predicate match against its local subscription table. Avoids N² traffic when N connections watch the same stream. - -### `09e-cross-shard-txn.md` — 2PC for transactions that span shards - -`fn checkout(customer, product, qty)` might touch shards A (customer), B (product), and C (order) if they hash differently. The transaction coordinator (already designed in [`../runtime/database/02-wo-language.md`](../runtime/database/02-wo-language.md) § Cross-Paradigm Transaction Coordinator) generalises to cross-shard: `begin(snapshot_ts)` broadcasts to all participating shards, `prepare()` collects votes, `commit(wal_lsn)` atomically flips markers, `abort()` if any participant refuses. The per-shard WAL entries carry the 2PC state machine. - -### `09f-observability-and-rebalance.md` — ops - -Per-shard metrics (connections, ops/s, p99, WAL lag), Prometheus scrape endpoint on one well-known thread. A `WO_RESHARD` admin command migrates a contiguous customer-id range from shard K to shard K′ via state snapshot → replay → cutover. For a fixed-core deployment this is rare; matters when `WO_THREADS` changes between runs. - -## Verification targets (after `09f` lands) - -Ecommerce sample on an 8-core box with `WO_THREADS=8`: - -| Metric | Target | How measured | -| --- | --- | --- | -| Concurrent WS subscribers | **10,000** | `websocat` fan-out against `/api/orders/live` + persistent count | -| Checkout throughput | **1,000/s** sustained | Load driver fires `POST /api/fn/checkout` with per-customer key distribution | -| Read p99 | **< 50 ms** | `GET /api/orders?customer=X` under 10k-subscriber background load | -| Commit p99 | **< 100 ms** | Measured from `POST /api/fn/checkout` acceptance to HTTP ack | -| Memory steady-state | **< 2 GB RSS** | 10k connections × 2 KB/conn + engine working set | -| Dep count | **1** (`libc`) | `crates/rt/Cargo.toml` still has only libc after all this | -| `wo run docs/examples/blog` | still boots and serves | phase 02–08 regression test, unchanged | - -## Non-scope - -- **No Go-style goroutines, even after this phase.** Adding M:N scheduling is not on the roadmap. When one core runs out, add more cores (more threads) — horizontally, thread-per-core. -- **No distributed (multi-node) sharding.** This phase is single-box only. Redis-Cluster-style network sharding is a separate future phase; the in-process shard bus (`09a`'s mailboxes) is not the same thing as a cluster membership protocol. -- **No dynamic thread count at runtime.** `WO_THREADS` is set at boot and pinned. Adding/removing a thread means a rolling restart. Acceptable for a database; fundamental to the zero-contention model. -- **No work-stealing.** A slow handler on thread A does not get rebalanced to thread B. Back-pressure is the thread-local queue filling up. If one thread hot-spots because of a bad shard key, the fix is to reshard — not to steal. -- **No `std::thread::available_parallelism` on exotic hosts.** `WO_THREADS` override covers kubernetes CFS-bound pods, NUMA partitioning, and single-core debug runs. -- **No new external Rust dependencies.** Same stance as phases 02–08 — `libc` only. Message passing, atomics, affinity, futex — all through libc or `std::sync::atomic`. - -## Escape hatch - -If the "single core per process, shard across processes" argument ([Redis Cluster model](../runtime/database/02-wo-language.md#concurrency-model)) turns out to be more operationally attractive than a single multi-threaded process, **every decision in this plan translates**. Per-thread shards become per-process shards; `SO_REUSEPORT` inside the kernel becomes a reverse proxy in front; in-process mailboxes become Unix domain sockets. The phase-02 event loop is the reusable atom regardless. - -## Cross-references - -- [`./exploration/c-runtime/00-plan.md`](./exploration/c-runtime/00-plan.md) — the C prototype's phased evolution (threads → arena → io_uring → WAL → recovery); the executable proving ground for 09a's thread-per-core skeleton and 09c's per-shard WAL before the Rust work starts. -- [`./08-sendfile-static-assets.md`](./08-sendfile-static-assets.md) — last prerequisite phase; feature-complete single-threaded runtime. -- [`./assembly/02-writeonce-stance.md`](./exploration/assembly/02-writeonce-stance.md) — updated to reference this phase's thread-per-core model; still no asm. -- [`../runtime/database/02-wo-language.md#concurrency-model`](../runtime/database/02-wo-language.md#concurrency-model) — the stance this plan refines. -- [`reference/go/src/runtime/proc.go`](../../.dev/reference/go/src/runtime/proc.go) — Go's scheduler, for contrast. -- [`reference/go/src/runtime/netpoll_epoll.go`](../../.dev/reference/go/src/runtime/netpoll_epoll.go) — per-P netpoller, the idea we borrow. -- [`reference/linux/net/core/sock_reuseport.c`](../../.dev/reference/linux/net/core/sock_reuseport.c) — kernel load balancer. -- [`reference/linux/kernel/sched/core.c`](../../.dev/reference/linux/kernel/sched/core.c) — affinity syscalls. diff --git a/docs/plan/10-storage-foundations.md b/docs/plan/10-storage-foundations.md deleted file mode 100644 index eb52c88..0000000 --- a/docs/plan/10-storage-foundations.md +++ /dev/null @@ -1,148 +0,0 @@ -# 10 — Storage Foundations: on-disk row codec + segment append path - -> **Status: ⬜ not started (scope reduced)** — WAL framing/fallocate/CRC landed early via plan 09c; the `@table(name:, index:)` storage-config surface and in-RAM secondary indexes (`Engine::find_by`) landed via the plan-13 follow-up (spec: [`02-wo-language.md § Type-Level Annotations`](../runtime/database/02-wo-language.md)) — this plan inherits the surface and gives indexes their on-disk form. Board: [00-status.md](../00-status.md) - -**Context sources:** [`./done/04-cutover-remove-tokio-axum.md`](./done/04-cutover-remove-tokio-axum.md), [`../runtime/database/03-inmemory-engine.md`](../runtime/database/03-inmemory-engine.md), [`../runtime/database/07-wo-seg-migration.md`](../runtime/database/07-wo-seg-migration.md), [`./exploration/postgresql/smgr-and-md.md`](./exploration/postgresql/smgr-and-md.md), [`./exploration/postgresql/page-format.md`](./exploration/postgresql/page-format.md), [`./exploration/linux/12-pwrite-fsync.md`](./exploration/linux/12-pwrite-fsync.md), [`./exploration/linux/09-fallocate.md`](./exploration/linux/09-fallocate.md), [`reference/crates/wo-seg/src/`](../../.dev/reference/crates/wo-seg/src/). - -## Goal - -Every engine mutation appends a typed record to a per-type segment file on disk. **Reads still hit the in-memory `HashMap` — no behaviour change visible to clients yet.** Killing the process after a write leaves a real `data/.seg` on disk; restart re-creates an empty `HashMap` and ignores the segment (recovery is phase 11). This phase only proves the **on-disk row format**. - -Lays the codec + filesystem layout that phase 11 (WAL + recovery) and phase 12 (disk-backed engine) build on top of. - -## Design decisions (locked) - -1. **One segment file per type.** `data/.seg`. No per-record file proliferation, no per-database tablespaces, no relfilenode indirection (per [`./exploration/postgresql/smgr-and-md.md`](./exploration/postgresql/smgr-and-md.md) — Postgres' multi-file model exists for multi-tenant ops; writeonce binds to one data dir per `wo run`). -2. **Append-only with tombstone byte.** Updates and deletes append a new record (with the old one's id + a `TOMBSTONE` flag); compaction is a follow-on phase. Same model as v1 wo-seg. -3. **Length-prefix framing with CRC32C trailer.** `[u32 length LE][u8 flags][u8 record_kind][u64 LSN][payload bytes][u32 CRC32C]`. The CRC trailer is the **one design point where writeonce diverges from v1 wo-seg**: wo-seg skipped checksums; we don't. -4. **Payload codec is `serde_json` for now.** Phase 05 (hand-rolled JSON) swaps it; the codec slot is a single `RowCodec` trait so the swap is mechanical. -5. **`posix_fallocate` to 1 MiB at file creation.** Doubles when full. Avoids `ENOSPC` mid-write and minimizes filesystem-level fragmentation. Per [`./exploration/linux/09-fallocate.md`](./exploration/linux/09-fallocate.md). -6. **`pwrite` for the append, no fsync yet.** This phase does not commit a durability barrier — the bytes land in the OS page cache and that's it. Phase 11 adds the fsync. Lets us validate the format without conflating it with fsync semantics. -7. **Module at `crates/db/`, not extracted from `rt`.** `crates/db/` has been a placeholder since the scaffolding phase — this phase populates it. Other crates (`engine`, `value`, `wal`, `txn`) stay placeholders until their phases activate. - -## Scope - -### New files inside `crates/db/src/` - -| File | Responsibility | Approx LOC | -| --- | --- | --- | -| `lib.rs` | Re-exports `SegStore`, `Frame`, `Flags`, `RecordKind`, `RowCodec`, `LSN`. Replaces today's empty `lib.rs` doc-comment. | ~30 | -| `frame.rs` | `Frame` struct + `encode(payload, flags, kind, lsn) -> Vec` + `decode(bytes) -> Result` with CRC verification. | ~150 | -| `crc.rs` | CRC32C via the SSE 4.2 `crc32c.h` algorithm. Software fallback for older CPUs. ~80 lines hand-rolled vs. pulling a crate. | ~80 | -| `codec.rs` | `trait RowCodec { fn encode(&self, row: &Row, buf: &mut Vec); fn decode(&self, bytes: &[u8]) -> Result; }` + `JsonCodec` impl backed by today's `serde_json`. | ~50 | -| `seg.rs` | `SegStore { dir: PathBuf, fds: HashMap, tails: HashMap }`. `open(dir)`, `append(ty, &Row) -> Result`, `read(ty, offset) -> Result` (used by phase 11 recovery, not by the engine yet). | ~250 | - -Total: ~560 LOC. The framing math + fallocate + pwrite plumbing is ported from [`reference/crates/wo-seg/src/{writer.rs,reader.rs,header.rs}`](../../.dev/reference/crates/wo-seg/src/) with the CRC trailer added. - -### File layout written under `/` - -``` -docs/examples/blog/ -├── app.wo -├── ui/... -└── data/ ← created by phase 10 - ├── Article.seg - ├── Author.seg - ├── Comment.seg - └── Tag.seg -``` - -`data/` is gitignored (already covered by `/data` and `/docs/examples/*/data` in `.gitignore`). Empty when no rows exist; created lazily on first write. - -### Record framing (illustrated) - -```text - ┌─ length excludes itself; covers flags..CRC. - ▼ -[u32 length LE][u8 flags][u8 kind][u64 LSN][payload bytes ...][u32 CRC32C] - │ │ - │ └─ 0x00 = ROW, 0x01 = TOMBSTONE, others reserved - └─ 0x00 = ACTIVE, 0x01 = DELETED (per-record live bit) -``` - -`flags` is a per-record live bit — flip it to `DELETED` to soft-delete in place without rewriting the payload. `kind` is the discriminator for upcoming record kinds (phase 11 introduces `WAL_BEGIN`, `WAL_COMMIT`); for phase 10 every record is `ROW`. `LSN` is `0` until phase 11 starts assigning real LSNs — it's a placeholder slot now so phase 11 doesn't reshape the format. - -### Engine integration - -The `Engine::create / update / delete` methods in `crates/rt/src/engine.rs` get a `seg_store: Arc>` field plumbed through `Engine::new`. After every successful in-memory mutation: - -```rust -self.seg_store.lock().unwrap() - .append(ty, &row) - .map_err(|e| anyhow!("seg append: {e}"))?; -``` - -Failure aborts the whole mutation — the in-memory write is rolled back. This phase does NOT introduce a "best-effort persistence" mode. - -`Engine::list / get` remain unchanged; reads stay in-memory. - -### `Cargo.toml` delta - -```diff - [dependencies] - anyhow = "1" - serde = { version = "1", features = ["derive"] } - serde_json = "1" - libc = "0.2" -+ -+[dependencies.db] -+path = "../db" -``` - -`crates/db/Cargo.toml` itself stays at `libc + serde_json` (the latter via `RowCodec`'s `JsonCodec`). When phase 05 lands, the `serde_json` import collapses into the runtime's hand-rolled `Value`. - -The root workspace member list also activates: `crates/db` joins `crates/rt` as a non-empty member. - -## Exit criteria - -1. **`cargo build`** at root — both `crates/rt` and `crates/db` compile. Five direct deps (`anyhow`, `serde`, `serde_json`, `libc`, `db`). -2. **`cargo test --lib`** — all existing 37 `rt` tests still green; new `db` tests cover: - - `frame_roundtrip` — encode then decode produces the same `Frame`. - - `crc_detects_corruption` — flipping one byte in the payload makes `decode` return `CrcMismatch`. - - `seg_append_writes_to_disk` — `append` then re-`open` reads the same row back. - - `seg_grows_when_full` — appending past the initial 1 MiB triggers a fallocate-grow without losing existing records. -3. **End-to-end** — `cargo run --bin wo -- run docs/examples/blog`, `curl -X POST /api/articles` with a body, then `xxd docs/examples/blog/data/Article.seg | head -3` — output shows the magic length prefix and the JSON payload. -4. **`reference/rest/blog.rest`** — 20-assertion battery still passes byte-identically. -5. **Restart leaves the segment on disk but ignores it.** `wo run`, write 5 rows, ctrl-C, `wo run` again, `GET /api/articles` returns `[]`. The segment file still exists. Phase 11 will start replaying it. - -## Non-scope - -- **No fsync.** Pure write path; durability barrier is phase 11. -- **No WAL.** Mutations go straight to the segment. Phase 11 introduces a separate WAL log; segments become the post-checkpoint home for replayed records. -- **No reads from disk.** `Engine::get` stays in-memory. Phase 12 cuts over. -- **No secondary indexes.** Phase 12 introduces a primary `id` BTree on disk; secondary indexes (`unique`, `index` schema attributes) are a later phase. -- **No compaction.** Tombstoned records pile up. Compaction lands when a benchmark says it has to. -- **No cross-type transactions / RETURNING aliases.** The locked schema design (`02-wo-language.md`) names cross-paradigm transactions; the runtime gets there in a later phase. -- **No `crates/db` API stability.** Internal-only until `crates/db/Cargo.toml` declares `[lib]`-level external surfaces. - -## Verification - -```bash -cargo build # rt + db both compile -cargo test --lib # rt + db unit tests -cargo test -p db # db-only - -# manual end-to-end -cargo run --bin wo -- run docs/examples/blog & -PID=$! -sleep 1 -curl -s -X POST http://127.0.0.1:8080/api/articles \ - -H 'Content-Type: application/json' \ - -d '{"slug":"a","title":"A","author":1,"published":true,"meta":{"excerpt":"e","body_md":"b"}}' -ls -la docs/examples/blog/data/ -xxd docs/examples/blog/data/Article.seg | head -5 -kill -INT $PID - -# restart sanity — phase 10 is "format-only", no replay -cargo run --bin wo -- run docs/examples/blog & -PID=$! -sleep 1 -curl -s http://127.0.0.1:8080/api/articles # expect [] -kill -INT $PID - -cd reference/crates && cargo build && cargo test # v1 untouched -``` - -## After this phase - -The on-disk format exists but is dead weight — written, never read. Phase 11 brings it to life: introduces a separate WAL log, fsync at commit, group commit per loop tick, and a recovery loop that replays the WAL into the in-memory `HashMap` on startup. Phase 12 then cuts the engine over to read from segments instead of from RAM, completing the transition from in-memory to durable storage. diff --git a/docs/plan/11-wal-and-recovery.md b/docs/plan/11-wal-and-recovery.md deleted file mode 100644 index cb03269..0000000 --- a/docs/plan/11-wal-and-recovery.md +++ /dev/null @@ -1,211 +0,0 @@ -# 11 — WAL + crash recovery - -> **Status: ⬜ not started (scope reduced)** — replay + ack-after-fsync + group commit landed via 09c and its follow-ups; remaining here: snapshots (`.data`), compaction, WAL rotation. Board: [00-status.md](../00-status.md) - -**Context sources:** [`./10-storage-foundations.md`](./10-storage-foundations.md), [`../runtime/database/02-wo-language.md#concurrency-model`](../runtime/database/02-wo-language.md#concurrency-model), [`../runtime/database/03-inmemory-engine.md`](../runtime/database/03-inmemory-engine.md), [`./exploration/postgresql/wal.md`](./exploration/postgresql/wal.md), [`./exploration/postgresql/buffer-and-checkpoint.md`](./exploration/postgresql/buffer-and-checkpoint.md), [`./exploration/linux/12-pwrite-fsync.md`](./exploration/linux/12-pwrite-fsync.md), `../02-recovery.md`. - -## Goal - -`kill -9` mid-write loses nothing acknowledged. On restart, recovery replays the WAL into the in-memory `HashMap` and the engine serves traffic exactly as if nothing had happened. **The engine is still `HashMap`-backed in this phase** — phase 12 changes that. Phase 11 wires durability without changing the engine's read shape. - -This is the phase where the locked architecture statement from [`02-wo-language.md` § Concurrency Model](../runtime/database/02-wo-language.md#concurrency-model) becomes code: - -> "On `COMMIT`, the WAL record must be `fsync`'d before the client gets acknowledgment. Group commit drains many pending commits into one fsync SQE per tick." - -## Design decisions (locked) - -1. **WAL is separate from the segment files.** Segments hold post-recovery row data; the WAL is the durability log we replay from. Per [`./exploration/postgresql/wal.md`](./exploration/postgresql/wal.md). Different fsync cadence (every commit for WAL, every checkpoint for segments). -2. **LSN = monotonic byte offset across all WAL segments.** Postgres convention. 64-bit. Simple `<` comparisons. Matches the placeholder slot phase 10 reserved in the frame header. -3. **Group commit via the loop tick.** No separate writer thread. Every loop tick: drain all pending commits, **one** `fdatasync` covers all of them, then ack each request. Same effect as Postgres' group-commit fence; the fence is the tick boundary. -4. **`fdatasync`, not `fsync`, for the WAL.** WAL files are `posix_fallocate`'d up front to a fixed segment size — writes never extend them, so the inode metadata doesn't change and `fdatasync` is sufficient. Per [`./exploration/linux/12-pwrite-fsync.md`](./exploration/linux/12-pwrite-fsync.md). -5. **Control file via rename-on-write.** `data/control.tmp` → `fsync` → `rename` → parent-dir `fsync`. Atomic across crashes. -6. **Replay is idempotent.** Recovery replays records starting at `last_durable_lsn`; partial replay (crash mid-recovery) re-replays from the same anchor with the same effect. -7. **Module at `crates/wal/`.** `crates/wal/`'s placeholder doc-comment names this work. Populated here. - -## Scope - -### New files inside `crates/wal/src/` - -| File | Responsibility | Approx LOC | -| --- | --- | --- | -| `lib.rs` | Re-exports `Wal`, `Lsn`, `Replay`, `ControlFile`. | ~30 | -| `lsn.rs` | `pub struct Lsn(pub u64)` — newtype with `Display`, ordering, segment-id + offset accessors. | ~50 | -| `wal.rs` | `Wal { dir, active_fd, active_seg_id, tail_lsn, pending: Vec }`. `append(rec) -> Lsn`, `commit() -> io::Result` (issues `fdatasync`), `enqueue_ack(fd) / drain_acks() -> Vec`. | ~250 | -| `segment.rs` | WAL-file rollover: open new `.wal`, `posix_fallocate` to 16 MiB, switch active fd, retire the previous segment. | ~120 | -| `control.rs` | `ControlFile { magic, version, last_durable_lsn, crc }` — read on startup, write on checkpoint. Rename-on-write. | ~120 | -| `replay.rs` | `Replay::from(dir, last_lsn) -> Iterator>` — walks WAL forward, yields decoded records to the caller. | ~150 | - -Total: ~720 LOC. No v1 precedent — wo-wal doesn't exist (despite being named in the database series). New territory. - -### File layout - -``` -docs/examples/blog/ -└── data/ - ├── control ← 32-byte fixed-size; updated atomically - ├── Article.seg ← phase 10's segment files - ├── ... - └── wal/ - ├── 0000000000000001.wal ← active WAL segment (16 MiB fallocated) - └── 0000000000000002.wal ← created at rollover -``` - -### Control file format (32 bytes) - -```text -[u8 magic[4] = b"WOCT"] -[u8 version = 1] -[u8 _pad[3] = 0] -[u64 last_durable_lsn LE] -[u64 created_unix_seconds LE] -[u32 crc32c] -``` - -Atomic update sequence (per [`./exploration/postgresql/buffer-and-checkpoint.md`](./exploration/postgresql/buffer-and-checkpoint.md)): - -```rust -fs::write("control.tmp", &bytes)?; -let f = File::open("control.tmp")?; -unsafe { libc::fsync(f.as_raw_fd()); } -fs::rename("control.tmp", "control")?; -let dfd = unsafe { libc::open(data_dir.as_ptr(), libc::O_RDONLY) }; -unsafe { libc::fsync(dfd); libc::close(dfd); } -``` - -### Engine integration - -`Engine::create / update / delete` — each calls into the WAL after the in-memory mutation succeeds and the segment append (phase 10) succeeds: - -```rust -let lsn = self.wal.append(WalRecord::Mutation { ty, op, row })?; -self.wal.enqueue_ack(/* request fd */ fd); -// loop tick later: drain_acks() runs after commit() fsyncs. -``` - -The HTTP handler in `crates/rt/src/server.rs` becomes: - -```rust -fn create_h(engine: &Shared, ty: &str, req: &Request, _params: &RouteParams) -> Response { - let body = parse_json_body(req); - let mut eng = engine.lock().unwrap(); - match eng.create(ty, body) { - Ok(row) => { - // Engine's create now returns *after* the WAL append, but BEFORE - // the fsync. The response is held until the next tick's group commit. - Response::deferred(Status::CREATED, json!(row)) - } - Err(e) => Response::status(Status::BAD_REQUEST).text(e.to_string()), - } -} -``` - -`Response::deferred` is a new variant — the response object is stashed on the connection, but the wire bytes aren't sent until `wal.drain_acks()` returns this fd. Phase 11 introduces this concept; phase 12 keeps it. - -### Recovery on startup - -```rust -fn recover(data_dir: &Path) -> Result { - let ctl = ControlFile::read_or_initialize(data_dir)?; - let mut engine = Engine::new(catalog); - let mut max_seen = ctl.last_durable_lsn; - for rec in Replay::from(data_dir.join("wal"), ctl.last_durable_lsn)? { - let rec = rec?; - match rec.payload { - WalRecord::Mutation { ty, op, row } => engine.apply_replay(ty, op, row)?, - } - max_seen = rec.lsn; - } - // Don't advance the control file yet — checkpoint (phase 12+) does that. - println!("[wo] recovered {} records, tail LSN {}", count, max_seen); - Ok(engine) -} -``` - -`Engine::apply_replay` is `Engine::create / update / delete` minus the `wal.append` callback (already-replayed records re-applied don't get re-WAL'd). - -### Group commit — the loop integration - -`crates/rt/src/bin/wo.rs`'s `serve_loop` gains: - -```rust -'outer: loop { - let events = eloop.wait_once(Some(Duration::from_secs(60)))?; - for ev in events { /* dispatch as before */ } - - // Group commit fence — runs once per tick after request dispatch. - if engine.lock().unwrap().wal.pending_commits() > 0 { - let _ = engine.lock().unwrap().wal.commit(); // one fdatasync - for fd in engine.lock().unwrap().wal.drain_acks() { - // Mark the connection writable; its queued response now flushes. - eloop.modify(fd, Interest::READ_WRITE, Token(fd as u64))?; - } - } -} -``` - -### `Cargo.toml` delta - -`crates/rt/Cargo.toml` adds `wal` as a path dep alongside `db`. Workspace adds `crates/wal` to the members list. - -## Exit criteria - -1. **`cargo build`** at root — `rt`, `db`, `wal` all compile. -2. **Unit tests in `crates/wal/src/`:** - - `wal_append_assigns_monotonic_lsn` — successive appends produce strictly increasing LSNs. - - `wal_rollover_at_segment_cap` — appending past `WAL_SEG_SIZE` opens segment 2 without losing tail. - - `replay_yields_records_in_order` — write 100, replay returns 100 in LSN order. - - `control_file_rename_on_write` — kill -9 between tmp-write and rename leaves old control intact. - - `crc_mismatch_aborts_replay` — corrupting one byte in WAL aborts replay with `CrcMismatch`. -3. **Integration test `crates/rt/tests/wal_recovery.rs`:** - - Starts `wo run docs/examples/blog` with `WO_LISTEN=127.0.0.1:0`. - - POSTs 100 articles via the http stack. - - Sends `kill -9` to the binary. - - Restarts; GETs all 100 back. -4. **The api.rest 20-assertion battery still passes byte-identically** under `WO_FSYNC=on` (default). -5. **`WO_FSYNC=off` env var** — when set, skips the `fdatasync` for tests that don't care about durability. Drops cold-restart-recovery latency to zero. -6. **`strace -e fdatasync,fsync,rename`** during a 5-commit run shows ~5 `fdatasync` calls (one per tick), no `fsync` (no rollover, no checkpoint yet), no `rename` (no control update yet — that's phase 12+). - -## Non-scope - -- **No checkpoint loop yet.** Phase 11 reads the control file at startup and writes it at clean shutdown only; periodic checkpoint lands with phase 12 or shortly after. Until then, recovery walks the entire WAL on every restart — fine for a sample workload, expensive for production. -- **No segment compaction.** Old WAL segments stay on disk forever in this phase. A future phase truncates after a checkpoint advances the control file past them. -- **No `io_uring`.** Synchronous `pwrite` + `fdatasync`. `io_uring` becomes interesting when the loop drains many fds per tick; phase 11's commit cadence doesn't need it. Layered on later. -- **No partial-record handling on torn writes.** A WAL segment is `posix_fallocate`'d up front, so partial-write torn-record on the leading edge of the file is the only scenario; the CRC trailer detects it and replay stops cleanly. -- **No multi-process recovery.** Single-binary invariant. -- **No engine cutover to disk reads.** `Engine::list / get` still walk the in-memory `HashMap` — phase 12. - -## Verification - -```bash -cargo build # rt + db + wal -cargo test --lib # all unit tests green -cargo test --test wal_recovery # the kill-9 integration test - -# durability smoke (manual) -WO_LISTEN=127.0.0.1:8765 cargo run --bin wo -- run docs/examples/blog & -PID=$! -sleep 1 -for i in 1 2 3 4 5; do - curl -sf -X POST http://127.0.0.1:8765/api/articles \ - -H 'Content-Type: application/json' \ - -d "{\"slug\":\"a$i\",\"title\":\"A$i\",\"author\":1,\"published\":true,\"meta\":{\"excerpt\":\"\",\"body_md\":\"\"}}" \ - > /dev/null -done -kill -9 $PID -WO_LISTEN=127.0.0.1:8765 cargo run --bin wo -- run docs/examples/blog & -PID=$! -sleep 1 -curl -s http://127.0.0.1:8765/api/articles | python3 -c "import json,sys;print(len(json.load(sys.stdin)))" -# expect: 5 -kill -INT $PID - -# strace check -strace -e fdatasync,fsync,rename -f -p $(pgrep -f 'target/debug/wo run') 2>&1 | head -20 - -# v1 untouched -cd reference/crates && cargo build && cargo test -``` - -## After this phase - -Durability is real but the engine is still `HashMap>` — every row, in full, lives in RAM. Phase 12 swaps the in-memory `Row` for an offset into the segment file and adds checkpoints, completing the transition to a durable, RAM-bounded engine. After phase 12 the runtime can serve a 10× larger dataset than fits in RAM without a redesign. diff --git a/docs/plan/12-engine-disk-cutover.md b/docs/plan/12-engine-disk-cutover.md deleted file mode 100644 index fbe6f42..0000000 --- a/docs/plan/12-engine-disk-cutover.md +++ /dev/null @@ -1,175 +0,0 @@ -# 12 — Engine cutover: rows live on disk - -> **Status: ⬜ not started** — the C prototype's phase B (mmap arena) is the proving ground. Board: [00-status.md](../00-status.md) - -**Context sources:** [`./10-storage-foundations.md`](./10-storage-foundations.md), [`./11-wal-and-recovery.md`](./11-wal-and-recovery.md), [`../runtime/database/03-inmemory-engine.md`](../runtime/database/03-inmemory-engine.md), [`../runtime/database/07-wo-seg-migration.md`](../runtime/database/07-wo-seg-migration.md), [`./exploration/postgresql/buffer-and-checkpoint.md`](./exploration/postgresql/buffer-and-checkpoint.md), [`./exploration/postgresql/page-format.md`](./exploration/postgresql/page-format.md), [`./exploration/linux/12-pwrite-fsync.md`](./exploration/linux/12-pwrite-fsync.md). - -## Goal - -`Engine`'s row payload is no longer in RAM. The in-memory map is `HashMap>`. Reads `pread` against the segment file and verify the CRC. RAM footprint is bounded by **id count + per-id overhead**, independent of row payload size — the runtime can now serve a dataset 10× larger than RAM. - -A periodic checkpoint flushes dirty segments and advances the control-file LSN, bounding recovery time on restart. - -## Design decisions (locked) - -1. **In-memory index is `BTreeMap`.** Keeps the existing `Engine::list` insertion-order iteration. Roughly 24 bytes per entry (i64 key + u64 value + tree node overhead) — a million rows fits in 24 MiB regardless of row size. -2. **Reads via `pread` + decode + CRC verify.** No user-space buffer pool — the OS page cache is the cache (per [`./exploration/postgresql/buffer-and-checkpoint.md`](./exploration/postgresql/buffer-and-checkpoint.md)). Hot rows hit cached pages and the syscall returns memcpy-fast. -3. **Bounded LRU on top of pread.** Optional small `HashMap<(ty, offset), Row>` capped at `WO_CACHE_ROWS=10000` (configurable). Avoids re-decoding on hot reads. Eviction on insert when full. **Phase 12 ships without it** if the bench numbers are fine; included here as a follow-on hatch. -4. **Tombstoned offsets stay in the BTreeMap until compaction.** A delete writes a tombstone to the segment + marks the BTreeMap entry as `SegmentOffset::Tombstone`. List skips them. Counts as wasted space until a future compaction phase rewrites the segment. -5. **Checkpoint = `fsync` every active segment fd + advance control file.** Runs every `CHECKPOINT_INTERVAL_SECS=60` (configurable) and at clean shutdown. -6. **MVCC stays out of scope.** Subscriber pre-commit views are tick-boundary semantics, not version chains (per [`../runtime/database/03-inmemory-engine.md`](../runtime/database/03-inmemory-engine.md)). When a future phase adds LIVE subscriber predicate matching, version chains may join the engine — until then, the single-thread invariant gives us the same visibility guarantees for free. -7. **Secondary indexes deferred.** Phase 12 ships only the primary `id` BTree. `unique` + `index` schema attributes get their own follow-on phase. - -## Scope - -### Files rewritten inside `crates/rt/src/` - -| File | Change | -| --- | --- | -| `engine.rs` | `BTreeMap` → `BTreeMap`. `Engine::get` becomes `seg_store.read(ty, offset)?`. `Engine::list` walks the BTreeMap and `pread`s each record (sequential — page cache makes it fast for the sample workload). `Engine::create / update / delete` keep the phase-10 segment append + phase-11 WAL append, but no longer keep the `Row` in memory. | -| `bin/wo.rs` | After WAL recovery, populate the BTreeMap with `(id → offset)` pairs by walking the recovered records. Also: spawn a `TimerFd::periodic(CHECKPOINT_INTERVAL_SECS)` registered on the event loop; the checkpoint step runs when the timer fires. | - -### New file inside `crates/db/src/` - -| File | Responsibility | Approx LOC | -| --- | --- | --- | -| `checkpoint.rs` | `Checkpoint::run(seg_store, wal, control)` — fsync every segment fd, write `last_durable_lsn = wal.tail_lsn` to the control file (rename-on-write), prune retired WAL segments older than the new LSN. | ~150 | - -### What `SegmentOffset` looks like - -```rust -#[derive(Debug, Clone, Copy)] -enum SegmentOffset { - Live(u64), // byte offset in the segment file - Tombstone(u64), // ditto, but the row is logically deleted -} -``` - -A `BTreeMap` consumes ~24 B per entry (key + 16-byte enum). 10M rows → ~240 MiB index. Order-of-magnitude bigger than `O(rowcount × pointer)` because the enum carries a discriminant; collapse to `u64` with a high-bit tombstone flag if memory pressure justifies it later. - -### Recovery (phase 11) becomes - -```rust -fn recover(data_dir: &Path) -> Result { - let ctl = ControlFile::read_or_initialize(data_dir)?; - let seg_store = SegStore::open(data_dir)?; - let wal = Wal::open(data_dir.join("wal"), ctl.last_durable_lsn)?; - let mut engine = Engine::new(catalog); - engine.attach(seg_store, wal); - - // Walk the segments first to populate the offset index from durable rows. - for ty in engine.catalog().order.iter() { - for (id, offset, flags) in seg_store.iter(ty)? { - engine.index_mut(ty).insert(id, match flags { - Flags::ACTIVE => SegmentOffset::Live(offset), - Flags::TOMBSTONE => SegmentOffset::Tombstone(offset), - }); - } - } - // Then replay any WAL records past the last checkpoint to catch up. - for rec in Replay::from(data_dir.join("wal"), ctl.last_durable_lsn)? { - engine.apply_replay(rec?)?; - } - Ok(engine) -} -``` - -The WAL replay still runs but covers a much smaller range — only what's been written since the last checkpoint. Recovery time is bounded by WAL volume between checkpoints, not by the entire history. - -### Checkpoint as a loop step - -```rust -let cp_timer = TimerFd::periodic(Duration::from_secs(60))?; -eloop.register(cp_timer.as_raw_fd(), Interest::READABLE, Token(cp_timer.as_raw_fd() as u64))?; - -// In serve_loop: -fd if fd == cp_timer.as_raw_fd() => { - let _ = cp_timer.read(); // drain timerfd's expirations - let mut eng = engine.lock().unwrap(); - Checkpoint::run(&eng.seg_store, &eng.wal, &mut eng.control)?; - println!("[wo] checkpoint at LSN {}", eng.control.last_durable_lsn); -} -``` - -The phase-02 `TimerFd::periodic` already exists; this is the first runtime caller for it. - -### Bench - -A small criterion-style microbench in `crates/rt/benches/engine_disk.rs`: - -| Test | Target | -| --- | --- | -| Insert 100k rows (50-byte payload) | < 5 s wall, < 50 MiB RSS at end | -| Random read 100k rows under steady-state load | < 5 µs p50, < 100 µs p99 (page cache hot) | -| Cold-cache read 100k rows | < 200 µs p50 (one disk seek per read) | -| Recovery time after kill -9 mid-bench | < WAL_volume / disk_throughput, dominated by `fdatasync` round-trips | - -`criterion` is normally an external crate; we're not adding deps. The bench is a `#[test]` with a `--release` runner — coarse but enough to catch regressions. - -### `Cargo.toml` delta - -None — `db` and `wal` are already in from phases 10 and 11. - -## Exit criteria - -1. **`cargo build`** at root, four direct deps unchanged (`anyhow`, `serde`, `serde_json`, `libc`). -2. **All existing unit tests still pass** after the engine rewrite. The two heaviest are `engine::tests::crud_roundtrip_auto_id` (port to verify offset semantics) and `server::tests::*` (HTTP-level CRUD — should be unaffected). -3. **End-to-end api.rest battery passes byte-identically** — same status codes, same JSON bodies, same key ordering. -4. **Integration test `crates/rt/tests/disk_engine.rs`:** - - Seed 10k rows of a 1 KiB payload type. Memory after seed (`/proc/self/status` `VmRSS`) is bounded by `id_count × 24 B + listener_overhead`, NOT by `10000 × 1024`. Specifically: less than 40 MiB. - - Restart with kill -9 mid-write; recovery completes in < 1 s for a 16-MiB-WAL-segment workload. - - GET random ids — every read returns the right row, CRC verified. -5. **Checkpoint smoke** — start the binary, write 5 rows, wait `CHECKPOINT_INTERVAL_SECS+1` seconds, verify `data/control` is updated (mtime moved, `last_durable_lsn` advanced). `strace -e fsync,rename` during the wait shows the checkpoint sequence. -6. **Cold start with no `data/`** — a fresh `wo run` on an empty data dir just works (creates the dir, no replay needed). Same for `data/` + empty WAL. - -## Non-scope - -- **No secondary indexes.** `unique` and `index` schema attributes still trigger no extra storage. Future phase. -- **No compaction.** Tombstoned offsets and old segment bytes accumulate. Trigger compaction is a separate phase keyed on a `dead-bytes / live-bytes` ratio. -- **No MVCC.** Subscriber pre-commit views are tick-boundary semantics (`docs/runtime/database/03-inmemory-engine.md`). Version chains land alongside the cross-shard subscription work in `09c-per-shard-wal` / `09d-cross-shard-subscriptions`. -- **No `O_DIRECT`.** Page cache is the cache. Per [`./exploration/linux/12-pwrite-fsync.md`](./exploration/linux/12-pwrite-fsync.md). -- **No `io_uring` reads.** `pread` syscalls are short and the loop has no other work waiting; an async batched read API isn't worth its own complexity at this size. -- **No streaming list.** `Engine::list` returns all rows for a type in one call. Pagination + cursor support is a future phase keyed on a real workload that hits the wall. - -## Verification - -```bash -cargo build -cargo test --lib # rt + db + wal unit tests -cargo test --test disk_engine # the new integration test -cargo test --release --test disk_engine -- --nocapture # bench numbers visible - -# manual end-to-end -cargo run --release --bin wo -- run docs/examples/blog & -PID=$! -sleep 1 -# Seed 10000 rows -for i in $(seq 1 10000); do - curl -sf -X POST http://127.0.0.1:8080/api/articles \ - -H 'Content-Type: application/json' \ - -d "{\"slug\":\"s$i\",\"title\":\"T$i\",\"author\":1,\"published\":true,\"meta\":{\"excerpt\":\"\",\"body_md\":\"\"}}" \ - > /dev/null -done -# Memory check -ps -o rss= -p $PID # expect under ~50 MiB even with 10k rows × 1 KiB each -# Wait for checkpoint -sleep 65 -ls -la docs/examples/blog/data/control # mtime should be recent -kill -INT $PID - -# Cold restart -cargo run --release --bin wo -- run docs/examples/blog & -PID=$! -sleep 1 -curl -s 'http://127.0.0.1:8080/api/articles' | python3 -c 'import json,sys;print(len(json.load(sys.stdin)))' -# expect: 10000 -kill -INT $PID - -cd reference/crates && cargo build && cargo test # v1 untouched -``` - -## After this phase - -The single-thread runtime is durable, RAM-bounded, and recovery-fast. Phases 13+ pivot to layering features on top: secondary indexes, compaction, query-layer integration, then the `09a-09f` scaleout sequence which lifts the same primitives into per-shard form. The empty `crates/{value, engine, txn}` skeletons get populated as their phases activate; `wal/` and `db/` are now real code, used by `rt/`. - -The `crates/rt/Cargo.toml` direct dep list at the end of phase 12 is `anyhow + serde + serde_json + libc + db + wal`. Phase 05 collapses `serde + serde_json` into the hand-rolled JSON module; phase 06 collapses `anyhow` into a bespoke error type. The `libc` + path-deps end state from [`./done/01-scafolding-crates.md`](./done/01-scafolding-crates.md) is reachable in two more phases past 12. diff --git a/docs/plan/13-class-model-live-pricing.md b/docs/plan/13-class-model-live-pricing.md deleted file mode 100644 index 3bcb7dc..0000000 --- a/docs/plan/13-class-model-live-pricing.md +++ /dev/null @@ -1,125 +0,0 @@ -# 13 — Class model + live pricing: state and methods, no inheritance - -> **Status: 🔄 in progress** — 13a ✅ shipped; 13b ✅ shipped (methods execute over RPC); 13c (LIVE push) is next; 13d ⏸ parked (frontend); 13e ⬜. Board: [00-status.md](../00-status.md) - -**Context sources:** [`../runtime/database/02-wo-language.md`](../runtime/database/02-wo-language.md) (schema layer, § Schema-Layer DML brace disambiguation, § Cross-Paradigm Transaction Coordinator), [`../runtime/database/04-client-api.md`](../runtime/database/04-client-api.md) (subscription engine), [`./09-concurrency-scaleout.md`](./09-concurrency-scaleout.md) (thread-per-core scale-out), `./exploration/ui/00-overview.md` + `./exploration/ui/01-htmlx-format-spec.md` (live UI), [`../examples/pricing/`](../examples/pricing/) (the demo this phase makes real), [`../examples/ecommerce/shared/logic/checkout.wo`](../examples/ecommerce/shared/logic/checkout.wo) (the existing `fn … in txn snapshot` signature style methods reuse). - -## Context - -writeonce is declarative by design — `type`, `service`, `policy`, `on ` — and the docs explicitly reject OO. But developers arriving from OO languages keep reaching for "a class with methods", and the request has a legitimate core: **behavior that belongs to a row** (`product.set_price(amount)`) is today only expressible as a free `fn` or a trigger. This phase adds the smallest class model that satisfies it: - -> **`class` = state + methods. No inheritance, no override, no polymorphic dispatch — ever.** Composition via `ref` / `multi`, exactly like `type`. Go-style encapsulation, not Java-style hierarchies. - -The driving workload is the [`pricing` demo](../examples/pricing/): a `Price` class and a `Product` class with methods, products owning prices, a `##ui pricing` screen showing live prices of selected products — RAM-resident data, one product readable by millions of customers at once, price updates pushed live to millions of subscribers, all I/O on kernel primitives. - -This doc is a **master plan** in the style of [`09-concurrency-scaleout.md`](./09-concurrency-scaleout.md): sub-phases 13a–13e at a high level, each landing as its own plan doc when implementation starts. No code changes in this pass — the demo project ships as a design artifact alongside this doc. - -## Goal - -`cargo run --bin wo -- run docs/examples/pricing` serves the demo fully live: `class` declarations parse and store like types, methods execute as row-scoped transactions over RPC, every `set_price` commit pushes a delta to all subscribed clients, the `##ui pricing` screen patches price cells in place, and the read path scales per the phase-09 architecture. - -## Design decisions (locked) - -1. **`class` is the behavior-bearing sibling of `type`.** Identical field grammar — scalars, defaults, `@unique`/`@check`, embedded docs, `ref`, `multi`, `backlink`, unions, plus the same attachable blocks (`service`, `policy`, `on `). One addition: `fn` methods. -2. **Methods are row-scoped transactional functions.** `fn name(args) -> Ret [in txn [snapshot]]` with an implicit `self` bound to the receiving row. Same signature grammar and execution machinery as the free-standing `fn checkout(...) in txn snapshot` already in the ecommerce sample — a method is a free `fn` with a hidden first parameter. No new transaction semantics. -3. **No inheritance.** No `extends`, no `override`, no virtual dispatch, no abstract classes. This kills the table-per-class/single-table storage mapping problem before it exists: a class IS one table (+ its doc/graph parts), exactly like a type. "Is-a" modelling uses tagged unions (already in the language); "has-a" uses `ref`/`multi`. -4. **`self` stays an identifier in the lexer.** Same gotcha as `subscribe`/`receive`/`me` (CLAUDE.md): it must remain usable as a plain name in expose lists and expressions. The parser binds it positionally inside method bodies. -5. **Storage and REST are class-blind.** `Catalog::from_schemas` treats a class exactly like a type; `service rest` blocks generate the same CRUD routes. Methods add RPC routes on top (13b). A migration from `type` to `class` (or back, if no methods) is a no-op for stored data. -6. **Spec wording is amended in 13a, not before.** `docs/runtime/wo-language.md` ("Isn't OO") and `docs/writeonce-pl.md` ("no class model") change in the same commit that makes the parser accept the syntax, so docs never describe an unparseable language. - -## The class surface (normative example) - -```wo -@table(name: "prices", index: [product, at]) -class Price { - id: Id - product: ref Product - amount: Money -- minor units, stdlib scalar - currency: Text = "EUR" - at: Timestamp = now() - - -- Pure method: computes from self, touches nothing else. No txn needed. - fn discounted(pct: Int) -> Money { - return self.amount * (100 - pct) / 100; - } -} - - -@table -class Product { - id: Id - sku: SKU @unique - name: Text - prices: multi Price -- products have prices (append-only history) - - fn current_price() -> Money in txn { - return latest(self.prices).amount; - } - - fn set_price(amount: Money) in txn { - insert Price { product: self.id, amount: amount }; - } - - service rest "/api/products" - expose list, get, create, update, delete, subscribe -} -``` - -**`@table` — storage configuration, never storage declaration** (✅ shipped, with DML support). Every `type`/`class` IS a table (decisions 3/5 stand); the optional type-level `@table(...)` annotation *configures* it: `name: "prices"` sets the storage/table name (catalog-unique; the surface plans 10–12 and the SQL layer consume — WAL records keep the type name as the stable identifier), and each `index: [product, at]` declares a composite secondary index that the engine maintains on every mutation path (CRUD, method-txn undo, WAL replay) and that accelerates `self.prices` relation reads, the schema-layer `select Price{ product == self.id }` expression in method bodies, and `GET /api/prices?product=1` REST filters — all through one `Engine::find_by` path with scan fallback. Bare `@table` is a legal no-op. Unknown keys (`shard_key`, `retention`) are parse errors until their phases land; unknown annotation *names* skip silently. Spec: [`02-wo-language.md § Type-Level Annotations`](../runtime/database/02-wo-language.md). - -## Sub-phase sequence - -Each lands as its own numbered plan doc (`13a-…`, `13b-…`) when ready. The blog + ecommerce smoke stays green after every sub-phase. - -### `13a-class-surface.md` — lexer, parser, AST, spec amendments — ✅ shipped - -`class` joins the keyword map (`crates/rt/src/lexer.rs` keyword match, ~line 202 — note `self` stays an ident per decision 4). `parse_type` (`crates/rt/src/parser.rs:120`) takes the leading keyword as a parameter and serves both constructs; `fn` members inside the body parse-and-discard through the existing brace-depth skip — the same mechanism that already swallows `on update … do { … }` triggers. `ast::TypeDecl` gains `is_class: bool`; `Catalog::from_schemas` ignores it (decision 5), so REST CRUD works the moment parsing does. Docs amended in the same change: a "Class Model" subsection in [`02-wo-language.md`](../runtime/database/02-wo-language.md) next to § Schema-Layer DML, the "Isn't OO" paragraph in `wo-language.md`, the class line in `writeonce-pl.md`, and `just pricing` / `just pricing-demo` recipes. -**Exit (met):** `wo run docs/examples/pricing` parses 2 classes, serves `/api/products` CRUD; parser unit tests (`parses_class_with_methods`, `class_method_braces_do_not_truncate_body`) green; blog/ecommerce/hello unchanged. - -### `13b-method-execution.md` — methods over RPC — ✅ shipped - -Method bodies compile into a real AST (`ast::{MethodDecl, Stmt, Expr}` — `let`, `insert Type{…}` construction, `return`, `assert … otherwise abort`, `if/else`, arithmetic/comparison/boolean expressions, `self.` resolution, built-ins `latest`/`count`/`now`) and execute in `crates/rt/src/method.rs` on the shard that owns the receiving row (`run_on(owner_of(id))` — inserts mint locally there, so a method owns everything it writes). Every call runs inside an engine **method transaction** (`Engine::{begin,commit,abort}_txn`): mutations journal undo entries and defer their WAL records; commit emits **one `WalRec::Txn` frame** — the frame's CRC makes a method's mutations replay whole-or-not-at-all, so a crash mid-method can never half-apply — and abort reverts RAM in reverse order. Group-commit acks gate on the batch fsync exactly like CRUD (`Response.gate` / parked replies). Routes: `POST /:id/` for every method of a class with a `service rest` block; errors map NoSuchRow→404, BadArgs→400, Abort→409, Exec→500. Lowercase `insert` stays an identifier in the lexer (same rule as `subscribe`/`me`/`self`); only the class-side `fn` arm parses bodies — `fn` inside a plain `type` keeps the 13a skip. -**Exit (met):** `curl -X POST /api/products/1/set_price -d '{"amount": 4999}'` inserts a Price atomically (verified on 2 shards incl. the cross-shard hop, in both WAL modes, and across restart — the Txn frame replays); `current_price` returns it; `set_price {"amount": 0}` trips the sample's assert → 409 and rolls back completely (price unchanged). 12 new unit tests (parser AST, executor, WAL atomicity, route statuses); blog/ecommerce/hello unchanged; `just pricing-demo` runs the whole sequence. - -### `13c-live-pricing-push.md` — LIVE deltas on commit - -The subscription registry from [`04-client-api.md`](../runtime/database/04-client-api.md): keyed by type + predicate, matched on commit, deltas framed over WebSocket. Replaces the 501 stub at `/api//live` (`crates/rt/src/server.rs`) with a real upgrade for the pricing demo's needs — `LIVE select Product{ name, prices }` and the `subscribe` expose. This is the Stage 3 milestone scoped to one workload; the full wire protocol stays in Phase 4. -**Exit:** two terminals — `websocat /api/products/live` in one, `set_price` via curl in the other — the delta frame arrives on the open socket within one commit tick, no polling. - -### `13d-pricing-ui.md` — the `/pricing` screen, MVC - -The screen ships as an **MVC triplet** per `exploration/ui/08-mvc-structure.md`, built in the sub-phase sequence of `14-mvc-ui-implementation.md`: model = the classes themselves, view = [`pricing.htmlx`](../examples/pricing/ui/pricing/pricing.htmlx) (plain htmlx, logic-free) + external [`pricing.scss`](../examples/pricing/ui/pricing/pricing.scss) (strict SCSS subset compiled at `wo build`, no external deps), controller = [`pricing.wo`](../examples/pricing/ui/pricing/pricing.wo) (`route:`/`view:`/`styles:`, `model:` bindings, `actions:` calling the 13b class methods). SSR per `exploration/ui/01-htmlx-format-spec.md`, compiler glue per `02-ui-compiler.md`, and the vanilla-JS client runtime (`03-client-runtime.md`) patches the price cell when the 13c delta lands. The controller's `model:` block is the M→V binding; the watchlist narrows the subscription predicate server-side. -**Exit:** browser at `/pricing` shows selected products; a `set_price` commit from curl changes the price cell in every open browser without reload. - -### `13e-pricing-at-scale.md` — millions of readers, millions of live updates - -No new architecture — this sub-phase wires the demo to [`09-concurrency-scaleout.md`](./09-concurrency-scaleout.md) and adds one mechanism: - -- **RAM-resident:** the engine is the in-memory design of [`03-inmemory-engine.md`](../runtime/database/03-inmemory-engine.md); disk (phases 10–12) is durability behind it, never the read path. -- **Read fan-out:** thread-per-core, shared-nothing shards behind `SO_REUSEPORT` (09a/09b). A single hot product row is owned by one shard but **read-replicated to every shard**: each thread keeps a read-only copy of hot rows, refreshed by the same per-thread broadcast that 09d uses for subscriber fan-out. Millions of concurrent `GET /api/products/1` spread across all cores and never contend — writes still serialize on the owning shard, preserving the single-writer model. -- **Live fan-out to millions:** one `set_price` commit → one delta message per thread (09d, not per subscriber) → each thread predicate-matches its local subscription table and batches socket writes on its own `io_uring` ring ([`exploration/linux/07-io_uring.md`](./exploration/linux/07-io_uring.md)). Kernel primitives only: epoll today ([`done/02-event-loop-epoll.md`](./done/02-event-loop-epoll.md), edge-triggered + group commit), io_uring per phase 09. - -**Exit: verification targets defined and a load harness scripted** (not necessarily met on dev hardware): - -| Metric | Target | How measured | -| ------------------------------------------- | ---------------------------------------------- | ----------------------------------------------------------------- | -| Concurrent readers of one product | 1 M req/s aggregate on 16 cores | `wrk -c 10000` against `GET /api/products/1`, hot-row replicas on | -| Live subscribers receiving one price update | 1 M open sockets, delta delivered p99 < 250 ms | `websocat` fan-out harness, timestamped frames | -| Commit→first-delta latency | p99 < 10 ms | in-process timestamp at commit vs first socket write | -| Memory | ~2 KB/connection + engine working set | RSS under subscriber load | -| Dep count | 1 (`libc`) | `crates/rt/Cargo.toml` unchanged by this phase | - -## Non-scope - -- **No inheritance, ever, under this plan.** If hierarchy modelling pressure appears, the answer is tagged unions and composition; a future interfaces/traits proposal would be its own phase with its own doc. -- **No method overloading, no statics, no constructors.** Row creation stays `insert` / REST `create`; one method name per class. -- **No client-side method stubs** — `wo gen sdk` method support belongs to Phase 5 (Go SDK), not here. -- **No multi-node distribution.** Same stance as phase 09: single box, threads-as-shards. - -## Cross-references - -- [`../examples/pricing/`](../examples/pricing/) — the demo project this plan makes real, file-by-file phase map in its README. -- [`../runtime/database/02-wo-language.md`](../runtime/database/02-wo-language.md) — schema layer the class grammar extends; transaction coordinator methods reuse. -- [`../runtime/database/04-client-api.md`](../runtime/database/04-client-api.md) — subscription engine 13c scopes down. -- [`./09-concurrency-scaleout.md`](./09-concurrency-scaleout.md) — the scale architecture 13e instantiates. -- [`../examples/hello/main.wo`](../examples/hello/main.wo) — the minimal example whose `Revision`-trigger pattern is the declarative ancestor of methods. diff --git a/docs/plan/15-mcp-streamable-http.md b/docs/plan/15-mcp-streamable-http.md deleted file mode 100644 index be6b145..0000000 --- a/docs/plan/15-mcp-streamable-http.md +++ /dev/null @@ -1,124 +0,0 @@ -# 15 — MCP over Streamable HTTP: every writeonce app is an MCP server - -> **Status: ⬜ not started (Track 4 — Language & API)** — board: [00-status.md](../00-status.md) - -**Context sources:** [MCP specification 2025-06-18 — Transports](https://modelcontextprotocol.io/specification/2025-06-18/basic/transports) (the normative Streamable HTTP contract this plan implements, verified 2026-07-12), [`reference/mcp-python-sdk/`](../../.dev/reference/README.md) (symlink to the official MCP Python SDK — grep `src/mcp/server/streamable_http.py` + `streamable_http_manager.py` for the reference server behaviour, `src/mcp/client/streamable_http.py` for what a conforming client expects; behaviour is ported, code is not), [`../runtime/database/04-client-api.md`](../runtime/database/04-client-api.md) (the wire-protocol design; its "REST + SSE gateway" row is what this plan makes concrete for agents), [`./13-class-model-live-pricing.md`](./13-class-model-live-pricing.md) (13b methods become MCP tools; 13c's subscription registry carries 15e), [`./09-concurrency-scaleout.md`](./09-concurrency-scaleout.md) (thread-per-core + shard bus the endpoint rides; 09d fan-out gates 15e), `crates/rt/src/server.rs` + `crates/rt/src/http/` (the keep-alive HTTP layer and router this lands in). - -## Context - -**MCP** (Model Context Protocol) is the open JSON-RPC 2.0 protocol LLM agents use to discover and call external capabilities: **tools** (typed functions), **resources** (readable content addressed by URI), and change **notifications**. **Streamable HTTP** is its HTTP transport (spec 2025-06-18, replacing the 2024-11-05 HTTP+SSE pair): one endpoint path serving POST and GET, where every client→server JSON-RPC message is a POST, the server answers each *request* with either a single `application/json` body **or** a `text/event-stream` (SSE) that may carry server messages before the final response, a GET opens a server→client SSE stream for unsolicited notifications, and an `Mcp-Session-Id` header carries optional stateful sessions. - -The fit with writeonce is unusually direct. The runtime already is the database, the schema authority, and the HTTP server in one binary; REST routes are *generated* from `type` declarations and their `service … expose` lists. MCP is the same generation problem with a different wire shape: the catalog becomes `tools/list` and `resources/templates/list`, the engine's CRUD paths become `tools/call`, and — once Stage 3/13c lands — committed deltas become `notifications/resources/updated`. A `.wo` app then serves browsers (REST/htmlx), programs (REST), and agents (MCP) from one catalog, one engine, one port. - -What exists today that this plan builds on: the hand-rolled epoll HTTP layer with keep-alive and pipelining (plan 03 + the 09 follow-up), thread-local routers per `SO_REUSEPORT` worker, sharded engine with owner-hop/fan-out on the shard bus (09b), and durable ack gating on the io_uring group commit (`Response.gate` — a mutation's response is parked until its fsync CQE). What does **not** exist yet: any non-buffered response (every `Response` is a full `Vec` with `Content-Length`), sessions, and the subscription registry (13c). - -## Transport contract (normative summary) - -The rules the sub-phases implement, condensed from the spec — each MUST below is the spec's, not ours: - -| # | Rule | -| --- | --- | -| T1 | One endpoint path (`/mcp`) MUST support POST and GET. Body of a POST is a **single** JSON-RPC message (batching is gone in 2025-06-18). | -| T2 | POSTed *request* → server returns `Content-Type: application/json` (one object) **or** `text/event-stream` (SSE stream that eventually carries the response, then SHOULD close). The server chooses; clients MUST support both. | -| T3 | POSTed *notification*/*response* → `202 Accepted`, no body (or 4xx if rejected). | -| T4 | GET → SSE stream for server-initiated messages, **or** `405 Method Not Allowed`. No JSON-RPC *responses* on a GET stream except when resuming. | -| T5 | Sessions: server MAY return `Mcp-Session-Id` on the `InitializeResult` response; clients MUST echo it on all subsequent requests; missing → `400`; terminated/unknown → `404` (client then re-initializes); client DELETE terminates a session (server MAY answer `405`). | -| T6 | `MCP-Protocol-Version` header required on post-initialize requests; absent → assume `2025-03-26`; invalid/unsupported → `400`. | -| T7 | Resumability: SSE events MAY carry `id:` (unique per stream, acting as a per-stream cursor); client reconnects with `Last-Event-ID`; server MAY replay messages from that stream only. | -| T8 | Security: server MUST validate `Origin` (DNS-rebinding defence), SHOULD bind localhost when local, SHOULD authenticate. | - -## Goal - -`cargo run --bin wo -- run docs/examples/blog` serves `POST /mcp` alongside `/api/*`: an MCP client (MCP Inspector, Claude Code, or a curl script) performs `initialize` → `tools/list` → `tools/call article_create` → `tools/call article_list` and sees its write — with the ack held for the fsync CQE exactly as REST does. After 15e (with 13c + 09d): `resources/subscribe` on `wo://product/1`, a `set_price` commit in another terminal, and `notifications/resources/updated` arrives on the open SSE stream — the agent-shaped twin of the 13d browser demo. - -## Design decisions (locked) - -1. **One endpoint, same workers.** `/mcp` is a route in the existing thread-local `Router` — no second listener, no port, no dedicated thread. It scales the way `/api/*` does: `SO_REUSEPORT` spreads connections, shard bus routes data ownership. -2. **Catalog-driven and class-blind.** Tools and resources are generated from the catalog + expose lists, never hand-registered — the 13a doctrine (storage/REST class-blind) extends to MCP. Until 15d, the existing `service rest … expose` list governs what MCP exposes; 15d adds `service mcp` for independent control. -3. **JSON first, streaming second.** 15a–15b answer every POSTed request in `application/json` mode — spec-legal per T2 — so the MCP surface is useful before any streaming machinery exists. SSE (T2's other arm, T4, T7) is additive in 15c/15e. -4. **The durable-ack rule is transport-independent.** A `tools/call` that mutates parks its JSON-RPC response on the group-commit gate exactly like a REST POST (`Response.gate` / `Parked` machinery from 09c). An MCP client never observes a result for a non-durable write. -5. **Sessions are worker-owned.** The worker that serves `initialize` mints `Mcp-Session-Id = w-<128-bit hex>`; the embedded worker index lets any other worker forward session-scoped work over the shard bus (`run_on`, the existing point-op machinery). Stateless until 15c — no session header is issued, which the spec permits. -6. **Protocol version `2025-06-18`.** Negotiated at `initialize`; absent header → assume `2025-03-26` (T6 — identical for the surface served here); anything else → `400`. -7. **`serde_json` for now.** Same dependency posture as the rest of `rt`; migrates when phase [05](05-hand-rolled-json.md) lands. **No MCP SDK crates** — the protocol layer is hand-rolled like the HTTP layer, per the zero-deps north star. -8. **Origin validated on every `/mcp` request** (T8): allow absent-Origin (non-browser clients) and a `WO_MCP_ORIGINS` allowlist defaulting to localhost origins; anything else → `403`. The localhost-bind guidance is already satisfied — `WO_LISTEN` defaults to `127.0.0.1:8080`. Authentication is deferred (non-scope; ties to the policy phase). - -## Dependency graph - -``` -15a JSON-RPC core + tools ──→ 15b resources ──→ 15c SSE + sessions ──→ 15e LIVE subscriptions - │ (first streaming (needs 13c + 09d) - │ response in rt) - └──→ 15d `service mcp` surface (parser-only; any time after 15a) -15a needs nothing that isn't shipped: router, sharded engine, group commit. -``` - -## Sub-phase sequence - -### `15a-jsonrpc-core-and-tools.md` — the endpoint speaks MCP, tools work - -- **Endpoint + envelope**: `POST /mcp` in `server.rs`; parse a single JSON-RPC 2.0 message (T1); protocol errors as JSON-RPC errors (`-32700` parse, `-32600` invalid request, `-32601` method not found, `-32602` invalid params). Notifications/responses → `202` empty (T3). `GET /mcp` and `DELETE /mcp` → `405` (T4/T5 — legal until 15c). -- **Header plumbing**: surface `Accept`, `Origin`, `MCP-Protocol-Version` (and later `Mcp-Session-Id`, `Last-Event-ID`) on `http::Request`; enforce decisions 6 and 8. -- **Lifecycle**: `initialize` (version negotiation; capabilities `{tools: {listChanged: false}}`; `serverInfo` from the app directory name + crate version), `notifications/initialized`, `ping`. -- **Tool generation**: per exposed type×op → `_list`, `_get`, `_create`, `_update`, `_delete`, with `inputSchema` (JSON Schema) derived from catalog field types (unions → `enum`, embedded structs → nested `object`) — same source of truth as `describe_routes`. -- **`tools/call` dispatch** through the *same* handler paths REST uses: creates local, point ops `run_on(owner_of(id))`, lists fan out — no second data path. Engine/validation failures return `isError: true` inside the tool *result* (the MCP rule: execution errors are results, protocol errors are JSON-RPC errors). Mutations park on the WAL gate (decision 4). - -**Exit:** scripted flow (checked in beside [`reference/rest/`](../../.dev/reference/rest/README.md)) against the blog sample passes: `initialize` → `202` for `initialized` → `tools/list` enumerates exactly the exposed ops → `article_create` → `article_list` shows the row; runs green with `WO_GROUP_COMMIT` on and off; `GET`→405, `DELETE`→405, bad version→400, disallowed Origin→403; unit tests in the `server.rs` style cover envelope errors and gate parking. - -### `15b-resources.md` — the schema and rows become addressable - -- URI scheme: `wo://schema/` (field/shape listing as JSON) and `wo:///` (one row). `resources/list` returns the schema resources (bounded); `resources/templates/list` returns `wo:///{id}` per exposed type; `resources/read` resolves both forms (row reads owner-hop like REST GET). Opaque id-based `nextCursor` pagination on list endpoints. -- Capabilities gain `resources: {subscribe: false, listChanged: false}` (flips in 15e). - -**Exit:** `resources/read wo://articles/1` body-equals `GET /api/articles/1`; templates enumerate every exposed type; unknown URI → resource-not-found error (`-32002`); cursor walks a 3-page listing without duplication or loss. - -### `15c-sse-and-sessions.md` — the "streamable" half - -- **First streaming response in the runtime**: a streaming variant beside the buffered `Response` (`ConnState::Streaming`) that writes SSE frames (`event: message\ndata: \n\n`) incrementally under epoll writability, honours backpressure (a slow reader parks on `EPOLLOUT`, never blocks the worker), and holds the connection out of keep-alive reuse until the stream closes. This is the piece 15e and Stage 3 inherit. -- **POST answering mode**: requests that will emit interim server messages answer in `text/event-stream` mode (response as the final SSE event, then close — T2); plain requests stay JSON. In 15c itself only long `tools/call`s use it; the machinery is the deliverable. -- **Sessions** (T5, decision 5): `Mcp-Session-Id` minted at `initialize`; missing on later requests → `400`; unknown → `404`; `DELETE /mcp` terminates → `200`. Per-worker session table; cross-worker requests forward via the worker index in the id. -- **`GET /mcp`** opens the session's server→client SSE stream (heartbeat comments to keep intermediaries happy; never carries responses — T4). - -**Exit:** the spec's own sequence diagram replayed end-to-end by script (init+session → 202 → JSON answer → GET stream stays open across ≥2 heartbeats); 400/404/DELETE conformance matrix green; a deliberately unread client stalls only its own connection (other connections' p99 unaffected, measured). - -### `15d-service-mcp-surface.md` — the language names the capability - -- Parser: `ServiceKind::Mcp` + an ident arm for `mcp` in `parse_service` (ident, not keyword — the `expose` gotcha stands); `service mcp "/mcp" expose list, get, set_price` inside a `type`/`class` controls generation independently of REST. Precedence: `service mcp` present → it alone governs MCP exposure; absent → fall back to the `service rest` list (15a behaviour, now documented in the spec doc [`02-wo-language.md`](../runtime/database/02-wo-language.md)). -- **Methods become tools**: a 13b class method in an `expose` list generates `_` with `inputSchema` from the method's parameter list — the agent-facing twin of `POST /api//:id/`. (Parses and lists from this phase; round-trips once 13b ships.) - -**Exit:** parser tests for the new arm and precedence; the pricing sample gains a `service mcp` block; `tools/list` reflects it (method tools listed; callable gated on 13b). - -### `15e-live-subscriptions.md` — commits push to agents - -- Capabilities flip to `resources: {subscribe: true}`. `resources/subscribe {uri: wo:///}` registers a keyed (O(1)) subscription in the 13c registry, bound to the session's GET stream; commit → `notifications/resources/updated {uri}` pushed as an SSE event; cross-shard commits reach the session's worker via 09d fan-out; `resources/unsubscribe` and session teardown free registry slots (the `EPOLLHUP` → unsubscribe philosophy of the v1 datalayer). -- **Resumability** (T7): per-stream monotonic SSE `id:`s; a bounded per-session ring buffer of undelivered notifications; reconnect `GET` with `Last-Event-ID` replays from the cursor, stream continues. - -**Exit:** two-terminal demo — subscribe to `wo://products/1` over the GET stream, `set_price` via REST curl in the other terminal, the notification arrives without polling; kill the client mid-stream, reconnect with `Last-Event-ID`, the missed notification is replayed exactly once. **Requires 13c + 09d.** - -## Verification targets (after 15e) - -| Check | Target | How | -| --- | --- | --- | -| Spec conformance | T1–T8 matrix green (status codes, headers, content types) | scripted curl flow checked in beside `reference/rest/` | -| Interop | MCP Inspector connects, lists tools/resources, calls a tool | manual check, noted per release | -| Parity | `tools/call _get` ≡ `GET /api//:id` byte-for-byte on the row payload | unit test | -| Durability | mutation results never precede their fsync CQE (`WO_GROUP_COMMIT` on) | gate test in `server.rs` style | -| Latency | `tools/call` read p99 within 1 ms of the REST equivalent under the plan-09 bench load | bench harness rerun | -| Dep budget | no new crates; `serde_json` only, dropped with phase 05 | `Cargo.toml` review | - -## Non-scope - -- **No 2024-11-05 HTTP+SSE backwards compatibility.** Only Streamable HTTP; old-transport clients are not served. -- **No stdio transport.** A `wo mcp-stdio` subcommand would be cheap later; out of scope here. -- **No authorization.** The MCP auth spec (OAuth 2.1) waits for the policy phase; until then the endpoint trusts what the Origin check and bind address admit. -- **No prompts capability, no client-feature counterparts** (sampling, elicitation, roots) — server capabilities only. -- **No JSON-RPC batching** — removed from the protocol in 2025-06-18; single message per POST, enforced. -- **No WebSocket.** MCP rides SSE only; the 13c browser WebSocket at `/api//live` is a separate surface sharing the same registry. - -## Cross-references - -- [`../runtime/database/04-client-api.md`](../runtime/database/04-client-api.md) — the protocol-tier survey; this plan implements its "REST + SSE gateway" row for agents, on the same subscription registry it specifies. -- [`./13-class-model-live-pricing.md`](./13-class-model-live-pricing.md) — 13b gates method tools (15d); 13c gates 15e; 13d's demo has an agent-shaped twin in 15e's exit. -- [`./09-concurrency-scaleout.md`](./09-concurrency-scaleout.md) — 09d gates cross-shard notification fan-out (15e); the shard-bus ownership rules 15a/15c reuse. -- [`./05-hand-rolled-json.md`](05-hand-rolled-json.md) — removes this plan's `serde_json` use when it lands. -- [`./07-inotify-content-watcher.md`](07-inotify-content-watcher.md) — a future `notifications/tools/list_changed` on hot reload would pair with it (not scheduled). -- [MCP specification 2025-06-18](https://modelcontextprotocol.io/specification/2025-06-18/basic/transports) — the normative transport text summarized in T1–T8. diff --git a/docs/plan/16-postgres-mirror.md b/docs/plan/16-postgres-mirror.md deleted file mode 100644 index 1283ad5..0000000 --- a/docs/plan/16-postgres-mirror.md +++ /dev/null @@ -1,92 +0,0 @@ -# 16 — PostgreSQL mirror: RAM-authoritative database, Postgres as the backup - -> **Status: 🔄 in progress (Track 3 — Storage & durability)** — 16a ✅, 16b ✅ shipped; 16c–16f ⬜. Board: [00-status.md](../00-status.md) - -**Context sources:** [`README.md` § persistent database](../../README.md) (the product goal this implements: *"reads and writes database to RAM, persist data to postgres SQL"*), [`../runtime/database/03-inmemory-engine.md`](../runtime/database/03-inmemory-engine.md) (RAM-resident doctrine: disk sits behind the read path, never in front), [`./09-concurrency-scaleout.md`](./09-concurrency-scaleout.md) (per-shard WAL + ack-after-fsync this rides behind), [`./13-class-model-live-pricing.md`](./13-class-model-live-pricing.md) (the Product/Price worked example; `@table(name: "prices")` names the mirrored table), [`../runtime/database/07-wo-seg-migration.md`](../runtime/database/07-wo-seg-migration.md) (the dual-write precedent), `reference/postgresql/` (research symlink — `src/include/libpq/` for the wire protocol), PostgreSQL docs *Frontend/Backend Protocol*. - -## Context - -The engine is RAM-resident by design and durable through its own per-shard WAL (09c). What's missing is an **external, queryable, operator-friendly copy** of the data — something a DBA can point `psql`, Grafana, or a nightly `pg_dump` at. That's what Postgres is here: **a backup mechanism**, not a storage engine. - -The contract, in one line each: - -- **The entire database lives in RAM.** Reads never touch Postgres — ever. -- **Writes go to RAM (+ WAL) and to Postgres** — but the Postgres write is asynchronous: the client's ack gates on the WAL fsync exactly as before; the mirror follows behind. -- **Postgres is disposable.** RAM is authoritative, so backup repair is always "re-push RAM state" — never a merge. - -Worked example throughout: `Product.set_price(amount)` from the [pricing demo](../examples/pricing/) → a `Price` row in RAM → the same row visible in `psql` as `SELECT * FROM prices`. - -## Design decisions (locked) - -1. **Async mirror, never a commit path.** The `wo-pg` thread is downstream of the commit: shard engines clone committed mutations onto a bounded channel (`try_send` — a full channel drops loudly, it never blocks a worker). Postgres being down costs clients nothing. -2. **Mirror what RAM holds.** The tap emits full post-merge rows (`MirrorRec::Upsert{ty, id, row}` — unlike `WalRec::Update`, which carries only the merge body), so a Postgres row is always byte-equivalent to its RAM row. Method transactions (13b) mirror as one `MirrorRec::Txn` → one `BEGIN…COMMIT` — an aborted method never reaches the channel at all. -3. **Hand-rolled wire client, zero new crates.** `crates/rt/src/pg.rs` speaks protocol v3 (startup, auth `trust`/`password`/`md5` with a hand-rolled MD5 — the CRC32 precedent; SCRAM is 16f) over a blocking `std::net::TcpStream`, **simple query protocol only**. Blocking is fine: the only caller is the dedicated mirror thread. -4. **One `wo-pg` thread per process.** N shard senders → one receiver; per-row ordering is preserved because a row's mutations always come from its owner shard (one FIFO sender). Batches drain opportunistically; statement errors are isolated per record (logged, skipped), socket errors reconnect with capped backoff. -5. **Boot = full resync.** Engines attach the mirror AFTER WAL replay and push their entire state as upserts (`mirror_sync_all`). Consequence: restarting `wo` against a fresh/empty/behind Postgres converges it — verified live (a record dropped during an outage reappeared after restart). -6. **Schema 16b: one JSONB table per type** — `"" (id BIGINT PRIMARY KEY, row JSONB NOT NULL)`, named by `@table(name: "prices")` (default: the type name). Upsert = `INSERT … ON CONFLICT (id) DO UPDATE`. Typed columns are 16c. -7. **Config: `WO_PG=postgres://user[:pass]@host[:port]/db`** env var (the `WO_DATA`/`WO_LISTEN` convention). Unset = mirror off, zero cost. -8. **The WAL stays the recovery source; Postgres is the backup of last resort.** Boot replay reads the WAL as today; restoring FROM Postgres (WAL lost) is 16e's explicit opt-in. - -## Sub-phase sequence - -### `16a` — hand-rolled wire client — ✅ shipped - -`crates/rt/src/pg.rs` (~450 lines, stdlib only): `PgConfig::from_url`, `Conn::connect` (startup + auth trust/cleartext/md5, RFC-1321 MD5 hand-rolled with test vectors), `simple_query` (RowDescription/DataRow/CommandComplete/ErrorResponse/ReadyForQuery; a backend error drains to ready so the connection stays usable), `escape_literal`/`escape_ident`. -**Exit (met):** unit tests (MD5 vectors, URL forms, escaping) green; gated integration test (`WO_PG_TEST=…`) round-trips DDL/upsert/select/error-recovery/multi-statement against `postgres:16`; md5-auth container connects with the right password and fails cleanly with the wrong one. - -### `16b` — the mirror pipeline — ✅ shipped - -`crates/rt/src/mirror.rs`: `MirrorRec{Upsert, Delete, Txn}`, `spawn(cfg, rx, tables)` → the `wo-pg` thread (DDL bootstrap per (re)connect, batched apply, per-record error isolation, capped-backoff reconnect, loud drop accounting). Engine tap (`crates/rt/src/engine.rs`): `attach_mirror`, `mirror_send` (txn-buffered like the WAL buffer — dispatched as one `Txn` on commit, dropped on abort), `mirror_sync_all` boot push; taps sit AFTER `wal_log` accepts in `create`/`update`/`delete`/`commit_txn`. Wiring in `bin/wo.rs` behind `WO_PG`. -**Exit (met, verified live on 2 shards):** `set_price` rows appear in `psql` under the `@table` name `prices` with full JSONB; the abort case (`amount: 0` → 409) mirrors **nothing**; `docker stop` mid-writes → writes keep acking 200, reads unaffected; fresh empty container → reconnect + DDL bootstrap + new writes flow; `wo` restart → WAL replay + bulk sync **heals the gap** (the row written during the outage appeared). 69 unit tests green; blog/ecommerce/hello unchanged without `WO_PG`. - -### `16c` — typed schema projection - -Catalog scalar fields become real columns (`Id`/`Int`/`Money`/`ref` → `BIGINT`, `Text`/`Timestamp`/unions → `TEXT`, `Bool` → `BOOLEAN`; arrays/structs stay in a residual `row JSONB`); `@table(index: [product, at])` → `CREATE INDEX IF NOT EXISTS`; schema evolution via `ADD COLUMN IF NOT EXISTS`. -**Exit:** `SELECT avg((row->>'amount')::bigint)` becomes `SELECT avg(amount) FROM prices WHERE product = 1` in psql, using the mirrored index. - -### `16d` — failure & lossless resync - -Replace drop-and-log with dirty-flag repair: channel overflow or reconnect marks shards dirty; workers re-enqueue their tables at tick (the 09b mail-eventfd mechanism), so convergence no longer waits for a process restart. Mirror lag + queue depth + drop counters surface on `GET /`. -**Exit:** kill Postgres under sustained write load, restart it → row counts converge with zero client errors and no `wo` restart. - -### `16e` — restore from Postgres - -Boot source of last resort when the WAL is gone: `WO_PG_RESTORE=1` makes each shard `SELECT id, row FROM …` its own partition (`(id-1) % n = shard`) before arming accept, seeding RAM and re-logging a fresh WAL. -**Exit:** `rm -rf wo-data` → boot with restore → `current_price` answers from the restored state; id high-water marks keep the interleave. - -### `16f` — SCRAM-SHA-256 auth - -Hand-rolled SHA-256 + HMAC + PBKDF2 (RFC 7677 exchange) so stock `postgres:16` works without `pg_hba` changes. -**Exit:** connects to an out-of-the-box scram-auth server; wrong password fails with the server's error. - -## Verification (16a/16b, reproducible) - -```bash -docker run -d --rm --name wo-pg -e POSTGRES_HOST_AUTH_METHOD=trust -e POSTGRES_DB=wo -p 54329:5432 postgres:16 -WO_PG_TEST=postgres://postgres@127.0.0.1:54329/wo cargo test --lib -- pg_ mirror_ # integration tests -just pricing-pg-demo # scripted end-to-end -``` - -| Check | Result | -| --- | --- | -| `set_price 4999/5999` → `psql: SELECT * FROM prices` | rows present, full JSONB, `@table` name honoured | -| Aborted method (`amount: 0` → 409) | nothing in Postgres — only committed txns mirror | -| Postgres stopped mid-writes | writes ack 200, reads unaffected, mirror retries with backoff | -| Fresh empty database on reconnect | DDL bootstrap recreates tables, stream resumes | -| `wo` restart against behind/empty Postgres | WAL replay + boot sync converges it (outage gap healed) | -| No `WO_PG` | zero behavioural change; 69 unit tests green | - -## Non-scope - -- **No reads from Postgres on any serving path** — doctrine; even 16e's restore happens before accept is armed. -- **No TLS** to Postgres (mirror a local/private endpoint; revisit with 16f). -- **No extended query protocol / prepared statements** — simple protocol is enough for a backup writer; revisit only if 16c profiling demands it. -- **No two-way sync / conflict resolution.** Postgres is write-only from writeonce's perspective (16e restore excepted); external writes to the mirrored tables are unsupported and will be overwritten. -- **No dependency creep.** `crates/rt` gains no crates for this plan — the wire client is part of the same hand-rolled surface as the HTTP layer. - -## Cross-references - -- [`./10-storage-foundations.md`](./10-storage-foundations.md) / [`11`](11-wal-and-recovery.md) / [`12`](12-engine-disk-cutover.md) — the native disk engine; the mirror is orthogonal (external queryable backup vs. native durability) and both sit behind the RAM read path. -- [`./13-class-model-live-pricing.md`](./13-class-model-live-pricing.md) — `@table(name:)` names the mirrored tables; 13b method txns map to Postgres txns; the pricing demo is the acceptance workload. -- [`../runtime/database/07-wo-seg-migration.md`](../runtime/database/07-wo-seg-migration.md) — the dual-write pattern precedent. -- [`./15-mcp-streamable-http.md`](./15-mcp-streamable-http.md) — the other "speak an established protocol, hand-rolled" track; 16a is to Postgres what 15a is to MCP. diff --git a/docs/plan/done/01-scafolding-crates.md b/docs/plan/done/01-scafolding-crates.md deleted file mode 100644 index 01f9269..0000000 --- a/docs/plan/done/01-scafolding-crates.md +++ /dev/null @@ -1,70 +0,0 @@ -# 01 — Scaffolding the Crate Tree - -> **Status: ✅ done** (Rust Stage 2 — shipped, maintained, not advancing) — 15 crates in the workspace tree. Board: [00-status.md](../../00-status.md) - -**Context sources:** [`../../CLAUDE.md`](../../../CLAUDE.md), [`../../README.md`](../../../README.md), [`../../crates/README.md`](../../../crates/README.md), [`../runtime/database.md`](../../runtime/database.md), [`../runtime/database/07-wo-seg-migration.md`](../../runtime/database/07-wo-seg-migration.md). - -## Goal - -Lay out the full `crates/` directory tree that the 7-phase `.wo` runtime design implies, **without moving any active code**. Every target crate the design docs name gets an empty-but-compilable home so that: - -- every doc reference to `ql`, `wal`, `policy`, etc. resolves to a real directory -- each phase's extraction work (move X out of `rt` into its target crate) is a mechanical copy into an existing skeleton rather than a net-new crate creation -- IDE workspace views, dependency graphs, and `cargo doc` show the project's intended shape on day one - -## Design decisions (locked) - -1. **Scope: all 7 phases.** 14 new empty library crates covering Phases 2 → 6 land in one pass. `rt` (already shipping Stage 2) is the 15th. -2. **`rt` stays monolithic.** Today's Stage 2 code — lexer / parser / AST / compile / engine / server — stays inside `crates/rt/` and continues to satisfy the 14 existing unit tests. Code migrates into the new crates as each phase activates, not in this pass. -3. **Contents: `Cargo.toml` + `src/lib.rs` doc-comment only.** Each `lib.rs` is one module-level `//!` doc block pointing at the phase doc, naming the responsibilities, and flagging which modules in `rt` migrate here later. No placeholder types, no stub traits. -4. **No `wo-` prefix.** New runtime crates are `ql`, `value`, `engine`, etc. — not `wo-ql`, `wo-value`. The prefix is redundant inside the project's own `wo` namespace and noisy in imports (`use ql::Parser` beats `use wo_ql::Parser`). The v1 crates in `reference/crates/` keep their `wo-` prefix — the distinct prefix makes the v1/v2 split visible at a glance. -5. **Workspace membership: root `Cargo.toml` lists every new crate as a member.** `reference/crates` stays `exclude`-d (nested workspace, separate v1 code). - -Rationale and alternatives considered: see [`../../CLAUDE.md`](../../../CLAUDE.md) "What's in `rt` today vs. what the empty crates promise" and the recorded `AskUserQuestion` answers that preceded this plan. - -## Crate map - -All names are stable — documented in [`../runtime/database/07-wo-seg-migration.md`](../../runtime/database/07-wo-seg-migration.md) (Phase 2–5) and derived from [`../runtime/database/06-lowcode-fullstack.md`](../../runtime/database/06-lowcode-fullstack.md) component tables (Phase 6). - -| Phase | Crate | One-line purpose | -| --- | --- | --- | -| 2 | `ql` | `.wo` grammar — lexer, parser, AST | -| 2 | `value` | tagged `Value` + dotted-path helpers | -| 2 | `engine` | in-memory executor (rel / doc / graph) + schema catalog | -| 2 | `txn` | transaction coordinator — MVCC, `RETURNING` alias table | -| 2 | `db` | top-level facade — `open()`, `Tx`, `Query`, `Subscribe` | -| 3 | `wal` | write-ahead log — io_uring + fsync + recovery | -| 4 | `sub` | live subscriptions — delta frames on commit | -| 4 | `http` | wire protocol — REST / GraphQL-over-WS / native codec | -| 5 | `gen` | codegen — `.wo type` → Go / TS / Rust / Python clients | -| 6 | `policy` | RBAC + row-level rules compiled into planner rewrites | -| 6 | `logic` | `on ` triggers + `fn ... in txn` interpreter | -| 6 | `service` | `service rest/graphql/native` endpoint dispatch | -| 6 | `ui` | `##ui` screens → SSR HTML + client runtime | -| 6 | `app` | `##app` route manifest + startup hooks | -| — | `rt` | **existing** — Stage-2 monolith + the `wo` binary | - -## Status - -✅ **Done.** All 15 crates exist, the root workspace members list includes them, `cargo build` and `cargo test --lib` both pass, and `cargo run --bin wo -- run docs/examples/blog` still serves the blog sample (Stage 2 behaviour unchanged). - -| Artifact | Status | -| --- | --- | -| 14 new crate skeletons (`Cargo.toml` + `src/lib.rs`) | ✅ | -| Root `Cargo.toml` lists all 15 crates as members | ✅ | -| `crates/README.md` with the phase-mapped inventory | ✅ | -| `cargo build` at root (compiles 15 crates) | ✅ | -| `cargo test --lib` at root (14 existing `rt` tests) | ✅ | -| `cd reference/crates && cargo build && cargo test` (v1 still green) | ✅ | -| `cargo run --bin wo -- run docs/examples/blog` (Stage 2 still serves) | ✅ | - -## Non-scope - -- **No code extraction.** Moving `lexer.rs` / `parser.rs` / `engine.rs` out of `rt` into `ql` / `engine` is explicitly deferred. That happens incrementally as each phase activates. -- **No `gen` binary target.** `gen` ships as a library-only crate in this pass. The `[[bin]]` lands when Phase 5 starts. -- **No test scaffolding.** The empty crates don't get unit-test stubs. When a crate gets real code, it gets real tests. -- **No cross-crate `pub use` reexports from `db`.** The facade crate documents its future surface in its doc comment but doesn't import anything yet. - -## After this plan lands - -Next planning documents in this directory should describe the first real extraction — likely `02-extract-ql.md` when Stage 3 begins and the subscription engine needs the parser from a second call site. Until then, the 14 placeholders sit unmodified alongside `rt`. diff --git a/docs/plan/done/02-event-loop-epoll.md b/docs/plan/done/02-event-loop-epoll.md deleted file mode 100644 index f0106f4..0000000 --- a/docs/plan/done/02-event-loop-epoll.md +++ /dev/null @@ -1,101 +0,0 @@ -# 02 — Event Loop on `epoll` - -> **Status: ✅ done** (Rust Stage 2 — shipped, maintained, not advancing) — `runtime/netpoll_epoll.rs`: the hand-rolled `epoll` loop that replaced the async runtime. Board: [00-status.md](../../00-status.md) - -**Context sources:** [`../01-problem.md`](../../01-problem.md), `../02-recovery.md`, [`./linux/00-linux.md`](../exploration/linux/00-linux.md), [`./done/01-scafolding-crates.md`](01-scafolding-crates.md). - -## Goal - -Land a hand-rolled single-threaded event loop inside `crates/rt/` that wraps Linux's file-descriptor primitives directly — **without** touching `tokio`, `axum`, or any other async runtime crate. This is the foundation every later phase builds on: phase 03 puts an HTTP server on top of it, phase 04 retires tokio + axum, phase 07 registers `inotify` watches on it, phase 08 drives `sendfile` through it. - -Nothing is removed in this phase. The module sits alongside the tokio-backed axum server, unused by the `wo` binary until phase 04 flips the switch. - -## Design decisions (locked) - -1. **`epoll`, not `io_uring`, on day one.** `epoll` is ubiquitous (Linux 2.6+), well-understood, and every primitive we need (eventfd, timerfd, signalfd, inotify, accepted sockets) already integrates with it via `epoll_ctl`. `io_uring` is a natural follow-on phase once the event-loop abstraction exists — [`00-linux.md`](../exploration/linux/00-linux.md) calls it out for that role. -2. **Single-threaded, edge-triggered.** Matches [02-wo-language.md § Concurrency Model](../../runtime/database/02-wo-language.md#concurrency-model). Every fd registered with `EPOLLET`; the loop reads until `EAGAIN`. No worker pool, no cross-thread state. -3. **`libc` is the only new dependency.** `libc = "0.2"` added to `crates/rt/Cargo.toml`. No `nix`, no `mio`. Direct `unsafe extern "C"` calls against the kernel surface. -4. **Module, not crate (yet).** Lives at `crates/rt/src/runtime/` so phase 03 can call into it cheaply. Extraction to the empty `crates/event/` sibling is deferred until a second caller appears outside `rt` — likely when [`sub`](../../crates/sub/) starts consuming the loop for subscription delivery. - -## Scope - -### New files inside `crates/rt/src/runtime/` - -| File | Responsibility | Port source | -| --- | --- | --- | -| `mod.rs` | Re-exports `EventLoop`, `Event`, `Interest`, `Token`, `EventFd`, `TimerFd`, `SignalFd` | [`reference/crates/wo-event/src/lib.rs`](../../../.dev/reference/crates/wo-event/src/lib.rs) (9 LOC) | -| `netpoll_epoll.rs` | `EventLoop { fd, events }` — `new()`, `register(raw_fd, interest, token)`, `wait_once(timeout) -> &[Event]`, `deregister(raw_fd)` | [`reference/crates/wo-event/src/epoll.rs`](../../../.dev/reference/crates/wo-event/src/epoll.rs) (183 LOC); [`reference/go/src/runtime/netpoll_epoll.go`](../../../.dev/reference/go/src/runtime/netpoll_epoll.go) for idiom | -| `eventfd.rs` | `EventFd { fd }` — counter semaphore for cross-fd wake-up (subscription dispatch, shutdown signal) | [`reference/crates/wo-event/src/eventfd.rs`](../../../.dev/reference/crates/wo-event/src/eventfd.rs) (66 LOC) | -| `timerfd.rs` | `TimerFd { fd }` — oneshot + periodic timers as fds for the loop | [`reference/crates/wo-event/src/timerfd.rs`](../../../.dev/reference/crates/wo-event/src/timerfd.rs) (91 LOC) | -| `signalfd.rs` | `SignalFd { fd }` — SIGINT / SIGTERM / SIGHUP delivered as fd reads for graceful shutdown without a tokio signal handler | [`reference/crates/wo-event/src/signalfd.rs`](../../../.dev/reference/crates/wo-event/src/signalfd.rs) (62 LOC) | - -Total: ~410 LOC lifted and adapted. The v1 code already compiles standalone in `reference/crates/wo-event/` and has unit tests; the port is near-verbatim plus namespace cleanups. - -### Why `runtime/` not `event/` - -Go's equivalent code lives at [`reference/go/src/runtime/netpoll_epoll.go`](../../../.dev/reference/go/src/runtime/netpoll_epoll.go) alongside siblings like `netpoll_kqueue.go` (macOS/BSD), `netpoll_io_uring.go` (if/when Go adds it), and the shared `netpoll.go` interface. The directory name "runtime" signals that this is the layer beneath user code — scheduler / netpoll / syscall shims — and the filename prefix `netpoll_` makes each implementation alternative visible at a glance. Adopting the same convention in writeonce makes porting ideas bidirectional: a reader who knows Go's layout can find the writeonce equivalent by trimming the `.go` extension and swapping it for `.rs`. When Phase 3's io_uring arrives it'll land as `netpoll_io_uring.rs` next to the epoll one; a cross-platform stub would be `netpoll.rs`. Module boundary and naming both match. See [`docs/plan/assembly/00-overview.md`](../exploration/assembly/00-overview.md) for why we stop short of mirroring Go's assembly conventions. - -### `Cargo.toml` change - -```toml -[dependencies] -anyhow = "1" -serde = { version = "1", features = ["derive"] } -serde_json = "1" -tokio = { version = "1", features = ["rt", "macros", "net", "signal", "sync", "time"] } -axum = "0.7" -tower = "0.4" -libc = "0.2" # NEW — see docs/plan/02-event-loop-epoll.md -``` - -## API shape (target — validate against v1 when porting) - -```rust -use rt::runtime::{EventLoop, EventFd, Interest, Token}; - -let mut loop_ = EventLoop::new()?; -let ev = EventFd::new()?; -loop_.register(ev.as_raw_fd(), Interest::READABLE, Token(0))?; -ev.write(1)?; // wake the loop from another flow -for event in loop_.wait_once(Some(Duration::from_millis(100)))? { - match event.token() { - Token(0) => { let n = ev.read()?; /* ... */ } - _ => unreachable!(), - } -} -``` - -## Exit criteria - -1. `cargo build` at root compiles cleanly. -2. A new unit test in `crates/rt/src/runtime/netpoll_epoll.rs`: - - create an `EventLoop`, - - register an `EventFd`, - - `write(1)` to the eventfd from the same thread, - - `wait_once(timeout)` returns an `Event` for the correct token, - - `read()` on the eventfd returns `1`. -3. A second unit test validates `TimerFd::oneshot(100ms)` fires within a `wait_once(500ms)` window. -4. All 14 existing `rt` tests still pass. `cargo run --bin wo -- run docs/examples/blog` still serves (tokio path unchanged). -5. `cd reference/crates && cargo build && cargo test` still green (nothing touched). - -## Non-scope - -- **No cutover.** The `wo` binary keeps calling `tokio::runtime::Builder::new_current_thread()`. That happens in phase 04. -- **No HTTP.** Accepting connections is phase 03's problem. This phase is pure kernel-primitive plumbing. -- **No subscription dispatch.** The `sub` crate doesn't exist yet as real code; phase 07 (inotify) is the first real loop consumer after phase 03. -- **No crate extraction.** Stays at `crates/rt/src/runtime/`. Pulling to `crates/event/` waits for a second consumer. -- **No `io_uring`.** Separate follow-on once the abstraction solidifies. - -## Verification - -```bash -cargo build -cargo test --lib runtime # new tests in crates/rt/src/runtime/ -cargo test --lib # all 14 existing + new epoll/eventfd/timerfd tests green -cargo run --bin wo -- run docs/examples/blog # axum path unchanged, still serves -cd reference/crates && cargo build && cargo test # v1 untouched -``` - -## After this phase - -Phase 03 puts a non-blocking HTTP/1.1 listener on top of the `EventLoop` and proves end-to-end I/O without tokio. The two phases together give phase 04 everything it needs to delete the tokio + axum dependencies. diff --git a/docs/plan/done/03-hand-rolled-http.md b/docs/plan/done/03-hand-rolled-http.md deleted file mode 100644 index 0265dc2..0000000 --- a/docs/plan/done/03-hand-rolled-http.md +++ /dev/null @@ -1,102 +0,0 @@ -# 03 — Hand-Rolled HTTP/1.1 - -> **Status: ✅ done** (Rust Stage 2 — shipped, maintained, not advancing) — hand-rolled HTTP/1.1, plus keep-alive and pipelining. Board: [00-status.md](../../00-status.md) - -**Context sources:** [`./02-event-loop-epoll.md`](./02-event-loop-epoll.md), [`./linux/00-linux.md`](../exploration/linux/00-linux.md), `../02-recovery.md`. - -## Goal - -A non-blocking HTTP/1.1 server module that accepts connections, parses requests, and writes responses, **driven by the phase-02 `EventLoop`** — no axum, no hyper, no tokio. Still additive: the existing axum router keeps serving `wo run` until phase 04 cuts over. - -## Design decisions (locked) - -1. **HTTP/1.1 only, keep-alive supported.** HTTP/2 and HTTP/3 are not on the roadmap for Stage 2 — they need ALPN / TLS support we don't have a plan for yet. HTTP/1.1 covers every endpoint the blog + ecommerce samples exercise. -2. **Per-connection state machine.** Each accepted socket fd is registered on the event loop with its own `Connection { state: Reading | Writing | Idle, parser, pending_response }`. Edge-triggered `EPOLLIN`/`EPOLLOUT` drive state transitions. Matches [v1 wo-http](../../../.dev/reference/crates/wo-http/src/connection.rs)'s model verbatim. -3. **Router is pattern-matched at registration.** `Router::new().route("/api/articles/:id", Method::GET, handler)` resolves to a trie at boot. Per-request dispatch is a single trie walk — no axum-style type-erased layers. -4. **Handlers are `fn(&Request, &Engine) -> Response`.** Synchronous. The single-threaded event loop means a handler blocking is a bug; each handler must be a pure transformation over engine state. -5. **Module, not crate (yet).** Lives at `crates/rt/src/http/` with the same "extract when a second consumer shows up" rule as phase 02. The eventual home is the empty [`crates/http/`](../../crates/http/) sibling — but not in this phase. Paired with [phase 02's `crates/rt/src/runtime/`](./02-event-loop-epoll.md) (Go-style naming — `netpoll_epoll.rs`, `eventfd.rs`, …) which this module depends on for the `EventLoop` + raw syscall shims. Go's `src/net/http/` and `src/runtime/` split is the layout precedent; see [`reference/go/src/net/http/`](../../../.dev/reference/go/src/net/http/). - -## Scope - -### New files inside `crates/rt/src/http/` - -| File | Responsibility | Port source | -| --- | --- | --- | -| `mod.rs` | Re-exports `Listener`, `Connection`, `Request`, `Response`, `Router`, `Method`, `Status` | [`reference/crates/wo-http/src/lib.rs`](../../../.dev/reference/crates/wo-http/src/lib.rs) (4 LOC) | -| `listener.rs` | `Listener { fd }` wrapping `socket + bind + listen + accept4(SOCK_NONBLOCK \| SOCK_CLOEXEC)`; integrates with `EventLoop` | [`reference/crates/wo-http/src/listener.rs`](../../../.dev/reference/crates/wo-http/src/listener.rs) (202 LOC) | -| `connection.rs` | Per-fd state machine: drain request bytes, parse, dispatch, drain response bytes, keep-alive or close | [`reference/crates/wo-http/src/connection.rs`](../../../.dev/reference/crates/wo-http/src/connection.rs) (202 LOC) | -| `request.rs` | Incremental HTTP/1.1 request parser: request line, headers, optional body. `Content-Length` only (no chunked request bodies in Stage 2 — they don't appear in the samples) | [`reference/crates/wo-http/src/request.rs`](../../../.dev/reference/crates/wo-http/src/request.rs) (158 LOC) | -| `response.rs` | Response builder + writer: status line, headers, body (fixed or chunked) | [`reference/crates/wo-http/src/response.rs`](../../../.dev/reference/crates/wo-http/src/response.rs) (110 LOC) | -| `route.rs` | Trie-based router: static paths + `:param` segments. `Router::route(method, path, handler) -> Router` | [`reference/crates/wo-route/src/router.rs`](../../../.dev/reference/crates/wo-route/src/router.rs) (127 LOC) + [`pattern.rs`](../../../.dev/reference/crates/wo-route/src/pattern.rs) (146 LOC) | - -Total: ~949 LOC ported. Most of it is mechanical adaptation from v1; the namespace + the `Interest` enum change from phase 02 are the only non-trivial edits. - -### `Cargo.toml` change - -None. `libc` already in from phase 02 covers the raw syscalls. - -## API shape (target) - -```rust -use rt::event::EventLoop; -use rt::http::{Listener, Router, Method, Status, Response}; - -let mut loop_ = EventLoop::new()?; -let router = Router::new() - .route(Method::GET, "/healthz", |_req, _eng| Response::ok().body("ok")) - .route(Method::GET, "/api/articles", list_articles) - .route(Method::GET, "/api/articles/:id", get_article) - .route(Method::POST,"/api/articles", create_article); - -let listener = Listener::bind("127.0.0.1:8080")?; -loop_.register(listener.as_raw_fd(), Interest::READABLE, Token::LISTENER)?; - -let mut conns: HashMap = HashMap::new(); -loop { - for ev in loop_.wait_once(None)? { - match ev.token() { - Token::LISTENER => { - while let Some(stream) = listener.accept_nonblocking()? { - let fd = stream.as_raw_fd(); - loop_.register(fd, Interest::READABLE, Token::CONN(fd))?; - conns.insert(fd, Connection::new(stream)); - } - } - Token::CONN(fd) => { - conns.get_mut(&fd).unwrap().drive(ev, &router, &engine)?; - if conns[&fd].is_closed() { conns.remove(&fd); } - } - _ => {} - } - } -} -``` - -## Exit criteria - -1. A test binary `crates/rt/src/bin/http-smoke.rs` (`[[bin]] name = "http-smoke"` in `rt/Cargo.toml`) binds on `127.0.0.1:0` (auto-assigned port), registers routes for `/healthz`, `/echo/:name`, `/counter`, and services them via the phase-02 event loop. -2. An integration test (also in `crates/rt/tests/http_smoke.rs` or similar) spawns the binary, sends three `curl` equivalents using `std::net::TcpStream`, validates status codes and bodies. -3. All 14 existing `rt` tests still pass. -4. `wo run docs/examples/blog` unchanged — axum path still drives the real CLI. -5. `cargo build` at root; `cd reference/crates && cargo build` still green. - -## Non-scope - -- **No TLS.** Deferred. When it lands, it's a wrapper around `Connection` that `read`/`write`s through `rustls` or (ideally) kTLS. Not this phase. -- **No HTTP/2.** See design decision 1. -- **No chunked request bodies.** Every sample's `POST /api/X` uses `Content-Length`. If a future sample needs chunked, it's a small extension to `request.rs`. -- **No middleware.** axum's `tower::Layer` idiom has no direct analog. Cross-cutting concerns (logging, auth) live in the handler or in a wrapper fn — phase 04 re-integrates with the existing axum state handling when the cutover happens. -- **Not wired into the `wo` binary yet.** That's phase 04. - -## Verification - -```bash -cargo build # root workspace compiles -cargo test --bin http-smoke # the bundled test binary -cargo test --lib # 14 existing rt tests still green -cargo run --bin wo -- run docs/examples/blog # axum path unchanged -``` - -## After this phase - -Phase 04 takes the same in-memory `Engine` that the axum router serves and points the phase-03 router at it instead. Removing `tokio`, `axum`, `tower` is a consequence; the behaviour visible to `reference/rest/blog.rest` does not change. diff --git a/docs/plan/done/04-cutover-remove-tokio-axum.md b/docs/plan/done/04-cutover-remove-tokio-axum.md deleted file mode 100644 index d1cd02c..0000000 --- a/docs/plan/done/04-cutover-remove-tokio-axum.md +++ /dev/null @@ -1,121 +0,0 @@ -# 04 — Cutover: Remove tokio, axum, tower - -> **Status: ✅ done** (Rust Stage 2 — shipped, maintained, not advancing) — tokio, axum and tower cut over and removed; dependencies now anyhow, serde, serde_json, libc. Board: [00-status.md](../../00-status.md) - -**Context sources:** [`./02-event-loop-epoll.md`](./02-event-loop-epoll.md), [`./03-hand-rolled-http.md`](./03-hand-rolled-http.md), [`../01-problem.md`](../../01-problem.md). - -## Goal - -Flip the `wo` binary off the tokio + axum stack and onto the phase-02 event loop + phase-03 HTTP server. Delete three dependencies from `crates/rt/Cargo.toml`. REST behaviour visible to [`reference/rest/blog.rest`](../../../.dev/reference/rest/blog.rest) does not change — same status codes, same response bodies, same endpoint paths. - -This is the first phase where the dependency count goes *down*. Phases 02 and 03 were additive; this one is the switch. - -## Design decisions (locked) - -1. **Atomic swap, single commit.** Don't run tokio and the new loop in parallel in production. Flip the binary's `main()` in one change. Phase 03 already gave us confidence the new stack works end-to-end via `http-smoke`. -2. **Preserve the `Engine` trait surface.** `Arc>` stays exactly as `crates/rt/src/engine.rs` has it today. The routing layer in `crates/rt/src/server.rs` — the function that maps `service rest` blocks to axum `MethodRouter` — gets rewritten to emit phase-03 `Router::route(...)` calls instead. Same data flow, different transport. -3. **No tokio — no async.** Handlers become synchronous `fn(&Request, &Engine) -> Response`. The single-threaded event loop [already assumes this](../../runtime/database/02-wo-language.md#concurrency-model); removing `async fn` plumbing simplifies the code. `tokio::sync::Mutex` becomes `std::sync::Mutex` (fine in a single-threaded loop since lock contention is impossible). -4. **`signalfd` replaces `tokio::signal::ctrl_c()`.** Registered as another fd on the loop; reading a SIGINT cleanly exits the loop and closes outstanding connections. -5. **`WO_LISTEN` env var semantics unchanged.** The `127.0.0.1:8080` default + the `WO_LISTEN=...` override stays exactly as today. Operators don't notice the change. - -## Scope - -### Files rewritten inside `crates/rt/` - -| File | Change | Notes | -| --- | --- | --- | -| `src/bin/wo.rs` | Replace `tokio::runtime::Builder::new_current_thread` + `axum::serve` with `EventLoop` + `Listener` + `Router` wiring | The `run()` / `serve()` fns fuse into a single synchronous `run()` that drives the loop | -| `src/server.rs` | Replace `axum::Router` construction + axum handlers (`async fn list_h(State(st): State) -> impl IntoResponse`) with phase-03 `Router::route(...)` + sync handlers | Handler bodies are otherwise untouched: `engine.lock().list(&ty).map(Json)` logic flows through | -| `src/engine.rs` | `tokio::sync::Mutex` → `std::sync::Mutex`; `.lock().await` → `.lock().unwrap()` | Only the wrapper changes; row logic intact | - -### `Cargo.toml` delta - -```diff - [dependencies] - anyhow = "1" - serde = { version = "1", features = ["derive"] } - serde_json = "1" --tokio = { version = "1", features = ["rt", "macros", "net", "signal", "sync", "time"] } --axum = "0.7" --tower = "0.4" - libc = "0.2" -``` - -Three deps gone. Four remaining: `anyhow`, `serde`, `serde_json`, `libc`. - -### Files deleted - -- None. The phase-02 `runtime/` module and phase-03 `http/` module stay in place and now become the primary code path. - -## Handler signature change - -**Before (axum + tokio):** - -```rust -async fn list_h(State(st): State) -> impl IntoResponse { - let eng = st.engine.lock().await; - match eng.list(&st.ty) { - Ok(rows) => (StatusCode::OK, Json(json!(rows))).into_response(), - Err(e) => (StatusCode::INTERNAL_SERVER_ERROR, e.to_string()).into_response(), - } -} -``` - -**After (sync, event loop):** - -```rust -fn list_h(_req: &Request, st: &TypeState) -> Response { - let eng = st.engine.lock().unwrap(); - match eng.list(&st.ty) { - Ok(rows) => Response::ok().json(&serde_json::json!(rows)), - Err(e) => Response::status(Status::INTERNAL_SERVER_ERROR).body(e.to_string()), - } -} -``` - -Twelve handlers total — one pair per `{list, get, create, update, delete}` × four types. Mechanical rewrite. - -## Exit criteria - -1. **`cargo build`** at root — compiles with four deps (not seven). -2. **`cargo test --lib`** — all 14 existing `rt` unit tests still pass. A new test in `src/server.rs` exercises the router build from a compiled catalog (no HTTP, just static registration). -3. **End-to-end REST smoke — the 20-assertion battery from [`reference/rest/blog.rest`](../../../.dev/reference/rest/blog.rest)** must pass byte-identical to Stage 2 today. Script: - ```bash - WO_LISTEN=127.0.0.1:8765 cargo run --bin wo -- run docs/examples/blog & - # ... curl each block, check expected status - ``` -4. **Graceful shutdown.** SIGINT on the process exits cleanly (no panic, no orphan fds). Validate with `strace -f -e signalfd4,close` on shutdown. -5. **`cd reference/crates && cargo build && cargo test`** still green. -6. **Dep audit.** `cargo tree -p rt --depth 1` shows `libc` as the only non-transitive external dep beyond `anyhow`, `serde`, `serde_json`. - -## Non-scope - -- **No JSON replacement.** `serde` + `serde_json` are still imported and used. Phase 05 removes them. -- **No inotify / sendfile.** Stage 3 capabilities. Phases 07 and 08. -- **No `io_uring`.** The `EventLoop` keeps using `epoll` here; swapping is a later phase. -- **No crate extraction.** `runtime/` and `http/` stay inside `crates/rt/src/`. The empty `crates/http/` sibling crate waits for a second consumer; `runtime/` doesn't have a sibling slot (it's the binary's private kernel-primitive layer, analogous to Go's `src/runtime/` staying internal to the toolchain). - -## Risk - -The `.rest` files are the safety net — 20 assertions that every Stage 2 endpoint returns the expected status. If one breaks after the cutover, the fix is almost always in the handler rewrite (sync semantics + the new `Response::json(...)` helper). No transport-layer regression should survive phase 03's `http-smoke` passing. - -## Verification - -```bash -cargo build # 4 deps, no tokio/axum/tower -cargo test --lib # 14 + any new server.rs tests green - -# end-to-end -WO_LISTEN=127.0.0.1:8765 cargo run --bin wo -- run docs/examples/blog & -PID=$! -sleep 2 -# every block in reference/rest/blog.rest, via curl, checking %{http_code} -# (copy-paste the 20-assertion script from the Stage 2 turn that verified blog.rest) -kill $PID - -cd reference/crates && cargo build && cargo test # v1 untouched -``` - -## After this phase - -Phase 05 removes `serde` + `serde_json` by writing a minimal JSON parser + emitter against the new `http::Response::json()` surface. At the end of phase 06, `crates/rt/Cargo.toml` is down to `libc` alone — the stated end goal. diff --git a/docs/runtime/async.md b/docs/runtime/async.md deleted file mode 100644 index 045926d..0000000 --- a/docs/runtime/async.md +++ /dev/null @@ -1,338 +0,0 @@ -# Creating an Async Runtime Environment - -## Goal - -Build a minimal async runtime from scratch using Linux kernel primitives. The runtime is not specific to writeonce — it is a general-purpose event loop that any project can use to multiplex I/O without threads, without tokio, and without any external async framework. - -The writeonce project uses this runtime (the `wo-event` and `wo-rt` crates), but the concepts apply to any server, daemon, or event-driven application on Linux. - -## What Is a Runtime? - -A runtime is the loop that decides **what code runs next**. In a synchronous program, the OS scheduler picks the next thread. In an async runtime, a single thread asks the kernel: "which of my file descriptors are ready?" — and runs the corresponding handler. - -``` -loop { - ready_fds = ask_kernel_which_fds_are_ready() - for fd in ready_fds { - run_handler(fd) - } -} -``` - -That's the entire concept. Everything else — epoll, tokens, interest flags — is implementation detail around this loop. - -## Explaining It with C - -Before Rust, before abstractions, here is a minimal async runtime in C that watches two file descriptors on one thread. - -### Step 1: Create an epoll instance - -```c -#include -#include -#include - -int main() { - // Create the event loop - int epoll_fd = epoll_create1(0); - - // This single fd is the "runtime" — all other fds register on it - printf("epoll fd: %d\n", epoll_fd); -} -``` - -`epoll_create1` returns a file descriptor. This fd *is* the runtime. Every other fd in the system registers itself on this one fd, and the kernel tracks readiness for all of them. - -### Step 2: Register file descriptors - -```c -#include -#include -#include -#include - -int main() { - int epoll_fd = epoll_create1(0); - - // Create an eventfd (like a semaphore as a file descriptor) - int event_fd = eventfd(0, EFD_NONBLOCK); - - // Create a timerfd (fires every 2 seconds) - int timer_fd = timerfd_create(CLOCK_MONOTONIC, TFD_NONBLOCK); - struct itimerspec spec = { - .it_interval = { .tv_sec = 2, .tv_nsec = 0 }, - .it_value = { .tv_sec = 2, .tv_nsec = 0 } - }; - timerfd_settime(timer_fd, 0, &spec, NULL); - - // Register both on epoll - struct epoll_event ev1 = { .events = EPOLLIN, .data.fd = event_fd }; - epoll_ctl(epoll_fd, EPOLL_CTL_ADD, event_fd, &ev1); - - struct epoll_event ev2 = { .events = EPOLLIN, .data.fd = timer_fd }; - epoll_ctl(epoll_fd, EPOLL_CTL_ADD, timer_fd, &ev2); -} -``` - -Two different kinds of fd — an event signal and a timer — both registered on the same epoll instance. The kernel will wake us when either is ready. - -### Step 3: The event loop - -```c -#include -#include -#include -#include -#include -#include -#include - -int main() { - int epoll_fd = epoll_create1(0); - - int event_fd = eventfd(0, EFD_NONBLOCK); - int timer_fd = timerfd_create(CLOCK_MONOTONIC, TFD_NONBLOCK); - - struct itimerspec spec = { - .it_interval = { .tv_sec = 2, .tv_nsec = 0 }, - .it_value = { .tv_sec = 2, .tv_nsec = 0 } - }; - timerfd_settime(timer_fd, 0, &spec, NULL); - - struct epoll_event ev1 = { .events = EPOLLIN, .data.fd = event_fd }; - epoll_ctl(epoll_fd, EPOLL_CTL_ADD, event_fd, &ev1); - - struct epoll_event ev2 = { .events = EPOLLIN, .data.fd = timer_fd }; - epoll_ctl(epoll_fd, EPOLL_CTL_ADD, timer_fd, &ev2); - - printf("Runtime started. Timer fires every 2s.\n"); - printf("Write to eventfd to trigger it: echo 1 > /proc/%d/fd/%d\n", - getpid(), event_fd); - - // The event loop — this IS the runtime - struct epoll_event events[10]; - while (1) { - int n = epoll_wait(epoll_fd, events, 10, -1); // block until ready - - for (int i = 0; i < n; i++) { - int fd = events[i].data.fd; - - if (fd == timer_fd) { - uint64_t expirations; - read(timer_fd, &expirations, sizeof(expirations)); - printf("[timer] fired (%lu expirations)\n", expirations); - } - else if (fd == event_fd) { - uint64_t val; - read(event_fd, &val, sizeof(val)); - printf("[event] signaled (value: %lu)\n", val); - } - } - } - - close(epoll_fd); - close(event_fd); - close(timer_fd); - return 0; -} -``` - -Compile and run: - -```bash -gcc -o runtime runtime.c -./runtime -``` - -Output: - -``` -Runtime started. Timer fires every 2s. -[timer] fired (1 expirations) -[timer] fired (1 expirations) -[timer] fired (1 expirations) -... -``` - -One thread. Two fd types. One loop. The kernel does the scheduling. - -### Step 4: Add a TCP server to the same loop - -```c -#include -#include -#include -#include -#include -#include -#include - -int main() { - int epoll_fd = epoll_create1(0); - - // Create a non-blocking TCP listener - int listen_fd = socket(AF_INET, SOCK_STREAM | SOCK_NONBLOCK, 0); - int opt = 1; - setsockopt(listen_fd, SOL_SOCKET, SO_REUSEADDR, &opt, sizeof(opt)); - - struct sockaddr_in addr = { - .sin_family = AF_INET, - .sin_port = htons(8080), - .sin_addr.s_addr = INADDR_ANY - }; - bind(listen_fd, (struct sockaddr*)&addr, sizeof(addr)); - listen(listen_fd, 128); - - // Register listener on epoll - struct epoll_event ev = { .events = EPOLLIN, .data.fd = listen_fd }; - epoll_ctl(epoll_fd, EPOLL_CTL_ADD, listen_fd, &ev); - - printf("Listening on port 8080\n"); - - struct epoll_event events[64]; - while (1) { - int n = epoll_wait(epoll_fd, events, 64, -1); - - for (int i = 0; i < n; i++) { - int fd = events[i].data.fd; - - if (fd == listen_fd) { - // Accept new connection - int client_fd = accept4(listen_fd, NULL, NULL, SOCK_NONBLOCK); - if (client_fd >= 0) { - struct epoll_event cev = { .events = EPOLLIN, .data.fd = client_fd }; - epoll_ctl(epoll_fd, EPOLL_CTL_ADD, client_fd, &cev); - printf("[accept] client fd=%d\n", client_fd); - } - } else { - // Read from client - char buf[4096]; - int nbytes = read(fd, buf, sizeof(buf)); - if (nbytes <= 0) { - // Client disconnected - epoll_ctl(epoll_fd, EPOLL_CTL_DEL, fd, NULL); - close(fd); - printf("[close] fd=%d\n", fd); - } else { - // Echo response - const char *response = - "HTTP/1.1 200 OK\r\n" - "Content-Length: 13\r\n" - "\r\n" - "Hello, world!"; - write(fd, response, strlen(response)); - epoll_ctl(epoll_fd, EPOLL_CTL_DEL, fd, NULL); - close(fd); - } - } - } - } -} -``` - -This is a complete HTTP server — no threads, no framework, no library. One `epoll_wait` drives accept, read, write, and close for every connection. - -## From C to Rust: The writeonce Runtime - -The writeonce `wo-event` crate wraps these same syscalls in safe Rust: - -| C syscall | Rust wrapper | Crate | -|-----------|-------------|-------| -| `epoll_create1` | `EventLoop::new()` | wo-event | -| `epoll_ctl(ADD)` | `EventLoop::register(fd, interest, token)` | wo-event | -| `epoll_ctl(MOD)` | `EventLoop::modify(fd, interest, token)` | wo-event | -| `epoll_ctl(DEL)` | `EventLoop::deregister(fd)` | wo-event | -| `epoll_wait` | `EventLoop::poll(timeout)` | wo-event | -| `eventfd` | `EventFd::new()` | wo-event | -| `timerfd_create` | `TimerFd::new()` | wo-event | -| `signalfd` | `SignalFd::new()` | wo-event | - -The key difference from the C examples: instead of matching on raw fd numbers, the Rust runtime assigns a **token** (u64) to each fd. The event loop returns tokens, and the runtime dispatches on them: - -```rust -let events = event_loop.poll(Some(Duration::from_millis(500)))?; - -for event in events { - match event.token { - TOKEN_WATCHER => handle_file_change(), - TOKEN_SIGNAL => handle_shutdown(), - TOKEN_TIMER => handle_periodic_task(), - TOKEN_LISTENER => handle_new_connection(), - token if token >= 10000 => handle_http(token), - _ => {} - } -} -``` - -## Why Not tokio? - -tokio is a production-grade async runtime. It handles everything — epoll, thread pools, work stealing, timers, I/O drivers. So why build a custom one? - -| Concern | tokio | Custom runtime | -|---------|-------|---------------| -| Binary size | Adds ~2-3 MB | Zero — just libc syscalls | -| Dependencies | 50+ transitive crates | 1 crate (libc) | -| Complexity | Work-stealing scheduler, multi-threaded executor | Single-threaded loop, ~200 lines | -| Control | Opaque — runtime internals hidden behind `.await` | Every fd, every syscall, every state transition is explicit | -| Learning | Abstracts the kernel away | Forces understanding of what the kernel actually does | - -For a content platform serving markdown files, the workload is: accept connection, read request, query in-memory index, render template, write response. This is microseconds of work per request. A single-threaded event loop handles thousands of concurrent connections without the complexity of a multi-threaded executor. - -## Reusing the Runtime in Other Projects - -The `wo-event` crate has no dependency on writeonce. It provides: - -- `EventLoop` — epoll wrapper with register/deregister/poll -- `EventFd` — lightweight signaling -- `TimerFd` — periodic timers as fds -- `SignalFd` — SIGINT/SIGTERM as fd events -- `Event` — readable/writable/hangup status -- `Token` — u64 identifier for dispatch - -Any project that needs non-blocking I/O on Linux can use it: - -```rust -use wo_event::{EventLoop, EventFd, Interest}; -use std::time::Duration; - -fn main() { - let eloop = EventLoop::new().unwrap(); - let efd = EventFd::new().unwrap(); - - eloop.register(efd.fd(), Interest::Readable, 1).unwrap(); - - // Signal from another thread - std::thread::spawn(move || { - std::thread::sleep(Duration::from_secs(1)); - efd.write(42).unwrap(); - }); - - // Wait for the signal - let events = eloop.poll(Some(Duration::from_secs(5))).unwrap(); - assert_eq!(events[0].token, 1); - println!("Event received!"); -} -``` - -## The Mental Model - -``` - ┌─────────────────────────────────┐ - │ your code │ - │ (handlers, business logic) │ - └──────────────┬──────────────────┘ - │ dispatches on token - ┌──────────────┴──────────────────┐ - │ event loop │ - │ epoll_wait → Vec │ - └──────────────┬──────────────────┘ - │ registered fds - ┌──────────────┴──────────────────┐ - │ Linux kernel │ - │ tracks readiness for all fds │ - │ inotify, sockets, timers, │ - │ signals, eventfds — all fds │ - └─────────────────────────────────┘ -``` - -The kernel is the scheduler. The event loop is the dispatcher. Your code is the handler. Everything in the system — files, sockets, timers, signals — is a file descriptor. One loop to rule them all. diff --git a/docs/runtime/database.md b/docs/runtime/database.md deleted file mode 100644 index 95398a8..0000000 --- a/docs/runtime/database.md +++ /dev/null @@ -1,60 +0,0 @@ -# Document & Graph Database — Design Series - -A seven-phase design series that starts with "should writeonce use a document or graph database?" and arrives at a full-stack declarative application platform — then plans the migration from writeonce's current flat-file store to that platform. - -> **Start here if you're new:** wo-language.md — the user-facing overview of what writeonce *is* (a programming language with DB + HTTP in its runtime, Go-style toolchain). This series is the engineering plan that gets you there. - -Each phase is self-contained and shippable on its own. Every phase after Phase 1 builds on the previous ones. - -## Phases - -| Phase | Document | Summary | -| --- | --- | --- | -| **1** | [Database Evaluation](./database/01-evaluation.md) | Evaluate CouchDB, Postgres, Neo4j, petgraph against writeonce's needs. Decision: no external DB; add petgraph-backed `mappings.idx`. | -| **2** | [The `.wo` Language & ACID Engine](./database/02-wo-language.md) | Design a two-layer `.wo` language (unified `type` schema layer + hybrid SQL/Cypher query layer with fixed glue) for an e-commerce platform with ACID transactions across relational, document, and graph storage. | -| **3** | [In-Memory Engine](./database/03-inmemory-engine.md) | RAM-primary, SSD-durable storage engine using `io_uring`, `mlockall`, group commit. 64 GB Linux machine. | -| **4** | [Client API: Wire Protocol & Subscriptions](./database/04-client-api.md) | Native binary protocol + GraphQL over WebSocket. Subscription engine inside the transaction coordinator — no polling anywhere. | -| **5** | Go Client SDK | Typed Go client with subscription-first design. `gen` codegen from `.wo` schema. Subscribe to a live query in 5 lines. | -| **6** | [Low-Code Full-Stack](./database/06-lowcode-fullstack.md) | Expand `.wo` into an application language (like SAP CDS): `##ui`, `##logic`, `##policy`, `##service` blocks compiled into a single binary. | -| **7** | [Replacing `wo-seg`](./database/07-wo-seg-migration.md) | Phased coexistence plan: abstract the article store behind a trait, stand up the `.wo` engine as a second impl, dual-run, cut over, decommission `wo-seg`. | - -## Build Order - -``` -Phase 1: Evaluation ← writeonce today (blog, flat files) - │ - ▼ -Phase 2: .wo Language ← query language + ACID engine design - │ - ▼ -Phase 3: In-Memory Engine ← RAM-primary storage, io_uring, WAL - │ - ▼ -Phase 4: Client API ← wire protocol, subscriptions, GraphQL - │ - ▼ -Phase 5: Go SDK ← typed client, codegen, subscription channels - │ - ▼ -Phase 6: Low-Code Full-Stack ← .wo as application DSL, UI generation, CLI - │ - ▼ -Phase 7: Replace wo-seg ← trait abstraction, dual-run, cutover, decommission -``` - -## Cumulative Scope - -| After Phase | What exists | Rough cumulative effort | -| --- | --- | --- | -| 1 | petgraph `mappings.idx` in writeonce | ~200 lines | -| 2 | `.wo` parser + planner + ACID engine prototype | 5–10 months | -| 3 | In-memory engine with io_uring durability | 8–16 months | -| 4 | Wire protocol + subscription engine | 14–28 months | -| 5 | Go SDK with typed subscriptions | 16–32 months | -| 6 | Full-stack low-code platform | 28–56 months | -| 7 | `wo-seg` replaced by the `.wo` engine in writeonce | +2–4 months on top of Phase 2 arrival | - -## Related Documents - -- [surreal-case-study.md](./surreal-case-study.md) — SurrealDB runtime analysis; why writeonce does not use a multi-model DB for live queries -- [async.md](./async.md) — custom async runtime using Linux kernel primitives diff --git a/docs/runtime/database/01-evaluation.md b/docs/runtime/database/01-evaluation.md deleted file mode 100644 index 2e03fe1..0000000 --- a/docs/runtime/database/01-evaluation.md +++ /dev/null @@ -1,270 +0,0 @@ -# Phase 1 — Database Evaluation - -> Why build a custom database instead of adopting CouchDB, Postgres, Neo4j, or an in-memory graph library? - -**Next**: [Phase 2 — The `.wo` Language & ACID Engine](./02-wo-language.md) | **Index**: [database.md](../database.md) - ---- - -Reference repositories: - -- [github.com/apache/couchdb](https://github.com/apache/couchdb) — document DB with Mango query language -- [github.com/postgres/postgres](https://github.com/postgres/postgres) — relational DB with JSONB -- [github.com/neo4j/neo4j](https://github.com/neo4j/neo4j) — native graph DB, Cypher query language -- [github.com/petgraph/petgraph](https://github.com/petgraph/petgraph) — Rust in-memory graph library (NetworkX analogue) - -## The Question - -writeonce articles have two shapes at once: they are **documents** (per-article JSON metadata + markdown body, per 06-markdown-render.md) and they form a **graph** (the `mappings` field — `related`, `prerequisite`, `series`, `supersedes`, `references` — per ai-agents-content-management.md). - -Should writeonce adopt an off-the-shelf document or graph database to back these two shapes, or keep the flat-file `.seg` + `.idx` storage already implemented in 05-datalayer.md? - -**Short answer: no external DB.** The dataset is small (hundreds of articles, not millions of rows), single-writer (author commits), and read-heavy. A full rebuild on change is cheap. The `mappings` graph fits entirely in RAM. External databases would add a process, a protocol, a driver, and a failure mode — none of which writeonce needs. - -This doc walks through each option, what it buys, and why writeonce does not adopt it. - -## The Data Shape - -Every article is a pair of files in `content/{sys_title}/`: - -```json -{ - "sys_title": "linux-misc", - "title": "Linux Miscellaneous", - "published": true, - "author": "Shoney Arickathil", - "tags": ["Linux"], - "published_on": 1740950884, - "mappings": { - "related": ["auto-scale-gitlab-runner-using-aws-spot-instance"], - "prerequisite": ["linux-misc"], - "series": { "name": "gitlab-runner", "order": 2 } - } -} -``` - -Plus `{sys_title}.md` with the full article body. - -The access patterns, per 05-datalayer.md: - -| Pattern | Frequency | Current Implementation | -| --- | --- | --- | -| `get_by_title(sys_title)` | Every `/blog/:sys_title` hit | `title.idx` hash — O(1) | -| `list_published(skip, limit)` | Homepage, pagination | `date.idx` sorted array — O(log n) | -| `list_by_tag(tag)` | Tag pages | `tags.idx` inverted index — O(1) + scan | -| `list_by_date_range` | Archive views | `date.idx` binary search | -| Graph traversal (`mappings`) | Agent workflows, "related" widget | **Not yet implemented** | - -The last row is the gap. Everything else already works on `.seg` + `.idx`. - -## Option 1: CouchDB + Mango - -CouchDB is an HTTP-native document store. Each article JSON would be a document; Mango queries are JSON selectors expressive enough for most writeonce reads: - -```json -{ "selector": { "published": true, "tags": { "$in": ["rust"] } }, "sort": [{ "published_on": "desc" }] } -``` - -What it buys: - -- Built-in revisions (MVCC) — each edit gets a `_rev` -- Multi-master replication, which matters if content is authored from multiple machines -- A _changes feed, conceptually similar to writeonce's subscription model - -What it costs: - -- Separate Erlang process, HTTP driver, JSON over the wire on every read -- No native graph traversal — `mappings` would require client-side joins or view functions -- Duplicates what inotify + `.seg` already do (change feed, indexing) - -### Comparison - -| Aspect | CouchDB | writeonce | -| --- | --- | --- | -| Transport | HTTP per query | In-process function call | -| Storage | B-tree per database, append-only | `.seg` append-only, tombstoned records | -| Change feed | `_changes` HTTP long-poll | inotify → epoll → fd write | -| Index | View functions (JavaScript map/reduce), Mango indexes | `title.idx`, `date.idx`, `tags.idx` rebuilt on change | -| Revisions | Every write creates `_rev` | Git already does this for `content/` | - -CouchDB's replication and revision model are attractive, but git already covers revision history for `content/`, and the dataset is too small to justify a separate daemon. - -## Option 2: PostgreSQL + JSONB - -Postgres with a `jsonb` column gives you SQL over the article metadata, GIN indexes for tag queries, and a real query planner. Schema would be roughly: - -```sql -CREATE TABLE articles ( - sys_title TEXT PRIMARY KEY, - metadata JSONB NOT NULL, - body_md TEXT NOT NULL, - updated TIMESTAMPTZ DEFAULT now() -); -CREATE INDEX ON articles USING GIN ((metadata->'tags')); -CREATE INDEX ON articles (((metadata->>'published_on')::bigint)); -``` - -What it buys: - -- Mature: WAL, point-in-time recovery, replication, tooling -- `jsonb_path_query` and `@>` containment make most writeonce queries one-liners -- Recursive CTEs (`WITH RECURSIVE`) can traverse `mappings` — adequate for shallow graphs - -What it costs: - -- A Postgres process, a driver (tokio-postgres or raw libpq), connection pooling — writeonce's 05-datalayer.md explicitly removed all of this -- JSONB query planning is excellent but still pays per-query cost that an in-process hash does not -- Recursive CTEs on deep mapping chains are slower than a RAM graph walk - -### Comparison - -| Aspect | Postgres JSONB | writeonce | -| --- | --- | --- | -| Process model | External daemon, TCP or unix socket | Single binary, single process | -| Query language | SQL + JSONB operators | Rust method calls on `Store` | -| Graph traversal | `WITH RECURSIVE` CTE | (Future) in-memory adjacency list | -| Backup | `pg_dump`, WAL archive | `content/` directory + git | -| Failure modes | Connection loss, pool exhaustion, vacuum stalls | File not found | - -Postgres is the default reflex for "I have structured data." But writeonce's structured data is ~500 rows that change when the author saves a file. The mismatch is two orders of magnitude on the dataset size and one process on the deployment surface. - -## Option 3: Neo4j — Native Graph DB - -Neo4j models articles as nodes and `mappings` as typed edges. Cypher makes the traversals natural: - -```cypher -MATCH (a:Article {sys_title: 'linux-misc'})-[:PREREQUISITE*1..3]->(p:Article) -RETURN p.sys_title -``` - -What it buys: - -- Native graph storage — constant-time edge traversal regardless of dataset size -- Cypher is the right query language for the `mappings` problem -- Useful when the graph is the primary shape - -What it costs: - -- JVM process, Bolt protocol, driver — heaviest option on this list -- Document storage is secondary (properties on nodes), so article body and metadata are awkwardly split -- Dataset is tiny — the graph fits in a few KB of RAM; Neo4j's disk-backed adjacency is overkill - -### Comparison - -| Aspect | Neo4j | writeonce | -| --- | --- | --- | -| Storage | Native adjacency on disk | (Future) `HashMap>` in memory | -| Traversal | Cypher over Bolt | Rust iteration over adjacency | -| Transactions | ACID with MVCC | Full rebuild on change | -| Deployment | JVM daemon + driver | Statically linked binary | -| Fit for writeonce dataset | Overprovisioned by 3+ orders of magnitude | Right-sized | - -## Option 4: In-Memory Graph — NetworkX / petgraph - -NetworkX (Python) and petgraph (Rust) are _libraries_, not databases. You load the graph into process memory and traverse it directly. This is the model gestured at in ai-agents-content-management.md line 181 — "traversable knowledge graphs available on RAM." - -For writeonce, petgraph is the right shape: - -```rust -use petgraph::graph::DiGraph; - -let mut g: DiGraph = DiGraph::new(); -// nodes: one per sys_title -// edges: one per mapping (related, prerequisite, series, ...) -``` - -What it buys: - -- Zero extra processes — compiles into the `wo-store` crate -- Constant-time neighbor lookup, standard BFS / Dijkstra / SCC algorithms included -- Rebuilt cheaply on any `content/` change by walking the `.seg` and following `mappings` - -What it costs: - -- No persistence layer — but the graph is derived from JSON, same as `title.idx`, so it rebuilds on cold start for free -- No query language — but the traversals agents need (`all prerequisites of X`, `next article in series Y`) are short Rust functions - -### Comparison - -| Aspect | petgraph (in-memory) | Neo4j | -| --- | --- | --- | -| Location | Same process as `wo-store` | External JVM | -| Build time | Single pass over `.seg` | Bulk import via CSV / Cypher | -| Cost per traversal | Pointer chase in RAM | Network round trip + disk I/O | -| Query language | Rust | Cypher | -| Scales to | Millions of nodes in RAM (plenty of headroom) | Billions on disk | - -For writeonce's hundreds of articles, petgraph is the correct answer. It fits the existing architecture — derived from `content/`, rebuildable, no external process — the same shape as `title.idx` already has. - -## Decision Matrix - -| Option | Dataset Fit | Process Count | Graph Support | Matches writeonce Philosophy | -| --- | --- | --- | --- | --- | -| CouchDB Mango | Overprovisioned | +1 (Erlang) | Manual joins | No — external daemon | -| Postgres JSONB | Overprovisioned | +1 (Postgres) | Recursive CTE | No — external daemon | -| Neo4j | Overprovisioned | +1 (JVM) | Native, excellent | No — external daemon | -| petgraph in-memory | Right-sized | 0 | Native, in-process | **Yes** | -| Current `.seg` + `.idx` | Right-sized | 0 | None yet | Already here | - -## Proposed Addition: `mappings.idx` Backed by petgraph - -Per 05-datalayer.md, indexes live alongside `.seg`. Add a fourth index: - -``` -data/ - articles.seg - index/ - title.idx # existing — O(1) sys_title lookup - date.idx # existing — sorted by published_on - tags.idx # existing — inverted index by tag - mappings.idx # NEW — serialized petgraph adjacency -``` - -New `wo-graph` crate (or extension to `wo-index`): - -```rust -pub struct MappingGraph { - graph: DiGraph, - by_title: HashMap, -} - -impl MappingGraph { - pub fn neighbors(&self, sys_title: &str, kind: MappingKind) -> Vec<&str> { ... } - pub fn prerequisites_transitive(&self, sys_title: &str) -> Vec<&str> { ... } - pub fn series(&self, name: &str) -> Vec<&str> { ... } // ordered by .order -} -``` - -Rebuild on every `Store::rebuild()`. Drop and recompute on any `ContentChange` — the graph is small enough that incremental updates are not worth the bug surface. - -## Key Takeaways - -1. **A document+graph database is the right model — but not the right dependency.** writeonce's articles really are documents with graph edges. That does not mean importing CouchDB, Postgres, or Neo4j. It means writing ~200 lines that give you the document and graph operations the product actually uses. - -2. **Dataset size dictates architecture.** SurrealDB, Postgres, and Neo4j all assume millions of rows and concurrent writers. writeonce has hundreds of articles and one author. The gap is where external databases become overhead, not infrastructure. - -3. **Git already solves the problems CouchDB's revisions solve.** Content lives in `content/` under version control. `_rev`, replication, and change history are the author's git history. - -4. **Query languages are a cost, not a benefit, at this scale.** Mango selectors, JSONB operators, and Cypher exist because production queries are written by humans against large evolving datasets. writeonce's queries are fixed (`get_by_title`, `list_by_tag`, `list_published`) and written once in Rust. - -5. **The graph belongs in RAM.** Article `mappings` form a small, mostly static DAG. petgraph holds it with zero protocol overhead and rebuilds from `content/` on cold start — same lifecycle as `title.idx`. - -## Reference - -If the graph side of this ever grows beyond what petgraph comfortably handles, the reference points are: - -```bash -git submodule add https://github.com/neo4j/neo4j.git references/neo4j -git submodule add https://github.com/apache/couchdb.git references/couchdb -git submodule add https://github.com/petgraph/petgraph.git references/petgraph -``` - -Key files to study: - -- `petgraph/src/graph_impl/` — adjacency list implementation, the smallest viable graph backend -- `couchdb/src/mango/` — Mango query compilation, if a selector-style query API ever becomes useful -- `neo4j/community/cypher/` — Cypher planner, for how a real graph query language is structured - -See also: - -- [surreal-case-study.md](../surreal-case-study.md) — why writeonce does not use a multi-model DB for live queries diff --git a/docs/runtime/database/02-wo-language.md b/docs/runtime/database/02-wo-language.md deleted file mode 100644 index e3b7ad4..0000000 --- a/docs/runtime/database/02-wo-language.md +++ /dev/null @@ -1,508 +0,0 @@ -# Phase 2 — The `.wo` Language & ACID Engine - -> A two-layer multi-paradigm language for an e-commerce platform with ACID transactions across relational, document, and graph storage. - -> **Query layer superseded (2026-08-15):** the SQL + Cypher query layer below -> is design history — on the C stack, programs query their tables through the -> language-integrated surface specified in -> [`2026-08-15-table-relations-query-design.md`](../../superpowers/specs/2026-08-15-table-relations-query-design.md) -> (fork 1's decision record). This document remains the reference for the -> schema layer's vocabulary and for the `wo-db` prototype's engine semantics; -> SQL text is at most a future export format, never a program surface. - -**Previous**: [Phase 1 — Database Evaluation](./01-evaluation.md) | **Next**: [Phase 3 — In-Memory Engine](./03-inmemory-engine.md) | **Index**: [database.md](../database.md) - ---- - -## Creating Your Own Runtime Language for Database - -- language parser -- file saved as `database.wo` - -Early sketch of the schema — three paradigms in one file: - -```wo -##sql -#users - id bigint - name varchar - -#article - title varchar - meta article-meta - -##doc -#article-meta - -##graph -#context-map - ##sql-article -``` - -queries (relational) - -```wo -SELECT name FROM users - -SELECT meta.sys_title FROM article -``` - -queries (graph) - -```wo -MATCH (a:article), (b:article) -WHERE a.meta.sys_title = "Alice" AND b.meta.sys_title = "Bob" -CREATE (a)-[:DEFINES]->(b) -``` - -That three-paradigm sketch is **the execution substrate** — the thing the engine actually parses and runs. But it is not the right *authoring* surface for a full-stack platform. The language is designed as two layers, described next. - -## The Two-Layer Design - -``` -┌────────────────────────────────────────────────────────┐ -│ Schema Layer — unified type DSL (SAP-CDS-style) │ ← source of truth -│ type User { ... } / type Purchase link … / policy … │ (Phases 5 + 6) -└──────────────────────┬─────────────────────────────────┘ - │ compiled to ↓ -┌──────────────────────▼─────────────────────────────────┐ -│ Query Layer — hybrid SQL + Cypher with fixed glue │ ← execution substrate -│ SELECT / UPDATE / MATCH / CREATE / BEGIN … COMMIT │ (Phase 2 prototype) -└────────────────────────────────────────────────────────┘ -``` - -Both layers are `.wo` files — same extension, same tooling, same parser front-end. They differ in role: - -- The **schema layer** names the data model once. One `type` declaration per entity covers what the three paradigm blocks cover today (relational columns, embedded documents, graph edges) plus constraints, computed fields, policies, and triggers. It is the source of truth for codegen (Phase 5) and for the full-stack blocks ([Phase 6](./06-lowcode-fullstack.md)). -- The **query layer** is the operational surface. SQL and Cypher stay as-is — they are universally legible, every backend developer already reads them — but five things are tightened so the three grammars share semantics (parameters, `RETURNING`, dotted paths, transactions, `LIVE`). - -The two layers ship on different timelines. The query layer is Phase 2 (already prototyped at [`prototypes/wo-db/`](../../../prototypes/wo-db/)). The schema layer enters when Phase 5 codegen needs a single authoritative input. - -### Why Not One Layer? - -Two alternatives were considered and rejected: - -- **Unified query language only** (EdgeQL-style path algebra replacing SQL and Cypher). Loses the Phase 2 adoption property — SQL+Cypher are universally legible; a novel path language is not. Reinvents 6+ years of EdgeDB planner work. -- **Three paradigm blocks only** (the original sketch). Works for Phase 2 but hits a wall at Phase 5: the SDK codegen has to invent a schema-above-schema layer anyway, because "a relational row with an embedded doc column plus an inverse graph edge" is one Go struct, not three. Also hits a wall at Phase 6: `##ui` / `##policy` / `##logic` want to attach to *entities*, not to tables-vs-collections-vs-edges. - -The two-layer split captures the Phase 2 adoption win and the Phase 5/6 coherence win without committing to a single novel query grammar. - -## Schema Layer — Unified Type DSL - -One `type` construct describes an entity. The compiler chooses physical storage (relational page, document LSM, graph node) from the field declarations. - -```wo -type User { - id: Id - email: Email @unique - meta: { name: Text, avatar: Url? } -- inline document - friends: multi User @edge(:FOLLOWS) -- anonymous edge (tag only) - purchased: multi Product via Purchase -- link type with props - orders: backlink Order.user -- inverse scalar ref -} - -type Purchase link User -> Product { -- edge with properties - order: ref Order - qty: Int @check(> 0) - at: Timestamp = now() -} - -type Order { - id: Id - user: ref User - status: Pending | Paid | Shipped | Refunded -- tagged union - line_items: [{ product: ref Product, qty: Int, unit: Money }] - total: Money = sum(line_items.*.qty * line_items.*.unit) -- computed -} - -type Product { - id: Id - sku: SKU @unique - price: Money - meta: { title: Text, description: Markdown, images: [Url], reviews: [Review] } - inventory: { on_hand: Int @check(>= 0), reserved: Int = 0, reorder_at: Int } -} - -type Review { - user: ref User - stars: Int @check(between 1 and 5) - body: Markdown - written_at: Timestamp = now() -} -``` - -**Type-system primitives:** - -| Primitive | Example | Compiles to | -| --- | --- | --- | -| Scalar | `Int Float Text Bool Timestamp Id Url Email Markdown Money SKU Slug` | relational column | -| Optional | `Url?` | nullable column | -| Array | `[Url]` | relational array or doc array | -| Struct | `{ k: V, ... }` | embedded document column (doc engine) | -| Scalar ref | `ref User` | foreign-key column | -| Zero-prop edge | `multi User @edge(:FOLLOWS)` | graph edge, no properties | -| Link-with-props | `multi Product via Purchase` | graph edge + linked row (`type Purchase link ...`) | -| Inverse | `backlink Order.user` | computed inverse of `Order.user: ref User` | -| Tagged union | `Pending \| Paid \| Shipped` | enum column | -| Computed | `total: Money = sum(...)` | view or materialized view | - -**Annotations:** `@unique @check(...) @default(...) @index @search @immutable`. - -**Full-stack blocks attach to types** — no separate `##ui`/`##logic`/`##policy`/`##service` paradigm markers. See [Phase 6](./06-lowcode-fullstack.md) for details; shape: - -```wo -type Article { - slug: Slug @unique - title: Text - body: Markdown - author: ref User - - policy read anyone - policy write when author == $session.user - - on update when old.status != "published" and new.status == "published" - do emit "article.published"(self) - - service rest "/articles" expose list, get, subscribe -} -``` - -### Schema-Layer DML — Brace Disambiguation - -The schema-layer DML used in `main` blocks and triggers (`insert` / `update` / `select` / `delete` over `Type{ ... }`) reuses one brace syntax for three roles. The parser decides per entry, with one token of lookahead: - -| Entry shape inside `Type{ ... }` | Meaning | Example | -| --- | --- | --- | -| `name: value` pairs | construction (insert only) | `insert Note { title: "hello" }` | -| expression containing an operator | predicate (filter) | `select Note{ title == "hello" }` | -| bare identifier / dotted path | projection (shape) | `select Note{ title }` — rows shaped `{ title }` | - -Rules: - -- **A bare identifier is always a projection; a predicate always requires an operator.** Boolean fields are the case this rule exists for: `select Note{ pinned }` projects the `pinned` column; filtering on it must be written `pinned == true`. -- **Entries mix.** `select Note{ title, pinned == true }` filters on `pinned` and projects `title`. Entry order inside the braces is irrelevant. -- **Cardinality.** A `select` with no predicate (or a non-unique one) returns a set; dotted access distributes over it — `select Note{ title }` followed by `.title` yields the set of all titles, consistent with the query layer's one-path rule below and with Cypher projections. A predicate over a `@unique` field returns at most one row, and dotted access yields a scalar. -- **Bindings are snapshots.** `let n = select ...` copies values at read time; a later `update` does not retro-change `n`. Live observation is the `LIVE` prefix's job, not `let`'s. - -Precedent: EdgeDB shapes (`select Movie { title }`). The projection form is also symmetric with construction — braces build a shape on `insert`, and select the same shape on read. - -### Class Model — state + methods, no inheritance - -`class` is the behavior-bearing sibling of `type` ([plan 13](../../plan/13-class-model-live-pricing.md)). Identical field grammar and attachable blocks (`service`, `policy`, `on `), plus `fn` methods: - -```wo -class Product { - id: Id - sku: SKU @unique - prices: multi Price -- composition, not inheritance - - fn current_price() -> Money in txn { -- implicit `self` = the receiving row - return latest(self.prices).amount; - } - - service rest "/api/products" expose list, get, create, update, delete, subscribe -} -``` - -Rules: - -- **No inheritance.** No `extends`, no override, no virtual dispatch. "Is-a" is a tagged union; "has-a" is `ref`/`multi`. This also kills the table-per-class storage-mapping problem: a class IS one table (+ doc/graph parts), exactly like a type. -- **Methods are row-scoped transactional functions.** `fn name(args) -> Ret [in txn [snapshot]]` — the same signature grammar and coordinator as a free-standing `fn` (see `fn checkout` in the ecommerce sample); `self` binds to the receiving row. Bodies are the schema-layer DML of the section above. Exposed as RPC: `POST /api/products/:id/current_price` (plan 13b). -- **Storage and REST are class-blind.** The catalog treats `class` exactly like `type`; converting between them is a no-op for stored data. `self` stays a plain identifier in the lexer (same rule as `subscribe`/`me`). -- **Status:** parsing + CRUD shipped (13a, `ast::TypeDecl::is_class`); method execution shipped (13b, `crates/rt/src/method.rs`) — bodies compile to `ast::{Stmt, Expr}` (`let`, `insert`, `select` expressions, `return`, `assert … otherwise abort`, `if/else`) and run on the row's owning shard, committing as one atomic `WalRec::Txn` frame; an abort rolls back completely (HTTP 409). `fn` inside a plain `type` is still parsed-and-discarded. - -### Type-Level Annotations — `@table` - -An optional annotation list may precede a `type`/`class` declaration. The first one is `@table` — **storage configuration, never storage declaration**: every `type`/`class` IS a table regardless (see Class Model rule 3 above and plan 13 decisions 3/5); `@table` only configures how. - -```wo -@table(name: "prices", index: [product, at], index: [sku]) -class Price { ... } -``` - -| Argument | Meaning | -| --- | --- | -| `name: "…"` | Storage/table name (default: the type name). Catalog-unique, enforced at compile. Consumed by the SQL layer and the storage phases (plans 10–12); the WAL keeps the *type name* as the stable identifier, so renaming a table is replay-safe. | -| `index: [a, b]` | One composite secondary index per entry (repeatable). Columns must be stored scalar columns — scalars, unions, and `ref` FKs; `multi`/`backlink` have no column and are compile errors. | - -Rules: - -- **Bare `@table` is a legal no-op.** Annotating changes nothing by itself. -- **Indexes are engine-maintained on every mutation path** — CRUD, method-transaction undo, and WAL replay — and serve three DML surfaces through one lookup (`Engine::find_by`, longest-prefix index selection, scan fallback): relation reads (`self.prices`), schema-layer `select Type{ field == expr }` expressions in method bodies, and REST list filters (`GET /api/prices?product=1`; unknown field → 400). -- **Unknown `@table` keys are parse errors** (`shard_key:` and `retention:` are reserved for later phases — no silent passthrough on owned surface). Unknown annotation *names* (`@foo`) skip silently, like unknown field annotations. -- Indexes are per-shard structures over that shard's rows (shared-nothing, plan 09); cross-shard list filters fan out and merge exactly like unfiltered lists. - -## Query Layer — Hybrid SQL + Cypher, Fixed Glue - -Keep the syntax developers already know. Fix five things so the three grammars share semantics: - -1. **One parameter rule.** `$name` everywhere — SQL, Cypher, document path expressions. Typed at prepare time from the surrounding function signature or session context. -2. **Cross-paradigm `RETURNING`.** `INSERT … RETURNING id AS oid` binds the alias into subsequent statements in the same `BEGIN … COMMIT`. Replaces `LAST_INSERT_ID()`. -3. **One path rule.** `a.b.c[i].d` reads the same inside SQL expressions, document `UPDATE … SET`, and Cypher projections (`RETURN u.meta.name`). -4. **One transaction block.** `BEGIN [SNAPSHOT|SERIALIZABLE] … [SAVEPOINT name; …] … COMMIT|ROLLBACK [TO name]`. No dialect split for stored procedures. -5. **One subscription prefix.** `LIVE ` returns a subscription handle. Same semantics on both sides. - -Every e-commerce query mixes at least two paradigms: - -```wo --- Relational + document: "products under $50 with >4 stars" -SELECT id, meta.title, price_cents -FROM products -WHERE price_cents < 5000 - AND AVG(meta.reviews[].stars) > 4.0; - --- Graph + relational: "top sellers among products similar to what I bought" -MATCH (me:user {id: $uid})-[:PURCHASED]->(p:product)-[:SIMILAR_TO]->(rec:product) -WHERE rec.inventory.on_hand > 0 -ORDER BY rec.meta.reviews.count DESC -LIMIT 20; - --- All three: atomic checkout, with RETURNING replacing LAST_INSERT_ID() -BEGIN SNAPSHOT - UPDATE products - SET inventory.on_hand = inventory.on_hand - $qty, - inventory.reserved = inventory.reserved + $qty - WHERE id = $pid AND inventory.on_hand >= $qty - RETURNING id AS pid; - - INSERT INTO orders (user_id, total_cents, status, line_items) - VALUES ($uid, $total, 'pending', - [{product_id: $pid, qty: $qty, unit_cents: $unit}]) - RETURNING id AS oid; - - MATCH (u:user {id: $uid}), (p:product {id: $pid}) - CREATE (u)-[:PURCHASED {order_id: $oid, qty: $qty, at: now()}]->(p); -COMMIT; - --- Subscription: same predicate language, LIVE prefix -LIVE SELECT id, status, total_cents FROM orders WHERE user_id = $uid; -``` - -The checkout query is the whole argument. Those three statements **must** commit together or not at all. `RETURNING id AS oid` threads the inserted order's id into the Cypher `CREATE` without inventing a special function call. No off-the-shelf system executes that atomically across SQL + JSONB + a graph store without stitching multiple transaction managers together (or adopting SurrealDB, which is the existence proof that this can be built). - -## Option C — Full `.wo` Language for an E-commerce Platform with ACID - -The [evaluation phase](./01-evaluation.md) argued against external databases for a small, single-writer content project. An **e-commerce platform inverts every one of those assumptions**: - -| writeonce (blog) | E-commerce platform | -| --- | --- | -| Hundreds of articles | Millions of products, orders, sessions | -| Single writer (author) | Thousands of concurrent writers (customers, fulfillment, admin) | -| Read-heavy | Write-heavy on the hot paths (cart, checkout, inventory) | -| Full rebuild on change is free | Full rebuild is impossible — mutations must commit in milliseconds | -| No transactions needed | ACID is the product | -| Data is one shape (article) | Data is genuinely three shapes: relational (orders, inventory), document (product descriptions, reviews), graph (recommendations, categories, affiliations) | - -The three-paradigm case that was marginal for a blog becomes **the correct design** for e-commerce. And ACID is not a nice-to-have — an order that decrements inventory but loses the payment record is a lawsuit. - -This is Option C: a full `.wo` language backed by a real storage engine, a real transaction manager, and a real planner. It is a database, and writeonce/the parent project becomes the vehicle for building it. - -### ACID — What Each Letter Requires - -**Atomicity.** The checkout above touches three storage areas. The engine needs a single transaction coordinator that owns writes to all three. On abort, every partial write rolls back. Implementation: one **write-ahead log (WAL)** records intents across all paradigms; commit flips a single on-disk marker; crash recovery replays or discards based on the marker. - -**Consistency.** Domain invariants that span paradigms must hold: -- `inventory.on_hand >= 0` (document field, relational-style constraint — expressed as `@check(>= 0)` in the schema layer) -- Every `PURCHASED` edge must reference an existing `orders.id` (graph-to-relational FK — enforced because `Purchase.order: ref Order` in the schema layer) -- `orders.total_cents == sum(line_items[].qty * line_items[].unit_cents)` (denormalization check — expressed as a computed field) - -The schema layer captures these as declarations; the compiler emits the planner-level checks. The query layer also supports explicit `CONSTRAINT` statements for invariants that don't fit the type system. - -**Isolation.** Thousands of concurrent carts racing for the last unit of inventory. Two realistic models: - -| Model | How it works | Trade-off | -| --- | --- | --- | -| **MVCC** (Postgres, SurrealDB, CockroachDB) | Each transaction sees a snapshot; conflicts detected at commit | Readers never block writers; abort rate rises under contention | -| **2PL with row/edge locks** (MySQL InnoDB) | Locks acquired on read/write, released at commit | Lower abort rate; deadlocks must be detected | - -MVCC is the modern default and what you'd target. Minimum isolation level for e-commerce: **Snapshot Isolation** (Postgres `REPEATABLE READ`). Anything weaker (`READ COMMITTED`) permits write skew — two customers each passing the "inventory >= 1" check and both succeeding on the last unit. - -**Durability.** On `COMMIT`, the WAL record must be `fsync`'d before the client gets acknowledgment. Lose fsync and you lose paid orders on power failure. The WAL is the primary storage commitment; data files are derived and can be rebuilt by replaying the log from the last checkpoint. - -### Storage Engine - -`.seg` + `.idx` rebuilt-on-change does not survive contact with e-commerce. The engine needs genuine mutable on-disk structures: - -| Component | Responsibility | Reference | -| --- | --- | --- | -| **WAL** | Ordered, fsynced log of every committed mutation | Postgres `pg_wal/`, RocksDB `*.log` | -| **Relational pages** | Fixed-size pages (4–16 KB) with slot-directory row layout, B+ tree indexes | Postgres heap + btree, SQLite | -| **Document store** | LSM tree (SSTables + memtable + compaction) for append-friendly JSON blobs | RocksDB, SurrealKV | -| **Graph store** | Adjacency list on disk — doubly-linked edge records per node for O(1) traversal | Neo4j's native store | -| **Buffer pool** | Shared page cache with LRU/CLOCK eviction, dirty page tracking | Postgres `shared_buffers` | -| **Checkpointer** | Periodically flushes dirty pages, truncates WAL | Every major DB has one | -| **Vacuum / compaction** | Reclaim space from MVCC dead tuples / LSM tombstones | Postgres autovacuum, RocksDB compaction | - -Three storage backends, one WAL, one transaction coordinator. That is the core of the project. The [In-Memory Engine](./03-inmemory-engine.md) phase details the RAM-primary variant of this. - -### Cross-Paradigm Transaction Coordinator - -The novel piece — nobody ships this exactly the way `.wo` would need it: - -``` - ┌──────────────────────────────────────┐ - │ .wo Query Planner │ - └──────────────────────────────────────┘ - │ - ▼ - ┌──────────────────────────────────────┐ - │ Transaction Coordinator (MVCC) │ - │ - txn_id allocation │ - │ - snapshot timestamp │ - │ - commit ordering │ - │ - WAL append + fsync │ - │ - RETURNING alias table per txn │ - └──────────────────────────────────────┘ - │ │ │ - ▼ ▼ ▼ - ┌────────┐ ┌──────────┐ ┌──────────┐ - │ Rel │ │ Doc │ │ Graph │ - │ Engine │ │ Engine │ │ Engine │ - │ (B+) │ │ (LSM) │ │ (adj) │ - └────────┘ └──────────┘ └──────────┘ -``` - -Every engine exposes the same transaction hooks: `begin(snapshot_ts)`, `stage(mutation)`, `prepare()`, `commit(wal_lsn)`, `abort()`. The coordinator drives a **two-phase commit internally** (not distributed 2PC — it's one process, so prepare+commit is cheap and deterministic). - -Snapshot read across paradigms: each engine stores per-record `(created_txn_id, deleted_txn_id)` visibility info. A read at snapshot timestamp `T` sees only records visible at `T` — same rule in all three engines. - -`RETURNING` aliases live in the transaction's scoped name table; each subsequent statement inside the same `BEGIN … COMMIT` resolves `$oid` etc. against it. This is how SQL results thread into Cypher `CREATE` atomically, without round-tripping to the client. - -### Concurrency Model - -**One process, one thread, one event loop.** Redis-style. The entire engine — connection accept, parser, planner, executor, buffer pool, subscription registry — runs on a single userland thread; the only non-userland thread is the kernel-owned io_uring SQPOLL helper, which is invisible to engine code. - -| Subsystem | Where it runs | -| --- | --- | -| Connection accept | Event loop — non-blocking `accept()` via io_uring | -| Query execution | Event loop — parse, plan, execute inline | -| WAL fsync | Event loop submits SQEs to io_uring; kernel-owned SQPOLL thread drains; loop parks on CQE | -| Subscription dispatch | Event loop — predicates matched on commit, deltas pushed to per-subscription ring buffers | -| Compaction / checkpoint / vacuum | Event loop — scheduled as low-priority tasks between client work | - -**Why single-threaded.** Three precedents: - -- **Redis** ran single-threaded for its first decade (and still runs command execution single-threaded; 6.0 added threaded I/O only). It hits hundreds of thousands of ops/second on one core. -- **TigerBeetle** is single-threaded by design because determinism beats concurrency for financial workloads. -- **Node.js** is the proof that the event-loop model scales for I/O-bound work at web-scale. - -Single-threaded execution removes entire failure modes: no lock ordering, no cross-thread memory ordering hazards, no MVCC visibility logic for reads racing with writes, no torn-page atomics. The transaction coordinator trivially serializes commits because there is only one of them at a time. Snapshot isolation reduces to a logical versioning scheme for live-query delta computation — not a multi-threaded correctness mechanism. - -**Group commit still applies.** The loop drains many pending commits into one fsync SQE per tick — same amortization as Postgres's group-commit path, without a dedicated WAL-writer thread. Expected throughput on a modern NVMe server: ~200–500k simple-transaction commits per second per core. More than enough for every writeonce-sized workload. - -**Scaling past one core.** The path is **sharding**: partition types across independent single-threaded engine processes (Redis Cluster is the reference). Each shard owns a disjoint set of types; cross-shard transactions use two-phase commit between shards. Revisit only when a production workload actually saturates the core — the multi-threaded single-engine alternative is years of work for the wrong kind of gain. - -### Query Language Scope (`.wo` Full Spec) - -The minimum grammar for an e-commerce workload is split by layer: - -**Schema layer (enters at Phase 5):** - -- `type Name { ... }` declarations with scalar fields, embedded structs, `ref`, `multi @edge`, `multi … via LinkType`, `backlink` -- `type Name link A -> B { ... }` for edge-with-properties -- Tagged unions `A | B | C` -- Annotations `@unique @check @default @index @search @immutable` -- Computed fields `total: Money = sum(…)` -- Per-type `policy` / `on ` / `service` blocks (full semantics in [Phase 6](./06-lowcode-fullstack.md)) - -**Query layer (Phase 2 prototype):** - -- **DDL (interim)**: the existing `##sql / ##doc / ##graph` block syntax, as the compiler target for the schema layer and the prototype's direct authoring surface -- **DML relational**: `INSERT`, `UPDATE`, `DELETE`, `UPSERT` — all supporting `RETURNING col AS alias` -- **DML document**: path updates (`SET meta.reviews[3].stars = 5`), array operations (append, remove, splice) -- **DML graph**: `CREATE`, `MERGE`, `DELETE` on nodes and edges; variable-length path (`*1..5`) -- **Queries**: `SELECT` with joins, `MATCH` with traversal, subqueries, aggregations (`COUNT`, `SUM`, `AVG`) -- **Transactions**: `BEGIN [SNAPSHOT|SERIALIZABLE]` / `COMMIT` / `ROLLBACK`, `SAVEPOINT name` / `ROLLBACK TO name` -- **Parameters**: `$name` everywhere, typed at prepare -- **Cross-paradigm expressions**: dotted paths descend relational → document (`order.line_items[0].qty`); graph bindings resolve to relational rows (`(u:user {id: $uid})`); `RETURNING` aliases resolve across SQL → Cypher boundary -- **Subscriptions / live queries**: `LIVE SELECT` / `LIVE MATCH` — SurrealDB-style; predicate known at registration time for O(1) lookup matching on commit -- **Prepared statements + parameter binding**: `$name` with typed signatures; required for SQL injection resistance -- **Role-based auth + row-level policies**: schema-layer `policy` declarations compile to planner rewrite rules AND'd into every query at registration time - -### Components to Build - -``` -.wo engine -├── parser Pratt or LALR — handles both layers and all 3 paradigms -├── analyzer name resolution, type checking, cross-paradigm dispatch -├── schema compiler type-DSL → physical schema (Phase 5) -├── planner cost-based: reorder joins, push predicates, choose index, -│ merge policy predicates at registration time -├── executor vectorized over relational, iterator over graph -├── txn manager MVCC, snapshot isolation, group commit, RETURNING aliases -├── wal ordered log, fsync, checkpoint, recovery -├── rel engine B+ tree heap, btree indexes, vacuum -├── doc engine LSM with bloom filters, compaction -├── graph engine native adjacency, property store -├── buffer pool shared page cache, dirty tracking -├── catalog schema metadata (from type DSL), evolvable at runtime -├── wire protocol client connections (pick: Postgres wire, or custom) -├── auth + rbac users, roles, row-level policies from type declarations -├── live queries incremental view maintenance for subscriptions -├── backup / replication physical log shipping; logical streaming for read replicas -└── observability query stats, slow log, lock waits, WAL lag -``` - -Implementation language is genuinely open. Reasonable picks and what each implies: - -| Language | Why | Precedent | -| --- | --- | --- | -| **Rust** | Memory safety without GC, good async story, zero-cost abstractions, fits writeonce's existing stack | SurrealDB, TiKV, Materialize, sled | -| **C++** | Lowest overhead, decades of mature DB internals literature | Postgres (C), MySQL, RocksDB, DuckDB, ClickHouse | -| **Zig** | C-like control with safer semantics, compile-time metaprogramming useful for query planner | TigerBeetle | -| **Go** | Fastest to productive, excellent concurrency primitives, some GC cost on hot paths | CockroachDB, InfluxDB, Dgraph | -| **OCaml / Haskell** | Query planner is a compiler; ML-family languages are excellent at compilers | Irmin, some research DBs | - -For e-commerce ACID specifically, **Rust or C++** — GC pauses during a checkout fsync batch are the kind of latency spike that loses money. Go works and has the fastest developer velocity, but CockroachDB has spent years tuning around GC; budget for that. - -### Reference Implementations to Study - -- **SAP CDS** — . The canonical declaration-first application language; direct inspiration for the schema layer. -- **EdgeDB / EdgeQL** — path-based query language over a typed schema; studies the corner cases of computed fields, link properties, and tagged unions. Worth reading their planner learnings before re-implementing. -- **Postgres** — the canonical ACID RDBMS. Read `src/backend/access/transam/` for WAL, `src/backend/storage/buffer/` for the buffer pool, `src/backend/storage/lmgr/` for locking. Decades of battle-tested code. -- **SurrealDB** — closest living example of the query layer. Multi-model, multi-paradigm query language, Rust, embeddable. Already referenced in [surreal-case-study.md](../surreal-case-study.md). -- **SQLite** — smallest complete ACID database in the open-source world. `src/btree.c`, `src/pager.c`, `src/wal.c` are worth reading front-to-back. -- **CockroachDB** — distributed SQL with serializable isolation. Go, but the transaction protocol (Parallel Commits) is well-documented. -- **TigerBeetle** — financial-grade ACID, deterministic, written in Zig. Essay on why they rewrote: . -- **DuckDB** — analytics-focused but single-file embeddable C++ engine, excellent reference for a modern vectorized executor. -- **Neo4j** — for the graph-storage side. Native adjacency, transaction log, lock manager. -- **Papers**: - - *Architecture of a Database System* (Hellerstein, Stonebraker, Hamilton, 2007) — the canonical survey - - *The Log-Structured Merge-Tree* (O'Neil 1996) — for the LSM document backend - - *Serializable Snapshot Isolation in PostgreSQL* (Ports & Grittner 2012) — making SI safe - - *A Critique of ANSI SQL Isolation Levels* (Berenson et al. 1995) — so you pick the right default - -### Realistic Scope - -This is a multi-person-year project. Concrete gates: - -| Milestone | What's usable | Rough effort | -| --- | --- | --- | -| Parser + analyzer + in-memory executor | Single-user prototype, no durability (**shipped** at `prototypes/wo-db/`) | 2–4 months | -| Fixed-glue query layer (`$name`, `RETURNING`, `BEGIN/SAVEPOINT/COMMIT`, `LIVE` keyword) | Multi-statement cross-paradigm transactions threadable | +1–2 months | -| WAL + crash recovery + single-table B+ tree | ACID on relational only, one writer | +3–6 months | -| MVCC + concurrent transactions | Multi-writer relational | +3–6 months | -| Document engine (LSM) | Relational + document, transactional | +4–8 months | -| Graph engine + cross-paradigm txns | All three paradigms ACID | +6–12 months | -| Schema-layer compiler (type DSL → physical schema) | Single source of truth for codegen and full-stack blocks | +2–4 months | -| Wire protocol + RBAC + observability | Deployable to production | +3–6 months | -| Replication + backup | Survive a node loss | +6–12 months | - -Two to four years for a small team to reach something a real e-commerce business would trust with payment data. Adopting Postgres (with JSONB for the document side and the Apache AGE extension or a separate graph store for the graph side) gets you there in a week. - -### Honest Decision Framing - -Build `.wo` for an e-commerce platform **only if** at least one of the following is true: - -1. The cross-paradigm query atomicity is a business-critical feature the founders want to sell ("our DB does what Postgres + Neo4j glued together cannot"). This is the SurrealDB and EdgeDB thesis. -2. Building the database **is** the product — the e-commerce platform is the test harness, not the goal. -3. You have a team comfortable with the papers listed above and the patience to ship a toy for 18 months before it's useful. - -Otherwise, the pragmatic stack for a multi-paradigm ACID e-commerce platform is: - -- **Postgres** for relational + document (JSONB) + row-level security. Handles 99% of the workload. Apache AGE extension adds Cypher-compatible graph queries in the same transaction. -- **Redis** for cart/session/rate-limit (expiring, non-durable). -- **Search** (OpenSearch/Meilisearch/Typesense) for product search — specialized workload. -- **Event log** (Kafka/Redpanda) for order events, downstream analytics, fulfillment. - -That stack is boring and it works. `.wo` as described is interesting and would take years. Pick based on whether the goal is to ship e-commerce or to ship a database. diff --git a/docs/runtime/database/03-inmemory-engine.md b/docs/runtime/database/03-inmemory-engine.md deleted file mode 100644 index 10eafcc..0000000 --- a/docs/runtime/database/03-inmemory-engine.md +++ /dev/null @@ -1,205 +0,0 @@ -# Phase 3 — In-Memory Engine - -> RAM-primary, SSD-durable storage using io_uring, designed for OLTP e-commerce workloads on a 64 GB Linux machine. - -**Previous**: [Phase 2 — The `.wo` Language & ACID Engine](./02-wo-language.md) | **Next**: [Phase 4 — Client API](./04-client-api.md) | **Index**: [database.md](../database.md) - ---- - -Seed constraints: - -- 64 GB RAM — the entire live dataset fits in memory; no page eviction on the hot path -- Dual write — every mutation goes to an in-RAM structure **and** to an on-SSD durable log simultaneously -- Linux-only — Linux kernel APIs are fair game, no portability obligation -- `io_uring` for asynchronous read/write to SSD - -This is a **RAM-primary, SSD-durable** engine — the modern OLTP architecture used by TigerBeetle, ScyllaDB (via Seastar), VoltDB/H-Store, SAP HANA, and Redis-with-AOF. Reads never touch disk. Writes touch RAM immediately and SSD asynchronously, with fsync gating commit acknowledgment. - -For the e-commerce workload in [Phase 2](./02-wo-language.md), this is the right physical design: checkout latency is dominated by the durability path, not lookup; and a 64 GB live set comfortably holds millions of products + recent orders + active sessions + the entire recommendation graph. - -## Memory Layout - -All three paradigm engines live in one address space: - -``` -┌──────────────────────────────────────────────────────────────┐ -│ Process address space │ -├──────────────────────────────────────────────────────────────┤ -│ Relational heap B+ tree pages (row-store) ~20 GB │ -│ Document store LSM memtable + sorted runs ~15 GB │ -│ Graph store Node + edge arenas, adjacency ~10 GB │ -│ Index shards Hash, sorted, inverted ~8 GB │ -│ Buffer for WAL Ring buffer staged for SSD ~2 GB │ -│ MVCC version chains Per-record visibility history ~5 GB │ -│ Connection / query Per-session scratch ~2 GB │ -│ Headroom / OS Free for kernel, page tables ~2 GB │ -└──────────────────────────────────────────────────────────────┘ -``` - -Key moves: - -- **`mlockall(MCL_CURRENT | MCL_FUTURE)`** — pin all pages, guarantee no swap-out. A single swap-in during checkout is a latency catastrophe. -- **`MAP_HUGETLB` / Transparent Huge Pages** — 2 MB pages reduce TLB pressure on hot indexes. For 64 GB of data, 4 KB pages mean 16 M TLB entries; 2 MB pages mean 32 K. -- **`/proc/sys/vm/swappiness = 0`** — belt and braces with `mlockall`. -- **NUMA awareness** — optional on multi-socket hosts. Since the engine is single-threaded ([Phase 2 concurrency model](./02-wo-language.md#concurrency-model)), pin the one event-loop thread and bind the arena to that socket (`numactl --membind=0 --cpunodebind=0`). Cross-socket memory access is 2–3× slower; a single-socket deployment sidesteps it entirely. -- **Slab / arena allocators** — avoid `malloc` in the hot path. Pre-size arenas per paradigm at startup. - -## Dual-Write Durability Path - -A write is committed only when its WAL record is on the SSD with `fsync` confirmed. The in-memory structure is updated first (fast), then the log write is awaited (slow enough to matter): - -``` - Transaction COMMIT - │ - ▼ - ┌──────────────────────┐ - │ 1. Stage mutations │ apply to RAM structures under MVCC - │ to in-memory │ (version chains; readers unaffected) - │ engines │ - └──────────────────────┘ - │ - ▼ - ┌──────────────────────┐ - │ 2. Serialize WAL │ append-only ring buffer in RAM - │ record │ header + paradigm deltas + LSN - └──────────────────────┘ - │ - ▼ - ┌──────────────────────┐ - │ 3. io_uring submit │ IORING_OP_WRITE with O_DIRECT - │ WAL write to SSD │ batched with concurrent commits - └──────────────────────┘ - │ - ▼ - ┌──────────────────────┐ - │ 4. io_uring submit │ IORING_OP_FSYNC - │ fsync (barrier) │ linked SQE after the write - └──────────────────────┘ - │ - ▼ - ┌──────────────────────┐ - │ 5. CQE received │ commit marker flipped - │ → ack client │ MVCC snapshot published - └──────────────────────┘ -``` - -Steps 1–2 are synchronous; 3–5 are asynchronous. The event loop submits the SQEs and moves on to the next client; it reaps the CQE on a later tick — so the loop can have thousands of commits in flight without parking on any fsync syscall. - -**Group commit**: the loop drains the commit queue into one fsync SQE per tick. If 500 transactions all committed within a 100 μs window, one fsync durable-s the batch. Amortizes SSD latency (~50–100 μs on NVMe) across the batch — throughput approaches `batch_size / fsync_latency`, which on a good NVMe is 500K+ commits/sec. - -**What "dual write" means here**: it is *not* a two-database write where both must succeed independently. It is one logical commit that updates RAM (the query surface) and appends to the SSD WAL (the recovery record). On crash, RAM is gone; recovery replays the WAL to rebuild RAM state. The SSD is the source of truth for *durability*; RAM is the source of truth for *reads*. - -## io_uring Mechanics - -`io_uring` (Linux 5.1+, mature by 5.11) is the replacement for `epoll` + `libaio` for storage I/O. Two lock-free ring buffers shared between user-space and kernel: - -| Ring | Direction | Contents | -| --- | --- | --- | -| **SQ** (Submission Queue) | Userland → Kernel | SQEs: `IORING_OP_WRITE`, `IORING_OP_FSYNC`, `IORING_OP_READ`, etc. | -| **CQ** (Completion Queue) | Kernel → Userland | CQEs: result code + user_data pointer back to the request | - -Configuration knobs that matter for a database: - -- **`IORING_SETUP_SQPOLL`** — a kernel thread polls the SQ. Userland writes SQEs without any syscall. Read/write submission becomes a memory write + memory barrier. Cost: one pinned kernel thread per ring. -- **`IORING_SETUP_IOPOLL`** — busy-poll for completions on the device instead of interrupt-driven. Lower latency on NVMe, higher CPU. Requires `O_DIRECT`. -- **`IORING_REGISTER_BUFFERS`** — pre-register WAL ring-buffer pages with the kernel. Skips per-I/O page pinning. -- **`IORING_REGISTER_FILES`** — pre-register the WAL fd. Skips fd table lookups per I/O. -- **Linked SQEs (`IOSQE_IO_LINK`)** — enforce ordering: write-then-fsync, or WAL-then-commit-marker. Kernel guarantees link order without userland waiting on the intermediate CQE. -- **`O_DIRECT`** on the WAL file — bypass the kernel page cache. The database manages its own buffering; double-caching wastes the 64 GB. - -Per-commit path with full optimization: no syscalls at all for submission (SQPOLL), one memory read for completion (IOPOLL), zero page-pinning cost (registered buffers), zero fd-table lookup (registered files). The commit loop is effectively as fast as the NVMe firmware allows. - -## Recovery - -RAM is volatile; on restart the engine is empty. Recovery rebuilds it: - -1. **Open WAL**. Scan forward from the last checkpoint LSN. -2. **Replay committed records.** Apply each to the in-memory engines in LSN order. Skip incomplete transactions (no commit marker). -3. **Load checkpoint snapshot** (optional but standard). Periodically, the engine dumps a consistent snapshot of the RAM state to SSD. On recovery, load snapshot → replay WAL from snapshot LSN forward. Avoids replaying hours of log. -4. **Rebuild indexes.** Indexes are derived from heap data — rebuilt during replay or lazily on first access. -5. **Open for traffic.** - -Recovery target: 60 GB of data + a few million WAL records = seconds to a minute on NVMe, not hours. A good checkpointer runs in the background every 5–15 minutes; recovery only replays the delta since the last checkpoint. - -## Concurrency in RAM - -No page eviction, no buffer pool locks — and because the engine is [single-threaded](./02-wo-language.md#concurrency-model), no cross-thread races either. The concurrency story collapses to "there is no concurrency within the engine; there is a queue of clients being served sequentially by one loop". Every data structure is owned by that one loop: - -| Structure | Primitive | -| --- | --- | -| B+ tree (relational) | Plain owned tree; no latches, no optimistic locks | -| LSM memtable (document) | Plain skiplist; sealed memtables still immutable for background compaction SQEs | -| Graph adjacency | Plain hashmap per node label | -| MVCC version chain | Plain singly-linked version list; no CAS | -| WAL ring buffer | Single-producer, single-consumer ring | -| Txn coordinator | Plain `u64` counter — incremented without atomics | - -**Readers still see snapshots.** MVCC remains useful but its purpose changes: instead of "readers don't block writers on another thread", it's "a live-query subscriber reading in the same tick sees the pre-commit view; the post-commit delta arrives on the next tick". That semantic is cheap to implement when there is only one mutator. - -**The model is sequential.** Clients are served round-robin by the loop; nothing races because nothing runs concurrently inside the engine. When one core isn't enough, [shard](./02-wo-language.md#concurrency-model) rather than bolting multi-threading onto this design. - -## Capacity Planning - -64 GB is a budget, not a guarantee. Three failure modes to design around: - -1. **Working set exceeds RAM.** Solution path: add a tier (warm SSD-backed region for cold rows), or shard across nodes. Neither is in the Phase 2 scope — flag when live data approaches 50 GB. -2. **MVCC version chains bloat.** Long-running transactions hold old versions alive. Solution: aggressive vacuum, transaction timeouts, snapshot horizon tracking. Postgres hits this same wall. -3. **Sudden write bursts flood the WAL.** Solution: admission control — if SSD write queue depth exceeds a threshold, slow down `COMMIT` acknowledgment. Better than OOM'ing the WAL buffer. - -## Comparison With Alternatives - -| Aspect | In-memory + WAL (this design) | Disk-primary (Postgres) | Pure in-memory (Redis w/o AOF) | -| --- | --- | --- | --- | -| Read latency | ~100 ns (RAM) | ~10 μs (buffer cache hit) to ms (miss) | ~100 ns (RAM) | -| Write latency | ~50–100 μs (fsync) | ~50–100 μs (fsync) | ~100 ns (none) | -| Durability | Full — WAL fsync before ack | Full — WAL fsync before ack | Window of loss (AOF every-sec) | -| Dataset size | Bounded by RAM | Bounded by disk | Bounded by RAM | -| Restart time | Seconds to minutes (WAL replay) | Seconds | Immediate (empty) or minutes (AOF) | -| Ideal workload | OLTP with small-to-medium dataset | General-purpose, large datasets | Cache, session, ephemeral | - -This design keeps the durability of Postgres and the read speed of Redis. - -## Linux Tuning Checklist - -Before production benchmarking: - -- `echo 0 > /proc/sys/vm/swappiness` -- `echo never > /sys/kernel/mm/transparent_hugepage/enabled` (databases typically prefer explicit hugepages over THP's defragmentation stalls) -- `vm.nr_hugepages = ` -- `ulimit -l unlimited` (for `mlockall`) -- `blk-mq` scheduler: `none` or `mq-deadline` on NVMe (not `cfq`/`bfq`) -- `IORING_SETUP_SINGLE_ISSUER` — always, since the engine is single-threaded (Linux 6.0+) -- NUMA: `numactl --membind=0 --cpunodebind=0` to pin the loop + its arena to one socket. Multi-socket deployments should shard across sockets rather than sharing one engine -- Disable CPU frequency scaling (`cpupower frequency-set -g performance`) — saves microseconds that add up across group-commit batches -- Disable Meltdown/Spectre mitigations only if you control the hardware and understand the trade-off — they cost 10–30% on syscall-heavy paths, but io_uring with SQPOLL largely sidesteps them anyway - -## Reference Implementations - -- **TigerBeetle** — Zig, in-memory, io_uring end-to-end, deterministic, designed for financial OLTP. The closest living example of this exact architecture. -- **ScyllaDB / Seastar** — C++, io_uring (and SPDK), shared-nothing per core, NUMA-aware. Seastar is the framework underneath. -- **VoltDB (H-Store)** — Java, in-memory OLTP, command-logging for durability. The academic ancestor of this design pattern. -- **Redis (`appendonly yes` + `appendfsync always`)** — simpler but exact same shape: RAM-primary, log-durable. -- **LMDB** — memory-mapped B+ tree; reads are literal pointer chases into mmap'd pages. Not WAL-based but worth studying for RAM-resident read paths. -- **SingleStore (formerly MemSQL)** — commercial in-memory row-store with columnar on-disk secondary. Hybrid of this design and disk-primary. -- **Readings**: - - *The End of an Architectural Era* (Stonebraker et al., 2007) — the H-Store paper that argued disk-primary databases were legacy for OLTP. - - *Efficient Lock-Free Durable Sets* (Zuriel et al.) — for lock-free structures that persist. - - *io_uring by Example* (Jens Axboe) and the `liburing` documentation — the authoritative guide. - -## Where This Fits in Phase 2 - -This replaces the storage-engine block in the [Phase 2](./02-wo-language.md) component list. Specifically: - -| Phase 2 component | Becomes (with in-memory design) | -| --- | --- | -| Relational pages + buffer pool | RAM-resident B+ tree / Masstree, no page eviction | -| Document engine (LSM on disk) | LSM memtable in RAM; sealed memtables spilled to SSD only for checkpointing | -| Graph engine (disk adjacency) | RAM adjacency arena; checkpointed, not paged | -| WAL | `io_uring` + `O_DIRECT` append-only log on NVMe | -| Checkpointer | Periodic snapshot of RAM arenas to SSD for fast recovery | -| Buffer pool | **Removed** — all data is in RAM | -| Vacuum | Still needed, but for MVCC chain pruning, not for reclaiming disk pages | - -The cross-paradigm transaction coordinator sketched in Phase 2 stays the same — it just drives in-memory engines instead of disk-paged ones, and the WAL append it depends on is the one `io_uring` path. - -Net effect: **shorter read paths, identical durability story, same ACID guarantees**, at the cost of a hard dataset ceiling set by RAM. diff --git a/docs/runtime/database/04-client-api.md b/docs/runtime/database/04-client-api.md deleted file mode 100644 index 3d25010..0000000 --- a/docs/runtime/database/04-client-api.md +++ /dev/null @@ -1,279 +0,0 @@ -# Phase 4 — Client API: Wire Protocol and Subscriptions - -> How remote clients connect, query, and subscribe to live changes — no polling anywhere in the chain. - -**Previous**: [Phase 3 — In-Memory Engine](./03-inmemory-engine.md) | **Next**: Phase 5 — Go Client SDK | **Index**: [database.md](../database.md) - ---- - -Once the engine is complete it is a **traditional database server**: remote clients connect over the network, issue queries, receive results, and (critically for e-commerce UX) **subscribe to changes without polling**. - -The client API has three problems to solve: - -1. **Wire protocol** — how bytes move between client and server. -2. **Query surface** — what queries look like from the client's perspective (raw `.wo`, SQL, GraphQL, REST). -3. **Subscriptions** — how the server pushes change notifications to clients when matching data mutates, with no client-side polling. - -## The Polling Problem This Must Avoid - -Every naive realtime system reaches for polling first. For an e-commerce platform it is disqualifying: - -| Polling | Subscriptions | -| --- | --- | -| Client asks "anything new?" every N ms | Server tells client "here's what changed" when it changes | -| Wasted RTTs when nothing changed | Zero traffic when nothing changes | -| Stale data up to N ms old | Sub-ms latency after commit | -| `O(clients × poll_rate)` server load | `O(mutations × matched_subscribers)` — scales with real change, not client count | -| Inventory display lies for up to N ms | Inventory display reflects the commit | -| "Order shipped" email triggered by cron | Fired by a committed status-change | - -Every subscription-based design in this section follows the same rule already set by writeonce in 05-datalayer.md and 03-data.md: **the client registers a query once, the server pushes deltas on commit, the client never asks again.** - -## Protocol Layer — Pick One or Both - -Two protocol tiers make sense: a **native binary protocol** for app servers and ORMs that want every microsecond, and a **GraphQL-over-WebSocket layer** for browsers, mobile apps, and third parties. They are not alternatives — they share the same planner and subscription registry underneath. - -| Option | Best for | Trade-off | -| --- | --- | --- | -| **Custom binary over TCP** | App servers, in-house clients, highest throughput | Need to ship client libs in every language | -| **Postgres wire protocol** (libpq) | Reuse the Postgres client ecosystem (psql, pgx, node-postgres, JDBC) | Locked into Postgres's shape — no native graph/live-query verbs | -| **gRPC (HTTP/2)** | Cross-language, well-tooled, server-streaming RPC covers subscriptions | Protobuf schema overhead; HTTP/2 stack cost | -| **GraphQL over HTTP + WebSocket** | Web/mobile clients, schema-aware tooling, built-in `subscription` operation | Parser/resolver overhead, N+1 risks | -| **REST + SSE** | Simplest to integrate (curl, browser `fetch`) | Verb-per-endpoint sprawl, SSE is unidirectional | - -**Recommended combination:** - -- **Native binary protocol** for first-party app servers (cart service, checkout, fulfillment). -- **GraphQL over WebSocket** for everything else (web, mobile, partner APIs). - -Both terminate at the same **session layer** inside the server, which delegates to the `.wo` planner. - -## Native Binary Protocol — Shape - -A minimal framing that's compatible with io_uring on both ends: - -``` -┌────────┬────────┬──────────┬─────────────────────────────┐ -│ opcode │ req_id │ len │ payload │ -│ u8 │ u64 │ u32 │ bincode / msgpack │ -└────────┴────────┴──────────┴─────────────────────────────┘ -``` - -Opcode set: - -| Opcode | Direction | Purpose | -| --- | --- | --- | -| `HELLO` | C → S | Protocol version + auth credentials | -| `WELCOME` | S → C | Session id + server capabilities | -| `PREPARE` | C → S | Compile a `.wo` query, cache plan on server | -| `EXECUTE` | C → S | Run prepared plan with bound parameters | -| `RESULT` | S → C | Full result set for one query | -| `BEGIN` / `COMMIT` / `ROLLBACK` | C → S | Explicit transaction control | -| `SUBSCRIBE` | C → S | Register a live query, get a subscription id | -| `UNSUBSCRIBE` | C → S | Cancel a subscription | -| `DELTA` | S → C | Pushed change matching a subscription | -| `COMPLETE` | S → C | Subscription terminated server-side (schema change, etc.) | -| `ERROR` | S → C | Typed error with query context | -| `PING` / `PONG` | bidirectional | Dead connection detection (no polling for data — just keepalive) | - -Multiplexed: many in-flight `req_id`s per connection, responses interleaved. Matches io_uring's async nature naturally — a connection never blocks on a slow query. - -## GraphQL — Schema, Queries, Mutations, Subscriptions - -GraphQL has the three verbs e-commerce actually uses: - -| GraphQL operation | `.wo` mapping | -| --- | --- | -| `query` | `SELECT` / `MATCH` over the engine, single response | -| `mutation` | `INSERT` / `UPDATE` / `DELETE` / `CREATE` inside an implicit transaction | -| `subscription` | `LIVE SELECT` / `LIVE MATCH` — server pushes on match | - -**Schema generation.** The `.wo` DDL is the source of truth; the GraphQL SDL is generated from it: - -``` -##sql #products (id, sku, price_cents, meta, inventory) -##doc #product-meta (title, description, reviews, ...) -##graph (user)-[:PURCHASED]->(product) - │ - ▼ generator - │ -type Product { - id: ID! - sku: String! - priceCents: Int! - meta: ProductMeta! - inventory: InventoryLevel! - similarTo(limit: Int = 10): [Product!]! # graph traversal - purchasedBy: [User!]! # graph traversal -} -type Subscription { - productUpdated(id: ID!): Product! - inventoryChanged(sku: String!): InventoryLevel! - orderStatus(orderId: ID!): Order! -} -``` - -**Subscription example (e-commerce checkout feedback loop):** - -```graphql -subscription CartInventory($skus: [String!]!) { - inventoryChanged(sku_in: $skus) { - sku - onHand - reserved - } -} -``` - -A web client opens this WebSocket subscription when the cart renders. The server only pushes when a committed transaction changes any of those SKUs' inventory — the cart's "2 left!" badge is always live, no polling. - -**Transport: `graphql-ws` protocol over WebSocket.** Standard, well-tooled (Apollo, urql, Relay, Hasura all speak it). Falls back to HTTP POST for plain queries and mutations. - -## Subscription Engine — How Push Actually Works - -This is the mechanism that makes polling unnecessary. It lives inside the transaction coordinator from [Phase 2](./02-wo-language.md): - -``` - ┌─────────────────────────────────────────────────────┐ - │ Transaction Coordinator (MVCC) │ - │ │ - │ on COMMIT(txn): │ - │ delta = collect_changes(txn) │ - │ matched = subscription_registry.match(delta) │ - │ for (sub, rows) in matched: │ - │ sub.writer.push(DELTA { sub.id, rows }) │ - └─────────────────────────────────────────────────────┘ - │ │ - ▼ ▼ - ┌─────────────────────┐ ┌──────────────────────────┐ - │ Subscription │ │ Session Writer (per conn)│ - │ Registry │ │ - native: io_uring send │ - │ │ │ - graphql: ws frame │ - │ predicate → [subs] │ │ - grpc: server stream │ - └─────────────────────┘ └──────────────────────────┘ -``` - -**Matching strategies**, in order of cost: - -| Subscription shape | Matching cost | Example | -| --- | --- | --- | -| Keyed (primary key) | O(1) hash lookup on commit | `productUpdated(id: 42)` | -| Tag / secondary index | O(1) index lookup + scan of matched rows | `orderStatusByUser(userId: 17)` | -| Range | O(log n) index range + filter | `ordersPlaced(between: [start, end])` | -| Graph traversal | O(edges visited) — bound by depth/limit | `recommendationsFor(userId: 17)` | -| Arbitrary predicate | O(subs) — evaluate each against the delta | `LIVE SELECT ... WHERE complex` | - -The engine indexes subscriptions by their shape so the common cases (keyed, tag-based) don't pay the arbitrary-predicate price. This is **incremental view maintenance** — the same idea that SurrealDB live queries, Materialize, Hasura, and Feldera all implement at different levels of generality. - -## Connection I/O — io_uring All the Way - -The same `io_uring` that drives the WAL (per [Phase 3](./03-inmemory-engine.md)) also drives client sockets. One scheduler, not a mix of epoll for networking and io_uring for storage: - -| Operation | io_uring opcode | -| --- | --- | -| Accept new client | `IORING_OP_ACCEPT` | -| Read request frame | `IORING_OP_RECV` (with registered buffers) | -| Write result / delta | `IORING_OP_SEND` (with `IOSQE_IO_LINK` to chain writes) | -| TLS handshake | Userland ring integrated with `IORING_OP_RECV`/`SEND` (e.g., rustls or BoringSSL in non-blocking mode) | -| Close | `IORING_OP_CLOSE` | -| Keepalive | `IORING_OP_TIMEOUT` per connection | - -A subscription push is one SQE: `SEND(client_fd, delta_frame)`. Thousands of in-flight pushes across thousands of subscribers is just thousands of SQEs — the kernel batches the actual NIC writes. No thread-per-connection, no blocking send. - -## Session State - -Each connected client has server-side state: - -| State | Lifetime | Notes | -| --- | --- | --- | -| Identity / principal | Session | JWT or mTLS validated at `HELLO` | -| Current transaction | One txn at a time per session | Auto-rollback on disconnect | -| Prepared statements | Session | Plan cached, re-parameterized per `EXECUTE` | -| Active subscriptions | Session | All torn down on disconnect (free registry slots, stop pushing) | -| Role / RBAC context | Session | Feeds row-level policies into the planner | -| Back-pressure credits | Per-subscription | Client advertises how many outstanding `DELTA` frames it can buffer | - -On disconnect (TCP close, keepalive failure, `EPOLLHUP`-equivalent from io_uring completion): all sessions state is freed, all subscriptions unregistered. Same philosophy as `wo-sub`'s `EPOLLHUP` → automatic `unsubscribe(fd)` from 05-datalayer.md, scaled up to a real server. - -## Back-Pressure - -A slow client cannot be allowed to stall commits. The push path must never block on a socket write: - -1. Each subscription has a **bounded outbound queue** (say, 1024 deltas). -2. Writer thread drains the queue via `io_uring_send`. -3. On queue overflow, the engine has three policies: - - **Drop + resync**: mark the subscription as "behind", push a single `RESYNC` marker, client re-requests current state. - - **Coalesce**: fold consecutive deltas for the same key into one (last-writer-wins). - - **Disconnect**: close the connection; clients with a stale subscription reconnect. -4. The coordinator never waits on a subscription — it hands the delta to the writer and moves on. - -This is the same trade-off Kafka makes with consumer lag: fast producers, independent consumers, bounded buffer, spillover policy. - -## Authentication and Authorization - -Covered briefly in [Phase 2](./02-wo-language.md); the wire protocol is where it bites: - -- **Transport**: TLS mandatory for any non-loopback connection. Offload to `rustls` / `boringssl` userland; io_uring handles only the underlying sockets. -- **Authentication** at `HELLO`: JWT (stateless), API key (server-validated), or mTLS (cert-based). -- **Authorization**: RBAC + row-level policies evaluated inside the planner. A subscription's registered predicate is **intersected with the user's access policy at registration time** — if the policy says user 17 only sees their own orders, the subscription's effective predicate becomes `(original) AND user_id = 17`. Enforced once, not per push. -- **Rate limiting**: per-session token bucket enforced before any query work. Cheap to implement in the io_uring accept/recv path. - -## Comparison: This Design vs. Existing Products - -| Aspect | This design | Postgres + Hasura | Supabase Realtime | SurrealDB | Firebase | -| --- | --- | --- | --- | --- | --- | -| Transport | Custom binary + GraphQL/WS | SQL wire + GraphQL/WS | Postgres WAL → WS | HTTP + WS | Custom WS | -| Subscriptions | Native, planner-integrated | Live queries via polling Postgres | Logical replication fan-out | Native live queries | Native | -| Storage coupling | In-process | External Postgres | External Postgres | In-process | Proprietary | -| Cross-paradigm | Yes (`.wo`: rel + doc + graph) | Partial (JSONB, no graph) | Relational only | Yes (rel + doc + graph) | Doc only | -| Polling internally? | No | **Yes** (Hasura polls Postgres) | No (uses WAL) | No | No | -| io_uring throughout | Yes | No | No | Partial | No | - -Hasura is the instructive one — it gives clients push subscriptions, but internally it polls Postgres because Postgres has no commit-time subscription hook. Building the subscription engine *inside* the database (as this design does) is what eliminates polling end-to-end. - -## Reference Implementations - -- **SurrealDB** — the tightest match: custom engine, WebSocket transport, native `LIVE SELECT`. . Also in [surreal-case-study.md](../surreal-case-study.md). -- **Hasura GraphQL Engine** — production-quality GraphQL over Postgres with subscriptions. Read their `graphql-engine/server/src-lib/Hasura/GraphQL/Transport/` for subscription multiplexing. -- **Supabase Realtime** — Phoenix/Elixir server that tails Postgres logical replication and fans out over WebSocket. Cleanest demo of "subscriptions as a layer over an existing DB." -- **PostgREST** — auto-generated REST from Postgres schema. Simpler than GraphQL, same spirit. -- **EdgeDB** — custom binary protocol, custom query language (EdgeQL), compiles to Postgres underneath. Good reference for protocol framing. -- **Materialize** — incremental view maintenance as a product; every query is implicitly a subscription. -- **Phoenix Channels** (Elixir) — mature pub/sub-over-WebSocket with presence, back-pressure, and reconnection baked in. Worth reading even if the server is Rust/C++. -- **graphql-ws** protocol — . The WebSocket sub-protocol every modern GraphQL client speaks. -- **Apollo Router** — GraphQL gateway with subscription multiplexing, federation. - -## Scope Addition to Phase 2 - -The client API is a sizable addition to the [Phase 2](./02-wo-language.md) component list: - -| Component | New work | -| --- | --- | -| Native wire codec | Binary framing, opcode dispatch, session lifecycle | -| Postgres-wire compatibility (optional) | libpq protocol v3 parser — reuse clients | -| GraphQL layer | SDL generation from `.wo`, resolver dispatch, `graphql-ws` subscriptions | -| REST/SSE gateway (optional) | Thin translation to native protocol | -| Subscription registry | Indexed by subscription shape; matched on commit | -| Push writer pool | io_uring-backed, per-connection outbound queues, back-pressure policy | -| TLS / auth | rustls or boringssl, JWT/mTLS at connection open | -| Connection manager | Accept, keepalive, graceful shutdown, fd limits | -| Observability | Per-session stats, slow query log, subscription lag, push-queue depth | - -Rough incremental effort on top of Phase 2: **6–12 months** for a production-quality client layer with both native and GraphQL protocols, assuming the engine underneath is working. - -## Why This Matters for E-commerce - -Every hot user-facing screen is a subscription in disguise: - -| Screen | Subscription | -| --- | --- | -| Product page | `productUpdated(id)` — price/stock changes reflect instantly | -| Cart | `inventoryChanged(sku_in: cartSkus)` — "out of stock!" appears the moment it's true | -| Order status | `orderStatus(orderId)` — pending → paid → shipped, no refresh | -| Admin dashboard | `LIVE SELECT COUNT(*) FROM orders WHERE placed_at > NOW() - 1h` | -| Recommendations sidebar | `LIVE MATCH (me)-[:VIEWED]->-[:SIMILAR_TO]->(p)` | -| Seller notifications | `LIVE MATCH (order)-[:CONTAINS]->(p) WHERE p.seller_id = $me` | - -Each of these is `O(1)` server work per commit — the matching subscription is indexed by the thing that changed. Without subscriptions, every one of those screens would be a polling loop hammering the database. With subscriptions, server load scales with **actual state change**, not with client count × poll rate. - -That is the whole argument for building the subscription engine into the database rather than bolting a message bus onto the side: **the engine already knows when something committed. Publishing the delta is a function call, not another system.** diff --git a/docs/runtime/database/06-lowcode-fullstack.md b/docs/runtime/database/06-lowcode-fullstack.md deleted file mode 100644 index 4bd8932..0000000 --- a/docs/runtime/database/06-lowcode-fullstack.md +++ /dev/null @@ -1,389 +0,0 @@ -# Phase 6 — Low-Code Full-Stack: `.wo` as an Application Language - -> Expand `.wo` from a query language into a declarative application DSL — schema, services, UI, business logic, and authorization in one language, compiled into a single binary. - -**Previous**: Phase 5 — Go Client SDK | **Index**: [database.md](../database.md) - ---- - -Up to this point `.wo` is a **query language**. The next move is to expand it into an **application language** — a declarative, low-code/no-code DSL in the shape of **SAP Core Data Services (CDS)**: one language, one file extension, one compilation pipeline that produces database schema, service endpoints, UI screens, and business logic from the same source tree. - -The reference precedent is SAP CDS, where a small amount of `.cds` code declares: - -- Entities (tables), types, associations (relationships) -- Services that project entities to OData/REST -- UI annotations (`@UI.LineItem`, `@UI.Facet`) that drive SAP Fiori rendering -- Actions, functions, and authorization rules - -From those declarations, SAP generates a full running application — data model, REST API, CRUD UI, authorization layer — with the developer writing almost no imperative code. `.wo` aims at the same target, for the same reason: **most enterprise and e-commerce applications are 90% CRUD on structured data with live views; declaring what you want and letting the runtime generate the rest is faster than hand-writing it**. - -## Two Authoring Styles — Type-Attached vs Standalone - -Behavior declarations fall into two groups: - -| Group | Blocks | Authoring style | -| --- | --- | --- | -| **Behavior ON an entity** | `policy`, `on `, `service` | **Type-attached** — declared inside the `type` block they govern. One name-resolution rule, zero cross-file coupling for single-entity behavior. | -| **Cross-entity composition** | `##ui`, `##app`, `##logic` that spans entities | **Standalone** — a screen composes multiple entities via `source:`; an app manifest names routes; a workflow that touches both orders and inventory needs its own block. | - -Type-attached is the default — and the preferred authoring style because the schema layer ([Phase 2](./02-wo-language.md)) already names entities, and `policy read when author == $session.user` is most legible next to the `author: ref User` declaration it references. - -Both forms compile to the same runtime model. A type-attached `policy` block is de-sugared into the same planner rewrite rule as a standalone `##policy` block — splitting is an authoring convenience. - -## Project Layout - -Convention over configuration. A `.wo` project is a tree of `.wo` files, each in a role-specific directory. The compiler discovers files by path. - -``` -myapp/ -├── app/ -│ ├── database/ # schema — `type` declarations (with inline -│ │ ├── article.wo # policy, on, service blocks per type) -│ │ ├── user.wo -│ │ └── order.wo -│ ├── ui/ # standalone screens — ##ui -│ │ ├── list.wo -│ │ ├── detail.wo -│ │ ├── form.wo -│ │ └── dashboard.wo -│ ├── logic/ # cross-entity workflows — ##logic -│ │ └── order-workflow.wo -│ ├── auth/ # cross-entity / session-level policies — -│ │ └── policies.wo # ##policy for things that don't fit on a type -│ ├── api/ # service bundles that expose many types — -│ │ └── services.wo # ##service for multi-entity APIs -│ └── app.wo # root: name, routes, theme, i18n -├── migrations/ # generated, versioned schema migrations -├── static/ # hand-written assets (images, custom CSS) -└── wo.toml # project metadata -``` - -Most single-entity behavior lives next to the `type` in `database/`; `logic/` and `auth/` and `api/` are for the cross-entity cases. Every file contributes to a **single compiled model**. Splitting is for humans; the runtime sees one graph of declarations. - -## File Type Examples - -**`app/database/article.wo`** — one `type` declaration covers relational fields, embedded document, graph edges, **and** the per-entity policy/triggers/service: - -```wo -type Article { - id: Id - sys_title: Slug @unique - title: Text - published: Bool = false - author: ref User -- foreign key - meta: { -- embedded document - tags: [Text] - excerpt: Text - body_md: Markdown - reviews: [{ user: ref User, stars: Int, body: Markdown, at: Timestamp }] - } - created_at: Timestamp = now() - published_at: Timestamp? - - related: multi Article @edge(:RELATED_TO) - prerequisites: multi Article @edge(:PREREQUISITE) - - -- type-attached policy — replaces a separate ##policy block for - -- the single-entity case - policy read when published == true - policy read for role editor - policy read for role owner when author == $session.user - policy write for role editor - policy write for role owner when author == $session.user - policy delete for role admin - - -- type-attached trigger — fires inside the transaction on commit - on update when old.published == false and new.published == true - do set self.published_at = now() - do emit "article.published"(self) - do enqueue "send-subscriber-emails" with { article_id: self.id } - - -- type-attached service — exposes CRUD + subscribe on a REST path - service rest "/api/articles" expose list, get, create, update, delete, subscribe -} - -type User { ... } -- authors; AUTHORED is derivable from Article.author via backlink -``` - -The compiler emits the underlying `##sql` relational row, `##doc` embedded structure, and `##graph` edges from the single `type` declaration. `multi Article @edge(:RELATED_TO)` declares a zero-property graph edge whose direction and label match the original graph sketch; `ref User` emits a foreign-key column in the relational store. Inverses (`User.articles: backlink Article.author`) generate an `AUTHORED` edge — or a plain inverse column, at the planner's discretion. - -**`app/ui/list.wo`** — a live list view; renders to HTML, wires subscriptions automatically: - -```wo -##ui -#article-list - title: "Articles" - source: article - live: true -- auto-subscribes via LIVE query - - filter: - published = true - - columns: - - sys_title label: "Slug" - - title label: "Title" searchable - - meta.tags label: "Tags" renderer: tag-chips - - created_at label: "Created" renderer: relative-date - - author.name label: "Author" join: author_id -> user - - sort: - default: created_at desc - - actions: - row-click: /article/:sys_title - create: /article/new role: editor - row-edit: /article/edit/:id role: editor | owner - row-delete: delete role: editor confirm: true - - pagination: 20 -``` - -**`app/ui/detail.wo`** — a detail view composed of nested renderers, including a graph traversal: - -```wo -##ui -#article-detail - title: $article.title - source: article - key: sys_title - live: true - - sections: - - header: - fields: [title, author.name, created_at] - - body: - renderer: markdown - source: meta.body_md - - related: - title: "Related Articles" - renderer: list - source: Article{ sys_title == $key }.related -- schema-layer path - columns: [title, meta.excerpt] - live: true -``` - -Screens stay standalone because they compose data from multiple types. The `source:` expression is a schema-layer path (preferred) or a raw `MATCH`/`SELECT` from the query layer — both resolve to the same planner input. - -**`app/logic/order-workflow.wo`** — `##logic` is reserved for **cross-entity** triggers that don't belong on a single type. The on-article-published trigger lives on `type Article` (shown above); the on-order-placed trigger touches orders **and** every product in the line items, so it stays standalone: - -```wo -##logic -#on-order-placed - when: insert(Order) - do: - - validate: self.total == sum(self.line_items.*.qty * self.line_items.*.unit) - - for-each item in self.line_items: - - update: Product{ id == item.product.id } - set inventory.on_hand -= item.qty - assert inventory.on_hand >= 0 -``` - -**`app/auth/policies.wo`** — reserved for **session-level** or **cross-entity** rules that don't fit on a single type. Most RBAC lives type-attached (see the `policy read ...` block on `type Article` above). A standalone `##policy` is useful for things like "admins bypass all row filters": - -```wo -##policy -#admin-bypass - applies_to: Article, Order, User - when: role == admin - effect: skip-row-filters -``` - -**`app/api/services.wo`** — reserved for **multi-entity** API bundles. Single-entity services live type-attached (see `service rest "/api/articles"` on `type Article` above). A bundle endpoint that exposes a curated subset or a custom aggregation goes here: - -```wo -##service -#storefront - path: /api/storefront - protocols: [rest, graphql] - expose: - - Product as products operations: [list, get, subscribe] - - Article as articles operations: [list, get] - - categories: Product{ featured == true }.category -- custom path -``` - -**`app/app.wo`** — root manifest: - -```wo -##app -name: "writeonce" -version: 1 -theme: "light" -i18n: [en, de] - -routes: - / -> ui.article-list { filter: { published: true } } - /article/:slug -> ui.article-detail { key: $slug } - /admin/articles -> ui.article-list { role: editor } -``` - -## Compilation Pipeline - -The `.wo` compiler loads every `.wo` file in the tree and emits a single runtime bundle: - -``` - app/**/*.wo - │ - ▼ - ┌─────────────┐ - │ Parser │ one grammar, all block types - └─────────────┘ - │ - ▼ - ┌─────────────┐ - │ Analyzer │ name resolution across files, type check, policy check - └─────────────┘ - │ - ▼ - ┌─────────────────────────────────────────┐ - │ Unified Application Model (AST) │ - └─────────────────────────────────────────┘ - │ │ │ │ │ - ▼ ▼ ▼ ▼ ▼ - ┌────────┐ ┌──────────┐ ┌─────────┐ ┌────────┐ ┌──────────┐ - │ Schema │ │ Services │ │ UI │ │ Logic │ │ Policies │ - │ (DDL) │ │(endpoints)│ │(widgets)│ │(hooks) │ │ (rbac) │ - └────────┘ └──────────┘ └─────────┘ └────────┘ └──────────┘ - │ │ │ │ │ - ▼ ▼ ▼ ▼ ▼ - Migrations HTTP/GraphQL HTML / JSON Commit- Planner - applied to / native manifest time filters - engine dispatch (SSR or triggers merged into - client) every query -``` - -Each leaf maps to a runtime component from the earlier phases: - -- **Schema → engine**: the `.wo` DDL goes to the in-memory engine from [Phase 3](./03-inmemory-engine.md). Migrations rebuild the schema; data is preserved where possible. -- **Services → endpoints**: HTTP/GraphQL/native dispatch via the wire-protocol layer from [Phase 4](./04-client-api.md). -- **UI → widgets**: a new component — UI declarations compile to a render tree. Default renderer is server-rendered HTML (SSR) with a thin client runtime for subscription wiring. `live: true` on any UI node issues a `LIVE SELECT`/`LIVE MATCH` through the subscription engine and the client runtime swaps DOM fragments on each delta. -- **Logic → hooks**: triggers compile to server-side procedures invoked by the transaction coordinator on matching commits. Same transaction as the mutation — ACID across the hook's writes. -- **Policies → planner**: predicates intersected with every query/subscription at registration time (already covered in [Phase 4](./04-client-api.md)). - -## How Subscriptions Wire Themselves - -The key low-code payoff: a UI developer never writes subscription code. They write `live: true` on a list or detail, and: - -1. The UI compiler inspects the view's `source` (table, document query, or graph `MATCH`). -2. It generates a `LIVE` query that returns exactly the fields the UI displays. -3. It emits a subscription handle in the rendered page. -4. The client runtime opens a WebSocket, registers the subscription, and binds incoming deltas to DOM fragments by key. -5. When a row changes in the engine, the delta flows: engine → subscription registry → client runtime → DOM patch. - -Zero hand-written subscription code. Zero polling. Adding a new live column to a list is one line in a `.wo` file. - -## Generated Application Stack - -For the writeonce schema above, `sa build` produces: - -| Layer | Generated From | Output | -| --- | --- | --- | -| Database schema | `app/database/*.wo` | In-memory engine arenas + migrations | -| REST / GraphQL / native endpoints | `app/api/*.wo` + `app/database/*.wo` | HTTP handlers, OpenAPI spec, GraphQL SDL | -| SSR HTML | `app/ui/*.wo` + routes in `app.wo` | Per-route renderers compiled into the server binary | -| Client runtime | `app/ui/*.wo` | Small JS bundle: subscription client + DOM patcher + form binding | -| Admin UI | All of the above | Auto-generated CRUD screens for every `##sql`/`##doc` entity (override any with a `##ui` block) | -| Typed SDKs | `app/database/*.wo` | Go/TypeScript/Rust clients per Phase 5 | -| Migrations | Schema diff vs. current database | Versioned forward/backward migrations in `migrations/` | -| Observability | Everything | Structured logs, query metrics, subscription lag dashboards | - -Equivalent hand-written stack: schema (SQL), ORM models, REST controllers, GraphQL schema + resolvers, HTML templates, client JS, subscription plumbing, migrations, admin CRUD, SDKs. Likely **10,000–50,000 lines** for a small e-commerce site. `.wo` target: **~500 lines** across the `app/` tree. - -## Developer Experience — The CLI - -```bash -sa init myapp # scaffold with sensible defaults -cd myapp -sa dev # live-reload server — edit .wo, see changes instantly -sa build --target linux-amd64 # single static binary with everything inside -sa migrate --plan # preview schema migrations -sa migrate --apply # apply migrations -sa gen sdk --lang go --out ./sdk # emit typed client -sa deploy # upload to a running runtime node -``` - -`sa dev` is the make-or-break command. It must: - -- Detect `.wo` changes via inotify (same mechanism as `wo-watch`) -- Recompile incrementally (~50 ms for a single-file change) -- Hot-swap UI renderers without losing client state -- Run schema migrations in a sandbox, surface conflicts before applying -- Keep open subscriptions alive across reloads (re-register on connect) - -## Comparison With Other Declarative Full-Stack Systems - -| System | Schema | UI | Logic | Live queries | Single binary | -| --- | --- | --- | --- | --- | --- | -| **SAP CDS** | `.cds` entities | `@UI` annotations → Fiori | Actions, functions | No (request-response) | No (Java/Node runtime) | -| **Hasura** | Reads from Postgres | No (external UI) | Actions, event triggers | Yes (polling-based internally) | No | -| **Supabase** | Postgres schema | Auto-admin UI only | Edge functions, triggers | Yes (logical replication) | No | -| **Retool / AppSmith / Budibase** | External DB | Visual drag-and-drop | JS snippets | Partial | No | -| **Wasp** (`wasp-lang.org`) | `.wasp` + Prisma | React components | JS functions | No | No (Node + React) | -| **RedwoodJS** | `.sdl` + Prisma | React | JS | No | No | -| **Django + Admin** | Python models | Auto-admin only | Python views | No | No | -| **Phoenix LiveView** | Ecto schemas | HEEx templates | Elixir functions | **Yes** (native) | No (BEAM runtime) | -| **Anvil** (`anvil.works`) | Proprietary | Python drag-and-drop | Python | No | No | -| **`.wo`** (this design) | `type` DSL (unified) over `##sql`+`##doc`+`##graph` substrate | `##ui` declarations + type-attached | type-attached `on ` + `##logic` | **Yes** (native) | **Yes** | - -The differentiators: **cross-paradigm schema** (no competitor unifies SQL + document + graph in one DDL), **engine-native subscriptions** (most bolt on a replication layer or poll), and **single static binary** as the deployment unit (no separate database process, no separate UI server, no separate message bus). - -Closest philosophical precedents: - -- **SAP CDS** for the declaration-first application language — the explicit inspiration. -- **Phoenix LiveView** for the subscription-wired UI model. -- **Wasp** for the `.wasp` → full stack compilation pipeline. -- **Django Admin** for the "generate CRUD from the model" reflex. - -## Scope Addition - -This is a compiler + a renderer + a UI toolkit on top of Phases 2–4. Rough new components: - -| Component | Work | -| --- | --- | -| `type` DSL parser + schema compiler | Entity declarations → `##sql`/`##doc`/`##graph` physical schema; type-attached `policy`/`on`/`service` → same runtime components as standalone blocks | -| `##ui` grammar + analyzer | UI widget tree, field bindings, renderer dispatch | -| Trigger compiler | Type-attached `on ` + standalone `##logic`; both run inside txn coordinator | -| Policy compiler + planner integration | Type-attached `policy` + standalone `##policy` both intersected with every query at registration (some of this exists in Phase 2 already) | -| Service compiler + dispatch table | Type-attached `service` + standalone `##service`; endpoint registration at startup | -| UI render tree → SSR HTML | Template engine, layout system, component library (table, form, chart, etc.) | -| Client runtime (~50 KB JS) | Subscription client, DOM patcher, form binding, validation | -| Auto-admin UI | Generic CRUD screens per entity, override-able with `##ui` | -| Migration engine | Schema diffing, forward/backward migrations, online reshape for the in-memory engine | -| CLI (`sa` binary) | `init`, `dev`, `build`, `migrate`, `gen`, `deploy` | -| Dev-mode live reload | inotify + incremental compiler + client hot-swap | -| Hosted runtime | Optional — for `sa deploy` to work without self-hosting | - -Rough effort on top of Phases 2–4: **12–24 months** with a small team, most of it in the UI compiler and client runtime — that's where the complexity lives, not in the language spec. - -## Honest Framing - -This section is the endgame, not the next step. The sensible build order: - -1. Ship the engine ([Phase 2](./02-wo-language.md), in-memory + io_uring durability via [Phase 3](./03-inmemory-engine.md)) — query layer (`##sql`/`##doc`/`##graph`) only. -2. Ship the wire protocol + Go SDK ([Phase 4](./04-client-api.md) + Phase 5). -3. Ship the **schema-layer `type` DSL** that compiles to the three paradigm blocks. From this point forward, authoring happens against types; the paradigm blocks become an artifact the compiler emits. -4. Add type-attached `service` (and standalone `##service` for bundles) — declarative endpoints. -5. Add type-attached `policy` (and standalone `##policy` for cross-entity rules) — declarative authorization. -6. Add type-attached `on ` triggers (and standalone `##logic` for cross-entity workflows) — declarative triggers. -7. Add `##ui` — declarative rendering. **This is where `.wo` becomes a low-code platform.** -8. Add the CLI, live-reload, admin UI. - -Each step is shippable on its own. Every step after (3) converts imperative code developers are writing by hand into declarative code they write once. The value compounds: by step (7), a small e-commerce app is a ~500-line `.wo` tree instead of a ~50 KLOC TypeScript/Go/SQL repository. - -The risk is the same as every low-code platform: the 20% of use cases outside the declarative model have to have an escape hatch. `.wo` reserves one: any `##ui` node can point to a custom server-rendered template, and any `##logic` block can call out to a host-language plugin (Go/Rust/Wasm). Without that escape hatch, low-code becomes no-code in the pejorative sense — you can build 80% of the app and the rest is impossible. - -## Reference Implementations Worth Studying - -- **SAP CDS** — . Read the CDS Language Reference cover to cover before designing `##ui`. Thirty years of ERP app-generation is condensed into that spec. -- **Wasp** — . Open-source `.wasp` → React + Node + Prisma compiler. Closest living sibling to `.wo`. -- **Phoenix LiveView** — . The rendering-subscription loop done right in Elixir. -- **HTMX + Hyperscript** — . Tiny client runtime that consumes server-rendered HTML fragments on events. Good model for `.wo`'s client bundle. -- **Retool / Budibase / Appsmith (open source)** — , . Visual low-code; useful to see which UI primitives users actually ask for. -- **SurrealDB `define` syntax** — SurrealQL includes declarative `DEFINE TABLE`, `DEFINE FIELD`, `DEFINE EVENT` that are partway toward an application DSL. Worth studying for how much declaration fits inside a query language. -- **Django admin source** — `django/contrib/admin/`. The canonical "CRUD from models" implementation; read how it introspects schema to generate list/detail/edit views. -- **PocketBase** — . Single-binary Go app with SQLite, auto-admin UI, realtime subscriptions. Proves the single-binary low-code model is buildable. `.wo` is what PocketBase would look like if its data model were multi-paradigm and its query engine were custom. - -## Why This Belongs in This Series - -The question that opened [Phase 1](./01-evaluation.md) — "should writeonce use a document or graph database?" — has now inverted. The answer drove past "no, use flat files" through "build your own query language" and "build your own ACID engine" and "build your own wire protocol" to arrive here: **a full-stack declarative application platform where the database, the subscriptions, the UI, and the business logic are one artifact compiled from one language**. - -That is the actual ambition. Every earlier phase is a subcomponent of this one. Decide honestly whether the project is a blog engine that needed a graph index, or an application platform that happens to start as a blog engine. The answer determines which phases are scope and which are cautionary. diff --git a/docs/runtime/database/07-wo-seg-migration.md b/docs/runtime/database/07-wo-seg-migration.md deleted file mode 100644 index 74d6bab..0000000 --- a/docs/runtime/database/07-wo-seg-migration.md +++ /dev/null @@ -1,207 +0,0 @@ -# Phase 7 — Replacing `wo-seg` with the writeonce Database - -> A phased coexistence plan: abstract the article store behind a trait, stand up the `.wo` engine as a second implementation, dual-run, cut over, decommission. - -**Previous**: [Phase 6 — Low-Code Full-Stack](./06-lowcode-fullstack.md) | **Index**: [database.md](../database.md) - ---- - -## Context - -Today's writeonce runtime stores articles in a hand-rolled append-only file format: - -- **`crates/wo-seg`** (~475 LOC) — `.seg` binary file: magic + header + `[u32 length][u8 flags][payload]` records serialized with `bincode`. `SegWriter::append()` returns a byte offset usable as an index pointer. Tombstoning flips a flag byte. No transactions, no MVCC, no concurrent writers. -- **`crates/wo-index`** — sidecar `title.idx`, `date.idx`, `tags.idx` files built from the `.seg`. Queries hit the index to resolve to a byte offset, then the `.seg` to load the record. -- **`crates/wo-store`** — composes the two above, owns cold-start (rebuild from `content/`), and exposes the query API the rest of the system calls: `get_by_title`, `list_published`, `list_by_tag`, `list_by_date_range`, `count_published`, `ingest_article`, `article_version`. - -This is the Phase 1 answer ("no external DB — add a petgraph-backed `mappings.idx`") made flesh. It is correct for a single-writer blog and wrong for everything Phases 2–6 want to deliver: no ACID across multiple shapes, no live subscriptions, no cross-paradigm queries, no declarative schema, no codegen. - -The six-phase `.wo` design is the replacement. This doc plans the migration — how to swap wo-seg for the `.wo` engine **without halting writeonce** while the engine is built over multiple quarters. - -## Intended Outcome - -- `crates/wo-seg` is deleted. -- `crates/wo-store` either (a) becomes a thin facade over the `.wo` engine or (b) disappears, with callers depending directly on the engine's Rust SDK. -- Writeonce's serving path is unchanged from the user's perspective throughout the migration. -- The `.wo` engine reaches production-ready status incrementally; each milestone is independently shippable. - -## Strategy — Phased Coexistence - -Do **not** big-bang. The seg-based store works; replacing it takes many months. Instead: - -1. **Abstract** the existing store behind a Rust trait — one weekend of mechanical refactor, zero behavior change. -2. **Build** the `.wo` engine crates next to seg, not in its place. Port the C++ prototype (`prototypes/wo-db/`) to Rust so the engine lives in the same Cargo workspace as the blog. -3. **Dual-run** — writes go to both backends, reads to seg. Compare results in CI and on production data. This surfaces engine bugs without user impact. -4. **Cut over** reads once the engine passes dual-run. Writes still hit seg as a cold standby. -5. **Decommission** seg when enough time has passed without rollback and a restore-from-seg fallback is no longer load-bearing. - -``` -wo-seg + wo-store ─────────► trait-abstracted ─────► dual-write, read seg ─────► read wo-db, write both ─────► wo-db only, delete wo-seg - (today) (phase A) (phase B) (phase C) (phase D) -``` - -Each transition is reversible — flip one feature flag or swap one trait object back. - -## Proposed Crate Layout - -Port the C++ prototype (`prototypes/wo-db/src/*`) to Rust, split along the natural seams. New-runtime crates are unprefixed; v1 crates keep `wo-` in `reference/crates/`. - -| Crate | Purpose | Prototype source | Phase | -| --- | --- | --- | --- | -| `ql` | `.wo` grammar: lexer, parser, AST | `src/lexer.*`, `src/parser.*`, `src/ast.hpp` | 2 | -| `value` | tagged `Value` + path utilities | `src/value.hpp`, path helpers in `src/storage.cpp` | 2 | -| `engine` | in-memory executor (sql / doc / graph), schema catalog | `src/storage.*`, `src/executor.*` | 2 | -| `txn` | MVCC, snapshot isolation, `RETURNING` alias table | new (Phase 2 milestone 3) | 2 | -| `wal` | write-ahead log + fsync + crash recovery | new ([Phase 3](./03-inmemory-engine.md)) | 3 | -| `sub` | live subscriptions — delta frames on commit | new ([Phase 4](./04-client-api.md)) | 4 | -| `http` | wire protocol — REST / GraphQL-over-WS / native codec | new ([Phase 4](./04-client-api.md)) | 4 | -| `db` | top-level facade: `open()`, `Tx`, `Query`, `Subscribe` — the Rust SDK | integrates the above | 2–4 | -| `gen` | codegen: `.wo type` → Rust structs, Go structs, TypeScript | `sa-gen`/`wo-gen` in Phase 5 | 5 | - -All 15 crates (these 14 plus the existing `rt` binary crate) now exist as empty skeletons in `crates/`. See [`crates/README.md`](../../../crates/README.md) and [`docs/plan/done/01-scafolding-crates.md`](../../plan/done/01-scafolding-crates.md) for the scaffolding plan that landed them. - -Today's `reference/crates/wo-seg` and `reference/crates/wo-index` remain in the v1 nested workspace for the entire migration window. They disappear only at the end of Phase D. - -`reference/crates/wo-store` evolves but survives — it becomes the writeonce-specific glue layer (trait, article domain model, content-directory cold-start) whose backend is swappable. - -## Phase A — Abstract the Article Store - -**Goal**: every caller depends on a trait, not on `wo_store::Store` directly. Zero behavior change. - -**Work**: - -- Define `trait ArticleStore` in `wo-store/src/lib.rs` with the existing public API: - ```rust - pub trait ArticleStore: Send + Sync { - fn get_by_title(&self, sys_title: &str) -> io::Result>; - fn list_published(&self, skip: usize, limit: usize) -> io::Result>; - fn list_by_tag(&self, tag: &str) -> io::Result>; - fn list_by_date_range(&self, start: i64, end: i64) -> io::Result>; - fn count_published(&self) -> io::Result; - fn ingest_article(&mut self, json_path: &Path) -> io::Result; - fn article_version(&self, sys_title: &str) -> Option; - fn content_dir(&self) -> &Path; - } - ``` -- Rename the existing `Store` struct to `SegStore` and implement `ArticleStore` for it. Re-export `SegStore as Store` for one release to avoid churn at call sites. -- Change `wo-route`, `wo-serve`, `wo-sub`, `wo-htmlx` to take `&dyn ArticleStore` (or generic ``). The trait import stays in `wo-store`; concrete impls move to sibling crates. -- Add a tiny `wo-store::open(content_dir, data_dir) -> Arc` factory that picks the backend based on a config env var (`WO_STORE_BACKEND=seg|db|dual`). - -**Exit criteria**: `cargo test` passes; `wo serve` boots unchanged; git log shows one PR. - -## Phase B — Stand Up `wo-db` in Rust - -**Goal**: a Rust `wo-db` crate that speaks the full `.wo` grammar from the prototype, stored in memory, with an `ArticleStore` impl mapping writeonce's `Article` onto the relational paradigm. - -**Work**: - -- Port `prototypes/wo-db/` (C++) to Rust crates per the layout table above. The `wo` namespace becomes the `wo_*` crate family; the test suites (`tests/smoke.wo`, `tests/checkout.wo`) run as Rust integration tests. -- Define a `.wo` schema for the writeonce domain (in a new file, `crates/wo-store/schema.wo`): - ```wo - type Article { - id: Id - sys_title: Slug @unique - title: Text - published: Bool = false - published_at: Timestamp? - author: Text - tags: [Text] - meta: { excerpt: Text, body_md: Markdown } - } - ``` -- Add `DbStore` — a second `ArticleStore` impl that translates calls into `.wo` queries: - - `get_by_title(t)` → `SELECT * FROM Article WHERE sys_title = $t` (one row) - - `list_published(skip, limit)` → `SELECT * FROM Article WHERE published = true ORDER BY published_at DESC LIMIT $limit OFFSET $skip` - - `list_by_tag(t)` → `SELECT * FROM Article WHERE $t IN tags` - - `list_by_date_range(a, b)` → `SELECT * FROM Article WHERE published_at BETWEEN $a AND $b` - - `count_published` → `SELECT COUNT(*) FROM Article WHERE published = true` - - `ingest_article(path)` → load JSON → `INSERT INTO Article (…)` -- Cold-start path: when the data dir is empty, `DbStore::open` loads all `content/*.json` the same way `SegStore::open` does today and inserts into the engine. -- Gate behind `#[cfg(feature = "db-backend")]` so seg-only builds keep working until Phase C. - -**Exit criteria**: `DbStore` passes the same unit tests as `SegStore` (rename `Store` → `ArticleStore` in test assertions). Memory footprint and per-query latency measured against seg; both within an order of magnitude. - -## Phase C — Dual-Write, Read Seg - -**Goal**: every mutation hits both backends; reads stay on seg; a differ flags mismatches. - -**Work**: - -- Add `DualStore` — a third `ArticleStore` impl that forwards writes to both `SegStore` and `DbStore` and returns `SegStore` results for reads. -- Add a background task (`wo-store::differ`) that on every ingest runs every query method against both backends and compares results. Mismatches → structured log entry (`store_mismatch` event) + a Prometheus counter. -- Set `WO_STORE_BACKEND=dual` on staging for two weeks, then on prod behind a rollout flag. - -**Exit criteria**: zero `store_mismatch` events for 14 consecutive days on production traffic. - -## Phase D — Cut Over Reads, Keep Seg as Fallback - -**Goal**: reads served from `DbStore`; seg still receives writes and is kept queryable as a cold standby. - -**Work**: - -- Invert `DualStore`: writes to both, reads from `DbStore`. -- Add an admin command `wo db verify --against seg` that re-runs the differ on demand (for post-incident checks). -- After a stable month, remove `DualStore` entirely. `WO_STORE_BACKEND=db` becomes the only supported value. - -**Exit criteria**: one month with no read-path regressions; no active rollback capability needed for routine ops. - -## Phase E — Decommission `wo-seg` - -**Goal**: delete `crates/wo-seg`, shrink `crates/wo-store` to the trait + content-directory cold-start. - -**Work**: - -- Delete `crates/wo-seg`. Remove `wo-seg` from `Cargo.toml` workspace members and from `wo-store/Cargo.toml` deps. -- Delete `SegStore` from `wo-store`. The trait `ArticleStore` and `DbStore` remain. -- Delete the `.seg` file from production data directories (via a migration: verify `DbStore` has every record, then `rm`). -- Delete `crates/wo-index` **if and only if** `DbStore` has replaced its indexes with the engine's internal ones. If the LSM/graph indexes inside `wo-db` cover the three sidecar indexes (title, date, tags) — expected — then wo-index goes too. If any index is still load-bearing outside the engine, keep it. - -**Exit criteria**: CI is green with the deletions; production runs a release cycle without rollback; `rg "wo-seg\|wo_seg"` returns zero hits. - -## Integration Touchpoints - -These crates reference the store today and will need light updates for Phase A (trait swap): - -| Crate | Current coupling | Change | -| --- | --- | --- | -| `wo-store` | owns `Store`, depends on `wo-seg` + `wo-index` | gains trait + factory + dual-write impl (A–C); shrinks to facade in E | -| `wo-route` | likely consumes `&Store` | accept `&dyn ArticleStore` | -| `wo-serve` | HTTP handlers read the store | accept `Arc` | -| `wo-sub` | subscription layer | later — see below | -| `wo-htmlx` | may read article state during render | accept trait or projection | -| `wo-watch` | inotify-driven ingest | unchanged; still calls `ingest_article` | -| `wo-rt` | runtime glue | pass the trait object through | - -`wo-sub` is a special case. Today it likely polls or reacts to `article_version` monotonic counters. When Phase 4 activates `LIVE` queries inside `wo-db`, `wo-sub` should stop doing its own diffing and become a pass-through for engine-emitted deltas. That transition happens in Phase C/D, not Phase A — it's not required for the trait refactor. - -## Risks - -1. **Cold-start cost.** `SegStore` builds its indexes in one pass over `.seg`. `DbStore` has to parse JSON from `content/` the same way but also commit through the engine's WAL. If this is slow, add a `wo db import --from-seg ` shortcut that bulk-loads from an existing `.seg` without going through the ingest path. -2. **Memory footprint.** Today's seg-based path `mmap`s the file; the `.wo` engine is RAM-primary. For a blog with hundreds of articles, immaterial; for a larger dataset, Phase 3's SSD-backed variant is what's needed. -3. **Article → `.wo` type drift.** `wo_model::Article` is the canonical domain type today. The `.wo` schema mirrors it, but if the two diverge (a new field is added to `Article` but not to the schema), queries silently drop that field. Mitigation: `wo-gen` should include a `--verify wo_model::Article` mode in Phase 5 that fails CI on drift. -4. **Dual-write contention.** If ingest becomes the bottleneck during Phase C, time-box dual-write: drop it after 14 clean days rather than running it indefinitely. -5. **Feature flag sprawl.** `WO_STORE_BACKEND` should be the only config knob. Resist per-method flags. - -## What's Out of Scope for This Doc - -- The engine internals themselves — those live in Phases 2–4. -- The Phase 5 SDK (`wo-gen`, typed Go client) — wo-store callers are Rust, and Rust codegen is part of `wo-gen` but not a blocker. -- `##ui` / `##policy` / `##logic` / `##service` — those are Phase 6 and assume the engine is already running. -- Any graph-first features (mappings, `RELATED_TO` traversal) — they become trivially available once `DbStore` is live, but don't need to gate the seg → db cutover. - -## Verification - -Each phase has its own exit criteria above. End-to-end verification for the whole migration: - -1. **Parity** — After Phase B: a shadow script replays one week of production ingest through `DbStore` in a sandbox; every query from the shadow matches seg. -2. **Latency** — After Phase D: p50/p95/p99 of `get_by_title`, `list_published`, `list_by_tag` are at or below the seg baseline. Measured by the existing request-timing middleware, not synthetic benchmarks. -3. **Crash safety** — After Phase 3 WAL ships: `kill -9` during write, reopen, confirm the committed state matches and uncommitted writes are gone. Automated test. -4. **Decommission audit** — After Phase E: `rg 'wo-seg|wo_seg|\.seg\b' crates/` returns zero; `cargo deny check` passes; production restart ingests from `content/` with no `.seg` file present. - -## Related Documents - -- [02-wo-language.md](./02-wo-language.md) — the two-layer `.wo` language the engine speaks -- [03-inmemory-engine.md](./03-inmemory-engine.md) — the storage engine behind `wo-db` -- [04-client-api.md](./04-client-api.md) — wire protocol and `LIVE` subscriptions -- [01-evaluation.md](./01-evaluation.md) — why writeonce built `wo-seg` in the first place, and why that choice still looks right for the blog even as the platform grows past it -- `prototypes/wo-db/` — the C++ prototype of the `.wo` engine, the reference implementation the Rust port follows diff --git a/docs/runtime/fibers.md b/docs/runtime/fibers.md deleted file mode 100644 index 9bf3eca..0000000 --- a/docs/runtime/fibers.md +++ /dev/null @@ -1,704 +0,0 @@ -# Runtime Fibers - -Runtime fibers are lightweight, user-space threads managed by an application's runtime system rather than the OS kernel. They enable massive concurrency — millions per machine — because they don't carry the overhead of kernel thread stacks and scheduling. Unlike pre-emptive kernel threads, fibers use cooperative multitasking: they yield control voluntarily at known suspension points. - -## Context Switching - -A context switch is saving the state of one execution unit and restoring another so it can continue running. The cost of this switch is what separates kernel threads from fibers. - -### Kernel Thread Context Switch - -When the OS switches between threads: - -1. Save all CPU registers (general purpose, floating point, SIMD) to kernel memory -2. Save the thread's stack pointer -3. Flush the TLB (translation lookaside buffer) if switching processes -4. Update scheduler data structures -5. Restore the next thread's registers and stack pointer -6. Return to userspace - -Cost: **1-10 microseconds**, involves a kernel trap (syscall boundary crossing), cache pollution from TLB flush. - -### Fiber Context Switch - -When a runtime switches between fibers: - -1. Save a few registers (stack pointer, instruction pointer, callee-saved registers) -2. Swap the stack pointer to the next fiber's stack -3. Jump to the next fiber's saved instruction pointer - -Cost: **~10-100 nanoseconds**, entirely in userspace, no kernel involvement, no TLB flush, cache stays warm. - -``` -Kernel thread switch: ~1,000-10,000 ns (kernel trap + TLB flush) -Fiber switch: ~10-100 ns (register swap in userspace) - 100x cheaper -``` - -## Types of Multitasking - -### Pre-emptive (Kernel Threads) - -The OS scheduler interrupts threads at arbitrary points using timer interrupts. The thread does not choose when to yield — the kernel forces it. - -``` -Thread A: ████████──┐ (interrupted by OS) - │ -Thread B: └──████████──┐ (interrupted by OS) - │ -Thread A: └──████████ -``` - -- Threads can be interrupted mid-instruction -- Requires locks/mutexes to protect shared state -- Fairness guaranteed by the scheduler -- Used by: pthreads, std::thread, OS processes - -### Cooperative (Fibers / Green Threads) - -Fibers explicitly yield at known points (I/O boundaries, channel sends, `.await` in Rust). The runtime only switches when the fiber says "I'm done for now." - -``` -Fiber A: ████████ yield ──┐ - │ -Fiber B: └── ████████ yield ──┐ - │ -Fiber A: └── ████████ -``` - -- Fibers are never interrupted mid-computation -- No locks needed for single-threaded runtimes — only one fiber runs at a time -- Starvation possible if a fiber never yields (compute-heavy work blocks the loop) -- Used by: Go goroutines, Erlang processes, Lua coroutines, Rust async/await - -### Comparison - -| Property | Pre-emptive (Threads) | Cooperative (Fibers) | -|----------|----------------------|---------------------| -| Scheduling | OS kernel decides | Runtime decides at yield points | -| Context switch cost | ~1-10 us | ~10-100 ns | -| Stack size | 1-8 MB per thread (fixed) | Bytes to KB per fiber (growable) | -| Max concurrency | ~10,000 threads | ~1,000,000+ fibers | -| Synchronization | Locks, mutexes, atomics | Not needed in single-threaded runtime | -| Interruption | Any point (timer interrupt) | Only at yield points | -| Fairness | Guaranteed by scheduler | Must be designed (fiber must yield) | - -## How Fibers Work Internally - -A fiber needs three things: - -1. **A stack** — a block of memory for local variables and call frames -2. **A saved context** — the register state at the point it yielded -3. **A function** — the code to run when resumed - -### Minimal Fiber in C - -```c -#include -#include - -static ucontext_t main_ctx, fiber_ctx; -static char fiber_stack[8192]; - -void fiber_fn() { - printf("Fiber: running\n"); - // Yield back to main - swapcontext(&fiber_ctx, &main_ctx); - printf("Fiber: resumed\n"); - // Yield again - swapcontext(&fiber_ctx, &main_ctx); -} - -int main() { - // Set up fiber context - getcontext(&fiber_ctx); - fiber_ctx.uc_stack.ss_sp = fiber_stack; - fiber_ctx.uc_stack.ss_size = sizeof(fiber_stack); - fiber_ctx.uc_link = &main_ctx; - makecontext(&fiber_ctx, fiber_fn, 0); - - printf("Main: starting fiber\n"); - swapcontext(&main_ctx, &fiber_ctx); // switch to fiber - - printf("Main: fiber yielded\n"); - swapcontext(&main_ctx, &fiber_ctx); // resume fiber - - printf("Main: fiber yielded again\n"); - swapcontext(&main_ctx, &fiber_ctx); // resume — fiber finishes - - printf("Main: done\n"); - return 0; -} -``` - -Output: -``` -Main: starting fiber -Fiber: running -Main: fiber yielded -Fiber: resumed -Main: fiber yielded again -Main: done -``` - -`swapcontext` saves the current register state to one context struct and loads another — that's the entire fiber switch. No kernel involved. - -### Minimal Fiber in Rust (unsafe) - -Rust's async/await compiles to state machines, not stack-swapping fibers. But you can build raw fibers with inline assembly: - -```rust -use std::arch::asm; - -struct Fiber { - stack: Vec, - sp: *mut u8, // saved stack pointer -} - -impl Fiber { - fn new(func: fn()) -> Self { - let mut stack = vec![0u8; 8192]; - let sp = unsafe { - let top = stack.as_mut_ptr().add(stack.len()); - let aligned = (top as usize & !0xF) as *mut u8; - // Push the function pointer as the return address - let sp = aligned.sub(8); - *(sp as *mut fn()) = func; - sp - }; - Fiber { stack, sp } - } - - unsafe fn switch_to(&mut self, from: &mut *mut u8) { - // Save callee-saved registers and swap stack pointers - asm!( - "push rbx", - "push rbp", - "push r12", - "push r13", - "push r14", - "push r15", - "mov [{from}], rsp", // save current sp - "mov rsp, [{to}]", // load fiber sp - "pop r15", - "pop r14", - "pop r13", - "pop r12", - "pop rbp", - "pop rbx", - from = in(reg) from, - to = in(reg) &self.sp, - ); - } -} -``` - -This is what runtimes like Go and Erlang do internally — allocate a small stack, save/restore a handful of registers, and jump. The cost is a few nanoseconds. - -## Fibers vs Rust async/await - -Rust chose a different approach than fibers for its async model: - -| Property | Fibers (Go, Erlang) | Rust async/await | -|----------|-------------------|-----------------| -| Implementation | Stack swapping at runtime | Compiler generates state machines | -| Stack | Each fiber has its own stack | No extra stack — state stored in Future struct | -| Memory per task | ~2-8 KB minimum (stack) | Bytes — only the live variables at yield points | -| Yield mechanism | `swapcontext` / assembly | `.await` compiles to `Poll::Pending` | -| Overhead | Stack allocation + register swap | Zero-cost — state machine is a regular struct | -| Debuggability | Separate stacks in debugger | State machine is harder to trace | -| Preemption | Runtime can preempt (Go does this) | Never preempted — cooperative only | - -Rust's approach is called **stackless coroutines** — no extra stack per task. The compiler transforms each `async fn` into a state machine enum where each variant holds the local variables alive across an `.await` point. - -## Relation to writeonce - -The writeonce runtime (`wo-event`, `wo-rt`) uses neither fibers nor Rust async/await. It uses a **single-threaded event loop with callbacks** — the simplest model: - -``` -loop { - events = epoll_wait() - for event in events { - match event.token { - WATCHER => handle_file_change(), - LISTENER => handle_accept(), - HTTP_CONN => handle_request(), - ... - } - } -} -``` - -This is the same model as nginx, Redis, and Node.js (before libuv's thread pool). It works because: - -- Each handler runs to completion quickly (microseconds) -- No handler blocks — all I/O is non-blocking -- No concurrent access to shared state — one thing runs at a time -- Blocking work (index rebuild) offloads to a thread pool and signals back via eventfd - -If writeonce ever needed millions of concurrent long-lived tasks (not just connections), fibers would be the next step. But for a content platform serving articles, the event loop is sufficient — and far simpler to reason about. - -## Summary - -``` -Kernel threads: OS-managed, pre-emptive, expensive switch, ~10K max -Fibers: Runtime-managed, cooperative, cheap switch, ~1M+ max -Async/await: Compiler-managed, cooperative, zero-cost, ~1M+ max -Event loop: No tasks at all — just fd readiness + callbacks - -Complexity: event loop < fibers < async/await < threads -Concurrency: event loop = fibers = async/await >> threads -``` - -The right choice depends on the workload. For writeonce — an event loop. For a database with millions of queries in flight — fibers or async. For CPU-bound parallel work — kernel threads. - -What are fibers? # -Fibers are a lightweight thread of execution similar to OS threads. However, unlike OS threads, they’re cooperatively scheduled as opposed to preemptively scheduled. What this means in plain English is that fibers yield themselves to allow another fiber to run. You may have used something similar to this in your programming language of choice where it’s typically called a coroutine, there’s no real distinction between coroutines and fibers other than that coroutines are usually a language-level construct, while fibers tend to be a systems-level concept. - -Other names for fibers you may have heard before include: - -green threads -user-space threads -coroutines -tasklets -microthreads -There are very few and minor differences between fibers and the above list. For the purposes of this document, we should consider them equivalent as the distinctions don’t quite matter. - -Scheduling # -At any given moment the OS is running multiple processes all with their own OS threads. All of those OS threads need to be making forward progress. There’s two classes of thought when it comes to how you solve this problem. - -Cooperative scheduling -Preemptive scheduling -It’s important to note that while you may observe that all processes and OS threads are running in parallel, scheduling is really providing the illusion of that. Not all threads are running in parallel, the scheduler is just switching between them quickly enough that it appears everything is running in parallel. That is they’re concurrent. Threads start, run, and complete in an interleaved fashion. - -It is possible for multiple OS threads to be running in parallel with symmetric multiprocessing (SMP) where they’re mapped to multiple hardware threads, but only as many hardware threads as the CPU physically has. - -Premptive scheduling # -Most people familiar with threads know that you don’t have to yield to other threads to allow them to run. This is because most operating systems (OS) schedule threads preemptively. - -The points at which the OS may decide to preempt a thread include: - -IO -sleeps -waits (seen in locking primitives) -interrupts (hardware events mostly) -The first three in particular are often expressed by an application as a system call. These system calls cause the CPU to cease executing the current code and execute the OS’s code registered for that system call. This allows the OS to service the request then resume execution of your application’s calling thread, or another thread entierly. - -This is possible because the OS will decide at one of the points listed above to save all the relevant state of that thread then resume some other thread, the idea being that when this thread can run again, the OS can reinstate that thread and continue executing it like nothing ever happened. These transition points where the OS switches a thread are called context switches. - -There’s a cost associated with this context switching and all modern operating systems have made great deals of effort to reduce this cost as much as possible. Unfortunately, that overhead begins to show itself when you have a lot of threads. In addition, recent cache side channel attacks like: Spectre, Meltdown, Spoiler, Foreshadow, and Microarchitectural Data Sampling on modern processors has led to a series of both user-space and kernel-space mitigation strategies, some of which increased the overhead of context switches significantly. - -You can read more about context switching overhead in this paper. - -Cooperative scheduling # -This idea of fibers yielding to each other is what is known as cooperative scheduling. Fibers effectively move the idea of context switching from kernel-space to user-space and then make those switches a fundamental part of computation, that is, they’re a deliberate and explicitly done thing, by the fibers themselves. The benefit of this is that a lot of the previously mentioned overhead can be entierly eliminated while still permitting an excess count of threads of execution, just in the form of these fibers now. - -The problem with multi-threading # -There’s many problems related to multi-threading, most obviously that it’s difficult to get right. Most proponents of fibers make false claims about how this problem goes away when you use fibers because you don’t have parallel threads of execution. Instead, you have these cooperatively scheduled fibers which yield to each other. This means it’s not possible to race data, dead lock, live lock, etc. While this statement is true when you look at fibers as a N:1 proposition, the story is entierly different when you introduce M:N. - -N:1 and what it means # -Most documentation, libraries, and tutorials on fibers are almost exclusively based around using a single thread given to you by the OS, then sharing it among multiple fibers that cooperatively yield and run all your asynchronous code. This is called N:1 (“N to one”). N fibers to 1 thread, and it’s the most prevalent form of fibers. This is how Lua coroutines work, how Javascript’s and Python’s async/await work, and it’s not what you’re interested in doing if you actually want to take advantage of hardware threads. What you’re interested in is M:N, (“M to N”) M fibers to N threads. - -M:N and what it means # -The idea behind M:N is to take the model given to us by N:1 and map it to multiple actual OS threads. Just like we’re familiar to the concept of thread pools where we execute tasks, here we have a pool of threads where we execute fibers and those fibers get to yield more of themselves on that thread. - -I should stress that M:N fibers have all the usual problems of multi-threading. You still need to syncronize access to resources shared between multiple fibers because there’s still multiple threads. - -The problem with thread pools # -A lot of you may be wondering how this is different from traditional task based parallelism choices seen in many game engines and applications. The model where you have a fixed-size pool of threads you queue tasks on to be executed at some point in the future by one of those threads. - -Locality of reference # -The first problem is locality of reference. The data-oriented / cache-aware programmers reading this will have to mind my overloading of that phrase because what I’m really talking about is the resources that a job needs access to are usually local. The job isn’t going to be executed immediately, but rather when the thread pool has a chance to. Any resource that job needs access to, needs to be available for the job at some point in the future. This means local values need their lifetime’s extended for an undefined amount of time. - -There’s many ways to solve this lifetime problem, except they all have overhead. Consider this trivial example where I want to asynchronously upload a local file to a webserver. - -## Async Runtime for the `.wo` Database - -The writeonce event loop (`wo-event` / `wo-rt`, described in [async.md](./async.md)) is a **single-threaded epoll loop with callbacks**. It works because the content workload is microseconds per request and single-writer. The `.wo` database (described in the [database series](./database/02-wo-language.md)) inverts every one of those assumptions: - -| writeonce (blog) | `.wo` database | -| --- | --- | -| Single-threaded — one handler at a time | Thousands of queries in flight concurrently | -| epoll — kernel wakes you on fd readiness | io_uring — userland submits I/O, kernel completes asynchronously | -| Callbacks — run to completion, return to loop | Tasks — query execution spans multiple I/O waits (WAL write, network send) | -| No blocking — handlers are microsecond-fast | Graph traversals and join plans can be milliseconds of CPU | -| Single writer — no contention | MVCC — many readers and writers touching shared data structures | -| Read-heavy | Write-heavy on hot paths (checkout, inventory) | - -A new runtime is needed. This section designs it, building on the fiber/async/event-loop theory above. - -### What the Runtime Must Schedule - -Every subsystem in the `.wo` engine produces work with a different shape: - -| Subsystem | Work shape | Blocking? | Parallelizable? | -| --- | --- | --- | --- | -| **Client accept** | Wait for incoming TCP connection | I/O-bound | No — one listener fd | -| **Request decode** | Parse binary frame or GraphQL | CPU, microseconds | Yes — per connection | -| **Query planning** | Compile `.wo` to execution plan | CPU, microseconds | Yes — per query | -| **Query execution** | Traverse B+ tree / LSM / graph | CPU, microseconds to milliseconds | Yes — per query | -| **WAL append** | Serialize + io_uring write + fsync | I/O-bound (NVMe) | Batched — group commit | -| **MVCC bookkeeping** | Version chain prepend, snapshot management | Atomic CAS, nanoseconds | Yes — per record | -| **Subscription matching** | Evaluate deltas against registered predicates | CPU, microseconds | Yes — per subscription | -| **Client push (DELTA)** | io_uring send to client socket | I/O-bound | Yes — per connection | -| **Checkpoint** | Snapshot RAM arenas to SSD | I/O-bound, background | Single background task | -| **Vacuum** | Prune old MVCC versions | CPU, background | Parallelizable by range | - -The runtime’s job is to keep all of these progressing concurrently — mixing CPU-bound query execution with I/O-bound disk and network operations — without blocking any subsystem on another. - -### Four Candidate Architectures - -#### 1. Thread-Per-Core / Shared-Nothing (Seastar / ScyllaDB) - -Each CPU core runs its own independent event loop. No shared memory between cores. Communication is message-passing over lock-free queues. - -``` -Core 0 Core 1 Core 2 Core 3 -┌──────────┐ ┌──────────┐ ┌──────────┐ ┌──────────┐ -│ io_uring │ │ io_uring │ │ io_uring │ │ io_uring │ -│ ring │ │ ring │ │ ring │ │ ring │ -│ │ │ │ │ │ │ │ -│ local │ ←msg→ │ local │ ←msg→ │ local │ ←msg→ │ local │ -│ shard of │ │ shard of │ │ shard of │ │ shard of │ -│ data │ │ data │ │ data │ │ data │ -└──────────┘ └──────────┘ └──────────┘ └──────────┘ -``` - -**How it works**: data is hash-partitioned across cores. A query for `product.id = 42` routes to the core that owns shard `hash(42) % num_cores`. That core executes the query entirely locally — no locks, no contention. - -**Pros**: zero contention on the hot path; each core runs flat-out; NUMA-friendly; no lock overhead. - -**Cons**: cross-shard queries (joins, graph traversals spanning partitions) require inter-core messaging; range scans hit all shards; transaction coordination across shards is complex; programming model is unfamiliar. - -**Used by**: ScyllaDB/Seastar (C++), Redpanda (C++), Glommio (Rust). - -#### 2. Work-Stealing Async Runtime (tokio / Rust `async/await`) - -M:N scheduling of `async` tasks across a thread pool. Tasks (futures) yield at `.await` points; the scheduler steals work from busy threads. - -``` -┌─────────────────────────────────────────┐ -│ tokio runtime (M:N) │ -│ ┌────────┐ ┌────────┐ ┌────────┐ │ -│ │ worker │ │ worker │ │ worker │ │ -│ │ thread │ │ thread │ │ thread │ │ -│ │ │ │ │ │ │ │ -│ │ task │ │ task │ │ task │ │ -│ │ task │ │ task │←steal──┐ │ │ -│ │ task │ │ │ │ task │ -│ └────────┘ └────────┘ └────────┘ │ -│ shared task queues │ -└─────────────────────────────────────────┘ -``` - -**How it works**: each query becomes an `async fn` that `.await`s I/O (network reads, WAL writes). The executor multiplexes thousands of tasks across a small thread pool. Shared state is accessed through `Arc>` or lock-free structures. - -**Pros**: familiar Rust async model; excellent library ecosystem; handles mixed I/O and CPU work; tokio has 7+ years of production hardening. - -**Cons**: `async/await` infects the entire codebase (everything must be async); shared state needs locks or lock-free structures; work-stealing adds scheduling overhead; harder to reason about NUMA locality. - -**Used by**: SurrealDB (Rust + tokio), TiKV (Rust + tokio), Materialize (Rust + tokio). - -#### 3. M:N Fiber Runtime (Go / Erlang) - -User-space fibers with their own stacks, scheduled cooperatively across OS threads. Each query is a fiber; yield points are implicit at I/O boundaries. - -``` -┌────────────────────────────────────────┐ -│ Go-style runtime │ -│ │ -│ goroutine goroutine goroutine │ -│ goroutine goroutine goroutine │ -│ goroutine goroutine goroutine │ -│ ↓ ↓ ↓ │ -│ ┌────────┐ ┌────────┐ ┌────────┐ │ -│ │ OS thr │ │ OS thr │ │ OS thr │ │ -│ └────────┘ └────────┘ └────────┘ │ -│ work-stealing scheduler │ -└────────────────────────────────────────┘ -``` - -**How it works**: each query spawns a fiber. The fiber runs synchronously from its own perspective — blocking calls (network read, disk write) are transparently converted to yields by the runtime. The scheduler multiplexes fibers across OS threads. - -**Pros**: synchronous programming model (no `async`/`.await` annotations); lightweight (2–8 KB per fiber vs. the state-machine size of a Rust future); Go and Erlang prove this works at massive scale; preemption possible (Go preempts goroutines at function calls since Go 1.14). - -**Cons**: stack allocation per fiber (minor); GC pressure if using Go (major for database latency); less ecosystem in Rust (no production-quality M:N fiber runtime); context switch is ~10–100 ns vs. ~1 ns for a Rust future poll. - -**Used by**: CockroachDB (Go), Dgraph (Go), Erlang/OTP databases, may_minihttp (Rust). - -#### 4. Deterministic io_uring Loop + Thread Pool (TigerBeetle) - -A single main thread drives an io_uring instance for all I/O. CPU-bound work (query execution) is dispatched to a thread pool. Results return to the main loop via the io_uring completion queue. - -``` -┌────────────────────────────────────────────┐ -│ Main thread (deterministic) │ -│ │ -│ io_uring ring: │ -│ - ACCEPT (new connections) │ -│ - RECV (query frames) │ -│ - WRITE (WAL append) │ -│ - FSYNC (WAL durable) │ -│ - SEND (results / deltas) │ -│ │ -│ On CQE(RECV, query_bytes): │ -│ dispatch to thread pool │ -│ │ -│ On CQE(thread_pool_result): │ -│ submit SEND(client_fd, result_bytes) │ -└────────────────────────────────────────────┘ - │ ▲ - ▼ │ -┌────────────────────┐ ┌─────────┐ -│ Worker threads │ │ eventfd │ -│ (query exec) │──→│ signal │ -│ │ │ back │ -└────────────────────┘ └─────────┘ -``` - -**How it works**: all I/O flows through one io_uring instance on the main thread. When a query arrives (CQE for RECV), the main thread dispatches it to a worker. The worker executes the query (CPU-bound, touching in-memory data structures), then signals completion back to the main loop via eventfd. The main loop submits the SEND SQE to return the result. - -**Pros**: fully deterministic — the main loop processes events in a fixed order, which enables replay-based testing and debugging; io_uring handles both disk and network in one scheduler; minimal coordination (workers are fire-and-forget); easiest to reason about. - -**Cons**: single main thread is a potential bottleneck for very high connection rates; worker pool dispatch adds one context switch per query; all state mutation in the main loop limits write throughput to one core. - -**Used by**: TigerBeetle (Zig), LMAX Disruptor (Java, similar philosophy). - -### Decision: Which Model for `.wo` - -> **Canonical model: single-threaded event loop.** The rest of this section catalogues a hybrid multi-threaded architecture that was the original proposal; it now lives here as the **scale-out reference** for when one core isn't enough. The Phase 2 spec ([02-wo-language.md § Concurrency Model](./database/02-wo-language.md#concurrency-model)) pins the initial runtime to a single userland thread driving io_uring directly (option 4 below, deterministic io_uring loop), without a worker pool, lock-free shared state, or a dedicated WAL-writer thread. Group commit still happens — the loop drains pending commits into one fsync SQE per tick. Scale past one core by **sharding** independent engine processes, not by reintroducing the multi-threaded hybrid here. - -The multi-threaded architecture below remains useful as (a) a reference for the sharded model's internal coordination when cross-shard 2PC is added, and (b) the fallback if a workload profile ever justifies abandoning the single-threaded invariant. Keep reading if you want the full trade-off map; skip to the next section if you only care about what ships. - -The architecture maps to the database subsystems (multi-threaded variant — not the current target): - -| Subsystem | Best fit | Why | -| --- | --- | --- | -| **I/O scheduling** (accept, recv, send, WAL, fsync) | io_uring on a dedicated I/O thread | One submission queue; batched syscalls; handles disk + network | -| **Query execution** (plan, traverse, filter) | Worker thread pool | CPU-bound; parallelizable per query; no I/O in the hot path (data is in RAM) | -| **WAL writer** | Single dedicated thread | Ordered writes; group commit batching; sequential fsync | -| **Subscription matching** | Worker thread pool (same as query execution) | CPU-bound predicate evaluation; parallelizable per commit | -| **Client push** | io_uring SEND from I/O thread | I/O-bound; batched with other sends | -| **Checkpoint / vacuum** | Background threads | Long-running, low-priority, can yield to production traffic | -| **Transaction coordination** | Lock-free shared state (AtomicU64 for LSN, CAS for version chains) | Must be accessible from any worker | - -This multi-threaded variant is a **hybrid** — closest to **option 4 (TigerBeetle) extended with a worker pool and lock-free shared state**. The actual Phase 2 design stops at plain option 4 (main-thread-only) — no worker pool, no cross-thread lock-free structures: - -``` -┌───────────────────────────────────────────────────────────┐ -│ .wo Runtime │ -│ │ -│ ┌───────────────────────────────────────────────────┐ │ -│ │ I/O Thread │ │ -│ │ io_uring ring: │ │ -│ │ ACCEPT → new session │ │ -│ │ RECV → dispatch query to worker pool │ │ -│ │ SEND → results / subscription deltas │ │ -│ │ WRITE → WAL (via WAL writer thread) │ │ -│ │ FSYNC → WAL barrier │ │ -│ │ TIMEOUT → keepalive / session expiry │ │ -│ └───────────────────────────────────────────────────┘ │ -│ │ dispatch ▲ results │ -│ ▼ │ │ -│ ┌───────────────────────────────────────────────────┐ │ -│ │ Worker Pool (N = num_cores - 2) │ │ -│ │ │ │ -│ │ worker 0: parse → plan → execute → match subs │ │ -│ │ worker 1: parse → plan → execute → match subs │ │ -│ │ worker 2: parse → plan → execute → match subs │ │ -│ │ ... │ │ -│ │ │ │ -│ │ Shared (lock-free): │ │ -│ │ - B+ tree (optimistic lock coupling) │ │ -│ │ - LSM memtable (crossbeam-skiplist) │ │ -│ │ - Graph adjacency (dashmap) │ │ -│ │ - MVCC version chains (atomic CAS) │ │ -│ │ - Subscription registry (sharded RwLock) │ │ -│ └───────────────────────────────────────────────────┘ │ -│ │ WAL records ▲ fsync ack │ -│ ▼ │ │ -│ ┌───────────────────────────────────────────────────┐ │ -│ │ WAL Writer Thread │ │ -│ │ │ │ -│ │ collect WAL records from worker batch queue │ │ -│ │ serialize → io_uring WRITE + FSYNC (linked SQE) │ │ -│ │ on CQE: signal workers that commit is durable │ │ -│ │ │ │ -│ │ Group commit: batch N commits into one fsync │ │ -│ └───────────────────────────────────────────────────┘ │ -│ │ -│ ┌────────────────────┐ ┌────────────────────┐ │ -│ │ Checkpoint thread │ │ Vacuum thread │ │ -│ │ (background) │ │ (background) │ │ -│ └────────────────────┘ └────────────────────┘ │ -└───────────────────────────────────────────────────────────┘ -``` - -### Thread Roles — Fixed, Not Dynamic - -Each thread has a single role for its lifetime. No work-stealing across roles. This gives predictability and avoids cache-pollution: - -| Thread | Count | Role | I/O model | -| --- | --- | --- | --- | -| **I/O thread** | 1 | Accept, recv, send via io_uring; dispatch queries; push subscription deltas | io_uring SQ/CQ poll | -| **Worker threads** | `num_cores - 3` | Parse, plan, execute queries; match subscriptions; stage MVCC mutations | Pure CPU; no I/O; no syscalls | -| **WAL writer** | 1 | Collect committed WAL records; batch; io_uring write + fsync; signal durability | Dedicated io_uring ring for WAL fd | -| **Checkpoint** | 1 | Periodic arena snapshot to SSD | io_uring or plain `pwrite` | -| **Vacuum** | 1 | Prune MVCC version chains; reclaim memtable space | CPU-bound scan | - -Total: `num_cores` threads, one per core, pinned via `pthread_setaffinity_np`. No oversubscription, no context switches between roles. - -### Query Lifecycle Through the Runtime - -A checkout query flows through the runtime like this: - -``` -1. Client sends EXECUTE frame over TCP - → io_uring CQE(RECV) on I/O thread - -2. I/O thread deserializes frame, identifies session + prepared plan - → pushes (session, plan, params) onto worker dispatch queue - -3. Worker thread picks up the query - → Acquires MVCC snapshot (AtomicU64 read — nanoseconds) - → Executes plan against in-memory structures: - UPDATE products SET inventory.on_hand = ... WHERE id = 42 - INSERT INTO orders ... - MATCH ... CREATE ... - → Stages mutations in a local write-set (no global visibility yet) - → Evaluates constraints - → Pushes WAL record to WAL writer’s batch queue - -4. WAL writer collects this record + records from other workers - → Serializes batch - → Submits io_uring WRITE + linked FSYNC to NVMe - → On CQE(FSYNC): marks all records in the batch as durable - → Signals each worker’s commit-complete channel - -5. Worker receives durability signal - → Publishes mutations to global MVCC (atomic pointer swaps) - → Evaluates subscription registry against the delta - → Pushes DELTA frames to the I/O thread’s send queue - -6. I/O thread submits io_uring SEND for each DELTA + the RESULT frame - → Client receives commit acknowledgment - → Subscribers receive live inventory update - -Total wall time: ~100–200 μs (dominated by step 4: NVMe fsync) -``` - -### Why Not Pure Async (tokio)? - -Tokio is production-proven and SurrealDB uses it. But for this database, the hybrid is better: - -| Concern | tokio | Hybrid (this design) | -| --- | --- | --- | -| **io_uring integration** | `tokio-uring` exists but is experimental; tokio’s core is epoll-based | io_uring is the primary I/O model; no epoll fallback | -| **Determinism** | Work-stealing introduces non-deterministic scheduling | Fixed thread roles; deterministic dispatch; replay-testable | -| **GC / allocator pressure** | Futures allocate on the heap; many small allocations per query | Workers use arena allocators; pre-allocated per-query scratch space | -| **Cache locality** | Tasks migrate between cores via work-stealing | Threads are pinned; data stays cache-hot per core | -| **Debugging** | Async stack traces are notoriously hard to read | Each thread has a clear role; stack traces are synchronous | -| **Dependency weight** | tokio + tower + hyper + … | libc + io_uring syscalls | - -The trade-off: less library reuse, more manual plumbing. For a database where every microsecond on the commit path matters, that trade-off is correct. - -### Why Not Pure Fibers (Go)? - -Go’s goroutine scheduler is excellent and CockroachDB proves databases can be built on it. But: - -| Concern | Go goroutines | Hybrid (this design) | -| --- | --- | --- | -| **GC pauses** | Stop-the-world pauses during checkout fsync batches lose money | No GC — Rust or C++ with arena allocators | -| **io_uring** | Go has no native io_uring support; falls back to epoll + thread pool for disk I/O | io_uring is first-class | -| **Memory control** | Cannot pin arenas, control huge pages, or use `mlockall` idiomatically | Full control via `libc` bindings | -| **Lock-free structures** | Possible but `sync/atomic` is more limited than Rust’s `crossbeam` | `crossbeam`, `dashmap`, `arc-swap` — mature ecosystem | - -If the database were written in Go, goroutines would be the right model. Since the [Phase 2 language analysis](./database/02-wo-language.md) argues for Rust or C++, fibers are not the natural fit. - -### Synchronization Between Threads - -Workers touch shared data structures. The synchronization budget is strict — any contention on the commit path adds latency to every checkout: - -| Shared structure | Accessed by | Primitive | Contention | -| --- | --- | --- | --- | -| **LSN counter** | All workers + WAL writer | `AtomicU64::fetch_add` | One atomic increment per commit — cheapest possible | -| **MVCC version chains** | All workers (read + write) | Atomic CAS to prepend | Per-record; independent records don’t contend | -| **B+ tree internal nodes** | All workers (read) | Optimistic lock coupling (version + retry) | Read-dominant; writes hold latches for microseconds | -| **Worker dispatch queue** | I/O thread (push) + workers (pop) | Lock-free MPSC queue (`crossbeam-channel`) | One producer, N consumers — no contention | -| **WAL batch queue** | Workers (push) + WAL writer (drain) | Lock-free MPMC queue | Drained in bulk every fsync batch (~100 μs) | -| **Subscription registry** | Workers (match) + I/O thread (register/unregister) | Sharded `RwLock` | Readers (match on commit) never block each other | -| **Send queue** | Workers (push DELTA) + I/O thread (drain to io_uring) | Lock-free MPSC per connection | One consumer per connection fd | - -No mutex on the commit path. The only serialization point is the LSN counter — a single `fetch_add`. - -### Handling Compute-Heavy Queries - -Graph traversals and analytical queries can consume milliseconds of CPU. In a fiber model, a long-running fiber starves others (cooperative scheduling). In this hybrid: - -- Workers are pre-emptible at the OS level (kernel threads, not fibers). -- Each worker runs one query at a time to completion. If a query takes 5 ms, that worker is busy for 5 ms — but the other N-1 workers continue serving other queries. -- If the pool is saturated, the I/O thread applies back-pressure: it stops reading from client sockets (io_uring RECV is not re-submitted), TCP flow control kicks in, and clients experience latency — which is the correct behavior under overload. -- Optional: a per-query CPU budget (checked at loop iteration points in graph traversal) that yields the worker back to the pool and resumes the query later. This is partial preemption — fiber-like semantics within a thread pool. - -### Lifecycle of a Subscription - -Subscriptions are long-lived — they span many commits. The runtime handles them without dedicated threads or fibers: - -``` -1. Client sends SUBSCRIBE frame - → I/O thread registers (predicate, client_fd) in subscription registry - -2. A commit happens on a worker thread - → Worker evaluates subscription registry against the commit’s delta - → For each match: serialize DELTA frame, push to that client’s send queue - -3. I/O thread drains send queues - → Submits io_uring SEND for each queued DELTA - -4. Client disconnects (CQE reports EPOLLHUP-equivalent) - → I/O thread removes all subscriptions for that fd -``` - -No thread or fiber per subscription. No polling. The cost of a subscription is: one entry in the registry (a few hundred bytes) + O(1) evaluation per commit (if keyed) or O(subs) per commit (if arbitrary predicate). Thousands of active subscriptions add microseconds to each commit, not threads. - -### Comparison With Real Database Runtimes - -| Database | Language | Runtime model | I/O | Query scheduling | -| --- | --- | --- | --- | --- | -| **Postgres** | C | Process-per-connection | epoll + blocking I/O on worker processes | One process per query | -| **MySQL** | C++ | Thread-per-connection or thread pool | epoll | One thread per query | -| **SurrealDB** | Rust | tokio (M:N async) | epoll (tokio) + Rayon for CPU | async tasks + Rayon parallel iterators | -| **ScyllaDB** | C++ | Thread-per-core (Seastar) | io_uring / epoll / SPDK | Futures on per-core reactor | -| **TigerBeetle** | Zig | Single-threaded io_uring loop | io_uring | Deterministic, single-threaded | -| **CockroachDB** | Go | M:N goroutines | epoll (Go netpoller) | Goroutine per query | -| **DuckDB** | C++ | Thread pool | Blocking I/O | Morsel-driven parallelism | -| **`.wo`** (this design) | Rust/C++ | **Hybrid: io_uring I/O thread + pinned worker pool + dedicated WAL thread** | io_uring | Worker per query, lock-free shared state | - -### What the Runtime Does NOT Do - -Keeping the scope honest: - -- **No work-stealing**. Workers are pinned, queries are assigned round-robin or by shard affinity. Work-stealing adds scheduling complexity for marginal throughput gains when the data is in RAM and queries are sub-millisecond. -- **No async/await in the engine codebase**. Workers run synchronous code against in-memory structures. The only async code is the I/O thread’s io_uring event loop. -- **No fiber stacks**. No `swapcontext`, no stack allocation per query. Each worker has one OS stack, runs one query at a time. -- **No M:N scheduling**. N queries map to N worker invocations, but each invocation is a plain function call on a fixed thread — not a scheduled task on a shared executor. - -The runtime is deliberately simpler than tokio, Go’s scheduler, or Seastar. The bet is that with all data in RAM, query execution is fast enough that a fixed-size thread pool with lock-free shared state is sufficient — and far easier to debug, profile, and reason about. - -### Reference Implementations - -- **Seastar** — . The thread-per-core framework. Read `seastar/core/reactor.cc` for the io_uring event loop and `seastar/core/smp.cc` for inter-core messaging. -- **Glommio** — . Rust thread-per-core runtime built on io_uring. Closest Rust analogue to Seastar. -- **tokio** — . Work-stealing async runtime. Read `tokio/src/runtime/scheduler/` for the work-stealing logic. -- **TigerBeetle** — . Deterministic single-threaded io_uring loop. Read `src/io.zig` for the I/O ring and `src/state_machine.zig` for deterministic processing. -- **crossbeam** — . Lock-free data structures for Rust: channels, skiplist, deque, epoch-based GC. -- **LMAX Disruptor** — . Lock-free ring buffer for inter-thread communication. The intellectual ancestor of the WAL batch queue. -- **Readings**: - - *The Seastar Tutorial* — thread-per-core explained from first principles. - - *Fibers under the magnifying glass* (Vyukov) — M:N scheduling in practice. - - *io_uring and networking in 2023* (Axboe) — io_uring for both storage and networking. - - *LMAX Architecture* (Fowler) — single-writer, mechanical sympathy, lock-free coordination. - -### Where This Fits in the Database Series - -This runtime design is the missing piece between [Phase 2 (ACID engine)](./database/02-wo-language.md) and [Phase 3 (in-memory storage)](./database/03-inmemory-engine.md). Phase 2 described *what* the engine does (ACID transactions across three paradigms). Phase 3 described *where* data lives (RAM, with io_uring WAL). This section describes *how* work is scheduled — the thread architecture that connects client I/O, query execution, WAL durability, and subscription push into a single coherent runtime. diff --git a/docs/runtime/garbage-collection.md b/docs/runtime/garbage-collection.md deleted file mode 100644 index 08422df..0000000 --- a/docs/runtime/garbage-collection.md +++ /dev/null @@ -1,248 +0,0 @@ -# Why Rust Does Not Use Fibers or Garbage Collection - -## The Question - -Most runtime-heavy languages ship with two things: a garbage collector (Go, Java, Python, C#, Erlang) and fiber-like concurrency (Go goroutines, Erlang processes, Java virtual threads). Rust ships with neither. Why? - -The answer is the same for both: **Rust pushes the cost to compile time so there is zero cost at runtime.** - -## Garbage Collection - -### What a GC Does - -A garbage collector tracks which objects in memory are still reachable from the program. Periodically (or continuously), it scans the heap, finds objects nothing points to, and frees them. - -``` -Allocate object A -Allocate object B -A.ref = B // B is reachable through A -drop(A) // A is unreachable — GC will free A - // B is now also unreachable — GC will free B -``` - -### How It Works (Simplified) - -**Mark-and-sweep** (Go, Java): - -``` -1. Pause the program (or run concurrently) -2. Start from "roots" (stack variables, globals) -3. Mark every object reachable from roots -4. Sweep: free every object NOT marked -``` - -**Reference counting** (Python, Swift, Objective-C): - -``` -1. Every object has a counter -2. When a reference is created: counter++ -3. When a reference is dropped: counter-- -4. When counter == 0: free immediately -``` - -### The Costs - -| Cost | Mark-and-sweep GC | Reference counting | -|------|-------------------|-------------------| -| Pause time | Stop-the-world pauses (Go: ~1ms, Java: varies) | No pauses, but slower per-operation | -| Memory overhead | 2x heap needed (live objects + garbage until collected) | Counter per object (8 bytes) | -| CPU overhead | GC thread scanning heap (10-30% throughput loss) | Increment/decrement on every pointer operation | -| Predictability | Unpredictable latency spikes | Predictable but cycles leak (need cycle collector) | -| Cache impact | GC walks heap → cache pollution | Counters spread across memory → cache misses | - -For a content platform that needs predictable low-latency responses, GC pauses are the enemy. Even Go's ~1ms pauses compound under load — if a GC pause hits during `epoll_wait`, every pending connection stalls. - -### What Rust Does Instead: Ownership - -Rust replaces garbage collection with a compile-time ownership system: - -```rust -fn main() { - let s = String::from("hello"); // s owns the string, allocated on heap - let t = s; // ownership moves to t — s is invalid - // println!("{}", s); // compile error: s was moved - println!("{}", t); // ok -} // t goes out of scope → String freed here. Deterministic. No GC. -``` - -The rules: - -1. **Every value has exactly one owner** -2. **When the owner goes out of scope, the value is dropped (freed)** -3. **Ownership can be moved or borrowed, but never duplicated** - -The compiler enforces these rules at compile time. At runtime, there is: -- No GC thread -- No mark phase -- No sweep phase -- No reference counters -- No heap scanning -- No pauses - -Memory is freed at the exact point it is no longer needed — deterministically, at the closing brace. - -### Lifetimes: The Compile-Time GC - -References (borrows) have lifetimes — the compiler tracks how long each reference lives and ensures no reference outlives its data: - -```rust -fn longest<'a>(x: &'a str, y: &'a str) -> &'a str { - if x.len() > y.len() { x } else { y } -} -``` - -The `'a` lifetime annotation tells the compiler: "the returned reference lives as long as both inputs." If you try to return a reference to a local variable, the compiler rejects it — at compile time, not at runtime. - -```rust -fn bad() -> &str { - let s = String::from("hello"); - &s // compile error: s is dropped at end of function, reference would dangle -} -``` - -This is what a GC does at runtime (detect unreachable memory). Rust does it at compile time (detect impossible references). Zero runtime cost. - -### The Tradeoff - -| Aspect | GC languages | Rust | -|--------|-------------|------| -| Developer effort | Low — just allocate, GC handles cleanup | Higher — must think about ownership and lifetimes | -| Compile time | Fast | Slower (borrow checker analysis) | -| Runtime cost | GC pauses, heap scanning, memory overhead | Zero — deterministic drop at scope exit | -| Latency | Unpredictable (GC can pause anytime) | Predictable — no hidden pauses | -| Memory usage | 2x+ (garbage accumulates between collections) | Tight — freed immediately when unused | - -Rust trades developer convenience for runtime performance. For systems software (databases, runtimes, web servers), this is the right trade. - -## Fibers - -### Why Other Languages Use Fibers - -Go has goroutines. Erlang has processes. Java 21 has virtual threads. These are all fibers — lightweight user-space threads that the runtime schedules cooperatively (or semi-preemptively in Go's case). - -They exist because these languages need to: -1. Handle millions of concurrent I/O tasks -2. Let developers write synchronous-looking code (`result = fetch(url)`) that blocks the fiber, not the OS thread -3. Manage scheduling without exposing the event loop - -```go -// Go: goroutine blocks on I/O — runtime suspends it and runs another -go func() { - resp, _ := http.Get("https://example.com") // blocks this goroutine, not the thread - fmt.Println(resp.Status) -}() -``` - -### Why Rust Does Not Use Fibers - -**1. Fibers require a runtime that allocates stacks.** - -Each fiber needs its own stack (Go starts at 2-8 KB, grows dynamically). This means: -- A heap allocation per fiber -- Stack overflow checks on every function call -- A runtime that manages stack growth and shrinking -- Memory overhead proportional to number of concurrent tasks - -Rust's goal is zero-cost abstractions. Allocating stacks at runtime is a cost. - -**2. Fibers are hard to optimize across FFI boundaries.** - -Rust interoperates with C libraries extensively. Fibers with tiny stacks can't safely call into C code (which expects a full OS stack). Go solves this by switching to a system stack for cgo calls — adding complexity and overhead. - -**3. Async/await achieves the same concurrency without stacks.** - -Rust's async/await compiles each async function into a state machine — a regular struct stored inline, no heap allocation needed: - -```rust -async fn fetch_article(title: &str) -> Article { - let data = read_from_seg(title).await; // suspend point 1 - let html = render_markdown(&data).await; // suspend point 2 - Article { title, html } -} -``` - -The compiler transforms this into something like: - -```rust -enum FetchArticle { - Start { title: String }, - AfterRead { title: String, data: Vec }, - AfterRender { title: String, html: String }, - Done, -} -``` - -Each `.await` becomes a variant transition. The "stack" is just the live variables in the current variant — bytes, not kilobytes. No allocation, no stack, no runtime overhead. - -### Fiber vs Async/Await: Memory Per Task - -``` -Go goroutine: ~2,048 bytes minimum (stack) -Erlang process: ~2,688 bytes minimum (stack + heap + mailbox) -Rust async task: size_of::() — often 32-128 bytes -``` - -For a million concurrent connections: -- Go: ~2 GB just for goroutine stacks -- Rust: ~128 MB for state machines (and often less, since the executor batches them) - -### When Fibers Would Be Better - -Fibers have one advantage: **deeply nested call stacks that suspend at arbitrary points.** If a function 20 calls deep needs to yield, a fiber just swaps the stack pointer. With async/await, every function in the chain must be `async` and every call must be `.await`ed — the "async infection" problem. - -``` -Fibers: yield anywhere in the call stack — transparent to callers -Async: yield only at .await points — every caller must be async -``` - -For database engines with complex query execution plans that suspend mid-evaluation, fibers are compelling. For an HTTP server that suspends at I/O boundaries, async/await is strictly better. - -## How This Applies to writeonce - -writeonce uses neither fibers nor async/await. It uses a plain event loop: - -```rust -loop { - events = epoll_wait(); - for event in events { - handle(event); // runs to completion, no suspension - } -} -``` - -This is the simplest model — no GC, no fibers, no async state machines. Each handler reads from the `.seg` file, renders a template, writes to the socket, and returns. Nothing suspends mid-handler. - -The memory model: - -| What | How it's managed | -|------|-----------------| -| Article data in .seg | Owned by `Store`, freed when `Store` drops | -| Template ASTs | Owned by `TemplateRegistry`, live for the process lifetime | -| HTTP connections | Owned by `HashMap`, freed on close/hangup | -| Subscription table | Owned by `SubscriptionManager`, entries removed on `EPOLLHUP` | - -No garbage. No fibers. No async. Just ownership, scopes, and the kernel's event notification. The Rust compiler guarantees at compile time that every allocation is freed exactly once, at exactly the right time. - -## Summary - -``` -GC languages (Go, Java): runtime scans heap → frees unreachable objects - cost: pauses, memory overhead, CPU overhead - -Reference counting (Python): counter per object → free at zero - cost: per-operation overhead, cycle leaks - -Rust ownership: compiler tracks ownership → free at scope exit - cost: zero at runtime, developer thinks harder - -Fibers (Go, Erlang): runtime manages stacks → swap on yield - cost: stack allocation, stack checks, runtime - -Async/await (Rust): compiler generates state machines → no stack - cost: zero allocation, async must propagate - -Event loop (writeonce): no tasks, no suspension → handlers run to completion - cost: nothing — simplest possible model -``` - -Rust's answer to both GC and fibers is the same: **make the compiler do the work so the runtime doesn't have to.** diff --git a/docs/runtime/surreal-case-study.md b/docs/runtime/surreal-case-study.md deleted file mode 100644 index 7cf8e59..0000000 --- a/docs/runtime/surreal-case-study.md +++ /dev/null @@ -1,157 +0,0 @@ -# SurrealDB — Runtime Case Study - -Reference repository: [github.com/surrealdb/surrealdb](https://github.com/surrealdb/surrealdb) - -```bash -git submodule add https://github.com/surrealdb/surrealdb.git references/surrealdb -``` - -## The Question - -Does SurrealDB only rely on async/await for concurrency? - -**No.** SurrealDB uses a layered concurrency model — async/await is one layer, but it also uses OS thread pools, CPU-affinity-pinned workers, lock-free data structures, and parallel computation frameworks. Each layer serves a different purpose. - -## Architecture Overview - -SurrealDB is a single Rust binary that ships a multi-model database (documents, graphs, key-value) with real-time live queries. It supports multiple deployment modes: - -- **Server**: `surreal start` runs HTTP/WebSocket API via Axum + storage engine -- **Embedded**: the library crate embeds directly in Rust applications -- **WASM**: runs in the browser with IndexedDB backend - -## Concurrency Layers - -### Layer 1: Tokio — Async I/O and Request Handling - -The primary runtime. Handles: -- HTTP/WebSocket connections (via Axum) -- Network I/O (accept, read, write) -- Timer-based operations -- Task scheduling (M:N scheduling of futures onto OS threads) - -``` -Client request → Axum handler (async) → parse query → execute → respond -``` - -Every request handler is an async function. Tokio's multi-threaded executor distributes tasks across OS threads using work-stealing. - -### Layer 2: Rayon — Parallel CPU-Bound Computation - -For operations that are compute-heavy, not I/O-bound: -- Query plan execution across partitions -- Data processing and transformation -- Parallel iteration over result sets - -Rayon provides `par_iter()` — automatic parallelism across CPU cores. It has its own thread pool, separate from tokio's. - -### Layer 3: affinitypool — CPU-Pinned Storage I/O - -SurrealDB's custom crate. Runs blocking storage operations on a dedicated thread pool where **each thread is pinned to a specific CPU core** via `libc` CPU affinity syscalls. - -Used by: -- RocksDB backend (blocking disk I/O) -- SurrealKV embedded storage -- In-memory engine for heavy operations - -This bridges the async world (tokio) and the blocking world (disk I/O) without polluting the tokio thread pool with blocking calls. - -### Layer 4: Lock-Free Data Structures - -The hot path in storage engines uses concurrent data structures that avoid locks entirely: - -| Crate | Data Structure | Used For | -|-------|---------------|----------| -| `crossbeam-skiplist` | Concurrent skip list | Index structures in SurrealKV and surrealmx | -| `crossbeam-deque` | Work-stealing deque | Task distribution | -| `crossbeam-queue` | Lock-free queue | Message passing | -| `papaya` | Concurrent HashMap | In-memory engine (surrealmx) | -| `dashmap` | Sharded concurrent map | Pub/sub routing for live queries | -| `arc-swap` | Atomic pointer swap | Hot-swapping data structures without locks | -| `parking_lot` | Fast mutex/rwlock | Where locking is needed (faster than std) | - -## No Fibers - -SurrealDB does **not** use fibers, green threads, or any custom scheduling mechanism. The concurrency model is: - -``` -Tokio tasks (async/await) — for I/O-bound work -Rayon threads (par_iter) — for CPU-bound work -affinitypool threads (pinned) — for blocking storage I/O -Lock-free structures — for concurrent data access -``` - -This is pragmatic — each concurrency mechanism is used where it fits, rather than forcing everything through one model. - -## Live Queries / Real-Time Subscriptions - -SurrealDB's live query system pushes changes to connected clients in real-time: - -1. Client registers a live query via WebSocket: `LIVE SELECT * FROM person WHERE age > 21` -2. Server tracks the query in a `dashmap` (concurrent map) -3. When a transaction commits changes to `person`, the engine evaluates which live queries are affected -4. Matching subscribers receive the diff via their WebSocket connection -5. Transport: `tokio-tungstenite` for WebSocket, `async-channel` for internal pub/sub routing - -### Comparison with writeonce Subscriptions - -| Aspect | SurrealDB | writeonce | -|--------|-----------|-----------| -| Transport | WebSocket (tokio-tungstenite) | Raw socket fd (kernel-level write) | -| Query registration | SQL-like live query over WebSocket | `register!` macro binding fd to content pattern | -| Change detection | Transaction commit triggers evaluation | inotify detects file change | -| Notification routing | `dashmap` + `async-channel` | `SubscriptionManager` HashMap + direct `write(fd)` | -| Runtime | Tokio multi-threaded executor | Single-threaded epoll event loop | -| Protocol framing | WebSocket frames | Length-prefixed payloads (no protocol) | - -SurrealDB's live queries are the architectural inspiration for writeonce's subscription model (as noted in 03-data.md), but the implementation is fundamentally different — SurrealDB uses a full async runtime with WebSocket transport, while writeonce uses kernel fd notifications with no protocol layer. - -## Storage Engine Architecture - -SurrealDB supports 5 backends: - -| Backend | Type | Concurrency | -|---------|------|------------| -| **surrealmx** | In-memory | Lock-free (papaya, crossbeam-skiplist, arc-swap) | -| **surrealkv** | Embedded persistent | Tokio async + crossbeam + parking_lot | -| **RocksDB** | Embedded persistent | affinitypool (CPU-pinned blocking threads) | -| **TiKV** | Distributed | Async TiKV client over gRPC | -| **IndxDB** | Browser/WASM | IndexedDB via wasm-bindgen-futures | - -### Comparison with writeonce Storage - -| Aspect | SurrealDB | writeonce | -|--------|-----------|-----------| -| Storage format | Key-value entries in LSM trees (RocksDB) or custom B-trees (SurrealKV) | `.seg` files with length-prefixed bincode records | -| Index | Built into storage engine | Separate `.idx` files (title hash, date sorted, tags inverted) | -| Concurrency | Multi-threaded with locks/lock-free structures | Single-threaded, positional I/O (pread/pwrite) | -| Transaction | ACID with MVCC | Full rebuild on change (article count is small) | -| Complexity | ~100K+ lines across storage crates | ~300 lines (wo-seg + wo-index) | - -writeonce's storage is intentionally simple — the dataset is small (hundreds of articles, not millions of rows), so a full rebuild on change is fast enough and avoids the complexity of concurrent transactions. - -## Key Takeaways - -1. **Async/await alone is not enough for a database.** SurrealDB uses four concurrency mechanisms, each for a different workload profile. - -2. **Blocking I/O needs its own thread pool.** The affinitypool pattern — CPU-pinned threads for storage operations — keeps blocking work off the async executor. writeonce avoids this entirely by using `pread` (non-blocking positional reads) in a single-threaded loop. - -3. **Lock-free data structures matter at scale.** SurrealDB's hot path avoids mutexes. writeonce doesn't need this — single-threaded access means no contention. - -4. **Live queries are the hard problem.** Both SurrealDB and writeonce solve "push changes to subscribers," but at vastly different scales. SurrealDB handles arbitrary SQL predicates over millions of rows. writeonce handles content queries over hundreds of articles. - -5. **The right amount of complexity depends on the problem.** SurrealDB is a general-purpose database — it needs the complexity. writeonce is a content platform — the event loop model is sufficient and far simpler. - -## Reference - -Add SurrealDB as a submodule for code reference: - -```bash -git submodule add https://github.com/surrealdb/surrealdb.git references/surrealdb -``` - -Key files to study: -- `crates/core/src/kvs/` — storage engine abstraction and transaction handling -- `crates/sdk/src/api/engine/` — live query subscription routing -- `lib/affinitypool/` — CPU-pinned thread pool for blocking I/O -- `crates/core/src/sql/` — query parser and execution engine diff --git a/justfile b/justfile index 69a8542..f2de9cd 100644 --- a/justfile +++ b/justfile @@ -1,27 +1,5 @@ # writeonce — task runner. `just --list` shows all recipes. -# C runtime reference (prototypes/wo-rt-c): build, serve, CRUD round-trip, shut down -# Phase A: thread-per-core — each connection hashes to one shard (SO_REUSEPORT), -# so a list may land on a different shard than the create. The counters on / -# show the spread. WO_THREADS=4 keeps the demo output readable. -rt-c-demo port="8085" threads="4": - #!/usr/bin/env bash - set -euo pipefail - make -C prototypes/wo-rt-c - data=$(mktemp -d /tmp/wo-demo-XXXXXX) - WO_PORT={{port}} WO_THREADS={{threads}} WO_DATA=$data ./prototypes/wo-rt-c/wo-rt & - server=$! - trap 'kill $server 2>/dev/null; sleep 0.3; rm -rf $data' EXIT - base=http://127.0.0.1:{{port}} - for _ in $(seq 1 40); do curl -s "$base/healthz" >/dev/null && break; sleep 0.25; done - echo - echo "--- runtime:"; curl -s "$base/"; echo - echo "--- create x4 (each connection may hash to a different shard):" - for i in 1 2 3 4; do curl -s -X POST "$base/api/notes" -d '{"title":"note '$i'"}'; echo; done - echo "--- list (the connection's own shard only — shared-nothing):" - curl -s "$base/api/notes"; echo - echo "--- spread:"; curl -s "$base/"; echo - # woc compiler front (compiler/): build the executable woc-build: dune build --root compiler @@ -143,21 +121,3 @@ oop-accept: echo echo "oop-accept: ALL CRITERIA MET" - -# phase-F benchmark: reads, durable writes, 10k idle conns (scaled geometry) -rt-c-bench port="8085" threads="8" conns="64": - #!/usr/bin/env bash - set -euo pipefail - make -C prototypes/wo-rt-c clean >/dev/null - make -C prototypes/wo-rt-c CFLAGS="-O2 -Wall -Wextra -std=c11 -DSLOTS_PER_SHARD=262144" wo-rt bench >/dev/null - data=$(mktemp -d /tmp/wo-bench-XXXXXX) - WO_PORT={{port}} WO_THREADS={{threads}} WO_DATA=$data ./prototypes/wo-rt-c/wo-rt >/dev/null 2>&1 & - server=$! - trap 'kill $server 2>/dev/null; sleep 0.3; rm -rf $data; make -C prototypes/wo-rt-c clean >/dev/null; make -C prototypes/wo-rt-c wo-rt bench >/dev/null' EXIT - base=127.0.0.1; for _ in $(seq 1 40); do curl -s "http://$base:{{port}}/healthz" >/dev/null && break; sleep 0.25; done - B=./prototypes/wo-rt-c/bench/bench - echo "wo-rt-c ({{threads}} shards, durable WAL):" - $B $base {{port}} {{conns}} 5 /healthz - $B $base {{port}} {{conns}} 5 / - $B $base {{port}} {{conns}} 3 /api/notes '{"title":"bench"}' - $B $base {{port}} 10000 0 /healthz diff --git a/prototypes/wo-db/README.md b/prototypes/wo-db/README.md deleted file mode 100644 index eb7d8ac..0000000 --- a/prototypes/wo-db/README.md +++ /dev/null @@ -1,140 +0,0 @@ -# wo-db — `.wo` language prototype (C++) - -Phase 2 milestone 1 from [docs/runtime/database.md](../../docs/runtime/database.md): -**parser + analyzer + in-memory executor, single-user, no durability.** - -Implements a cut-down `.wo` dialect spanning all three paradigms described in -[02-wo-language.md](../../docs/runtime/database/02-wo-language.md): relational, -document, and graph — all in one grammar, one process, one in-RAM store. - -## Build & run - -```bash -make # builds build/wo-db -make test # runs tests/smoke.wo + tests/checkout.wo -./build/wo-db # interactive REPL -./build/wo-db < file.wo # batch mode -``` - -Requires `g++` with C++20 (tested on 13.3). No external dependencies. - -## Supported grammar - -### Schema - -```wo -##sql -#users - id int - name string - email string - -##doc -#article_meta - id int - title string - -##graph -#recommendations - (user)-[:PURCHASED]->(product) -``` - -Paradigms: `sql`, `doc`, `graph`. Types are parsed but not enforced (prototype). -Graph schema blocks are documentation only — nodes and edges are created at -runtime by `CREATE`. - -### Queries - -```wo --- relational -INSERT INTO users (id, name) VALUES (1, 'Alice'); -SELECT id, name FROM users WHERE id > 1 AND name != 'Carol'; -UPDATE users SET email = 'a@b.com' WHERE id = 1; -DELETE FROM users WHERE id = 3; - --- document (same syntax as sql; dotted columns write nested objects) -INSERT INTO article_meta (id, title) VALUES (100, 'Intro'); -SELECT title FROM article_meta; - --- graph -CREATE (u:user {id: 1, name: 'Alice'}); -CREATE (u:user {id: 1})-[:PURCHASED {qty: 2}]->(p:product {id: 10}); -MATCH (u:user {id: 1})-[:PURCHASED]->(p:product) RETURN u, p; -``` - -### Fixed-glue: parameters, RETURNING, transactions, LIVE - -The five things from the two-layer design ([docs/runtime/database/02-wo-language.md](../../docs/runtime/database/02-wo-language.md)) that tie the three grammars together: - -```wo --- $name parameters work everywhere (SQL, Cypher, expressions) -INSERT INTO users (email) VALUES ('a@b.com') RETURNING id AS uid; -SELECT * FROM users WHERE id = $uid; - --- cross-paradigm RETURNING threads ids from SQL into Cypher -BEGIN SNAPSHOT; - INSERT INTO orders (user_id, status) VALUES ($uid, 'pending') RETURNING id AS oid; - CREATE (u:user {id: $uid})-[:PURCHASED {order_id: $oid}]->(p:product {id: $pid}); -COMMIT; - --- SAVEPOINT / ROLLBACK TO are parsed but not yet enforced in the prototype -BEGIN; - SAVEPOINT s1; - INSERT INTO users (email) VALUES ('typo@example.com'); - ROLLBACK TO s1; -COMMIT; - --- LIVE prefix reserved — inner query runs now, subscription activation in Phase 3 -LIVE SELECT id, email FROM users; -LIVE MATCH (u:user)-[:PURCHASED]->(p:product) RETURN u, p; -``` - -Auto-populated `id` columns: when a table declares an `id int` column and `INSERT` omits it, the engine mints a fresh id from a per-table counter and fills it in. Combined with `RETURNING id AS alias`, this is how ids thread from SQL into Cypher inside a transaction without `LAST_INSERT_ID()`. - -### Expressions - -Literals (`int`, `string`, `true`/`false`/`null`, `[array]`, `{object}`), -dotted paths (`meta.title`), comparisons (`= != < <= > >=`), boolean -(`AND OR NOT`). - -### REPL meta-commands - -``` -.tables list sql/doc tables -.schema dump schema -.exit -``` - -## What's intentionally missing - -This is a Phase 2 *prototype*, not a database. It does not implement: - -- **real** transactions — `BEGIN/COMMIT/SAVEPOINT/ROLLBACK` parse and are acknowledged, but no atomic rollback, no MVCC, no WAL -- **real** `LIVE` subscriptions — the inner query runs; no delta frames, no push (Phase 3/4) -- durability, crash recovery (Phase 3) -- real document operations (array push/splice, deep path updates beyond simple dotted SET) -- joins, aggregations, subqueries -- the LSM document engine, B+ tree relational pages, native graph adjacency (Phase 3) -- wire protocol, auth (Phase 4+) - -See [docs/runtime/database.md](../../docs/runtime/database.md) for the full phase plan -and [docs/runtime/database/02-wo-language.md](../../docs/runtime/database/02-wo-language.md) -for the target language spec. - -## Source layout - -``` -src/ - value.hpp tagged Value (null/bool/int/string/array/object) - lexer.{hpp,cpp} tokenizer - ast.hpp AST node types - parser.{hpp,cpp} recursive-descent parser - storage.{hpp,cpp} in-memory Database (sql tables + doc tables + graph) - executor.{hpp,cpp} walks AST, returns ResultSet - main.cpp REPL + batch driver -tests/ - smoke.wo core statement coverage - checkout.wo cross-paradigm RETURNING + BEGIN/SAVEPOINT/COMMIT + LIVE -``` - -Roughly 1.1k lines of C++ in the `wo` namespace.