diff --git a/docs/plan/exploration/assembly/00-overview.md b/docs/plan/exploration/assembly/00-overview.md index 71235b6..db845d1 100644 --- a/docs/plan/exploration/assembly/00-overview.md +++ b/docs/plan/exploration/assembly/00-overview.md @@ -1,14 +1,14 @@ # 00 — The role of assembly in a runtime -Why does a runtime ship hand-written assembly at all? Three reasons — each one a place where a higher-level language literally cannot express the operation it needs, so the compiler is bypassed and machine instructions are written directly. Go's [`src/runtime/`](../../../reference/go/src/runtime/) is the canonical example; this doc names the three reasons and points at the Go files that embody each. +Why does a runtime ship hand-written assembly at all? Three reasons — each one a place where a higher-level language literally cannot express the operation it needs, so the compiler is bypassed and machine instructions are written directly. Go's [`src/runtime/`](../../../.dev/reference/go/src/runtime/) is the canonical example; this doc names the three reasons and points at the Go files that embody each. ## 1 — Operations that violate the language's own calling convention The biggest category. The language's calling convention — how arguments are passed, who saves which registers, how the stack grows — is the contract every compiled function obeys. A few runtime operations *have* to break it because they ARE the mechanism by which control flow enters and exits that contract. -**Goroutine stack switching.** When Go's scheduler switches from one goroutine to another, it's literally rewriting the stack pointer mid-function — jumping from one goroutine's stack to another's. The language compiler can't emit this safely because every function assumes its stack is the one it got called on. See [`reference/go/src/runtime/asm_amd64.s`](../../../reference/go/src/runtime/asm_amd64.s) for `TEXT runtime·gogo(SB)`, `TEXT runtime·mcall(SB)`, `TEXT runtime·systemstack(SB)` — all unavoidable. +**Goroutine stack switching.** When Go's scheduler switches from one goroutine to another, it's literally rewriting the stack pointer mid-function — jumping from one goroutine's stack to another's. The language compiler can't emit this safely because every function assumes its stack is the one it got called on. See [`.dev/reference/go/src/runtime/asm_amd64.s`](../../../.dev/reference/go/src/runtime/asm_amd64.s) for `TEXT runtime·gogo(SB)`, `TEXT runtime·mcall(SB)`, `TEXT runtime·systemstack(SB)` — all unavoidable. -**Signal-handler entry.** When a signal arrives, the kernel drops the process onto an alternate stack with preserved registers. Returning to normal code means restoring everything the handler touched plus switching stacks back. Go's `runtime·sigtramp` in [`reference/go/src/runtime/sys_linux_amd64.s`](../../../reference/go/src/runtime/sys_linux_amd64.s) handles this. +**Signal-handler entry.** When a signal arrives, the kernel drops the process onto an alternate stack with preserved registers. Returning to normal code means restoring everything the handler touched plus switching stacks back. Go's `runtime·sigtramp` in [`.dev/reference/go/src/runtime/sys_linux_amd64.s`](../../../.dev/reference/go/src/runtime/sys_linux_amd64.s) handles this. **Cgo boundary crossing.** Calling C from Go means switching to the OS thread's "real" stack (C expects contiguous stacks; Go uses segmented). Going back means the inverse. Entirely asm-driven. @@ -16,17 +16,17 @@ The biggest category. The language's calling convention — how arguments are pa Atomics, memory barriers, and some hardware-accelerated primitives need specific instruction sequences. A compiler that sees `a = *b` can't know whether you wanted a relaxed load or an acquire fence without annotation — and the *right* instruction on x86 vs ARM vs RISC-V is different. -**Atomic CAS / load-acquire / store-release.** On x86 it's `LOCK CMPXCHG`; on ARM it's `LDXR` / `STXR` with a retry loop; on RISC-V it's `LR.W.AQ` / `SC.W.RL`. Go emits these from [`reference/go/src/runtime/atomic_amd64.s`](../../../reference/go/src/runtime/atomic_amd64.s) (and its per-arch siblings) because a portable compiler can't. +**Atomic CAS / load-acquire / store-release.** On x86 it's `LOCK CMPXCHG`; on ARM it's `LDXR` / `STXR` with a retry loop; on RISC-V it's `LR.W.AQ` / `SC.W.RL`. Go emits these from [`.dev/reference/go/src/runtime/atomic_amd64.s`](../../../.dev/reference/go/src/runtime/atomic_amd64.s) (and its per-arch siblings) because a portable compiler can't. **Memory barriers.** `MFENCE`, `LFENCE`, `SFENCE` on x86; `DMB` / `DSB` / `ISB` on ARM. Used by Go's `publicationBarrier`, `procyield`, and friends. Per-arch asm files carry them. -**Optimised `memmove` / `memequal` / `memclr`.** The compiler knows how to emit `rep movsb`, but a runtime sometimes ships a *better* version than the compiler's — wider vector loads, prefetch hints, alignment-aware loops. Go ships its own in [`asm_amd64.s`](../../../reference/go/src/runtime/asm_amd64.s) using AVX/SSE paths. +**Optimised `memmove` / `memequal` / `memclr`.** The compiler knows how to emit `rep movsb`, but a runtime sometimes ships a *better* version than the compiler's — wider vector loads, prefetch hints, alignment-aware loops. Go ships its own in [`asm_amd64.s`](../../../.dev/reference/go/src/runtime/asm_amd64.s) using AVX/SSE paths. ## 3 — Syscall trampolines Every raw syscall to the kernel is an asm stub. The kernel expects arguments in specific registers (on x86_64: `rdi`, `rsi`, `rdx`, `r10`, `r8`, `r9`, with the syscall number in `rax`), a `syscall` instruction, and return-value unpacking from `rax` (including `-errno` convention). A high-level language's calling convention doesn't match that layout — you need a thin asm wrapper per syscall. -See [`reference/go/src/runtime/sys_linux_amd64.s`](../../../reference/go/src/runtime/sys_linux_amd64.s) — 43 `TEXT` functions, one per syscall family: `runtime·write`, `runtime·read`, `runtime·futex`, `runtime·clone`, `runtime·rt_sigaction`, `runtime·rt_sigprocmask`, `runtime·rt_sigreturn`, `runtime·sched_yield`, `runtime·mmap`, `runtime·munmap`, `runtime·madvise`, `runtime·epollcreate1`, `runtime·epollctl`, `runtime·epollwait`, etc. +See [`.dev/reference/go/src/runtime/sys_linux_amd64.s`](../../../.dev/reference/go/src/runtime/sys_linux_amd64.s) — 43 `TEXT` functions, one per syscall family: `runtime·write`, `runtime·read`, `runtime·futex`, `runtime·clone`, `runtime·rt_sigaction`, `runtime·rt_sigprocmask`, `runtime·rt_sigreturn`, `runtime·sched_yield`, `runtime·mmap`, `runtime·munmap`, `runtime·madvise`, `runtime·epollcreate1`, `runtime·epollctl`, `runtime·epollwait`, etc. Go does these in asm because it cannot rely on libc — Go's scheduler needs to enter/exit syscalls at exactly controlled points (`runtime·entersyscall`, `runtime·exitsyscall`) so the M (OS thread) can be parked or reused without losing the goroutine. Going through `libc::write` would sidestep the scheduler's accounting. diff --git a/docs/plan/exploration/assembly/01-go-runtime-asm.md b/docs/plan/exploration/assembly/01-go-runtime-asm.md index e6b3a9f..e3c737f 100644 --- a/docs/plan/exploration/assembly/01-go-runtime-asm.md +++ b/docs/plan/exploration/assembly/01-go-runtime-asm.md @@ -2,11 +2,11 @@ The Go runtime ships ~72 `TEXT` functions in `asm_amd64.s` alone, ~43 in `sys_linux_amd64.s`, and per-architecture variants of both for `386`, `arm`, `arm64`, `loong64`, `mips(64)x`, `ppc64x`, `riscv64`, `s390x`, `wasm`. This doc inventories them by purpose so a reader can map each Go asm concern to the writeonce equivalent (spoiler: usually "Rust stdlib does it"). Follow-on reading: [`02-writeonce-stance.md`](./02-writeonce-stance.md). -All paths are inside [`reference/go/src/runtime/`](../../../reference/go/src/runtime/). +All paths are inside [`.dev/reference/go/src/runtime/`](../../../.dev/reference/go/src/runtime/). ## Scheduler & stack switching — `asm_.s` -One file per arch, everything that has to break Go's calling convention. The x86_64 version lives at [`asm_amd64.s`](../../../reference/go/src/runtime/asm_amd64.s). +One file per arch, everything that has to break Go's calling convention. The x86_64 version lives at [`asm_amd64.s`](../../../.dev/reference/go/src/runtime/asm_amd64.s). | Go symbol | What | | --- | --- | @@ -26,7 +26,7 @@ One file per arch, everything that has to break Go's calling convention. The x86 ## Atomics & barriers — `internal/runtime/atomic/atomic_.s` -Lives at [`internal/runtime/atomic/atomic_amd64.s`](../../../reference/go/src/internal/runtime/atomic/atomic_amd64.s) (and arch variants). Wrappers around arch-specific instructions: +Lives at [`internal/runtime/atomic/atomic_amd64.s`](../../../.dev/reference/go/src/internal/runtime/atomic/atomic_amd64.s) (and arch variants). Wrappers around arch-specific instructions: | Go symbol | x86 instruction | Purpose | | --- | --- | --- | @@ -41,7 +41,7 @@ Lives at [`internal/runtime/atomic/atomic_amd64.s`](../../../reference/go/src/in ## Syscall trampolines — `sys__.s` -On Linux-x86_64 that's [`sys_linux_amd64.s`](../../../reference/go/src/runtime/sys_linux_amd64.s) — 43 `TEXT` functions. Each is a short wrapper: move args into the kernel's register layout, execute `SYSCALL`, convert `rax` into a Go return value + error. +On Linux-x86_64 that's [`sys_linux_amd64.s`](../../../.dev/reference/go/src/runtime/sys_linux_amd64.s) — 43 `TEXT` functions. Each is a short wrapper: move args into the kernel's register layout, execute `SYSCALL`, convert `rax` into a Go return value + error. | Go symbol | Linux syscall | | --- | --- | @@ -70,7 +70,7 @@ On Linux-x86_64 that's [`sys_linux_amd64.s`](../../../reference/go/src/runtime/s ## Cgo bridge — `cgo__.s` -Files like [`cgo/asm_amd64.s`](../../../reference/go/src/runtime/cgo/asm_amd64.s). Machine-code marshalling between Go's register convention and C's SysV AMD64 ABI. Needed because Go's calling convention uses stack slots differently from C's register passing. +Files like [`cgo/asm_amd64.s`](../../../.dev/reference/go/src/runtime/cgo/asm_amd64.s). Machine-code marshalling between Go's register convention and C's SysV AMD64 ABI. Needed because Go's calling convention uses stack slots differently from C's register passing. **Writeonce doesn't cross language boundaries** — Rust is the only language in the binary; `libc` is already in Rust's register convention via `extern "C"`. No cgo bridge needed. diff --git a/docs/plan/exploration/blue-green-vm/00-vision.md b/docs/plan/exploration/blue-green-vm/00-vision.md new file mode 100644 index 0000000..27faee5 --- /dev/null +++ b/docs/plan/exploration/blue-green-vm/00-vision.md @@ -0,0 +1,180 @@ +# Blue/Green VMs — a self-hosting, agent-managed runtime + +> **Partially superseded (2026-08-03):** the deployment subsystem (§5–§6 here) +> is now specified in +> [`docs/superpowers/specs/2026-08-03-blue-green-vm-design.md`](../../../superpowers/specs/2026-08-03-blue-green-vm-design.md) +> — developer + `wo` CLI as the management client (agent/MCP becomes a later +> wrapper), schema migration folded into the approval step (additive-only +> auto-diff in v1), fixed slots with alternating activity, HTTP+JSON+SSE. +> §1–§4 (transports, recipe box, fibers, source-in-binary) remain current +> thinking feeding plans 3/4/6. + +> Thought-process capture (2026-08-02). Not a phase plan yet — the vision that +> shapes how the wovm runtime grows past milestone 1, recorded before the +> details harden. Related: the OOP spec +> ([`../../../superpowers/specs/2026-08-01-oop-compiler-vm-design.md`](../../../superpowers/specs/2026-08-01-oop-compiler-vm-design.md)), +> plan 4 (shard-actor runtime), plan 15 (MCP streamable HTTP), and the +> single-binary trailer already shipped by `woc build`. + +## The idea, in five sentences + +The writeonce executable is a **systemd service that never stops**. It embeds +its own **source code**, not just its bytecode. An external **Claude agent** +reads and edits that source through a managed channel; an approved change is +compiled **inside the runtime** and loaded into the idle VM slot. The runtime +holds **exactly two VMs — Blue (active) and Green (previous version)** — and +deployment is an atomic switch between them. Rollback is the same switch in +reverse, because the previous version never left memory. + +## 1. A runtime is not a port + +The runtime core is the VM pair + engine + scheduler — it must run with zero +listeners. Ports are **transports**, attached at boot like modules: an HTTP +listener, a unix socket, an MCP endpoint, stdio. Consequences: + +- The same binary serves as web app, CLI batch runner, or agent-managed + service depending on which transports the deployment attaches — one of the + recipes a custom web framework builds from (§2). +- **systemd socket activation** fits exactly: the unit owns the socket + (`LISTEN_FDS`), the runtime accepts on whatever fds it inherits. The + "always running" property (§6) and the "no port of its own" property come + from the same mechanism. + +## 2. The runtime is a recipe box for web frameworks + +Everything a custom web framework needs in later phases must exist as a +separable runtime capability, not a monolith: transports (§1), fibers (§3), +routing surface (plan 6), the subscription registry (plan 7), the DB engine +(plan 5), and the deploy/rollback machinery (§5). A "framework" in a later +phase is a `.wo` library that composes these recipes — the runtime itself +stays framework-agnostic. + +## 3. Fibers (green threads) + +Concurrency inside a shard is **cooperative fibers scheduled by the VM**, not +OS threads — the Erlang shape on the wovm substrate: + +- A fiber is exactly the execution state `wo_vm` already isolates: a register + window stack + frame stack + a current pc. Making that state per-fiber + instead of per-VM turns the interpreter into a fiber scheduler almost for + free. +- **Preemption by reduction budget**: the dispatch loop decrements a counter + per instruction (or per call/back-edge); at zero, the fiber parks and the + scheduler picks the next runnable one. No signals, no stack switching + tricks, deterministic and debuggable. +- Fibers **park on I/O**: a blocked read hands the fd to the shard's event + loop (`wo-rt.c`'s epoll/io_uring machinery) and the fiber resumes when the + completion arrives. One OS thread per core (plan 4's shard), thousands of + fibers per shard. +- Fits the ownership model: a fiber is an actor mailbox owner; cross-fiber + sends follow the same ownership-move rule as cross-shard sends. + +## 4. The binary contains its source + +`woc build` already appends the `.wob` image to a copy of `wovm` with an +offset trailer. The trailer grows one more section: **the `.wo` source tree** +(paths + contents, compressed). Why: + +- The deployed artifact is self-describing — no "which commit is prod + running?" class of question. `wovm --dump-source` can always reproduce + exactly what is executing. +- The agent workflow (§5) needs a source of truth that travels with the + binary, not a checkout that can drift from it. +- After a deployment, the runtime rewrites its own source section (write to + temp, fsync, rename) so the artifact on disk always matches the Blue VM. + +## 5. Agent-managed source — how Claude fits + +The runtime exposes a **management transport** (MCP over streamable HTTP — +plan 15's machinery, localhost + bearer token, the log-watcher posture). +Claude Code connects as an MCP client. Tools the runtime serves: + +| Tool | What it does | +| --- | --- | +| `source_list` / `source_read` | browse the embedded source tree of the running (Blue) version | +| `source_propose` | submit a changed file set as a **proposal** — staged, never applied | +| `proposal_diff` | render the pending proposal against Blue's source | +| `proposal_check` | run `woc check` on the proposal inside the runtime — diagnostics come back to the agent | +| `proposal_approve` | **human-only gate** (separate credential or out-of-band confirmation) — approval triggers compile + green-slot load | +| `deploy_switch` | atomic Blue↔Green switch after health checks | +| `deploy_rollback` | the same switch back — Green still holds the previous version | +| `deploy_status` | which version is Blue, which is Green, in-flight drain state | + +Properties worth pinning now: + +- **The agent proposes; a human approves.** `proposal_approve` is not + reachable with the agent's token. Approval is the compile trigger, not the + edit. +- **Every step is WAL-logged** — proposals, diagnostics, approvals, switches, + rollbacks form an audit trail that survives crashes like any other commit. +- **The compiler lives with the runtime** for this loop to work: either + `woc` embedded in the binary (adds OCaml runtime weight) or shipped beside + it in the service directory (lighter; the systemd unit owns both files). + Open question in §8 — start with "beside it". + +## 6. Blue/Green VM lifecycle + +Exactly **two VM slots** per runtime, never more: + +- **Blue** — the active VM: all new requests/fibers dispatch into it. +- **Green** — the previous version, loaded and warm: the instant-rollback + target. After a successful deploy the roles swap; the old Blue becomes the + new Green. + +The critical separation: **VMs own code, the engine owns data.** Tables, +WAL, subscriptions, and the arena slabs live in the engine layer beneath both +VMs; a switch swaps which bytecode handles requests, never the data. That is +what makes the switch cheap and rollback safe — no state migration on the +happy path (and schema changes are exactly the hard part, §8). + +Deploy sequence: + +1. Approved proposal compiles (`woc emit`) — failure ends the deploy, + Blue untouched. +2. New image loads + validates into the idle slot (loader is the same + validation battery as always — a bad image cannot boot). +3. Health gate: entry smoke / conformance subset runs against the idle VM. +4. **Switch at the dispatch boundary**: new work enters the new Blue; + in-flight fibers on the old VM drain to completion (bounded timeout). +5. Old Blue becomes Green (rollback target); the binary's source section is + rewritten to match (§4). +6. `deploy_rollback` at any later point is step 4 in reverse — no compile, + no load, the code is already resident. + +## 7. Always running + +The executable maps to a **systemd service**: `Restart=always`, socket +activation for the transports (§1), the hardening posture proven in the +log-watcher units (unprivileged user, read-only system, `StateDirectory` +for WAL/data). Deployment never restarts the unit — that is the whole point +of the VM pair. The unit restarting (crash, host reboot) boots Blue from the +binary's current source/bytecode section and reloads Green only when the +next deploy happens. + +## 8. Open questions (deliberately unresolved here) + +1. **Schema migrations.** Code switches atomically; data does not. A + proposal that changes a class's fields needs a migration story between + Green-shaped and Blue-shaped rows — the wo-seg migration doc's + dual-write thinking applies inside one process. Hardest problem in this + vision; needs its own exploration. +2. **Live subscriptions across a switch.** Do WebSocket subscribers survive + a deploy (registry lives in the engine layer → yes, by design), and what + do they see mid-drain? +3. **`woc` placement** — beside the binary vs embedded (§5). +4. **Fiber preemption granularity** — per-instruction counter vs + call/back-edge only (cheaper, coarser). +5. **Does Green count against the heap budget** (two arenas resident) or + does Green hibernate (bytecode resident, heap lazily rebuilt on + rollback)? + +## 9. Where this lands in the plan sequence + +- Fibers (§3): extends **plan 4** (shard-actor runtime) — same scheduler + work, one more scheduling unit. +- Transports-not-ports (§1): shapes **plan 6** (HTTP/service layer) — the + listener becomes one attachable transport among several. +- Management MCP (§5): builds on **plan 15**'s streamable-HTTP machinery. +- Source-in-binary (§4): extends plan 3's `woc build` trailer. +- Blue/Green switch (§6) + agent loop (§5): a new phase after those land — + needs spec + plan of its own once this vision stabilizes. diff --git a/docs/plan/exploration/c-runtime/00-plan.md b/docs/plan/exploration/c-runtime/00-plan.md index e51655f..e7bec73 100644 --- a/docs/plan/exploration/c-runtime/00-plan.md +++ b/docs/plan/exploration/c-runtime/00-plan.md @@ -2,11 +2,11 @@ > **Kanban: ✅ done** — phases A–F all shipped with measured exit evidence below. Board: [../../00-kanban.md](../../00-kanban.md) -**Context sources:** [`prototypes/wo-rt-c/wo-rt.c`](../../../../prototypes/wo-rt-c/wo-rt.c) (phase 0 — the single-threaded epoll baseline), [`../../09-concurrency-scaleout.md`](../../09-concurrency-scaleout.md) (the thread-per-core doctrine every phase here miniaturizes), [`../../10-storage-foundations.md`](../../10-storage-foundations.md) / [`11-wal-and-recovery.md`](../../11-wal-and-recovery.md) / [`12-engine-disk-cutover.md`](../../12-engine-disk-cutover.md) (the storage track), kernel reference cards [`../linux/07-io_uring.md`](../linux/07-io_uring.md), [`08-mmap.md`](../linux/08-mmap.md), [`09-fallocate.md`](../linux/09-fallocate.md), [`12-pwrite-fsync.md`](../linux/12-pwrite-fsync.md), [`02-eventfd.md`](../linux/02-eventfd.md). +**Context sources:** [`runtime/wo-rt.c`](../../../../runtime/wo-rt.c) (phase 0 — the single-threaded epoll baseline), [`../../09-concurrency-scaleout.md`](../../09-concurrency-scaleout.md) (the thread-per-core doctrine every phase here miniaturizes), [`../../10-storage-foundations.md`](../../10-storage-foundations.md) / [`11-wal-and-recovery.md`](../../11-wal-and-recovery.md) / [`12-engine-disk-cutover.md`](../../12-engine-disk-cutover.md) (the storage track), kernel reference cards [`../linux/07-io_uring.md`](../linux/07-io_uring.md), [`08-mmap.md`](../linux/08-mmap.md), [`09-fallocate.md`](../linux/09-fallocate.md), [`12-pwrite-fsync.md`](../linux/12-pwrite-fsync.md), [`02-eventfd.md`](../linux/02-eventfd.md). ## Goal -Evolve the [`prototypes/wo-rt-c/`](../../../../prototypes/wo-rt-c/) prototype from a single-threaded epoll reference into a **multi-threaded runtime environment for writeonce applications**: thread-per-core io_uring event loops at million-scale read/write concurrency, the whole database resident in RAM (one mmap arena, addressed per shard — no duplication), **ACID** commits that dual-write RAM-first-then-disk, and a boot path that loads the hard drive's state back into RAM before serving. Still one C file's worth of honesty per concern, still **zero dependencies beyond libc** — raw io_uring syscalls, no liburing. +Evolve the [`runtime/`](../../../../runtime/) prototype from a single-threaded epoll reference into a **multi-threaded runtime environment for writeonce applications**: thread-per-core io_uring event loops at million-scale read/write concurrency, the whole database resident in RAM (one mmap arena, addressed per shard — no duplication), **ACID** commits that dual-write RAM-first-then-disk, and a boot path that loads the hard drive's state back into RAM before serving. Still one C file's worth of honesty per concern, still **zero dependencies beyond libc** — raw io_uring syscalls, no liburing. Each phase is the executable proving ground for the matching Rust plan (09–12): get the syscall sequence right here in a few hundred lines, then port with confidence. @@ -66,9 +66,9 @@ Boot, before any listener opens: each thread replays its own WAL into its arena ### Phase F — million-scale harness + ACID verification — ✅ shipped *Maps to [plan 09's verification-targets table](../../09-concurrency-scaleout.md).* -`setrlimit(RLIMIT_NOFILE)` raised at boot. A small C load client under `prototypes/wo-rt-c/bench/` (keep-alive, pipelined GETs, latency timestamps — `wrk` would be an external dep). Measure honestly on the dev box and commit the numbers to the prototype README: aggregate read req/s across cores (goal order 10⁶/s on 8–16 cores), concurrent open connections (goal order 10⁵–10⁶; ~8 KB/conn + fd limits are the ceiling), commits/s under group fsync, p99 read latency under write load. ACID scripts: torn-WAL injection (atomicity), single-shard interleaving probe (isolation), the phase-D crash test under load (durability). A `just rt-c-bench` recipe runs it all. +`setrlimit(RLIMIT_NOFILE)` raised at boot. A small C load client under `runtime/bench/` (keep-alive, pipelined GETs, latency timestamps — `wrk` would be an external dep). Measure honestly on the dev box and commit the numbers to the prototype README: aggregate read req/s across cores (goal order 10⁶/s on 8–16 cores), concurrent open connections (goal order 10⁵–10⁶; ~8 KB/conn + fd limits are the ceiling), commits/s under group fsync, p99 read latency under write load. ACID scripts: torn-WAL injection (atomicity), single-shard interleaving probe (isolation), the phase-D crash test under load (durability). A `just rt-c-bench` recipe runs it all. -**Exit (met):** measured on a 20-core box (table in the [prototype README](../../../../prototypes/wo-rt-c/README.md)): **908,916 reads/s p99 154 µs and 643,250 fsync-acked commits/s p99 177 µs** on 8 shards — vs Go `net/http` on 20 cores at 495k/355k with ~8× worse p99 and no durability (.NET unavailable on the box); 10k idle connections, 0 errors; only 2xx counted (the client tracks status codes). **The crash-under-load test found two real durability bugs the phase-D test missed** — an ack-armed-before-fsync race in `conn_continue` (route parks the response *during* `try_process`; the pre-check missed it) and an fd-reuse ABA hazard in batch ack-parking (fixed with per-connection generation stamps). After both fixes, three `kill -9`-mid-bench rounds at ~1–2M commits each showed **WAL records ≥ acked, every round** (one exact). Isolation: 300 concurrent commits → 300 distinct interleaved ids. Geometry scaling via `-DSLOTS_PER_SHARD` (bitmap region generalized to multi-page); 512 MB arena verified mlocked. +**Exit (met):** measured on a 20-core box (table in the [prototype README](../../../../runtime/README.md)): **908,916 reads/s p99 154 µs and 643,250 fsync-acked commits/s p99 177 µs** on 8 shards — vs Go `net/http` on 20 cores at 495k/355k with ~8× worse p99 and no durability (.NET unavailable on the box); 10k idle connections, 0 errors; only 2xx counted (the client tracks status codes). **The crash-under-load test found two real durability bugs the phase-D test missed** — an ack-armed-before-fsync race in `conn_continue` (route parks the response *during* `try_process`; the pre-check missed it) and an fd-reuse ABA hazard in batch ack-parking (fixed with per-connection generation stamps). After both fixes, three `kill -9`-mid-bench rounds at ~1–2M commits each showed **WAL records ≥ acked, every round** (one exact). Isolation: 300 concurrent commits → 300 distinct interleaved ids. Geometry scaling via `-DSLOTS_PER_SHARD` (bitmap region generalized to multi-page); 512 MB arena verified mlocked. ## Non-scope @@ -82,7 +82,7 @@ Boot, before any listener opens: each thread replays its own WAL into its arena - [`../../09-concurrency-scaleout.md`](../../09-concurrency-scaleout.md) — the doctrine; this prototype is its executable proving ground (A↔09a, C↔09 decision 4, D↔09c). - [`../../10-storage-foundations.md`](../../10-storage-foundations.md), [`11-wal-and-recovery.md`](../../11-wal-and-recovery.md), [`12-engine-disk-cutover.md`](../../12-engine-disk-cutover.md) — the storage track phases B/D/E miniaturize. -- [`../../../../prototypes/wo-rt-c/README.md`](../../../../prototypes/wo-rt-c/README.md) — current state and module map (phase 0). +- [`../../../../runtime/README.md`](../../../../runtime/README.md) — current state and module map (phase 0). - [`./01-architecture.md`](./01-architecture.md) — the target architecture traced through one memory address at million-connection concurrency, plus improvement proposals (seqlock reads, registered buffers, SEND_ZC, SQPOLL) that slot into phases C/F. - [`./02-single-binary.md`](./02-single-binary.md) — the end goal: how the `wo build` single binary runs on this runtime environment (Go model, not JVM — the kernel is statically linked into every app; the embedding contract between compiler payload and runtime kernel). - [`../../../../prototypes/wo-db/`](../../../../prototypes/wo-db/) — the query-layer sibling; one day a phase-G could splice its engine on top of this runtime. diff --git a/docs/plan/exploration/c-runtime/01-architecture.md b/docs/plan/exploration/c-runtime/01-architecture.md index 01867ed..b5206e3 100644 --- a/docs/plan/exploration/c-runtime/01-architecture.md +++ b/docs/plan/exploration/c-runtime/01-architecture.md @@ -1,6 +1,6 @@ # wo-rt-c architecture — one memory address, two spaces, a million connections -This document defines the runtime's architecture by following **one memory address** through user space, kernel space, and hardware, under a million connections reading and writing it concurrently — then suggests improvements. Companion docs: [`00-plan.md`](./00-plan.md) (the phases that build this), [`README.md`](../../../../prototypes/wo-rt-c/README.md) (phase-0 module map). +This document defines the runtime's architecture by following **one memory address** through user space, kernel space, and hardware, under a million connections reading and writing it concurrently — then suggests improvements. Companion docs: [`00-plan.md`](./00-plan.md) (the phases that build this), [`README.md`](../../../../runtime/README.md) (phase-0 module map). ## The cast: one address diff --git a/docs/plan/exploration/colibri/00-colibri-and-mixtral.md b/docs/plan/exploration/colibri/00-colibri-and-mixtral.md new file mode 100644 index 0000000..8943dd4 --- /dev/null +++ b/docs/plan/exploration/colibri/00-colibri-and-mixtral.md @@ -0,0 +1,213 @@ +# Colibrì — reference analysis, and running Mistral's MoE models locally + +Analysis of the vendored reference tree at [`.dev/reference/colibri/`](../../../../.dev/reference/colibri) (Apache-2.0, upstream ), and a grounded, hands-on answer to the follow-on question: **what does it take to run Mistral's Mixture-of-Experts models (Mixtral 8x22B / 8x7B) locally?** — including a working demonstration of colibrì's "dense resident, stream the experts from disk" idea using the vendored llama.cpp (§7). + +> **TL;DR** +> - Colibrì is a **single-file, zero-dependency C inference engine** that runs a **744B-parameter MoE (GLM-5.2)** on a ~25 GB-RAM consumer box by **streaming routed experts from disk** and treating VRAM/RAM/disk as one managed memory hierarchy. It is here as a *runtime-engineering* reference: it does all its I/O with the exact kernel primitives writeonce's north star is built on (`pread`, `posix_fadvise`, `io_uring`, `mmap`, `mlock`, `O_DIRECT`). +> - Colibrì supports **exactly two model architectures today: GLM-5.2 (`c/glm.c`) and OLMoE (`c/olmoe.c`)**. **There is no Mixtral/Mistral code in the tree** (`grep -ri mixtral` → 0 hits). +> - **To run any Mistral MoE locally right now, don't wait on colibrì** — use a runtime that already supports it. The repo now also vendors a full **llama.cpp** checkout at `.dev/reference/llama-cpp` with **verified, first-class Mistral/Mixtral support** (§6): GGUF + `--n-cpu-moe`. Other options: **KTransformers** (CPU/GPU hybrid, the closest philosophical cousin) or **vLLM/SGLang** on a multi-GPU box. See §5–§6. +> - Mixtral 8x22B is actually a **much easier** target for the colibrì streaming trick than GLM-5.2 — 8 coarse experts/layer instead of 256 fine-grained ones, so the whole int4 expert set (~67 GB) fits in commodity RAM and the disk-streaming stops mattering. A `mixtral.c` port modelled on `olmoe.c` is small and plausible (§4), but it does not exist yet. +> - You can **reproduce and observe** colibrì's core mechanism on a small machine with the [`prototypes/llama-moe-stream/`](../../../../prototypes/llama-moe-stream) demo (§7): run an MoE (default **Qwen3-Coder-30B-A3B**) under a `MemoryMax` cap so the small dense part stays resident while the experts stream from disk on demand — the model still answers correctly on far less RAM than its size. **Gotcha found in practice:** in-circulation Mixtral GGUFs use the pre-2024 per-expert layout and **won't load** on current llama.cpp, so the demo uses a modern fused-format MoE. + +--- + +## 1. What colibrì is + +**"Tiny engine, immense model."** Colibrì is a lightweight, quality-preserving Mixture-of-Experts *inference runtime* written in pure C with no external libraries (no BLAS, no Python at runtime, no GPU required). Its thesis: + +> A 744B MoE activates only ~40B params per token, and only ~11 GB of those (the *routed experts*) change from token to token. So keep the **dense part resident** and **stream the experts from disk on demand.** + +Concretely, for GLM-5.2 at int4: + +| Component | Size | Placement | +|---|---|---| +| Dense (attention, shared experts, embeddings — ~17B params) | ~9.9 GB | **resident in RAM** at int4 | +| 19,456 routed experts (75 MoE layers × 256 + MTP head, ~19 MB each) | ~370 GB | **on disk**, streamed on demand | + +The engine treats **VRAM → RAM → disk as one memory hierarchy** with a per-layer LRU expert cache, an optional pinned hot-store (the hottest experts stay in spare RAM/VRAM), and the OS page cache as a free L2. Insufficient fast memory reduces *speed*, never *precision or router semantics* — the default policy is lossless. + +This is not fast (0.05–2 tok/s depending on disk/RAM/CPU — see the community benchmark table in the upstream README), but it runs a **frontier-class 744B model correctly on hardware that costs less than one H100 fan.** + +## 2. Why it lives in `.dev/reference/` + +writeonce's north star (see root `CLAUDE.md`, `docs/01-problem.md`, `docs/plan/linux/00-linux.md`) is **one binary, zero external crates, all I/O driven directly by Linux kernel primitives.** Colibrì is a working, production-shaped proof of exactly that discipline in a different domain (ML inference rather than a database): + +- **One binary, `libc`-only.** The engine is `c/glm.c` (~348 KB) plus small headers. Python appears *only* in the one-time offline weight converter, never at runtime — the same "transitional tooling is allowed, the runtime is not" line writeonce draws. +- **The kernel *is* the async runtime and the storage tier.** Colibrì's expert streaming is built from the same primitives `crates/rt/src/runtime/` is being built on: + + | Primitive | Colibrì use | writeonce analogue | + |---|---|---| + | `pread` | read one expert slab at a known offset | WAL / segment reads | + | `posix_fadvise(WILLNEED/DONTNEED)` | async readahead of the next expert block; evict used slabs | page-cache management | + | `io_uring` (`URING=1`, `c/uring.h`) | batched, queued cold expert reads via `IOSQE_ASYNC` | the target event loop (`docs/plan/02`) | + | `O_DIRECT` (`DIRECT=1`) | bypass page cache for sustained NVMe | direct segment I/O | + | `mmap` (`COLI_MMAP=1`) | map weights instead of `read()` into slabs | `sendfile`/mmap static assets (`docs/plan/08`) | + | `mlock` (`MLOCK=1`) | wire the hot expert cache into physical RAM | pinning hot pages | + + It even has a portability story writeonce will need: `c/compat.h` maps every POSIX call to the Win32 API (`pread`→`ReadFile`+`OVERLAPPED`, etc.) so the engine source stays platform-clean. + +So colibrì is a reference for **how to engineer a disk/RAM/VRAM memory hierarchy on raw syscalls in one C binary** — read `c/uring.h`, `c/tier.h`, `c/st.h` (the safetensors mmap reader), and `c/compat.h` when designing writeonce's I/O layer. It is *not* a database and shares no code; the value is the technique. + +## 3. Use case — who runs colibrì, and when + +**Use it when:** you want to run a *very large* open-weight MoE (hundreds of billions of params) **locally, offline, at full quality**, on hardware that cannot hold the model in VRAM (or even in RAM), and you can tolerate low-but-usable token rates. Typical: a single workstation or a homelab NVMe box, privacy-sensitive or air-gapped inference, model-behaviour research, or squeezing a frontier model onto a laptop. + +**Don't use it when:** you need interactive throughput on a small model (llama.cpp/Ollama are simpler and faster there), or you have enough VRAM to hold your model outright (use vLLM/SGLang/ExLlamaV2). + +**Surface area** (all via the `coli` Python CLI, which just sets env vars and launches the C engine): + +| `coli ` | What it does | +|---|---| +| `convert` | offline FP8→int4 converter; downloads the HF checkpoint one ~5 GB shard at a time so the full 756 GB never lands on disk at once (resumable) | +| `plan` | read-only: reports the dense/expert footprint and the planned VRAM/RAM/disk tiers (`--json`) | +| `doctor` | read-only readiness check (model dir, tokenizer, RAM budget, CUDA linkage, GPU devices) | +| `chat` | interactive REPL | +| `run` | one-shot prompt | +| `serve` | OpenAI-compatible HTTP API (`/v1/chat/completions`, SSE streaming) — stdlib-only gateway (`c/openai_server.py`), one model process, FIFO admission queue | +| `web` | serves the React dashboard in `web/` (live token metrics, hardware panel, the "Brain" expert-heat view) | +| `bench` | MMLU/HellaSwag/ARC quality benchmarks | + +Also shipped: a **Tauri desktop shell** (`desktop/`) and a **Nix flake** (`flake.nix`, gcc + OpenMP + gmp; Python env for the converter only). + +**Runtime environment colibrì itself needs:** +- **OS:** Linux (or WSL2), macOS, or native Windows 11 (MinGW-w64). +- **CPU:** gcc with OpenMP; AVX2 baseline (`x86-64-v3`), with faster paths on AVX-VNNI (Alder Lake+) and ARM NEON/i8mm/SVE2 (Apple Silicon, Grace). `make ARCH=native` enables the best kernel for the host. +- **GPU (optional):** CUDA backend for NVIDIA (resident/pinned expert tier; on Windows a runtime-loaded `coli_cuda.dll`), Metal backend for Apple Silicon. Both are opt-in accelerators — the CPU path is the reference and stays byte-exact. +- **RAM:** ≥16 GB minimum; more RAM = more experts stay hot = higher tok/s (auto-budgeted from `MemAvailable`). +- **Disk:** the int4 model on a **local** NVMe (ext4/NTFS — never a network/9p mount). Random-read bandwidth is the cold-decode ceiling. + +Feature depth worth noting (all in `c/glm.c`): MLA attention with a 57×-compressed KV cache, DeepSeek-V3-style sigmoid router, native **MTP speculative decoding** (int8 draft head), grammar-forced drafts (`GRAMMAR=*.gbnf`), int8/int4/int2 packed quant kernels, DSA sparse attention, crash-safe KV-cache persistence, and cache-aware routing. Every knob is an env var — see [`.dev/reference/colibri/docs/ENVIRONMENT.md`](../../../../.dev/reference/colibri/docs/ENVIRONMENT.md). + +## 4. The Mixtral gap — and what a port would take + +**Colibrì does not support Mixtral / any Mistral model.** The only architectures implemented are: + +- **`c/glm.c`** — GLM-5.2 (`glm_moe_dsa`): 744B, 256 experts/layer top-8, MLA, DSA, MTP. The flagship target. +- **`c/olmoe.c`** — OLMoE-1B-7B (`allenai/OLMoE-1B-7B-0125-Instruct`): 7B total / 1B active, 64 experts/layer top-8. Its header states its purpose plainly: *"validate the streaming core before scaling to GLM-5.2."* **This is the template for adding a new architecture.** + +Adding Mixtral would mean writing the same two pieces OLMoE has: + +1. **`c/mixtral.c`** — a faithful forward pass. Good news: Mixtral is *architecturally simpler* than either existing engine — plain GQA + RoPE attention (no MLA, no DSA, no q/k-norm), RMSNorm, SwiGLU experts, no shared expert, no MTP head. It is closer to `olmoe.c` than to `glm.c`, and smaller. +2. **`c/tools/convert_mixtral.py`** — modelled on `convert_olmoe.py`: keep dense weights as f16/f32, row-wise-quantize the expert matrices to the int8/int4 container. Only the expert-key regex changes — Mixtral names them `model.layers.{L}.block_sparse_moe.experts.{E}.(w1|w2|w3).weight` and the router is `block_sparse_moe.gate`. + +**Why Mixtral is an *easier* streaming target than GLM-5.2** (int4, from its config — 56 layers, hidden 6144, intermediate 16384, 8 experts/layer, top-2): + +- Each expert = 3 matrices of 6144×16384 ≈ 302M params → **~151 MB at int4** (vs GLM's 19 MB fine-grained experts). +- Total experts = 8 × 56 = **448 experts ≈ 67 GB at int4** (vs GLM's 19,456 experts ≈ 370 GB). +- Cold cost/token = top-2 × 56 = **112 expert-loads ≈ 17 GB/token** — but with only 8 experts/layer, **any 96 GB+ machine caches the entire expert set in RAM**, giving ~100 % hit rate and *zero* disk streaming after warmup. The engine becomes RAM-bandwidth / matmul bound, not disk bound. + +In other words, the whole "stream from disk" apparatus that colibrì needs for GLM-5.2 is mostly *unnecessary* for Mixtral 8x22B — the model is small enough (at int4) to just live in RAM. That is exactly why the practical answer below does not require colibrì at all. + +## 5. Running Mixtral 8x22B locally — the ready paths + +### 5.1 The model + +| Config (`Mixtral-8x22B-v0.1`) | Value | +|---|---| +| Total / active params | ~141B / ~39B | +| Layers | 56 | +| hidden_size | 6144 | +| intermediate_size (per expert) | 16384 | +| attention heads / KV heads (GQA) | 48 / 8 (head_dim 128) | +| experts / top-k | 8 / 2 | +| vocab | 32768 | +| rope_theta / context | 1,000,000 / 65,536 | + +Approximate on-disk sizes (GGUF): **FP16 ≈ 281 GB · Q8_0 ≈ 149 GB · Q5_K_M ≈ 100 GB · Q4_K_M ≈ 86 GB · Q3_K ≈ 65 GB · Q2_K ≈ 52 GB.** For decent-quality local use, **Q4_K_M (~86 GB) or Q5** is the sweet spot; Q2/Q3 fit smaller boxes with quality loss. + +### 5.2 Runtime options, from most-consumer to most-datacenter + +| Runtime | How it runs Mixtral 8x22B locally | Hardware reality | Closeness to colibrì | +|---|---|---|---| +| **Ollama** | `ollama run mixtral:8x22b` (wraps llama.cpp, pulls a Q4 GGUF) | ~90 GB RAM for Q4 CPU-only, or GPU+CPU split | Same tiering idea, turnkey | +| **llama.cpp (GGUF)** | Load a Q4/Q5 GGUF; offload expert layers to CPU RAM and keep attention/dense on GPU with **`--n-cpu-moe N`** (or `-ot`/`--override-tensor` regex for per-tensor control) | Runs CPU-only with ~90 GB RAM, *or* a 16–24 GB GPU + system RAM hybrid | **Closest mainstream analog** — same "experts in slow memory, dense on fast" split colibrì automates | +| **KTransformers** | CPU/GPU **hybrid MoE** — attention + shared/hot experts on GPU, the parameter-heavy routed experts in system RAM with AMX/AVX-512 CPU kernels. Explicitly lists **Mixtral 8x7B and 8x22B** as supported. | One consumer GPU + a big-RAM host; higher throughput than llama.cpp on large MoE | **Philosophically closest** — it is colibrì's heterogeneous-tiering idea as a Python/CUDA framework | +| **vLLM / SGLang** | GPU-native, high-throughput serving (AWQ/GPTQ 4-bit or FP16) | Realistically **2× A100-80GB** (4-bit) to 4–8× for FP16 — a local *server*, not a desktop | Different niche (VRAM-resident, batch throughput) | +| **ExLlamaV2 (EXL2)** | 4-bit EXL2 quant, GPU-only | ~4× 24 GB consumer GPUs for a low-bpw quant | GPU-resident, no disk tier | +| **LM Studio / text-generation-webui** | Desktop front-ends over llama.cpp/GGUF | Same as llama.cpp | GUI convenience layer | + +### 5.3 Recommendation + +- **Single consumer/workstation box (one GPU + 64–128 GB RAM):** **llama.cpp or Ollama** with a **Q4_K_M GGUF** and **`--n-cpu-moe`** to push experts into RAM while attention stays on the GPU. Simplest and proven. If you have AMX/AVX-512 and want more speed on the same hardware, try **KTransformers** — it is the closest thing to "colibrì for Mixtral" that exists today. +- **Local multi-GPU server:** **vLLM or SGLang** with a 4-bit quant for real throughput. +- **If you specifically want the colibrì engine to run it:** that requires writing `c/mixtral.c` + `c/tools/convert_mixtral.py` against the `c/olmoe.c` template (§4). Feasible and not large, but it is net-new work — and because Mixtral's int4 expert set fits in RAM, it would buy little over the paths above except staying inside the pure-C, zero-dep runtime that makes colibrì interesting to writeonce in the first place. + +## 6. Verified: `.dev/reference/llama-cpp` already runs Mistral/Mixtral + +The repo also vendors a full, recent **llama.cpp** checkout at `.dev/reference/llama-cpp` (a symlink to a local clone; HEAD `635cdd5fc`). Unlike colibrì, it has **first-class Mistral/Mixtral support**, confirmed across the whole stack: + +- **Architecture** (`src/llama-arch.{h,cpp}`): Mistral 7B and Mixtral 8x7B/8x22B load under `LLM_ARCH_LLAMA` — llama-arch MoE, driven by the `expert_count` / `expert_used_count` GGUF keys. Dedicated `LLM_ARCH_MISTRAL3` / `LLM_ARCH_MISTRAL4` cover the newer Mistral Small / Mistral 4 families; Pixtral / Mistral-Small-3.1 handle the vision variants. +- **Conversion** (`conversion/` package — the refactored `convert_hf_to_gguf.py`): registers `MistralForCausalLM` / `MixtralForCausalLM` (→ llama arch), plus dedicated `MistralModel`, `MistralMoeModel` (remapped onto DeepSeek-V2), `Mistral3Model`, `Ministral3Model`, `Mistral4Model`, `PixtralModel`. +- **Tokenizer + chat templates**: native `mistral-common` (Tekken / SentencePiece) tokenizers, a `TEKKEN` pre-type, and five built-in templates — `mistral-v1`, `mistral-v3`, `mistral-v3-tekken`, `mistral-v7`, `mistral-v7-tekken` (`src/llama-chat.cpp`). +- **MoE-offload flags** (`common/arg.cpp`): `-cmoe`/`--cpu-moe` and `-ncmoe N`/`--n-cpu-moe N` — the colibrì-style "experts on the slow tier, dense on the fast tier" split, built in (with `--n-cpu-moe-draft` variants for speculative decoding). + +So on this repo the runnable path for any Mistral MoE is **llama.cpp**, not colibrì. Of the two vendored inference references: **colibrì = GLM-5.2 + OLMoE only; llama.cpp = full Mistral/Mixtral.** + +## 7. Hands-on: understand MoE experts, and stream them from disk + +The demo lives at [`prototypes/llama-moe-stream/`](../../../../prototypes/llama-moe-stream) (`run-moe.sh` + a teaching README). It runs an MoE and *forces* the streaming behavior with a RAM cap so the mechanism is observable — the same idea colibrì applies to GLM-5.2. + +> **Format-wall gotcha (found the hard way).** The demo originally targeted Mixtral 8x7B, but **every in-circulation Mixtral GGUF (TheBloke Dec-2023, MaziyarPanahi Feb-2024) uses the pre-2024 *per-expert* tensor layout** (`blk.0.ffn_gate.0.weight` … `.7.weight`). Current llama.cpp (HEAD `635cdd5fc`) only loads the **fused** layout (`blk.0.ffn_gate_exps.weight`) and dies with `missing tensor 'blk.0.ffn_down_exps.weight'`. Re-downloading another old quant does not help. So the demo defaults to **Qwen3-Coder-30B-A3B-Instruct** — a modern MoE whose GGUF is fused-format (verified), and which doubles as a capable local coding model. The Mixtral analysis in §1–§6 stands; only the *runnable demo* switched models. + +### 7.1 What a "Mixture of Experts" is (the concept) + +A **dense** transformer runs every weight for every token. An **MoE** replaces each layer's feed-forward block with **N expert FFNs + a small router**; per token the router routes through only the **top-k** experts, and the rest stay idle. That splits the weights into two classes — and the split is the whole point: + +| | what it is | touched per token? | share of the weights | +|---|---|---|---| +| **Dense part** | attention, embeddings, norms, the routers | **always** — every token, every layer | small → keep **resident** | +| **Experts** (routed) | the N expert FFNs in each layer | **only top-k of N** | the bulk → **stream from disk** | + +**Qwen3-Coder-30B-A3B** (the demo model): 30B total but only **~3.3B active per token** — the router fires a small top-k of many experts each layer. **Mixtral 8x7B** is the same idea at 46.7B total / ~12.9B active (32 layers, 8 experts/layer, top-2; the name misleads — experts share one attention stack, so it is 46.7B not 56B), and **Mixtral 8x22B** at 141B / ~39B active. + +**Why this enables streaming:** the dense part is small and hit constantly → keep it **resident** in fast memory. The experts are the majority of the bytes but each is hit rarely → they can live **on disk** and be pulled in exactly when routed to. A *dense* model of the same size could not do this (all of it every token); an MoE reads only the slice it routes to. That is colibrì's thesis. + +### 7.2 Realizing "dense resident, experts streamed" with llama.cpp + +- **mmap (on by default)** memory-maps the GGUF; the kernel demand-pages weights and evicts under pressure, backed by the file. This is the streaming engine, for free. **Never `--no-mmap`** on a >RAM model — it forces a full allocation and thrashes. +- **`--cpu-moe`** keeps expert tensors on the CPU/mmap (disk-backed) side. On a **CUDA** build you pair it with `-ngl` to put the dense part in the GPU (resident) while experts stream on the CPU — the textbook split. The vendored build is **CPU-only** (no CUDA backend compiled), so `-ngl`/`--cpu-moe` are GPU no-ops; the split is instead realized by a RAM cap. +- **`MemoryMax` (cgroup v2)** — `systemd-run --user --scope -p MemoryMax=6G` caps the process below the model size. The kernel then keeps the small dense part + hot experts resident and evicts cold expert pages, re-reading them from disk on demand. This turns the OS page cache into colibrì's tiering (hot resident, cold on disk); colibrì just makes it *smart* — per-layer LRU, `fadvise` readahead, pinning the measured-hottest experts. + +### 7.3 Run it + +```bash +cd prototypes/llama-moe-stream +./run-moe.sh # baseline: 17 GB model fits in RAM → all resident +MEM_CAP=6G ./run-moe.sh # capped: dense stays hot, cold experts stream from disk +``` + +The proof: under a 6 GB cap the 17 GB model **still answers correctly** — the missing experts are served from disk on demand — and tok/s drops vs the baseline; that gap is the disk-streaming cost. `-hf` downloads + caches the GGUF (`~/.cache/llama.cpp`) on first run. + +### 7.4 Privacy — does local inference leak your data? + +**No.** llama.cpp inference is fully on-device: it reads a local GGUF, has **no telemetry**, and opens **no outbound connections** while generating — prompts and code never leave the machine (weights are inert data, not code that can "phone home"). The *only* network is the one-time `-hf` weight download (inbound; HuggingFace sees your IP + which file, not your data). To be certain, run air-gapped from the cached file: + +```bash +GGUF=$(find ~/.cache/llama.cpp -name 'Qwen3-Coder-30B-A3B-Instruct-Q4_K_M.gguf' | head -1) +MODEL_PATH="$GGUF" OFFLINE=1 MEM_CAP=6G ./run-moe.sh # HF_HUB_OFFLINE=1, no -hf, zero network +``` + +`ss -tnp` during the run shows **no established connections** from `llama-cli`. The real leak surface is the *client* (an editor plugin misconfigured to a cloud model, or plugin telemetry) — not the engine. + +### 7.5 Measured on the dev box + +Host: i7-13700H (20 threads, AVX2+VNNI), **31 GB RAM**, RTX 4050 Laptop (6 GB), NVMe; vendored llama.cpp is a **CPU-only** build. Model: `unsloth/Qwen3-Coder-30B-A3B-Instruct-GGUF:Q4_K_M` (~17.3 GB, single file, fused-expert layout). + +| Run | RAM available to process | Expected behavior | Measured tok/s | +|---|---|---|---| +| baseline (uncapped) | whole 17 GB can stay resident | RAM/matmul-bound after warm-up | _pending run_ | +| `MEM_CAP=6G` + offline | 6 GB — dense + hot experts only | cold experts stream from NVMe; **no network** | _pending run_ | + +_Numbers are filled in from the in-progress background run; the qualitative result — correct output under a cap far below model size, with zero outbound connections — is the point regardless of the exact rate._ + +--- + +## Sources + +- `.dev/reference/colibri/` — vendored source (README, `docs/ENVIRONMENT.md`, `c/glm.c`, `c/olmoe.c`, `c/tools/convert_olmoe.py`, `c/uring.h`, `c/compat.h`, `flake.nix`), upstream +- `.dev/reference/llama-cpp/` — vendored llama.cpp checkout (HEAD `635cdd5fc`), Mistral support verified in `src/llama-arch.{h,cpp}`, `conversion/{llama,mistral,mistral3,pixtral}.py`, `src/llama-chat.cpp`, `common/arg.cpp` +- `prototypes/llama-moe-stream/` — the hands-on demo (`run-moe.sh`, README) added by this work +- Model GGUFs: [`unsloth/Qwen3-Coder-30B-A3B-Instruct-GGUF`](https://huggingface.co/unsloth/Qwen3-Coder-30B-A3B-Instruct-GGUF) (fused-format, the demo default); old per-expert-layout examples that **fail** to load on current llama.cpp: [`TheBloke/Mixtral-8x7B-Instruct-v0.1-GGUF`](https://huggingface.co/TheBloke/Mixtral-8x7B-Instruct-v0.1-GGUF), [`MaziyarPanahi/Mixtral-8x22B-Instruct-v0.1-GGUF`](https://huggingface.co/MaziyarPanahi/Mixtral-8x22B-Instruct-v0.1-GGUF) +- [Mixtral 8x22B — Prompt Engineering Guide](https://www.promptingguide.ai/models/mixtral-8x22b) and [Ollama library: mixtral:8x22b](https://ollama.com/library/mixtral:8x22b) +- [Performant local MoE CPU inference with GPU acceleration in llama.cpp](https://huggingface.co/blog/Doctor-Shotgun/llamacpp-moe-offload-guide) (the `--n-cpu-moe` / `--override-tensor` guide) +- [KTransformers](https://github.com/kvcache-ai/ktransformers) — CPU/GPU hybrid MoE inference (lists Mixtral 8x7B/8x22B support) diff --git a/docs/plan/exploration/fibers/00-fibers.md b/docs/plan/exploration/fibers/00-fibers.md new file mode 100644 index 0000000..ec00485 --- /dev/null +++ b/docs/plan/exploration/fibers/00-fibers.md @@ -0,0 +1,96 @@ +# Fibers — green threads on the wovm shard scheduler + +> Research note (2026-08-08) feeding story iteration 11. Expands +> [the blue-green vision §3](../blue-green-vm/00-vision.md) with the +> kernel's-eye evidence and the precedent survey. Prose only; the spec and +> plan follow the brainstorming → writing-plans path when the iteration +> starts. + +## Why the kernel cannot do this for us + +Reference: [Threads and the OS kernel's view](https://learn.padho.ai/wiki/threads-and-the-os-kernels-view). +The facts that matter, condensed: + +- A Linux thread IS a `task_struct`: own TID, own kernel stack (8–16 KiB), + own scheduling slot; `clone()` flags decide what is shared. There is no + cheaper kernel thread to ask for. +- Costs per thread: ~8 MiB stack VMA, ~9 KiB `task_struct`, + ~20 µs creation, 1–3 µs per context switch. Measured against a userspace + runtime spawn (~58 ns): **~342× creation cost, ~70× memory**. +- CFS keys a red-black tree by `vruntime` — O(log N) per pick. At tens of + thousands of runnable tasks the scheduling slice approaches the + context-switch cost and the kernel scheduler becomes the bottleneck + before application code does. +- M:N threading died *in the kernel* (NPTL, 2003) and was reborn in + userspace (goroutines, BEAM processes, Tokio tasks, Loom virtual + threads) — because only the runtime knows its own blocking points and + can keep per-task state tiny. + +Conclusion the industry already reached and we adopt: **one kernel task +per core** (plan 09 / iteration 8's pinned shard — already doctrine), +**userspace tasks above it**. The kernel schedules cores; the runtime +schedules work. + +## Precedent survey — who has green threads, and how + +| Runtime | Task state | Preemption | I/O integration | Lesson for wovm | +| --- | --- | --- | --- | --- | +| **Erlang/BEAM** | interpreter state per process, private heap | **reduction budget** (~2k reductions, checked at calls) | park on scheduler, poll set | the closest shape: interpreter = scheduler; deterministic, signal-free preemption | +| **Go** | native stack, 8 KiB grown by copy | async preemption via signals (since 1.14) + safepoints | netpoller parks goroutines | stack copying needs precise pointer maps — heavy machinery; signals are what the vision explicitly avoids | +| **Java Loom** | continuation frames on heap, unmounted from carrier | cooperative at yield points | blocking calls in the JDK park the virtual thread | "same blocking API, runtime parks underneath" — exactly our stdlib posture | +| **Tokio/Rust** | stackless state machines (`async fn`) | cooperative at `.await` | reactor + waker | rejected surface: writeonce has **no async/await keyword** (systems-track spec Part 2); function coloring is the disease | +| **Lua** | coroutine = own Lua stack (interpreter state) | none (pure cooperative) | up to the host | proof that interpreter-state fibers are nearly free; but no preemption = one hot loop starves the shard | +| **boost.context / libco** | native stack + hand-rolled register switch | none | none | what we do NOT need: wovm executes bytecode, so no native stack switching, no asm, no guard pages | + +## The wovm design (vision §3, confirmed by this survey) + +A fiber is **the execution state the VM already isolates**: register-window +stack + frame stack + pc. Making that per-fiber instead of per-VM turns the +interpreter into a scheduler almost for free — the BEAM/Lua insight, minus +Lua's starvation problem: + +- **Preemption by reduction budget** — the dispatch loop decrements a + counter per instruction (or per call/back-edge); at zero the fiber parks + and the next runnable one is picked. No signals, no safepoint asm, no + stack copying; deterministic and debuggable. (Erlang has run this design + in telecom production for three decades.) +- **Park on I/O** — a blocking stdlib builtin on a server shard hands its + fd to the shard's epoll/io_uring loop and parks the fiber; the completion + resumes it. Same typed builtins, blocking look, no coloring — the + systems-track "one API, two execution disciplines" doctrine gains its + third discipline: program mode blocks the thread, server shards park the + fiber, the source text is identical. +- **Ownership fits** — a fiber owns its objects like a shard owns its heap; + cross-fiber sends move ownership exactly like iteration 8's cross-shard + sends (same rule, cheaper path: same heap, no copy). `@gc` references + stay shard-local either way, so the per-shard cycle collector (staged to + iteration 8) needs fiber stacks as additional roots — the one real + collector interaction to spec. +- **Cost target** — fiber creation is an arena allocation of a small + context (hundreds of bytes, the ~58 ns class), park/resume is a pointer + swap in the dispatch loop; thousands of fibers per shard where the + kernel tops out at hundreds of threads per core. + +## What must be specified before implementation (open questions) + +1. Surface: `spawn` returns what — a fiber handle, an actor address, or + nothing (fire-and-forget)? Iteration 8's `spawn` and mailbox surface + should be the same word; fibers refine its granularity. +2. Reduction budget size and where it is checked (per instruction vs per + call/back-edge) — measure both in the interpreter before choosing. +3. Parked-fiber lifetime: what drops a fiber blocked forever (shard + shutdown, blue-green drain)? Unwinding a parked fiber must run its drop + maps — same machinery as trap unwind (OOP spec §6). +4. Fairness: run queue is FIFO per shard in v1; priorities/timers are a + later capability (the recipe-box rule — no framework policy in the + runtime). +5. Program mode: stays fiber-free in v1 (one thread, blocking legal) or + gains the same scheduler? Default: fiber-free — log-watcher needs none. + +## Doctrine check + +Kernel primitives only (epoll/io_uring already ours; no new syscalls +needed) · no signals · no async/await keyword · ownership moves, never +shares · per-shard everything (heap, GC, now run queue) · plain +diagnostics (a starved-fiber warning names the hot function via the line +table). No principle bends. diff --git a/docs/plan/exploration/linux/00-linux.md b/docs/plan/exploration/linux/00-linux.md index 9d332ee..4669652 100644 --- a/docs/plan/exploration/linux/00-linux.md +++ b/docs/plan/exploration/linux/00-linux.md @@ -4,7 +4,7 @@ Kernel primitives that the writeonce binary can leverage, mapped to the architec ### Per-primitive reference cards -Each primitive has its own numbered file with the kernel source path (into [`reference/linux/`](../../../reference/linux/)), Rust FFI signature via `libc`, a minimal direct-syscall example, and the v1 port source. Use these when implementing the phase docs under [`docs/plan/`](../). +Each primitive has its own numbered file with the kernel source path (into [`.dev/reference/linux/`](../../../.dev/reference/linux/)), Rust FFI signature via `libc`, a minimal direct-syscall example, and the v1 port source. Use these when implementing the phase docs under [`docs/plan/`](../). | # | Primitive | Used by | | --- | --- | --- | @@ -143,4 +143,4 @@ Each pattern resolves to a set of inotify watch descriptors. When the watched se ## Related: the assembly policy -Every primitive above is reached via `libc::` or `libc::syscall(SYS_*, ...)` — no custom assembly. The reasoning lives in [`../assembly/`](../assembly/) — three files covering why runtimes use asm at all ([`00-overview.md`](../assembly/00-overview.md)), what Go's [`reference/go/src/runtime/*.s`](../../../reference/go/src/runtime/) actually contains ([`01-go-runtime-asm.md`](../assembly/01-go-runtime-asm.md)), and the writeonce policy that all of it is replaced by Rust stdlib + libc ([`02-writeonce-stance.md`](../assembly/02-writeonce-stance.md)). +Every primitive above is reached via `libc::` or `libc::syscall(SYS_*, ...)` — no custom assembly. The reasoning lives in [`../assembly/`](../assembly/) — three files covering why runtimes use asm at all ([`00-overview.md`](../assembly/00-overview.md)), what Go's [`.dev/reference/go/src/runtime/*.s`](../../../.dev/reference/go/src/runtime/) actually contains ([`01-go-runtime-asm.md`](../assembly/01-go-runtime-asm.md)), and the writeonce policy that all of it is replaced by Rust stdlib + libc ([`02-writeonce-stance.md`](../assembly/02-writeonce-stance.md)). diff --git a/docs/plan/exploration/linux/01-epoll.md b/docs/plan/exploration/linux/01-epoll.md index d891ca4..c4db9c5 100644 --- a/docs/plan/exploration/linux/01-epoll.md +++ b/docs/plan/exploration/linux/01-epoll.md @@ -6,8 +6,8 @@ Event-driven I/O multiplexing. One `epoll_fd` watches many fds for readiness; `e | Path | What | | --- | --- | -| [`reference/linux/fs/eventpoll.c`](../../../reference/linux/fs/eventpoll.c) | All three syscalls (`epoll_create1`, `epoll_ctl`, `epoll_wait`) live here. Grep for `SYSCALL_DEFINE`. | -| [`reference/linux/include/uapi/linux/eventpoll.h`](../../../reference/linux/include/uapi/linux/eventpoll.h) | `struct epoll_event`, `EPOLL_*` flags, the userspace-facing ABI. | +| [`.dev/reference/linux/fs/eventpoll.c`](../../../.dev/reference/linux/fs/eventpoll.c) | All three syscalls (`epoll_create1`, `epoll_ctl`, `epoll_wait`) live here. Grep for `SYSCALL_DEFINE`. | +| [`.dev/reference/linux/include/uapi/linux/eventpoll.h`](../../../.dev/reference/linux/include/uapi/linux/eventpoll.h) | `struct epoll_event`, `EPOLL_*` flags, the userspace-facing ABI. | ## Man pages @@ -72,4 +72,4 @@ Every runtime phase that touches I/O: [`02-event-loop-epoll.md`](../02-event-loo ## v1 port source -[`reference/crates/wo-event/src/epoll.rs`](../../../reference/crates/wo-event/src/epoll.rs) (183 LOC) — already wraps all three syscalls with a safe `EventLoop { register, deregister, wait_once }` facade. +[`.dev/reference/crates/wo-event/src/epoll.rs`](../../../.dev/reference/crates/wo-event/src/epoll.rs) (183 LOC) — already wraps all three syscalls with a safe `EventLoop { register, deregister, wait_once }` facade. diff --git a/docs/plan/exploration/linux/02-eventfd.md b/docs/plan/exploration/linux/02-eventfd.md index 236d8e1..a7f0a38 100644 --- a/docs/plan/exploration/linux/02-eventfd.md +++ b/docs/plan/exploration/linux/02-eventfd.md @@ -6,8 +6,8 @@ Counter as a file descriptor. `write(fd, &n, 8)` adds `n` to the counter; `read( | Path | What | | --- | --- | -| [`reference/linux/fs/eventfd.c`](../../../reference/linux/fs/eventfd.c) | `SYSCALL_DEFINE2(eventfd, ...)`, `struct eventfd_ctx`, read/write handlers. | -| [`reference/linux/include/uapi/linux/eventfd.h`](../../../reference/linux/include/uapi/linux/eventfd.h) | `EFD_*` flags. | +| [`.dev/reference/linux/fs/eventfd.c`](../../../.dev/reference/linux/fs/eventfd.c) | `SYSCALL_DEFINE2(eventfd, ...)`, `struct eventfd_ctx`, read/write handlers. | +| [`.dev/reference/linux/include/uapi/linux/eventfd.h`](../../../.dev/reference/linux/include/uapi/linux/eventfd.h) | `EFD_*` flags. | ## Man pages @@ -63,4 +63,4 @@ unsafe { ## v1 port source -[`reference/crates/wo-event/src/eventfd.rs`](../../../reference/crates/wo-event/src/eventfd.rs) (66 LOC) — `EventFd { new, write, read, as_raw_fd }`. +[`.dev/reference/crates/wo-event/src/eventfd.rs`](../../../.dev/reference/crates/wo-event/src/eventfd.rs) (66 LOC) — `EventFd { new, write, read, as_raw_fd }`. diff --git a/docs/plan/exploration/linux/03-timerfd.md b/docs/plan/exploration/linux/03-timerfd.md index e523a6b..21d5695 100644 --- a/docs/plan/exploration/linux/03-timerfd.md +++ b/docs/plan/exploration/linux/03-timerfd.md @@ -6,8 +6,8 @@ Timers as file descriptors. Set an expiry with `timerfd_settime`, `read` the fd | Path | What | | --- | --- | -| [`reference/linux/fs/timerfd.c`](../../../reference/linux/fs/timerfd.c) | All three syscalls (`timerfd_create`, `timerfd_settime`, `timerfd_gettime`). | -| [`reference/linux/include/uapi/linux/timerfd.h`](../../../reference/linux/include/uapi/linux/timerfd.h) | `TFD_*` flags. | +| [`.dev/reference/linux/fs/timerfd.c`](../../../.dev/reference/linux/fs/timerfd.c) | All three syscalls (`timerfd_create`, `timerfd_settime`, `timerfd_gettime`). | +| [`.dev/reference/linux/include/uapi/linux/timerfd.h`](../../../.dev/reference/linux/include/uapi/linux/timerfd.h) | `TFD_*` flags. | ## Man pages @@ -74,4 +74,4 @@ unsafe { ## v1 port source -[`reference/crates/wo-event/src/timerfd.rs`](../../../reference/crates/wo-event/src/timerfd.rs) (91 LOC) — `TimerFd { oneshot(dur), periodic(dur), disarm, read_expirations }`. +[`.dev/reference/crates/wo-event/src/timerfd.rs`](../../../.dev/reference/crates/wo-event/src/timerfd.rs) (91 LOC) — `TimerFd { oneshot(dur), periodic(dur), disarm, read_expirations }`. diff --git a/docs/plan/exploration/linux/04-signalfd.md b/docs/plan/exploration/linux/04-signalfd.md index e364870..052cced 100644 --- a/docs/plan/exploration/linux/04-signalfd.md +++ b/docs/plan/exploration/linux/04-signalfd.md @@ -6,9 +6,9 @@ Unix signals as file descriptors. `signalfd(fd, mask)` installs a mask on the pr | Path | What | | --- | --- | -| [`reference/linux/fs/signalfd.c`](../../../reference/linux/fs/signalfd.c) | `SYSCALL_DEFINE4(signalfd4, ...)` + `signalfd_dequeue`. | -| [`reference/linux/include/uapi/linux/signalfd.h`](../../../reference/linux/include/uapi/linux/signalfd.h) | `struct signalfd_siginfo`, `SFD_*` flags. | -| [`reference/linux/kernel/signal.c`](../../../reference/linux/kernel/signal.c) | Background: `sigprocmask`, pending-signal dequeue. | +| [`.dev/reference/linux/fs/signalfd.c`](../../../.dev/reference/linux/fs/signalfd.c) | `SYSCALL_DEFINE4(signalfd4, ...)` + `signalfd_dequeue`. | +| [`.dev/reference/linux/include/uapi/linux/signalfd.h`](../../../.dev/reference/linux/include/uapi/linux/signalfd.h) | `struct signalfd_siginfo`, `SFD_*` flags. | +| [`.dev/reference/linux/kernel/signal.c`](../../../.dev/reference/linux/kernel/signal.c) | Background: `sigprocmask`, pending-signal dequeue. | ## Man pages @@ -74,4 +74,4 @@ unsafe { ## v1 port source -[`reference/crates/wo-event/src/signalfd.rs`](../../../reference/crates/wo-event/src/signalfd.rs) (62 LOC) — `SignalFd::new(&[SIGINT, SIGTERM]) -> SignalFd` with a safe `read_signo()` helper. +[`.dev/reference/crates/wo-event/src/signalfd.rs`](../../../.dev/reference/crates/wo-event/src/signalfd.rs) (62 LOC) — `SignalFd::new(&[SIGINT, SIGTERM]) -> SignalFd` with a safe `read_signo()` helper. diff --git a/docs/plan/exploration/linux/05-inotify.md b/docs/plan/exploration/linux/05-inotify.md index a282628..3c24b6f 100644 --- a/docs/plan/exploration/linux/05-inotify.md +++ b/docs/plan/exploration/linux/05-inotify.md @@ -6,9 +6,9 @@ Filesystem event notifications as a file descriptor. `inotify_add_watch(dir, mas | Path | What | | --- | --- | -| [`reference/linux/fs/notify/inotify/inotify_user.c`](../../../reference/linux/fs/notify/inotify/inotify_user.c) | `SYSCALL_DEFINE1(inotify_init1, ...)`, `SYSCALL_DEFINE3(inotify_add_watch, ...)`, `SYSCALL_DEFINE2(inotify_rm_watch, ...)`. | -| [`reference/linux/fs/notify/inotify/inotify_fsnotify.c`](../../../reference/linux/fs/notify/inotify/inotify_fsnotify.c) | The fsnotify backend that feeds events into the fd. | -| [`reference/linux/include/uapi/linux/inotify.h`](../../../reference/linux/include/uapi/linux/inotify.h) | `struct inotify_event`, `IN_*` masks. | +| [`.dev/reference/linux/fs/notify/inotify/inotify_user.c`](../../../.dev/reference/linux/fs/notify/inotify/inotify_user.c) | `SYSCALL_DEFINE1(inotify_init1, ...)`, `SYSCALL_DEFINE3(inotify_add_watch, ...)`, `SYSCALL_DEFINE2(inotify_rm_watch, ...)`. | +| [`.dev/reference/linux/fs/notify/inotify/inotify_fsnotify.c`](../../../.dev/reference/linux/fs/notify/inotify/inotify_fsnotify.c) | The fsnotify backend that feeds events into the fd. | +| [`.dev/reference/linux/include/uapi/linux/inotify.h`](../../../.dev/reference/linux/include/uapi/linux/inotify.h) | `struct inotify_event`, `IN_*` masks. | ## Man pages @@ -85,4 +85,4 @@ unsafe { ## v1 port source -[`reference/crates/wo-watch/src/lib.rs`](../../../reference/crates/wo-watch/src/lib.rs) (280 LOC) — already does recursive watch setup, event parsing, and path resolution via a `wd → PathBuf` map. +[`.dev/reference/crates/wo-watch/src/lib.rs`](../../../.dev/reference/crates/wo-watch/src/lib.rs) (280 LOC) — already does recursive watch setup, event parsing, and path resolution via a `wd → PathBuf` map. diff --git a/docs/plan/exploration/linux/06-sendfile.md b/docs/plan/exploration/linux/06-sendfile.md index 7e91e42..bda7cb2 100644 --- a/docs/plan/exploration/linux/06-sendfile.md +++ b/docs/plan/exploration/linux/06-sendfile.md @@ -6,8 +6,8 @@ Zero-copy transfer from a file fd to a socket fd. The kernel splices pages direc | Path | What | | --- | --- | -| [`reference/linux/fs/read_write.c`](../../../reference/linux/fs/read_write.c) | `SYSCALL_DEFINE4(sendfile, ...)` and `SYSCALL_DEFINE4(sendfile64, ...)`. Modern glibc aliases the first to the second; the syscalls are distinguished by the offset type. | -| [`reference/linux/fs/splice.c`](../../../reference/linux/fs/splice.c) | Internally `sendfile` delegates to `splice_direct_to_actor`. Related — see [07-splice.md](./07-splice.md) if you ever need the more general fd-to-fd pipe path. | +| [`.dev/reference/linux/fs/read_write.c`](../../../.dev/reference/linux/fs/read_write.c) | `SYSCALL_DEFINE4(sendfile, ...)` and `SYSCALL_DEFINE4(sendfile64, ...)`. Modern glibc aliases the first to the second; the syscalls are distinguished by the offset type. | +| [`.dev/reference/linux/fs/splice.c`](../../../.dev/reference/linux/fs/splice.c) | Internally `sendfile` delegates to `splice_direct_to_actor`. Related — see [07-splice.md](./07-splice.md) if you ever need the more general fd-to-fd pipe path. | ## Man pages @@ -75,4 +75,4 @@ unsafe { ## v1 port source -[`reference/crates/wo-serve/src/sendfile.rs`](../../../reference/crates/wo-serve/src/sendfile.rs) (109 LOC) — `send_file(sock, path) -> Result` wrapping the loop + `EAGAIN` handling. +[`.dev/reference/crates/wo-serve/src/sendfile.rs`](../../../.dev/reference/crates/wo-serve/src/sendfile.rs) (109 LOC) — `send_file(sock, path) -> Result` wrapping the loop + `EAGAIN` handling. diff --git a/docs/plan/exploration/linux/07-io_uring.md b/docs/plan/exploration/linux/07-io_uring.md index b6e2f9a..fe6ee75 100644 --- a/docs/plan/exploration/linux/07-io_uring.md +++ b/docs/plan/exploration/linux/07-io_uring.md @@ -8,9 +8,9 @@ Ring-buffer based async I/O (Linux 5.1+, mature 5.11+). Two lock-free SPSC rings | Path | What | | --- | --- | -| [`reference/linux/io_uring/`](../../../reference/linux/io_uring/) | Whole subsystem. Start with `io_uring.c` (ring setup + submission/completion) and `fs.c` (fsync op). | -| [`reference/linux/io_uring/io_uring.c`](../../../reference/linux/io_uring/io_uring.c) | `SYSCALL_DEFINE2(io_uring_setup, ...)`, `SYSCALL_DEFINE6(io_uring_enter, ...)`, `SYSCALL_DEFINE4(io_uring_register, ...)`. | -| [`reference/linux/include/uapi/linux/io_uring.h`](../../../reference/linux/include/uapi/linux/io_uring.h) | `struct io_uring_sqe`, `io_uring_cqe`, `io_uring_params`, every `IORING_*` flag. | +| [`.dev/reference/linux/io_uring/`](../../../.dev/reference/linux/io_uring/) | Whole subsystem. Start with `io_uring.c` (ring setup + submission/completion) and `fs.c` (fsync op). | +| [`.dev/reference/linux/io_uring/io_uring.c`](../../../.dev/reference/linux/io_uring/io_uring.c) | `SYSCALL_DEFINE2(io_uring_setup, ...)`, `SYSCALL_DEFINE6(io_uring_enter, ...)`, `SYSCALL_DEFINE4(io_uring_register, ...)`. | +| [`.dev/reference/linux/include/uapi/linux/io_uring.h`](../../../.dev/reference/linux/include/uapi/linux/io_uring.h) | `struct io_uring_sqe`, `io_uring_cqe`, `io_uring_params`, every `IORING_*` flag. | ## Man pages diff --git a/docs/plan/exploration/linux/08-mmap.md b/docs/plan/exploration/linux/08-mmap.md index 8156ff3..ed7e2f8 100644 --- a/docs/plan/exploration/linux/08-mmap.md +++ b/docs/plan/exploration/linux/08-mmap.md @@ -8,9 +8,9 @@ Central to Phase 3's storage engine: segment files are `mmap`ed read-only for O( | Path | What | | --- | --- | -| [`reference/linux/mm/mmap.c`](../../../reference/linux/mm/mmap.c) | VMA creation, `SYSCALL_DEFINE6(mmap, ...)`, `SYSCALL_DEFINE2(munmap, ...)`. | -| [`reference/linux/mm/madvise.c`](../../../reference/linux/mm/madvise.c) | `SYSCALL_DEFINE3(madvise, ...)` + every `MADV_*` handler. | -| [`reference/linux/include/uapi/linux/mman.h`](../../../reference/linux/include/uapi/linux/mman.h) | `MAP_*` flags, huge-page sizing macros. | +| [`.dev/reference/linux/mm/mmap.c`](../../../.dev/reference/linux/mm/mmap.c) | VMA creation, `SYSCALL_DEFINE6(mmap, ...)`, `SYSCALL_DEFINE2(munmap, ...)`. | +| [`.dev/reference/linux/mm/madvise.c`](../../../.dev/reference/linux/mm/madvise.c) | `SYSCALL_DEFINE3(madvise, ...)` + every `MADV_*` handler. | +| [`.dev/reference/linux/include/uapi/linux/mman.h`](../../../.dev/reference/linux/include/uapi/linux/mman.h) | `MAP_*` flags, huge-page sizing macros. | | POSIX `` | The other half of the constants (`PROT_*`, `MADV_*`). Usually folded into `linux/mman.h` by libc. | ## Man pages diff --git a/docs/plan/exploration/linux/09-fallocate.md b/docs/plan/exploration/linux/09-fallocate.md index 252200c..f16edbd 100644 --- a/docs/plan/exploration/linux/09-fallocate.md +++ b/docs/plan/exploration/linux/09-fallocate.md @@ -8,9 +8,9 @@ Together they form the backbone of the storage engine's on-disk layout: segment | Path | What | | --- | --- | -| [`reference/linux/fs/open.c`](../../../reference/linux/fs/open.c) | `SYSCALL_DEFINE4(fallocate, ...)`. The syscall delegates to `file->f_op->fallocate` — per-filesystem. | -| [`reference/linux/fs/read_write.c`](../../../reference/linux/fs/read_write.c) | `SYSCALL_DEFINE4(pread64, ...)`, `SYSCALL_DEFINE4(pwrite64, ...)`, `SYSCALL_DEFINE6(pwritev2, ...)`. | -| [`reference/linux/include/uapi/linux/falloc.h`](../../../reference/linux/include/uapi/linux/falloc.h) | `FALLOC_FL_*` flags. | +| [`.dev/reference/linux/fs/open.c`](../../../.dev/reference/linux/fs/open.c) | `SYSCALL_DEFINE4(fallocate, ...)`. The syscall delegates to `file->f_op->fallocate` — per-filesystem. | +| [`.dev/reference/linux/fs/read_write.c`](../../../.dev/reference/linux/fs/read_write.c) | `SYSCALL_DEFINE4(pread64, ...)`, `SYSCALL_DEFINE4(pwrite64, ...)`, `SYSCALL_DEFINE6(pwritev2, ...)`. | +| [`.dev/reference/linux/include/uapi/linux/falloc.h`](../../../.dev/reference/linux/include/uapi/linux/falloc.h) | `FALLOC_FL_*` flags. | ## Man pages diff --git a/docs/plan/exploration/linux/10-pidfd.md b/docs/plan/exploration/linux/10-pidfd.md index f4514b6..3b655be 100644 --- a/docs/plan/exploration/linux/10-pidfd.md +++ b/docs/plan/exploration/linux/10-pidfd.md @@ -8,10 +8,10 @@ Not on the runtime's critical path today; useful when the runtime grows a superv | Path | What | | --- | --- | -| [`reference/linux/kernel/pid.c`](../../../reference/linux/kernel/pid.c) | `SYSCALL_DEFINE2(pidfd_open, ...)`, `pidfd_create`, `pidfd_pid`. | -| [`reference/linux/kernel/signal.c`](../../../reference/linux/kernel/signal.c) | `SYSCALL_DEFINE4(pidfd_send_signal, ...)`. | -| [`reference/linux/kernel/fork.c`](../../../reference/linux/kernel/fork.c) | `clone3` — the only way to get a pidfd atomically with spawn. | -| [`reference/linux/include/uapi/linux/pidfd.h`](../../../reference/linux/include/uapi/linux/pidfd.h) | `PIDFD_*` flags. | +| [`.dev/reference/linux/kernel/pid.c`](../../../.dev/reference/linux/kernel/pid.c) | `SYSCALL_DEFINE2(pidfd_open, ...)`, `pidfd_create`, `pidfd_pid`. | +| [`.dev/reference/linux/kernel/signal.c`](../../../.dev/reference/linux/kernel/signal.c) | `SYSCALL_DEFINE4(pidfd_send_signal, ...)`. | +| [`.dev/reference/linux/kernel/fork.c`](../../../.dev/reference/linux/kernel/fork.c) | `clone3` — the only way to get a pidfd atomically with spawn. | +| [`.dev/reference/linux/include/uapi/linux/pidfd.h`](../../../.dev/reference/linux/include/uapi/linux/pidfd.h) | `PIDFD_*` flags. | ## Man pages diff --git a/docs/plan/exploration/linux/11-memfd_create.md b/docs/plan/exploration/linux/11-memfd_create.md index f046a5f..06eea06 100644 --- a/docs/plan/exploration/linux/11-memfd_create.md +++ b/docs/plan/exploration/linux/11-memfd_create.md @@ -8,9 +8,9 @@ Useful for the storage engine's transient work: building an index in memory befo | Path | What | | --- | --- | -| [`reference/linux/mm/memfd.c`](../../../reference/linux/mm/memfd.c) | `SYSCALL_DEFINE2(memfd_create, ...)` + seal ops. | -| [`reference/linux/include/uapi/linux/memfd.h`](../../../reference/linux/include/uapi/linux/memfd.h) | `MFD_*` flags. | -| [`reference/linux/include/uapi/linux/fcntl.h`](../../../reference/linux/include/uapi/linux/fcntl.h) | `F_ADD_SEALS`, `F_GET_SEALS`, `F_SEAL_*` constants. Sealing is a `fcntl(F_ADD_SEALS, ...)` operation on the memfd. | +| [`.dev/reference/linux/mm/memfd.c`](../../../.dev/reference/linux/mm/memfd.c) | `SYSCALL_DEFINE2(memfd_create, ...)` + seal ops. | +| [`.dev/reference/linux/include/uapi/linux/memfd.h`](../../../.dev/reference/linux/include/uapi/linux/memfd.h) | `MFD_*` flags. | +| [`.dev/reference/linux/include/uapi/linux/fcntl.h`](../../../.dev/reference/linux/include/uapi/linux/fcntl.h) | `F_ADD_SEALS`, `F_GET_SEALS`, `F_SEAL_*` constants. Sealing is a `fcntl(F_ADD_SEALS, ...)` operation on the memfd. | ## Man pages diff --git a/docs/plan/exploration/linux/12-pwrite-fsync.md b/docs/plan/exploration/linux/12-pwrite-fsync.md index f2c3799..6d77474 100644 --- a/docs/plan/exploration/linux/12-pwrite-fsync.md +++ b/docs/plan/exploration/linux/12-pwrite-fsync.md @@ -15,10 +15,10 @@ The previous cards cover positional I/O ([`09-fallocate.md`](./09-fallocate.md)) | Postgres call | Wraps | Where | | --- | --- | --- | -| `pg_pwrite()` | `pwrite64` | [`storage/file/fd.c`](../../../../reference/postgresql/src/backend/storage/file/fd.c) — every block-aligned write. | -| `pg_fsync()` | `fsync` (or platform variant) | [`storage/file/fd.c`](../../../../reference/postgresql/src/backend/storage/file/fd.c) — wraps `wal_sync_method` GUC dispatch. | +| `pg_pwrite()` | `pwrite64` | [`storage/file/fd.c`](../../../../.dev/reference/postgresql/src/backend/storage/file/fd.c) — every block-aligned write. | +| `pg_fsync()` | `fsync` (or platform variant) | [`storage/file/fd.c`](../../../../.dev/reference/postgresql/src/backend/storage/file/fd.c) — wraps `wal_sync_method` GUC dispatch. | | `pg_fdatasync()` | `fdatasync` | Same. Selected when `wal_sync_method = fdatasync`. | -| Async writeback | `sync_file_range` | [`access/transam/xlog.c`](../../../../reference/postgresql/src/backend/access/transam/xlog.c) — `issue_xlog_fsync` calls `sync_file_range(SYNC_FILE_RANGE_WRITE)` to start I/O on the WAL ahead of the durability barrier. | +| Async writeback | `sync_file_range` | [`access/transam/xlog.c`](../../../../.dev/reference/postgresql/src/backend/access/transam/xlog.c) — `issue_xlog_fsync` calls `sync_file_range(SYNC_FILE_RANGE_WRITE)` to start I/O on the WAL ahead of the durability barrier. | The Postgres GUC matrix (`wal_sync_method`) lets the operator pick between `fsync`, `fdatasync`, `open_sync`, `open_datasync`, `fsync_writethrough`. **Writeonce picks one** — `fdatasync` for the WAL, `fsync` for control files and segment rollovers — and ships it. @@ -26,10 +26,10 @@ The Postgres GUC matrix (`wal_sync_method`) lets the operator pick between `fsyn | Path | What | | --- | --- | -| [`reference/linux/fs/read_write.c`](../../../reference/linux/fs/read_write.c) | `SYSCALL_DEFINE4(pread64, ...)`, `SYSCALL_DEFINE4(pwrite64, ...)`, `SYSCALL_DEFINE6(pwritev2, ...)`. | -| [`reference/linux/fs/sync.c`](../../../reference/linux/fs/sync.c) | `SYSCALL_DEFINE1(fsync, ...)`, `SYSCALL_DEFINE1(fdatasync, ...)`, `SYSCALL_DEFINE4(sync_file_range, ...)`. | -| [`reference/linux/include/uapi/asm-generic/fcntl.h`](../../../reference/linux/include/uapi/asm-generic/fcntl.h) | `O_SYNC`, `O_DSYNC`, `O_DIRECT`. | -| [`reference/linux/Documentation/filesystems/ext4/journal.rst`](../../../reference/linux/Documentation/filesystems/ext4/journal.rst) | What ext4's journal commits when `fsync` runs. Worth understanding what the kernel actually does on the durability path. | +| [`.dev/reference/linux/fs/read_write.c`](../../../.dev/reference/linux/fs/read_write.c) | `SYSCALL_DEFINE4(pread64, ...)`, `SYSCALL_DEFINE4(pwrite64, ...)`, `SYSCALL_DEFINE6(pwritev2, ...)`. | +| [`.dev/reference/linux/fs/sync.c`](../../../.dev/reference/linux/fs/sync.c) | `SYSCALL_DEFINE1(fsync, ...)`, `SYSCALL_DEFINE1(fdatasync, ...)`, `SYSCALL_DEFINE4(sync_file_range, ...)`. | +| [`.dev/reference/linux/include/uapi/asm-generic/fcntl.h`](../../../.dev/reference/linux/include/uapi/asm-generic/fcntl.h) | `O_SYNC`, `O_DSYNC`, `O_DIRECT`. | +| [`.dev/reference/linux/Documentation/filesystems/ext4/journal.rst`](../../../.dev/reference/linux/Documentation/filesystems/ext4/journal.rst) | What ext4's journal commits when `fsync` runs. Worth understanding what the kernel actually does on the durability path. | ## Man pages @@ -149,4 +149,4 @@ Pair with [`postgresql/wal.md`](../postgresql/wal.md), [`postgresql/buffer-and-c ## v1 port source -**Partial.** `reference/crates/wo-seg/src/writer.rs:92` calls `file.sync_all()` (Rust stdlib's `fsync` wrapper). Phase 11 replaces with explicit `libc::fdatasync` for the WAL path; segment files keep `fsync` semantics for rollover events. +**Partial.** `.dev/reference/crates/wo-seg/src/writer.rs:92` calls `file.sync_all()` (Rust stdlib's `fsync` wrapper). Phase 11 replaces with explicit `libc::fdatasync` for the WAL path; segment files keep `fsync` semantics for rollover events. diff --git a/docs/plan/exploration/postgresql/00-postgresql.md b/docs/plan/exploration/postgresql/00-postgresql.md index 3124a94..6370778 100644 --- a/docs/plan/exploration/postgresql/00-postgresql.md +++ b/docs/plan/exploration/postgresql/00-postgresql.md @@ -1,14 +1,14 @@ # PostgreSQL — storage subsystem reference -These cards exist to make the Postgres backend a useful **library of patterns** for writeonce's persistent-storage phases (10–12) without inviting a multi-process port. Each card pulls one subsystem out of [`reference/postgresql/src/backend/`](../../../../reference/postgresql/src/backend/) — paths into the Postgres tree, the underlying *idea*, and the writeonce translation. +These cards exist to make the Postgres backend a useful **library of patterns** for writeonce's persistent-storage phases (10–12) without inviting a multi-process port. Each card pulls one subsystem out of [`.dev/reference/postgresql/src/backend/`](../../../../.dev/reference/postgresql/src/backend/) — paths into the Postgres tree, the underlying *idea*, and the writeonce translation. The symlink is user-specific: ```bash -ln -s /home/shoney/projects/postgresql reference/postgresql +ln -s /home/shoney/projects/postgresql .dev/reference/postgresql ``` -Gitignored — see [`.gitignore`](../../../../.gitignore). Pair it with [`reference/linux`](../../../../reference/linux) and [`reference/go`](../../../../reference/go) if not already linked. +Gitignored — see [`.gitignore`](../../../../.gitignore). Pair it with [`.dev/reference/linux`](../../../../.dev/reference/linux) and [`.dev/reference/go`](../../../../.dev/reference/go) if not already linked. ## Per-subsystem cards diff --git a/docs/plan/exploration/postgresql/buffer-and-checkpoint.md b/docs/plan/exploration/postgresql/buffer-and-checkpoint.md index 5496d88..b758ba9 100644 --- a/docs/plan/exploration/postgresql/buffer-and-checkpoint.md +++ b/docs/plan/exploration/postgresql/buffer-and-checkpoint.md @@ -13,12 +13,12 @@ No separate process. No shared-buffer pinning. No dynamic-shared-memory coordina | File | Responsibility | | --- | --- | -| [`storage/buffer/bufmgr.c`](../../../../reference/postgresql/src/backend/storage/buffer/bufmgr.c) | Page cache front-door: `ReadBuffer`, `BufferGetPage`, `MarkBufferDirty`, `FlushBuffer`. Tracks dirty bit per buffer; pinning prevents eviction. | -| [`storage/buffer/freelist.c`](../../../../reference/postgresql/src/backend/storage/buffer/freelist.c) | Clock-sweep eviction policy. Buffers with `usage_count = 0` and `pin_count = 0` are eviction candidates; usage decremented on every sweep pass, incremented on access. | -| [`storage/buffer/buf_table.c`](../../../../reference/postgresql/src/backend/storage/buffer/buf_table.c) | Hash table from `(file, block)` → buffer slot. The lookup that `ReadBuffer` does. | -| [`postmaster/checkpointer.c`](../../../../reference/postgresql/src/backend/postmaster/checkpointer.c) | The checkpointer process. Triggered by time (`checkpoint_timeout`), WAL volume (`max_wal_size`), or signal. Runs `BufferSync()` to flush dirty buffers, then `CreateCheckPoint()` to update the control file. | -| [`postmaster/bgwriter.c`](../../../../reference/postgresql/src/backend/postmaster/bgwriter.c) | Continuously trickles dirty pages to disk between checkpoints. Smooths the I/O burst the checkpointer would cause. | -| [`storage/buffer/README`](../../../../reference/postgresql/src/backend/storage/buffer/README) | Overview of the pinning, locking, and replacement policy. Worth reading. | +| [`storage/buffer/bufmgr.c`](../../../../.dev/reference/postgresql/src/backend/storage/buffer/bufmgr.c) | Page cache front-door: `ReadBuffer`, `BufferGetPage`, `MarkBufferDirty`, `FlushBuffer`. Tracks dirty bit per buffer; pinning prevents eviction. | +| [`storage/buffer/freelist.c`](../../../../.dev/reference/postgresql/src/backend/storage/buffer/freelist.c) | Clock-sweep eviction policy. Buffers with `usage_count = 0` and `pin_count = 0` are eviction candidates; usage decremented on every sweep pass, incremented on access. | +| [`storage/buffer/buf_table.c`](../../../../.dev/reference/postgresql/src/backend/storage/buffer/buf_table.c) | Hash table from `(file, block)` → buffer slot. The lookup that `ReadBuffer` does. | +| [`postmaster/checkpointer.c`](../../../../.dev/reference/postgresql/src/backend/postmaster/checkpointer.c) | The checkpointer process. Triggered by time (`checkpoint_timeout`), WAL volume (`max_wal_size`), or signal. Runs `BufferSync()` to flush dirty buffers, then `CreateCheckPoint()` to update the control file. | +| [`postmaster/bgwriter.c`](../../../../.dev/reference/postgresql/src/backend/postmaster/bgwriter.c) | Continuously trickles dirty pages to disk between checkpoints. Smooths the I/O burst the checkpointer would cause. | +| [`storage/buffer/README`](../../../../.dev/reference/postgresql/src/backend/storage/buffer/README) | Overview of the pinning, locking, and replacement policy. Worth reading. | ## The page-cache idea worth porting diff --git a/docs/plan/exploration/postgresql/page-format.md b/docs/plan/exploration/postgresql/page-format.md index fe06369..f82fbd3 100644 --- a/docs/plan/exploration/postgresql/page-format.md +++ b/docs/plan/exploration/postgresql/page-format.md @@ -8,11 +8,11 @@ Writeonce's phase 10 starts simpler — variable-length records, no pages. Phase | File | Responsibility | | --- | --- | -| [`storage/page/bufpage.c`](../../../../reference/postgresql/src/backend/storage/page/bufpage.c) | Page initialization (`PageInit`), line-pointer manipulation, free-space accounting. | -| [`storage/page/checksum.c`](../../../../reference/postgresql/src/backend/storage/page/checksum.c) | The page checksum algorithm — CRC32C-style with a Postgres-specific finalization. Optional, enabled at cluster init. | -| [`storage/page/itemptr.c`](../../../../reference/postgresql/src/backend/storage/page/itemptr.c) | Item pointer (`ItemPointerData`) — `(block_number, offset_within_page)` 6-byte tuple address. The on-disk equivalent of writeonce's `(TypeName, SegmentOffset)`. | -| [`include/storage/bufpage.h`](../../../../reference/postgresql/src/include/storage/bufpage.h) | The header-file definition. Read this first — it's the spec. | -| [`storage/page/README`](../../../../reference/postgresql/src/backend/storage/page/README) | One-page overview of the slotted-page model and how checksums interact with WAL. | +| [`storage/page/bufpage.c`](../../../../.dev/reference/postgresql/src/backend/storage/page/bufpage.c) | Page initialization (`PageInit`), line-pointer manipulation, free-space accounting. | +| [`storage/page/checksum.c`](../../../../.dev/reference/postgresql/src/backend/storage/page/checksum.c) | The page checksum algorithm — CRC32C-style with a Postgres-specific finalization. Optional, enabled at cluster init. | +| [`storage/page/itemptr.c`](../../../../.dev/reference/postgresql/src/backend/storage/page/itemptr.c) | Item pointer (`ItemPointerData`) — `(block_number, offset_within_page)` 6-byte tuple address. The on-disk equivalent of writeonce's `(TypeName, SegmentOffset)`. | +| [`include/storage/bufpage.h`](../../../../.dev/reference/postgresql/src/include/storage/bufpage.h) | The header-file definition. Read this first — it's the spec. | +| [`storage/page/README`](../../../../.dev/reference/postgresql/src/backend/storage/page/README) | One-page overview of the slotted-page model and how checksums interact with WAL. | ## The Postgres page header (24 bytes) diff --git a/docs/plan/exploration/postgresql/smgr-and-md.md b/docs/plan/exploration/postgresql/smgr-and-md.md index d70ff99..9600ac1 100644 --- a/docs/plan/exploration/postgresql/smgr-and-md.md +++ b/docs/plan/exploration/postgresql/smgr-and-md.md @@ -8,10 +8,10 @@ The writeonce equivalent is **per-type segment files** (`data/.seg`). | File | Responsibility | | --- | --- | -| [`storage/smgr/smgr.c`](../../../../reference/postgresql/src/backend/storage/smgr/smgr.c) | Front-door API. `smgropen`, `smgrread`, `smgrwrite`, `smgrextend`, `smgrdounlink`. Holds the `SMgrRelation` cache. | -| [`storage/smgr/md.c`](../../../../reference/postgresql/src/backend/storage/smgr/md.c) | The actual implementation against the kernel. Manages `MdfdVec` (open file descriptor handles per segment number), opens missing segments lazily. | -| [`storage/smgr/bulk_write.c`](../../../../reference/postgresql/src/backend/storage/smgr/bulk_write.c) | Optimized path for bulk-loading: writes directly to `smgrwrite` without going through shared buffers. Useful for `COPY` / `CREATE INDEX` + the recovery path's wal-replay-rebuilds-pages flow. | -| [`storage/smgr/README`](../../../../reference/postgresql/src/backend/storage/smgr/README) | Brief but worth reading — explains the relfilenode → file naming convention and how `RELSEG_SIZE` interacts with 32-bit-fs-size historical limits. | +| [`storage/smgr/smgr.c`](../../../../.dev/reference/postgresql/src/backend/storage/smgr/smgr.c) | Front-door API. `smgropen`, `smgrread`, `smgrwrite`, `smgrextend`, `smgrdounlink`. Holds the `SMgrRelation` cache. | +| [`storage/smgr/md.c`](../../../../.dev/reference/postgresql/src/backend/storage/smgr/md.c) | The actual implementation against the kernel. Manages `MdfdVec` (open file descriptor handles per segment number), opens missing segments lazily. | +| [`storage/smgr/bulk_write.c`](../../../../.dev/reference/postgresql/src/backend/storage/smgr/bulk_write.c) | Optimized path for bulk-loading: writes directly to `smgrwrite` without going through shared buffers. Useful for `COPY` / `CREATE INDEX` + the recovery path's wal-replay-rebuilds-pages flow. | +| [`storage/smgr/README`](../../../../.dev/reference/postgresql/src/backend/storage/smgr/README) | Brief but worth reading — explains the relfilenode → file naming convention and how `RELSEG_SIZE` interacts with 32-bit-fs-size historical limits. | ## What `md.c` actually does diff --git a/docs/plan/exploration/postgresql/wal.md b/docs/plan/exploration/postgresql/wal.md index e908570..b2ca7d2 100644 --- a/docs/plan/exploration/postgresql/wal.md +++ b/docs/plan/exploration/postgresql/wal.md @@ -8,11 +8,11 @@ Writeonce mirrors the algorithm. The single-thread loop replaces multi-process c | File | Responsibility | | --- | --- | -| [`access/transam/xlog.c`](../../../../reference/postgresql/src/backend/access/transam/xlog.c) | Top-level WAL machinery: insertion locks, segment rollover, flush coordination, control-file rendezvous. | -| [`access/transam/xloginsert.c`](../../../../reference/postgresql/src/backend/access/transam/xloginsert.c) | Build a WAL record (header + payload + backup-block deltas) and place it into the in-memory WAL buffer. | -| [`access/transam/xlogreader.c`](../../../../reference/postgresql/src/backend/access/transam/xlogreader.c) | Decode WAL records during recovery — pure parser, no I/O. Useful as the read-side spec. | -| [`access/transam/xlogrecovery.c`](../../../../reference/postgresql/src/backend/access/transam/xlogrecovery.c) | The replay loop. Walks the WAL from the last-checkpoint LSN, replays each record into shared buffers, advances the redo pointer. | -| [`postmaster/walwriter.c`](../../../../reference/postgresql/src/backend/postmaster/walwriter.c) | Background process that flushes the WAL buffer to disk asynchronously. Writeonce does this **inline in the loop tick**. | +| [`access/transam/xlog.c`](../../../../.dev/reference/postgresql/src/backend/access/transam/xlog.c) | Top-level WAL machinery: insertion locks, segment rollover, flush coordination, control-file rendezvous. | +| [`access/transam/xloginsert.c`](../../../../.dev/reference/postgresql/src/backend/access/transam/xloginsert.c) | Build a WAL record (header + payload + backup-block deltas) and place it into the in-memory WAL buffer. | +| [`access/transam/xlogreader.c`](../../../../.dev/reference/postgresql/src/backend/access/transam/xlogreader.c) | Decode WAL records during recovery — pure parser, no I/O. Useful as the read-side spec. | +| [`access/transam/xlogrecovery.c`](../../../../.dev/reference/postgresql/src/backend/access/transam/xlogrecovery.c) | The replay loop. Walks the WAL from the last-checkpoint LSN, replays each record into shared buffers, advances the redo pointer. | +| [`postmaster/walwriter.c`](../../../../.dev/reference/postgresql/src/backend/postmaster/walwriter.c) | Background process that flushes the WAL buffer to disk asynchronously. Writeonce does this **inline in the loop tick**. | ## The five Postgres WAL ideas writeonce keeps @@ -57,9 +57,9 @@ Same effect as Postgres' group-commit fence (one `fsync` flushes many commits) w ## Pointers when implementing phase 11 -- [`xloginsert.c:XLogInsert()`](../../../../reference/postgresql/src/backend/access/transam/xloginsert.c) — entry point for "insert this record into the WAL." Read the prologue + the LSN-assignment loop, ignore the buffer-juggling. -- [`xlog.c:XLogFlush()`](../../../../reference/postgresql/src/backend/access/transam/xlog.c) — "make this LSN durable on disk." Read the early-out for "already flushed" and the group-commit waiter logic. -- [`xlogrecovery.c:PerformWalRecovery()`](../../../../reference/postgresql/src/backend/access/transam/xlogrecovery.c) — the replay loop. Read the redo-pointer advance logic; ignore the multi-process startup signaling. +- [`xloginsert.c:XLogInsert()`](../../../../.dev/reference/postgresql/src/backend/access/transam/xloginsert.c) — entry point for "insert this record into the WAL." Read the prologue + the LSN-assignment loop, ignore the buffer-juggling. +- [`xlog.c:XLogFlush()`](../../../../.dev/reference/postgresql/src/backend/access/transam/xlog.c) — "make this LSN durable on disk." Read the early-out for "already flushed" and the group-commit waiter logic. +- [`xlogrecovery.c:PerformWalRecovery()`](../../../../.dev/reference/postgresql/src/backend/access/transam/xlogrecovery.c) — the replay loop. Read the redo-pointer advance logic; ignore the multi-process startup signaling. ## Used by diff --git a/docs/plan/exploration/ui/00-overview.md b/docs/plan/exploration/ui/00-overview.md index 9af5520..d47fd5c 100644 --- a/docs/plan/exploration/ui/00-overview.md +++ b/docs/plan/exploration/ui/00-overview.md @@ -1,12 +1,12 @@ # UI track — `.htmlx` live templates + Angular-style monorepo -**Context sources:** [`docs/examples/ecommerce/ui/`](../../examples/ecommerce/ui/) (current `##ui` screens — storefront, order_tracker, admin_orders), [`docs/examples/ecommerce/types/`](../../examples/ecommerce/types/) + [`docs/examples/ecommerce/logic/`](../../examples/ecommerce/logic/) (the shared-schema + shared-fn anchor), [`reference/crates/wo-htmlx/`](../../../reference/crates/wo-htmlx/) (v1 template engine — `{{path}}`, `{{#each}}`, `{{> partial}}`, `data-bind` attributes), [`templates/`](../../../templates/) (v1 blog's concrete `.htmlx` usage), [`docs/runtime/database/06-lowcode-fullstack.md`](../../runtime/database/06-lowcode-fullstack.md) (Phase 6's `##ui` + `##app` block spec). +**Context sources:** [`docs/examples/ecommerce/ui/`](../../examples/ecommerce/ui/) (current `##ui` screens — storefront, order_tracker, admin_orders), [`docs/examples/ecommerce/types/`](../../examples/ecommerce/types/) + [`docs/examples/ecommerce/logic/`](../../examples/ecommerce/logic/) (the shared-schema + shared-fn anchor), [`.dev/reference/crates/wo-htmlx/`](../../../.dev/reference/crates/wo-htmlx/) (v1 template engine — `{{path}}`, `{{#each}}`, `{{> partial}}`, `data-bind` attributes), [`templates/`](../../../templates/) (v1 blog's concrete `.htmlx` usage), [`docs/runtime/database/06-lowcode-fullstack.md`](../../runtime/database/06-lowcode-fullstack.md) (Phase 6's `##ui` + `##app` block spec). ## Context Three threads converge into one plan: -1. **`##ui` needs a concrete output format.** Phase 6's spec says screens "compile to a render tree" served as SSR HTML with a thin client runtime, but the actual template format isn't named. The v1 `.htmlx` engine at [`reference/crates/wo-htmlx/`](../../../reference/crates/wo-htmlx/) already speaks `{{bindings}}`, `{{#each}}`, `{{> partials}}`, and `data-bind` attributes — it's 90% of what the new runtime needs and already has a working parser + renderer. Adopting it (and extending it with live-subscription semantics) is cheaper than inventing a new format. +1. **`##ui` needs a concrete output format.** Phase 6's spec says screens "compile to a render tree" served as SSR HTML with a thin client runtime, but the actual template format isn't named. The v1 `.htmlx` engine at [`.dev/reference/crates/wo-htmlx/`](../../../.dev/reference/crates/wo-htmlx/) already speaks `{{bindings}}`, `{{#each}}`, `{{> partials}}`, and `data-bind` attributes — it's 90% of what the new runtime needs and already has a working parser + renderer. Adopting it (and extending it with live-subscription semantics) is cheaper than inventing a new format. 2. **The samples want a home that matches how real frontends are organised.** The ecommerce sample today is one flat directory with `types/`, `logic/`, and `ui/` beside each other. A real deployment has *multiple apps* against the same data: a customer storefront, an admin dashboard, a fulfillment console, maybe a read-only analytics viewer. Each has its own routes, its own policies, its own ideal binary shape. Angular (via Nx / Angular CLI workspaces) solved this with `apps/*` + `libs/*` on top of a shared root config — writeonce adopts the same shape. @@ -61,7 +61,7 @@ Read before writing each sub-phase: | Source | Why | | --- | --- | -| [`reference/crates/wo-htmlx/src/parser.rs`](../../../reference/crates/wo-htmlx/src/parser.rs) + [`render.rs`](../../../reference/crates/wo-htmlx/src/render.rs) | The v1 template engine's exact surface — what parses, what renders, what the AST looks like. ~500 LOC total. | +| [`.dev/reference/crates/wo-htmlx/src/parser.rs`](../../../.dev/reference/crates/wo-htmlx/src/parser.rs) + [`render.rs`](../../../.dev/reference/crates/wo-htmlx/src/render.rs) | The v1 template engine's exact surface — what parses, what renders, what the AST looks like. ~500 LOC total. | | [`templates/article.htmlx`](../../../templates/article.htmlx), [`templates/home.htmlx`](../../../templates/home.htmlx) | Concrete usage of the v1 format — how `{{path}}` and `data-bind` actually read in real templates. | | [`docs/examples/ecommerce/ui/{storefront,order_tracker,admin_orders}.wo`](../../examples/ecommerce/ui/) | The `##ui` side — what the declarative DSL promises to produce. These screens are the target of the first compiler pass. | | [`docs/runtime/database/06-lowcode-fullstack.md`](../../runtime/database/06-lowcode-fullstack.md) | Phase 6's full-stack block spec — `##ui`, `##app`, `##policy`, `##service`, `##logic` — already designed but not yet compiled. | @@ -209,5 +209,5 @@ After all seven sub-phases land: - [`../../runtime/database/06-lowcode-fullstack.md`](../../runtime/database/06-lowcode-fullstack.md) — Phase 6's full-stack block spec that this track implements. - [`../../runtime/database/04-client-api.md`](../../runtime/database/04-client-api.md) — the wire protocol per-app binaries speak to the shared DB over. - [`../../examples/ecommerce/ui/admin_orders.wo`](../../examples/ecommerce/ui/admin_orders.wo) — the motivating workload: a live ops table bound to the order stream. -- [`reference/crates/wo-htmlx/`](../../../reference/crates/wo-htmlx/) — the template engine ~90% of this track will reuse. +- [`.dev/reference/crates/wo-htmlx/`](../../../.dev/reference/crates/wo-htmlx/) — the template engine ~90% of this track will reuse. - [`templates/`](../../../templates/) — v1 blog's actual `.htmlx` files; the format this track extends. diff --git a/docs/plan/exploration/ui/01-htmlx-format-spec.md b/docs/plan/exploration/ui/01-htmlx-format-spec.md index aee9e1f..f6e2a75 100644 --- a/docs/plan/exploration/ui/01-htmlx-format-spec.md +++ b/docs/plan/exploration/ui/01-htmlx-format-spec.md @@ -1,6 +1,6 @@ # 01 — `.htmlx` format spec -**Context sources:** [`./00-overview.md`](./00-overview.md) §§ "`.htmlx` with live subscriptions — target format" (L127–166), "Design decisions" (L28–37), [`reference/crates/wo-htmlx/`](../../../reference/crates/wo-htmlx/) (the v1 template engine that 90% of this phase ports), [`templates/article.htmlx`](../../../templates/article.htmlx) and [`templates/home.htmlx`](../../../templates/home.htmlx) (v1 concrete usage), [`docs/examples/ecommerce/shared/components/order-row.htmlx`](../../examples/ecommerce/shared/components/order-row.htmlx) (the live-binding workload this format must serve). +**Context sources:** [`./00-overview.md`](./00-overview.md) §§ "`.htmlx` with live subscriptions — target format" (L127–166), "Design decisions" (L28–37), [`.dev/reference/crates/wo-htmlx/`](../../../.dev/reference/crates/wo-htmlx/) (the v1 template engine that 90% of this phase ports), [`templates/article.htmlx`](../../../templates/article.htmlx) and [`templates/home.htmlx`](../../../templates/home.htmlx) (v1 concrete usage), [`docs/examples/ecommerce/shared/components/order-row.htmlx`](../../examples/ecommerce/shared/components/order-row.htmlx) (the live-binding workload this format must serve). ## Goal @@ -8,7 +8,7 @@ Lock the exact `.htmlx` grammar — every v1 Mustache construct unchanged plus t ## Design decisions (locked) -1. **Mustache constructs carry through unchanged.** `{{path}}`, `{{#each xs as y}}…{{/each}}`, `{{#if cond}}…{{/if}}`, `{{#when cond}}…{{/when}}`, `{{> partial arg=val}}`. The v1 parser already handles all of these; the new parser inherits them verbatim. See [`reference/crates/wo-htmlx/src/parser.rs`](../../../reference/crates/wo-htmlx/src/parser.rs) (173 LOC) and the AST in [`ast.rs`](../../../reference/crates/wo-htmlx/src/ast.rs) (18 LOC). +1. **Mustache constructs carry through unchanged.** `{{path}}`, `{{#each xs as y}}…{{/each}}`, `{{#if cond}}…{{/if}}`, `{{#when cond}}…{{/when}}`, `{{> partial arg=val}}`. The v1 parser already handles all of these; the new parser inherits them verbatim. See [`.dev/reference/crates/wo-htmlx/src/parser.rs`](../../../.dev/reference/crates/wo-htmlx/src/parser.rs) (173 LOC) and the AST in [`ast.rs`](../../../.dev/reference/crates/wo-htmlx/src/ast.rs) (18 LOC). 2. **`` is a parsed structured node, not HTML passthrough.** The parser recognises the `` contains another ``. 3. **`wo:bind="field"` is an HTML attribute, parsed but emitted verbatim.** SSR writes the attribute through; the consumer is the client runtime. The parser records each `(element, field)` pair into the manifest; nothing else changes about element rendering. 4. **Helpers are a closed Rust enum.** v1 invocation forms (`{{relative ts}}`, `{{#if (eq for "ops")}}`, `{{> money amount=x}}`) carry through. The registered set is fixed for this phase: `relative`, `eq`, `markdown`, `code`, `money`, `tag-chips`, `pill`, `image`, `stock-badge`, `list`. No author extensibility. @@ -20,12 +20,12 @@ Lock the exact `.htmlx` grammar — every v1 Mustache construct unchanged plus t | File | Responsibility | Port source | | --- | --- | --- | -| `mod.rs` | Re-exports `Template`, `Manifest`, `LiveSubscription`, `BindSite`, `ParseError`, `RenderError` | [`reference/crates/wo-htmlx/src/lib.rs`](../../../reference/crates/wo-htmlx/src/lib.rs) (11 LOC) | -| `ast.rs` | Adds `Node::Live { attrs, body }` and `wo_bind: Option` on element nodes | [`reference/crates/wo-htmlx/src/ast.rs`](../../../reference/crates/wo-htmlx/src/ast.rs) (18 LOC) — extend by ~50 LOC | -| `parser.rs` | Adds `` body in `
` for the runtime | [`reference/crates/wo-htmlx/src/render.rs`](../../../reference/crates/wo-htmlx/src/render.rs) (140 LOC) — extend by ~70 LOC | +| `mod.rs` | Re-exports `Template`, `Manifest`, `LiveSubscription`, `BindSite`, `ParseError`, `RenderError` | [`.dev/reference/crates/wo-htmlx/src/lib.rs`](../../../.dev/reference/crates/wo-htmlx/src/lib.rs) (11 LOC) | +| `ast.rs` | Adds `Node::Live { attrs, body }` and `wo_bind: Option` on element nodes | [`.dev/reference/crates/wo-htmlx/src/ast.rs`](../../../.dev/reference/crates/wo-htmlx/src/ast.rs) (18 LOC) — extend by ~50 LOC | +| `parser.rs` | Adds `` body in `
` for the runtime | [`.dev/reference/crates/wo-htmlx/src/render.rs`](../../../.dev/reference/crates/wo-htmlx/src/render.rs) (140 LOC) — extend by ~70 LOC | | `manifest.rs` | Walks the AST, collects subscriptions + bind sites, serialises JSON | new (~150 LOC) | Total: ~835 LOC (585 ported + ~250 new). @@ -103,7 +103,7 @@ cargo test -p ui --test manifest # manifest emission cargo run --bin wo -- run docs/examples/blog & PID=$!; sleep 1; curl -fsS http://127.0.0.1:8080/ >/dev/null; kill $PID -cd reference/crates && cargo build && cargo test +cd .dev/reference/crates && cargo build && cargo test ``` ## After this phase diff --git a/docs/plan/exploration/ui/02-ui-compiler.md b/docs/plan/exploration/ui/02-ui-compiler.md index bf33e9c..0fbc82a 100644 --- a/docs/plan/exploration/ui/02-ui-compiler.md +++ b/docs/plan/exploration/ui/02-ui-compiler.md @@ -60,7 +60,7 @@ for (path, src) in outputs { fs::write(path, src)?; } 3. **Hand-written fallback honoured.** With a hand-written `apps/admin/ui/orders/orders.htmlx` present, the compiler returns its source unchanged but still emits the manifest. 4. **Manifest cross-check fires.** Renaming `body` to `text` in a hand-written template that the `##ui` block expects under `wo:bind="body"` produces a `CompileError::HandWrittenMissingField` diagnostic. 5. **Parser change is non-breaking.** `crates/rt`'s 14 unit tests still pass; `cargo run --bin wo -- run docs/examples/blog` boots and serves REST as before. -6. `cd reference/crates && cargo build && cargo test`. +6. `cd .dev/reference/crates && cargo build && cargo test`. ## Non-scope @@ -87,7 +87,7 @@ head -1 target/wo/storefront/ui/orders.htmlx # starts with ` body with the fresh snapshot. No diff, no replay buffer. 5. **Backpressure = drop all but latest update per `data-key`.** A coalescing queue keyed by `(subscription_id, key)` collapses queued `update` frames; the latest wins. New frames of other kinds (`insert`/`delete`) flush the queue. @@ -73,7 +73,7 @@ ws.send_text(serde_json::to_string(&frame)?)?; 3. **DOM-patch test (jsdom).** `node crates/ui/runtime-tests/run.mjs` loads a stub HTML containing one `` block and a manifest, fakes a WebSocket emitting `snapshot` → `insert` → `update` → `delete` frames, and asserts each patch hits the right element. 4. **Reconnect test.** Killing the fake WS triggers exponential backoff; on resume the runtime re-issues subscriptions and replaces the body with the new snapshot. 5. **Asset served.** Once phase 05 lands, `curl http://127.0.0.1:8080/_wo/runtime.js` returns the file with a stable `ETag` matching `sha256(RUNTIME_JS)`. -6. `cd reference/crates && cargo build && cargo test`. +6. `cd .dev/reference/crates && cargo build && cargo test`. ## Non-scope @@ -101,7 +101,7 @@ test "$(wc -c < crates/ui/assets/wo-runtime.js)" -le 25600 cargo run --bin wo -- run docs/examples/blog & PID=$!; sleep 1; curl -fsS http://127.0.0.1:8080/ >/dev/null; kill $PID -cd reference/crates && cargo build && cargo test +cd .dev/reference/crates && cargo build && cargo test ``` ## After this phase diff --git a/docs/plan/exploration/ui/04-workspace-layout.md b/docs/plan/exploration/ui/04-workspace-layout.md index e7214a4..df56dcc 100644 --- a/docs/plan/exploration/ui/04-workspace-layout.md +++ b/docs/plan/exploration/ui/04-workspace-layout.md @@ -102,12 +102,12 @@ assert_eq!(blog.apps().len(), 1); 3. **Component resolution.** `storefront.resolve_component("money")` returns the path to `shared/components/money.htmlx`. `storefront.resolve_component("nonsense")` errors as `ResolverError::NotFound`. 4. **App-local override.** Adding `apps/storefront/ui/components/money.htmlx` makes `resolve_component("money")` return the app-local path; removing it falls back to the shared one. 5. **Degenerate form.** `Workspace::load(docs/examples/blog)` loads as a one-app workspace; `wo run docs/examples/blog` continues to start unchanged. -6. `cd reference/crates && cargo build && cargo test`. +6. `cd .dev/reference/crates && cargo build && cargo test`. ## Non-scope - **No semver, no registry, no lockfile.** Path references only. -- **No `wo dev` hot-reload.** File watching against `apps/*/ui/` is deferred (would consume `reference/crates/wo-watch/`). +- **No `wo dev` hot-reload.** File watching against `apps/*/ui/` is deferred (would consume `.dev/reference/crates/wo-watch/`). - **No cross-workspace symlinks.** `shared = […]` paths must resolve under the workspace root. - **No build-time enforcement that an app touches only its declared shared dirs.** That's an integrity check for a later hardening phase. - **No env-var interpolation in `wo.toml`.** `${VAR}` syntax stays out; runtime config comes through env vars at startup, not manifest time. @@ -131,7 +131,7 @@ cargo run --bin wo -- ls-apps docs/examples/blog cargo run --bin wo -- run docs/examples/blog & PID=$!; sleep 1; curl -fsS http://127.0.0.1:8080/ >/dev/null; kill $PID -cd reference/crates && cargo build && cargo test +cd .dev/reference/crates && cargo build && cargo test ``` ## After this phase diff --git a/docs/plan/exploration/ui/05-per-app-binaries.md b/docs/plan/exploration/ui/05-per-app-binaries.md index 5547b02..1620ecd 100644 --- a/docs/plan/exploration/ui/05-per-app-binaries.md +++ b/docs/plan/exploration/ui/05-per-app-binaries.md @@ -85,7 +85,7 @@ fn main() -> Result<()> { 3. **Storefront boots.** `WO_DB=wo://127.0.0.1:5555 STOREFRONT_DB_KEY=test ./target/wo/storefront &` then `curl -fsS http://127.0.0.1:8080/healthz` returns `200`. (The DB daemon from phase 06 is mocked or stubbed for this test if 06 hasn't landed yet — refuse-to-start without DB is the contract; the test verifies refuse-to-start when `WO_DB` is unset.) 4. **Admin builds separately.** `wo build apps/admin` produces a *different* binary with a disjoint route table. Diffing the two `app_config.rs` files shows different route lists. 5. **Refuse-to-start without DB.** `./target/wo/storefront` with no `WO_DB` and no manifest URL exits non-zero with a clear error. -6. `cd reference/crates && cargo build && cargo test`. +6. `cd .dev/reference/crates && cargo build && cargo test`. ## Non-scope @@ -116,7 +116,7 @@ test -x target/wo/admin cargo run --bin wo -- run docs/examples/blog & PID=$!; sleep 1; curl -fsS http://127.0.0.1:8080/ >/dev/null; kill $PID -cd reference/crates && cargo build && cargo test +cd .dev/reference/crates && cargo build && cargo test ``` ## After this phase diff --git a/docs/plan/exploration/ui/06-shared-db-daemon.md b/docs/plan/exploration/ui/06-shared-db-daemon.md index ea47d5a..c3b221b 100644 --- a/docs/plan/exploration/ui/06-shared-db-daemon.md +++ b/docs/plan/exploration/ui/06-shared-db-daemon.md @@ -1,6 +1,6 @@ # 06 — Shared database daemon (`wo db serve`) -**Context sources:** [`./00-overview.md`](./00-overview.md) §§ "Goal" (L23), "Design decisions" 3 (L32), "Non-scope" (L201–203), [`./03-client-runtime.md`](./03-client-runtime.md) (the wire frames this daemon emits), [`./05-per-app-binaries.md`](./05-per-app-binaries.md) (the apps that connect), [`reference/crates/wo-sub/src/lib.rs`](../../../reference/crates/wo-sub/src/lib.rs) (the v1 subscription registry, 470 LOC, that needs generalising past `ByTitle`/`ByTag`/`All`), [`../../runtime/database/04-client-api.md`](../../runtime/database/04-client-api.md) (the wire-protocol owner). +**Context sources:** [`./00-overview.md`](./00-overview.md) §§ "Goal" (L23), "Design decisions" 3 (L32), "Non-scope" (L201–203), [`./03-client-runtime.md`](./03-client-runtime.md) (the wire frames this daemon emits), [`./05-per-app-binaries.md`](./05-per-app-binaries.md) (the apps that connect), [`.dev/reference/crates/wo-sub/src/lib.rs`](../../../.dev/reference/crates/wo-sub/src/lib.rs) (the v1 subscription registry, 470 LOC, that needs generalising past `ByTitle`/`ByTag`/`All`), [`../../runtime/database/04-client-api.md`](../../runtime/database/04-client-api.md) (the wire-protocol owner). ## Goal @@ -10,7 +10,7 @@ Stand up a headless daemon — `wo db serve` — that runs the engine + WAL + su 1. **Daemon = `crates/db` thin entrypoint + `crates/engine` + the wire acceptor.** No HTTP, no `.htmlx`, no `##ui`. The shared DB process knows nothing about the UI layer. 2. **API-key table is in-memory, env-seeded.** On startup the daemon reads `WO_DB_KEY_=` for each app declared in the workspace and builds an `AuthTable: HashMap`. A `--keys ` flag is accepted but treated as a future hook. -3. **Generalise `wo-sub`** from `Subscription::ByTitle/ByTag/All` to `Subscription::ByPredicate(TypeRef, Expr, SortKey)`. The v1 variants stay as legacy aliases (`ByTitle(t)` ⇒ `ByPredicate(Article, sys_title == t, _)`) for the blog regression test. Anchored in [`reference/crates/wo-sub/src/lib.rs`](../../../reference/crates/wo-sub/src/lib.rs) L8–17. +3. **Generalise `wo-sub`** from `Subscription::ByTitle/ByTag/All` to `Subscription::ByPredicate(TypeRef, Expr, SortKey)`. The v1 variants stay as legacy aliases (`ByTitle(t)` ⇒ `ByPredicate(Article, sys_title == t, _)`) for the blog regression test. Anchored in [`.dev/reference/crates/wo-sub/src/lib.rs`](../../../.dev/reference/crates/wo-sub/src/lib.rs) L8–17. 4. **Connection scope = `Principal { app, roles }` stored on the connection.** Every query evaluator reads it; phase 07 wires it into policy AND-composition. 5. **One data dir, one engine, many connections.** Snapshot isolation by default (per `[database].isolation = "snapshot"` in the workspace `wo.toml`). 6. **Foreground-only this phase.** No daemonisation, no PID file, no signal handling beyond `SIGTERM` graceful shutdown. A future ops doc can add `wo db daemonize`. @@ -24,7 +24,7 @@ Stand up a headless daemon — `wo db serve` — that runs the engine + WAL + su | `crates/db/src/main.rs` | Entrypoint, arg parsing, env-key loading | new (~100 LOC) | | `crates/db/src/server.rs` | Wire-protocol acceptor (TCP listener + per-conn handler) | new (~250 LOC) | | `crates/db/src/auth.rs` | `AuthTable`, `Principal`, key handshake | new (~120 LOC) | -| `crates/sub/src/lib.rs` | Generalised subscription manager | port [`reference/crates/wo-sub/src/lib.rs`](../../../reference/crates/wo-sub/src/lib.rs) (470 LOC) + ~150 new | +| `crates/sub/src/lib.rs` | Generalised subscription manager | port [`.dev/reference/crates/wo-sub/src/lib.rs`](../../../.dev/reference/crates/wo-sub/src/lib.rs) (470 LOC) + ~150 new | | `crates/sub/src/predicate.rs` | Predicate evaluation against a row (uses `crates/ql` if available, else minimal subset) | new (~150 LOC) | Total: ~1240 LOC (470 ported + ~770 new). @@ -81,7 +81,7 @@ let id = subs.register(conn_fd, sub)?; 3. **Two principals.** Two clients connect, one with each API key; each receives a distinct `Principal` in the `WELCOME` frame. 4. **Predicate subscription.** Client registers `Subscription::ByPredicate(Order, "status != Cancelled", "placed_at desc")`; the manager returns a fresh `subscription_id`; on a stub `Order` insert, the matching client receives an `insert` frame. 5. **v1 regression.** A connection running the legacy `Subscription::ByTitle("hello-world")` against the blog corpus still produces notifications via the legacy alias. -6. `cd reference/crates && cargo build && cargo test`. +6. `cd .dev/reference/crates && cargo build && cargo test`. ## Non-scope @@ -112,7 +112,7 @@ kill $DB_PID # legacy v1 path cargo test -p sub --test legacy_by_title -cd reference/crates && cargo build && cargo test +cd .dev/reference/crates && cargo build && cargo test ``` ## After this phase diff --git a/docs/plan/exploration/ui/07-per-app-policies.md b/docs/plan/exploration/ui/07-per-app-policies.md index 9608e26..cab1027 100644 --- a/docs/plan/exploration/ui/07-per-app-policies.md +++ b/docs/plan/exploration/ui/07-per-app-policies.md @@ -65,7 +65,7 @@ assert!(rs.contains(Role::Ops)); 3. **Build-time domain check fires.** A test workspace where `apps/storefront/app.wo` declares `role: Anonymous` against a type whose global policy does not define `Anonymous` — `wo build apps/storefront` exits non-zero with `PolicyDomainError`. 4. **Cross-app integration.** Two storefront customers issue the same `GET /api/orders` against the daemon; each sees only their own rows (storefront app-scope narrows global). Admin sees both. Test runs against the phase-06 daemon. 5. **v1 regression.** `wo run docs/examples/blog` boots; the global `policy read for anyone when published == true` on the blog `Article` type continues to gate anonymous reads as it does today. -6. `cd reference/crates && cargo build && cargo test`. +6. `cd .dev/reference/crates && cargo build && cargo test`. ## Non-scope @@ -97,7 +97,7 @@ curl -fsS http://127.0.0.1:8080/api/articles # only published r test -z "$(curl -fsS http://127.0.0.1:8080/api/articles | grep '"published":false')" kill $PID -cd reference/crates && cargo build && cargo test +cd .dev/reference/crates && cargo build && cargo test ``` ## After this phase diff --git a/docs/plan/exploration/ui/08-mvc-structure.md b/docs/plan/exploration/ui/08-mvc-structure.md index adfccdc..f177e44 100644 --- a/docs/plan/exploration/ui/08-mvc-structure.md +++ b/docs/plan/exploration/ui/08-mvc-structure.md @@ -1,6 +1,6 @@ # 08 — MVC structure: model = class, view = htmlx + scss, controller = .wo -**Context sources:** [`reference/writeonce-app/src/app/`](../../../../reference/writeonce-app/src/app/) (the v1 Angular app whose component anatomy this formalizes), [`./00-overview.md`](./00-overview.md) ("Angular-component-style layout" — `home/{home.wo, home.htmlx, home.css}`), [`./01-htmlx-format-spec.md`](./01-htmlx-format-spec.md) (the view grammar: Mustache + `` + `wo:bind`), [`./02-ui-compiler.md`](./02-ui-compiler.md), [`./03-client-runtime.md`](./03-client-runtime.md), [`../../13-class-model-live-pricing.md`](../../13-class-model-live-pricing.md) (the class methods controllers call), [`../../../examples/pricing/ui/pricing/`](../../../examples/pricing/ui/pricing/) (the reference screen). +**Context sources:** [`.dev/reference/writeonce-app/src/app/`](../../../../.dev/reference/writeonce-app/src/app/) (the v1 Angular app whose component anatomy this formalizes), [`./00-overview.md`](./00-overview.md) ("Angular-component-style layout" — `home/{home.wo, home.htmlx, home.css}`), [`./01-htmlx-format-spec.md`](./01-htmlx-format-spec.md) (the view grammar: Mustache + `` + `wo:bind`), [`./02-ui-compiler.md`](./02-ui-compiler.md), [`./03-client-runtime.md`](./03-client-runtime.md), [`../../13-class-model-live-pricing.md`](../../13-class-model-live-pricing.md) (the class methods controllers call), [`../../../examples/pricing/ui/pricing/`](../../../examples/pricing/ui/pricing/) (the reference screen). ## Goal @@ -15,7 +15,7 @@ ui/pricing/ ## The mapping, against the v1 Angular app -| MVC role | v1 Angular (`reference/writeonce-app/src/app/`) | writeonce | +| MVC role | v1 Angular (`.dev/reference/writeonce-app/src/app/`) | writeonce | | --- | --- | --- | | **Model** | `models/article.ts` (interface) + `services/article.service.ts` (HTTP fetch) | the `class` / `type` declaration itself (`types/product.wo`). No service layer: the database is in-process, and a model binding **is** a query — `LIVE select` for push, `select` for snapshot | | **View** | `article.component.html` + `article.component.css` | `pricing.htmlx` + `pricing.scss`. Plain markup; the only dynamic constructs are Mustache paths and `` / `wo:bind` from [`01-htmlx-format-spec.md`](./01-htmlx-format-spec.md) | @@ -72,7 +72,7 @@ browser action wo:action="set-price" ## Migration note -The two existing screen specs (`docs/examples/ecommerce/apps/*/ui/*/`, single-file `##ui` shorthand) stay valid under decision 5. New screens — starting with [`docs/examples/pricing/ui/pricing/`](../../../examples/pricing/ui/pricing/) — use the triplet. The v1 Angular app stays archived; its components are the *shape* reference, not a port source (the htmlx port source remains `reference/crates/wo-htmlx`). +The two existing screen specs (`docs/examples/ecommerce/apps/*/ui/*/`, single-file `##ui` shorthand) stay valid under decision 5. New screens — starting with [`docs/examples/pricing/ui/pricing/`](../../../examples/pricing/ui/pricing/) — use the triplet. The v1 Angular app stays archived; its components are the *shape* reference, not a port source (the htmlx port source remains `.dev/reference/crates/wo-htmlx`). ## Exit criteria (implementation sequenced in [plan 14](../../14-mvc-ui-implementation.md), landing with plan 13d)