Commit graph

28 commits

Author SHA1 Message Date
258222c3a1 feat(lang42): proc.run parks — pidfd + epoll bundle + child registry
- deadlock proven first: chatty child (200 KB stdout, stderr held open)
  hung the old sequential drain 5.0 s into the alarm, code -1, stdout
  truncated at 8192; the leg demands completion under 4 s
- rework: nonblocking pipe read ends + pidfd_open behind one epoll fd the
  fiber parks on (the _dl retry mould); both pipes drain on readiness, so
  the deadlock is gone structurally — leg passes in 15 ms
- wo_child slot table in wo_vm (32/shard) carries cross-park state; caps
  refuse by name (kill + WO_T_IO), deadline armed via dl_active/dl_at,
  defaults 30 s / 1 MiB / 64 KiB
- WO_B_PROC_RUN_DL = 96 shares the case (per-call deadline_ms/out_cap/
  err_cap; compiler row lands in a later task)
- fib_reap kills a reaped fiber's child; wo_vm_destroy sweeps the table
- raw syscalls for pidfd_open/pidfd_send_signal: glibc 2.35 build floor
  has no wrappers
- all 19 suites green under ASan+UBSan

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-09-01 22:19:25 +02:00
e16d4896f8 feat(db2-keys): inserts and boot — payload dropped after the barrier
databasev2 2, task 5c. The write and boot halves. Still not exposed: the
loader refuses `resident: keys` until 5d rewires the readers.

GROUP COMMIT FORCED THE DESIGN. A keys-resident payload can only be dropped
once its record is durable, but databasev2 4 deferred the barrier to the drain
— so at append time the bytes are still in the staging buffer and the recorded
offset would pread ZEROS. Dropping at append would have produced rows that
read as garbage, intermittently, only under multi-shard load.

So the drop is recorded, not performed:

- wo_wal gains a pending-drop list, the same shape as the drain's held replies
  and for the same reason
- both write paths take the offset BEFORE the append (wo_wal_next_offset) and
  record it; the inline path flushes right after its own commit, the request
  path's flush runs in the drain immediately after the barrier
- if the process dies before the barrier the list dies with it, which is
  correct: nothing was dropped and nothing was lost
- an out-of-memory pend is ignored on purpose — the row simply stays resident,
  which is safe

Boot: replay now leaves a keys-resident table pointing at the LOG. Each record
is applied normally, so indexes and uniqueness are built exactly as for any
other table, and the payload is then dropped with THAT record's offset. For an
update the later record wins, because each apply overwrites the map in order —
the rule replay already follows.

Tests: the round trip (insert, commit, drop, read back with Text intact) and
now BOOT — a fresh wo_db replays the store and every row materialises from the
log, count intact, nothing in a slab.

Verified: just wovm-test — 36 suites 0 fail, test_wal 4301 pass, cli_smoke OK.

REMAINING (5d), and precise: every reader still goes through wo_row_ptr, which
for a keys table would index a freed slot. The scans in db.c walk the BITMAP,
and a keys table's bitmap is empty by construction — so a query over one would
today return no rows at all. That, FK restrict, and the @unique shadow are 5d,
and the loader refusal stays until they land.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
(cherry picked from commit 08abd09bf88918f2582e74713dc7903beb8aaeb8)
2026-08-30 20:36:31 +02:00
aa89fb6c97 feat(db): the checkpoint trigger, and compaction is wired to BOTH write paths — T3
databasev2 3, task 3.

- wo_wal_should_compact is a PURE decision (used bytes, last compaction's
  measured output, floor, ratio) so it is testable without a store —
  which is the only way a policy like this gets tested at all. Denominator
  is the last compaction's real output, not an estimate of the live set:
  estimating would mean estimating Text
- 8 boundary assertions incl. "exactly 3x is not MORE than 3x" and a zero
  ratio disabling the policy rather than dividing by nothing
- MUTATION-TESTED instead of observing RED: implementation and test were
  written together, so removing the floor check was verified to fail
  exactly the two floor assertions. Equivalent evidence, stated plainly
- WO_CHECKPOINT_BYTES / WO_CHECKPOINT_RATIO at boot beside WO_MAILBOX.
  The knobs are what make the policy testable — a gate sets a tiny floor
  and forces compaction in a few writes instead of megabytes
- NO timer, per the spec: Postgres' CheckPointTimeout bounds loss from
  unflushed buffers; our records are durable at commit and an idle log
  does not grow
- the ordering rule is now asserted, not trusted: a test stages a record,
  requests compaction, and requires REFUSAL with the log untouched and
  the staged record still committable afterwards

FOUND AND FIXED a gap in my own wiring. The plan said to call the check
"after the drain's barrier", and I did — but a statement running ON the
owner shard never enters that drain, so WO_SHARDS=1 never compacted and
its log grew forever: measured 536086 bytes where the multi-shard run
held 446024. Now checked after the inline path's commit too (db.c
maybe_compact), where the buffer is equally empty. WO_SHARDS=1 went
536086 -> 260657 bytes. For a checkpoint this mattered more than part A's
equivalent gap: an unbounded log is an operational failure, not just lost
throughput.

Also corrected a measurement of my own: multi-shard logs looked unbounded
(448KB -> 1013KB -> 1647KB across 8k/24k/48k updates). They are not.
Instrumentation showed compaction ran 25 times with zero failures, each
writing MORE than the last, because the live set genuinely grows — wmix's
hist_dump and done-markers are themselves durable inserts. Final log
1631040 against a last compaction of 866432 is a ratio of 1.88, just under
the 2x threshold: the policy holding exactly.

Replies are released BEFORE compaction runs, deliberately: their records
are already durable, and holding them across a stop-the-world rewrite
would add its full duration to their latency for nothing.

Verified: wovm-test 36 suites 0 fail, test_wal 360 pass; db-bench-quick
crash.s1/crash.sN and both restart legs green, and part A still batches
(sN mean 4.16, peak 30).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-28 18:43:44 +02:00
0b618ace19 docs+fix(db): T6 closeout — and reads no longer wait for the barrier
databasev2 4 part A, task 6. Mostly documentation, plus one real fix the
full battery caught.

THE FIX. The drain held EVERY DB reply until the barrier — including
reads, which stage nothing and have no stake in durability. That parked
readers behind an fsync for no reason: durable.sN.mixread.p99 rose from
~1043us to 4057us. Only a statement that actually staged a record now has
its reply held. Caught by the gate, not by review.

THE TRADE, recorded rather than smoothed over. What remains is inherent: a
barrier blocks the owner shard LONGER (more records per fsync) though LESS
OFTEN, so anything queued behind one waits. Three full runs of the same
build gave durable.sN.mixread.p99 of 1043 / 2318 / 4147us and wmix.p99 of
8758 / 20000us — a 2-4x spread with the box near idle. So part A buys ~3x
write throughput at the cost of a longer, noisier tail on the owner shard,
and that is the strongest argument for part B (submit and keep serving).

- durable.sN.*.p99us tolerance widened to 100% WITH the reason in the
  code: a 2-4x-variable tail gated at 50% gates the disk, not the engine.
  The floor is the real guard and is not slack — mixread's (4172us) came
  within 25us of tripping on the worst run. Baseline refreshed; a fresh
  full run then passed 106 checks 0 failures

EXIT STATUS MOVED 3 -> 74 (sysexits EX_IOERR). 3 and 4 are already used by
SAMPLES for their own meanings — db-bench's own `verify` exits 3 on a
checksum mismatch, and it is the gate that exercises durability, so a
durability abort exiting 3 would have been indistinguishable from the
mismatch it should help diagnose. The low range belongs to programs.

Docs:

- story: progress, the payoff measured two ways, the cost side, criteria
  split met/outstanding, and a "part B — its premise changed" section:
  it was justified by "close the 66x gap", but that gap is two problems
  and only the concurrent one was a batching problem
- board: standup entry in the six-question shape; both databasev2 4 rows
  rewritten. They had said "close the 66x gap" — recorded as MIS-STATED
  rather than quietly renumbered
- 00-wob-format.md and 04-db-binding.md: the normative failure contract
  ("a failed WAL commit traps WO_T_IO after un-applying the row") was
  false; corrected, along with the tick-scoped group commit that never
  happened
- database/src/CODE-LOGIC.md: where the barrier runs and why there, why
  replies are held, why the inline path is asymmetric, the one failure
  rule, and how to measure it
- db-bench README: the wmix mode, the env knobs, and the tmpfs warning

Battery: wovm-test 36 suites 0 fail, woc-test, oop-e2e 119/0,
db-bench 106/0, linkcheck clean.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-28 16:48:23 +02:00
76d80cc027 feat(db): one barrier per drain, replies held — T2
databasev2 4 part A, task 2. The core change, and mostly deletion.

- the REQUEST path (wo_db_exec_req) no longer commits after each append.
  Applying to RAM and staging stay exactly where they were
- wo_vm_adopt holds each DB reply envelope in a local FIFO instead of
  pushing it as the statement finishes. Pushing there would unpark the
  requester before its record is durable — the ack contract this
  iteration exists to make literally true rather than true by accident
  of every batch having one member
- at the end of the drain: ONE wo_wal_commit_fatal for everything staged,
  then every held reply. Locals rather than per-shard state: nothing
  needs to outlive the batch it describes
- "did this statement stage anything" is asked of the buffer, not guessed
  from the opcode, and that count is what the failure diagnostic reports
- the drain commits unconditionally when anything is staged, because the
  inline path relies on finding the buffer empty (task 3 documents that)
- staging failure on the request path is now FATAL via wo_wal_stage_fatal:
  the row is already in RAM and of the three verbs only insert could undo
  itself, so continuing means RAM ahead of disk. One rule
- wal_die is now shared by both fatal points

Verified — the ack contract is the thing that could break, so it is what
was tested:

- just wovm-test: 36 suites (18 x both dispatch flavors) 0 fail, cli_smoke OK
- just db-bench-quick: 85 checks, 0 failures. The legs that matter:
  crash.sN.0 — 612 acked rows all present after kill -9, which is the
  BATCHING path (multi-shard requests, held replies, one barrier);
  crash.s1.0 — 800 acked rows; restart.s1 and restart.sN replay byte-true

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-28 09:25:10 +02:00
60414a1754 feat(runtime): the shutdown drain guarantee — iteration 40
A message sent before the stop flag is observed must be delivered and run
before the engine stops. One rule; a spin count could never express it.

- root cause in `shard_main` (runtime/src/vm.c): NEXT_RUNNABLE() already
  stated the contract — "a WORKER on stop keeps DRAINING ... so queued
  shutdown messages (close frames!) still run" — but the IDLE branch
  contradicted it, calling fib_reap_all and breaking on WO_IO_STOP,
  abandoning its inbox for wo_engine_stop() to free wholesale
- an actor between messages is exactly that idle case, which is why a WARM
  soak server hid it: warm shards held live fibers and took the right path
- fix: while the primary's drain window is open, an idle worker adopts its
  inbox and runs what arrives; sched_yield on an empty poll so a drain
  cannot burn a core per shard and starve the actors it exists to let run
- unreachable at WO_SHARDS=1: wo_engine_stop returns early at nshards <= 1

Measured:

- fresh-server SIGTERM drain: 5 of 16 failing before, 20 of 20 clean after
- `just chat` at the FULL 1000-client soak: 11 checks, 0 failures, both
  WO_IO backends, ASan clean with zero leaks
- the fd leg settled at scale too: 1000 connections left the count at 44,
  unchanged after 20 more — lazy per-shard init, not a leak
- runtime battery 36 suites (18 x both dispatch flavors) 0 fail;
  compiler 556 checks 0 fail

- story: docs/stories/language-runtime-database/40-shutdown-drain-guarantee.md
  (chain 3 with 31, status done), board row, slice marker updated
- outstanding and named: a pin below the gate needs new multithreaded test
  infrastructure — nothing in runtime/test/ drives wo_engine_start/stop and
  no corpus fixture can trigger a stop

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-27 23:43:45 +02:00
735fd270db feat: chat sample + gate (T8/T9, IN PROGRESS) + stop-drain semantics
- docs/examples/chat: registry (call consumer) / room / reader+writer
  actor pair per connection over ws_accept + wsframe; presence,
  broadcast, cross-room isolation, mailbox-full = drop-from-room;
  reader tail sends hardened (a full writer no longer orphans the fd)
- RUNTIME SEMANTICS CHANGE (the drain): SIGTERM no longer kills parked
  fibers from outside — the plane WAKES them and each wait RESOLVES
  (deadline'd waits answer their timeout result, sleeps return early,
  plain waits answer WO_SYS_STOPPED and unwind THAT fiber alone; main's
  STOPPED still ends the program). Workers keep adopting their inboxes
  after stop until eng_shutdown. This is what lets a program drain:
  chat's close frames now reach clients (byte-verified 0x88), then
  main returns and the reap runs
- also: SIGPIPE ignored process-wide (EPIPE trap instead of death);
  two-phase engine teardown (real drops while arenas+routing live,
  settle passes for routed frees) — fixes the registry-map leak and
  the drain UAF ASan found
- gate scripts/chat-accept.sh + just chat: handshake independently
  verified, functional matrix on BOTH backends, 1k-hot-room soak
  (1000/1000 in ~35ms), drain close-frames, SIGTERM exit 0, ASan leg
  clean. OPEN: soak-fds check (18 fds settle slower than the window)
  + full battery after the semantics change — NOT yet run
- committed for manual testing at the user's request

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-23 09:53:21 +02:00
4092074201 feat: monitor + time.after (ids 89/90) — the lifecycle slice completes
- monitor(watched, observer, msg): registration lives on the watched
  actor's home thread (kind-7 envelope cross-shard); actor_die walks
  the list; already-dead fires NOW; the notice msg moves; a full
  observer's notice drops with a stderr line (no fiber to trap)
- time.after(ms, addr, msg): per-shard timer list riding the deadline
  machinery (uring tick min + epoll timeout both include timers;
  fired from the same sweep); ms <= 0 delivers now; NO cancel — the
  generation-counter idiom is pinned by run/timer-generation
- runtime_notify: one runtime-sourced delivery path (notices, timers) —
  reserve-or-drop, cross-shard via kind-0 envelopes
- compiler: monitor typed as a bespoke free fn (notice typed against
  the OBSERVER's mailbox — the three-argument deviation, disclosed);
  time.after as a stdlib row whose msg arg is EXEMPT from the module-
  call fresh-arg drop (it moves — the double-own bug the timer fixture
  caught); owner move slots for both
- corpus: run/monitor-death (trap-death + already-dead notices),
  run/timer-delivery (armed + immediate), run/timer-generation
- teardown drops undelivered notices and unfired timers; battery 13/13

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-23 08:56:48 +02:00
a32550d968 feat: iteration 35 — net seams + the serving slice (fiber-per-connection)
- runtime ids 91-95: net.read_dl/accept_dl/write_dl (per-call deadline,
  nil/false = the EXPECTED timeout; ms<=0 = old behavior bit for bit),
  net.listen_unix (unlink-before-bind, O_NONBLOCK on the listener —
  probe-found: accept4's flag covers accepted sockets only), net.peer
- plane: one-op-per-park stays law — deadlines ride one per-shard
  TIMEOUT tick (sentinel user_data) + post-CQE expiry sweep +
  POLL_REMOVE tombstone; epoll's deadline scan grew the fd-park case;
  fibers POOL instead of freeing mid-run (stale-CQE UAF); plain parks
  zero park_deadline (no stale sleep deadlines)
- probe: all five seams verified on BOTH WO_IO backends (timeout
  timing exact, peer round-trip, unix rebind)
- framework: parse_request grows first_ms/read_ms; serve_conn — the
  keep-alive loop with deadlines where parked idle conns are LEGAL
  (close-when-idle RETIRED); App.handle_conn exposes it; plain serve()
  unchanged for simple apps
- web-app: app-owned accept_dl loop + ConnWorker actor per connection
  (each builds its own App; cross-shard placement rides the DB actor);
  WA_IDLE_MS knob; gate grows to 41 checks — two slow requests served
  in PARALLEL, stalled client evicted at the idle deadline, slow-loris
  torn at the read deadline (400)
- docs: story 35 -> done with banner; SQE/CQE design spec LANDED (was
  the review doc); ledger rows (timeouts/unix/keep-alive/peer), graph
  (NETSEAM cleared, KEEPAL done), builtin-surface rows, runtime
  CODE-LOGIC section, board entry
- battery 13/13 fresh-built

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-23 08:04:32 +02:00
9a9e4927f6 feat: call/reply — send that waits (id 88, envelope kinds 5/6, WO-E226)
- runtime: mailbox slots grow caller metadata (wo_msg), call parks on
  WO_PARK_INBOX (the DB-RPC protocol) and the resume consumes a SCALAR
  reply; FIBER_DONE ships the receive's return value home (same-shard
  unpark or kind-6 envelope); kind-5 carries cross-shard calls
- actor death is real now: a receive trapping uncaught marks the actor
  dead, error-unparks the in-flight caller AND every queued caller,
  drops queued payloads + state, releases cap slots; send-to-dead
  drops silently, call-to-dead traps — a call never hangs. Fixes the
  pre-existing leak/dangle in TRAPF's fiber-death path (cur_msg leaked,
  a->active dangled, the mailbox rotted)
- compiler: reply typing through actor-M erasure — every receive(M)
  program-wide must agree on one return type and it must be a copyable
  scalar (v1); WO-E226 names disagreeing classes / void receives /
  non-scalar replies; call's message moves exactly like send's (owner)
- corpus: run/call-echo (park + ordered replies), run/call-dead-trap
  (mid-call + to-dead, both catchable), compile-fail/call-void-receive,
  compile-fail/call-reply-disagree; cross-shard call proof rides the
  chat gate next
- battery 12/12 fresh-built (ASan+TSan lanes in fibers/db-actor green)

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-23 06:22:24 +02:00
dd7dd42bd1 feat: bounded mailboxes + WO_T_ACTOR (trap 13); try-arm place-copy fix
- cap 1024 (WO_MAILBOX override at boot): sender-side atomic
  reserve/release on every path — same-shard, cross-shard envelope,
  OOM rollbacks; full mailbox traps the SENDER catchably; delivery
  pop releases; overshoot bounded by in-flight sends (disclosed)
- test_mailbox 12/0: exact cap single-threaded, two racing senders win
  exactly cap slots, drain/refill clean
- corpus run/mailbox-full-trap: parked sleeper, send loop catches
  "actor mailbox full" after >= 1024 sends
- pre-existing compiler bug found + fixed: a try ARM yielding a Text
  PLACE (bare e.msg, try box.field) aliased a register the arm's scope
  end freed — ASan use-after-free, SEGV on the next unwind's
  double-walk; emit_try now applies copy_place_text to both arm
  results; pinned by corpus run/catch-msg-place
- db-bench driver: msgrate keeps iteration 22's unbounded-flood
  contract via WO_MAILBOX=MSG_N (the cap is 24's policy, not 22's)
- battery 12/12 fresh-built

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-23 01:05:13 +02:00
4ec01c4c43 Merge branch 'concurrency-arc-stage3' (iteration 8 complete)
- brings arc stage 3: transparent DB actor (T7) + closeout (T8),
  db-actor sample + gate, board/story reorg (stories/00-status.md)
- conflicts resolved: runtime CODE-LOGIC (kept iteration 36 bitwise
  section AND stage-3 DB-actor section), board in-progress table
  (kept iteration 36 + framework rows AND db-bench row, links fixed
  for the moved board path)
- full battery fresh-built 11/11: woc-build wovm-build woc-test
  wovm-test oop-e2e deps-accept web-app log-watcher employee fibers
  db-actor (first run hit a stale pre-merge wovm — gates require
  built binaries and never rebuild; builds now run first)

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-22 22:50:18 +02:00
c2c9b240f1 feat: iteration 36 task 4 — runtime bitwise opcodes, .wob v6
- WOP_BAND..WOP_SHR = 42..46 (wob.h), WOP_MAX 46, WOB_VERSION 6
- trap kind WO_T_SHIFT=12: shift count outside 0..63 traps (DIV0
  precedent, never x86's silent count%64); SHR arithmetic
- vm.c: one shared case-body serves both dispatch flavors; SHL shifts
  the unsigned word (wrapping), SHR casts int64_t (sign extends)
- loader.c + test runner battery + disasm: v6 accepted, new opcodes
  validated three-register, rendered BAND/BOR/BXOR/SHL/SHR
- woc-test 543/0, wovm-test ASan both flavors green, cli_smoke OK

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-22 21:38:22 +02:00
07c1f7b257 feat: arc stage 3 T7 — transparent DB actor (WO_T_DB hole closed)
- worker DB builtins marshal to shard 0: requester-side slot encode
  (VM heaps never read cross-shard), owner executes serialized in
  adopt, reply unparks via new WO_PARK_INBOX park + envelope 3/4
- engine gains thread-agnostic slot entry points (insert_slots,
  update_field_slot, val_encode/clone, wo_db_exec_req); traps and
  messages byte-identical to the local path
- main.c: engine + replay boot BEFORE shards spawn; workers assert
  rt.db/rt.wal NULL; busy shard adopts inbox once per slice
- latent stage-1 bug fixed: shared io_uring params static raced by
  lazy worker init lost park wakes (~1/20 hangs); params per-vm,
  short submit now fails loud
- new sample docs/examples/db-actor + just db-actor gate 8/0 (multi
  x3, uring/epoll forced, single byte-exact, WAL replay pair);
  ASan+TSan 6/6; full battery green

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-21 13:10:26 +02:00
d24c705860 feat: iterations 19 + 17 — Float/Bytes scalars (.wob v5), library kind + internal/
- Float full stack: literals (fraction/exponent; `0..10` still a range), f64
  opcodes 34-41, @table column, WAL bit-exact replay, json fractions in and
  shortest-round-trip out. IEEE-quiet — FDIV never traps where DIV does.
- Bytes: a wo_str with its own class id, so alloc/free/copy are shared but no
  Text builtin accepts one; len/at/slice/eq/concat, base64 both ways, json
  boundary as base64; TEXT_COPY preserves the kind.
- No implicit Int/Float mixing (WO-E201 in the typechecker, not the emitter,
  which picks the opcode from one side and would misread the other).
- One IEEE deviation: float_cmp total order (NaN last, -0.0 == +0.0) for
  indexes and order-by, keys canonicalized to match. `?Float` nil is a
  reserved quiet NaN — the zero word is +0.0, WO_NIL_SCALAR's bits are -2.0.
- Renderer prefers fixed over exponential in 1e-6..1e21: pure shortest makes
  a price of 900.0 read `9e+02`. One renderer for interp/json/float_to_text.
- Fixed en route: lexer double-counted the leading digit; is_scalar_shaped
  took Float/Bytes as Int-shaped; Bytes ownership needed a shared heap-scalar
  predicate or temps never dropped; order-by bit-compared negatives backwards.
- Iteration 17: `kind = "library"` (absent = program; bad value = WO-E109),
  entry-less check mode retiring the `--emit` workaround, Go's `internal/` as
  WO-E108 at the consumer's `use`. Driver-only; VM/.wob/GC untouched.
- Framework reorg: internal/{parse,serve}.wo; http/form.wo split out to keep
  media_type/form_values public (parse.wo had grown public surface).
- Docs: link audit (97 -> 88 broken, conflict markers resolved, 2 duplicate
  stories removed), 00-code-review verified 26/27, iterations re-sequenced.
- Also carries the pre-staged pub(read)/using/#if work from the index.
- Gates: corpus 103/0, test_wal 156/0, web-app 26/0, oop-accept ALL MET.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-20 19:24:15 +02:00
ec9d264355 feat: cross-shard actors — placement, envelopes, home-routed frees, WO-E222 (arc T6)
- spawn placement: round-robin across shards (same-shard when the
  engine is absent/single); the actor's mailbox and delivery belong to
  its HOME thread — a spawn to another shard travels as an ADOPT
  envelope, a send as a SEND envelope (mutex-guarded inbox + eventfd
  wake; the spec's lock-free rings stay a disclosed deviation until
  9e measures the mutex)
- workers: first envelope triggers lazy full-vm init UNDER the inbox
  mutex (TSan caught the memset racing a concurrent push, twice — the
  second was inbox_push reading wake_efd outside the lock; both fixed,
  gate x8 + battery clean); serve loop = adopt -> run to drained ->
  wait on the plane (the wake eventfd is watched by io_uring POLL_ADD
  oneshot / epoll level-triggered on BOTH backends)
- ownership across heaps: every allocation stamps rt->shard_id into
  the header (the field reserved since iteration 2); a drop on the
  wrong shard routes home as a FREE envelope — the owner's arena stays
  single-threaded by construction; at teardown routed frees become
  no-ops (arenas die wholesale) which is what un-danced the freed-mutex
  ASan SEGV the first ordering had
- WO-E222: an actor's state or message type that is (or transitively
  contains) an inferred-traced class refuses at the spawn/send — with
  round-robin every actor is potentially remote; corpus-pinned
  (compile-fail/traced-send, inference-aware: Box contains ?Node)
- determinism narrowed per spec: oop-e2e pins WO_SHARDS=1 (exact
  outputs); the fibers gate grows multi-shard SET assertions + a TSan
  run (wovm-tsan target; setarch -R fallback for kernel 6.5+ ASLR)
- NEXT_RUNNABLE honors engine shutdown for parked workers (deadlock
  hole closed); io_wait's adopt-wake (rc 1) no longer reads as fatal
- battery: oop-e2e 93/0, fibers 10/0 x8 (+WO_IO=epoll), log-watcher
  7/0, employee 8/0, web-app 21/0, deps 8/0, runtime tests 16/16

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-20 12:06:13 +02:00
5bd8813b9b feat(runtime): shard engine — pinned worker threads, all cores default (arc T5)
- wo_engine + wo_engine_start/stop: one pinned pthread per extra core
  (pthread_setaffinity_np); shard 0 is the primary (entry + database);
  workers idle on a wake eventfd until fibers arrive (T6) or shutdown
- default = all cores (the arc's brave landing), WO_SHARDS=1..64
  overrides; N=1 spawns no threads — byte-identical to stage 1
- worker vms are LAZY: identity + wake fd only until their first fiber
  arrives — 20 idle shards must not cost 1.25 GiB of eager arenas
  (they did: the web-app gate flaked on exactly that before the fix;
  3 consecutive green runs after)
- engine stops (join + destroy) before the primary's teardown
- battery at the 20-core default: oop-e2e 92/0, fibers 8/0,
  log-watcher 7/0 (+WO_SHARDS=1 identical), web-app 21/0 x3,
  employee 8/0, deps 8/0, runtime tests 16/16 files green

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-20 11:46:28 +02:00
897c8442f0 feat(runtime): the per-shard I/O plane — io_uring-first fiber parking (arc T4)
- park.c/park.h: one event loop per shard. io_uring PRIMARY (raw
  io_uring_setup/io_uring_enter, uapi structs mirrored, 5.4-floor ops:
  POLL_ADD for fd readiness, TIMEOUT for sleeps, user_data = the fiber);
  epoll+deadline-scan FALLBACK behind the startup probe; WO_IO=
  uring|epoll forces either so CI proves both on one kernel
- park protocol: a blocking builtin fills cur->park_* and returns
  WO_SYS_PARKED; resume either RE-EXECUTES it (fd readiness: accept/
  read/write retry) or continues PAST it (sleep: result preset,
  park_done=1 — re-executing would restart the full duration)
- sysio: listener + accepted fds nonblocking (accept4 SOCK_NONBLOCK);
  accept/read park on EAGAIN; write parks on EAGAIN with its partial
  progress carried across the retry in park_wr_at; sleep parks on a
  deadline — with ONE fiber the plane's wait IS the blocking call,
  program mode is the degenerate case, not a special one
- scheduler: NEXT_RUNNABLE waits on the plane when the queue empties;
  a stop interrupting the wait reaps EVERY fiber (queued and parked)
  and returns the clean-stop status; parked fibers are GC roots and
  fib_reap_all drains them
- proof: full battery green on the uring path (oop-e2e 92/0,
  log-watcher 7/0 incl. the mcp accept/read/write loop, employee 8/0,
  web-app 21/0, deps 8/0), WO_IO=epoll battery green (log-watcher 7/0,
  web-app 21/0), WO_IO=uring forced green, LW_SOAK=8 10/0 (fd + RSS
  flatness holds over parked I/O)

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-20 09:58:04 +02:00
d2d1721e05 feat: spawn / send / actor M — the unified actor surface (arc T3)
- language: `spawn Cls { fields }` expression (ctor semantics — fields
  MOVE; result is the address); `actor M` parametric field type
  (contextual like multi/map — actor stays a legal identifier); `send`
  is a builtin free-fn name, not a keyword (shadowing rule applies)
- typing: M inferred from Cls's receive(msg: M); WO-E221 when receive
  is missing, mis-armed, or M is not a class/record/union; send checks
  addr is `actor M` and the message IS an M (silent when underivable);
  ctor half of spawn delegates to the Ctor arm (completeness, ?T, E207)
- ownership: send's message TRANSFERS (sender's later use = WO-E301,
  corpus-pinned); spawn's fields move via the ctor machinery; an
  address is Copy
- emit: spawn lowers to ctor + LOADK receive's method index + BUILTIN
  68; send is BUILTIN 69 with the message excluded from fresh-arg drops
  (the runtime owns it now)
- runtime: wo_actor (moved-in instance, receive idx, growable FIFO
  mailbox, one delivery fiber at a time); delivery reuses the fiber
  context across messages and re-queues per message (fairness — an
  actor never monopolizes); the runtime drops each message after its
  receive returns; actor state/queued/in-flight messages are GC roots;
  teardown drops everything (main-return reap included); loader knows
  the two arities
- corpus: run/actor-echo (typed spawn/send, one-at-a-time delivery
  interleaved with main by budget — output exact, ASan-clean),
  compile-fail/spawn-no-receive (WO-E221), send-after-move (WO-E301)
- battery green: oop-e2e 92/0, woc-test, wovm-test, log-watcher 7/0,
  employee 8/0, web-app 21/0, deps-accept 8/0

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-20 07:21:24 +02:00
f88aa11cef feat(runtime): fiber run queue + reduction budget (arc T2)
- fiber states (RUNNABLE/PARKED/DONE), intrusive FIFO run queue,
  wo_vm_spawn_fiber (calloc'd context, frame 0 set up like wo_vm_call)
- reduction budget: WO_REDUCTIONS (default 4000), checked at loop
  BACK-EDGES AFTER the jump lands so the saved pc is the loop head —
  a pre-instruction save at budget 1 re-executes the jump into the
  same decrement and livelocks (found by reasoning, pinned by the
  budget-1 test; deviation from the spec's three-site wording,
  recorded in the yield macro's comment)
- FIBER_DONE: main returning ends the program and reaps every
  remaining fiber through vm_unwind (drop maps run); a spawned fiber
  ending frees silently; its return value is discarded by contract
- TRAPF: an uncaught trap in a spawned fiber kills that fiber ALONE
  (stderr report, program lives); in main it stays the program's death
- WO_SYS_STOPPED reaps all fibers wherever it lands (main unlinked
  from the queue and unwound if a spawned fiber caught the stop)
- vm_gc_roots walks the live fiber plus every queued one
- test_fiber (45 checks, ASan): EXACT round-robin interleave at budget
  1 across three fibers pushing tags into one shared multi;
  main-return reaps a spinning fiber holding an owned Big (ASan proves
  the free); a DIV0 fiber dies alone, main answers 0
- full battery green: wovm-test, oop-e2e 89/0, woc-test, log-watcher
  7/0, employee 8/0, web-app 21/0, deps-accept 8/0 (scheduler dormant
  = one branch per back-edge)

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-20 07:03:40 +02:00
58e37cc19e refactor(runtime): extract the fiber context from wo_vm (arc T1)
- wo_fiber = the interpreter state wo_vm held inline: register window,
  frame stack, catch stack, caught-error slot; wo_vm keeps module,
  runtime, the embedded fiber 0 (main) and the cur pointer every
  interpreter access now reads through
- vm_gc_roots split into a per-fiber walker + the all-fibers caller
  (one fiber today; the loop is where stage 1 T2 adds the rest)
- PURE refactor, no functional change to hide behind: full battery
  byte-identical — wovm-test (test, test-iso, cli_smoke) green,
  oop-e2e 89/0, woc-test green, log-watcher 7/0, employee 8/0,
  web-app 21/0, deps-accept 8/0
- plan: docs/superpowers/plans/2026-08-20-shard-fiber-arc.md task 1
  (plan/spec docs live on branch language-surface-strictness)

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-20 06:40:58 +02:00
092b528081 feat: retire RC from the emitter and the format — .wob v4 (7b Phase 3b)
The compiler no longer emits reference-counting ops anywhere, and the format
reserves them. With Phase 3a's collector this completes the runtime half of
iteration 7b: spec success criteria 3 (no RC ops in any image, opcodes
reserved) and 6 (corpus ASan-clean) are met — `just oop-accept` is fully green.

- owner.ml: the rc machinery is deleted outright — rc_site/rc_op types, the
  rcs table, fn_rcs/rc_groups/rc_escaped, record_rc, release_gc, gc_escape,
  resolve_rc, and the clobber rule (its only consumer was elision). The
  `push`-of-a-gc-value RC_INC special case is gone (the bug class cannot recur
  without RC). Drop tables (owned + LGc kinds) are untouched — the gc mask is
  what feeds the collector's root maps.
- emit.ml: emit_rc, the v_rc view, the escape-acquire anchor, and every
  caller deleted; assignment displacing a traced value emits nothing (the VM's
  store barrier owns it); scope-ended LGc handles clear their gc-mask bit so
  root maps stay precise.
- .wob v4: WOB_VERSION 3 -> 4 in wob.h + emit.ml + disasm.ml + the runner's
  loader battery; opcodes 27-28 removed from the enum/jump table/interpreter
  and REJECTED by the loader like any unknown opcode.
- dump.ml: the == RC == owner-dump section is gone; 6 goldens re-blessed
  (owner dumps lose the section, elision.wo's bc dump loses its RC ops).
- runner.ml: rc-table/ELIDED assertions deleted; the elision test now asserts
  the WHOLE image contains no RC op; the table-contract sweep asserts rc ops
  never appear.
- test_unwind.c: the rc-opcodes test becomes two — the loader rejects reserved
  opcode 27, and an abandoned traced instance is freed by rt_destroy
  (ASan-proven).

Verified: woc-test 540/0 + test_diag 14/0; runtime test + test-iso all suites
ASan/UBSan (test_unwind 12/0); cli_smoke; oop-e2e 79/0 (v4 images end to end);
employee 8/0; log-watcher 7/0; ring runs + reclaimed (freed=3) with zero RC
ops in its image; `just oop-accept` ALL CRITERIA MET.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-19 17:07:43 +02:00
841cb41c6b feat(runtime): incremental tri-color mark-sweep replaces RC (7b Phase 3a)
Reference counting and Bacon-Rajan trial deletion are gone from the runtime.
Traced (inferred-gc) objects now die only by the collector; owned values keep
deterministic drops exactly as before.

- wo_hdr: rc retired; borrow and the freed 4 bytes become a union — non-traced
  values keep the borrow word, traced objects use the 8 bytes as the intrusive
  sweep-list link. Header stays exactly 16 bytes. WHITE is now the all-zero
  color (allocations born white by memset); WO_F_BUF retired.
- gc.c rewritten: snapshot-at-beginning tri-color mark-sweep. Roots (frames'
  gc+owned masks) shaded atomically at cycle start; Yuasa deletion barrier
  shades the OLD target of every gcref edge deleted while marking (SETF
  overwrites + every owned-death path, which all funnel through wo_drop_kind's
  GCREF case); allocations mid-cycle born black. Mark AND sweep budgeted
  (WO_GC_BUDGET objects/slice), sweep resumes via a cursor; gray-worklist OOM
  degrades to a blacken-all cycle (frees nothing, never wrong). Owned interiors
  walked eagerly (single-owner trees), pruned by a per-class may-gcref bit
  computed at rt_init (fixpoint over kinds + v2 field_class/field_elem;
  conservative when metadata is absent).
- vm.c: safepoints at NEW (the heap-goal trigger), CALL, and backward JMP;
  root scan follows vm_unwind's governing-pc convention. Unwind's gc-mask
  branch just nulls the register. RC_INC/RC_DEC are accepted as no-ops until
  the emitter stops producing them (next commit) — which also deletes the old
  RC_DEC-on-nil trap that broke `?Node` gcref field stores.
- main.c pump: post-exit, a rootless cycle frees everything unreachable in
  budgeted slices; the trace line moved into wo_gc_slice (one format for pump
  and in-program slices). rt_destroy frees traced remnants (trap paths, tests).
- WO_GC_GOAL joins WO_GC_BUDGET/WO_GC_TRACE as an rt-owned knob (default 256
  KiB; a tiny goal forces mid-program cycles for testing).
- tests: test_cycle.c rewritten (abandoned cycle freed, rooted cycle survives,
  slices bounded, cycle-through-multi, repeated-cycle leak-freedom, and the
  spec's load-bearing DELETION-BARRIER test: an object hidden behind a black
  object mid-mark must survive). test_rc.c re-pinned to owned drops + the
  owned/traced boundary; test_obj.c asserts tracked-white-linked instead of
  rc=1.

Verified: make test + test-iso (all suites, ASan/UBSan; test_cycle 42/0,
test_rc 14/0) + cli_smoke; oop-e2e 79/0 (gc corpus traces unchanged: the new
slice math reproduces steps=1/freed=2 and steps=2/freed=4); employee 8/0;
log-watcher 7/0. THE RING RUNS: docs/examples/gc-cycle prints
`ring a -> b -> c -> a`, is reclaimed post-exit (freed=3 remaining=0), is ASan
clean, and survives an in-program cycle while rooted (WO_GC_GOAL=64: mid-run
slice frees 0, post-exit frees 3).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-19 16:52:48 +02:00
22910e3974 fix: a stopping program stops (executable plan, Task 4)
- blocking stdlib calls that PARK (net.accept, socket read/write,
  time.sleep, a child wait) no longer restart the syscall when the
  stop flag is set on an interruption: a server sitting in accept
  ignored SIGTERM and only `kill -9` ended it
- a stop is NOT a trap -- builtin.h's WO_SYS_STOPPED carries no error
  record and no catch handler sees it (`try` must not swallow
  SIGTERM); the VM unwinds the whole stack through the same drop
  machinery an uncaught trap uses, so nothing leaks on the way out
- wo_vm_call gained a third outcome (1 = stopped); the CLI maps it to
  the status the program's own `return 0` would have given, and a
  regular-file read keeps its plain EINTR retry -- it does not park
- an ASSIGNMENT was not an ownership boundary: `api_key =
  j.mcp.apiKey` moved the field pointer into the local, so the local
  aliased the record and the first unwind freed the same string twice
  (SIGSEGV in class_free). `let` copied a Text place, assignment now
  does too -- the same double free was latent on the normal exit path,
  hidden by the order the compiler happens to emit drops in
- log-watcher-accept is 7 checks: the seventh is the stop itself, with
  the hard kill demoted to a fallback whose use is the failure
- measured under ASan: mcp parked, mcp after traffic, watch and run
  all exit rc 0 with zero leaks; SIGINT behaves as SIGTERM
- gates: oop-accept ALL CRITERIA MET, oop-e2e 71/0, woc-test 565/0,
  wovm-test green, log-watcher 7/0

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-14 23:58:21 +02:00
eb0095428d fix: Text is an owned value copied at every boundary (executable plan, Task 1)
Measured on the workload's supervisor mode, eight seconds, clean SIGTERM exit:
run 1 051 040 B in 24 allocations -> 2 112 B in 19; watch 128 B in 2 -> 64 B in
1. corpus 71/0, woc runtest 565/0, wovm unit gates green, just log-watcher 6/0.

- owner.ml: `oclass_of` called `Text` a builtin scalar, so it was Copy and NO
  Text local was ever dropped — that, not the missing stdlib table, was the
  leak. Text is now Owned, which forces an answer for what it does at an
  ownership boundary, and the answer is uniform: it is COPIED. Into a
  container (push/set/`m[i] = v`, already true), into a field (SETF), out of a
  function (return), into a binding (`let s = other`), and into a loop cursor.
  The source keeps its value; a freshly built Text stays the caller's and is
  dropped at the site
- owner.ml: resolve_callee answers for three shapes it never knew — reserved
  stdlib members, builtins, and a class's `static` members — so their results
  get a type, an owner and a drop
- vm/builtin: WO_B_TEXT_COPY, the one new builtin the rule needs; SETF copies a
  TEXT field in; emit copies a Text read out of a container, bound from a
  place, returned from a place, or loaded into a cursor, and drops a freshly
  built one after a copying store
- sysio.c: fs.read_all/net.read allocated their cap then relabelled the buffer
  with the short length — but wo_str_free sizes a block by its len (no size
  headers, obj.h), so a 1 MiB buffer wearing a 30-byte length went onto a
  32-byte free list and never came back. They copy out at the true size now
- two regressions the corpus caught, fixed in the same pass: a @gc value read
  out of a container is a plain borrow, not an rc-counted alias; and push's @gc
  escape is keyed on "push is not a user-declared fn" rather than "the callee
  did not resolve", which stopped being true once builtins resolved
- docs: Task 1 closed in the executable plan with its before/after numbers, and
  the status board's item 1 records the deeper root cause

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-14 22:45:03 +02:00
993a540d7a fix: nullable scalars need their own nil word; EQS accepts nil
Found by running the compiled log-watcher, not by reading code: the supervisor
rejected every cron line ("malformed schedule: * * * * *") because a `*` field
expands to 0 and `?Int`'s nil was also 0, so `a == nil` was true for a real
value. Both log-watcher subcommands now behave: `watch` alerts on a live file,
`run` reports SCHEDULE /var/log/backup.log: * * * * *. corpus 71/0, woc 565/0,
wovm gates green.

- a nullable SCALAR (?Int/?Bool/?Timestamp/?Id) spells nil as WO_NIL_SCALAR
  (-2^62), not the zero word. Heap-shaped optionals keep 0 — a null pointer is
  unambiguous. The value is -2^62 and NOT INT64_MIN on purpose: the compiler's
  integers are OCaml's 63-bit natives, so INT64_MIN is not expressible there
  (and `min_int * 2` silently wraps to 0 — the first attempt did exactly that)
- the class table marks such fields (WOB_FIELD_NIL_SCALAR in field_class), so
  the runtime writes the right absence where it produces absence itself:
  json.decode leaving a key absent or seeing `null`, and parse_int on
  unparseable input (so parse_int("0") is now distinguishable from a failure).
  json.encode renders a nil scalar as JSON null
- emit.ml: `nil` takes its word from its destination (annotation, field,
  return type); a comparison against `nil` emits the literal with the other
  operand's type, so ?scalar compares against the sentinel and ?heap against 0
- vm.c: EQS accepts a nil operand — two `?Text` values compare with it, and the
  answer is "both absent is equal, one absent is not". Trapping there made
  `a != b` on optionals unusable (it was trapping BOUNDS "null text" in the
  supervisor's rescan). A non-nil operand must still be a real Text
- docs: both normative docs now state the heap-vs-scalar nil split and the EQS
  rule; the stale duplicate vm_unwind comment is gone

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-14 17:28:50 +02:00
2bb39d6b6b feat: try/catch over the trap system (haxe-parity plan 8, Task 5)
VM catch frames + expression-form try/catch in the compiler. Uncaught traps
keep byte-for-byte today's surface. log-watcher parse errors 18 -> 7;
corpus 71/0, woc runtest 565/0, wovm unit gates green (both dispatch flavors).

- wob.h: WOP_TRY (A sBx: push catch frame, handler at pc+sBx) / WOP_ENDTRY;
  WO_B_ERR_FILL builtin (fills the catch record: 0 code, 1 line, 2 method,
  3 msg — the field-order contract with the compiler)
- vm.h/vm.c: catch stack (depth, handler pc, error reg) + the caught error;
  vm_unwind takes a stop depth, so a caught trap kills every frame above the
  catching one exactly as an uncaught trap would, then releases only what the
  try region owned in the catching frame (drop-entry diff against the handler
  pc) and resumes at the handler; RET/RET0 drop the catch frames of the frame
  they leave; TRAPF resumes instead of returning when the trap was caught
- builtin.c: err_fill allocates the method/msg Texts into the record the
  compiler owns, so the pending error never has to outlive the landing
- loader.c: TRY's handler target validated like a jump, error register like
  any register operand; err_fill arity
- lexer/token/ast/parser: `try`/`catch` keywords; `try expr catch (e) expr`
  and `catch (e) { block }`, newline allowed before `catch`; try binds looser
  than every operator, so `try a / b catch (e) 0` catches the division
- types.ml: predeclared `Error` record (merged table only), catch binding,
  arm-type agreement reported only when both arms are confidently typed
- owner.ml: analyze_try — the catch arm is an alternate flow join off the
  entry state, the error record is an owned handler-scope local
- emit.ml: TRY/body/ENDTRY/JMP + handler prologue (NEW Error, err_fill),
  join drops on both arms, `Error` class entry only for programs that catch

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-14 16:35:36 +02:00
bab88a53b9 feat(runtime): wovm VM core (Iteration 2)
- 16 tasks complete: arena, object model, borrow word, containers,
  RC + budgeted cycle collector, wob_build, validating loader,
  interpreter core (dual dispatch), object opcodes, drop-map unwinding,
  builtins + DB_STUB + TRAP, ICALL, wovm CLI + just recipes
- 13 test suites × 2 dispatch flavors (ASan+UBSan) + CLI smoke, all green
- .wob v1 format pinned in src/wob.h + docs/plan/oop-vm/00-wob-format.md
- wo-rt.c reference event-loop preserved for sub-project 2

This is Iteration 2 of the OOP milestone; compiler front (Iteration 3)
is in progress on this branch. They meet at Iteration 4 (emitter+e2e).
2026-08-10 09:35:55 +02:00