Commit graph

10 commits

Author SHA1 Message Date
803ff0b790 feat(rt2): proc.spawn/wait_dl/signal — the streaming child
- a child is fds: Child {id, stdin, stdout, stderr}, driven by the
  existing net verbs (echo leg proves cat round-trip through write_dl/
  read_dl); caller owns the fds, the runtime owns pid + pidfd
- wait_dl parks on the pidfd: code on exit, nil at the deadline with the
  child untouched; one waiter per id, a second refuses by name; stale
  ids refused via a generation counter in the handle
- proc.signal through pidfd_send_signal; actor_die kills the streaming
  children the dying actor owns; dead fibers cannot linger as waiters
- ids 97-107 registered wholesale (wob.h, loader arities, dispatch
  bound); Child + Signal predeclared records in types.ml; unimplemented
  ids trap at the default case until their task lands
- test_proc 168/0 (echo, wait trio, one-waiter refusal, 200-round churn
  fd-flat), suite ASan clean, woc-test green

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
(cherry picked from commit 9be87f159f1bf9cdd509ceed160e7ea518fde46c)
2026-09-15 01:15:30 +02:00
258222c3a1 feat(lang42): proc.run parks — pidfd + epoll bundle + child registry
- deadlock proven first: chatty child (200 KB stdout, stderr held open)
  hung the old sequential drain 5.0 s into the alarm, code -1, stdout
  truncated at 8192; the leg demands completion under 4 s
- rework: nonblocking pipe read ends + pidfd_open behind one epoll fd the
  fiber parks on (the _dl retry mould); both pipes drain on readiness, so
  the deadlock is gone structurally — leg passes in 15 ms
- wo_child slot table in wo_vm (32/shard) carries cross-park state; caps
  refuse by name (kill + WO_T_IO), deadline armed via dl_active/dl_at,
  defaults 30 s / 1 MiB / 64 KiB
- WO_B_PROC_RUN_DL = 96 shares the case (per-call deadline_ms/out_cap/
  err_cap; compiler row lands in a later task)
- fib_reap kills a reaped fiber's child; wo_vm_destroy sweeps the table
- raw syscalls for pidfd_open/pidfd_send_signal: glibc 2.35 build floor
  has no wrappers
- all 19 suites green under ASan+UBSan

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-09-01 22:19:25 +02:00
735fd270db feat: chat sample + gate (T8/T9, IN PROGRESS) + stop-drain semantics
- docs/examples/chat: registry (call consumer) / room / reader+writer
  actor pair per connection over ws_accept + wsframe; presence,
  broadcast, cross-room isolation, mailbox-full = drop-from-room;
  reader tail sends hardened (a full writer no longer orphans the fd)
- RUNTIME SEMANTICS CHANGE (the drain): SIGTERM no longer kills parked
  fibers from outside — the plane WAKES them and each wait RESOLVES
  (deadline'd waits answer their timeout result, sleeps return early,
  plain waits answer WO_SYS_STOPPED and unwind THAT fiber alone; main's
  STOPPED still ends the program). Workers keep adopting their inboxes
  after stop until eng_shutdown. This is what lets a program drain:
  chat's close frames now reach clients (byte-verified 0x88), then
  main returns and the reap runs
- also: SIGPIPE ignored process-wide (EPIPE trap instead of death);
  two-phase engine teardown (real drops while arenas+routing live,
  settle passes for routed frees) — fixes the registry-map leak and
  the drain UAF ASan found
- gate scripts/chat-accept.sh + just chat: handshake independently
  verified, functional matrix on BOTH backends, 1k-hot-room soak
  (1000/1000 in ~35ms), drain close-frames, SIGTERM exit 0, ASan leg
  clean. OPEN: soak-fds check (18 fds settle slower than the window)
  + full battery after the semantics change — NOT yet run
- committed for manual testing at the user's request

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-23 09:53:21 +02:00
9661c08696 feat: TE rejection + listen backlog 1024 + the 1k soak gate
- parse.wo: ANY Transfer-Encoding header is 400-and-close (RFC 9112
  §6.1) — silently treating chunked as body-less was the smuggling
  door the dup-CL fix left open
- net.listen/listen_unix backlog 64 -> 1024: the soak's connect bursts
  overflowed the kernel accept queue and BLACK-HOLED clients (three-way
  handshake done, server never sees the conn — 35-70 stuck per run,
  fully reproduced then gone at 1024; kernel clamps via somaxconn)
- web-app gate grows to 46 checks: TE-reject; the 1k soak — 500 real
  conns all served + 500 idle conns all evicted, server fds home
  (45 -> 45), RSS 24MB, healthy after
- battery green (site restart + fibers-TSan legs flaked under parallel
  battery load, both clean serially — the standing flake pair)

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-23 08:23:55 +02:00
a32550d968 feat: iteration 35 — net seams + the serving slice (fiber-per-connection)
- runtime ids 91-95: net.read_dl/accept_dl/write_dl (per-call deadline,
  nil/false = the EXPECTED timeout; ms<=0 = old behavior bit for bit),
  net.listen_unix (unlink-before-bind, O_NONBLOCK on the listener —
  probe-found: accept4's flag covers accepted sockets only), net.peer
- plane: one-op-per-park stays law — deadlines ride one per-shard
  TIMEOUT tick (sentinel user_data) + post-CQE expiry sweep +
  POLL_REMOVE tombstone; epoll's deadline scan grew the fd-park case;
  fibers POOL instead of freeing mid-run (stale-CQE UAF); plain parks
  zero park_deadline (no stale sleep deadlines)
- probe: all five seams verified on BOTH WO_IO backends (timeout
  timing exact, peer round-trip, unix rebind)
- framework: parse_request grows first_ms/read_ms; serve_conn — the
  keep-alive loop with deadlines where parked idle conns are LEGAL
  (close-when-idle RETIRED); App.handle_conn exposes it; plain serve()
  unchanged for simple apps
- web-app: app-owned accept_dl loop + ConnWorker actor per connection
  (each builds its own App; cross-shard placement rides the DB actor);
  WA_IDLE_MS knob; gate grows to 41 checks — two slow requests served
  in PARALLEL, stalled client evicted at the idle deadline, slow-loris
  torn at the read deadline (400)
- docs: story 35 -> done with banner; SQE/CQE design spec LANDED (was
  the review doc); ledger rows (timeouts/unix/keep-alive/peer), graph
  (NETSEAM cleared, KEEPAL done), builtin-surface rows, runtime
  CODE-LOGIC section, board entry
- battery 13/13 fresh-built

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-23 08:04:32 +02:00
146d0e27f3 feat: time.ticks builtin — CLOCK_MONOTONIC us Int (id 84)
- iteration 22's honest clock: monotone, never wall time; time.now
  stays ms. One types.ml row, sysio case, explicit dispatch arm
  (sys range gate stops at PROC_RUN)
- corpus run/time-ticks (RED WO-E406 first); oop-e2e 104/0, full
  battery green; surface doc row (07-systems-stdlib absent, disclosed)

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-21 16:13:10 +02:00
897c8442f0 feat(runtime): the per-shard I/O plane — io_uring-first fiber parking (arc T4)
- park.c/park.h: one event loop per shard. io_uring PRIMARY (raw
  io_uring_setup/io_uring_enter, uapi structs mirrored, 5.4-floor ops:
  POLL_ADD for fd readiness, TIMEOUT for sleeps, user_data = the fiber);
  epoll+deadline-scan FALLBACK behind the startup probe; WO_IO=
  uring|epoll forces either so CI proves both on one kernel
- park protocol: a blocking builtin fills cur->park_* and returns
  WO_SYS_PARKED; resume either RE-EXECUTES it (fd readiness: accept/
  read/write retry) or continues PAST it (sleep: result preset,
  park_done=1 — re-executing would restart the full duration)
- sysio: listener + accepted fds nonblocking (accept4 SOCK_NONBLOCK);
  accept/read park on EAGAIN; write parks on EAGAIN with its partial
  progress carried across the retry in park_wr_at; sleep parks on a
  deadline — with ONE fiber the plane's wait IS the blocking call,
  program mode is the degenerate case, not a special one
- scheduler: NEXT_RUNNABLE waits on the plane when the queue empties;
  a stop interrupting the wait reaps EVERY fiber (queued and parked)
  and returns the clean-stop status; parked fibers are GC roots and
  fib_reap_all drains them
- proof: full battery green on the uring path (oop-e2e 92/0,
  log-watcher 7/0 incl. the mcp accept/read/write loop, employee 8/0,
  web-app 21/0, deps 8/0), WO_IO=epoll battery green (log-watcher 7/0,
  web-app 21/0), WO_IO=uring forced green, LW_SOAK=8 10/0 (fd + RSS
  flatness holds over parked I/O)

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-20 09:58:04 +02:00
22910e3974 fix: a stopping program stops (executable plan, Task 4)
- blocking stdlib calls that PARK (net.accept, socket read/write,
  time.sleep, a child wait) no longer restart the syscall when the
  stop flag is set on an interruption: a server sitting in accept
  ignored SIGTERM and only `kill -9` ended it
- a stop is NOT a trap -- builtin.h's WO_SYS_STOPPED carries no error
  record and no catch handler sees it (`try` must not swallow
  SIGTERM); the VM unwinds the whole stack through the same drop
  machinery an uncaught trap uses, so nothing leaks on the way out
- wo_vm_call gained a third outcome (1 = stopped); the CLI maps it to
  the status the program's own `return 0` would have given, and a
  regular-file read keeps its plain EINTR retry -- it does not park
- an ASSIGNMENT was not an ownership boundary: `api_key =
  j.mcp.apiKey` moved the field pointer into the local, so the local
  aliased the record and the first unwind freed the same string twice
  (SIGSEGV in class_free). `let` copied a Text place, assignment now
  does too -- the same double free was latent on the normal exit path,
  hidden by the order the compiler happens to emit drops in
- log-watcher-accept is 7 checks: the seventh is the stop itself, with
  the hard kill demoted to a fallback whose use is the failure
- measured under ASan: mcp parked, mcp after traffic, watch and run
  all exit rc 0 with zero leaks; SIGINT behaves as SIGTERM
- gates: oop-accept ALL CRITERIA MET, oop-e2e 71/0, woc-test 565/0,
  wovm-test green, log-watcher 7/0

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-14 23:58:21 +02:00
eb0095428d fix: Text is an owned value copied at every boundary (executable plan, Task 1)
Measured on the workload's supervisor mode, eight seconds, clean SIGTERM exit:
run 1 051 040 B in 24 allocations -> 2 112 B in 19; watch 128 B in 2 -> 64 B in
1. corpus 71/0, woc runtest 565/0, wovm unit gates green, just log-watcher 6/0.

- owner.ml: `oclass_of` called `Text` a builtin scalar, so it was Copy and NO
  Text local was ever dropped — that, not the missing stdlib table, was the
  leak. Text is now Owned, which forces an answer for what it does at an
  ownership boundary, and the answer is uniform: it is COPIED. Into a
  container (push/set/`m[i] = v`, already true), into a field (SETF), out of a
  function (return), into a binding (`let s = other`), and into a loop cursor.
  The source keeps its value; a freshly built Text stays the caller's and is
  dropped at the site
- owner.ml: resolve_callee answers for three shapes it never knew — reserved
  stdlib members, builtins, and a class's `static` members — so their results
  get a type, an owner and a drop
- vm/builtin: WO_B_TEXT_COPY, the one new builtin the rule needs; SETF copies a
  TEXT field in; emit copies a Text read out of a container, bound from a
  place, returned from a place, or loaded into a cursor, and drops a freshly
  built one after a copying store
- sysio.c: fs.read_all/net.read allocated their cap then relabelled the buffer
  with the short length — but wo_str_free sizes a block by its len (no size
  headers, obj.h), so a 1 MiB buffer wearing a 30-byte length went onto a
  32-byte free list and never came back. They copy out at the true size now
- two regressions the corpus caught, fixed in the same pass: a @gc value read
  out of a container is a plain borrow, not an rc-counted alias; and push's @gc
  escape is keyed on "push is not a user-declared fn" rather than "the callee
  did not resolve", which stopped being true once builtins resolved
- docs: Task 1 closed in the executable plan with its before/after numbers, and
  the status board's item 1 records the deeper root cause

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-14 22:45:03 +02:00
fb91085bdb feat: systems stdlib OS half — fs, time, env, net, proc
log-watcher diagnostics 129 -> 55 (json is what is left: 13 encode sites,
3 `as` parses and their cascade). corpus 71/0, woc 565/0, wovm gates green.

- runtime/src/sysio.c (new): 17 builtins behind the reserved module names —
  fs.exists/list/stat/read_all/read_at/append, time.sleep/local/iso,
  env.get/stopping, net.listen/accept/read/write/close, proc.run. Thin
  blocking libc calls; a failed syscall traps the new WO_T_IO with errno's
  own message, which `try ... catch` is how a program handles
  - record-returning members (fs.stat, time.local, proc.run) take their
    result record's CLASS ID as the last argument, so the VM allocates what
    it fills without knowing any source type name (the err_fill pattern)
  - absence is the zero word: a missing path from fs.stat and an unset
    env.get are nil, not traps
  - env.stopping installs SIGTERM/SIGINT handlers on first use only
- types.ml: predeclared Stat/TimeParts/Proc records (field order is the
  contract with sysio.c) + the stdlib member table (arity, builtin id,
  return type, result record) + stdlib return types in confident_typ
- emit.ml: stdlib member calls lower to their builtin with the record class
  id appended; WO-E406 now means "no such member", not "not linked";
  predeclared records enter the class table only when a program needs them
- emit.ml: fstate carries the method's declared return type, so a tail
  `return []` / `return {}` gets its element kinds; a non-empty list literal
  falls back to its own element type when there is no declared destination
- types.ml: confident_typ chases a container read (`c[i]`), which is what
  makes a switch over a value pulled out of a map resolve; a void `try` arm
  no longer demands its catch arm agree
- corpus: lang-use-stdlib-not-linked now pins WO-E406 for an unknown MEMBER
  (fs.slurp) — the "not linked" premise is gone now that fs is linked

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-14 16:54:31 +02:00