- four forks settled with KISS defaults grounded in runtime/src: counters+gauges only (profiling split out), Prometheus text rendered in .wo from a map<Text, Int>, pull via proc.metrics(), stack trace on trap lands first - phases A (trace on trap at both trap sites) / B (proc.metrics from existing gc/arena/fiber fields) / C (porch mounts /metrics — consumer's phase) - builtin id to be confirmed against WO_B_MAX at build time (random_bytes claims 119 per porch 2's brief) - review_pending marker: forks auto-approved 2026-09-09, developer second review before code lands - board row: refine -> ready Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> (cherry picked from commit feb11c3aad613ed7f41b280b27c1f6c0dda92ec7)
7.8 KiB
| track | iteration | was_language_iteration | status | readiness | review_pending |
|---|---|---|---|---|---|
| runtime-v2 | 7 | 30 | pending | ready | forks auto-approved 2026-09-09 for autonomous execution — developer second review before code lands; four forks settled with KISS defaults grounded in runtime/src (existing gauges, existing line table) |
runtime-v2 7 — observability: trace on trap, then counters and gauges
Moved 2026-09-06 from the language track (was language iteration 30, the number a dozen docs still point at) into runtime-v2, whose builtin-sized-seam shape it fits. Brainstormed to
ready2026-09-09: the four forks below are settled, grounded in what the runtime already holds rather than in what an observability stack usually ships. Profiling is split out (fork 1) — this iteration is the cheap, high-value half.
Why this exists
The runtime has no observability surface. A healthcheck answers one bit (up/not-up); nothing exposes counters, gauges, latencies, memory, or a profile. Every attempt to reason about the system's behaviour at runtime hits the same wall, which is why the number is referenced from five directions at once:
- porch excludes
expvar,pprofand metrics endpoints by pointing here (story 8, 39) — a healthcheck is one bit, not observability. - databasev2 cannot observe table size without it
(bounded-tables 5) and leans on
it for per-change benchmark CI
(databasev2 story); the RAM-ceiling study read
RSS from
/procby hand (databasev2 1) precisely because this does not exist. - The rate limiter's ephemeral-row expiry is lazy "because porch has no timer and iteration 30 owns" the sweep story (that "iteration 30" is now this one — porch 1).
- Debugging a trap today is one stderr line —
trap CODE in METHOD at line N: MESSAGE— with no call stack, so a trap three calls deep names only the innermost frame.
Fiber ships expvar and pprof as middleware; the equivalents here are runtime
work, because the numbers they expose (allocations, fiber counts, shard load,
GC pauses) live in the C runtime, not in .wo.
Decisions locked (2026-09-09)
- Counters and gauges only; profiling is its own later iteration. The
gauges this iteration exposes already exist as runtime fields —
gc_traced_cnt,gc_alloc_bytes,gc_step_no(obj.h), the arena'sused,nfibers,nchildren,ntls(vm.h) — so exposing them is a read, not new machinery. CPU/heap profiling needs sampling infrastructure in the VM and is a different size of commitment; it gets its own runtime-v2 iteration when a consumer measures the need. (The story always allowed this split.) - Exposition is Prometheus text, rendered in
.wo. The runtime returns numbers; the format is the consumer's. The builtin hands back amap<Text, Int>(an existing container — no new record class), and porch renders the ops-standard Prometheus text (name valuelines) from it in a few.wolines. Noexpvar-style JSON in this iteration (ajson.encodeof the same map is a one-liner if a consumer asks); no format negotiation. - Pull, not push. A
proc.metrics() -> map<Text, Int>builtin —procbecause it is this process/shard's introspection, besideproc.run/spawn— and porch mounts/metricson it. Push to a collector is possible now thatnet.connect/net.connect_tlsexist, but it needs a collector protocol and a consumer; deferred by name. The snapshot is per shard (the calling shard's gauges); a cross-shard aggregate would need an inbox round-trip and is deferred — scrape each shard or sum in the app. - Stack trace on trap lands first, alone, in this iteration. Highest
debugging value per line and zero dependency on the metrics half. The trap
path already resolves method + line through the per-method line table
(loader.c validates it ascending);
vm_unwindalready walks the frame stack. Phase A walks the frames before unwinding and prints oneat METHOD line Nline per frame under the existing trap line — both the main-fiber trap (main.c) and the fiber trap (vm.c) sites. Stderr only, always on (a trap is already a stderr event); no new builtin, no format flag.
Builtin id: the next free after the ones being claimed ahead of it
(random_bytes takes 119 per porch 2's brief) — confirm against WO_B_MAX
in runtime/src/wob.h at build time; the id space is one shared enum
(wob.h + emit.ml/types.ml + loader.c arity), the lesson porch 2's brief
recorded.
Phases
- A — stack trace on trap. At both trap-report sites, walk the trapping
fiber's frames innermost-first and print
at <method> line <n>per frame (line from the method's line table at the frame's pc;?when a method has no table). Verify: a corpus/regress fixture that traps three calls deep shows threeatlines in order under the trap line; a top-level trap shows one; every existing fixture's first stderr line is byte-unchanged (the trace is appended, never prepended). - B —
proc.metrics(). The builtin fills amap<Text, Int>from the shard's existing fields — at leastarena_used_bytes,gc_traced_objects,gc_alloc_bytes_since_cycle,gc_slices,live_fibers,live_children,live_tls_conns, plusshard_id— registered in wob.h + types.ml + loader.c (module member,proc). Verify: a runtime test asserts every key is present and thatlive_fibersrises with spawned fibers and falls when they finish; the map's keys are stable names (they are the metric names porch will emit). - C — porch mounts
/metricsrendering Prometheus text from the map. This is the consumer's phase — pure.wo, owned by porch 8's lifecycle slice — and is named here so the builtin ships with its consumer visible, not built here.
Acceptance Criteria
- Given a program that traps three calls deep, when it traps, then
stderr carries the existing trap line unchanged followed by three
at METHOD line Nlines, innermost first. - Given every existing fixture and gate, when run, then the first stderr line of each trap is byte-identical to before (the trace only appends).
- Given a
.woprogram that spawns N fibers and callsproc.metrics(), when inspected, thenlive_fibersreflects them and every named key is present with anInt. - Given the map, when porch renders it, then the output is valid
Prometheus text (one
name valueline per key) — proven in porch's own gate when phase C lands.
Out of scope (named, owned elsewhere)
- CPU/heap profiling (
pprof's equivalent) — its own runtime-v2 iteration when a consumer measures the need (fork 1). - Push / OpenTelemetry export — needs a collector protocol and a consumer;
net.connect_tlsnow exists, so the transport is no longer the blocker. - Cross-shard aggregation of the gauges — an inbox round-trip; scrape or sum in the app for now.
expvar-style JSON, format flags, per-request latency histograms — later, on demand.- Per-change CI and fuzzing. Tooling and process, not a runtime surface; the benchmark harness already exists (language iteration 22).
- Alerting, dashboards. Downstream of exposition, not the runtime's job.
Info
Consumers exist and are named above, so this is not a primitive shipped as decoration. Ordering: phase A (trace on trap) has no dependency and leads; phase B (counters) follows; phase C is porch's. Nothing here depends on the porch track — the dependency runs the other way. Small: a frame walk at two trap sites, one module builtin reading fields that already exist.