- jarvis (00-story): 6th track, 2nd software built with writeonce — an AI assistant; direct-HTTPS design; blockers named (net.connect + TLS) - runtime-v2 7 observability + 8 symmetric cipher: moved from the language track (were 30/43); 9 in-process TLS: created from the gap jarvis surfaces, RETIRES the "TLS is the proxy's job" doctrine (both directions) - language 41 (arena hang): fix design to ready — marshal cross-shard messages (root), align the shard_id % nshards route/compare + assert bound; poison-on-free + minimal fixture as follow-ups - fiber scope-gap analysis (plan/exploration/fiber/01): porch vs fiber, what porch lacks, would developers prefer porch - board + dependency-graph synced (porch 2-8 ready; rv2 table; §5/§5a graphs) (cherry picked from commit 203470ceb2a151fe3584931cd4237af3f96a9f29)
4.3 KiB
| track | iteration | was_language_iteration | status | readiness |
|---|---|---|---|---|
| runtime-v2 | 7 | 30 | pending | refine |
runtime-v2 7 — observability: metrics, profiling, and traces on trap
Moved 2026-09-06 from the language track (was language iteration 30, the number a dozen docs still point at) into runtime-v2, whose builtin-sized-seam shape it fits. It stretches the track's original processes/terminals/signals charter — observability is runtime instrumentation of the VM, GC and shards — but the track already grew past its first five seams.
readiness: refine— the gap, its consumers and its forks are named here, nothing is brainstormed toreadyyet.
Why this exists
The runtime has no observability surface. A healthcheck answers one bit (up/not-up); nothing exposes counters, gauges, latencies, memory, or a profile. Every attempt to reason about the system's behaviour at runtime hits the same wall, which is why the number is referenced from five directions at once:
- porch excludes
expvar,pprofand metrics endpoints by pointing here (story 8, 39) — a healthcheck is one bit, not observability. - databasev2 cannot observe table size without it
(bounded-tables 5) and leans on
it for per-change benchmark CI
(databasev2 story); the RAM-ceiling study read
RSS from
/procby hand (databasev2 1) precisely because this does not exist. - The rate limiter's ephemeral-row expiry is lazy "because porch has no timer and iteration 30 owns" the sweep story (that "iteration 30" is now this one — porch 1).
Fiber ships expvar and pprof as middleware; the equivalents here are runtime
work, because the numbers they expose (allocations, fiber counts, shard load,
GC pauses) live in the C runtime, not in .wo.
What it should deliver (scope to be refined)
- Runtime counters and gauges — allocations, arena high-water, live fiber and actor counts, per-shard load, GC pause totals, request counters — exposed through one endpoint the app can mount.
- A profiling story — CPU and heap sampling, the
pprofequivalent, so a hot path can be found rather than guessed at. - A stack trace on trap — today a trap is a 500 and a line; a trace at the trap site is the cheapest debugging win and may be separable from the metrics work.
Forks the brainstorm must settle
- Counters only, or profiling too? Counters and gauges are a bounded, mostly
.wo-plus-a-few-builtins surface; CPU/heap profiling needs sampling machinery in the runtime and is a much larger commitment. Splitting profiling into its own iteration is a legitimate outcome. - Exposition format. Prometheus text (the ops-standard, scrape-friendly),
an
expvar-style JSON blob, or both. The format decides who can consume it without a translator. - Pull endpoint or push. A mounted
/metricsendpoint (pull) fits the proxy-fronted, single-binary model; a push to a collector needsnet.connect, which does not exist (iteration 38) — so pull is almost certainly the answer, but say so. - Is stack-trace-on-trap in this iteration at all? It is separable, it is the highest debugging value per line, and it touches the trap path rather than the metrics path — a candidate to land first and alone.
Out of scope (named, owned elsewhere)
- Per-change CI and fuzzing. Frequently lumped under the old "iteration 30" but they are tooling and process, not a runtime surface; they belong to a CI/ops story, not this one. The benchmark harness already exists (language iteration 22).
- Distributed tracing / OpenTelemetry export. Needs
net.connect(iteration 38) and a wire protocol; a later slice if a consumer appears. - Alerting, dashboards. Downstream of exposition, not the runtime's job.
Info
Consumers exist and are named above, so this is not a primitive shipped as decoration. Ordering: stack-trace-on-trap has no dependency and could lead; counters/gauges are next; profiling is the heaviest and most separable. Nothing here depends on the porch track — the dependency runs the other way.