Commit graph

103 commits

Author SHA1 Message Date
bd99da599d feat(tls): sans-io server handshake FSM (rv2 9 phase G2)
- wo_tls_server: the mirror of the client driver. parse ClientHello (pick
  suite, extract x25519 share, echo session id; reject no-x25519/no-1.3),
  build ServerHello, derive the role-symmetric keys, emit the encrypted
  flight (EncryptedExtensions + Certificate + a signed CertificateVerify +
  Finished), verify the client Finished, switch to application keys
- server_sign_cv signs the CertificateVerify with the phase-G1 primitives
  (RSA-PSS or ECDSA-P256 + a minimal DER SEQ{r,s} encoder); parse_client_hello
  + build helpers reuse the file's wire reader/writer
- wo_tls_server_start builds the Certificate message from a cert chain +
  private key (RSA n/d or EC scalar) + ephemeral; encrypt/decrypt over the
  application keys
- KAT: loopback — our client driver against our server driver, EC then RSA
  server identity, reaching ESTABLISHED with an app round-trip both ways.
  test_tls 123, ASan/UBSan clean

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
(cherry picked from commit 34d2b8f87cebe536cd2b1b33e6948251ec11684f)
2026-09-15 01:15:52 +02:00
df5054b07d feat(crypto): ECDSA-P256 signing, RFC 6979 nonce (rv2 9 phase G1b)
- wo_ecdsa_p256_sha256_sign: deterministic nonce (RFC 6979 HMAC-DRBG over
  the key + message — no RNG, no nonce-reuse/bias risk), then r = (k*G).x
  mod n and s = k^-1 (z + r*d) mod n
- constant-time in the secret: jmul_ct (double-and-add-always + point
  cmov) for k*G, and bn_modexp_ct for k^-1 mod n and the affine inversion.
  Known residual (documented): the ladder leaks k's leading-zero count (a
  bit-length hint, not the key) — a complete-formula/Montgomery-ladder
  upgrade is the named follow-up
- KAT: byte-for-byte vs the RFC 6979 A.2.5 P-256/SHA-256 vectors ("sample"
  + "test"), our sign verifies with our verify, determinism checked.
  test_crypto 115, ASan/UBSan clean

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
(cherry picked from commit 1bc6d04f9e18180dc7d0a6e5e7021dbd588e499a)
2026-09-15 01:15:52 +02:00
631506d36f feat(crypto): constant-time RSA-PSS signing (rv2 9 phase G1a)
- bn_modexp_ct: constant-time modexp for the SECRET exponent — squares and
  multiplies every bit, selects the product with a mask (bn_cmov), so the
  op sequence is independent of d (the existing bn_modexp branches on the
  bit, fine only for the public e)
- wo_rsa_pss_sha256_sign: EMSA-PSS-ENCODE (RFC 8017 §9.1.1) + modexp with d;
  caller supplies the salt (fresh in production; fixed makes the KAT
  deterministic). Private key (n,d)
- KAT: deterministic sign vs a python from-spec oracle byte-for-byte
  (fixed salt), our sign round-trips through our verify, tamper rejected.
  test_crypto 108, ASan/UBSan clean

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
(cherry picked from commit cf8fdfcc47b2b076b63b9c63bf56411615b0d6f9)
2026-09-15 01:15:52 +02:00
fe352b63c2 docs(runtime): CODE-LOGIC — the hand-rolled TLS 1.3 client (rv2 9)
- new "Hand-rolled TLS 1.3 client" section: the crypto ladder in crypto.c,
  the tls.c layers (record / key schedule / messages / sans-io driver /
  chain validation / PEM), and the net.*_tls builtins in sysio.c —
  per-shard no-lock slot table, deadline-bounded blocking handshake then a
  parked data plane, getrandom ephemeral (the runtime's first RNG),
  WO_CA_BUNDLE trust store, loud WO_T_IO failures, the just tls gate
- files table: crypto.c entry updated, tls.c/.h added

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
(cherry picked from commit 5670304d8a7a54e291b81e99e27aa16e29aa1d3e)
2026-09-15 01:15:52 +02:00
6d6c810695 feat(tls): net.connect_tls/read_tls/write_tls builtins (rv2 9 F3c-net)
The outbound TLS 1.3 client wired into the VM (ids 115-117, WO_B_MAX->117):

- net.connect_tls(host,port)->Int: DNS + non-blocking connect+poll bounded
  by WO_TLS_HANDSHAKE_MS (decision 5), then a blocking, SO_*TIMEO-bounded
  hand-rolled handshake over the sans-io driver, then wo_tls_verify_chain
  (chain + host + validity + basicConstraints/EKU) against the shard's
  lazily-loaded read-only CA bundle (decision 4). Any failure traps WO_T_IO
  loudly (decision 3). Returns the fd.
- net.read_tls / net.write_tls: application data over the parked data plane
  (decision 1) — O_NONBLOCK + park on POLLIN/POLLOUT like net.read/write,
  with record reassembly + leftover-plaintext + in-flight-record buffers in
  the per-fd slot so a park/retry never re-seals or loses progress.
- per-shard wo_tls_conn slot table keyed by fd, no locks (one thread per
  shard, the wo_child pattern; decision 2); net.close frees the slot;
  wo_vm_destroy reaps all slots + the CA bundle. getrandom ephemeral.
- driver keeps the whole Certificate message + wo_tls_client_chain() so the
  trust walk sees the full chain, not just the leaf.
- wiring: wob.h, loader.c arities, builtin.c dispatch (second net range),
  types.ml (net.connect_tls/read_tls/write_tls), sysio.c impl.

Builds; full runtime suite 0 fail; woc builds. Live behaviour is the
Phase-4 gate (next commit).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
(cherry picked from commit 9a922b3245eaa9134efb60bfd46521952ed12a39)
2026-09-15 01:15:31 +02:00
040ec9c189 feat(tls): PEM trust-anchor decoder (rv2 9 F3c-net decision 4)
- wo_tls_pem_to_ders: scan a PEM bundle for CERTIFICATE blocks, base64-decode
  each into a caller arena, record DER spans as trust anchors for
  wo_tls_verify_chain. Pure (caller reads the file + owns the arena) so it is
  offline-testable; the file read + per-shard cache land with the builtin
- b64_decode helper (standard alphabet, skips whitespace/newlines)
- KAT: decode the real /etc/ssl/certs/ca-certificates.crt (>100 anchors,
  each parses, first is a CA), garbage PEM -> 0 with no over-read,
  skip-if-absent for CI. test_tls 107 pass, ASan/UBSan clean

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
(cherry picked from commit 6445d55aa83dbed831d84fd3cca3a74fe4609a00)
2026-09-15 01:15:31 +02:00
3a1ba15851 feat(tls): X.509 basicConstraints + EKU chain hardening (rv2 9 F3c-net decision 6)
- crypto.c: x509_find_ext (generic extension walker) + wo_x509_basic_constraints
  (cA / pathLenConstraint, absent => not a CA) + wo_x509_eku_serverauth_ok
  (EKU absent, serverAuth, or anyEKU => usable; else not)
- wo_tls_verify_chain enforces decision 6: the leaf must be server-usable
  (EKU), every server-sent issuer and the signing anchor must be a CA
  (basicConstraints CA:TRUE) with a pathLenConstraint covering the
  intermediates below it — stops a leaf masquerading as a CA
- gen_x509.py extended (folds in the wildcard leaf, adds EKU clientAuth-only,
  EKU serverAuth, a non-CA intermediate + a leaf issued under it); vectors
  regenerated
- KATs: extractors (test_crypto 104) + chain enforcement (test_tls 103) —
  EKU serverAuth accepted, clientAuth-only rejected, leaf-under-non-CA
  rejected though every signature verifies; existing chains still pass.
  ASan/UBSan clean

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
(cherry picked from commit 3811418014c2c8d9bc3a9464256f52104261b1b9)
2026-09-15 01:15:31 +02:00
5f0ac2d434 feat(tls): certificate chain validation (rv2 9 phase F3c-net security core)
- wo_tls_verify_chain: leaf-first DER chain — each cert signed by the
  next, the top trusted (equal to, or signed by, a trust anchor), the leaf
  SAN matching host, every cert temporally valid. Any failure rejects;
  no partial trust. Pure over the phase-D/E verifiers, so offline-testable
- KAT with the phase-E RSA + EC chains: leaf trusted via its issuing CA
  anchor; wrong-anchor / wrong-host / expired / broken-link / no-anchor
  all rejected; two-cert chain with a byte-equal root anchor; NULL host
  skips the SAN check. test_tls 100 pass, ASan/UBSan clean
- remaining F3c-net (live-gated): CA-bundle PEM loader, random ephemeral,
  the net.connect_tls builtin driving the sans-io driver over a real fd

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
(cherry picked from commit 9d40055108d301be32df0e1d4140c23446baa7c6)
2026-09-15 01:15:31 +02:00
d7a6888931 feat(tls): SAN/hostname verification + driver enforcement (rv2 9 phase E/F3c)
- wo_x509_check_host: match a hostname against the cert subjectAltName
  dNSNames (RFC 6125) — case-insensitive, single left-most wildcard that
  covers exactly one label; no SAN => refused; no legacy CN fallback.
  Completes the phase-E deferred hostname check (walks the [3] extensions)
- wo_tls_client_set_host + driver enforcement: with a host set, a leaf
  whose SAN does not match is refused at the Certificate step (MITM
  defense); unset skips the check (offline testing only, documented unsafe)
- KAT: exact/case-insensitive/mismatch, no-SAN refused, wildcard one-label
  (not zero, not sub-label) via a wildcard-SAN cert; driver refuses the
  RFC 8448 leaf (no SAN) once a host is set. test_crypto 95, test_tls 91,
  ASan/UBSan clean

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
(cherry picked from commit 319ce8bfcf608030f958b36ab2c8f62fd76e1740)
2026-09-15 01:15:31 +02:00
d7e7eff6b5 feat(tls): sans-io TLS 1.3 client handshake driver (rv2 9 phase F3c-core)
- wo_tls_client: a pure state machine (no sockets). Caller frames
  records; driver runs ClientHello->ServerHello->flight->Finished and
  hands back bytes to send. Keeps all I/O out of the security-critical FSM
- start_with (inject CH + ephemeral priv), push_record, take_output,
  encrypt/decrypt (application traffic keys). Handshake-message reassembly
  across records; per-message transcript timing (CertVerify signs CH..Cert,
  Finished MACs CH..CertVerify); constant-time Finished compare; every
  failure lands in FAILED (no warn-and-continue)
- verifies server CertificateVerify (phase E+D) + server Finished, emits
  the client Finished, switches to application keys
- KAT: whole handshake driven offline against the RFC 8448 record trace —
  client Finished record byte-for-byte, first client app record
  byte-for-byte, NewSessionTicket + server app data decrypt to plaintext,
  tampered flight -> FAILED. test_tls 90 pass, ASan/UBSan clean
- SECURITY TODO before live use (documented in tls.h + story): chain walk
  to a trust anchor + SAN/hostname match; random ephemeral for production
  start; the net.connect_tls socket glue

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
(cherry picked from commit 74c332d7efdb8bbfdbe90bd2fcc3defa2fe8da00)
2026-09-15 01:15:31 +02:00
74b153aa0a feat(tls): offline handshake verification (rv2 9 phase F3b)
- wo_tls_verify_cert_verify: verifies a server CertificateVerify
  (RFC 8446 §4.4.3) — builds the 64-space || context || 0x00 ||
  transcript-hash content, parses the leaf SPKI (phase E) and dispatches
  to phase-D RSA-PSS / RSA-PKCS1 / ECDSA-P256; the scheme must match the
  leaf key type. ECDSA sig r/s pulled from its DER SEQ
- reuses wo_tls_finished_verify (phase F2) for server + client Finished
- KAT: the whole handshake crypto driven offline from the RFC 8448 §3
  recorded messages — CertificateVerify (RSA-PSS) VALID, wrong-transcript
  / tampered-sig / mismatched-scheme rejected, server Finished byte-exact,
  and the client Finished we would send byte-exact. test_tls 78 pass,
  ASan/UBSan clean

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
(cherry picked from commit afd9f23508c648322712df91329aa117c975ccab)
2026-09-15 01:15:31 +02:00
c375110aac feat(tls): TLS 1.3 handshake message layer (rv2 9 phase F3a)
- bounded wire reader/writer (malformation -> reject, overflow -> fail;
  no over-read on attacker-controlled bytes)
- wo_tls_parse_server_hello: extracts negotiated suite + server x25519
  key share; rejects HelloRetryRequest, unsupported suite/group,
  non-1.3 selected_version, and any truncation
- wo_tls_build_client_hello: ClientHello offering TLS 1.3 / x25519 /
  RSA-PSS+RSA-PKCS1+ECDSA-P256, SNI, 32-byte legacy session id
- KAT: ServerHello parser vs RFC 8448 recorded message (suite 0x1301 +
  server pubkey byte-exact), malformed rejected; ClientHello builder
  structural + SNI/keyshare present + too-small refused, and validated
  byte-for-byte spec-valid by an independent python parser. test_tls 71
  pass, ASan/UBSan clean

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
(cherry picked from commit 541c71bca1655f370b145621d272d2e8bdb6c7ce)
2026-09-15 01:15:31 +02:00
25dda4cb2f feat(tls): TLS 1.3 key schedule (rv2 9 phase F2)
- wo_tls_derive_handshake: Early/Handshake/Master secrets + client/server
  handshake-traffic secrets from the ECDHE shared secret and the
  ClientHello..ServerHello transcript hash (RFC 8446 §7.1)
- wo_tls_derive_application: client/server application-traffic secrets
  from master_secret + the ClientHello..server-Finished transcript hash
- wo_tls_traffic_keys: record key + IV via HKDF-Expand-Label "key"/"iv"
- wo_tls_finished_verify: finished_key = Expand-Label(base,"finished"),
  verify_data = HMAC(finished_key, transcript_hash)
- all over phase-B HKDF (Extract/Expand-Label) + Derive-Secret helper
- KAT vs RFC 8448 §3 "Simple 1-RTT Handshake" byte-for-byte: c/s hs
  traffic, master, c/s ap traffic, server hs key+iv. Also validates the
  phase-B "tls13 " Expand-Label. test_tls 58 pass, ASan/UBSan clean

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
(cherry picked from commit 417fcc16f80c1dd31e84a0336944573902fde06e)
2026-09-15 01:15:31 +02:00
7e71c1a171 feat(tls): TLS 1.3 record layer (rv2 9 phase F1)
- new tls.c/tls.h on the crypto ladder: wo_tls_record_seal/open
  (RFC 8446 §5.2) — TLSInnerPlaintext (content||type, no padding),
  5-byte header as AEAD additional-data, per-record nonce = iv XOR
  seq big-endian (§5.3)
- suite dispatch: TLS_AES_128_GCM_SHA256 (mandatory) +
  TLS_CHACHA20_POLY1305_SHA256 (AES-NI-less fallback), over phase-A AEAD
- open() strips trailing zero padding to recover the inner content type;
  rejects a length-field lie before the AEAD, and auth failure after
- KAT vs python AEAD oracle (test/gen_tls_record.py): sealed record
  byte-for-byte both suites, open() recovers it, 5-seq round-trip,
  tamper + wrong-seq + bad-suite rejected. test_tls 51 pass, ASan/UBSan

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
(cherry picked from commit 5021a99f8359f78a642bb0cc2b28ab66f6b624e2)
2026-09-15 01:15:31 +02:00
5f5d774d76 feat(crypto): X.509 chain-link verification (rv2 9 phase E core)
- defensive ASN.1/DER reader: every length/bound checked; malformation
  is rejection, never over-read (truncated input KAT-gated)
- x509_parse: tbsCertificate span, sig-alg OID, signature,
  SubjectPublicKeyInfo (RSA n/e or EC P-256 x/y), validity dates
- wo_x509_verify_one: one chain link's signature, dispatching to
  phase-D RSA-PKCS1/PSS + ECDSA-P256 by the issuer key type
- wo_x509_parse_spki + wo_x509_check_validity (caller supplies time)
- KAT against real python-generated chains (test/gen_x509.py):
  RSA CA+leaf (SHA256withRSA), EC P-256 CA+leaf (ecdsa-with-SHA256);
  leaf-vs-CA, self-signed CA, wrong-issuer/tampered/truncated reject,
  validity window, SPKI extraction. test_crypto 84 pass, ASan/UBSan clean
- deferred to phase F: SAN/hostname match + multi-cert chain walk to a
  system CA bundle (both need the target host / trust store, known at
  handshake time)

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
(cherry picked from commit 4ec1c75f889df3612e31b9a1a77c4c927fb546c1)
2026-09-15 01:15:31 +02:00
c085c7390c feat(crypto): ECDSA-P256 verification (rv2 9 phase D part 2)
- wo_ecdsa_p256_sha256_verify: NIST P-256 signature verify for EC-cert chains
  and TLS 1.3 CertificateVerify
- Jacobian point arithmetic (double a=-3, general add with the H==0 special
  cases), double-and-add scalar mult; field/scalar arithmetic reuses the
  bignum Montgomery multiply and modexp (Fermat inverses mod p and mod n)
- validates r,s in [1,n-1] and that Q is on the curve (invalid-curve guard)
- verify-only public data -> not constant-time by design
- renamed the P-256 field mul to fpmul to avoid the clash with X25519's fmul
- VERIFIED against a python ECDSA-P256 vector; tamper + wrong-hash rejected;
  test_crypto 69/0; ASan/UBSan clean; battery green
- phase D COMPLETE (RSA PKCS1+PSS + ECDSA-P256). Next E: ASN.1/X.509

(cherry picked from commit 92c996ba96d785d098c173d4ac6ace9052ef6a2f)
2026-09-15 01:15:31 +02:00
d0bd66c704 feat(crypto): RSA signature verification (rv2 9 phase D part 1, PKCS1 + PSS)
- wo_rsa_pkcs1_sha256_verify + wo_rsa_pss_sha256_verify (SHA-256), for the
  server cert chain and TLS 1.3 CertificateVerify
- bignum: Montgomery multiply (CIOS, 64-bit limbs, __int128), modexp with the
  public exponent (R^2 via 128k modular doublings, no division); MGF1-SHA256
- verification is public data only -> NOT constant-time by design (correct and
  much simpler than a private-key op)
- assumes a full-length modulus for PSS emBits (standard RSA-2048/3072/4096)
- VERIFIED against python cryptography RSA-2048 vectors (PKCS#1 v1.5 + PSS,
  salt 32); tamper + wrong-hash rejected; test_crypto 66/0; ASan/UBSan clean;
  battery green
- internal C, no builtin/compiler change. Remaining in D: ECDSA-P256 (D2)

(cherry picked from commit 9118177fbfd03eff9757defea6931afdd68830c4)
2026-09-15 01:15:31 +02:00
aee98661d2 feat(crypto): X25519 key exchange (rv2 9 phase C, RFC 7748)
- wo_x25519: constant-time Montgomery ladder + mask-based conditional swap,
  radix-2^51 field arithmetic with __int128 products (curve25519-donna-c64,
  public domain); scalar clamped, u-coord high bit masked per RFC 7748
- internal C (consumer is the TLS ECDHE handshake); no builtin/compiler change
- KAT-gated in test_crypto: RFC 7748 §5.2 both direct vectors AND the
  1000-iteration base-point test; test_crypto 61/0; ASan/UBSan clean; battery green
- fixed one transcription bug found via the KAT: crecip needs 5 final squarings
  (p-2 = 2^255-21 = (2^250-1)*2^5 + 11), not 3
- rv2 9 ladder: A (AEAD) + B (HKDF) + C (X25519) done; next D signatures/RSA

(cherry picked from commit f41b1c5f56caa841d1904382830baff0f75525d9)
2026-09-15 01:15:31 +02:00
36ce2332ef feat(crypto): HKDF-SHA256 for the TLS 1.3 key schedule (rv2 9 phase B)
- wo_hkdf_sha256_extract/expand (RFC 5869) + expand_label (RFC 8446 §7.1),
  internal C over the existing hmac_sha256; SHA-256 (mandatory-suite hash;
  SHA-384 a later add for the AES-256 suite)
- no builtin, no compiler change -- no .wo consumer yet (the TLS handshake
  is the consumer); exposed for the C unit test
- KAT-gated in test_crypto: RFC 5869 Test Case 1 (PRK + 42-byte OKM) and
  three HKDF-Expand-Label vectors (key/iv/derived-secret shape); 57/0,
  ASan/UBSan clean; runtime battery green
- rv2 9 ladder: A (AEAD, = rv2 8) and B (HKDF) now done; next C X25519

(cherry picked from commit c8d27b6c89a80cd97a996ff7d8b64ff4895e4b26)
2026-09-15 01:15:31 +02:00
94d5f176f7 feat(crypto): portable constant-time AES-GCM software fallback (rv2 8 phase C)
- no-intrinsics AES: S-box = GF(2^8) inverse via a fixed-exponent power ladder
  (constant-time in the input, no tables), constant-time gf8_mul, byte-oriented
  ShiftRows/MixColumns/key-expansion (AES-128 and AES-256)
- constant-time GHASH: bit-by-bit GF(2^128) multiply (mask-driven, no tables)
- aes_gcm_seal/open now dispatch: AES-NI path when present (and not forced
  software), else this portable fallback -> AES-GCM works on ANY CPU, so the
  phase-B no-AES-NI trap is retired
- wo_aes_force_software test hook; both hw and sw paths verified against NIST
  SP 800-38D cases 4 (AES-128) and 16 (AES-256) byte-for-byte; test_crypto 48/0;
  ASan/UBSan clean; full runtime battery green
- ARMv8 crypto-extension hardware path deferred (untestable on x86-64 host)

(cherry picked from commit dccf650899798401a9adac8489f34c85ed9304af)
2026-09-15 01:15:31 +02:00
1ef4e463ec feat(crypto): AES-GCM via AES-NI + PCLMULQDQ (rv2 8 phase B, ids 113/114)
- aes_gcm_seal/open, AES-128 and AES-256 (variant by key length 16/32),
  nonce 12 bytes, out = ciphertext||tag; open returns nil on auth failure
- hardware path only (phase B): AES-NI key schedule (128/256) + block, GHASH
  via PCLMULQDQ with the fast GF(2^128) reduction, GCM mode (J0, CTR from
  counter 2, GHASH over aad|pad|ct|pad|len, tag = GHASH ^ AES(J0))
- constant-time by hardware; target-attributed functions + __builtin_cpu_supports
  gate so the binary stays portable -- no AES-NI traps with a clear message
  (bitsliced software + ARMv8 paths are phase C)
- wiring: wob.h ids + WO_B_MAX 114; builtin.c crypto range; loader arity 4;
  emit.ml (ids/arity/return/name); types.ml (register + return type)
- VERIFIED: matches NIST SP 800-38D cases 4 (AES-128) and 16 (AES-256) and the
  python cryptography reference byte-for-byte; KAT-gated in test_crypto (36/0);
  ASan/UBSan clean; runtime battery + compiler 557/0 green

(cherry picked from commit f12a745a3c1313847f9d7f65e53bcd8093758af9)
2026-09-15 01:15:31 +02:00
ac52c3fdb5 feat(crypto): ChaCha20-Poly1305 AEAD (rv2 8 phase A, ids 111/112)
- hand-rolled ChaCha20 + poly1305-donna-32 + RFC 8439 §2.8 AEAD in crypto.c;
  constant-time (add/xor/rotate + limb math, no tables, no data-dep branches),
  constant-time tag compare
- two bare-name crypto-family builtins beside sha256/hmac:
  chacha20poly1305_seal(key,nonce,aad,pt) -> Bytes (ct||tag)
  chacha20poly1305_open(key,nonce,aad,ct||tag) -> ?Bytes (nil on auth fail)
  key 32B, nonce 12B (caller-supplied, per TLS's per-record nonce need)
- wiring: wob.h enum + WO_B_MAX 112; builtin.c crypto dispatch range; loader.c
  arity 4; emit.ml (ids, arity_of 4-case, return type, is_builtin_name,
  name->id); types.ml (registration + return type)
- VERIFIED: matches RFC 8439 §2.8.2 byte-for-byte (vs python cryptography +
  the RFC vector); test_crypto 24/0 (Poly1305 §2.5.2 + AEAD seal/open/tamper);
  ASan/UBSan clean; runtime battery + compiler 557/0 green
- first rung of the TLS ladder (rv2 9 phase A)

(cherry picked from commit 961854a8f4e9e632b6fa17f7f2e519e2d08f4936)
2026-09-15 01:15:31 +02:00
e91a3704fe feat(net): net.connect outbound TCP client (id 110)
- new builtin net.connect(host, port) -> Int: the outbound-socket gap
  language 38 named and jarvis surfaced; the client half of the net verbs
- getaddrinfo for DNS (v4/v6, numeric or hostname), blocking connect with
  the same EINTR/stop handling as net.connect_unix, then O_NONBLOCK for the
  park plane; returns the same fd-scalar accept yields
- wob.h enum + WO_B_MAX 110; types.ml registration; loader.c arity;
  builtin.c sysio dispatch range extended to WO_B_NET_CONNECT; sysio.c impl
- verified: numeric IP + hostname (DNS) connect to a local listener, closed
  port traps cleanly; ASan-clean; runtime battery 0 fail
- deferred (next slice): net.connect_dl deadline/park variant (no shard
  stall during handshake), on the accept_dl pattern

(cherry picked from commit 13c6f124428243b4956fbb4eceb1e7d0206d45f2)
2026-09-15 01:15:31 +02:00
1d78e0fa70 fix(runtime): don't deref a poisoned class's NULL fmap during migration
- wo_schema_diff poisons a class (fmap=NULL, new_cid=NONE) when a
  referenced/nested type changed or a field type is incompatible — the
  field-level map does not apply and wo_wal_migrate transcodes it instead.
- main.c's pre-migration "migrating `X`: +/-fields" print loop dereferenced
  fmap unconditionally, so a poisoned-but-field-added class (e.g. a wmux
  Window whose nested Vte gained fields) was a NULL read → SIGSEGV at boot,
  before the migrate call could refuse or transcode.
- guard the field detail on fmap != NULL; for a poisoned class print
  "(a referenced type changed — cannot migrate in place)" and let
  wo_wal_migrate proceed. It then transcodes cleanly when no record blocks
  it — so an additive nested change (Vte +oscbuf +title) now migrates and
  the session replays, instead of crashing serve.
- test_wal 5966/0; verified against the real WAL that crashed (recovers
  session `main`); wmux gate 52/0.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
(cherry picked from commit 35efa214d20dd7b055910ae66e4bfcc6201821c9)
2026-09-15 01:15:30 +02:00
5e8e0960bc feat(rt2): term.size + term.width — the wmux ladder's last runtime asks
- term.size(fd) -> ?TermSize{cols,rows}: TIOCGWINSZ, resize's read twin;
  nil = not a tty (expected answer, never a trap)
- term.width(cp): libc wcwidth under C.UTF-8 (LC_CTYPE set on first
  use, host-locale fallback): -1 control, 0 combining, 1, 2
- ids 108/109 (all four registrations); TermSize predeclared
- legs: PTY sized 77x33 from outside answers exactly that, pipe answers
  nil, widths a/CJK/combining/BEL = 1/2/0/-1; test_term 81/0, woc 557/0
- story runtime-v2 6 recorded done; board row appended

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
(cherry picked from commit 1514fb46c21c4318856cfcb3ca2b4d430caba72b)
2026-09-15 01:15:30 +02:00
6e997759e5 docs(rt2): close out runtime-v2 1-5
- five stories status: done; 00-story records the one-run landing
- spec History: three implementation amendments (Signal record not
  scalar, caller-owned stdio fds, handler-latch instead of signalfd)
- board NEXT PLAN entry with measured findings (zero transport code
  added; the tty-across-the-socket handover proven; the double-raw
  refusal restoring the terminal — the "bug" that was the design
  working); section rows flipped; graph nodes green
- CODE-LOGIC.md: the runtime-v2 section
- full belt quoted on the board: suites 0 fail both flavors (test_proc
  193/0, test_term 60/0), woc 557/0, subprocess 12/0, site 23/0

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
(cherry picked from commit bc1b4f070693eb755ad6a9fd0c853fb3e2bda347)
2026-09-15 01:15:30 +02:00
f52ee83ff4 feat(rt2): send_fd/recv_fd/connect_unix — an fd crosses the socket
- sendmsg/recvmsg with one SCM_RIGHTS fd and a sentinel byte; EAGAIN
  parks in the net mould; the received fd arrives nonblocking as a plain
  Int every fd verb accepts
- SO_DOMAIN gate: send_fd on anything but a unix socket refuses by name;
  plain bytes deliver nil from recv_fd
- net.connect_unix carried here (iteration 38 still pending)
- legs (single-fiber: unix connect completes while the listener holds
  the handshake): a pipe's read end crosses and still reads "ping"; a
  tty crosses, term.raw works on the RECEIVED copy and destroy restores
  it; refusal and nil legs verbatim. test_term 60/0
- full belt: all suites 0 fail both flavors, woc 557/0,
  subprocess-accept 12/0, site-accept 23/0

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
(cherry picked from commit 1d689027920d6814f87b97c216b0cb42f7eba3e9)
2026-09-15 01:15:30 +02:00
22ba51bfd3 feat(rt2): term.raw/restore — no wrecked tty, ever
- two verbs on any tty fd; saved termios in a per-shard 8-entry table;
  double-raw and restore-without-save refuse by name
- restore is a RUNTIME obligation: vm_unwind at depth 0 (uncaught trap,
  fiber reap) restores the dying fiber's entries newest-first, and
  wo_vm_destroy sweeps the rest — proven twice in the legs: a DIV0
  while raw restores, and even the double-raw REFUSAL (itself a trap)
  restores the first raw
- test_term 39/0 against a real PTY pair made by the test

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
(cherry picked from commit b439387dbf4f653e034c721bc3f08b83616c5e24)
2026-09-15 01:15:30 +02:00
c55d6e1d33 feat(rt2): signal.on — latched signals become Signal records for actors
- mechanics amendment to the spec (recorded at close-out): no signalfd —
  the stop-latch pattern generalized. An async-signal-safe handler
  latches the number, bumps a sequence and pokes shard 0's wake eventfd;
  wo_io_wait's loop head drains latches into fresh Signal{sig} records
  delivered via runtime_notify (exported as wo_actor_notify)
- payloads must be heap objects (vm.c drops them unconditionally) — the
  Signal record exists exactly for that; class id rides the call as the
  appended record operand (sm_record drives it even with no return)
- offerable: WINCH/CHLD/HUP/USR1/USR2; SIGTERM/SIGINT refused naming the
  stop latch; shard-0-only registration; coalescing disclosed
- stdlib_modules gains `signal` (and `term`, next task)
- test_term: a real child kills the test process with USR1; the actor's
  multi holds one coalesced delivery; refusal leg verbatim. 14/0

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
(cherry picked from commit 14e03a6a4a343b97c9b47fab1a5e3c4bb69d8201)
2026-09-15 01:15:30 +02:00
0c7d0530e9 feat(rt2): spawn_pty + resize — a child that believes it owns a terminal
- posix_openpt/grantpt/unlockpt/ptsname_r (plain libc, no -lutil); child
  setsid + opens the slave as its controlling terminal, initial
  TIOCSWINSZ from the call
- Child.stdin == Child.stdout = the master (caller's copy); the slot
  keeps a private dup so resize survives the caller closing theirs
- proc.resize -> TIOCSWINSZ; refuses by name on a pipe child
- legs: test -t proves a real tty; stty size reads "24 80" then "40 120"
  after a mid-sleep resize; refusal asserted; test_proc 193/0 ASan clean

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
(cherry picked from commit 9836c9cd5197153517c054b243cab3b453d130a0)
2026-09-15 01:15:30 +02:00
803ff0b790 feat(rt2): proc.spawn/wait_dl/signal — the streaming child
- a child is fds: Child {id, stdin, stdout, stderr}, driven by the
  existing net verbs (echo leg proves cat round-trip through write_dl/
  read_dl); caller owns the fds, the runtime owns pid + pidfd
- wait_dl parks on the pidfd: code on exit, nil at the deadline with the
  child untouched; one waiter per id, a second refuses by name; stale
  ids refused via a generation counter in the handle
- proc.signal through pidfd_send_signal; actor_die kills the streaming
  children the dying actor owns; dead fibers cannot linger as waiters
- ids 97-107 registered wholesale (wob.h, loader arities, dispatch
  bound); Child + Signal predeclared records in types.ml; unimplemented
  ids trap at the default case until their task lands
- test_proc 168/0 (echo, wait trio, one-waiter refusal, 200-round churn
  fd-flat), suite ASan clean, woc-test green

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
(cherry picked from commit 9be87f159f1bf9cdd509ceed160e7ea518fde46c)
2026-09-15 01:15:30 +02:00
b9ce271b3d fix(lang41): an unadopted shard must not impersonate shard 0
- root cause: a worker's runtime is initialised lazily on first fiber
  adoption, and rt.shard_id is stamped only there — but INBOX_READY[i]
  is set at thread creation. A shard that never adopts is still settled
  at shutdown, carrying rt.shard_id 0 from the memset
- it then impersonated shard 0: wo_drop_obj saw 0 == 0 for anything the
  primary allocated, took the "we are home" branch instead of routing,
  and called class_free against rt->classes, which lazy init never
  filled. &rt->classes[class_id] off a NULL base is the faulting read
- fix: stamp the runtime's real identity at thread creation. An
  uninitialised shard owns nothing, so its true id makes every payload
  correctly foreign and routes it to an owner that can free it
- ASan could not name this: the arena is one hand-managed malloc block,
  so intra-arena reuse is invisible and it surfaces as a bare SEGV
- pinned by tests/regress/lang-41, driven from db-actor-accept. Needs
  multiple shards (the corpus runner pins WO_SHARDS=1) and the ASan
  build. SEGVs twice per run unfixed, clean fixed
- the HANG is a separate defect and is NOT fixed: with this in place the
  harness stops losing whole sections, but idempotent-stop-2 still
  fires ~1 run in 6. The story records where to look

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
(cherry picked from commit 9dca0b4b4727b976d326b29cb4c6522b62d48a73)
2026-09-15 01:15:30 +02:00
b17b848403 docs(lang42): close out iteration 42
- story frontmatter status: done, Progress section records what landed
  vs the spec (everything, same day as the brainstorm)
- board: NEXT PLAN entry with the six standup answers (deadlock proven
  real: 5 s hang, 8192-byte truncation; 15 ms after; ping 2 ms during a
  parked child; 1000 spawns fd-flat; SIGTERM leaves no child); pending
  row flipped to DONE
- graph: node 42 class done, same change as the board row
- runtime/src/CODE-LOGIC.md: the bounded-subprocess section (bundle
  park, slot registry, ownership sweeps, raw pidfd syscalls)
- full belt at close: 19 runtime suites 0 fail (test_proc 128/0),
  woc-test 557/0, subprocess-accept 12/0, site-accept 23/0

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-09-01 22:19:25 +02:00
5d1c82bbd6 feat(lang42): deadline and output caps refuse by name, shard keeps scheduling
- proc.run_dl reachable: dispatch range extended to id 96 (builtin.c) and
  the loader arity table gains [WO_B_PROC_RUN_DL] = 6 — without both, the
  builtin answered "unknown stdlib builtin" (WO_T_EXPLICIT)
- deadline leg: sleep 10 vs 100 ms deadline traps WO_T_IO naming the
  deadline in ~120 ms; the pid is gone (waitpid -1 = ECHILD) and the fd
  count is flat; a worker fiber completes WHILE main is parked — the
  shard was never blocked
- cap legs: stdout and stderr caps trap naming "cap 1000", child dead
- argv multi carries a drop entry at the run pc: a trapping run frees it
  (LeakSanitizer caught the miss)

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-09-01 22:19:25 +02:00
258222c3a1 feat(lang42): proc.run parks — pidfd + epoll bundle + child registry
- deadlock proven first: chatty child (200 KB stdout, stderr held open)
  hung the old sequential drain 5.0 s into the alarm, code -1, stdout
  truncated at 8192; the leg demands completion under 4 s
- rework: nonblocking pipe read ends + pidfd_open behind one epoll fd the
  fiber parks on (the _dl retry mould); both pipes drain on readiness, so
  the deadlock is gone structurally — leg passes in 15 ms
- wo_child slot table in wo_vm (32/shard) carries cross-park state; caps
  refuse by name (kill + WO_T_IO), deadline armed via dl_active/dl_at,
  defaults 30 s / 1 MiB / 64 KiB
- WO_B_PROC_RUN_DL = 96 shares the case (per-call deadline_ms/out_cap/
  err_cap; compiler row lands in a later task)
- fib_reap kills a reaped fiber's child; wo_vm_destroy sweeps the table
- raw syscalls for pidfd_open/pidfd_send_signal: glibc 2.35 build floor
  has no wrappers
- all 19 suites green under ASan+UBSan

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-09-01 22:19:25 +02:00
27aecd3dc1 feat(db2-migrate): boot performs the migration, and refuses by name
- main.c builds the compiled schema (names out of the constant pool,
  which the database layer never sees), peeks the log head before
  replay, and diffs: match replays as-is, add/delete migrates through
  the transcode, poisons refuse naming class, field and what to do
- fixed en route: a poisoned class SKIPPED the identity check, so no
  transcode ran and replay greeted the shape mismatch with the generic
  "corruption" — the exact message this iteration exists to replace.
  A poison now forces the transcode, where it either bites with its
  text or passes harmlessly when the class has no records
- the schema head is written LAZILY, ahead of the first real record:
  an eager head broke the documented "durable: false writes ZERO
  bytes" contract by 75 bytes and the residency gate caught it
- end-to-end at the language level: fresh boot seeds, identity
  replays, +field migrates with "migrating `Note`: +flag" and reads 0,
  retype refuses naming `val`, and the refused log still boots the
  previous binary untouched
- gates: wovm-test all green (test_wal 5951/0), woc-test clean,
  residency-accept 14/0, site-accept 23/0

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
(cherry picked from commit b21943aa91152ebdcfc72bd4c9fba38730ab1c2f)
2026-08-31 21:54:27 +02:00
49a0a9d047 fix(db2-delta): refuse resident:keys with no WO_DATA at runtime
- loader stopped refusing durable:true+resident:keys once UPDATE
  landed; nothing replaced it at runtime
- rows for such a table live only in the WAL, so every read failed
  with a misleading "no such row" instead of naming the problem
- main.c now refuses at startup, names the class, exit(2)
- residency-accept.sh gains a leg: refuses without WO_DATA, still
  runs with it — verified failing before the fix, passing after

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
(cherry picked from commit 3ea6d6452f260d45f92045c0f300fa49faa1d810)
2026-08-30 20:38:03 +02:00
e08c26309a feat(db2-delta): lift the resident:keys refusal, prove it end to end
- loader.c: delete the INCOMPLETE-update BAIL; durable:false +
  resident:keys stays refused (nowhere to read from)
- table.c: root-cause fix for the Text-index gap — a keys-resident
  borrow now holds ENGINE values, matching wo_row_ptr's contract
  (table.h's "no VM pointer" doctrine), not a VM-decoded row. Fixes
  idx_hash/idx_cols_equal/wo_idx_probe AND db.c's GET_FIELD/PROBE
  arms with one change; reproduced pre-fix as an ASan
  heap-buffer-overflow
- docs/examples/residency: Product is genuinely resident:keys;
  residency-accept.sh's refusal leg replaced by proving the program
  runs and stock survives a restart (11/0)
- test_wal.c: oracle test drives resident:all and resident:keys
  through the same update sequence and asserts identical rows;
  Text-indexed-update test catches the representation bug; five
  pre-existing tests corrected to the fixed contract (4746/0)
- story, README, status board, CODE-LOGIC.md updated; three known
  limitations documented: mid-drain stale reads, O(N^2) replay in
  chain length, compaction blind to per-row chain length

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
(cherry picked from commit b87c68f950f01aa5e572fbb86a0f374adc83d813)
2026-08-30 20:37:49 +02:00
8311330531 docs(db2-keys): reconcile databasev2 and porch markdown with the code
- loader's resident:keys refusal said "rows are still fully resident"
  and "until tasks 5c/5d land". Both false since f606fc9. Corrected to
  name the real blocker: UPDATE needs read-modify-append
- databasev2 00-story: the sequence graph drew 2->3->4, which reads as
  3 needing 2 and 4 needing 3. Both backwards, and it still drew the
  2->5->6 path the 2026-08-27 amendment retired. Redrawn stating only
  real dependencies, with 4 and 3 shown as composing rather than
  ordered, and the execution order that actually happened
- databasev2 03: the hazard and its Outstanding entry both claimed
  nothing fails "because iteration 2's storage half is unimplemented".
  Marked discharged, and recorded that the hazard named only half the
  danger — the bitmap walk would have dropped keys rows outright
- databasev2 06: pending -> hold (largely superseded, revisit only on
  a measurement); dated its 5c/5d references
- porch 01: rewritten to the settled shape. readiness ready, status
  in-progress, phases B and C marked superseded with why
- porch 01 claimed time.after "is still a reserved builtin id". False —
  builtin 90, implemented. That claim is what made the iteration look
  cheaper than it is
- porch README gains honest ledger rows for both features (partial,
  being rebuilt), not shipped
- skill-catalog README pointed at a story path that moved tracks;
  linkcheck now 0 broken

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
(cherry picked from commit b3d8c403e1d19ac27ec966de85cb293e0765795c)
2026-08-30 20:36:55 +02:00
e16d4896f8 feat(db2-keys): inserts and boot — payload dropped after the barrier
databasev2 2, task 5c. The write and boot halves. Still not exposed: the
loader refuses `resident: keys` until 5d rewires the readers.

GROUP COMMIT FORCED THE DESIGN. A keys-resident payload can only be dropped
once its record is durable, but databasev2 4 deferred the barrier to the drain
— so at append time the bytes are still in the staging buffer and the recorded
offset would pread ZEROS. Dropping at append would have produced rows that
read as garbage, intermittently, only under multi-shard load.

So the drop is recorded, not performed:

- wo_wal gains a pending-drop list, the same shape as the drain's held replies
  and for the same reason
- both write paths take the offset BEFORE the append (wo_wal_next_offset) and
  record it; the inline path flushes right after its own commit, the request
  path's flush runs in the drain immediately after the barrier
- if the process dies before the barrier the list dies with it, which is
  correct: nothing was dropped and nothing was lost
- an out-of-memory pend is ignored on purpose — the row simply stays resident,
  which is safe

Boot: replay now leaves a keys-resident table pointing at the LOG. Each record
is applied normally, so indexes and uniqueness are built exactly as for any
other table, and the payload is then dropped with THAT record's offset. For an
update the later record wins, because each apply overwrites the map in order —
the rule replay already follows.

Tests: the round trip (insert, commit, drop, read back with Text intact) and
now BOOT — a fresh wo_db replays the store and every row materialises from the
log, count intact, nothing in a slab.

Verified: just wovm-test — 36 suites 0 fail, test_wal 4301 pass, cli_smoke OK.

REMAINING (5d), and precise: every reader still goes through wo_row_ptr, which
for a keys table would index a freed slot. The scans in db.c walk the BITMAP,
and a keys table's bitmap is empty by construction — so a query over one would
today return no rows at all. That, FK restrict, and the @unique shadow are 5d,
and the loader refusal stays until they land.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
(cherry picked from commit 08abd09bf88918f2582e74713dc7903beb8aaeb8)
2026-08-30 20:36:31 +02:00
7cc80405dd feat(db2-keys): storage — drop the payload, read it back from the log
databasev2 2, task 5c step 2. The storage half the accessor was left waiting
for. Not yet wired into insert (that and 5d remain), and the loader still
refuses `resident: keys`, so nothing is exposed to a program yet.

- wo_db gains an `rt` back-pointer, set in main.c beside VM.rt.db. wo_rt
  already carries `db` and `wal` as opaque handles, so this closes the loop
  and a borrow can reach the log WITHOUT threading a wal pointer through
  eleven call sites — which is the whole reason 5c is one accessor
- wo_row_drop_payload: the operation the plan recorded as MISSING. Frees the
  slot and its engine-owned values, then re-points the id map at the record's
  log offset (off + 1, reusing the same 0-is-empty trick as slot + 1). It
  deliberately does NOT touch the secondary indexes (they store row ids, so
  they stay correct), does NOT decrement count (the row is still live, only
  its backing moved), and does NOT remove the id (that is how it is found)
- wo_row_borrow materialises for a keys table: reads the offset from the id
  map, calls 5b's wo_wal_read_row_at into the per-table scratch, and checks
  the record actually holds the expected class and id — a compaction that
  moved records without rebuilding the map lands exactly there, which is the
  obligation recorded at wo_wal_compact
- fully-resident tables keep today's path and pay one predicate

A REAL BUG, exposed the first time the path was used: wo_row_release freed the
materialised values with the ENGINE's allocator. They are VM values —
wo_wal_read_row_at is the out-gate and always copies — so ASan reported a
bad-free immediately. It now drops them through the runtime. That stub was
written in 5c step 1 for a path that did not exist yet.

Recorded while implementing: wo_wal_next_offset's contract says to trust an
offset "only after the matching commit returns 0". Group commit (databasev2 4)
defers that barrier to the drain, so db.c can no longer check inline — but part
A also made a failed commit FATAL, so no execution can record an offset whose
record never became durable. Same guarantee, different mechanism.

Test: a heap-valued row is inserted, committed, has its payload dropped, and is
read back out of the log with its Text intact; count is unchanged (still live);
and a second borrow succeeds, which fails if release did not clear the scratch.

Verified: just wovm-test — 36 suites 0 fail, test_wal 4273 pass, cli_smoke OK.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
(cherry picked from commit 125bd09218d616b2a16b140de770d3f38b45f0ac)
2026-08-30 20:36:31 +02:00
02b4b13a52 Merge master into db-residency-doctrine — and close the two half-exposed features
The branch was 17 ahead / 25 behind with 11 conflicting files, and drifting
further: db.c had been rewritten twice on master since (group commit, then
compaction). Resolved rather than rebased so both histories stay legible.

Conflicts, and how each was settled:

- db.c: BOTH semantics kept. Master's fatal path and compaction check now sit
  behind the branch's `table_is_durable` predicate, in all three inline arms —
  a volatile table reaches neither the barrier nor the compaction check
- db-bench sample: every mode from both sides (growth, growth-verify, randread,
  replayseed, wmix) and ONE `boot` mode, which both sides had added
  independently
- db-bench.py: all six legs kept. Both sides had also grown the same
  WAL-size helper under different names; collapsed into one
- perf-targets: the branch's §5 (RAM ceiling) then master's §6/§7 — master's
  numbering had already assumed a §5 it did not have
- story frontmatter: master's `status` (the landing truth) plus the branch's
  `readiness` axis. 03 would have read `done` + `refine`, which is a
  contradiction — it was brainstormed and landed on master, so `ready`
- board: both standup blocks newest-first; master's chain rows (a superset);
  the branch's databasev2 1-2 rows with master's 3-4. Fixed a stray `|` in
  master's row 3
- baseline: master's, then REGENERATED from a full campaign — 143 metrics,
  132 checks, 0 failures with both sides' legs present

TWO HALF-EXPOSED FEATURES FIXED, because the merge rule is that master gets
no feature that is honoured in name only:

- `resident: keys` PARSED, set a .wob flag, and did nothing: rows stayed fully
  resident. A developer could declare a 120 GB table keys-resident, watch it
  compile, and be OOM-killed. The loader now REFUSES it with a message naming
  what to write instead, until tasks 5c/5d land. The compiler still parses it
  and its AST golden still passes, so the grammar work stays tested
- `durable: false` was honoured ONLY on the inline path. wo_db_exec_req had no
  guard at all, so a volatile table written from an actor on a worker shard
  would still be logged — precisely porch's session-table case, and precisely
  what iteration 2 exists to provide. All three request-path arms now carry the
  same predicate. Found by reading the merged code, not by a test: the obvious
  probe runs main() on the primary and therefore only exercises the inline path

Verified on the merged tree: wovm-test 0, woc-test 0, oop-e2e 122/0,
residency-accept 8/0, db-bench 132/0, linkcheck clean.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-29 10:14:25 +02:00
40d56c4664 perf(wal): checkpoint measured — 2.16x space, 1.78x boot, 2.7ms pause — T5
databasev2 3, task 5.

Full campaign, same workload twice, differing only in whether
checkpointing may fire:

- WAL used 1962358 -> 907094 bytes (2.16x reclaimed)
- boot 114 -> 64 ms (1.78x), median of 3
- stop-the-world pause max 2651us against a STATED 50ms budget

The budget is asserted, not assumed: 50ms is a stall a serving process
can absorb without a client seeing a timeout, and the leg fails if it is
exceeded. The pause is O(live rows) — at ~181 MB/s a 1GB live set implies
~5.5s, which is the number an incremental design must be bought against.
The spec deliberately did not buy it in advance.

FOUND BY MEASURING: the dump was 8x slower than it needed to be. It
flushed through wo_wal_commit, which fdatasyncs, so it paid one barrier
per 256 records. Intermediate durability there is worthless — the temp is
not authoritative until the rename and is fsynced once immediately before
it. With a single final barrier:

- ~107KB live: 23948us -> 2903us
- ~500KB live: 36361us -> 7526us
- ~1.98MB live: 107649us -> 13212us
- marginal ~22 MB/s -> ~181 MB/s, sync-bound to bandwidth-bound

Correctness re-proven after that change: wovm-test 36 suites 0 fail,
test_wal 760 pass including the 40-round kill-during-compaction battery.

Two measurement defects of my own, fixed rather than reported:

- boot measured through the driver's run() helper reported 251ms both
  with and without checkpointing — run() samples RSS on a 250ms poll, so
  every timing floors at the quantum. Measured directly instead, median
  of 3
- ckpt.reclaim_x was recorded as lower-is-better by the default detector,
  which would have PASSED "reclaimed nothing" and FAILED an improvement:
  the feature's central claim, gated backwards. Now higher-is-better,
  gated at 15% while the wall-clock metrics stay wide — waiving them all
  would have left the leg ungated, part A's task 4 mistake

- sample gains a `boot` mode that does nothing, so boot time is boot time
- walstats now reports compactions, pause max/total and compacted bytes
- baseline refreshed from the FULL campaign (N=20000, crash_reps=3), and
  a fresh full run passes 116 checks 0 failures
- gate bites: reclaim_x doctored to 1.0 -> FAIL on exactly that metric

One flake seen and checked, not papered over: durable.sN.query.ops_sec
failed once at 53% below baseline. It is a read-only metric that touches
no WAL code, and a re-run passed 116/0 with the box at load 1.85.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-29 06:20:13 +02:00
aa89fb6c97 feat(db): the checkpoint trigger, and compaction is wired to BOTH write paths — T3
databasev2 3, task 3.

- wo_wal_should_compact is a PURE decision (used bytes, last compaction's
  measured output, floor, ratio) so it is testable without a store —
  which is the only way a policy like this gets tested at all. Denominator
  is the last compaction's real output, not an estimate of the live set:
  estimating would mean estimating Text
- 8 boundary assertions incl. "exactly 3x is not MORE than 3x" and a zero
  ratio disabling the policy rather than dividing by nothing
- MUTATION-TESTED instead of observing RED: implementation and test were
  written together, so removing the floor check was verified to fail
  exactly the two floor assertions. Equivalent evidence, stated plainly
- WO_CHECKPOINT_BYTES / WO_CHECKPOINT_RATIO at boot beside WO_MAILBOX.
  The knobs are what make the policy testable — a gate sets a tiny floor
  and forces compaction in a few writes instead of megabytes
- NO timer, per the spec: Postgres' CheckPointTimeout bounds loss from
  unflushed buffers; our records are durable at commit and an idle log
  does not grow
- the ordering rule is now asserted, not trusted: a test stages a record,
  requests compaction, and requires REFUSAL with the log untouched and
  the staged record still committable afterwards

FOUND AND FIXED a gap in my own wiring. The plan said to call the check
"after the drain's barrier", and I did — but a statement running ON the
owner shard never enters that drain, so WO_SHARDS=1 never compacted and
its log grew forever: measured 536086 bytes where the multi-shard run
held 446024. Now checked after the inline path's commit too (db.c
maybe_compact), where the buffer is equally empty. WO_SHARDS=1 went
536086 -> 260657 bytes. For a checkpoint this mattered more than part A's
equivalent gap: an unbounded log is an operational failure, not just lost
throughput.

Also corrected a measurement of my own: multi-shard logs looked unbounded
(448KB -> 1013KB -> 1647KB across 8k/24k/48k updates). They are not.
Instrumentation showed compaction ran 25 times with zero failures, each
writing MORE than the last, because the live set genuinely grows — wmix's
hist_dump and done-markers are themselves durable inserts. Final log
1631040 against a last compaction of 866432 is a ratio of 1.88, just under
the 2x threshold: the policy holding exactly.

Replies are released BEFORE compaction runs, deliberately: their records
are already durable, and holding them across a stop-the-world rewrite
would add its full duration to their latency for nothing.

Verified: wovm-test 36 suites 0 fail, test_wal 360 pass; db-bench-quick
crash.s1/crash.sN and both restart legs green, and part A still batches
(sN mean 4.16, peak 30).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-28 18:43:44 +02:00
0b618ace19 docs+fix(db): T6 closeout — and reads no longer wait for the barrier
databasev2 4 part A, task 6. Mostly documentation, plus one real fix the
full battery caught.

THE FIX. The drain held EVERY DB reply until the barrier — including
reads, which stage nothing and have no stake in durability. That parked
readers behind an fsync for no reason: durable.sN.mixread.p99 rose from
~1043us to 4057us. Only a statement that actually staged a record now has
its reply held. Caught by the gate, not by review.

THE TRADE, recorded rather than smoothed over. What remains is inherent: a
barrier blocks the owner shard LONGER (more records per fsync) though LESS
OFTEN, so anything queued behind one waits. Three full runs of the same
build gave durable.sN.mixread.p99 of 1043 / 2318 / 4147us and wmix.p99 of
8758 / 20000us — a 2-4x spread with the box near idle. So part A buys ~3x
write throughput at the cost of a longer, noisier tail on the owner shard,
and that is the strongest argument for part B (submit and keep serving).

- durable.sN.*.p99us tolerance widened to 100% WITH the reason in the
  code: a 2-4x-variable tail gated at 50% gates the disk, not the engine.
  The floor is the real guard and is not slack — mixread's (4172us) came
  within 25us of tripping on the worst run. Baseline refreshed; a fresh
  full run then passed 106 checks 0 failures

EXIT STATUS MOVED 3 -> 74 (sysexits EX_IOERR). 3 and 4 are already used by
SAMPLES for their own meanings — db-bench's own `verify` exits 3 on a
checksum mismatch, and it is the gate that exercises durability, so a
durability abort exiting 3 would have been indistinguishable from the
mismatch it should help diagnose. The low range belongs to programs.

Docs:

- story: progress, the payoff measured two ways, the cost side, criteria
  split met/outstanding, and a "part B — its premise changed" section:
  it was justified by "close the 66x gap", but that gap is two problems
  and only the concurrent one was a batching problem
- board: standup entry in the six-question shape; both databasev2 4 rows
  rewritten. They had said "close the 66x gap" — recorded as MIS-STATED
  rather than quietly renumbered
- 00-wob-format.md and 04-db-binding.md: the normative failure contract
  ("a failed WAL commit traps WO_T_IO after un-applying the row") was
  false; corrected, along with the tick-scoped group commit that never
  happened
- database/src/CODE-LOGIC.md: where the barrier runs and why there, why
  replies are held, why the inline path is asymmetric, the one failure
  rule, and how to measure it
- db-bench README: the wmix mode, the env knobs, and the tmpfs warning

Battery: wovm-test 36 suites 0 fail, woc-test, oop-e2e 119/0,
db-bench 106/0, linkcheck clean.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-28 16:48:23 +02:00
5f9598af6a feat(db-bench): prove batches form — the write-concurrent leg, T4
databasev2 4 part A, task 4. Scope extended with developer approval: the
plan authorised touching the sample only for observability, but no
existing leg has enough concurrent durable writes to exercise group
commit at all, so the payoff was unevaluable either way.

The finding that forced it:

- `mix` writes on one op in ten with C=4 (all_mode calls mix_mode(n/10,
  4); Mixer writes on i % 10 == 9), so the quick run performs 20 writes
  total. Measured mean batch 1.01 over 3112 barriers, peak 3
- that is a property of the WORKLOAD, not the mechanism: peak 3 of a
  possible 4 shows batches form whenever writes actually coincide

- `wmix N C` added: every op a durable write, C at once. Updates rather
  than inserts, so it is comparable to mixwrite and the row count stays
  flat. Histogram kind 2 — a replayed store still holds the seeding run's
  kind-0/1 Hist rows and merging those would report someone else's
  latencies
- WO_WAL_STATS=1 prints one line at exit: batches, records, peak_batch,
  peak_staged. Opt-in, because it would otherwise pollute every durable
  program's output. Counters live in wo_wal; no builtin, the numbers are
  diagnostic and not part of the language

Measured, and it scales with concurrency exactly as designed:

- C = 4 / 16 / 64 -> mean batch 1.13 / 1.76 / 5.35, peak 3 / 10 / 39
- the gate's own legs: durable.s1 5412 records over 5412 barriers (mean
  1.0, peak 1 — the inline path, one barrier per statement BY DESIGN),
  durable.sN 7757 over 2296 (mean 3.38, peak 28) at 2x the throughput
- peak staged 1372 B settles the no-cap decision with a number: the batch
  is tiny, so the upstream mailbox bound is sufficient

- mean_batch/peak_batch are higher-is-better (the default detector would
  have called bigger batches worse)
- only the batch SHAPE metrics are waived to 100%; wmix throughput and
  latency keep real tolerances (15% s1, 50% sN) — a blanket waiver would
  have left the entire new leg ungated
- the live assertion `mean > 1.0` on the sN leg is what catches inertness
- gate bites: sN wmix ops_sec -60% -> FAIL on exactly that metric, 1 of 86

Verified: db-bench-quick 89 checks 0 failures; baseline refreshed (86
metrics).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-28 09:47:51 +02:00
76d80cc027 feat(db): one barrier per drain, replies held — T2
databasev2 4 part A, task 2. The core change, and mostly deletion.

- the REQUEST path (wo_db_exec_req) no longer commits after each append.
  Applying to RAM and staging stay exactly where they were
- wo_vm_adopt holds each DB reply envelope in a local FIFO instead of
  pushing it as the statement finishes. Pushing there would unpark the
  requester before its record is durable — the ack contract this
  iteration exists to make literally true rather than true by accident
  of every batch having one member
- at the end of the drain: ONE wo_wal_commit_fatal for everything staged,
  then every held reply. Locals rather than per-shard state: nothing
  needs to outlive the batch it describes
- "did this statement stage anything" is asked of the buffer, not guessed
  from the opcode, and that count is what the failure diagnostic reports
- the drain commits unconditionally when anything is staged, because the
  inline path relies on finding the buffer empty (task 3 documents that)
- staging failure on the request path is now FATAL via wo_wal_stage_fatal:
  the row is already in RAM and of the three verbs only insert could undo
  itself, so continuing means RAM ahead of disk. One rule
- wal_die is now shared by both fatal points

Verified — the ack contract is the thing that could break, so it is what
was tested:

- just wovm-test: 36 suites (18 x both dispatch flavors) 0 fail, cli_smoke OK
- just db-bench-quick: 85 checks, 0 failures. The legs that matter:
  crash.sN.0 — 612 acked rows all present after kill -9, which is the
  BATCHING path (multi-shard requests, held replies, one barrier);
  crash.s1.0 — 800 acked rows; restart.s1 and restart.sN replay byte-true

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-28 09:25:10 +02:00
62d29d6a77 docs(24): T10 closeout — stories done, board, graph, ledger, CODE-LOGIC
Iteration 24 closes, absorbing 31 and 34. No code in this commit.

- stories 24, 31, 34 -> `status: done`, each with a landing banner. 24's
  records the gate numbers and BOTH disclosed deviations: monitor takes
  three arguments (the caller may be `main`, which has no mailbox) and a
  v1 `call` reply is a typed scalar (which is what let the agreement be
  checked at compile time, WO-E226). 31's notes it landed INSIDE 24 and
  that a fifth mechanism it never anticipated came out of proving the
  gate — the drain guarantee (40). 34's names the gap it did NOT close:
  still no RNG, so CSRF/sessions stay blocked
- board: in-progress row cleared, marker doc deleted (convention), the
  standup entry in the six-question shape, chain note — next link is
  databasev2 4 (io_uring group-commit, chain 5)
- graph: PUBSUB2 (pub/sub + WebSockets, "rejected until here") -> done
- porch ledger: a WebSocket/pub-sub row added; the cancellation row now
  says what it actually waits on rather than repeating "the arc"; the
  README's "no WebSockets/SSE" limitation was stale — WebSockets are
  supported, SSE and chunked encoding are not
- CODE-LOGIC: runtime/src gains the actor-lifecycle section (call, death,
  the cap counter's sender/home-thread split, the monitor walk, the timer
  list), the drain guarantee, and the digest section; docs/examples/chat
  gains its own — actor topology, WHY two actors per connection, fd
  ownership, and the shutdown choreography

Battery after the doc edits: wovm-test 36 suites 0 fail, woc-test exit 0,
oop-e2e 119/0, chat 11/0, web-app 46/0, linkcheck clean.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-28 00:06:31 +02:00
60414a1754 feat(runtime): the shutdown drain guarantee — iteration 40
A message sent before the stop flag is observed must be delivered and run
before the engine stops. One rule; a spin count could never express it.

- root cause in `shard_main` (runtime/src/vm.c): NEXT_RUNNABLE() already
  stated the contract — "a WORKER on stop keeps DRAINING ... so queued
  shutdown messages (close frames!) still run" — but the IDLE branch
  contradicted it, calling fib_reap_all and breaking on WO_IO_STOP,
  abandoning its inbox for wo_engine_stop() to free wholesale
- an actor between messages is exactly that idle case, which is why a WARM
  soak server hid it: warm shards held live fibers and took the right path
- fix: while the primary's drain window is open, an idle worker adopts its
  inbox and runs what arrives; sched_yield on an empty poll so a drain
  cannot burn a core per shard and starve the actors it exists to let run
- unreachable at WO_SHARDS=1: wo_engine_stop returns early at nshards <= 1

Measured:

- fresh-server SIGTERM drain: 5 of 16 failing before, 20 of 20 clean after
- `just chat` at the FULL 1000-client soak: 11 checks, 0 failures, both
  WO_IO backends, ASan clean with zero leaks
- the fd leg settled at scale too: 1000 connections left the count at 44,
  unchanged after 20 more — lazy per-shard init, not a leak
- runtime battery 36 suites (18 x both dispatch flavors) 0 fail;
  compiler 556 checks 0 fail

- story: docs/stories/language-runtime-database/40-shutdown-drain-guarantee.md
  (chain 3 with 31, status done), board row, slice marker updated
- outstanding and named: a pin below the gate needs new multithreaded test
  infrastructure — nothing in runtime/test/ drives wo_engine_start/stop and
  no corpus fixture can trigger a stop

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-27 23:43:45 +02:00
ebc3522c40 Merge branch 'master' into chat-ws-lifecycle 2026-08-27 23:11:49 +02:00