Commit graph

12 commits

Author SHA1 Message Date
40d56c4664 perf(wal): checkpoint measured — 2.16x space, 1.78x boot, 2.7ms pause — T5
databasev2 3, task 5.

Full campaign, same workload twice, differing only in whether
checkpointing may fire:

- WAL used 1962358 -> 907094 bytes (2.16x reclaimed)
- boot 114 -> 64 ms (1.78x), median of 3
- stop-the-world pause max 2651us against a STATED 50ms budget

The budget is asserted, not assumed: 50ms is a stall a serving process
can absorb without a client seeing a timeout, and the leg fails if it is
exceeded. The pause is O(live rows) — at ~181 MB/s a 1GB live set implies
~5.5s, which is the number an incremental design must be bought against.
The spec deliberately did not buy it in advance.

FOUND BY MEASURING: the dump was 8x slower than it needed to be. It
flushed through wo_wal_commit, which fdatasyncs, so it paid one barrier
per 256 records. Intermediate durability there is worthless — the temp is
not authoritative until the rename and is fsynced once immediately before
it. With a single final barrier:

- ~107KB live: 23948us -> 2903us
- ~500KB live: 36361us -> 7526us
- ~1.98MB live: 107649us -> 13212us
- marginal ~22 MB/s -> ~181 MB/s, sync-bound to bandwidth-bound

Correctness re-proven after that change: wovm-test 36 suites 0 fail,
test_wal 760 pass including the 40-round kill-during-compaction battery.

Two measurement defects of my own, fixed rather than reported:

- boot measured through the driver's run() helper reported 251ms both
  with and without checkpointing — run() samples RSS on a 250ms poll, so
  every timing floors at the quantum. Measured directly instead, median
  of 3
- ckpt.reclaim_x was recorded as lower-is-better by the default detector,
  which would have PASSED "reclaimed nothing" and FAILED an improvement:
  the feature's central claim, gated backwards. Now higher-is-better,
  gated at 15% while the wall-clock metrics stay wide — waiving them all
  would have left the leg ungated, part A's task 4 mistake

- sample gains a `boot` mode that does nothing, so boot time is boot time
- walstats now reports compactions, pause max/total and compacted bytes
- baseline refreshed from the FULL campaign (N=20000, crash_reps=3), and
  a fresh full run passes 116 checks 0 failures
- gate bites: reclaim_x doctored to 1.0 -> FAIL on exactly that metric

One flake seen and checked, not papered over: durable.sN.query.ops_sec
failed once at 53% below baseline. It is a read-only metric that touches
no WAL code, and a re-run passed 116/0 with the box at load 1.85.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-29 06:20:13 +02:00
d432bc5301 test(wal): kill -9 DURING compaction — 40 rounds, mutation-proven — T4
databasev2 3, task 4. The plan called this the riskiest task because a race
can pass by luck, so it is argued with mutants rather than green runs.

The battery: a forked child inserts, acks, deletes the oldest so HISTORY
grows while the live set stays ~17, and compacts every 24 iterations. The
parent SIGKILLs at varied instants so kills land before, inside and after
rewrites, then replays and checks the ACKED LIVE SET. The existing
battery's "records >= acks" oracle cannot be reused: collapsing history is
exactly what compaction is for.

A REAL DEFECT IN MY FIRST VERSION, found by the failures and fixed in the
TEST, not by weakening it:

- the child acked deletes AFTER committing them, so a kill in between left
  the row legitimately gone on disk while the last ack still said
  "inserted" — the parent then demanded a row the engine was right to
  remove. Symptom was an acked insert missing near the end of the stream,
  ~1 run in 3
- deletes now announce INTENT BEFORE committing, so such a row's fate is
  simply UNKNOWN to the parent, which is the honest thing to assert. Every
  acked insert never marked for deletion must still be present with its
  acked value
- the stale-temp assertion was also wrong: it checked for absence after
  wo_wal_replay, which never opens the WAL. The guarantee is "removed AT
  OPEN", so the test now opens and then asserts. A temp surviving a kill
  is expected debris, not a defect

Proven to have teeth, which matters because assertions were softened:

- against the design's rejected alternative (in-place rewrite instead of
  the atomic rename) it fails EVERY run, reporting log_records=0 — the
  kill landed mid-copy and destroyed the log. That is the corruption
  rename exists to prevent
- on correct code: 10 consecutive runs x 40 rounds clean, plus the suite

Also carried the log's PREALLOCATION to the replacement. The WAL is
preallocated so appends never extend the file, which is what lets
fdatasync alone be the ack barrier; a replacement opened with prealloc 0
silently changes that property and the zero-padded tail the scan relies
on. Stated honestly: this is hygiene making the replacement equivalent to
what open() would have produced — I could NOT prove it was the cause of
the observed loss, and the ack race above explains it.

Verified: just wovm-test — 36 suites 0 fail, test_wal 760 pass, cli_smoke OK.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-29 05:55:58 +02:00
aa89fb6c97 feat(db): the checkpoint trigger, and compaction is wired to BOTH write paths — T3
databasev2 3, task 3.

- wo_wal_should_compact is a PURE decision (used bytes, last compaction's
  measured output, floor, ratio) so it is testable without a store —
  which is the only way a policy like this gets tested at all. Denominator
  is the last compaction's real output, not an estimate of the live set:
  estimating would mean estimating Text
- 8 boundary assertions incl. "exactly 3x is not MORE than 3x" and a zero
  ratio disabling the policy rather than dividing by nothing
- MUTATION-TESTED instead of observing RED: implementation and test were
  written together, so removing the floor check was verified to fail
  exactly the two floor assertions. Equivalent evidence, stated plainly
- WO_CHECKPOINT_BYTES / WO_CHECKPOINT_RATIO at boot beside WO_MAILBOX.
  The knobs are what make the policy testable — a gate sets a tiny floor
  and forces compaction in a few writes instead of megabytes
- NO timer, per the spec: Postgres' CheckPointTimeout bounds loss from
  unflushed buffers; our records are durable at commit and an idle log
  does not grow
- the ordering rule is now asserted, not trusted: a test stages a record,
  requests compaction, and requires REFUSAL with the log untouched and
  the staged record still committable afterwards

FOUND AND FIXED a gap in my own wiring. The plan said to call the check
"after the drain's barrier", and I did — but a statement running ON the
owner shard never enters that drain, so WO_SHARDS=1 never compacted and
its log grew forever: measured 536086 bytes where the multi-shard run
held 446024. Now checked after the inline path's commit too (db.c
maybe_compact), where the buffer is equally empty. WO_SHARDS=1 went
536086 -> 260657 bytes. For a checkpoint this mattered more than part A's
equivalent gap: an unbounded log is an operational failure, not just lost
throughput.

Also corrected a measurement of my own: multi-shard logs looked unbounded
(448KB -> 1013KB -> 1647KB across 8k/24k/48k updates). They are not.
Instrumentation showed compaction ran 25 times with zero failures, each
writing MORE than the last, because the live set genuinely grows — wmix's
hist_dump and done-markers are themselves durable inserts. Final log
1631040 against a last compaction of 866432 is a ratio of 1.88, just under
the 2x threshold: the policy holding exactly.

Replies are released BEFORE compaction runs, deliberately: their records
are already durable, and holding them across a stop-the-world rewrite
would add its full duration to their latency for nothing.

Verified: wovm-test 36 suites 0 fail, test_wal 360 pass; db-bench-quick
crash.s1/crash.sN and both restart legs green, and part A still batches
(sN mean 4.16, peak 30).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-28 18:43:44 +02:00
d7dde018ec feat(wal): a stale compaction temp is removed at open — T2
databasev2 3, task 2.

- wo_wal_open removes `<log>.compact` before reading anything. The only
  way one exists is a crash before the rename, which means its records
  were never authoritative
- deleted rather than ignored, deliberately: a file full of well-formed
  records sitting beside the log is exactly what a future reader mistakes
  for data

Test uses PLAUSIBLE content, not garbage — a byte copy of a real log —
because garbage would be rejected by the CRC anyway and would prove
nothing. It asserts the temp is present before the open, gone after, and
that the live log still replays to exactly what it said.

RED was an assertion failure (`access(tmp, F_OK) != 0` unmet), not a
compile error, so the test was proven to exercise the behaviour before the
behaviour existed.

Verified: just wovm-test — 36 suites 0 fail, test_wal 340 pass, cli_smoke OK.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-28 18:01:52 +02:00
0f652dd93f feat(wal): wo_wal_compact — rewrite the log, swap it in with rename — T1
databasev2 3, task 1.

- walks each class's live rows via the bitmap-over-slabs pattern db.c
  already uses in three places, appending one INSERT per live row through
  the EXISTING append path. No second encoder, no new format, and ids are
  preserved exactly because wo_wal_append_insert takes the id and reads
  the row from the store
- FLUSHES EVERY 256 RECORDS rather than staging the whole store: stage()
  grows the staging buffer by doubling and never shrinks it, so a
  one-buffer dump would hold the entire store in RAM on top of the store
  — the unbounded growth databasev2 1 identified as how this engine dies
- the switch, in order: fsync the temp file, rename over the live path,
  fsync the PARENT DIRECTORY (rename's atomicity is in-kernel; the
  directory entry is not durable until the parent is synced — Postgres
  does the same for the same reason), then reopen the descriptor, because
  the old one refers to an unlinked inode
- REFUSES when anything is staged: those records would land in a file
  about to be replaced. The caller-side guard is task 3; this is the
  backstop
- a failure leaves the ORIGINAL log intact and usable and returns -1. A
  failed checkpoint is a missed optimisation, not a durability event, so
  it deliberately does NOT take databasev2 4's fatal path
- records the bytes written, so task 3's trigger can compare against a
  measured denominator instead of estimating the live set (which would
  mean estimating Text)

Recovery is untouched — the result is an ordinary log in the ordinary
grammar, replayed from byte 0. Crash safety comes from rename, not from
code of ours.

Test asserts BOTH halves, on purpose:

- the log shrinks: 43 records (3 inserts + 40 updates of the SAME row, so
  history grows while the live set does not) -> 3 records, fewer bytes
- AND a fresh replay reproduces the store: every id present, and row 0
  carries the 40th update's value rather than its original. "It got
  shorter" is also true of a truncating bug, so the replay comparison is
  what actually proves it
- and the WAL stays usable after the swap: a further append lands after
  the compacted records, giving 4 on the next check

Verified: just wovm-test — 36 suites 0 fail, test_wal 315 pass (was 165),
cli_smoke OK.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-28 17:55:43 +02:00
5f9598af6a feat(db-bench): prove batches form — the write-concurrent leg, T4
databasev2 4 part A, task 4. Scope extended with developer approval: the
plan authorised touching the sample only for observability, but no
existing leg has enough concurrent durable writes to exercise group
commit at all, so the payoff was unevaluable either way.

The finding that forced it:

- `mix` writes on one op in ten with C=4 (all_mode calls mix_mode(n/10,
  4); Mixer writes on i % 10 == 9), so the quick run performs 20 writes
  total. Measured mean batch 1.01 over 3112 barriers, peak 3
- that is a property of the WORKLOAD, not the mechanism: peak 3 of a
  possible 4 shows batches form whenever writes actually coincide

- `wmix N C` added: every op a durable write, C at once. Updates rather
  than inserts, so it is comparable to mixwrite and the row count stays
  flat. Histogram kind 2 — a replayed store still holds the seeding run's
  kind-0/1 Hist rows and merging those would report someone else's
  latencies
- WO_WAL_STATS=1 prints one line at exit: batches, records, peak_batch,
  peak_staged. Opt-in, because it would otherwise pollute every durable
  program's output. Counters live in wo_wal; no builtin, the numbers are
  diagnostic and not part of the language

Measured, and it scales with concurrency exactly as designed:

- C = 4 / 16 / 64 -> mean batch 1.13 / 1.76 / 5.35, peak 3 / 10 / 39
- the gate's own legs: durable.s1 5412 records over 5412 barriers (mean
  1.0, peak 1 — the inline path, one barrier per statement BY DESIGN),
  durable.sN 7757 over 2296 (mean 3.38, peak 28) at 2x the throughput
- peak staged 1372 B settles the no-cap decision with a number: the batch
  is tiny, so the upstream mailbox bound is sufficient

- mean_batch/peak_batch are higher-is-better (the default detector would
  have called bigger batches worse)
- only the batch SHAPE metrics are waived to 100%; wmix throughput and
  latency keep real tolerances (15% s1, 50% sN) — a blanket waiver would
  have left the entire new leg ungated
- the live assertion `mean > 1.0` on the sN leg is what catches inertness
- gate bites: sN wmix ops_sec -60% -> FAIL on exactly that metric, 1 of 86

Verified: db-bench-quick 89 checks 0 failures; baseline refreshed (86
metrics).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-28 09:47:51 +02:00
76d80cc027 feat(db): one barrier per drain, replies held — T2
databasev2 4 part A, task 2. The core change, and mostly deletion.

- the REQUEST path (wo_db_exec_req) no longer commits after each append.
  Applying to RAM and staging stay exactly where they were
- wo_vm_adopt holds each DB reply envelope in a local FIFO instead of
  pushing it as the statement finishes. Pushing there would unpark the
  requester before its record is durable — the ack contract this
  iteration exists to make literally true rather than true by accident
  of every batch having one member
- at the end of the drain: ONE wo_wal_commit_fatal for everything staged,
  then every held reply. Locals rather than per-shard state: nothing
  needs to outlive the batch it describes
- "did this statement stage anything" is asked of the buffer, not guessed
  from the opcode, and that count is what the failure diagnostic reports
- the drain commits unconditionally when anything is staged, because the
  inline path relies on finding the buffer empty (task 3 documents that)
- staging failure on the request path is now FATAL via wo_wal_stage_fatal:
  the row is already in RAM and of the three verbs only insert could undo
  itself, so continuing means RAM ahead of disk. One rule
- wal_die is now shared by both fatal points

Verified — the ack contract is the thing that could break, so it is what
was tested:

- just wovm-test: 36 suites (18 x both dispatch flavors) 0 fail, cli_smoke OK
- just db-bench-quick: 85 checks, 0 failures. The legs that matter:
  crash.sN.0 — 612 acked rows all present after kill -9, which is the
  BATCHING path (multi-shard requests, held replies, one barrier);
  crash.s1.0 — 800 acked rows; restart.s1 and restart.sN replay byte-true

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-28 09:25:10 +02:00
ceea00e0b6 feat(wal): a failed barrier is detected, and fatal — T1
databasev2 4 part A, task 1.

- wo_wal gains `path`: the abort diagnostic is worthless without naming
  the file it could not write. strdup'd in open, freed in close; NULL is
  tolerated so the message degrades rather than crashes
- wo_wal_commit now reports WHICH half failed — WO_WAL_ERR_WRITE for
  pwrite, WO_WAL_ERR_SYNC for fdatasync. A short write and a device
  refusing the flush are different operational problems and the operator
  needs the right one named
- wo_wal_commit_fatal(w, nrec): commits, or prints one diagnostic naming
  the operation, path, errno and record count, then exits
  WO_EXIT_DURABILITY (3 — 1 is a trap, 2 is a refusal, so this takes a
  third of its own)
- retrying is not offered, deliberately: on Linux a failed fsync may have
  already discarded the dirty pages, so a second call can report success
  having written nothing. Replay is the recovery that works

- test_wal: a failed commit is DETECTED, reports the write error
  specifically, keeps the batch staged (a failed commit consumes
  nothing), and the WAL knows its own path. 165 pass (was 156)

DISCLOSED GAP: the exit path itself is not exercised. Forcing a real
fdatasync failure needs a full or read-only filesystem, which the gate
cannot arrange without mount privileges. No fault-injection switch was
added — shipping a binary that can be told to kill itself is the worse
trade, and the spec rejected it.

Verified: just wovm-test — 36 suites (18 x both dispatch flavors) 0 fail,
cli_smoke OK.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-28 09:16:35 +02:00
d24c705860 feat: iterations 19 + 17 — Float/Bytes scalars (.wob v5), library kind + internal/
- Float full stack: literals (fraction/exponent; `0..10` still a range), f64
  opcodes 34-41, @table column, WAL bit-exact replay, json fractions in and
  shortest-round-trip out. IEEE-quiet — FDIV never traps where DIV does.
- Bytes: a wo_str with its own class id, so alloc/free/copy are shared but no
  Text builtin accepts one; len/at/slice/eq/concat, base64 both ways, json
  boundary as base64; TEXT_COPY preserves the kind.
- No implicit Int/Float mixing (WO-E201 in the typechecker, not the emitter,
  which picks the opcode from one side and would misread the other).
- One IEEE deviation: float_cmp total order (NaN last, -0.0 == +0.0) for
  indexes and order-by, keys canonicalized to match. `?Float` nil is a
  reserved quiet NaN — the zero word is +0.0, WO_NIL_SCALAR's bits are -2.0.
- Renderer prefers fixed over exponential in 1e-6..1e21: pure shortest makes
  a price of 900.0 read `9e+02`. One renderer for interp/json/float_to_text.
- Fixed en route: lexer double-counted the leading digit; is_scalar_shaped
  took Float/Bytes as Int-shaped; Bytes ownership needed a shared heap-scalar
  predicate or temps never dropped; order-by bit-compared negatives backwards.
- Iteration 17: `kind = "library"` (absent = program; bad value = WO-E109),
  entry-less check mode retiring the `--emit` workaround, Go's `internal/` as
  WO-E108 at the consumer's `use`. Driver-only; VM/.wob/GC untouched.
- Framework reorg: internal/{parse,serve}.wo; http/form.wo split out to keep
  media_type/form_values public (parse.wo had grown public surface).
- Docs: link audit (97 -> 88 broken, conflict markers resolved, 2 duplicate
  stories removed), 00-code-review verified 26/27, iterations re-sequenced.
- Also carries the pre-staged pub(read)/using/#if work from the index.
- Gates: corpus 103/0, test_wal 156/0, web-app 26/0, oop-accept ALL MET.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-20 19:24:15 +02:00
44c77ff5bf feat: update-point + delete engine half (iteration 9, Task 5 engine)
- wo_row_update_field: encode new value, unique re-check against a
  shadow BEFORE any mutation (violating update leaves the row
  untouched, DB_ERR_UNIQUE), index entries moved old-hash -> new-hash,
  old engine value freed; proven by test_table (unique refusal keeps
  the row, released key becomes insertable)
- WAL UPDATE record: full-row re-log, replay = replace (remove +
  re-create same id); prefix/suffix delta recorded as later
  optimization; test_wal replays insert+update to the updated state
- builtins 62 DB_UPDATE_FIELD (cid,id,field,value) and 63 DB_DELETE
  (cid,id), commit-before-ack like insert, WO_T_UNIQUE/WO_T_DB/WO_T_IO
  mapping; dispatch range 61..63; loader arities; runner mirror
- plan Task 5 marked superseded-in-part with the recorded deviation:
  the language surface (reads, queries, row views, delete statement)
  is 9b's, where the comprehension design put it -- no interim brace-
  select grammar to retire later
- gates: test_table 839/0, test_wal 102/0, 15 suites, oop-e2e 73/0

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-15 13:01:05 +02:00
6f2f9b6f0a feat: secondary indexes + @unique trap (iteration 9, Task 4; wob v3)
- .wob v3: class records carry an index tail (flags bit0 = unique,
  col_cnt, columns) -- @table(index:[a,b]) entries plus one unique
  single-column entry per @unique field; loader validates columns in
  range and scalar/Text-kinded; emitter validates the declarations
  (unknown column, un-indexable kind => diagnostic)
- engine: db_index hash multimap per table, built from the class
  table at first touch, maintained ONLY inside wo_row_insert/
  wo_row_remove; unique checks re-compare actual column values (a
  hash is a hint); replay re-indexes via wo_row_raw_commit AFTER
  slots are filled, so recovered tables carry their indexes
- WO_T_UNIQUE = 10; a violating insert is un-applied whole (bitmap,
  hash, count, and the never-observable id reclaimed) and traps
  catchably -- the employee SEED-DUP pattern
- wo_row_insert gains err_kind so db.c maps UNIQUE/OOM/other to the
  right trap; test images and the runner's loader mirror speak v3
- fixtures: trap/db-unique-violation (code 10 exact) and
  run/db-unique-catch (catchable dup, composite index accepts
  duplicates, next id dense after a refusal)
- gates: oop-e2e 73/0, all 15 runtime suites, woc-test green,
  log-watcher 7/0

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-15 12:57:32 +02:00
c3741262cb feat(database): typed WAL + boot replay (iteration 9, Task 2)
- database/src/wal.{c,h}: framed records len|crc32|payload|mark
  ("WOL1" written last -- no mark, no record), typed-row payloads
  walking the class-table kinds (nested records, containers, nil
  encodings), little-endian like the loader
- commit order verbatim from the shipped phase-D pattern: RAM apply,
  stage, ONE pwrite + ONE fdatasync for the batch, ack after -- group
  commit is everything staged riding one sync
- replay decodes straight into engine-owned values (no VM at boot)
  and re-enters rows through the choke-point row API, so Task 4's
  indexes will rebuild for free; next_id advances past replayed ids
  this shard owns (wo_row_create_raw)
- torn tail = short/CRC-fail/no-mark/zero-len: intact prefix applies,
  tear dropped whole, wo_wal_open positions AT the tear so the next
  commit overwrites it; CRC-valid-but-undecodable = corruption, loud
- wo_wal_check: offline oracle, no engine needed -- the crash
  battery's verifier
- test_wal 90/0 ASan+UBSan incl. five crash-battery rounds (fork,
  insert/commit/ack-over-pipe, SIGKILL mid-stream: zero acked-but-
  missing, zero acked-but-wrong); all runtime suites green, oop-e2e
  71/0; binding doc WAL section + CODE-LOGIC + plan Task 2 checked

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-15 11:01:41 +02:00