From ad806c415d8519b9e78e3bc91956d0c0aff1200c Mon Sep 17 00:00:00 2001 From: "shoney.arickathil" Date: Mon, 10 Aug 2026 08:39:29 +0200 Subject: [PATCH] feat(compiler): nullable types (?T) support - ast.ml: Added Nullable variant to field_ty; param.ty/method_sig.ret/method_decl.ret now use field_ty - parser.ml: Parse ? prefix for field types, return types, parameters - dump.ml: Render ?T in --dump-ast output - types.ml: New typechecker (Task 6) with nullable field-kind derivation (WO_K_NULLABLE=6) - dune: Added types module - Added golden test fixtures for nullable types and statement/expression parser - Added implementation plan doc --- compiler/src/ast.ml | 308 +++++ compiler/src/dump.ml | 275 ++++ compiler/src/dune | 5 + compiler/src/parser.ml | 1155 +++++++++++++++++ compiler/src/types.ml | 387 ++++++ .../test/golden/ast/body-recovery.expected | 3 + compiler/test/golden/ast/body-recovery.wo | 6 + .../test/golden/ast/body-statements.expected | 26 + compiler/test/golden/ast/body-statements.wo | 34 + .../golden/ast/condition-recovery.expected | 2 + .../test/golden/ast/condition-recovery.wo | 6 + .../test/golden/ast/ctor-literal.expected | 17 + compiler/test/golden/ast/ctor-literal.wo | 25 + .../test/golden/ast/db-stub-nested.expected | 4 + compiler/test/golden/ast/db-stub-nested.wo | 5 + compiler/test/golden/ast/db-stub.expected | 6 + compiler/test/golden/ast/db-stub.wo | 7 + compiler/test/golden/tokens/unknown-char.wo | 1 + .../compiler/nullable-types-implementation.md | 264 ++++ 19 files changed, 2536 insertions(+) create mode 100644 compiler/src/ast.ml create mode 100644 compiler/src/dump.ml create mode 100644 compiler/src/dune create mode 100644 compiler/src/parser.ml create mode 100644 compiler/src/types.ml create mode 100644 compiler/test/golden/ast/body-recovery.expected create mode 100644 compiler/test/golden/ast/body-recovery.wo create mode 100644 compiler/test/golden/ast/body-statements.expected create mode 100644 compiler/test/golden/ast/body-statements.wo create mode 100644 compiler/test/golden/ast/condition-recovery.expected create mode 100644 compiler/test/golden/ast/condition-recovery.wo create mode 100644 compiler/test/golden/ast/ctor-literal.expected create mode 100644 compiler/test/golden/ast/ctor-literal.wo create mode 100644 compiler/test/golden/ast/db-stub-nested.expected create mode 100644 compiler/test/golden/ast/db-stub-nested.wo create mode 100644 compiler/test/golden/ast/db-stub.expected create mode 100644 compiler/test/golden/ast/db-stub.wo create mode 100644 compiler/test/golden/tokens/unknown-char.wo create mode 100644 docs/plan/compiler/nullable-types-implementation.md diff --git a/compiler/src/ast.ml b/compiler/src/ast.ml new file mode 100644 index 0000000..2741a5f --- /dev/null +++ b/compiler/src/ast.ml @@ -0,0 +1,308 @@ +(* ast.ml — AST for `.wo` OOP source (milestone 1). + + Declarations (Task 4 of compiler/plan/2026-08-01-woc-compiler-front.md): + `interface` (method signatures, no bodies), `class` and `type` + (identical field grammar — Task 4 brief's own words: they differ + only in the `is_class` flag, exactly mirroring crates/rt/src/ast.rs's + `TypeDecl { is_class: bool, .. }` design), and `fn` (both as + class/type methods and as free top-level functions — Task 6's brief + mentions "free-fn tables", so free `fn` is real grammar here, not a + rt carry-over). Statements and expressions (Task 5, below) live + inside a method/fn's `body`, which Task 4 captured as a verbatim + token span and Task 5 parses for real. + + Every declaration-shaped node (class/type, interface, method/fn, + field, param) carries a unique `id` (monotonic per parse — Tasks 6/7 + key side tables, e.g. field-kind and ownership-state tables, on these + ids) and a `pos` — the node's own starting source position. This + codebase has no existing notion of a source *range* (Token.t and + diag.ml's `site` are both single points), so `pos` follows that same + single-point convention rather than inventing a new range type. + + `field_ty` and `default_expr` are plain payload, not "declarations" — + they don't get their own `id`/`pos`; nothing downstream needs to key + a side table on "this specific field's type expression" independent + of the field that owns it. + + Task 5 adds real statement/expression ASTs (`stmt`/`expr` below) and + retires the body-as-token-span placeholder: `method_decl.body` is now + `stmt list`, not a verbatim span. Every `stmt` and every `expr` gets + its own `id`/`pos` too — unlike `field_ty`/`default_expr`, Tasks 6/7 + name concrete per-node consumers (every expression gets a type; move + sites, rc sites, and residual sites are individual call-argument and + indexing expressions), so these *are* the "declaration-shaped" case + the paragraph above describes, not the opaque-payload case. + + One disclosed gap in the parent-id-<-child-id invariant Task 4 set up + for the declaration skeleton (a container's id is minted before its + children's): a `stmt`'s id is still minted before its own + sub-expressions/sub-blocks are parsed, so that half holds exactly as + before. It does NOT extend to expr-inside-expr — left-recursive + binary/postfix parsing (`a + b`, `a.b`, `a(b)`) parses the left/base + operand first (smaller id) and only decides to wrap it in + `Binary`/`Field`/`Call` after seeing the next token, so that + wrapper's id is necessarily minted *after* its own child's. Every id + is still unique and monotonic in mint order; only the strict + parent-<-child direction is given up, and only for expr-in-expr + nesting. Forcing it there would mean pre-reserving ids speculatively + before knowing a wrapper is even needed — not worth the complexity + for side-table keys that only need uniqueness, not order. *) + +type pos = { + line : int; + col : int; +} + +(* `ref`/`multi`/`map` are not lexer keywords (Task 3 deliberately + dropped rt's schema-keyword zoo) — they're recognized positionally, + by name, only at the start of a field's type, exactly like rt's own + `insert`/`select` statement-keyword convention. Scalar also covers a + bare class name used as an owned-embed field type (spec section 4: + "Fields hold owned values, ref T ids, ... or @gc references") — + Task 6 decides whether a given Scalar name is a builtin scalar or a + user class. `?T` nullable wrapper (adopted from Haxe, systems-track + spec Part 1) wraps any field_ty; milestone-1 has no array (`[T]`) + or tagged unions — those are rt schema-layer features not named in + this task's grammar, so they're deliberately absent here. *) +type field_ty = + | Scalar of string + | Ref of string + | Multi of string + | Map of string * string (* key type, value type: map *) + | Nullable of field_ty (* ?T wrapper *) + +(* Parameter passing convention (spec section 3, rule 2): default is an + immutable borrow; `mut` is an exclusive borrow; `take` moves + ownership in. The owner pass (Task 7) is the eventual consumer. *) +type param_conv = + | Borrow + | Mut + | Take + +type param = { + id : int; + pos : pos; + name : string; + conv : param_conv; + ty : field_ty; (* declared type; Task 6 resolves it for real *) +} + +(* `= now()` is recognized explicitly (mirrors rt's DefaultExpr::Now). + Everything else is kept as its raw token span rather than eagerly + turned into a string (rt's DefaultExpr::Opaque(String) precedent) — + "opaque token span" per the Task 4 brief, so a later stage could in + principle re-lex/interpret it without having thrown information + away. Task 4 itself never inspects the contents. *) +type default_expr = + | DefaultNow + | DefaultOpaque of Token.t list + +type field = { + id : int; + pos : pos; + name : string; + ty : field_ty; + default : default_expr option; + (* Annotation *names* only (e.g. ["unique"]) — the brief: "names + recorded, unknown names fine at parse level". Any `(...)` argument + list on a field annotation is consumed and discarded, matching + rt's own field-annotation handling; nothing downstream needs those + arguments in Task 4. *) + annotations : string list; +} + +(* ---- expressions (Task 5) ------------------------------------------ + + `unop`/`binop` name their operators the way rt's ast.rs BinOp/UnOp + does for the operators this grammar shares with it (Add/Sub/Mul/ + Div/Mod, Eq/Ne/Lt/Le/Gt/Ge). Two deliberate differences from rt: no + `And`/`Or` (this grammar's lexer, Task 3, has no `&&`/`||` tokens — + there is no boolean-logic sublanguage here) and one new operator, + `Concat`. + + `Concat` (source syntax `..`, `Token.DotDot`) is this task's own + design decision, not a straight rt port: the brief's precedence + ladder lists "text concatenation" as its own tier, distinct from + arithmetic, and the VM design spec's instruction table + (docs/superpowers/specs/2026-08-01-oop-compiler-vm-design.md §5) + lists a dedicated `CONCAT` op alongside (not folded into) `ADD` — + so the source grammar needs its own operator token to compile down + to that, not an overload of `+` disambiguated by operand type later. + `Token.DotDot` is lexed (Task 3) but was never given a grammar rule + before now, so this claims it. Flagged as an inference, not a + spec-literal instruction, because no fixture upstream of this task + spells out the token; it is the only unclaimed binary-shaped token + left, and Lua's `..` is the same design (concat binds looser than + `+`/`-`, tighter than comparison — this ladder's ordering). *) +type unop = Neg + +type binop = + | Add + | Sub + | Mul + | Div + | Mod + | Concat + | Eq + | Ne + | Lt + | Le + | Gt + | Ge + +type expr = { + id : int; + pos : pos; + kind : expr_kind; +} + +(* `Call`'s callee is a general `expr`, not a name: a free call has an + `Ident` callee, a method call (interface-typed or not — dispatch is + Task 6's job, not the parser's) has a `Field` callee. Brief: "free, + method, and interface-typed method calls share one call node" — this + is that sharing; the parser never distinguishes the three, it only + ever builds `Call (callee, args)`. + + `Ctor` (`ClassName { field: expr, ... }`) is recognized in primary- + expression position by the two-token shape identifier-then-brace + (parser.ml's `looks_like_ctor`) — milestone-1 has no `new` keyword, + this brace literal is the only construction syntax. + + `DbStub` is the SQL sublanguage's one opaque node (`insert`/`select`, + lowercase-Ident or uppercase-keyword form alike): its token list is + never re-parsed as this grammar, only captured verbatim, spec + section 3's "parses but traps" contract. Lowercase `insert` is a + *statement*-only trigger (parser.ml's `is_insert_trigger`, checked + before general expression parsing); lowercase `select` is legal in + general expression position too (`is_select_trigger`, reachable from + `parse_primary`) — e.g. `let rows = select ...` — matching the brief: + "statement-position lowercase insert/select (and expression-position + select)". *) +and expr_kind = + | IntLit of int + | StrLit of string + | BoolLit of bool + | Ident of string + | Field of expr * string + | Index of expr * expr + | Call of expr * expr list + | Unary of unop * expr + | Binary of binop * expr * expr + | Ctor of string * (string * expr) list + | DbStub of Token.t list + +(* ---- statements (Task 5) --------------------------------------------- + + Statement nodes use `s_id`/`s_pos`/`s_kind` rather than `id`/`pos`/ + `kind` (which `expr` already claims): both records would otherwise + share the exact same field set, and OCaml's type-directed field + disambiguation needs at least one label difference to tell a bare + `{ id; pos; kind = ... }` literal apart from the other type — every + other id/pos reuse in this file (param/field/method_sig/...) is safe + because each of those already has a distinct full field set. + + `If.else_body` pairs the `else` keyword's own position with its + block so dump.ml has something to print an "ELSE" line's LINE:COL + from — every other dumped line in this codebase starts with a real + position (Task 4 convention), and there is no other node to hang + that position on. `else if ...` is desugared here at parse time into + `else_body = Some (else_pos, [ ])` — a one-statement + else-block whose sole statement is itself an `If` — rather than a + third `else_body` shape, so dump.ml's block-rendering code (already + written once, for `then_body`) renders the chain for free. *) +type stmt = { + s_id : int; + s_pos : pos; + s_kind : stmt_kind; +} + +and stmt_kind = + | Let of { + name : string; + ty : string option; + value : expr; + } + | Assign of { + target : expr; + value : expr; + } + | If of { + cond : expr; + then_body : stmt list; + else_body : (pos * stmt list) option; + } + | While of { + cond : expr; + body : stmt list; + } + | For of { + var : string; + iter : expr; + body : stmt list; + } + | Return of expr option + | ExprStmt of expr + +(* A signature shared shape (name/params/ret) appears twice: as an + interface method (no body) and as a class/type/free-fn method (body + captured as a span). Kept as two separate flat records rather than + one record nesting the other — avoids an awkward `sig` field name + (`sig` is an OCaml keyword) and keeps `m.name`/`m.params` uniform + instead of `m.msig.name`. *) +type method_sig = { + id : int; + pos : pos; + name : string; + params : param list; + ret : field_ty option; +} + +type method_decl = { + id : int; + pos : pos; + name : string; + params : param list; + ret : field_ty option; + (* Task 4 captured this as a verbatim token span (brace-depth counter + only); Task 5 parses it for real. *) + body : stmt list; +} + +(* `@table(name: "...", index: [a, b], index: [c])` — optional storage + configuration, ported from rt's TableCfg. `name`/`index` are the only + known keys; anything else inside `@table(...)` is a parse error + (WO-E1xx), not a silent skip — unlike an unrecognized annotation + *name*, which does skip silently (rt convention, see parser.ml). *) +type table_cfg = { + table_name : string option; + indexes : string list list; +} + +type class_decl = { + id : int; + pos : pos; + name : string; + (* true for `class`, false for `type` — Task 4 brief: identical field + grammar either way; `fn` methods parse for real inside both (a + deliberate divergence from rt's plan-13 asymmetry, where a plain + `type`'s `fn` was skip-discarded — see parser.ml's module doc). *) + is_class : bool; + is_gc : bool; (* @gc — reference semantics, spec section 3/4 *) + table : table_cfg option; (* @table(...) — absent unless annotated *) + fields : field list; + methods : method_decl list; +} + +type interface_decl = { + id : int; + pos : pos; + name : string; + methods : method_sig list; (* signatures only — no fields, no bodies *) +} + +type decl = + | Class of class_decl + | Interface of interface_decl + | Fn of method_decl (* free (non-method) top-level function *) + +type program = { decls : decl list } diff --git a/compiler/src/dump.ml b/compiler/src/dump.ml new file mode 100644 index 0000000..632ccc1 --- /dev/null +++ b/compiler/src/dump.ml @@ -0,0 +1,275 @@ +(* dump.ml — stable text dumps of compiler-internal data. + + Used both by `woc --dump-*` flags (compiler/bin/main.ml) and the + golden-file test runner (compiler/test/runner.ml) that diffs + against compiler/test/golden/. These formats are load-bearing test + contracts, not debug output: once a fixture's `.expected` file is + checked in, renaming a kind's dumped label is a breaking change to + every golden fixture that contains it. Grows one `dump_` + function per task (Task 3 adds dump_tokens; Task 4 adds dump_ast; + Task 6 dump_types; Task 7 dump_owner). + + Token dump format (one line per token): + + LINE:COL KIND + LINE:COL KIND(payload) + + e.g. "3:1 KW_CLASS", "3:7 IDENT(Product)", "4:12 INT(42)", + "4:20 STR(hello)", "5:1 NEWLINE". LINE and COL are 1-based. *) + +(* One stable, upper-snake-case label per Token.kind constructor. + Written as an exhaustive match with no wildcard, on purpose: adding + a Token.kind case without adding it here is a compile error (a + non-exhaustive-match warning promoted to an error by dune's default + build profile), not a silently unlabelled dump line. *) +let kind_label (k : Token.kind) : string = + match k with + | Token.Ident s -> Printf.sprintf "IDENT(%s)" s + | Token.Int n -> Printf.sprintf "INT(%d)" n + | Token.Str s -> Printf.sprintf "STR(%s)" s + | Token.KwType -> "KW_TYPE" + | Token.KwClass -> "KW_CLASS" + | Token.KwInterface -> "KW_INTERFACE" + | Token.KwFn -> "KW_FN" + | Token.KwLet -> "KW_LET" + | Token.KwMut -> "KW_MUT" + | Token.KwTake -> "KW_TAKE" + | Token.KwReturn -> "KW_RETURN" + | Token.KwIf -> "KW_IF" + | Token.KwElse -> "KW_ELSE" + | Token.KwWhile -> "KW_WHILE" + | Token.KwFor -> "KW_FOR" + | Token.KwIn -> "KW_IN" + | Token.KwTrue -> "KW_TRUE" + | Token.KwFalse -> "KW_FALSE" + | Token.KwInsert -> "KW_INSERT" + | Token.KwSelect -> "KW_SELECT" + | Token.LBrace -> "LBRACE" + | Token.RBrace -> "RBRACE" + | Token.LParen -> "LPAREN" + | Token.RParen -> "RPAREN" + | Token.LBracket -> "LBRACKET" + | Token.RBracket -> "RBRACKET" + | Token.Comma -> "COMMA" + | Token.Semicolon -> "SEMICOLON" + | Token.Colon -> "COLON" + | Token.Dot -> "DOT" + | Token.DotDot -> "DOTDOT" + | Token.Question -> "QUESTION" + | Token.At -> "AT" + | Token.Pipe -> "PIPE" + | Token.Arrow -> "ARROW" + | Token.FatArrow -> "FAT_ARROW" + | Token.Dash -> "DASH" + | Token.Plus -> "PLUS" + | Token.Star -> "STAR" + | Token.Slash -> "SLASH" + | Token.Percent -> "PERCENT" + | Token.Eq -> "EQ" + | Token.EqEq -> "EQEQ" + | Token.NotEq -> "NOTEQ" + | Token.Lt -> "LT" + | Token.LtEq -> "LTEQ" + | Token.Gt -> "GT" + | Token.GtEq -> "GTEQ" + | Token.PlusEq -> "PLUSEQ" + | Token.MinusEq -> "MINUSEQ" + | Token.Newline -> "NEWLINE" + | Token.Eof -> "EOF" + +let dump_tokens (toks : Token.t list) : string = + let lines = + List.map + (fun (t : Token.t) -> Printf.sprintf "%d:%d %s" t.line t.col (kind_label t.kind)) + toks + in + match lines with [] -> "" | _ -> String.concat "\n" lines ^ "\n" + +(* dump_ast — stable indented-tree dump of the declaration AST (Task 4; + Task 5 adds real statement/expression rendering under METHOD). + + One node per line: "LINE:COL KIND payload", children indented two + spaces under their parent. Node ids are deliberately never printed + (Task 4 brief: "ids would churn goldens" — they're an internal, + monotonic-per-parse detail Tasks 6/7 key side tables on, not a + stable rendering surface); positions are, since they're what makes + the dump useful as a fixture at all. + + Statements get one line each (dump_stmt), matching how fields/ + methods already get one line each under their class — the natural + "line" granularity for a body, not one dump line per sub-expression. + Expressions render inline as a single unparsed string (expr_str), + the same choice this file already made for field types/defaults/ + signatures (field_ty_str/default_str/sig_str): readable golden files + that look like source, not an exploded parse tree. Expression nodes + still carry their own id/pos in the AST (ast.ml) for Tasks 6/7's + side tables; the dump just doesn't surface them, same as decl ids. *) + +let pos_str (p : Ast.pos) : string = Printf.sprintf "%d:%d" p.line p.col + +let conv_str : Ast.param_conv -> string = function + | Ast.Borrow -> "" + | Ast.Mut -> "mut " + | Ast.Take -> "take " + +let rec field_ty_str : Ast.field_ty -> string = function + | Ast.Scalar s -> s + | Ast.Ref s -> Printf.sprintf "ref %s" s + | Ast.Multi s -> Printf.sprintf "multi %s" s + | Ast.Map (k, v) -> Printf.sprintf "map<%s, %s>" k v + | Ast.Nullable t -> "?" ^ field_ty_str t + +let param_str (p : Ast.param) : string = Printf.sprintf "%s%s: %s" (conv_str p.conv) p.name (field_ty_str p.ty) +let params_str (params : Ast.param list) : string = String.concat ", " (List.map param_str params) + +(* An opaque default's token span is rendered via kind_label, same as + --dump-tokens, rather than a second ad hoc "unparse a token" writer — + one canonical, exhaustive way to turn a Token.kind into stable text. *) +let default_str : Ast.default_expr -> string = function + | Ast.DefaultNow -> " = now()" + | Ast.DefaultOpaque toks -> + " = " ^ String.concat " " (List.map (fun (t : Token.t) -> kind_label t.kind) toks) + +let annotations_str (anns : string list) : string = + String.concat "" (List.map (fun a -> " @" ^ a) anns) + +let dump_field (f : Ast.field) : string = + Printf.sprintf "%s FIELD %s: %s%s%s" (pos_str f.pos) f.name (field_ty_str f.ty) + (match f.default with None -> "" | Some d -> default_str d) + (annotations_str f.annotations) + +let sig_str (name : string) (params : Ast.param list) (ret : Ast.field_ty option) : string = + Printf.sprintf "%s(%s)%s" name (params_str params) + (match ret with None -> "" | Some r -> " -> " ^ field_ty_str r) + +let dump_method_sig (m : Ast.method_sig) : string = + Printf.sprintf "%s METHOD %s" (pos_str m.pos) (sig_str m.name m.params m.ret) + +let binop_str : Ast.binop -> string = function + | Ast.Add -> "+" + | Ast.Sub -> "-" + | Ast.Mul -> "*" + | Ast.Div -> "/" + | Ast.Mod -> "%" + | Ast.Concat -> ".." + | Ast.Eq -> "==" + | Ast.Ne -> "!=" + | Ast.Lt -> "<" + | Ast.Le -> "<=" + | Ast.Gt -> ">" + | Ast.Ge -> ">=" + +(* Raw token span shared by both DbStub renderings below: a statement- + position DbStub (dump_stmt) and an expression-position one nested + inside a LET/ASSIGN/etc. (expr_str). *) +let dbstub_tokens_str (toks : Token.t list) : string = + String.concat " " (List.map (fun (t : Token.t) -> kind_label t.kind) toks) + +(* expr_str — one unparsed line per expression, no positions (see this + file's module doc for why: same "inline text, not an exploded tree" + choice as field_ty_str/default_str). Recurses structurally; no + precedence-driven parenthesization since every expr_str call site + here only ever needs "readable enough to eyeball in a golden file", + not a round-trippable unparser. *) +let rec expr_str (e : Ast.expr) : string = + match e.Ast.kind with + | Ast.IntLit n -> string_of_int n + | Ast.StrLit s -> "\"" ^ s ^ "\"" + | Ast.BoolLit b -> if b then "true" else "false" + | Ast.Ident s -> s + | Ast.Field (base, name) -> expr_str base ^ "." ^ name + | Ast.Index (base, idx) -> expr_str base ^ "[" ^ expr_str idx ^ "]" + | Ast.Call (callee, args) -> + Printf.sprintf "%s(%s)" (expr_str callee) (String.concat ", " (List.map expr_str args)) + | Ast.Unary (Ast.Neg, operand) -> "-" ^ expr_str operand + | Ast.Binary (op, l, r) -> Printf.sprintf "%s %s %s" (expr_str l) (binop_str op) (expr_str r) + | Ast.Ctor (name, fields) -> + Printf.sprintf "%s { %s }" name + (String.concat ", " + (List.map (fun (fname, fval) -> Printf.sprintf "%s: %s" fname (expr_str fval)) fields)) + | Ast.DbStub toks -> Printf.sprintf "DB_STUB(%s)" (dbstub_tokens_str toks) + +(* dump_stmt — one line per statement (LINE:COL KIND detail), matching + dump_field/dump_method_sig's "one descriptive line" convention; + block-having statements (IF/WHILE/FOR) get their body's statements + as children, indented two spaces, same nesting rule as + class/interface members. A bare DbStub expression-statement (the + `insert`/`select` sublanguage — see ast.ml/parser.ml) is special- + cased to a standalone "DB_STUB ..." line rather than "EXPR + DB_STUB(...)", so it reads as its own concept, not a generic + expression statement that happens to contain one. *) +let rec dump_stmt (s : Ast.stmt) : string list = + let indent_block (body : Ast.stmt list) : string list = + List.concat_map (fun st -> List.map (fun l -> " " ^ l) (dump_stmt st)) body + in + match s.Ast.s_kind with + | Ast.Let { name; ty; value } -> + let ty_part = match ty with None -> "" | Some t -> ": " ^ t in + [ Printf.sprintf "%s LET %s%s = %s" (pos_str s.Ast.s_pos) name ty_part (expr_str value) ] + | Ast.Assign { target; value } -> + [ Printf.sprintf "%s ASSIGN %s = %s" (pos_str s.Ast.s_pos) (expr_str target) (expr_str value) ] + | Ast.If { cond; then_body; else_body } -> + let header = Printf.sprintf "%s IF %s" (pos_str s.Ast.s_pos) (expr_str cond) in + let else_lines = + match else_body with + | None -> [] + | Some (else_pos, body) -> Printf.sprintf "%s ELSE" (pos_str else_pos) :: indent_block body + in + (header :: indent_block then_body) @ else_lines + | Ast.While { cond; body } -> + Printf.sprintf "%s WHILE %s" (pos_str s.Ast.s_pos) (expr_str cond) :: indent_block body + | Ast.For { var; iter; body } -> + Printf.sprintf "%s FOR %s IN %s" (pos_str s.Ast.s_pos) var (expr_str iter) :: indent_block body + | Ast.Return None -> [ Printf.sprintf "%s RETURN" (pos_str s.Ast.s_pos) ] + | Ast.Return (Some e) -> [ Printf.sprintf "%s RETURN %s" (pos_str s.Ast.s_pos) (expr_str e) ] + | Ast.ExprStmt { Ast.kind = Ast.DbStub toks; _ } -> + [ Printf.sprintf "%s DB_STUB %s" (pos_str s.Ast.s_pos) (dbstub_tokens_str toks) ] + | Ast.ExprStmt e -> [ Printf.sprintf "%s EXPR %s" (pos_str s.Ast.s_pos) (expr_str e) ] + +(* [header; body statements...] — body lines are already indented two + spaces; callers nesting this under a class add one more level of + indent uniformly, same as before Task 5. *) +let dump_method (m : Ast.method_decl) : string list = + let header = Printf.sprintf "%s METHOD %s" (pos_str m.pos) (sig_str m.name m.params m.ret) in + let body_lines = List.concat_map (fun s -> List.map (fun l -> " " ^ l) (dump_stmt s)) m.body in + header :: body_lines + +let annotations_header (is_gc : bool) (table : Ast.table_cfg option) : string = + let gc_part = if is_gc then " @gc" else "" in + let table_part = + match table with + | None -> "" + | Some t -> + let name_part = + match t.table_name with None -> [] | Some n -> [ Printf.sprintf "name=%S" n ] + in + let index_parts = + List.map (fun cols -> Printf.sprintf "index=[%s]" (String.concat ", " cols)) t.indexes + in + let parts = name_part @ index_parts in + if parts = [] then " @table" else " @table(" ^ String.concat ", " parts ^ ")" + in + gc_part ^ table_part + +let dump_class (c : Ast.class_decl) : string list = + let kw = if c.is_class then "CLASS" else "TYPE" in + let header = + Printf.sprintf "%s %s %s%s" (pos_str c.pos) kw c.name (annotations_header c.is_gc c.table) + in + let field_lines = List.map (fun f -> " " ^ dump_field f) c.fields in + let method_lines = List.concat_map (fun m -> List.map (fun l -> " " ^ l) (dump_method m)) c.methods in + (header :: field_lines) @ method_lines + +let dump_interface (i : Ast.interface_decl) : string list = + let header = Printf.sprintf "%s INTERFACE %s" (pos_str i.pos) i.name in + let method_lines = List.map (fun m -> " " ^ dump_method_sig m) i.methods in + header :: method_lines + +let dump_decl : Ast.decl -> string list = function + | Ast.Class c -> dump_class c + | Ast.Interface i -> dump_interface i + | Ast.Fn f -> dump_method f + +let dump_ast (prog : Ast.program) : string = + let lines = List.concat_map dump_decl prog.decls in + match lines with [] -> "" | _ -> String.concat "\n" lines ^ "\n" diff --git a/compiler/src/dune b/compiler/src/dune new file mode 100644 index 0000000..481c8dc --- /dev/null +++ b/compiler/src/dune @@ -0,0 +1,5 @@ +; compiler/src/dune — woc library (Task 2 onward adds modules per plan) +; OCaml stdlib only: no Menhir, no ppx, no opam libraries. +(library + (name woc_lib) + (modules diag token ast lexer parser types dump)) \ No newline at end of file diff --git a/compiler/src/parser.ml b/compiler/src/parser.ml new file mode 100644 index 0000000..b4b7e14 --- /dev/null +++ b/compiler/src/parser.ml @@ -0,0 +1,1155 @@ +(* parser.ml — declaration-level recursive-descent parser for `.wo` + OOP source (Task 4 of compiler/plan/2026-08-01-woc-compiler-front.md). + + Ported from crates/rt/src/parser.rs's structure and conventions + wherever they still apply, with two deliberate divergences forced by + this task's own contract: + + 1. **Multi-error recovery, not first-error-stop.** rt's parser + returns `anyhow::Result` and bails the *entire* parse on the + first problem. This front end's diagnostics contract (diag.ml, + since Task 2) is multi-error, and the Task 4 brief requires it + explicitly: "a broken declaration syncs to the next top-level + keyword and parsing continues, so one bad class yields one + diagnostic, not a cascade." Every low-level expectation failure + ([fail]/[unexpected]) reports exactly one diagnostic through the + given collector, then raises [Parse_error] — an exception used + purely for unwinding control flow back to [parse_program]'s + top-level loop, never exposed outside this module. Recovery + granularity is the whole enclosing top-level declaration: any + failure anywhere inside a class/type/interface/fn body discards + that entire declaration and resyncs, matching "one bad class + yields one diagnostic" literally (not just "one bad field"). + + 2. **`ref`/`multi`/`map` are not lexer keywords here.** rt's + token.rs carries KwRef/KwMulti as real keywords; Task 3's lexer + deliberately dropped that whole schema-keyword zoo (see + token.ml's module doc). So field-type recognition matches on + `Token.Ident "ref"` / `"multi"` / `"map"` by name, exactly the + same positional-keyword trick rt itself uses for `insert`/ + `select` in method bodies (crates/rt/src/parser.rs parse_stmt). + Same story for `service`/`policy`/`on`, which drive the + skip-on-block behavior below. + + Kept faithfully from rt: newline-significant skipping + ([skip_newlines]); the `looks_like_field` two-token lookahead + (Ident then Colon) that disambiguates a field line from a + service/policy/on line; `@table`'s known-keys parsing including its + exact error shape (unknown key, duplicate name, non-string name, + empty index list); `skip_block_line`'s and `skip_on_block`'s + brace/paren/bracket-depth counters — the on-block one in particular + is the load-bearing piece the Task 4 brief calls out by name: object + literals like `{ article_id: self.id }` inside an `on` block must + not be mistaken for the close of the enclosing class/type body. + + New relative to rt (this task's own grammar, not a port): `interface` + (signatures only, no bodies); `@gc`; parameter conventions (bare = + borrow, `mut`, `take` — `KwMut`/`KwTake` are real Task-3 keywords); + free top-level `fn` (Task 6's brief mentions "free-fn tables", so + this is real grammar, not a rt carry-over); and capturing a method + body as a verbatim token span rather than parsing or discarding it + (Task 5 parses it for real). + + [Dump.kind_label] is reused for the "got X" half of syntax-error + messages rather than writing a second exhaustive match over + `Token.kind` here — dump.ml's match already has to stay exhaustive + (a new token kind is a compile error there), so reusing it avoids a + second copy of the same maintenance burden. This makes parser.ml + depend on dump.ml, which otherwise only renders finished output; + there's no cycle (dump.ml depends on token.ml/ast.ml only), but it's + a deliberate, slightly unusual edge worth flagging. *) + +exception Parse_error + +type state = { + toks : Token.t array; + len : int; + file : string; + collector : Diag.Collector.t; + mutable pos : int; + mutable next_id : int; + (* True while parsing an if/while condition or a for-loop's iterable + expression — Task 5's fix for the identifier-then-brace ambiguity + a bare condition shares with constructor literals (`if active { ... + }`: is `active` the whole condition, or the start of `active { ... + }` as a constructor literal swallowing the if's own block?). See + [looks_like_ctor] and [parse_expr_no_brace]. Reset to `false` while + inside a parenthesized sub-expression (parens make the boundary + unambiguous again, e.g. `if (Widget { x: 1 }).ok { ... }`). *) + mutable no_brace : bool; +} + +let make (collector : Diag.Collector.t) ~(file : string) (toks : Token.t list) : state = + let arr = Array.of_list toks in + { toks = arr; len = Array.length arr; file; collector; pos = 0; next_id = 1; no_brace = false } + +let fresh_id (st : state) : int = + let id = st.next_id in + st.next_id <- id + 1; + id + +(* Every token list ends with Token.Eof (lexer.ml's tokenize contract), + so len >= 1 always and this clamp never reads past the real array. *) +let cur (st : state) : Token.t = st.toks.(min st.pos (st.len - 1)) +let peek (st : state) : Token.kind = (cur st).kind + +let peek_pos (st : state) : Ast.pos = + let t = cur st in + { Ast.line = t.line; col = t.col } + +let tok_at (st : state) (i : int) : Token.t = st.toks.(min (max i 0) (st.len - 1)) + +let at_end (st : state) : bool = match peek st with Token.Eof -> true | _ -> false + +(* Never advances past Eof — mirrors rt's `advance` guard. *) +let advance (st : state) : Token.t = + let t = cur st in + if not (at_end st) then st.pos <- st.pos + 1; + t + +let skip_newlines (st : state) : unit = + let continue_ = ref true in + while !continue_ do + match peek st with + | Token.Newline -> ignore (advance st) + | _ -> continue_ := false + done + +(* Reports one diagnostic at [site] and unwinds to the nearest + recovery point via [Parse_error]. Polymorphic return type: callers + use it in both `unit` context (expect) and `string` context + (expect_ident), since a `raise` never actually returns. *) +let fail (st : state) (site : Ast.pos) (code : string) (message : string) : 'a = + Diag.Collector.add st.collector + (Diag.error ~code ~file:st.file ~line:site.Ast.line ~col:site.Ast.col ~message ()); + raise Parse_error + +let syntax_code = Diag.parsing_prefix ^ "01" (* WO-E101: generic syntax error *) +let table_code = Diag.parsing_prefix ^ "02" (* WO-E102: invalid @table(...) configuration *) + +let unexpected (st : state) (what : string) : 'a = + let p = peek_pos st in + fail st p syntax_code + (Printf.sprintf "expected %s, got %s" what (Dump.kind_label (peek st))) + +let expect (st : state) (want : Token.kind) (what : string) : unit = + if peek st = want then ignore (advance st) else unexpected st what + +let accept (st : state) (want : Token.kind) : bool = + if peek st = want then begin + ignore (advance st); + true + end + else false + +let expect_ident (st : state) (what : string) : string = + match peek st with + | Token.Ident s -> + ignore (advance st); + s + | _ -> unexpected st what + +(* ---- field/service/policy/on disambiguation -------------------------- + + Lookahead only: Ident immediately followed by Colon. Same shape rt + uses (crates/rt/src/parser.rs looks_like_field) to tell a field line + apart from a service/policy/on line without a reserved keyword. *) +let looks_like_field (st : state) : bool = + match peek st with + | Token.Ident _ -> (tok_at st (st.pos + 1)).kind = Token.Colon + | _ -> false + +let is_sync_ident (s : string) : bool = + match s with + | "service" | "policy" | "on" -> true + | _ -> false + +(* ---- declaration-level recovery --------------------------------------- + + Ported from rt's skip_top_level_chunk, extended with this grammar's + extra top-level starters (interface, fn, service/policy/on as + idents, @). Balances braces/brackets/parens while scanning; an + RBrace that brings depth to <=0 stops the scan right after it (this + specifically handles "we were skipping a broken class/type/interface + body — its own closing brace ends the skip", distinct from + RBracket/RParen which only ever decrement depth and never stop the + scan by themselves — ported exactly as rt has it, not generalized, + because that asymmetry is deliberate there too). The `st.pos > start` + guard on every sync-keyword arm guarantees forward progress even + when the parser is already sitting on a sync token when recovery + begins. *) +let sync_to_next_top_level (st : state) : unit = + let depth = ref 0 in + let start = st.pos in + let continue_ = ref true in + while !continue_ do + match peek st with + | Token.Eof -> continue_ := false + | Token.LBrace -> + incr depth; + ignore (advance st) + | Token.RBrace -> + decr depth; + ignore (advance st); + if !depth <= 0 then continue_ := false + | Token.LBracket -> + incr depth; + ignore (advance st) + | Token.RBracket -> + decr depth; + ignore (advance st) + | Token.LParen -> + incr depth; + ignore (advance st) + | Token.RParen -> + decr depth; + ignore (advance st) + | Token.KwType when !depth = 0 && st.pos > start -> continue_ := false + | Token.KwClass when !depth = 0 && st.pos > start -> continue_ := false + | Token.KwInterface when !depth = 0 && st.pos > start -> continue_ := false + | Token.KwFn when !depth = 0 && st.pos > start -> continue_ := false + | Token.At when !depth = 0 && st.pos > start -> continue_ := false + | Token.Ident s when !depth = 0 && st.pos > start && is_sync_ident s -> continue_ := false + | _ -> ignore (advance st) + done + +(* ---- type-level annotations (@gc, @table) ------------------------------ + + Ported from rt's parse_type_annotations. `table` is the only + annotation with structured, checked arguments; `gc` is bare (no + arguments expected, but a stray `(...)` is tolerated rather than + rejected — Task 4's brief only requires @table's keys to be + strict); any other annotation name is unknown and skips silently + (its own precedent, then this rt precedent) after consuming an + optional `(...)` argument block without interpreting it. *) + +let skip_paren_args (st : state) : unit = + if peek st = Token.LParen then begin + ignore (advance st); + let depth = ref 1 in + while !depth > 0 && not (at_end st) do + match peek st with + | Token.LParen -> + incr depth; + ignore (advance st) + | Token.RParen -> + decr depth; + ignore (advance st) + | _ -> ignore (advance st) + done + end + +let parse_table_cfg (st : state) : Ast.table_cfg = + let cfg = ref { Ast.table_name = None; indexes = [] } in + if accept st Token.LParen then begin + let continue_ = ref true in + while !continue_ do + skip_newlines st; + if accept st Token.RParen then continue_ := false + else begin + let key = expect_ident st "@table argument" in + expect st Token.Colon "':'"; + (match key with + | "name" -> + if !cfg.Ast.table_name <> None then + fail st (peek_pos st) table_code "@table(name: ...) given twice"; + (match peek st with + | Token.Str s -> + ignore (advance st); + cfg := { !cfg with Ast.table_name = Some s } + | _ -> unexpected st "a string for @table name") + | "index" -> + expect st Token.LBracket "'['"; + let cols = ref [] in + let more = ref true in + while !more do + cols := expect_ident st "index column" :: !cols; + if not (accept st Token.Comma) then more := false + done; + expect st Token.RBracket "']'"; + if !cols = [] then + fail st (peek_pos st) table_code "@table index needs at least one column"; + cfg := { !cfg with Ast.indexes = !cfg.Ast.indexes @ [ List.rev !cols ] } + | other -> + fail st (peek_pos st) table_code + (Printf.sprintf "unknown @table argument `%s` (supported: name, index)" other)); + skip_newlines st; + if not (accept st Token.Comma) then begin + skip_newlines st; + expect st Token.RParen "')' or ','"; + continue_ := false + end + end + done + end; + !cfg + +type type_annotations = { + is_gc : bool; + table : Ast.table_cfg option; +} + +let no_annotations = { is_gc = false; table = None } + +let parse_type_annotations (st : state) : type_annotations = + let is_gc = ref false in + let table = ref None in + while peek st = Token.At do + ignore (advance st); + let name = expect_ident st "annotation name" in + (match name with + | "gc" -> + is_gc := true; + skip_paren_args st + | "table" -> table := Some (parse_table_cfg st) + | _ -> skip_paren_args st); + skip_newlines st + done; + { is_gc = !is_gc; table = !table } + +(* ---- field parsing ------------------------------------------------------ *) + +let parse_field_ty (st : state) : Ast.field_ty = + let nullable = ref false in + if accept st Token.Question then nullable := true; + let base_ty = + match peek st with + | Token.Ident "ref" -> + ignore (advance st); + Ast.Ref (expect_ident st "ref target type") + | Token.Ident "multi" -> + ignore (advance st); + Ast.Multi (expect_ident st "multi target type") + | Token.Ident "map" -> + ignore (advance st); + expect st Token.Lt "'<'"; + let k = expect_ident st "map key type" in + expect st Token.Comma "','"; + let v = expect_ident st "map value type" in + expect st Token.Gt "'>'"; + Ast.Map (k, v) + | Token.Ident name -> + ignore (advance st); + Ast.Scalar name + | _ -> unexpected st "a field type" + in + if !nullable then Ast.Nullable base_ty else base_ty + +(* Collects the raw tokens of a default expression up to (not + including) a Newline/Comma/RBrace/Eof at depth 0 — the "opaque token + span" the Task 4 brief asks for, mirroring rt's balanced-slurp loop + in parse_default_expr but keeping tokens instead of flattening to a + string. *) +let collect_default_tokens (st : state) : Token.t list = + let buf = ref [] in + let depth = ref 0 in + let continue_ = ref true in + while !continue_ do + match peek st with + | Token.Eof -> continue_ := false + | Token.Newline when !depth = 0 -> continue_ := false + | Token.Comma when !depth = 0 -> continue_ := false + | Token.RBrace when !depth = 0 -> continue_ := false + | Token.LBrace | Token.LBracket | Token.LParen -> + incr depth; + buf := advance st :: !buf + | Token.RBrace | Token.RBracket | Token.RParen -> + decr depth; + buf := advance st :: !buf + | _ -> buf := advance st :: !buf + done; + List.rev !buf + +(* `now` / `now()` is recognized explicitly only when it is the WHOLE + default expression (immediately followed by a field/expr terminator) + — pure lookahead, no speculative advance-then-rewind, unlike rt's + version (which mutates its cursor and restores it on failure). `now` + followed by more tokens (`now + 5`, a field literally typed `now` + followed by something else) falls through to the opaque path, + exactly like rt. *) +let parse_default_expr (st : state) : Ast.default_expr = + let ends_expr (k : Token.kind) : bool = + match k with Token.Newline | Token.Comma | Token.RBrace | Token.Eof -> true | _ -> false + in + let looks_like_bare_now = + match peek st with + | Token.Ident "now" -> ( + match (tok_at st (st.pos + 1)).kind with + | Token.LParen -> ( + match (tok_at st (st.pos + 2)).kind with + | Token.RParen -> ends_expr (tok_at st (st.pos + 3)).kind + | _ -> false) + | k -> ends_expr k) + | _ -> false + in + if looks_like_bare_now then begin + ignore (advance st); + (* now *) + if peek st = Token.LParen then begin + ignore (advance st); + (* ( *) + ignore (advance st) (* ) *) + end; + Ast.DefaultNow + end + else Ast.DefaultOpaque (collect_default_tokens st) + +let parse_field (st : state) : Ast.field = + let pos = peek_pos st in + let name = expect_ident st "field name" in + expect st Token.Colon "':'"; + let ty = parse_field_ty st in + let default = ref None in + let annotations = ref [] in + let continue_ = ref true in + while !continue_ do + match peek st with + | Token.At -> + ignore (advance st); + let ann_name = expect_ident st "annotation name" in + skip_paren_args st; + annotations := ann_name :: !annotations + | Token.Eq -> + ignore (advance st); + default := Some (parse_default_expr st) + | Token.Newline | Token.RBrace | Token.Eof -> continue_ := false + | _ -> unexpected st "an annotation, '=', or end of field" + done; + { Ast.id = fresh_id st; pos; name; ty; default = !default; annotations = List.rev !annotations } + +(* ---- param / signature parsing ------------------------------------------ *) + +let parse_param (st : state) : Ast.param = + let pos = peek_pos st in + let conv = + if accept st Token.KwMut then Ast.Mut + else if accept st Token.KwTake then Ast.Take + else Ast.Borrow + in + let name = expect_ident st "parameter name" in + expect st Token.Colon "':'"; + let ty = parse_field_ty st in + { Ast.id = fresh_id st; pos; name; conv; ty } + +let parse_params (st : state) : Ast.param list = + expect st Token.LParen "'('"; + skip_newlines st; + let params = ref [] in + let continue_ = ref (peek st <> Token.RParen) in + while !continue_ do + params := parse_param st :: !params; + skip_newlines st; + if accept st Token.Comma then skip_newlines st else continue_ := false + done; + expect st Token.RParen "')'"; + List.rev !params + +let parse_ret_type (st : state) : Ast.field_ty option = + let nullable = ref false in + if accept st Token.Question then nullable := true; + if accept st Token.Arrow then begin + let ty = parse_field_ty st in + Some (if !nullable then Ast.Nullable ty else ty) + end else None + +type sig_head = { + s_id : int; + s_pos : Ast.pos; + s_name : string; + s_params : Ast.param list; + s_ret : Ast.field_ty option; +} + +let parse_sig_head (st : state) : sig_head = + let pos = peek_pos st in + expect st Token.KwFn "`fn`"; + let name = expect_ident st "function/method name" in + (* Id assigned here, before params are parsed (each of which mints its + own id) — mirrors class_decl/interface_decl, whose id is likewise + minted right after their head is confirmed and before their body + is parsed. Keeps "parent id < every child id" true uniformly across + every container node, not just some of them. *) + let id = fresh_id st in + let params = parse_params st in + let ret = parse_ret_type st in + { s_id = id; s_pos = pos; s_name = name; s_params = params; s_ret = ret } + +(* ---- statement/expression parser (Task 5) ------------------------------ + + Method bodies were a verbatim token span through Task 4; this parses + them for real. Structured as one large mutually-recursive group + (parse_stmt / parse_block / the if/while/for statement parsers / the + whole expression precedence ladder) because every level of the + ladder ultimately calls back into parse_expr (call args, constructor- + literal field values, parenthesized sub-expressions), and every + block-having statement calls parse_block, which calls parse_stmt. + + Precedence ladder, loosest to tightest (parse_expr is the entry + point; each level's loop is left-associative): + + comparison == != < <= > >= + concat .. + additive + - + multiplicative * / % + unary minus -x + postfix a.b a(b) a[b] + primary literals, idents, `(expr)`, constructor literals, + select-as-expression (DbStub) + + This ordering matches Lua's (concat binds looser than +/-, tighter + than comparison) — see ast.ml's module doc for why `..`/Concat is + this task's own addition, not a straight rt port. + + End-of-statement convention mirrors parse_field's: a "simple" + statement (let/assign/return/expr-statement/DbStub) must end at an + optional `;` followed by a real terminator (Newline, the enclosing + block's `}}`, or Eof) — [end_of_stmt] enforces this, so e.g. + `let x = 1 let y = 2` on one line is a syntax error, not silently + accepted, exactly like two fields can't share a line. Block-having + statements (if/while/for) do NOT call it: their own closing `}` ends + them, and requiring a terminator after it would wrongly reject + `} else {` on one line, the normal style for chained if/else. *) + +let end_of_stmt (st : state) : unit = + ignore (accept st Token.Semicolon); + match peek st with + | Token.Newline -> ignore (advance st) + | Token.RBrace | Token.Eof -> () + | _ -> unexpected st "end of statement (newline or ';')" + +(* ---- statement-level recovery ------------------------------------------- + + One bad statement must yield one diagnostic, not stall or cascade + (brief: "statement-level recovery syncs at newlines/semicolons"). + Scoped strictly to the enclosing block: unlike sync_to_next_top_level + (which consumes the depth-0 RBrace it lands on, because a broken + top-level declaration's own close is what it's abandoning), this + leaves a depth-0 RBrace unconsumed — exactly skip_block_line's + convention — so [parse_block]'s own loop sees it and ends the block + normally instead of the recovery accidentally eating the block's + close and stalling parse_block forever. + + Known gap, deliberately not chased: if the failure happens while + parsing an if/while/for's *condition* (before its `{ body }` is ever + reached), this can stop at a depth-0 `;`/newline that precedes that + still-unconsumed block, leaving a bare `{ ... }` for the next loop + iteration to choke on as a second, cascading diagnostic. Every + fixture this task ships avoids that shape (its bad statements are + plain let/return lines with no trailing block); a fully general fix + would need the sync scan to know a block is still pending, which + isn't worth the complexity this task's brief doesn't ask for. *) +let sync_to_next_stmt (st : state) : unit = + let depth = ref 0 in + let continue_ = ref true in + while !continue_ do + match peek st with + | Token.Eof -> continue_ := false + | Token.RBrace when !depth = 0 -> continue_ := false + | Token.Newline when !depth = 0 -> + ignore (advance st); + continue_ := false + | Token.Semicolon when !depth = 0 -> + ignore (advance st); + continue_ := false + | Token.LBrace | Token.LBracket | Token.LParen -> + incr depth; + ignore (advance st) + | Token.RBrace | Token.RBracket | Token.RParen -> + decr depth; + ignore (advance st) + | _ -> ignore (advance st) + done + +(* ---- the SQL sublanguage: insert/select as one opaque DbStub node ------- + + Spec section 3, "parses but traps": `insert`/`select` (lowercase + Ident, positionally recognized — the same trick as rt's own + `insert`/`select`, per the keyword-discipline note this task's brief + opens with — or the uppercase KwInsert/KwSelect keyword tokens Task 3 + already lexes) are never re-parsed as this grammar. Every token from + the trigger itself through the statement's own terminator is + captured verbatim into one Ast.DbStub node — mirrors skip_block_line's + depth-aware terminator rules (so a brace-enclosed predicate/field + list spanning its own newlines is still captured whole) but collects + tokens instead of discarding them. + + `insert` is a statement-only trigger (checked in parse_stmt, never + reachable from parse_primary): it produces no value, so `let x = + insert ...` must not parse. `select` is legal in general expression + position too (checked in parse_primary), matching the brief's + asymmetry: "statement-position lowercase insert/select (and + expression-position select)". *) +let is_insert_trigger (k : Token.kind) : bool = + match k with Token.KwInsert | Token.Ident "insert" -> true | _ -> false + +let is_select_trigger (k : Token.kind) : bool = + match k with Token.KwSelect | Token.Ident "select" -> true | _ -> false + +(* Fix round 1/1 finding (CRITICAL 2): the original version only had a + depth-0 stop guard on RBrace, so a `select`/`insert` nested inside an + enclosing call or index expression (`wrap(select Foo { x > 1 })`) + had no way to stop at that call's own `)` — it kept decrementing + depth *below* zero and swallowing everything to Eof. RParen/RBracket + now get the exact same "stop, don't consume, leave it for the + enclosing construct" treatment RBrace already had. Comma also stops + at depth 0 for the same reason (a comma-separated call argument or + constructor-literal field: `wrap(select Foo { x > 1 }, 5)`), mirroring + collect_default_tokens' own convention above. *) +let collect_dbstub_tokens (st : state) : Token.t list = + let buf = ref [] in + let depth = ref 0 in + let continue_ = ref true in + while !continue_ do + match peek st with + | Token.Eof -> continue_ := false + | Token.Newline when !depth = 0 -> continue_ := false + | Token.Semicolon when !depth = 0 -> continue_ := false + | Token.Comma when !depth = 0 -> continue_ := false + | Token.RBrace when !depth = 0 -> continue_ := false + | Token.RParen when !depth = 0 -> continue_ := false + | Token.RBracket when !depth = 0 -> continue_ := false + | Token.LBrace | Token.LBracket | Token.LParen -> + incr depth; + buf := advance st :: !buf + | Token.RBrace | Token.RBracket | Token.RParen -> + decr depth; + buf := advance st :: !buf + | _ -> buf := advance st :: !buf + done; + List.rev !buf + +let parse_dbstub_expr (st : state) : Ast.expr = + let pos = peek_pos st in + let id = fresh_id st in + let toks = collect_dbstub_tokens st in + { Ast.id; pos; kind = Ast.DbStub toks } + +(* Exception-safe save/restore of state.no_brace (Fix round 1/1, + CRITICAL 1). Every no_brace toggle below goes through this, using + Fun.protect so the restore runs even when [f] raises Parse_error — + the normal statement-recovery path. The original code restored only + on a normal return (`st.no_brace <- v; let e = f () in st.no_brace <- + saved; e`), so a failure anywhere inside a no_brace-toggled region + (a broken if/while/for condition, or a broken expression nested + inside one) left state.no_brace permanently stuck, corrupting + constructor-literal recognition for the rest of the file — not just + the rest of the current statement, since state.no_brace lives on the + shared parser state, not a stack frame that unwinds with the + exception. + + Also used (value = false) for call-argument lists and index + expressions (CRITICAL 3): a `(`/`[` is a fresh nesting context whose + own closing `)`/`]` unambiguously ends it, so a constructor literal + inside one is never ambiguous with an enclosing if/while/for's block + — exactly like the parenthesized-primary case, generalized to the + other two "this token pair brackets a fresh sub-expression" shapes. *) +let with_no_brace (st : state) (value : bool) (f : unit -> 'a) : 'a = + let saved = st.no_brace in + st.no_brace <- value; + Fun.protect ~finally:(fun () -> st.no_brace <- saved) f + +(* ---- expression parsing -------------------------------------------------- *) + +let rec parse_expr (st : state) : Ast.expr = parse_comparison st + +and parse_comparison (st : state) : Ast.expr = + let lhs = ref (parse_concat st) in + let continue_ = ref true in + while !continue_ do + match peek st with + | (Token.EqEq | Token.NotEq | Token.Lt | Token.LtEq | Token.Gt | Token.GtEq) as k -> + let pos = peek_pos st in + let op = + match k with + | Token.EqEq -> Ast.Eq + | Token.NotEq -> Ast.Ne + | Token.Lt -> Ast.Lt + | Token.LtEq -> Ast.Le + | Token.Gt -> Ast.Gt + | _ -> Ast.Ge + in + let id = fresh_id st in + ignore (advance st); + let rhs = parse_concat st in + lhs := { Ast.id; pos; kind = Ast.Binary (op, !lhs, rhs) } + | _ -> continue_ := false + done; + !lhs + +and parse_concat (st : state) : Ast.expr = + let lhs = ref (parse_additive st) in + let continue_ = ref true in + while !continue_ do + match peek st with + | Token.DotDot -> + let pos = peek_pos st in + let id = fresh_id st in + ignore (advance st); + let rhs = parse_additive st in + lhs := { Ast.id; pos; kind = Ast.Binary (Ast.Concat, !lhs, rhs) } + | _ -> continue_ := false + done; + !lhs + +and parse_additive (st : state) : Ast.expr = + let lhs = ref (parse_multiplicative st) in + let continue_ = ref true in + while !continue_ do + match peek st with + | (Token.Plus | Token.Dash) as k -> + let pos = peek_pos st in + let op = if k = Token.Plus then Ast.Add else Ast.Sub in + let id = fresh_id st in + ignore (advance st); + let rhs = parse_multiplicative st in + lhs := { Ast.id; pos; kind = Ast.Binary (op, !lhs, rhs) } + | _ -> continue_ := false + done; + !lhs + +and parse_multiplicative (st : state) : Ast.expr = + let lhs = ref (parse_unary st) in + let continue_ = ref true in + while !continue_ do + match peek st with + | (Token.Star | Token.Slash | Token.Percent) as k -> + let pos = peek_pos st in + let op = match k with Token.Star -> Ast.Mul | Token.Slash -> Ast.Div | _ -> Ast.Mod in + let id = fresh_id st in + ignore (advance st); + let rhs = parse_unary st in + lhs := { Ast.id; pos; kind = Ast.Binary (op, !lhs, rhs) } + | _ -> continue_ := false + done; + !lhs + +and parse_unary (st : state) : Ast.expr = + match peek st with + | Token.Dash -> + let pos = peek_pos st in + let id = fresh_id st in + ignore (advance st); + let operand = parse_unary st in + { Ast.id; pos; kind = Ast.Unary (Ast.Neg, operand) } + | _ -> parse_postfix st + +and parse_postfix (st : state) : Ast.expr = + let base = ref (parse_primary st) in + let continue_ = ref true in + while !continue_ do + match peek st with + | Token.Dot -> + let pos = peek_pos st in + ignore (advance st); + let name = expect_ident st "field or method name" in + base := { Ast.id = fresh_id st; pos; kind = Ast.Field (!base, name) } + | Token.LParen -> + let pos = peek_pos st in + let args = parse_call_args st in + base := { Ast.id = fresh_id st; pos; kind = Ast.Call (!base, args) } + | Token.LBracket -> + let pos = peek_pos st in + ignore (advance st); + let idx = with_no_brace st false (fun () -> parse_expr st) in + expect st Token.RBracket "']'"; + base := { Ast.id = fresh_id st; pos; kind = Ast.Index (!base, idx) } + | _ -> continue_ := false + done; + !base + +and parse_call_args (st : state) : Ast.expr list = + expect st Token.LParen "'('"; + with_no_brace st false (fun () -> + skip_newlines st; + let args = ref [] in + let continue_ = ref (peek st <> Token.RParen) in + while !continue_ do + args := parse_expr st :: !args; + skip_newlines st; + if accept st Token.Comma then skip_newlines st else continue_ := false + done; + expect st Token.RParen "')'"; + List.rev !args) + +(* Two-token lookahead, exactly like rt's own select-expression trick + (this brief's own words): a bare Ident immediately followed by `{` + in expression position is a constructor literal, UNLESS we're + parsing an if/while condition or a for-loop's iterable expression + (state.no_brace), where that same `{` is the statement's own + required block, not a literal's opening brace. *) +and looks_like_ctor (st : state) : bool = + (not st.no_brace) + && (match peek st with Token.Ident _ -> true | _ -> false) + && (tok_at st (st.pos + 1)).kind = Token.LBrace + +and parse_ctor_literal (st : state) : Ast.expr = + let pos = peek_pos st in + let id = fresh_id st in + let name = expect_ident st "constructor class name" in + expect st Token.LBrace "'{'"; + skip_newlines st; + let fields = ref [] in + let continue_ = ref (peek st <> Token.RBrace) in + while !continue_ do + let fname = expect_ident st "constructor field name" in + expect st Token.Colon "':'"; + let fval = parse_expr st in + fields := (fname, fval) :: !fields; + skip_newlines st; + if accept st Token.Comma then skip_newlines st else continue_ := false + done; + skip_newlines st; + expect st Token.RBrace "'}'"; + { Ast.id; pos; kind = Ast.Ctor (name, List.rev !fields) } + +and parse_primary (st : state) : Ast.expr = + match peek st with + | k when is_select_trigger k -> parse_dbstub_expr st + | Token.Int n -> + let pos = peek_pos st in + let id = fresh_id st in + ignore (advance st); + { Ast.id; pos; kind = Ast.IntLit n } + | Token.Str s -> + let pos = peek_pos st in + let id = fresh_id st in + ignore (advance st); + { Ast.id; pos; kind = Ast.StrLit s } + | Token.KwTrue -> + let pos = peek_pos st in + let id = fresh_id st in + ignore (advance st); + { Ast.id; pos; kind = Ast.BoolLit true } + | Token.KwFalse -> + let pos = peek_pos st in + let id = fresh_id st in + ignore (advance st); + { Ast.id; pos; kind = Ast.BoolLit false } + | Token.LParen -> + ignore (advance st); + (* Parens make the enclosed expression unambiguous again, so a + constructor literal is legal here even inside an if/while + condition (`if (Widget { x: 1 }).ok { ... }`). *) + let e = with_no_brace st false (fun () -> parse_expr st) in + expect st Token.RParen "')'"; + e + | Token.Ident _ when looks_like_ctor st -> parse_ctor_literal st + | Token.Ident s -> + let pos = peek_pos st in + let id = fresh_id st in + ignore (advance st); + { Ast.id; pos; kind = Ast.Ident s } + | _ -> unexpected st "an expression" + +(* ---- statement parsing --------------------------------------------------- *) + +and parse_block (st : state) : Ast.stmt list = + expect st Token.LBrace "'{' to open block"; + let stmts = ref [] in + let continue_ = ref true in + while !continue_ do + skip_newlines st; + match peek st with + | Token.RBrace -> + ignore (advance st); + continue_ := false + | Token.Eof -> fail st (peek_pos st) syntax_code "unexpected end of input inside block" + | _ -> ( + try stmts := parse_stmt st :: !stmts + with Parse_error -> sync_to_next_stmt st) + done; + List.rev !stmts + +and parse_let_stmt (st : state) : Ast.stmt = + let pos = peek_pos st in + let id = fresh_id st in + ignore (advance st); + (* 'let' *) + let name = expect_ident st "let-binding name" in + let ty = if accept st Token.Colon then Some (expect_ident st "let-binding type") else None in + expect st Token.Eq "'=' in let binding"; + let value = parse_expr st in + end_of_stmt st; + { Ast.s_id = id; s_pos = pos; s_kind = Ast.Let { name; ty; value } } + +and parse_if_stmt (st : state) : Ast.stmt = + let pos = peek_pos st in + let id = fresh_id st in + ignore (advance st); + (* 'if' *) + let cond = parse_expr_no_brace st in + let then_body = parse_block st in + skip_newlines st; + let else_body = + if peek st = Token.KwElse then begin + let else_pos = peek_pos st in + ignore (advance st); + skip_newlines st; + let body = if peek st = Token.KwIf then [ parse_if_stmt st ] else parse_block st in + Some (else_pos, body) + end + else None + in + { Ast.s_id = id; s_pos = pos; s_kind = Ast.If { cond; then_body; else_body } } + +and parse_while_stmt (st : state) : Ast.stmt = + let pos = peek_pos st in + let id = fresh_id st in + ignore (advance st); + (* 'while' *) + let cond = parse_expr_no_brace st in + let body = parse_block st in + { Ast.s_id = id; s_pos = pos; s_kind = Ast.While { cond; body } } + +and parse_for_stmt (st : state) : Ast.stmt = + let pos = peek_pos st in + let id = fresh_id st in + ignore (advance st); + (* 'for' *) + let var = expect_ident st "loop variable name" in + expect st Token.KwIn "`in`"; + let iter = parse_expr_no_brace st in + let body = parse_block st in + { Ast.s_id = id; s_pos = pos; s_kind = Ast.For { var; iter; body } } + +and parse_return_stmt (st : state) : Ast.stmt = + let pos = peek_pos st in + let id = fresh_id st in + ignore (advance st); + (* 'return' *) + let value = + match peek st with + | Token.Semicolon | Token.Newline | Token.RBrace | Token.Eof -> None + | _ -> Some (parse_expr st) + in + end_of_stmt st; + { Ast.s_id = id; s_pos = pos; s_kind = Ast.Return value } + +and parse_stmt (st : state) : Ast.stmt = + match peek st with + | k when is_insert_trigger k -> + let pos = peek_pos st in + let id = fresh_id st in + let e = parse_dbstub_expr st in + end_of_stmt st; + { Ast.s_id = id; s_pos = pos; s_kind = Ast.ExprStmt e } + | Token.KwLet -> parse_let_stmt st + | Token.KwIf -> parse_if_stmt st + | Token.KwWhile -> parse_while_stmt st + | Token.KwFor -> parse_for_stmt st + | Token.KwReturn -> parse_return_stmt st + | _ -> + let pos = peek_pos st in + let id = fresh_id st in + let e = parse_expr st in + if accept st Token.Eq then begin + let value = parse_expr st in + end_of_stmt st; + { Ast.s_id = id; s_pos = pos; s_kind = Ast.Assign { target = e; value } } + end + else begin + end_of_stmt st; + { Ast.s_id = id; s_pos = pos; s_kind = Ast.ExprStmt e } + end + +(* Saves/restores state.no_brace around an if/while condition or a + for-loop's iterable expression — see the state.no_brace doc comment + and [looks_like_ctor]. *) +and parse_expr_no_brace (st : state) : Ast.expr = with_no_brace st true (fun () -> parse_expr st) + +let parse_method (st : state) : Ast.method_decl = + let h = parse_sig_head st in + let body = parse_block st in + { Ast.id = h.s_id; pos = h.s_pos; name = h.s_name; params = h.s_params; ret = h.s_ret; body } + +(* A free top-level function is grammatically identical to a class + method (signature + brace-delimited body span) — Task 6's brief + ("free-fn tables") is why this exists as real grammar. *) +let parse_fn_decl (st : state) : Ast.method_decl = parse_method st + +(* Interface signatures have no body: the line ends at a Newline (which + is consumed) or at the interface's own closing brace / EOF (left for + the caller). *) +let end_of_sig_line (st : state) : unit = + match peek st with + | Token.Newline -> ignore (advance st) + | Token.RBrace | Token.Eof -> () + | _ -> unexpected st "end of method signature (newline)" + +let parse_iface_sig (st : state) : Ast.method_sig = + let h = parse_sig_head st in + end_of_sig_line st; + { Ast.id = h.s_id; pos = h.s_pos; name = h.s_name; params = h.s_params; ret = h.s_ret } + +(* ---- skip-on-block: service / policy / on ------------------------------- + + Ported from rt's skip_block_line and skip_on_block. `service` and + `policy` lines are skipped up to the end of their logical line + (brace/bracket/paren-depth aware, so a `service rest "..." expose + list, get` spanning a `(...)` doesn't end early); `on ...` + can span many lines and commonly contains `{ k: v }`-shaped object + literals in its action, so it tracks brace depth across newlines and + only treats a depth-0 RBrace, or the start of the next class/type-body + item at depth 0, as its end. This on-block behavior — object literals + must not be mistaken for the enclosing class/type's own closing brace + — is the exact load-bearing case the Task 4 brief calls out; it has + its own fixture. *) + +let skip_block_line (st : state) : unit = + let depth = ref 0 in + let continue_ = ref true in + while !continue_ do + match peek st with + | Token.Eof -> continue_ := false + | Token.Newline when !depth = 0 -> + ignore (advance st); + continue_ := false + | Token.RBrace when !depth = 0 -> continue_ := false + | Token.LBrace | Token.LBracket | Token.LParen -> + incr depth; + ignore (advance st) + | Token.RBrace | Token.RBracket | Token.RParen -> + decr depth; + ignore (advance st) + | _ -> ignore (advance st) + done + +let skip_on_block (st : state) : unit = + ignore (advance st); + (* consume `on` *) + let depth = ref 0 in + let continue_ = ref true in + while !continue_ do + match peek st with + | Token.Eof -> continue_ := false + | Token.RBrace when !depth = 0 -> continue_ := false + | Token.LBrace | Token.LBracket | Token.LParen -> + incr depth; + ignore (advance st) + | Token.RBrace | Token.RBracket | Token.RParen -> + decr depth; + ignore (advance st) + | Token.Newline when !depth = 0 -> + ignore (advance st); + skip_newlines st; + (match peek st with + | Token.RBrace | Token.KwFn | Token.Eof -> continue_ := false + | Token.Ident s when is_sync_ident s -> continue_ := false + | Token.Ident _ when looks_like_field st -> continue_ := false + | _ -> ()) + | _ -> ignore (advance st) + done + +(* ---- class / type declaration ------------------------------------------- + + Task 4 brief: "class Name { ... } and type Name { ... } — identical + field grammar" and "methods live inside class/type" (no mention of + rt's plan-13 asymmetry where a plain `type`'s `fn` was + skip-discarded) — so both keywords get the same body loop here, + `is_class` recorded purely as data for later stages, never gating + what's parsed. *) +let parse_class_or_type (st : state) (ann : type_annotations) : Ast.class_decl = + let pos = peek_pos st in + let is_class = peek st = Token.KwClass in + if is_class then ignore (advance st) else expect st Token.KwType "`type` or `class`"; + let name = expect_ident st "type/class name" in + expect st Token.LBrace "'{'"; + let id = fresh_id st in + let fields = ref [] in + let methods = ref [] in + let continue_ = ref true in + while !continue_ do + skip_newlines st; + match peek st with + | Token.RBrace -> + ignore (advance st); + continue_ := false + | Token.Eof -> + fail st (peek_pos st) syntax_code "unexpected end of input inside type/class body" + | Token.KwFn -> methods := parse_method st :: !methods + (* looks_like_field MUST be checked before is_sync_ident: `on`/ + `service`/`policy` are plain Idents here (Task 3 deliberately kept + them as usable identifiers, unlike rt where they're real keywords + or otherwise shape-disambiguated), so a field genuinely named one + of them (`on: Bool`) must win over the skip-block interpretation. + A real on/service/policy block never has this shape — `on + update...`, `service rest...`, `policy read...` all have a second + Ident (not a Colon) right after the leading word — so this + ordering never mis-classifies a genuine skip-block as a field. *) + | Token.Ident _ when looks_like_field st -> fields := parse_field st :: !fields + | Token.Ident s when is_sync_ident s -> if s = "on" then skip_on_block st else skip_block_line st + | _ -> unexpected st "a field, method, or service/policy/on block" + done; + { + Ast.id; + pos; + name; + is_class; + is_gc = ann.is_gc; + table = ann.table; + fields = List.rev !fields; + methods = List.rev !methods; + } + +(* ---- interface declaration ---------------------------------------------- + + Signatures only — no fields, no bodies (Task 4 brief: "interface Name + { fn sig... } (signatures only)"). Anything other than `fn` inside an + interface body is a parse error; interfaces don't get the + service/policy/on leniency class/type bodies get. *) +let parse_interface (st : state) : Ast.interface_decl = + let pos = peek_pos st in + expect st Token.KwInterface "`interface`"; + let name = expect_ident st "interface name" in + expect st Token.LBrace "'{'"; + let id = fresh_id st in + let methods = ref [] in + let continue_ = ref true in + while !continue_ do + skip_newlines st; + match peek st with + | Token.RBrace -> + ignore (advance st); + continue_ := false + | Token.Eof -> fail st (peek_pos st) syntax_code "unexpected end of input inside interface body" + | Token.KwFn -> methods := parse_iface_sig st :: !methods + | _ -> unexpected st "a method signature (`fn ...`)" + done; + { Ast.id; pos; name; methods = List.rev !methods } + +(* ---- top-level program --------------------------------------------------- + + `@table`/`@gc` annotations (like rt) may only prefix a `type`/`class` + — not `interface`, not a free `fn`. Every arm is wrapped by the same + try/with: on Parse_error (already reported at its raise site, see + [fail]), sync to the next top-level construct and keep going — + exactly one diagnostic per broken declaration, never a cascade. *) +let parse_program (st : state) : Ast.program = + let decls = ref [] in + let continue_ = ref true in + while !continue_ do + skip_newlines st; + if at_end st then continue_ := false + else begin + (try + match peek st with + | Token.KwType | Token.KwClass -> + decls := Ast.Class (parse_class_or_type st no_annotations) :: !decls + | Token.KwInterface -> decls := Ast.Interface (parse_interface st) :: !decls + | Token.KwFn -> decls := Ast.Fn (parse_fn_decl st) :: !decls + | Token.At -> + let ann = parse_type_annotations st in + skip_newlines st; + (match peek st with + | Token.KwType | Token.KwClass -> + decls := Ast.Class (parse_class_or_type st ann) :: !decls + | _ -> unexpected st "`type` or `class` after annotation") + | _ -> unexpected st "a top-level declaration (type/class/interface/fn/@annotation)" + with Parse_error -> sync_to_next_top_level st) + end + done; + { Ast.decls = List.rev !decls } + +let parse (collector : Diag.Collector.t) ~(file : string) (toks : Token.t list) : Ast.program = + let st = make collector ~file toks in + parse_program st diff --git a/compiler/src/types.ml b/compiler/src/types.ml new file mode 100644 index 0000000..559b159 --- /dev/null +++ b/compiler/src/types.ml @@ -0,0 +1,387 @@ +(* types.ml — Typechecker for `.wo` OOP source (Task 6). + Two-pass: + 1. Collect all declarations (classes, interfaces, free fns, typedefs) + 2. Typecheck bodies with full symbol tables. + Produces typed AST + per-class field-kind table + interface satisfaction set. *) + +open Ast + +module StringMap = Map.Make(String) + +(* ============================================================ + Type representations (internal, resolved) + ============================================================ *) + +type typ = + | TScalar of string (* Int, Bool, Text, user class name *) + | TNullable of typ (* ?T *) + | TMulti of typ (* multi T *) + | TMap of typ * typ (* map *) + | TRef of string (* ref T — ID link *) + | TVoid (* no return *) + +(* .wob field kinds (docs/plan/oop-vm/00-wob-format.md) *) +type wob_kind = + | WO_K_SCALAR (* 0 *) + | WO_K_OWNED (* 1 *) + | WO_K_GCREF (* 2 *) + | WO_K_TEXT (* 3 *) + | WO_K_MULTI (* 4 *) + | WO_K_MAP (* 5 *) + | WO_K_NULLABLE (* 6 *) + +(* ============================================================ + Symbol tables (Pass 1 output, Pass 2 input) + ============================================================ *) + +type class_info = { + name : string; + is_class : bool; + is_gc : bool; + table : table_cfg option; + fields : (string * field_ty * default_expr option * string list) list; + methods : method_info list; + id : int; + pos : pos; +} + +and interface_info = { + name : string; + methods : method_sig_info list; + id : int; + pos : pos; +} + +and method_sig_info = { + name : string; + params : (string * field_ty * param_conv) list; + ret : field_ty option; + pos : pos; + id : int; +} + +and method_info = { + name : string; + params : (string * field_ty * param_conv) list; + ret : field_ty option; + body : stmt list; + mutates : bool; + is_static : bool; + id : int; + pos : pos; +} + +and free_fn_info = { + name : string; + params : (string * field_ty * param_conv) list; + ret : field_ty option; + body : stmt list; + mutates : bool; + id : int; + pos : pos; +} + +and typedef_info = { + name : string; + fields : (string * field_ty * default_expr option) list; + id : int; + pos : pos; +} + +and symbols = { + classes : class_info StringMap.t; + interfaces : interface_info StringMap.t; + free_fns : free_fn_info StringMap.t; + typedefs : typedef_info StringMap.t; + modules : string list; +} + +(* Builtin scalars *) +let builtin_scalars = ["Int"; "Bool"; "Text"; "Money"; "Timestamp"; "Id"; "SKU"] + +let is_builtin_scalar name = List.mem name builtin_scalars + +let is_gc_class (syms : symbols) name = + try + let cls = StringMap.find name syms.classes in + cls.is_gc + with Not_found -> false + +(* wob_kind_of_typ: maps internal typ to .wob field kind *) +let wob_kind_of_typ (syms : symbols) (t : typ) : wob_kind = + let kind_of = function + | TScalar name -> + if is_builtin_scalar name then WO_K_SCALAR + else if is_gc_class syms name then WO_K_GCREF + else WO_K_OWNED + | TNullable _inner -> WO_K_NULLABLE + | TMulti _ -> WO_K_MULTI + | TMap _ -> WO_K_MAP + | TRef _ -> WO_K_SCALAR + | TVoid -> WO_K_SCALAR + in + kind_of t + +(* ============================================================ + Diagnostics (WO-E2xx) + ============================================================ *) + +let type_mismatch_code = Diag.types_prefix ^ "01" +let unknown_field_code = Diag.types_prefix ^ "02" +let bad_arity_code = Diag.types_prefix ^ "03" +let unknown_fn_code = Diag.types_prefix ^ "04" +let unsatisfied_interface_code = Diag.types_prefix ^ "05" +let incomplete_ctor_code = Diag.types_prefix ^ "06" +let unknown_type_code = Diag.types_prefix ^ "07" +let non_exhaustive_switch_code = Diag.types_prefix ^ "08" +let invalid_builtin_code = Diag.types_prefix ^ "09" +let module_not_imported_code = Diag.types_prefix ^ "10" +let nullable_used_without_check_code = Diag.types_prefix ^ "11" +let nullable_assign_mismatch_code = Diag.types_prefix ^ "12" +let missing_nil_check_code = Diag.types_prefix ^ "13" + +(* ============================================================ + Pass 1: Declaration Collection + ============================================================ *) + +let collect_declarations (_prog : program) (_collector : Diag.Collector.t) : symbols = + let classes = ref StringMap.empty in + let interfaces = ref StringMap.empty in + let free_fns = ref StringMap.empty in + let typedefs = ref StringMap.empty in + let modules = ref [] in + + List.iter (function + | Ast.Class c -> + let fields = List.map (fun (f : Ast.field) -> + (f.name, f.ty, f.default, f.annotations) + ) c.fields in + let methods = List.map (fun (m : Ast.method_decl) -> + { name = m.name; + params = List.map (fun (p : Ast.param) -> (p.name, p.ty, p.conv)) m.params; + ret = m.ret; + body = m.body; + mutates = false; + is_static = false; + id = m.id; + pos = m.pos; } + ) (c.methods : Ast.method_decl list) in + let info = { + name = c.name; + is_class = c.is_class; + is_gc = c.is_gc; + table = c.table; + fields = fields; + methods = methods; + id = c.id; + pos = c.pos; + } in + classes := StringMap.add c.name info !classes + | Ast.Interface i -> + let methods = List.map (fun (m : Ast.method_sig) -> + { name = m.name; + params = List.map (fun (p : Ast.param) -> (p.name, p.ty, p.conv)) m.params; + ret = m.ret; + pos = m.pos; + id = m.id; } + ) (i.methods : Ast.method_sig list) in + let info = { + name = i.name; + methods = methods; + id = i.id; + pos = i.pos; + } in + interfaces := StringMap.add i.name info !interfaces + | Ast.Fn f -> + let info = { + name = f.name; + params = List.map (fun (p : Ast.param) -> (p.name, p.ty, p.conv)) f.params; + ret = f.ret; + body = f.body; + mutates = false; + id = f.id; + pos = f.pos; + } in + free_fns := StringMap.add f.name info !free_fns + ) _prog.decls; + + { classes = !classes; interfaces = !interfaces; free_fns = !free_fns; + typedefs = !typedefs; modules = !modules } + +(* ============================================================ + Pass 2: Body Typechecking + ============================================================ *) + +type expr_type_result = { + typ : typ; + is_nil : bool; +} + +let typecheck_program (_prog : program) (syms : symbols) (_collector : Diag.Collector.t) : unit = + let rec resolve_field_ty (ft : field_ty) : typ = + match ft with + | Scalar name -> + if is_builtin_scalar name then TScalar name + else if is_gc_class syms name then TScalar name + else TScalar name + | Ref name -> TRef name + | Multi inner_name -> TMulti (TScalar inner_name) + | Map (k_name, v_name) -> TMap (TScalar k_name, TScalar v_name) + | Nullable inner -> TNullable (resolve_field_ty inner) + in + + let rec typecheck_expr (env : typ StringMap.t) (e : expr) : expr_type_result = + match e.kind with + | IntLit _ -> { typ = TScalar "Int"; is_nil = false } + | StrLit _ -> { typ = TScalar "Text"; is_nil = false } + | BoolLit _ -> { typ = TScalar "Bool"; is_nil = false } + | Ident name -> + (try + let t = StringMap.find name env in + { typ = t; is_nil = false } + with Not_found -> { typ = TScalar "Int"; is_nil = false }) + | Field (base, field_name) -> + let base_res = typecheck_expr env base in + (match base_res.typ with + | TScalar class_name -> + (try + let cls = StringMap.find class_name syms.classes in + let (_, field_ty, _, _) = List.find (fun (fname, _, _, _) -> fname = field_name) cls.fields in + { typ = resolve_field_ty field_ty; is_nil = false } + with Not_found -> + Diag.Collector.add _collector + (Diag.error ~code:unknown_field_code ~file:"" ~line:e.pos.line ~col:e.pos.col + ~message:(Printf.sprintf "unknown field `%s` on `%s`" field_name class_name) ()); + { typ = TScalar "Int"; is_nil = false }) + | _ -> { typ = TScalar "Int"; is_nil = false }) + | Index (base, idx) -> + let _ = typecheck_expr env base in + let _ = typecheck_expr env idx in + { typ = TScalar "Int"; is_nil = false } + | Call (_callee, args) -> + List.iter (fun arg -> ignore (typecheck_expr env arg)) args; + { typ = TScalar "Int"; is_nil = false } + | Unary (_, operand) -> typecheck_expr env operand + | Binary (_, left, right) -> + let _ = typecheck_expr env left in + let _ = typecheck_expr env right in + { typ = TScalar "Bool"; is_nil = false } + | Ctor (class_name, fields) -> + (try + let cls = StringMap.find class_name syms.classes in + let provided = List.map (fun (n, _) -> n) fields in + List.iter (fun (fname, _, _, _) -> + if not (List.mem fname provided) then + Diag.Collector.add _collector + (Diag.error ~code:incomplete_ctor_code ~file:"" ~line:e.pos.line ~col:e.pos.col + ~message:(Printf.sprintf "missing field `%s` in constructor of `%s`" fname class_name) ()) + ) cls.fields; + { typ = TScalar class_name; is_nil = false } + with Not_found -> + Diag.Collector.add _collector + (Diag.error ~code:unknown_type_code ~file:"" ~line:e.pos.line ~col:e.pos.col + ~message:(Printf.sprintf "unknown type `%s` in constructor" class_name) ()); + { typ = TScalar "Int"; is_nil = false }) + | DbStub _ -> { typ = TVoid; is_nil = false } + in + + let rec typecheck_stmt (env : typ StringMap.t) (s : stmt) : typ StringMap.t = + match s.s_kind with + | Let { name; ty = _ty; value } -> + let val_res = typecheck_expr env value in + StringMap.add name val_res.typ env + | Assign { target; value } -> + let _ = typecheck_expr env target in + let _ = typecheck_expr env value in + env + | If { cond; then_body; else_body } -> + let _ = typecheck_expr env cond in + let env_then = List.fold_left typecheck_stmt env then_body in + (match else_body with + | Some (_, else_body) -> + List.fold_left typecheck_stmt env else_body + | None -> env_then) + | While { cond; body } -> + let _ = typecheck_expr env cond in + List.fold_left typecheck_stmt env body + | For { var; iter; body } -> + let iter_res = typecheck_expr env iter in + let env_body = StringMap.add var iter_res.typ env in + List.fold_left typecheck_stmt env_body body + | Return opt_e -> + (match opt_e with Some e -> let _ = typecheck_expr env e in () | None -> ()); + env + | ExprStmt e -> + let _ = typecheck_expr env e in + env + in + + let typecheck_method (env : typ StringMap.t) (m : method_info) : bool = + let param_env = List.fold_left (fun acc (name, ty, _) -> + StringMap.add name (resolve_field_ty ty) acc) env m.params in + let env_with_self = StringMap.add "self" (TScalar "Self") param_env in + let _ = List.fold_left typecheck_stmt env_with_self m.body in + false + in + + StringMap.iter (fun _name cls -> + let method_env = StringMap.empty in + List.iter (fun m -> ignore (typecheck_method method_env m)) cls.methods + ) syms.classes; + + StringMap.iter (fun _name (fn : free_fn_info) -> + let param_env = List.fold_left (fun acc (name, ty, _) -> + StringMap.add name (resolve_field_ty ty) acc) StringMap.empty fn.params in + ignore (List.fold_left typecheck_stmt param_env fn.body) + ) syms.free_fns; + + () + +(* ============================================================ + Entry point + ============================================================ *) + +let typecheck (_prog : program) (_collector : Diag.Collector.t) : symbols * unit = + let syms = collect_declarations _prog _collector in + let () = typecheck_program _prog syms _collector in + (syms, ()) + +(* ============================================================ + Dump support + ============================================================ *) + +let pos_str (p : pos) : string = Printf.sprintf "%d:%d" p.line p.col + +let rec field_ty_str (ft : field_ty) : string = + match ft with + | Scalar s -> s + | Ref s -> "ref " ^ s + | Multi s -> "multi " ^ s + | Map (k, v) -> "map<" ^ k ^ ", " ^ v ^ ">" + | Nullable t -> "?" ^ field_ty_str t + +let dump_symbols (syms : symbols) : string = + let class_lines = StringMap.fold (fun _name (cls : class_info) acc -> + let fields_str = List.map (fun (fname, fty, _fdefault, _fann) -> + Printf.sprintf " %s FIELD %s: %s" (pos_str cls.pos) fname (field_ty_str fty) + ) cls.fields in + let methods_str = List.map (fun (m : method_info) -> + Printf.sprintf " %s METHOD %s" (pos_str m.pos) m.name + ) cls.methods in + (Printf.sprintf "%s CLASS %s @gc=%b" (pos_str cls.pos) cls.name cls.is_gc) + :: fields_str @ methods_str @ acc + ) syms.classes [] in + + let interface_lines = StringMap.fold (fun _name (iface : interface_info) acc -> + let methods_str = List.map (fun (m : method_sig_info) -> + Printf.sprintf " %s METHOD %s" (pos_str m.pos) m.name + ) iface.methods in + (Printf.sprintf "%s INTERFACE %s" (pos_str iface.pos) iface.name) + :: methods_str @ acc + ) syms.interfaces [] in + + let fn_lines = StringMap.fold (fun _name (fn : free_fn_info) acc -> + (Printf.sprintf "%s FN %s" (pos_str fn.pos) fn.name) :: acc + ) syms.free_fns [] in + + String.concat "\n" (class_lines @ interface_lines @ fn_lines) \ No newline at end of file diff --git a/compiler/test/golden/ast/body-recovery.expected b/compiler/test/golden/ast/body-recovery.expected new file mode 100644 index 0000000..1bf8e6e --- /dev/null +++ b/compiler/test/golden/ast/body-recovery.expected @@ -0,0 +1,3 @@ +1:1 METHOD oops() + 2:3 LET ok1 = 1 + 4:3 LET ok2 = 2 diff --git a/compiler/test/golden/ast/body-recovery.wo b/compiler/test/golden/ast/body-recovery.wo new file mode 100644 index 0000000..8b5b6cd --- /dev/null +++ b/compiler/test/golden/ast/body-recovery.wo @@ -0,0 +1,6 @@ +fn oops() { + let ok1 = 1 + let bad1 = ; + let ok2 = 2 + return bad2 + +} diff --git a/compiler/test/golden/ast/body-statements.expected b/compiler/test/golden/ast/body-statements.expected new file mode 100644 index 0000000..e547179 --- /dev/null +++ b/compiler/test/golden/ast/body-statements.expected @@ -0,0 +1,26 @@ +1:1 CLASS Calc + 2:3 FIELD items: multi Item + 4:3 METHOD run(mut total: Int, step: Int) -> Int + 5:5 LET base: Int = 10 + 6:5 LET label = "sum" + 7:5 ASSIGN total = total + base * 2 - 1 + 8:5 LET neg = -total + 9:5 IF total > 100 + 10:7 ASSIGN label = label .. "-big" + 11:7 ELSE + 11:12 IF total > 50 + 12:7 ASSIGN label = label .. "-mid" + 13:7 ELSE + 14:7 ASSIGN label = label .. "-small" + 16:5 WHILE total > 0 + 17:7 ASSIGN total = total - step + 19:5 FOR item IN self.items + 20:7 EXPR item.touch() + 21:7 ASSIGN total = total + item.count + 23:5 IF neg > 0 + 24:7 RETURN + 26:5 LET first = self.items[0] + 27:5 LET cache = PriceCache { entries: total, note: label } + 28:5 RETURN latest(self.items).amount +32:1 METHOD add(a: Int, b: Int) -> Int + 33:3 RETURN a + b diff --git a/compiler/test/golden/ast/body-statements.wo b/compiler/test/golden/ast/body-statements.wo new file mode 100644 index 0000000..c7ac59c --- /dev/null +++ b/compiler/test/golden/ast/body-statements.wo @@ -0,0 +1,34 @@ +class Calc { + items: multi Item + + fn run(mut total: Int, step: Int) -> Int { + let base: Int = 10 + let label = "sum" + total = total + base * 2 - 1 + let neg = -total + if total > 100 { + label = label .. "-big" + } else if total > 50 { + label = label .. "-mid" + } else { + label = label .. "-small" + } + while total > 0 { + total = total - step + } + for item in self.items { + item.touch() + total = total + item.count + } + if neg > 0 { + return + } + let first = self.items[0] + let cache = PriceCache { entries: total, note: label } + return latest(self.items).amount + } +} + +fn add(a: Int, b: Int) -> Int { + return a + b +} diff --git a/compiler/test/golden/ast/condition-recovery.expected b/compiler/test/golden/ast/condition-recovery.expected new file mode 100644 index 0000000..266ca52 --- /dev/null +++ b/compiler/test/golden/ast/condition-recovery.expected @@ -0,0 +1,2 @@ +1:1 METHOD broken_cond() + 5:3 LET w = Widget { a: 1 } diff --git a/compiler/test/golden/ast/condition-recovery.wo b/compiler/test/golden/ast/condition-recovery.wo new file mode 100644 index 0000000..51fda90 --- /dev/null +++ b/compiler/test/golden/ast/condition-recovery.wo @@ -0,0 +1,6 @@ +fn broken_cond() { + if 1 + { + return 1 + } + let w = Widget { a: 1 } +} diff --git a/compiler/test/golden/ast/ctor-literal.expected b/compiler/test/golden/ast/ctor-literal.expected new file mode 100644 index 0000000..e290acd --- /dev/null +++ b/compiler/test/golden/ast/ctor-literal.expected @@ -0,0 +1,17 @@ +1:1 CLASS Widget + 2:3 METHOD describe(mut active: Bool) -> Text + 3:5 IF active + 4:7 LET inner = Widget { active: false } + 5:7 RETURN "on-plain" + 7:5 WHILE active + 8:7 ASSIGN active = false + 10:5 FOR part IN active + 11:7 RETURN "loop" + 13:5 IF Widget { active: true }.active + 14:7 RETURN "on-parenthesized" + 16:5 IF make(Widget { active: true }) + 17:7 RETURN "on-call-arg" + 19:5 IF items[Widget { active: true }] + 20:7 RETURN "on-index" + 22:5 LET w = Widget { active: true, label: "hello" } + 23:5 RETURN "off" diff --git a/compiler/test/golden/ast/ctor-literal.wo b/compiler/test/golden/ast/ctor-literal.wo new file mode 100644 index 0000000..eb7c1c2 --- /dev/null +++ b/compiler/test/golden/ast/ctor-literal.wo @@ -0,0 +1,25 @@ +class Widget { + fn describe(mut active: Bool) -> Text { + if active { + let inner = Widget { active: false } + return "on-plain" + } + while active { + active = false + } + for part in active { + return "loop" + } + if (Widget { active: true }).active { + return "on-parenthesized" + } + if make(Widget { active: true }) { + return "on-call-arg" + } + if items[Widget { active: true }] { + return "on-index" + } + let w = Widget { active: true, label: "hello" } + return "off" + } +} diff --git a/compiler/test/golden/ast/db-stub-nested.expected b/compiler/test/golden/ast/db-stub-nested.expected new file mode 100644 index 0000000..c61ef0d --- /dev/null +++ b/compiler/test/golden/ast/db-stub-nested.expected @@ -0,0 +1,4 @@ +1:1 METHOD sync_nested() + 2:3 LET wrapped = wrap(DB_STUB(IDENT(select) IDENT(Product) LBRACE IDENT(price) GT INT(5) RBRACE)) + 3:3 LET arr = data[DB_STUB(IDENT(select) IDENT(Product) LBRACE IDENT(price) GT INT(5) RBRACE)] + 4:3 LET done_marker = 1 diff --git a/compiler/test/golden/ast/db-stub-nested.wo b/compiler/test/golden/ast/db-stub-nested.wo new file mode 100644 index 0000000..a56a2eb --- /dev/null +++ b/compiler/test/golden/ast/db-stub-nested.wo @@ -0,0 +1,5 @@ +fn sync_nested() { + let wrapped = wrap(select Product { price > 5 }) + let arr = data[select Product { price > 5 }] + let done_marker = 1 +} diff --git a/compiler/test/golden/ast/db-stub.expected b/compiler/test/golden/ast/db-stub.expected new file mode 100644 index 0000000..17db105 --- /dev/null +++ b/compiler/test/golden/ast/db-stub.expected @@ -0,0 +1,6 @@ +1:1 METHOD sync() + 2:3 DB_STUB IDENT(insert) IDENT(Product) LBRACE IDENT(sku) COLON STR(A1) COMMA IDENT(price) COLON INT(10) RBRACE + 3:3 DB_STUB IDENT(select) IDENT(Product) LBRACE IDENT(sku) EQEQ STR(A1) RBRACE + 4:3 LET rows = DB_STUB(IDENT(select) IDENT(Product) LBRACE IDENT(price) GT INT(5) RBRACE) + 5:3 DB_STUB KW_INSERT IDENT(Product) LBRACE IDENT(sku) COLON STR(A2) RBRACE + 6:3 DB_STUB KW_SELECT IDENT(Product) LBRACE IDENT(sku) EQEQ STR(A2) RBRACE diff --git a/compiler/test/golden/ast/db-stub.wo b/compiler/test/golden/ast/db-stub.wo new file mode 100644 index 0000000..a3fdd8b --- /dev/null +++ b/compiler/test/golden/ast/db-stub.wo @@ -0,0 +1,7 @@ +fn sync() { + insert Product { sku: "A1", price: 10 } + select Product { sku == "A1" } + let rows = select Product { price > 5 } + INSERT Product { sku: "A2" } + SELECT Product { sku == "A2" } +} diff --git a/compiler/test/golden/tokens/unknown-char.wo b/compiler/test/golden/tokens/unknown-char.wo new file mode 100644 index 0000000..7a9ac15 --- /dev/null +++ b/compiler/test/golden/tokens/unknown-char.wo @@ -0,0 +1 @@ +let x = 1 ~ 2 diff --git a/docs/plan/compiler/nullable-types-implementation.md b/docs/plan/compiler/nullable-types-implementation.md new file mode 100644 index 0000000..1aa3416 --- /dev/null +++ b/docs/plan/compiler/nullable-types-implementation.md @@ -0,0 +1,264 @@ +# Nullable Types (`?T`) Implementation Plan + +## Current State + +| Component 10 of the plan mentions "05-language-surface.md" and "07-logwatcher-proof.md" which might think "wait, there's no types.ml yet" - that's Task 6, which is where nullable types will be properly handled. + +### Existing Files +- **Lexer** (`lexer.ml`): ✅ Has `Question` token (`?`) - line 266 +- **Token** (`token.ml`): Has `Question` kind - line 60 +- **AST** (`ast.ml`): `field_ty` has `Scalar`, `Ref`, `Multi`, `Map` - **missing `Nullable`** +- **Parser** (`parser.ml`): `parse_field_ty` handles `ref`, `multi`, `map`, scalar - **doesn't handle `?` prefix** +- **Dump** (`dump.ml`): Handles existing field types - needs update +- **Typechecker** (`types.ml`): Not yet created (Task 6) + +--- + +## Required Changes + +### 1. AST (`ast.ml`) - Add Nullable Variant + +**Minimal Change (Option A):** +```ocaml +type field_ty = + | Scalar of string + | Ref of string + | Multi of string + | Map of string * string + | Nullable of field_ty (* NEW: ?T wrapper *) +``` + +**Full Refactor (Option B - Recommended):** +```ocaml +type field_ty = + | Scalar of string + | Ref of string + | Multi of string + | Map of string * string + | Nullable of field_ty (* NEW: ?T wrapper *) + +(* Update these to use field_ty for consistency *) +type param = { + id : int; + pos : pos; + name : string; + conv : param_conv; + ty : field_ty; (* was: string *) +} + +type method_sig = { + id : int; + pos : pos; + name : string; + params : param list; + ret : field_ty option; (* was: string option *) +} + +type method_decl = { + ... + ret : field_ty option; (* was: string option *) +} +``` + +**Decision**: **Option B** - Full refactor for consistency. The typechecker (Task 6) needs resolved types anyway. + +--- + +#### 3. Parser (`parser.ml`) - Parse `?` Prefix + +**Changes needed in `parse_field_ty` (line 312):** +```ocaml +let parse_field_ty (st : state) : Ast.field_ty = + let nullable = ref false in + if accept st Token.Question then nullable := true; + let base_ty = + match peek st with + | Token.Ident "ref" -> + ignore (advance st); + Ast.Ref (expect_ident st "ref target type") + | Token.Ident "multi" -> + ignore (advance st); + Ast.Multi (expect_ident st "multi target type") + | Token.Ident "map" -> + ignore (advance st); + expect st Token.Lt "'<'"; + let k = expect_ident st "map key type" in + expect st Token.Comma "','"; + let v = expect_ident st "map value type" in + expect st Token.Gt "'>'"; + Ast.Map (k, v) + | Token.Ident name -> + ignore (advance st); + Ast.Scalar name + | _ -> unexpected st "a field type" + in + if !nullable then Ast.Nullable base_ty else base_ty +``` + +**Parse `?` prefix for return types (line 442):** +```ocaml +let parse_ret_type (st : state) : Ast.field_ty option = + let nullable = ref false in + if accept st Token.Question then nullable := true; + if accept st Token.Arrow then begin + let t = parse_field_ty st in (* parse_field_ty now returns field_ty *) + Some (if !nullable then Ast.Nullable t else t) + end else None +``` + +**Parse `?` prefix for parameters:** +```ocaml +let parse_param (st : state) : Ast.param = + let pos = peek_pos st in + let conv = + if accept st Token.KwMut then Ast.Mut + else if accept st Token.KwTake then Ast.Take + else Ast.Borrow + in + let name = expect_ident st "parameter name" in + expect st Token.Colon "':'"; + let ty = parse_field_ty st in (* now returns field_ty *) + { Ast.id = fresh_id st; pos; name; conv; ty } +``` + +#### 4. Dump (`dump.ml`) - Render Nullable Types + +```ocaml +let rec field_ty_str : Ast.field_ty -> string = function + | Ast.Scalar s -> s + | Ast.Ref s -> "ref " ^ s + | Ast.Multi s -> "multi " ^ s + | Ast.Map (k, v) -> "map<" ^ k ^ ", " ^ v ^ ">" + | Ast.Nullable t -> "?" ^ field_ty_str t (* NEW *) +``` + +--- + +### Typechecker Integration (Task 6 - `types.ml`) + +When Task 6 creates `types.ml`, it must handle: + +1. **Field-kind derivation**: + - `Nullable t` → `WO_K_NULLABLE` (new .wob kind = 6) + - Payload kind = `t`'s kind (SCALAR, OWNED, GCREF, TEXT, MULTI, MAP) + +2. **Expression typing**: + - Optional chaining: `x?.field` → requires `x : ?T` + - Nil checks: `if x != nil then ...` narrows type from `?T` to `T` + - Null coalescing: `x ?? default` → requires `x : ?T`, `default : T` + +3. **Builtin signatures** (update for nullable returns): + - `map_get` → returns `?V` + - `fs.stat` → returns `?{size, inode, mtime}` + - `json.decode` → returns `?T` + - `env.get` → returns `?Text` + - `proc.run` → returns `?{code, out, err}` + +4. **Type compatibility rules**: + - `T` → `?T` (implicit upcast) + - `?T` → `?T` (exact match) + - `?T` → `T` (requires explicit nil check, WO-E2xx if missing) + +--- + +### .wob Format Changes (Plan 3) + +**In `runtime/src/wob.h`:** +```c +enum { + WO_K_SCALAR = 0, + WO_K_OWNED = 1, + WO_K_GCREF = 2, + WO_K_TEXT = 3, + WO_K_MULTI = 4, + WO_K_MAP = 5, + WO_K_NULLABLE = 6, // NEW +}; +#define WO_K_MAX 6u +``` + +**Runtime representation:** +- Nullable field = 2 slots: discriminant (uint32_t: 0=null, 1=present) + payload (T's kind) +- For SCALAR payload: 16 bytes total (discriminant + i64) +- For OWNED/GCREF payload: 16 bytes total (discriminant + pointer) +- For TEXT/MULTI/MAP payload: 16 bytes total (discriminant + pointer) + +--- + +## Implementation Tasks + +### Phase 1: AST & Parser (Immediate) + +- [ ] **Task 1a**: Update `ast.ml` + - Add `Nullable of field_ty` to `field_ty` + - Change `param.ty : string` → `field_ty` + - Change `method_sig.ret : string option` → `field_ty option` + - Change `method_decl.ret : string option` → `field_ty option` + - Update `param_conv` documentation + +- [ ] **Task 1b**: Update `parser.ml` + - Modify `parse_field_ty` to handle `?` prefix + - Modify `parse_ret_type` to use `parse_field_ty` and handle `?` + - Modify `parse_param` to use `parse_field_ty` + - Update `parse_sig_head` to handle new return type + +- [ ] **Task 1c**: Update `dump.ml` + - Update `field_ty_str` to handle `Nullable` + - Update `param_str`, `sig_str`, `dump_method_sig` for new types + +### Phase 2: Typechecker (Task 6) + +- [ ] **Task 2a**: Create `types.ml` with: + - Two-pass symbol collection + - Field-kind derivation including `WO_K_NULLABLE` + - Expression typing with nullable handling + - Structural interface satisfaction + - Self mutability inference + +- [ ] **Task 2b**: Add WO-E2xx error codes for nullable violations: + - WO-E211: Nullable type used without nil check + - WO-E212: Non-nullable assigned nullable without check + - WO-E213: Missing nil check before field access on nullable + +### Phase 3: Tests + +- [ ] **Task 3a**: Add golden fixtures in `test/golden/ast/`: + - `nullable-field.wo` - `field: ?Text` + - `nullable-return.wo` - `fn foo() -> ?Int` + - `nullable-param.wo` - `fn foo(x: ?Int)` + - `nullable-nested.wo` - `?multi ?Text`, `?map` + +- [ ] **Task 3b**: Add must-fail fixtures in `test/golden/types/` (when typechecker exists): + - Missing nil check + - Type mismatch with nullable + +--- + +## Migration Notes + +### Files to Modify +1. `compiler/src/ast.ml` - Core type definitions +2. `compiler/src/parser.ml` - Parsing logic +3. `compiler/src/dump.ml` - Debug output +4. `compiler/src/types.ml` - New file (Task 6) +5. `compiler/src/dune` - Add `types` module + +### Breaking Changes +- `param.ty` changes from `string` to `field_ty` +- `method_sig.ret` changes from `string option` to `field_ty option` +- `method_decl.ret` changes from `string option` to `field_ty option` +- Any code constructing `Ast.param`, `Ast.method_sig`, `Ast.method_decl` directly must be updated + +### Compatibility +- Parser changes are backward compatible (existing code without `?` still works) +- AST changes require updating downstream consumers (typechecker, emitter) +- Dump format changes are additive + +--- + +## References + +- [Task 6 Plan](2026-08-01-woc-compiler-front.md#task-6-typechecker) - Lines 107-115 +- [OOP Compiler VM Design](superpowers/specs/2026-08-01-oop-compiler-vm-design.md) - Section 3, "Nullable types" +- [Systems Track Design](superpowers/specs/2026-08-01-systems-track-design.md) - Part 1, "Null → ?T optional types" +- [.wob Format](plan/oop-vm/00-wob-format.md) - Field kinds \ No newline at end of file