feat(compiler): nullable types (?T) support

- ast.ml: Added Nullable variant to field_ty; param.ty/method_sig.ret/method_decl.ret now use field_ty
- parser.ml: Parse ? prefix for field types, return types, parameters
- dump.ml: Render ?T in --dump-ast output
- types.ml: New typechecker (Task 6) with nullable field-kind derivation (WO_K_NULLABLE=6)
- dune: Added types module
- Added golden test fixtures for nullable types and statement/expression parser
- Added implementation plan doc
This commit is contained in:
shoney.arickathil 2026-08-10 08:39:29 +02:00
parent c8adc7230d
commit ad806c415d
19 changed files with 2536 additions and 0 deletions

308
compiler/src/ast.ml Normal file
View file

@ -0,0 +1,308 @@
(* ast.ml — AST for `.wo` OOP source (milestone 1).
Declarations (Task 4 of compiler/plan/2026-08-01-woc-compiler-front.md):
`interface` (method signatures, no bodies), `class` and `type`
(identical field grammar — Task 4 brief's own words: they differ
only in the `is_class` flag, exactly mirroring crates/rt/src/ast.rs's
`TypeDecl { is_class: bool, .. }` design), and `fn` (both as
class/type methods and as free top-level functions — Task 6's brief
mentions "free-fn tables", so free `fn` is real grammar here, not a
rt carry-over). Statements and expressions (Task 5, below) live
inside a method/fn's `body`, which Task 4 captured as a verbatim
token span and Task 5 parses for real.
Every declaration-shaped node (class/type, interface, method/fn,
field, param) carries a unique `id` (monotonic per parse — Tasks 6/7
key side tables, e.g. field-kind and ownership-state tables, on these
ids) and a `pos` — the node's own starting source position. This
codebase has no existing notion of a source *range* (Token.t and
diag.ml's `site` are both single points), so `pos` follows that same
single-point convention rather than inventing a new range type.
`field_ty` and `default_expr` are plain payload, not "declarations" —
they don't get their own `id`/`pos`; nothing downstream needs to key
a side table on "this specific field's type expression" independent
of the field that owns it.
Task 5 adds real statement/expression ASTs (`stmt`/`expr` below) and
retires the body-as-token-span placeholder: `method_decl.body` is now
`stmt list`, not a verbatim span. Every `stmt` and every `expr` gets
its own `id`/`pos` too — unlike `field_ty`/`default_expr`, Tasks 6/7
name concrete per-node consumers (every expression gets a type; move
sites, rc sites, and residual sites are individual call-argument and
indexing expressions), so these *are* the "declaration-shaped" case
the paragraph above describes, not the opaque-payload case.
One disclosed gap in the parent-id-<-child-id invariant Task 4 set up
for the declaration skeleton (a container's id is minted before its
children's): a `stmt`'s id is still minted before its own
sub-expressions/sub-blocks are parsed, so that half holds exactly as
before. It does NOT extend to expr-inside-expr — left-recursive
binary/postfix parsing (`a + b`, `a.b`, `a(b)`) parses the left/base
operand first (smaller id) and only decides to wrap it in
`Binary`/`Field`/`Call` after seeing the next token, so that
wrapper's id is necessarily minted *after* its own child's. Every id
is still unique and monotonic in mint order; only the strict
parent-<-child direction is given up, and only for expr-in-expr
nesting. Forcing it there would mean pre-reserving ids speculatively
before knowing a wrapper is even needed — not worth the complexity
for side-table keys that only need uniqueness, not order. *)
type pos = {
line : int;
col : int;
}
(* `ref`/`multi`/`map` are not lexer keywords (Task 3 deliberately
dropped rt's schema-keyword zoo) — they're recognized positionally,
by name, only at the start of a field's type, exactly like rt's own
`insert`/`select` statement-keyword convention. Scalar also covers a
bare class name used as an owned-embed field type (spec section 4:
"Fields hold owned values, ref T ids, ... or @gc references") —
Task 6 decides whether a given Scalar name is a builtin scalar or a
user class. `?T` nullable wrapper (adopted from Haxe, systems-track
spec Part 1) wraps any field_ty; milestone-1 has no array (`[T]`)
or tagged unions — those are rt schema-layer features not named in
this task's grammar, so they're deliberately absent here. *)
type field_ty =
| Scalar of string
| Ref of string
| Multi of string
| Map of string * string (* key type, value type: map<K, V> *)
| Nullable of field_ty (* ?T wrapper *)
(* Parameter passing convention (spec section 3, rule 2): default is an
immutable borrow; `mut` is an exclusive borrow; `take` moves
ownership in. The owner pass (Task 7) is the eventual consumer. *)
type param_conv =
| Borrow
| Mut
| Take
type param = {
id : int;
pos : pos;
name : string;
conv : param_conv;
ty : field_ty; (* declared type; Task 6 resolves it for real *)
}
(* `= now()` is recognized explicitly (mirrors rt's DefaultExpr::Now).
Everything else is kept as its raw token span rather than eagerly
turned into a string (rt's DefaultExpr::Opaque(String) precedent) —
"opaque token span" per the Task 4 brief, so a later stage could in
principle re-lex/interpret it without having thrown information
away. Task 4 itself never inspects the contents. *)
type default_expr =
| DefaultNow
| DefaultOpaque of Token.t list
type field = {
id : int;
pos : pos;
name : string;
ty : field_ty;
default : default_expr option;
(* Annotation *names* only (e.g. ["unique"]) — the brief: "names
recorded, unknown names fine at parse level". Any `(...)` argument
list on a field annotation is consumed and discarded, matching
rt's own field-annotation handling; nothing downstream needs those
arguments in Task 4. *)
annotations : string list;
}
(* ---- expressions (Task 5) ------------------------------------------
`unop`/`binop` name their operators the way rt's ast.rs BinOp/UnOp
does for the operators this grammar shares with it (Add/Sub/Mul/
Div/Mod, Eq/Ne/Lt/Le/Gt/Ge). Two deliberate differences from rt: no
`And`/`Or` (this grammar's lexer, Task 3, has no `&&`/`||` tokens —
there is no boolean-logic sublanguage here) and one new operator,
`Concat`.
`Concat` (source syntax `..`, `Token.DotDot`) is this task's own
design decision, not a straight rt port: the brief's precedence
ladder lists "text concatenation" as its own tier, distinct from
arithmetic, and the VM design spec's instruction table
(docs/superpowers/specs/2026-08-01-oop-compiler-vm-design.md §5)
lists a dedicated `CONCAT` op alongside (not folded into) `ADD` —
so the source grammar needs its own operator token to compile down
to that, not an overload of `+` disambiguated by operand type later.
`Token.DotDot` is lexed (Task 3) but was never given a grammar rule
before now, so this claims it. Flagged as an inference, not a
spec-literal instruction, because no fixture upstream of this task
spells out the token; it is the only unclaimed binary-shaped token
left, and Lua's `..` is the same design (concat binds looser than
`+`/`-`, tighter than comparison — this ladder's ordering). *)
type unop = Neg
type binop =
| Add
| Sub
| Mul
| Div
| Mod
| Concat
| Eq
| Ne
| Lt
| Le
| Gt
| Ge
type expr = {
id : int;
pos : pos;
kind : expr_kind;
}
(* `Call`'s callee is a general `expr`, not a name: a free call has an
`Ident` callee, a method call (interface-typed or not — dispatch is
Task 6's job, not the parser's) has a `Field` callee. Brief: "free,
method, and interface-typed method calls share one call node" — this
is that sharing; the parser never distinguishes the three, it only
ever builds `Call (callee, args)`.
`Ctor` (`ClassName { field: expr, ... }`) is recognized in primary-
expression position by the two-token shape identifier-then-brace
(parser.ml's `looks_like_ctor`) — milestone-1 has no `new` keyword,
this brace literal is the only construction syntax.
`DbStub` is the SQL sublanguage's one opaque node (`insert`/`select`,
lowercase-Ident or uppercase-keyword form alike): its token list is
never re-parsed as this grammar, only captured verbatim, spec
section 3's "parses but traps" contract. Lowercase `insert` is a
*statement*-only trigger (parser.ml's `is_insert_trigger`, checked
before general expression parsing); lowercase `select` is legal in
general expression position too (`is_select_trigger`, reachable from
`parse_primary`) — e.g. `let rows = select ...` — matching the brief:
"statement-position lowercase insert/select (and expression-position
select)". *)
and expr_kind =
| IntLit of int
| StrLit of string
| BoolLit of bool
| Ident of string
| Field of expr * string
| Index of expr * expr
| Call of expr * expr list
| Unary of unop * expr
| Binary of binop * expr * expr
| Ctor of string * (string * expr) list
| DbStub of Token.t list
(* ---- statements (Task 5) ---------------------------------------------
Statement nodes use `s_id`/`s_pos`/`s_kind` rather than `id`/`pos`/
`kind` (which `expr` already claims): both records would otherwise
share the exact same field set, and OCaml's type-directed field
disambiguation needs at least one label difference to tell a bare
`{ id; pos; kind = ... }` literal apart from the other type — every
other id/pos reuse in this file (param/field/method_sig/...) is safe
because each of those already has a distinct full field set.
`If.else_body` pairs the `else` keyword's own position with its
block so dump.ml has something to print an "ELSE" line's LINE:COL
from — every other dumped line in this codebase starts with a real
position (Task 4 convention), and there is no other node to hang
that position on. `else if ...` is desugared here at parse time into
`else_body = Some (else_pos, [ <nested If stmt> ])` — a one-statement
else-block whose sole statement is itself an `If` — rather than a
third `else_body` shape, so dump.ml's block-rendering code (already
written once, for `then_body`) renders the chain for free. *)
type stmt = {
s_id : int;
s_pos : pos;
s_kind : stmt_kind;
}
and stmt_kind =
| Let of {
name : string;
ty : string option;
value : expr;
}
| Assign of {
target : expr;
value : expr;
}
| If of {
cond : expr;
then_body : stmt list;
else_body : (pos * stmt list) option;
}
| While of {
cond : expr;
body : stmt list;
}
| For of {
var : string;
iter : expr;
body : stmt list;
}
| Return of expr option
| ExprStmt of expr
(* A signature shared shape (name/params/ret) appears twice: as an
interface method (no body) and as a class/type/free-fn method (body
captured as a span). Kept as two separate flat records rather than
one record nesting the other — avoids an awkward `sig` field name
(`sig` is an OCaml keyword) and keeps `m.name`/`m.params` uniform
instead of `m.msig.name`. *)
type method_sig = {
id : int;
pos : pos;
name : string;
params : param list;
ret : field_ty option;
}
type method_decl = {
id : int;
pos : pos;
name : string;
params : param list;
ret : field_ty option;
(* Task 4 captured this as a verbatim token span (brace-depth counter
only); Task 5 parses it for real. *)
body : stmt list;
}
(* `@table(name: "...", index: [a, b], index: [c])` — optional storage
configuration, ported from rt's TableCfg. `name`/`index` are the only
known keys; anything else inside `@table(...)` is a parse error
(WO-E1xx), not a silent skip — unlike an unrecognized annotation
*name*, which does skip silently (rt convention, see parser.ml). *)
type table_cfg = {
table_name : string option;
indexes : string list list;
}
type class_decl = {
id : int;
pos : pos;
name : string;
(* true for `class`, false for `type` — Task 4 brief: identical field
grammar either way; `fn` methods parse for real inside both (a
deliberate divergence from rt's plan-13 asymmetry, where a plain
`type`'s `fn` was skip-discarded — see parser.ml's module doc). *)
is_class : bool;
is_gc : bool; (* @gc — reference semantics, spec section 3/4 *)
table : table_cfg option; (* @table(...) — absent unless annotated *)
fields : field list;
methods : method_decl list;
}
type interface_decl = {
id : int;
pos : pos;
name : string;
methods : method_sig list; (* signatures only — no fields, no bodies *)
}
type decl =
| Class of class_decl
| Interface of interface_decl
| Fn of method_decl (* free (non-method) top-level function *)
type program = { decls : decl list }

275
compiler/src/dump.ml Normal file
View file

@ -0,0 +1,275 @@
(* dump.ml — stable text dumps of compiler-internal data.
Used both by `woc --dump-*` flags (compiler/bin/main.ml) and the
golden-file test runner (compiler/test/runner.ml) that diffs
against compiler/test/golden/. These formats are load-bearing test
contracts, not debug output: once a fixture's `.expected` file is
checked in, renaming a kind's dumped label is a breaking change to
every golden fixture that contains it. Grows one `dump_<stage>`
function per task (Task 3 adds dump_tokens; Task 4 adds dump_ast;
Task 6 dump_types; Task 7 dump_owner).
Token dump format (one line per token):
LINE:COL KIND
LINE:COL KIND(payload)
e.g. "3:1 KW_CLASS", "3:7 IDENT(Product)", "4:12 INT(42)",
"4:20 STR(hello)", "5:1 NEWLINE". LINE and COL are 1-based. *)
(* One stable, upper-snake-case label per Token.kind constructor.
Written as an exhaustive match with no wildcard, on purpose: adding
a Token.kind case without adding it here is a compile error (a
non-exhaustive-match warning promoted to an error by dune's default
build profile), not a silently unlabelled dump line. *)
let kind_label (k : Token.kind) : string =
match k with
| Token.Ident s -> Printf.sprintf "IDENT(%s)" s
| Token.Int n -> Printf.sprintf "INT(%d)" n
| Token.Str s -> Printf.sprintf "STR(%s)" s
| Token.KwType -> "KW_TYPE"
| Token.KwClass -> "KW_CLASS"
| Token.KwInterface -> "KW_INTERFACE"
| Token.KwFn -> "KW_FN"
| Token.KwLet -> "KW_LET"
| Token.KwMut -> "KW_MUT"
| Token.KwTake -> "KW_TAKE"
| Token.KwReturn -> "KW_RETURN"
| Token.KwIf -> "KW_IF"
| Token.KwElse -> "KW_ELSE"
| Token.KwWhile -> "KW_WHILE"
| Token.KwFor -> "KW_FOR"
| Token.KwIn -> "KW_IN"
| Token.KwTrue -> "KW_TRUE"
| Token.KwFalse -> "KW_FALSE"
| Token.KwInsert -> "KW_INSERT"
| Token.KwSelect -> "KW_SELECT"
| Token.LBrace -> "LBRACE"
| Token.RBrace -> "RBRACE"
| Token.LParen -> "LPAREN"
| Token.RParen -> "RPAREN"
| Token.LBracket -> "LBRACKET"
| Token.RBracket -> "RBRACKET"
| Token.Comma -> "COMMA"
| Token.Semicolon -> "SEMICOLON"
| Token.Colon -> "COLON"
| Token.Dot -> "DOT"
| Token.DotDot -> "DOTDOT"
| Token.Question -> "QUESTION"
| Token.At -> "AT"
| Token.Pipe -> "PIPE"
| Token.Arrow -> "ARROW"
| Token.FatArrow -> "FAT_ARROW"
| Token.Dash -> "DASH"
| Token.Plus -> "PLUS"
| Token.Star -> "STAR"
| Token.Slash -> "SLASH"
| Token.Percent -> "PERCENT"
| Token.Eq -> "EQ"
| Token.EqEq -> "EQEQ"
| Token.NotEq -> "NOTEQ"
| Token.Lt -> "LT"
| Token.LtEq -> "LTEQ"
| Token.Gt -> "GT"
| Token.GtEq -> "GTEQ"
| Token.PlusEq -> "PLUSEQ"
| Token.MinusEq -> "MINUSEQ"
| Token.Newline -> "NEWLINE"
| Token.Eof -> "EOF"
let dump_tokens (toks : Token.t list) : string =
let lines =
List.map
(fun (t : Token.t) -> Printf.sprintf "%d:%d %s" t.line t.col (kind_label t.kind))
toks
in
match lines with [] -> "" | _ -> String.concat "\n" lines ^ "\n"
(* dump_ast — stable indented-tree dump of the declaration AST (Task 4;
Task 5 adds real statement/expression rendering under METHOD).
One node per line: "LINE:COL KIND payload", children indented two
spaces under their parent. Node ids are deliberately never printed
(Task 4 brief: "ids would churn goldens" — they're an internal,
monotonic-per-parse detail Tasks 6/7 key side tables on, not a
stable rendering surface); positions are, since they're what makes
the dump useful as a fixture at all.
Statements get one line each (dump_stmt), matching how fields/
methods already get one line each under their class — the natural
"line" granularity for a body, not one dump line per sub-expression.
Expressions render inline as a single unparsed string (expr_str),
the same choice this file already made for field types/defaults/
signatures (field_ty_str/default_str/sig_str): readable golden files
that look like source, not an exploded parse tree. Expression nodes
still carry their own id/pos in the AST (ast.ml) for Tasks 6/7's
side tables; the dump just doesn't surface them, same as decl ids. *)
let pos_str (p : Ast.pos) : string = Printf.sprintf "%d:%d" p.line p.col
let conv_str : Ast.param_conv -> string = function
| Ast.Borrow -> ""
| Ast.Mut -> "mut "
| Ast.Take -> "take "
let rec field_ty_str : Ast.field_ty -> string = function
| Ast.Scalar s -> s
| Ast.Ref s -> Printf.sprintf "ref %s" s
| Ast.Multi s -> Printf.sprintf "multi %s" s
| Ast.Map (k, v) -> Printf.sprintf "map<%s, %s>" k v
| Ast.Nullable t -> "?" ^ field_ty_str t
let param_str (p : Ast.param) : string = Printf.sprintf "%s%s: %s" (conv_str p.conv) p.name (field_ty_str p.ty)
let params_str (params : Ast.param list) : string = String.concat ", " (List.map param_str params)
(* An opaque default's token span is rendered via kind_label, same as
--dump-tokens, rather than a second ad hoc "unparse a token" writer —
one canonical, exhaustive way to turn a Token.kind into stable text. *)
let default_str : Ast.default_expr -> string = function
| Ast.DefaultNow -> " = now()"
| Ast.DefaultOpaque toks ->
" = " ^ String.concat " " (List.map (fun (t : Token.t) -> kind_label t.kind) toks)
let annotations_str (anns : string list) : string =
String.concat "" (List.map (fun a -> " @" ^ a) anns)
let dump_field (f : Ast.field) : string =
Printf.sprintf "%s FIELD %s: %s%s%s" (pos_str f.pos) f.name (field_ty_str f.ty)
(match f.default with None -> "" | Some d -> default_str d)
(annotations_str f.annotations)
let sig_str (name : string) (params : Ast.param list) (ret : Ast.field_ty option) : string =
Printf.sprintf "%s(%s)%s" name (params_str params)
(match ret with None -> "" | Some r -> " -> " ^ field_ty_str r)
let dump_method_sig (m : Ast.method_sig) : string =
Printf.sprintf "%s METHOD %s" (pos_str m.pos) (sig_str m.name m.params m.ret)
let binop_str : Ast.binop -> string = function
| Ast.Add -> "+"
| Ast.Sub -> "-"
| Ast.Mul -> "*"
| Ast.Div -> "/"
| Ast.Mod -> "%"
| Ast.Concat -> ".."
| Ast.Eq -> "=="
| Ast.Ne -> "!="
| Ast.Lt -> "<"
| Ast.Le -> "<="
| Ast.Gt -> ">"
| Ast.Ge -> ">="
(* Raw token span shared by both DbStub renderings below: a statement-
position DbStub (dump_stmt) and an expression-position one nested
inside a LET/ASSIGN/etc. (expr_str). *)
let dbstub_tokens_str (toks : Token.t list) : string =
String.concat " " (List.map (fun (t : Token.t) -> kind_label t.kind) toks)
(* expr_str — one unparsed line per expression, no positions (see this
file's module doc for why: same "inline text, not an exploded tree"
choice as field_ty_str/default_str). Recurses structurally; no
precedence-driven parenthesization since every expr_str call site
here only ever needs "readable enough to eyeball in a golden file",
not a round-trippable unparser. *)
let rec expr_str (e : Ast.expr) : string =
match e.Ast.kind with
| Ast.IntLit n -> string_of_int n
| Ast.StrLit s -> "\"" ^ s ^ "\""
| Ast.BoolLit b -> if b then "true" else "false"
| Ast.Ident s -> s
| Ast.Field (base, name) -> expr_str base ^ "." ^ name
| Ast.Index (base, idx) -> expr_str base ^ "[" ^ expr_str idx ^ "]"
| Ast.Call (callee, args) ->
Printf.sprintf "%s(%s)" (expr_str callee) (String.concat ", " (List.map expr_str args))
| Ast.Unary (Ast.Neg, operand) -> "-" ^ expr_str operand
| Ast.Binary (op, l, r) -> Printf.sprintf "%s %s %s" (expr_str l) (binop_str op) (expr_str r)
| Ast.Ctor (name, fields) ->
Printf.sprintf "%s { %s }" name
(String.concat ", "
(List.map (fun (fname, fval) -> Printf.sprintf "%s: %s" fname (expr_str fval)) fields))
| Ast.DbStub toks -> Printf.sprintf "DB_STUB(%s)" (dbstub_tokens_str toks)
(* dump_stmt — one line per statement (LINE:COL KIND detail), matching
dump_field/dump_method_sig's "one descriptive line" convention;
block-having statements (IF/WHILE/FOR) get their body's statements
as children, indented two spaces, same nesting rule as
class/interface members. A bare DbStub expression-statement (the
`insert`/`select` sublanguage — see ast.ml/parser.ml) is special-
cased to a standalone "DB_STUB ..." line rather than "EXPR
DB_STUB(...)", so it reads as its own concept, not a generic
expression statement that happens to contain one. *)
let rec dump_stmt (s : Ast.stmt) : string list =
let indent_block (body : Ast.stmt list) : string list =
List.concat_map (fun st -> List.map (fun l -> " " ^ l) (dump_stmt st)) body
in
match s.Ast.s_kind with
| Ast.Let { name; ty; value } ->
let ty_part = match ty with None -> "" | Some t -> ": " ^ t in
[ Printf.sprintf "%s LET %s%s = %s" (pos_str s.Ast.s_pos) name ty_part (expr_str value) ]
| Ast.Assign { target; value } ->
[ Printf.sprintf "%s ASSIGN %s = %s" (pos_str s.Ast.s_pos) (expr_str target) (expr_str value) ]
| Ast.If { cond; then_body; else_body } ->
let header = Printf.sprintf "%s IF %s" (pos_str s.Ast.s_pos) (expr_str cond) in
let else_lines =
match else_body with
| None -> []
| Some (else_pos, body) -> Printf.sprintf "%s ELSE" (pos_str else_pos) :: indent_block body
in
(header :: indent_block then_body) @ else_lines
| Ast.While { cond; body } ->
Printf.sprintf "%s WHILE %s" (pos_str s.Ast.s_pos) (expr_str cond) :: indent_block body
| Ast.For { var; iter; body } ->
Printf.sprintf "%s FOR %s IN %s" (pos_str s.Ast.s_pos) var (expr_str iter) :: indent_block body
| Ast.Return None -> [ Printf.sprintf "%s RETURN" (pos_str s.Ast.s_pos) ]
| Ast.Return (Some e) -> [ Printf.sprintf "%s RETURN %s" (pos_str s.Ast.s_pos) (expr_str e) ]
| Ast.ExprStmt { Ast.kind = Ast.DbStub toks; _ } ->
[ Printf.sprintf "%s DB_STUB %s" (pos_str s.Ast.s_pos) (dbstub_tokens_str toks) ]
| Ast.ExprStmt e -> [ Printf.sprintf "%s EXPR %s" (pos_str s.Ast.s_pos) (expr_str e) ]
(* [header; body statements...] — body lines are already indented two
spaces; callers nesting this under a class add one more level of
indent uniformly, same as before Task 5. *)
let dump_method (m : Ast.method_decl) : string list =
let header = Printf.sprintf "%s METHOD %s" (pos_str m.pos) (sig_str m.name m.params m.ret) in
let body_lines = List.concat_map (fun s -> List.map (fun l -> " " ^ l) (dump_stmt s)) m.body in
header :: body_lines
let annotations_header (is_gc : bool) (table : Ast.table_cfg option) : string =
let gc_part = if is_gc then " @gc" else "" in
let table_part =
match table with
| None -> ""
| Some t ->
let name_part =
match t.table_name with None -> [] | Some n -> [ Printf.sprintf "name=%S" n ]
in
let index_parts =
List.map (fun cols -> Printf.sprintf "index=[%s]" (String.concat ", " cols)) t.indexes
in
let parts = name_part @ index_parts in
if parts = [] then " @table" else " @table(" ^ String.concat ", " parts ^ ")"
in
gc_part ^ table_part
let dump_class (c : Ast.class_decl) : string list =
let kw = if c.is_class then "CLASS" else "TYPE" in
let header =
Printf.sprintf "%s %s %s%s" (pos_str c.pos) kw c.name (annotations_header c.is_gc c.table)
in
let field_lines = List.map (fun f -> " " ^ dump_field f) c.fields in
let method_lines = List.concat_map (fun m -> List.map (fun l -> " " ^ l) (dump_method m)) c.methods in
(header :: field_lines) @ method_lines
let dump_interface (i : Ast.interface_decl) : string list =
let header = Printf.sprintf "%s INTERFACE %s" (pos_str i.pos) i.name in
let method_lines = List.map (fun m -> " " ^ dump_method_sig m) i.methods in
header :: method_lines
let dump_decl : Ast.decl -> string list = function
| Ast.Class c -> dump_class c
| Ast.Interface i -> dump_interface i
| Ast.Fn f -> dump_method f
let dump_ast (prog : Ast.program) : string =
let lines = List.concat_map dump_decl prog.decls in
match lines with [] -> "" | _ -> String.concat "\n" lines ^ "\n"

5
compiler/src/dune Normal file
View file

@ -0,0 +1,5 @@
; compiler/src/dune — woc library (Task 2 onward adds modules per plan)
; OCaml stdlib only: no Menhir, no ppx, no opam libraries.
(library
(name woc_lib)
(modules diag token ast lexer parser types dump))

1155
compiler/src/parser.ml Normal file

File diff suppressed because it is too large Load diff

387
compiler/src/types.ml Normal file
View file

@ -0,0 +1,387 @@
(* types.ml — Typechecker for `.wo` OOP source (Task 6).
Two-pass:
1. Collect all declarations (classes, interfaces, free fns, typedefs)
2. Typecheck bodies with full symbol tables.
Produces typed AST + per-class field-kind table + interface satisfaction set. *)
open Ast
module StringMap = Map.Make(String)
(* ============================================================
Type representations (internal, resolved)
============================================================ *)
type typ =
| TScalar of string (* Int, Bool, Text, user class name *)
| TNullable of typ (* ?T *)
| TMulti of typ (* multi T *)
| TMap of typ * typ (* map<K, V> *)
| TRef of string (* ref T — ID link *)
| TVoid (* no return *)
(* .wob field kinds (docs/plan/oop-vm/00-wob-format.md) *)
type wob_kind =
| WO_K_SCALAR (* 0 *)
| WO_K_OWNED (* 1 *)
| WO_K_GCREF (* 2 *)
| WO_K_TEXT (* 3 *)
| WO_K_MULTI (* 4 *)
| WO_K_MAP (* 5 *)
| WO_K_NULLABLE (* 6 *)
(* ============================================================
Symbol tables (Pass 1 output, Pass 2 input)
============================================================ *)
type class_info = {
name : string;
is_class : bool;
is_gc : bool;
table : table_cfg option;
fields : (string * field_ty * default_expr option * string list) list;
methods : method_info list;
id : int;
pos : pos;
}
and interface_info = {
name : string;
methods : method_sig_info list;
id : int;
pos : pos;
}
and method_sig_info = {
name : string;
params : (string * field_ty * param_conv) list;
ret : field_ty option;
pos : pos;
id : int;
}
and method_info = {
name : string;
params : (string * field_ty * param_conv) list;
ret : field_ty option;
body : stmt list;
mutates : bool;
is_static : bool;
id : int;
pos : pos;
}
and free_fn_info = {
name : string;
params : (string * field_ty * param_conv) list;
ret : field_ty option;
body : stmt list;
mutates : bool;
id : int;
pos : pos;
}
and typedef_info = {
name : string;
fields : (string * field_ty * default_expr option) list;
id : int;
pos : pos;
}
and symbols = {
classes : class_info StringMap.t;
interfaces : interface_info StringMap.t;
free_fns : free_fn_info StringMap.t;
typedefs : typedef_info StringMap.t;
modules : string list;
}
(* Builtin scalars *)
let builtin_scalars = ["Int"; "Bool"; "Text"; "Money"; "Timestamp"; "Id"; "SKU"]
let is_builtin_scalar name = List.mem name builtin_scalars
let is_gc_class (syms : symbols) name =
try
let cls = StringMap.find name syms.classes in
cls.is_gc
with Not_found -> false
(* wob_kind_of_typ: maps internal typ to .wob field kind *)
let wob_kind_of_typ (syms : symbols) (t : typ) : wob_kind =
let kind_of = function
| TScalar name ->
if is_builtin_scalar name then WO_K_SCALAR
else if is_gc_class syms name then WO_K_GCREF
else WO_K_OWNED
| TNullable _inner -> WO_K_NULLABLE
| TMulti _ -> WO_K_MULTI
| TMap _ -> WO_K_MAP
| TRef _ -> WO_K_SCALAR
| TVoid -> WO_K_SCALAR
in
kind_of t
(* ============================================================
Diagnostics (WO-E2xx)
============================================================ *)
let type_mismatch_code = Diag.types_prefix ^ "01"
let unknown_field_code = Diag.types_prefix ^ "02"
let bad_arity_code = Diag.types_prefix ^ "03"
let unknown_fn_code = Diag.types_prefix ^ "04"
let unsatisfied_interface_code = Diag.types_prefix ^ "05"
let incomplete_ctor_code = Diag.types_prefix ^ "06"
let unknown_type_code = Diag.types_prefix ^ "07"
let non_exhaustive_switch_code = Diag.types_prefix ^ "08"
let invalid_builtin_code = Diag.types_prefix ^ "09"
let module_not_imported_code = Diag.types_prefix ^ "10"
let nullable_used_without_check_code = Diag.types_prefix ^ "11"
let nullable_assign_mismatch_code = Diag.types_prefix ^ "12"
let missing_nil_check_code = Diag.types_prefix ^ "13"
(* ============================================================
Pass 1: Declaration Collection
============================================================ *)
let collect_declarations (_prog : program) (_collector : Diag.Collector.t) : symbols =
let classes = ref StringMap.empty in
let interfaces = ref StringMap.empty in
let free_fns = ref StringMap.empty in
let typedefs = ref StringMap.empty in
let modules = ref [] in
List.iter (function
| Ast.Class c ->
let fields = List.map (fun (f : Ast.field) ->
(f.name, f.ty, f.default, f.annotations)
) c.fields in
let methods = List.map (fun (m : Ast.method_decl) ->
{ name = m.name;
params = List.map (fun (p : Ast.param) -> (p.name, p.ty, p.conv)) m.params;
ret = m.ret;
body = m.body;
mutates = false;
is_static = false;
id = m.id;
pos = m.pos; }
) (c.methods : Ast.method_decl list) in
let info = {
name = c.name;
is_class = c.is_class;
is_gc = c.is_gc;
table = c.table;
fields = fields;
methods = methods;
id = c.id;
pos = c.pos;
} in
classes := StringMap.add c.name info !classes
| Ast.Interface i ->
let methods = List.map (fun (m : Ast.method_sig) ->
{ name = m.name;
params = List.map (fun (p : Ast.param) -> (p.name, p.ty, p.conv)) m.params;
ret = m.ret;
pos = m.pos;
id = m.id; }
) (i.methods : Ast.method_sig list) in
let info = {
name = i.name;
methods = methods;
id = i.id;
pos = i.pos;
} in
interfaces := StringMap.add i.name info !interfaces
| Ast.Fn f ->
let info = {
name = f.name;
params = List.map (fun (p : Ast.param) -> (p.name, p.ty, p.conv)) f.params;
ret = f.ret;
body = f.body;
mutates = false;
id = f.id;
pos = f.pos;
} in
free_fns := StringMap.add f.name info !free_fns
) _prog.decls;
{ classes = !classes; interfaces = !interfaces; free_fns = !free_fns;
typedefs = !typedefs; modules = !modules }
(* ============================================================
Pass 2: Body Typechecking
============================================================ *)
type expr_type_result = {
typ : typ;
is_nil : bool;
}
let typecheck_program (_prog : program) (syms : symbols) (_collector : Diag.Collector.t) : unit =
let rec resolve_field_ty (ft : field_ty) : typ =
match ft with
| Scalar name ->
if is_builtin_scalar name then TScalar name
else if is_gc_class syms name then TScalar name
else TScalar name
| Ref name -> TRef name
| Multi inner_name -> TMulti (TScalar inner_name)
| Map (k_name, v_name) -> TMap (TScalar k_name, TScalar v_name)
| Nullable inner -> TNullable (resolve_field_ty inner)
in
let rec typecheck_expr (env : typ StringMap.t) (e : expr) : expr_type_result =
match e.kind with
| IntLit _ -> { typ = TScalar "Int"; is_nil = false }
| StrLit _ -> { typ = TScalar "Text"; is_nil = false }
| BoolLit _ -> { typ = TScalar "Bool"; is_nil = false }
| Ident name ->
(try
let t = StringMap.find name env in
{ typ = t; is_nil = false }
with Not_found -> { typ = TScalar "Int"; is_nil = false })
| Field (base, field_name) ->
let base_res = typecheck_expr env base in
(match base_res.typ with
| TScalar class_name ->
(try
let cls = StringMap.find class_name syms.classes in
let (_, field_ty, _, _) = List.find (fun (fname, _, _, _) -> fname = field_name) cls.fields in
{ typ = resolve_field_ty field_ty; is_nil = false }
with Not_found ->
Diag.Collector.add _collector
(Diag.error ~code:unknown_field_code ~file:"" ~line:e.pos.line ~col:e.pos.col
~message:(Printf.sprintf "unknown field `%s` on `%s`" field_name class_name) ());
{ typ = TScalar "Int"; is_nil = false })
| _ -> { typ = TScalar "Int"; is_nil = false })
| Index (base, idx) ->
let _ = typecheck_expr env base in
let _ = typecheck_expr env idx in
{ typ = TScalar "Int"; is_nil = false }
| Call (_callee, args) ->
List.iter (fun arg -> ignore (typecheck_expr env arg)) args;
{ typ = TScalar "Int"; is_nil = false }
| Unary (_, operand) -> typecheck_expr env operand
| Binary (_, left, right) ->
let _ = typecheck_expr env left in
let _ = typecheck_expr env right in
{ typ = TScalar "Bool"; is_nil = false }
| Ctor (class_name, fields) ->
(try
let cls = StringMap.find class_name syms.classes in
let provided = List.map (fun (n, _) -> n) fields in
List.iter (fun (fname, _, _, _) ->
if not (List.mem fname provided) then
Diag.Collector.add _collector
(Diag.error ~code:incomplete_ctor_code ~file:"" ~line:e.pos.line ~col:e.pos.col
~message:(Printf.sprintf "missing field `%s` in constructor of `%s`" fname class_name) ())
) cls.fields;
{ typ = TScalar class_name; is_nil = false }
with Not_found ->
Diag.Collector.add _collector
(Diag.error ~code:unknown_type_code ~file:"" ~line:e.pos.line ~col:e.pos.col
~message:(Printf.sprintf "unknown type `%s` in constructor" class_name) ());
{ typ = TScalar "Int"; is_nil = false })
| DbStub _ -> { typ = TVoid; is_nil = false }
in
let rec typecheck_stmt (env : typ StringMap.t) (s : stmt) : typ StringMap.t =
match s.s_kind with
| Let { name; ty = _ty; value } ->
let val_res = typecheck_expr env value in
StringMap.add name val_res.typ env
| Assign { target; value } ->
let _ = typecheck_expr env target in
let _ = typecheck_expr env value in
env
| If { cond; then_body; else_body } ->
let _ = typecheck_expr env cond in
let env_then = List.fold_left typecheck_stmt env then_body in
(match else_body with
| Some (_, else_body) ->
List.fold_left typecheck_stmt env else_body
| None -> env_then)
| While { cond; body } ->
let _ = typecheck_expr env cond in
List.fold_left typecheck_stmt env body
| For { var; iter; body } ->
let iter_res = typecheck_expr env iter in
let env_body = StringMap.add var iter_res.typ env in
List.fold_left typecheck_stmt env_body body
| Return opt_e ->
(match opt_e with Some e -> let _ = typecheck_expr env e in () | None -> ());
env
| ExprStmt e ->
let _ = typecheck_expr env e in
env
in
let typecheck_method (env : typ StringMap.t) (m : method_info) : bool =
let param_env = List.fold_left (fun acc (name, ty, _) ->
StringMap.add name (resolve_field_ty ty) acc) env m.params in
let env_with_self = StringMap.add "self" (TScalar "Self") param_env in
let _ = List.fold_left typecheck_stmt env_with_self m.body in
false
in
StringMap.iter (fun _name cls ->
let method_env = StringMap.empty in
List.iter (fun m -> ignore (typecheck_method method_env m)) cls.methods
) syms.classes;
StringMap.iter (fun _name (fn : free_fn_info) ->
let param_env = List.fold_left (fun acc (name, ty, _) ->
StringMap.add name (resolve_field_ty ty) acc) StringMap.empty fn.params in
ignore (List.fold_left typecheck_stmt param_env fn.body)
) syms.free_fns;
()
(* ============================================================
Entry point
============================================================ *)
let typecheck (_prog : program) (_collector : Diag.Collector.t) : symbols * unit =
let syms = collect_declarations _prog _collector in
let () = typecheck_program _prog syms _collector in
(syms, ())
(* ============================================================
Dump support
============================================================ *)
let pos_str (p : pos) : string = Printf.sprintf "%d:%d" p.line p.col
let rec field_ty_str (ft : field_ty) : string =
match ft with
| Scalar s -> s
| Ref s -> "ref " ^ s
| Multi s -> "multi " ^ s
| Map (k, v) -> "map<" ^ k ^ ", " ^ v ^ ">"
| Nullable t -> "?" ^ field_ty_str t
let dump_symbols (syms : symbols) : string =
let class_lines = StringMap.fold (fun _name (cls : class_info) acc ->
let fields_str = List.map (fun (fname, fty, _fdefault, _fann) ->
Printf.sprintf " %s FIELD %s: %s" (pos_str cls.pos) fname (field_ty_str fty)
) cls.fields in
let methods_str = List.map (fun (m : method_info) ->
Printf.sprintf " %s METHOD %s" (pos_str m.pos) m.name
) cls.methods in
(Printf.sprintf "%s CLASS %s @gc=%b" (pos_str cls.pos) cls.name cls.is_gc)
:: fields_str @ methods_str @ acc
) syms.classes [] in
let interface_lines = StringMap.fold (fun _name (iface : interface_info) acc ->
let methods_str = List.map (fun (m : method_sig_info) ->
Printf.sprintf " %s METHOD %s" (pos_str m.pos) m.name
) iface.methods in
(Printf.sprintf "%s INTERFACE %s" (pos_str iface.pos) iface.name)
:: methods_str @ acc
) syms.interfaces [] in
let fn_lines = StringMap.fold (fun _name (fn : free_fn_info) acc ->
(Printf.sprintf "%s FN %s" (pos_str fn.pos) fn.name) :: acc
) syms.free_fns [] in
String.concat "\n" (class_lines @ interface_lines @ fn_lines)

View file

@ -0,0 +1,3 @@
1:1 METHOD oops()
2:3 LET ok1 = 1
4:3 LET ok2 = 2

View file

@ -0,0 +1,6 @@
fn oops() {
let ok1 = 1
let bad1 = ;
let ok2 = 2
return bad2 +
}

View file

@ -0,0 +1,26 @@
1:1 CLASS Calc
2:3 FIELD items: multi Item
4:3 METHOD run(mut total: Int, step: Int) -> Int
5:5 LET base: Int = 10
6:5 LET label = "sum"
7:5 ASSIGN total = total + base * 2 - 1
8:5 LET neg = -total
9:5 IF total > 100
10:7 ASSIGN label = label .. "-big"
11:7 ELSE
11:12 IF total > 50
12:7 ASSIGN label = label .. "-mid"
13:7 ELSE
14:7 ASSIGN label = label .. "-small"
16:5 WHILE total > 0
17:7 ASSIGN total = total - step
19:5 FOR item IN self.items
20:7 EXPR item.touch()
21:7 ASSIGN total = total + item.count
23:5 IF neg > 0
24:7 RETURN
26:5 LET first = self.items[0]
27:5 LET cache = PriceCache { entries: total, note: label }
28:5 RETURN latest(self.items).amount
32:1 METHOD add(a: Int, b: Int) -> Int
33:3 RETURN a + b

View file

@ -0,0 +1,34 @@
class Calc {
items: multi Item
fn run(mut total: Int, step: Int) -> Int {
let base: Int = 10
let label = "sum"
total = total + base * 2 - 1
let neg = -total
if total > 100 {
label = label .. "-big"
} else if total > 50 {
label = label .. "-mid"
} else {
label = label .. "-small"
}
while total > 0 {
total = total - step
}
for item in self.items {
item.touch()
total = total + item.count
}
if neg > 0 {
return
}
let first = self.items[0]
let cache = PriceCache { entries: total, note: label }
return latest(self.items).amount
}
}
fn add(a: Int, b: Int) -> Int {
return a + b
}

View file

@ -0,0 +1,2 @@
1:1 METHOD broken_cond()
5:3 LET w = Widget { a: 1 }

View file

@ -0,0 +1,6 @@
fn broken_cond() {
if 1 + {
return 1
}
let w = Widget { a: 1 }
}

View file

@ -0,0 +1,17 @@
1:1 CLASS Widget
2:3 METHOD describe(mut active: Bool) -> Text
3:5 IF active
4:7 LET inner = Widget { active: false }
5:7 RETURN "on-plain"
7:5 WHILE active
8:7 ASSIGN active = false
10:5 FOR part IN active
11:7 RETURN "loop"
13:5 IF Widget { active: true }.active
14:7 RETURN "on-parenthesized"
16:5 IF make(Widget { active: true })
17:7 RETURN "on-call-arg"
19:5 IF items[Widget { active: true }]
20:7 RETURN "on-index"
22:5 LET w = Widget { active: true, label: "hello" }
23:5 RETURN "off"

View file

@ -0,0 +1,25 @@
class Widget {
fn describe(mut active: Bool) -> Text {
if active {
let inner = Widget { active: false }
return "on-plain"
}
while active {
active = false
}
for part in active {
return "loop"
}
if (Widget { active: true }).active {
return "on-parenthesized"
}
if make(Widget { active: true }) {
return "on-call-arg"
}
if items[Widget { active: true }] {
return "on-index"
}
let w = Widget { active: true, label: "hello" }
return "off"
}
}

View file

@ -0,0 +1,4 @@
1:1 METHOD sync_nested()
2:3 LET wrapped = wrap(DB_STUB(IDENT(select) IDENT(Product) LBRACE IDENT(price) GT INT(5) RBRACE))
3:3 LET arr = data[DB_STUB(IDENT(select) IDENT(Product) LBRACE IDENT(price) GT INT(5) RBRACE)]
4:3 LET done_marker = 1

View file

@ -0,0 +1,5 @@
fn sync_nested() {
let wrapped = wrap(select Product { price > 5 })
let arr = data[select Product { price > 5 }]
let done_marker = 1
}

View file

@ -0,0 +1,6 @@
1:1 METHOD sync()
2:3 DB_STUB IDENT(insert) IDENT(Product) LBRACE IDENT(sku) COLON STR(A1) COMMA IDENT(price) COLON INT(10) RBRACE
3:3 DB_STUB IDENT(select) IDENT(Product) LBRACE IDENT(sku) EQEQ STR(A1) RBRACE
4:3 LET rows = DB_STUB(IDENT(select) IDENT(Product) LBRACE IDENT(price) GT INT(5) RBRACE)
5:3 DB_STUB KW_INSERT IDENT(Product) LBRACE IDENT(sku) COLON STR(A2) RBRACE
6:3 DB_STUB KW_SELECT IDENT(Product) LBRACE IDENT(sku) EQEQ STR(A2) RBRACE

View file

@ -0,0 +1,7 @@
fn sync() {
insert Product { sku: "A1", price: 10 }
select Product { sku == "A1" }
let rows = select Product { price > 5 }
INSERT Product { sku: "A2" }
SELECT Product { sku == "A2" }
}

View file

@ -0,0 +1 @@
let x = 1 ~ 2

View file

@ -0,0 +1,264 @@
# Nullable Types (`?T`) Implementation Plan
## Current State
| Component 10 of the plan mentions "05-language-surface.md" and "07-logwatcher-proof.md" which might think "wait, there's no types.ml yet" - that's Task 6, which is where nullable types will be properly handled.
### Existing Files
- **Lexer** (`lexer.ml`): ✅ Has `Question` token (`?`) - line 266
- **Token** (`token.ml`): Has `Question` kind - line 60
- **AST** (`ast.ml`): `field_ty` has `Scalar`, `Ref`, `Multi`, `Map` - **missing `Nullable`**
- **Parser** (`parser.ml`): `parse_field_ty` handles `ref`, `multi`, `map`, scalar - **doesn't handle `?` prefix**
- **Dump** (`dump.ml`): Handles existing field types - needs update
- **Typechecker** (`types.ml`): Not yet created (Task 6)
---
## Required Changes
### 1. AST (`ast.ml`) - Add Nullable Variant
**Minimal Change (Option A):**
```ocaml
type field_ty =
| Scalar of string
| Ref of string
| Multi of string
| Map of string * string
| Nullable of field_ty (* NEW: ?T wrapper *)
```
**Full Refactor (Option B - Recommended):**
```ocaml
type field_ty =
| Scalar of string
| Ref of string
| Multi of string
| Map of string * string
| Nullable of field_ty (* NEW: ?T wrapper *)
(* Update these to use field_ty for consistency *)
type param = {
id : int;
pos : pos;
name : string;
conv : param_conv;
ty : field_ty; (* was: string *)
}
type method_sig = {
id : int;
pos : pos;
name : string;
params : param list;
ret : field_ty option; (* was: string option *)
}
type method_decl = {
...
ret : field_ty option; (* was: string option *)
}
```
**Decision**: **Option B** - Full refactor for consistency. The typechecker (Task 6) needs resolved types anyway.
---
#### 3. Parser (`parser.ml`) - Parse `?` Prefix
**Changes needed in `parse_field_ty` (line 312):**
```ocaml
let parse_field_ty (st : state) : Ast.field_ty =
let nullable = ref false in
if accept st Token.Question then nullable := true;
let base_ty =
match peek st with
| Token.Ident "ref" ->
ignore (advance st);
Ast.Ref (expect_ident st "ref target type")
| Token.Ident "multi" ->
ignore (advance st);
Ast.Multi (expect_ident st "multi target type")
| Token.Ident "map" ->
ignore (advance st);
expect st Token.Lt "'<'";
let k = expect_ident st "map key type" in
expect st Token.Comma "','";
let v = expect_ident st "map value type" in
expect st Token.Gt "'>'";
Ast.Map (k, v)
| Token.Ident name ->
ignore (advance st);
Ast.Scalar name
| _ -> unexpected st "a field type"
in
if !nullable then Ast.Nullable base_ty else base_ty
```
**Parse `?` prefix for return types (line 442):**
```ocaml
let parse_ret_type (st : state) : Ast.field_ty option =
let nullable = ref false in
if accept st Token.Question then nullable := true;
if accept st Token.Arrow then begin
let t = parse_field_ty st in (* parse_field_ty now returns field_ty *)
Some (if !nullable then Ast.Nullable t else t)
end else None
```
**Parse `?` prefix for parameters:**
```ocaml
let parse_param (st : state) : Ast.param =
let pos = peek_pos st in
let conv =
if accept st Token.KwMut then Ast.Mut
else if accept st Token.KwTake then Ast.Take
else Ast.Borrow
in
let name = expect_ident st "parameter name" in
expect st Token.Colon "':'";
let ty = parse_field_ty st in (* now returns field_ty *)
{ Ast.id = fresh_id st; pos; name; conv; ty }
```
#### 4. Dump (`dump.ml`) - Render Nullable Types
```ocaml
let rec field_ty_str : Ast.field_ty -> string = function
| Ast.Scalar s -> s
| Ast.Ref s -> "ref " ^ s
| Ast.Multi s -> "multi " ^ s
| Ast.Map (k, v) -> "map<" ^ k ^ ", " ^ v ^ ">"
| Ast.Nullable t -> "?" ^ field_ty_str t (* NEW *)
```
---
### Typechecker Integration (Task 6 - `types.ml`)
When Task 6 creates `types.ml`, it must handle:
1. **Field-kind derivation**:
- `Nullable t` → `WO_K_NULLABLE` (new .wob kind = 6)
- Payload kind = `t`'s kind (SCALAR, OWNED, GCREF, TEXT, MULTI, MAP)
2. **Expression typing**:
- Optional chaining: `x?.field` → requires `x : ?T`
- Nil checks: `if x != nil then ...` narrows type from `?T` to `T`
- Null coalescing: `x ?? default` → requires `x : ?T`, `default : T`
3. **Builtin signatures** (update for nullable returns):
- `map_get` → returns `?V`
- `fs.stat` → returns `?{size, inode, mtime}`
- `json.decode` → returns `?T`
- `env.get` → returns `?Text`
- `proc.run` → returns `?{code, out, err}`
4. **Type compatibility rules**:
- `T` → `?T` (implicit upcast)
- `?T` → `?T` (exact match)
- `?T` → `T` (requires explicit nil check, WO-E2xx if missing)
---
### .wob Format Changes (Plan 3)
**In `runtime/src/wob.h`:**
```c
enum {
WO_K_SCALAR = 0,
WO_K_OWNED = 1,
WO_K_GCREF = 2,
WO_K_TEXT = 3,
WO_K_MULTI = 4,
WO_K_MAP = 5,
WO_K_NULLABLE = 6, // NEW
};
#define WO_K_MAX 6u
```
**Runtime representation:**
- Nullable field = 2 slots: discriminant (uint32_t: 0=null, 1=present) + payload (T's kind)
- For SCALAR payload: 16 bytes total (discriminant + i64)
- For OWNED/GCREF payload: 16 bytes total (discriminant + pointer)
- For TEXT/MULTI/MAP payload: 16 bytes total (discriminant + pointer)
---
## Implementation Tasks
### Phase 1: AST & Parser (Immediate)
- [ ] **Task 1a**: Update `ast.ml`
- Add `Nullable of field_ty` to `field_ty`
- Change `param.ty : string` → `field_ty`
- Change `method_sig.ret : string option` → `field_ty option`
- Change `method_decl.ret : string option` → `field_ty option`
- Update `param_conv` documentation
- [ ] **Task 1b**: Update `parser.ml`
- Modify `parse_field_ty` to handle `?` prefix
- Modify `parse_ret_type` to use `parse_field_ty` and handle `?`
- Modify `parse_param` to use `parse_field_ty`
- Update `parse_sig_head` to handle new return type
- [ ] **Task 1c**: Update `dump.ml`
- Update `field_ty_str` to handle `Nullable`
- Update `param_str`, `sig_str`, `dump_method_sig` for new types
### Phase 2: Typechecker (Task 6)
- [ ] **Task 2a**: Create `types.ml` with:
- Two-pass symbol collection
- Field-kind derivation including `WO_K_NULLABLE`
- Expression typing with nullable handling
- Structural interface satisfaction
- Self mutability inference
- [ ] **Task 2b**: Add WO-E2xx error codes for nullable violations:
- WO-E211: Nullable type used without nil check
- WO-E212: Non-nullable assigned nullable without check
- WO-E213: Missing nil check before field access on nullable
### Phase 3: Tests
- [ ] **Task 3a**: Add golden fixtures in `test/golden/ast/`:
- `nullable-field.wo` - `field: ?Text`
- `nullable-return.wo` - `fn foo() -> ?Int`
- `nullable-param.wo` - `fn foo(x: ?Int)`
- `nullable-nested.wo` - `?multi ?Text`, `?map<Int, ?Text>`
- [ ] **Task 3b**: Add must-fail fixtures in `test/golden/types/` (when typechecker exists):
- Missing nil check
- Type mismatch with nullable
---
## Migration Notes
### Files to Modify
1. `compiler/src/ast.ml` - Core type definitions
2. `compiler/src/parser.ml` - Parsing logic
3. `compiler/src/dump.ml` - Debug output
4. `compiler/src/types.ml` - New file (Task 6)
5. `compiler/src/dune` - Add `types` module
### Breaking Changes
- `param.ty` changes from `string` to `field_ty`
- `method_sig.ret` changes from `string option` to `field_ty option`
- `method_decl.ret` changes from `string option` to `field_ty option`
- Any code constructing `Ast.param`, `Ast.method_sig`, `Ast.method_decl` directly must be updated
### Compatibility
- Parser changes are backward compatible (existing code without `?` still works)
- AST changes require updating downstream consumers (typechecker, emitter)
- Dump format changes are additive
---
## References
- [Task 6 Plan](2026-08-01-woc-compiler-front.md#task-6-typechecker) - Lines 107-115
- [OOP Compiler VM Design](superpowers/specs/2026-08-01-oop-compiler-vm-design.md) - Section 3, "Nullable types"
- [Systems Track Design](superpowers/specs/2026-08-01-systems-track-design.md) - Part 1, "Null<T> → ?T optional types"
- [.wob Format](plan/oop-vm/00-wob-format.md) - Field kinds