From f465ad5751df30b9b50b97d0275868e644413204 Mon Sep 17 00:00:00 2001 From: "shoney.arickathil" Date: Tue, 25 Aug 2026 03:58:41 +0200 Subject: [PATCH] feat(lang+wo-html): raw text literals, component layer, MVC samples MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit - lexer: backtick raw text literal — content verbatim, no escape processing, common source margin removed at lex time; `${ }` raw and `{{ }}` auto-escaping holes - `{{ e }}` desugars to `esc(${e})` in parser.ml — a Call on the `esc` in scope, so types/owner/emit/.wob/VM are untouched - WO-E004 unterminated raw literal; WO-E005 newline inside "..." — closes a hole where a missing quote silently ate the rest of the file - wo-html: `Component` interface, `render_all`, `Layout`, README - framework: `ok_html` joins ok_text/ok_json in http/types.wo - site + shop restructured to one-feature-one-module MVC (view + controller per directory, model at the root, bootstrap-only main) - removed the filler `pad: Int` convention — verified unnecessary for plain classes, interface dispatch, containers and actors - corrected recorded claims: gap #1 blocks neither the build nor the layout; a class crosses module lines, only a free fn is scoped - docs/guides/language-surface.md — the full grammar inventory - story 37 landed and moved to done/ Gates: oop-accept MET, oop-e2e 116/0, woc-test 556/0, site 11/0, web-app 46/0, fibers 10/0, db-actor 8/0 Co-Authored-By: Claude Opus 5 (1M context) --- compiler/src/CODE-LOGIC.md | 59 ++++ compiler/src/dump.ml | 1 + compiler/src/emit.ml | 14 +- compiler/src/lexer.ml | 234 +++++++++++++++- compiler/src/parser.ml | 16 +- compiler/src/token.ml | 7 + compiler/test/golden/ast/raw-literal.expected | 2 + compiler/test/golden/ast/raw-literal.wo | 3 + .../test/golden/tokens/raw-literal.expected | 42 +++ compiler/test/golden/tokens/raw-literal.wo | 15 + compiler/test/runner.ml | 122 ++++++++ docs/examples/db-actor/main.wo | 5 +- docs/examples/fibers/main.wo | 3 +- docs/examples/shop/README.md | 112 ++++++-- docs/examples/shop/assets/style.css | 2 +- docs/examples/shop/layout/app.wo | 58 ++-- docs/examples/shop/layout/footer.wo | 2 +- docs/examples/shop/layout/header.wo | 10 +- docs/examples/shop/main.wo | 10 +- .../controller.wo} | 16 +- docs/examples/shop/orders/view.wo | 51 ++-- docs/examples/shop/product_list.controller.wo | 20 -- docs/examples/shop/product_list/controller.wo | 21 ++ docs/examples/shop/product_list/view.wo | 43 +-- .../controller.wo} | 7 +- docs/examples/shop/product_page/view.wo | 35 +-- docs/examples/site/CODE-LOGIC.md | 82 ++++-- docs/examples/site/README.md | 33 +++ docs/examples/site/admin/controller.wo | 33 +++ docs/examples/site/chapter/controller.wo | 23 ++ docs/examples/site/chapter/view.wo | 46 +++ docs/examples/site/content.wo | 35 ++- docs/examples/site/health/controller.wo | 9 + docs/examples/site/home/controller.wo | 19 ++ docs/examples/site/home/view.wo | 53 ++++ docs/examples/site/layout/app.wo | 48 ++++ docs/examples/site/layout/footer.wo | 6 + docs/examples/site/layout/header.wo | 9 + docs/examples/site/main.wo | 174 +----------- docs/examples/site/types.wo | 52 ++++ docs/examples/web-app/main.wo | 35 +-- docs/examples/wo-html/README.md | 102 +++++++ docs/examples/wo-html/html.wo | 105 +++++-- docs/examples/writeonce-framework/app.wo | 2 +- .../writeonce-framework/http/secure.wo | 1 - .../writeonce-framework/http/types.wo | 11 + .../writeonce-framework/router/router.wo | 6 +- docs/guides/language-surface.md | 244 ++++++++++++++++ docs/plan/oop-vm/01-error-catalog.md | 2 + docs/plan/oop-vm/08-builtin-surface.md | 2 + docs/stories/00-status.md | 47 +++- .../language-runtime-database/00-story.md | 2 +- .../done/37-wo-html-components.md | 265 ++++++++++++++++++ .../refine/37-wo-html-components.md | 147 ---------- .../call-reply-disagree/fixture.wo | 4 +- .../raw-literal-esc-missing/fixture.code | 1 + .../raw-literal-esc-missing/fixture.wo | 8 + .../raw-literal-unterminated/fixture.code | 1 + .../raw-literal-unterminated/fixture.wo | 10 + .../compile-fail/send-after-move/fixture.wo | 3 +- .../compile-fail/traced-send/fixture.wo | 3 +- tests/corpus/run/call-dead-trap/fixture.wo | 3 +- .../run/dotdot-continuation/fixture.out | 2 + .../corpus/run/dotdot-continuation/fixture.wo | 13 + tests/corpus/run/mailbox-full-trap/fixture.wo | 3 +- tests/corpus/run/raw-text-literal/fixture.out | 5 + tests/corpus/run/raw-text-literal/fixture.wo | 35 +++ 67 files changed, 2022 insertions(+), 572 deletions(-) create mode 100644 compiler/test/golden/ast/raw-literal.expected create mode 100644 compiler/test/golden/ast/raw-literal.wo create mode 100644 compiler/test/golden/tokens/raw-literal.expected create mode 100644 compiler/test/golden/tokens/raw-literal.wo rename docs/examples/shop/{orders.controller.wo => orders/controller.wo} (77%) delete mode 100644 docs/examples/shop/product_list.controller.wo create mode 100644 docs/examples/shop/product_list/controller.wo rename docs/examples/shop/{product_page.controller.wo => product_page/controller.wo} (74%) create mode 100644 docs/examples/site/admin/controller.wo create mode 100644 docs/examples/site/chapter/controller.wo create mode 100644 docs/examples/site/chapter/view.wo create mode 100644 docs/examples/site/health/controller.wo create mode 100644 docs/examples/site/home/controller.wo create mode 100644 docs/examples/site/home/view.wo create mode 100644 docs/examples/site/layout/app.wo create mode 100644 docs/examples/site/layout/footer.wo create mode 100644 docs/examples/site/layout/header.wo create mode 100644 docs/examples/site/types.wo create mode 100644 docs/examples/wo-html/README.md create mode 100644 docs/guides/language-surface.md create mode 100644 docs/stories/language-runtime-database/done/37-wo-html-components.md delete mode 100644 docs/stories/language-runtime-database/refine/37-wo-html-components.md create mode 100644 tests/corpus/compile-fail/raw-literal-esc-missing/fixture.code create mode 100644 tests/corpus/compile-fail/raw-literal-esc-missing/fixture.wo create mode 100644 tests/corpus/compile-fail/raw-literal-unterminated/fixture.code create mode 100644 tests/corpus/compile-fail/raw-literal-unterminated/fixture.wo create mode 100644 tests/corpus/run/dotdot-continuation/fixture.out create mode 100644 tests/corpus/run/dotdot-continuation/fixture.wo create mode 100644 tests/corpus/run/raw-text-literal/fixture.out create mode 100644 tests/corpus/run/raw-text-literal/fixture.wo diff --git a/compiler/src/CODE-LOGIC.md b/compiler/src/CODE-LOGIC.md index 4f8549e..238780e 100644 --- a/compiler/src/CODE-LOGIC.md +++ b/compiler/src/CODE-LOGIC.md @@ -295,3 +295,62 @@ a keyword, and visibility is name resolution at compile time. The `0x`/`0b` prefix commits only when a real base digit follows, so `0xg` stays `Int 0` + `Ident` — a parse error at its own position, no new lexer diagnostic. `_` separators are consumed only BETWEEN digits. + +## The raw text literal (iteration 37) + +Multi-line markup used to be impossible to write: a statement ends at a +newline, so a page was one `h = h .. "<...>"` statement per line, every +attribute single-quoted to dodge `\"`, and every piece of data wrapped +in a hand-written `esc()` call. Backtick literals replace all three. +Things worth knowing before editing them: + +- **It is a LEXER form, not a node.** A backtick literal emits exactly + the `Token.Str` (no holes) or `Token.InterpStr` (holes) a `"..."` + string emits, so `types.ml`, `owner.ml`, `emit.ml`, the `.wob` format + and the VM are all untouched — nothing downstream can tell the two + spellings apart. That is the whole reason the feature is small. A + design that introduced a `Markup`/`Element` AST variant instead would + have had to teach five files about it. +- **No escape processing at all inside.** Quotes and backslashes are + content, which is the point. The cost is that the form cannot express + a literal backtick, a literal `${`, or a literal `{{` — those are + written by concatenating an ordinary `"..."` string with `..`. One + greppable door beats inventing an escape character for the one form + whose selling point is not having any. (`docs/examples/site/content.wo` + keeps two `code_block` samples as escaped `"..."` strings for exactly + this reason: they contain `\${`.) +- **The margin is stripped at LEX time**, so the constant pool holds the + dedented text and there is no runtime cost. Java's text-block rule: + one newline right after the opening backtick is dropped, the smallest + leading whitespace run across non-blank lines is removed from every + line, and a whitespace-only closing line loses its whitespace but + keeps its newline. A literal with no newline is left alone — eating + the leading spaces of `` ` hi` `` would be a surprise, not a service. + The measuring pass runs over a SHADOW string where each hole is one + non-whitespace sentinel byte, so ` {{ x }}` counts as indent 4 and + as a non-blank line. +- **`{{ e }}` desugars to `esc(${e})`, resolved by ordinary name + lookup.** `desugar_interp` in `parser.ml` builds a `Call` on an + `Ident "esc"` — precisely what a developer wrote by hand before. The + compiler learns nothing about HTML, `esc` stays wo-html's ordinary + `pub fn`, a typo'd field inside the hole is a normal name/type error, + and a locally defined `esc` shadows deliberately (a custom escaper is + a feature). `${ }` inside the same literal stays raw — that is the + greppable door for markup you built yourself. The one place the + desugar leaks: with no `esc` in scope the program fails on a name it + never typed, so `emit.ml`'s WO-E403 message carries a hint for that + one name. +- **`{{` is special ONLY inside a backtick literal.** Inside `"..."` it + is still two braces, so CSS and JS text in existing samples lexes + byte-identically. +- **WO-E005 closed a real hole.** The string scanner's catch-all used to + append a raw newline like any other byte, so a forgotten closing quote + silently swallowed the rest of the file with no diagnostic. Now the + scan stops at the newline WITHOUT consuming it — the `Newline` token + still terminates the statement, so recovery costs one line instead of + the file. The rt-parity silence for a plain unterminated string with + no newline is untouched, and `runner.ml` still pins it. +- **The `..` line continuation stays.** A line ending in `..` still + swallows its newline. Raw literals took over the multi-line-markup job + that motivated it, but it remains the general way to spread a long + concatenation over several lines and has its own corpus fixture. diff --git a/compiler/src/dump.ml b/compiler/src/dump.ml index 95722d8..11ed6e2 100644 --- a/compiler/src/dump.ml +++ b/compiler/src/dump.ml @@ -34,6 +34,7 @@ let kind_label (k : Token.kind) : string = let part_str = function | Token.SText s -> Printf.sprintf "TEXT(%s)" s | Token.SExpr s -> Printf.sprintf "EXPR(%s)" s + | Token.SEsc s -> Printf.sprintf "ESC(%s)" s in Printf.sprintf "INTERP_STR(%s)" (String.concat "," (List.map part_str segs)) | Token.KwType -> "KW_TYPE" diff --git a/compiler/src/emit.ml b/compiler/src/emit.ml index 271f71f..01cc8d3 100644 --- a/compiler/src/emit.ml +++ b/compiler/src/emit.ml @@ -3258,8 +3258,20 @@ and emit_call (p : pctx) (f : fstate) (v : views) ~(dst : int) ?expected (e : As | None -> if is_builtin_name name then emit_builtin p f v ~dst ?expected e name args else begin + (* iteration 37: `{{ e }}` in a raw text literal desugars to + a call to `esc`, so a program using that hole without an + `esc` in scope lands here with a name it never typed. + The caret is already on the literal; this says why. *) + let hint = + if name = "esc" then + " (a `{{ ... }}` hole calls it -- add `use html`, or \ + declare your own `fn esc(t: Text) -> Text`)" + else "" + in err p ~code:cannot_lower_code ~file:f.f_file ~pos:e.pos - ~message:(Printf.sprintf "call to `%s`, which is not a declared fn or a builtin" name); + ~message: + (Printf.sprintf "call to `%s`, which is not a declared fn or a builtin%s" + name hint); put f (ins_abx op_loadk dst (const_int p 0)) end))) | Field (base, mname) -> ( diff --git a/compiler/src/lexer.ml b/compiler/src/lexer.ml index b9782ba..d5c4f72 100644 --- a/compiler/src/lexer.ml +++ b/compiler/src/lexer.ml @@ -55,6 +55,32 @@ let unknown_char_code = Diag.lexing_prefix ^ "01" (* WO-E001 *) let unterminated_escape_code = Diag.lexing_prefix ^ "02" (* WO-E002 *) let directive_code = Diag.lexing_prefix ^ "03" (* WO-E003: #if/#else/#end misuse *) +(* iteration 37, the raw text literal (backtick-delimited, verbatim + content, no backslash escapes). Two codes, because the two shapes + are genuinely different situations: + + WO-E004 — a raw literal that runs off the end of the file. Unlike a + plain "..." string (silent, rt parity, see above), this one IS + reported: multi-line is the raw literal's normal case, so a missing + closing backtick would otherwise swallow every remaining line of the + file with nothing to show for it. Reported at the OPENING backtick, + which is the only position that helps -- EOF tells the reader + nothing about which literal never closed. + + WO-E005 — a raw newline inside a "..." or '...' string. This used to + be accepted silently: the string scanner's catch-all appended the + newline like any other byte, so a forgotten closing quote ate the + rest of the file with no diagnostic at all. Nothing in the repo ever + relied on it (zero of the .wo sources span a line inside quotes) and + the backtick literal is now the spelling for multi-line text, so the + accident becomes an error. The scan stops at the newline WITHOUT + consuming it, so the Newline token is still emitted and the + statement terminates -- one diagnostic, and the next line parses + normally instead of being swallowed. The rt-parity silence for a + plain unterminated string with no newline is untouched. *) +let unterminated_raw_code = Diag.lexing_prefix ^ "04" (* WO-E004 *) +let newline_in_string_code = Diag.lexing_prefix ^ "05" (* WO-E005 *) + (* haxe-parity Task 8: build flags. `woc -D name` fills this before any tokenize call; undefined flags are false. A module-level ref because the compiler is a single-shot process — tests that care set it explicitly @@ -269,6 +295,126 @@ let preprocess (collector : Diag.Collector.t) ~(file : string) go toks; List.rev !out +(* ---- the raw literal's margin rule (iteration 37) ------------------- + + A render() body is written at its method's indentation, but that + indentation is an artifact of the SOURCE, not of the markup -- nobody + wants six leading spaces on every line of the served HTML. So the + common margin is removed here, at lex time: the constant pool holds + the dedented text, no downstream stage ever sees the source + indentation, and the whole rule costs nothing at run time. + + The rule (Java's text blocks, which solved exactly this): + - one newline immediately after the opening backtick is dropped, + so the first markup line can start on its own line; + - the smallest leading run of spaces/tabs across all non-blank + lines is removed from every line (characters counted, tabs NOT + expanded -- mixing them is the author's problem, and expanding + would need a tab width the language does not have); + - a whitespace-only final line (the usual case: the closing + backtick sits on its own line) loses its whitespace but keeps + its newline. + A literal with no newline in it is left completely alone -- there is + no margin to speak of, and silently eating the leading spaces of + ` hi` would be a surprise, not a service. + + Holes do not disturb any of this. A line's indentation is by + definition the run of whitespace at its start, and the only thing + that can split a line across segments is a hole, which ends that run + -- so an indentation run always lives whole inside one SText. The + measuring pass replaces each hole with a single non-whitespace + sentinel byte so that a line that is ` {{ x }}` correctly counts + as indent 4 and as NON-blank. *) + +let is_indent_char c = c = ' ' || c = '\t' + +let segments_shadow (segs : Token.str_part list) : string = + let b = Buffer.create 64 in + List.iter + (function + | Token.SText s -> Buffer.add_string b s + | Token.SExpr _ | Token.SEsc _ -> Buffer.add_char b '\001') + segs; + Buffer.contents b + +let min_indent (shadow : string) : int = + let m = ref max_int in + List.iter + (fun line -> + let n = String.length line in + let i = ref 0 in + while !i < n && is_indent_char line.[!i] do + incr i + done; + (* a blank (or whitespace-only) line never sets the margin *) + if !i < n && !i < !m then m := !i) + (String.split_on_char '\n' shadow); + if !m = max_int then 0 else !m + +let strip_margin (k : int) (segs : Token.str_part list) : Token.str_part list = + if k = 0 then segs + else begin + let at_line_start = ref true in + let one seg = + match seg with + | Token.SExpr _ | Token.SEsc _ -> + at_line_start := false; + seg + | Token.SText s -> + let n = String.length s in + let b = Buffer.create n in + let i = ref 0 in + while !i < n do + if !at_line_start then begin + let dropped = ref 0 in + while !dropped < k && !i < n && is_indent_char s.[!i] do + incr dropped; + incr i + done; + at_line_start := false + end + else begin + let c = s.[!i] in + Buffer.add_char b c; + if c = '\n' then at_line_start := true; + incr i + end + done; + Token.SText (Buffer.contents b) + in + (* fold_left, not List.map: `one` carries state across segments and + List.map's application order is unspecified. *) + List.rev (List.fold_left (fun acc seg -> one seg :: acc) [] segs) + end + +let drop_trailing_margin (segs : Token.str_part list) : Token.str_part list = + match List.rev segs with + | Token.SText s :: rest_rev -> + let n = String.length s in + let i = ref n in + while !i > 0 && is_indent_char s.[!i - 1] do + decr i + done; + (* only a run that directly follows a newline is a closing line *) + if !i < n && !i > 0 && s.[!i - 1] = '\n' then + List.rev (Token.SText (String.sub s 0 !i) :: rest_rev) + else segs + | _ -> segs + +let dedent (segs : Token.str_part list) : Token.str_part list = + let shadow = segments_shadow segs in + if not (String.contains shadow '\n') then segs + else begin + let segs = + match segs with + | Token.SText s :: rest when String.length s > 0 && s.[0] = '\n' -> + Token.SText (String.sub s 1 (String.length s - 1)) :: rest + | _ -> segs + in + let k = min_indent (segments_shadow segs) in + drop_trailing_margin (strip_margin k segs) + end + let tokenize (collector : Diag.Collector.t) ~(file : string) (src : string) : Token.t list = let lx = make src in @@ -279,6 +425,14 @@ let tokenize (collector : Diag.Collector.t) ~(file : string) (src : string) : | { Token.kind = Token.Newline; _ } :: _ -> true | _ -> false in + (* A line ending in `..` continues on the next line — the ONE newline + suppression in the language, so multi-line markup/text builds read + as one expression (the shop template's ask; story 37 rides it). *) + let last_is_dotdot () = + match !out with + | { Token.kind = Token.DotDot; _ } :: _ -> true + | _ -> false + in let report_unknown line col c = Diag.Collector.add collector (Diag.error ~code:unknown_char_code ~file ~line ~col @@ -289,6 +443,18 @@ let tokenize (collector : Diag.Collector.t) ~(file : string) (src : string) : (Diag.error ~code:unterminated_escape_code ~file ~line ~col ~message:"unterminated string escape" ()) in + let report_unterminated_raw line col = + Diag.Collector.add collector + (Diag.error ~code:unterminated_raw_code ~file ~line ~col + ~message:"unterminated raw text literal" ()) + in + let report_newline_in_string line col = + Diag.Collector.add collector + (Diag.error ~code:newline_in_string_code ~file ~line ~col + ~message: + "newline in string literal (use a `...` raw text literal for \ + multi-line text)" ()) + in let running = ref true in while !running do match peek lx with @@ -307,7 +473,8 @@ let tokenize (collector : Diag.Collector.t) ~(file : string) (src : string) : end else if c = '\n' then begin ignore (advance lx); - if not (last_is_newline ()) then emit Token.Newline line col + if not (last_is_newline ()) && not (last_is_dotdot ()) then + emit Token.Newline line col end else if c = ' ' || c = '\t' || c = '\r' then ignore (advance lx) else if c = '"' || c = '\'' then begin @@ -374,6 +541,13 @@ let tokenize (collector : Diag.Collector.t) ~(file : string) (src : string) : | None -> report_unterminated_escape esc_line esc_col; scanning := false) + | Some '\n' -> + (* WO-E005. Deliberately NOT consumed: the outer loop turns + it into the Newline token that terminates the statement, + so recovery is one bad line rather than the rest of the + file. *) + report_newline_in_string lx.line lx.col; + scanning := false | Some other -> ignore (advance lx); Buffer.add_char buf other @@ -384,6 +558,64 @@ let tokenize (collector : Diag.Collector.t) ~(file : string) (src : string) : | [ Token.SText s ] -> emit (Token.Str s) line col | segs -> emit (Token.InterpStr segs) line col) end + else if c = '`' then begin + (* iteration 37: the raw text literal. Everything up to the + closing backtick is content -- newlines included, and with NO + escape processing at all, which is the whole point: markup + carries quotes and backslashes verbatim. A literal backtick + (or a literal `{{`) is written by concatenating an ordinary + "..." string with `..`; that door is one greppable operator, + which beats inventing an escape character for the one form + whose selling point is not having any. + + Two hole forms, and ONLY here -- inside "..." a `{{` is still + two literal braces, so existing CSS/JS text is untouched: + ${ expr } raw, exactly like a "..." string's hole + {{ expr }} HTML-escaped (the parser wraps it in esc()) *) + ignore (advance lx); + let buf = Buffer.create 64 in + let parts = ref [] in + let flush_text () = + parts := Token.SText (Buffer.contents buf) :: !parts; + Buffer.clear buf + in + let scanning = ref true in + while !scanning do + match peek lx with + | None -> + report_unterminated_raw line col; + scanning := false + | Some '`' -> + ignore (advance lx); + scanning := false + | Some '$' when peek_at lx 1 = Some '{' -> + flush_text (); + ignore (advance lx); + ignore (advance lx); + parts := Token.SExpr (read_interp_expr lx) :: !parts + | Some '{' when peek_at lx 1 = Some '{' -> + flush_text (); + ignore (advance lx); + ignore (advance lx); + (* read_interp_expr stops at the first `}` at depth 0 and + consumes it -- the second one closes this hole. Reusing it + means brace depth and nested string literals are already + handled, so `{{ Point{x:1}.x }}` scans correctly. *) + let raw = read_interp_expr lx in + (match peek lx with + | Some '}' -> ignore (advance lx) + | _ -> report_unterminated_raw line col); + parts := Token.SEsc raw :: !parts + | Some other -> + ignore (advance lx); + Buffer.add_char buf other + done; + flush_text (); + (match dedent (List.rev !parts) with + | [] -> emit (Token.Str "") line col + | [ Token.SText s ] -> emit (Token.Str s) line col + | segs -> emit (Token.InterpStr segs) line col) + end else if is_digit c then begin (* iteration 19: one scanner for both numeric worlds. The integer run is scanned into a buffer as well as accumulated, because a fraction diff --git a/compiler/src/parser.ml b/compiler/src/parser.ml index c1ecff0..8734d82 100644 --- a/compiler/src/parser.ml +++ b/compiler/src/parser.ml @@ -1332,6 +1332,19 @@ and parse_primary (st : state) : Ast.expr = and desugar_interp (st : state) (pos : Ast.pos) (segs : Token.str_part list) : Ast.expr = let mk_str s = { Ast.id = fresh_id st; pos; kind = Ast.StrLit s } in let mk_interp inner = { Ast.id = fresh_id st; pos; kind = Ast.Interp inner } in + (* iteration 37: `{{ e }}` in a raw text literal IS `esc(${e})` -- the + desugar builds exactly the call a developer writes by hand today + (docs/examples/shop/**/view.wo used `${esc(...)}` throughout), so + every later stage sees only Call/Interp/StrLit/Concat nodes it + already handles. No new AST variant, no new builtin, no VM change, + and a typo'd field inside the hole is an ordinary name/type error. + `esc` resolves by ordinary lookup (wo-html's `pub fn esc`, in scope + after `use html`); a locally defined `esc` shadows it deliberately + -- a custom escaper is a feature, not a collision. *) + let mk_esc inner = + let callee = { Ast.id = fresh_id st; pos; kind = Ast.Ident "esc" } in + { Ast.id = fresh_id st; pos; kind = Ast.Call (callee, [ mk_interp inner ]) } + in let parse_segment_expr (raw : string) : Ast.expr = let sub_collector = Diag.Collector.create () in let sub_toks = Lexer.tokenize sub_collector ~file:st.file raw in @@ -1362,7 +1375,8 @@ and desugar_interp (st : state) (pos : Ast.pos) (segs : Token.str_part list) : A (function | Token.SText "" -> None | Token.SText s -> Some (mk_str s) - | Token.SExpr raw -> Some (mk_interp (parse_segment_expr raw))) + | Token.SExpr raw -> Some (mk_interp (parse_segment_expr raw)) + | Token.SEsc raw -> Some (mk_esc (parse_segment_expr raw))) segs in match parts with diff --git a/compiler/src/token.ml b/compiler/src/token.ml index 5f71a0c..bf9515b 100644 --- a/compiler/src/token.ml +++ b/compiler/src/token.ml @@ -23,6 +23,13 @@ type str_part = | SText of string (* literal text, escapes already applied *) | SExpr of string (* raw, unlexed source of one `${...}`'s body *) + (* iteration 37: the escaping half of the raw text literal. Same raw, + unlexed payload as SExpr -- what differs is only what the parser + wraps it in: `${...}` desugars to a bare Interp, `{{...}}` to an + `esc(Interp ...)` call. Produced ONLY by a backtick raw literal; + inside a "..." string `{{` stays two literal braces, so CSS and JS + text in existing samples lexes byte-identically. *) + | SEsc of string (* raw, unlexed source of one `{{...}}`'s body *) type kind = (* literals *) diff --git a/compiler/test/golden/ast/raw-literal.expected b/compiler/test/golden/ast/raw-literal.expected new file mode 100644 index 0000000..babebc8 --- /dev/null +++ b/compiler/test/golden/ast/raw-literal.expected @@ -0,0 +1,2 @@ +1:1 METHOD page(name: Text) -> Text + 2:3 RETURN "

" .. INTERP(name) .. esc(INTERP(name)) .. "

" diff --git a/compiler/test/golden/ast/raw-literal.wo b/compiler/test/golden/ast/raw-literal.wo new file mode 100644 index 0000000..2d095f8 --- /dev/null +++ b/compiler/test/golden/ast/raw-literal.wo @@ -0,0 +1,3 @@ +fn page(name: Text) -> Text { + return `

${name}{{ name }}

` +} diff --git a/compiler/test/golden/tokens/raw-literal.expected b/compiler/test/golden/tokens/raw-literal.expected new file mode 100644 index 0000000..163611a --- /dev/null +++ b/compiler/test/golden/tokens/raw-literal.expected @@ -0,0 +1,42 @@ +1:71 NEWLINE +4:1 KW_FN +4:4 IDENT(page) +4:8 LPAREN +4:9 IDENT(name) +4:13 COLON +4:15 IDENT(Text) +4:19 RPAREN +4:21 ARROW +4:24 IDENT(Text) +4:29 LBRACE +4:30 NEWLINE +5:3 KW_LET +5:7 IDENT(one) +5:11 EQ +5:13 STR(
hi
) +5:28 NEWLINE +6:3 KW_LET +6:7 IDENT(verbatim) +6:16 EQ +6:18 STR(a"b\n) +6:25 NEWLINE +7:3 KW_LET +7:7 IDENT(block) +7:13 EQ +7:15 STR(
+ many + line +
+) +12:4 NEWLINE +13:3 KW_LET +13:7 IDENT(holes) +13:13 EQ +13:15 INTERP_STR(TEXT(

),EXPR(name),TEXT(),ESC( name ),TEXT(

)) +13:41 NEWLINE +14:3 KW_RETURN +14:10 IDENT(block) +14:15 NEWLINE +15:1 RBRACE +15:2 NEWLINE +16:1 EOF diff --git a/compiler/test/golden/tokens/raw-literal.wo b/compiler/test/golden/tokens/raw-literal.wo new file mode 100644 index 0000000..5680827 --- /dev/null +++ b/compiler/test/golden/tokens/raw-literal.wo @@ -0,0 +1,15 @@ +-- iteration 37: the backtick raw text literal. One token per literal, +-- content verbatim (no escape processing), common margin removed at +-- lex time, two hole forms -- ${} raw and {{}} escaped. +fn page(name: Text) -> Text { + let one = `
hi
` + let verbatim = `a"b\n` + let block = ` +
+ many + line +
+ ` + let holes = `

${name}{{ name }}

` + return block +} diff --git a/compiler/test/runner.ml b/compiler/test/runner.ml index 61c8875..10e969d 100644 --- a/compiler/test/runner.ml +++ b/compiler/test/runner.ml @@ -236,6 +236,128 @@ let () = check "dash-continuation: a - b (spaced) lexes as Ident, Dash, Ident" (kinds = [ Token.Ident "a"; Token.Dash; Token.Ident "b"; Token.Eof ]) +(* ---- raw text literal, iteration 37 (not golden-diffed) ------------ + + golden/tokens/raw-literal.wo pins the token STREAM; these pin the + pieces a dump cannot show: that a backtick literal with no holes is + byte-identical to the Str a "..." string would have produced, that + the common margin is removed at LEX time (so no runtime cost and no + downstream stage ever sees the source indentation), and that the two + new diagnostics fire at the right position. *) + +let () = + let collector = Diag.Collector.create () in + let toks = Lexer.tokenize collector ~file:"raw.wo" "let t = `
hi
`" in + let kinds = List.map (fun (t : Token.t) -> t.kind) toks in + check "raw literal: no holes lexes as a plain Str" + (kinds + = [ Token.KwLet; Token.Ident "t"; Token.Eq; Token.Str "
hi
"; Token.Eof ]); + check_eq "raw literal: reports nothing" ~expected:0 + ~actual:(List.length (Diag.Collector.diagnostics collector)) + string_of_int + +let () = + (* Nothing between the backticks is escape-processed: a quote is a + quote and a backslash-n is two characters, which is the whole + point of the form: markup without quote-escape noise. *) + let collector = Diag.Collector.create () in + let toks = Lexer.tokenize collector ~file:"raw.wo" "`a\"b\\n`" in + let kinds = List.map (fun (t : Token.t) -> t.kind) toks in + check "raw literal: content is verbatim, no escape processing" + (kinds = [ Token.Str "a\"b\\n"; Token.Eof ]) + +let () = + (* The margin case, written the way a render() body actually is: + opening newline dropped, the 4-space common margin removed from + every line, the whitespace-only closing line reduced to nothing + while its newline survives (Java text-block behavior). *) + let src = "let t = `\n
\n many\n
\n `" in + let collector = Diag.Collector.create () in + let toks = Lexer.tokenize collector ~file:"raw.wo" src in + let kinds = List.map (fun (t : Token.t) -> t.kind) toks in + check "raw literal: common margin stripped, leading newline dropped" + (kinds + = [ + Token.KwLet; + Token.Ident "t"; + Token.Eq; + Token.Str "
\n many\n
\n"; + Token.Eof; + ]) + +let () = + (* Both hole forms in one literal. The payloads are raw and unlexed, + exactly as SExpr has always carried `${...}` -- the parser is what + tells them apart (SEsc gains the esc() wrapper). *) + let collector = Diag.Collector.create () in + let toks = Lexer.tokenize collector ~file:"raw.wo" "`

${a}{{ b }}

`" in + let kinds = List.map (fun (t : Token.t) -> t.kind) toks in + check "raw literal: ${} stays raw, {{}} becomes SEsc" + (kinds + = [ + Token.InterpStr + [ + Token.SText "

"; + Token.SExpr "a"; + (* the empty run between two adjacent holes, exactly as a + "..." string has always produced it -- the parser drops + empty SText segments in desugar_interp *) + Token.SText ""; + Token.SEsc " b "; + Token.SText "

"; + ]; + Token.Eof; + ]) + +let () = + (* Unlike a plain "..." string, an unterminated raw literal is an + error: multi-line is its normal case, so silently swallowing the + rest of the file would be a footgun, not rt parity. *) + let collector = Diag.Collector.create () in + let toks = Lexer.tokenize collector ~file:"raw.wo" "let t = `abc" in + let kinds = List.map (fun (t : Token.t) -> t.kind) toks in + check "unterminated raw literal: closes with what was collected" + (kinds = [ Token.KwLet; Token.Ident "t"; Token.Eq; Token.Str "abc"; Token.Eof ]); + let diags = Diag.Collector.diagnostics collector in + check_eq "unterminated raw literal: exactly one diagnostic reported" ~expected:1 + ~actual:(List.length diags) string_of_int; + match diags with + | [ d ] -> + check "unterminated raw literal: WO-E004 at the backtick (line 1, col 9)" + (d.code = "WO-E004" && d.site.line = 1 && d.site.col = 9) + | _ -> check "unterminated raw literal: diagnostic shape" false + +let () = + (* A raw newline inside "..." used to be accepted silently (the + scanner's catch-all appended it like any other byte), which meant a + forgotten closing quote ate the rest of the file with no + diagnostic. Now that the backtick literal is the blessed spelling + for multi-line text, that newline is an error and the scan stops + WITHOUT consuming it, so the Newline token still terminates the + statement and the next line parses normally. *) + let collector = Diag.Collector.create () in + let toks = Lexer.tokenize collector ~file:"nl.wo" "let t = \"ab\ncd" in + let kinds = List.map (fun (t : Token.t) -> t.kind) toks in + check "newline in string: scan stops at the newline, which still tokenizes" + (kinds + = [ + Token.KwLet; + Token.Ident "t"; + Token.Eq; + Token.Str "ab"; + Token.Newline; + Token.Ident "cd"; + Token.Eof; + ]); + let diags = Diag.Collector.diagnostics collector in + check_eq "newline in string: exactly one diagnostic reported" ~expected:1 + ~actual:(List.length diags) string_of_int; + match diags with + | [ d ] -> + check "newline in string: WO-E005 at the newline (line 1, col 12)" + (d.code = "WO-E005" && d.site.line = 1 && d.site.col = 12) + | _ -> check "newline in string: diagnostic shape" false + (* ---- direct parser/AST assertions (Task 4, not golden-diffed) ------------ golden/ast/*.wo fixtures already pin the AST *shape* via --dump-ast, diff --git a/docs/examples/db-actor/main.wo b/docs/examples/db-actor/main.wo index e879608..b4b47c1 100644 --- a/docs/examples/db-actor/main.wo +++ b/docs/examples/db-actor/main.wo @@ -21,7 +21,6 @@ class Job { -- round-robin, so with two writers at default shards one lands off the -- primary — the RPC path under test. class Writer { - pad: Int fn receive(msg: Job) { insert Note { tag: "w", val: msg.n }; let total = 0; @@ -33,8 +32,8 @@ class Writer { } fn main() -> Int { - let a: actor Job = spawn Writer { pad: 0 }; - let b: actor Job = spawn Writer { pad: 1 }; + let a: actor Job = spawn Writer {}; + let b: actor Job = spawn Writer {}; send(a, Job { n: 1 }); send(b, Job { n: 2 }); -- no request/response surface yet (iteration 31): poll until both rows diff --git a/docs/examples/fibers/main.wo b/docs/examples/fibers/main.wo index f7c4025..debb25d 100644 --- a/docs/examples/fibers/main.wo +++ b/docs/examples/fibers/main.wo @@ -31,7 +31,6 @@ class Counter { -- Part 2: an actor whose receive PARKS mid-message. The park releases -- the shard: main keeps ticking while this fiber sleeps. class Sleeper { - pad: Int fn receive(msg: Tick) { print("sleeper: down for ${msg.n}ms"); time.sleep(msg.n); @@ -61,7 +60,7 @@ fn main() -> Int { print("part1 done"); -- ---- part 2: a parked fiber blocks nobody ---- - let s: actor Tick = spawn Sleeper { pad: 0 }; + let s: actor Tick = spawn Sleeper {}; send(s, Tick { n: 150 }); let t = 0; while t < 8 { diff --git a/docs/examples/shop/README.md b/docs/examples/shop/README.md index 216ae42..0592262 100644 --- a/docs/examples/shop/README.md +++ b/docs/examples/shop/README.md @@ -17,6 +17,49 @@ Browse http://127.0.0.1:8080/ — products → product page → buy (stock checked and decremented) → confirmation → /orders. With `WO_DATA`, kill it and restart: the orders are still there (WAL replay). +The template builds and runs as written — the two `[deps]` resolve, the +seed lands, and every route answers. + +## The view form + +A `render()` body is one backtick raw text literal. Markup is markup: +real newlines, real double-quoted attributes, and the method's source +indentation removed at compile time, so the served bytes carry the +markup's own nesting and not the code's. + +``` +fn render() -> Text { + return ` +
+

{{ self.name }}

+

€ ${self.price}

+ ${stock} +
`; +} +``` + +Every `render()` makes its class a **component** — wo-html's structural +`Component` interface, satisfied by having the method, never declared. +Parent components hold children directly (`cards: multi Component`) and +render them with `render_all`, so `ProductListPage` knows nothing about +`ProductCard` beyond `render()`. The document itself is a component too: +`AppShell { title, content }` in `layout/app.wo`, which links a real +stylesheet rather than inlining one — that is why it is its own shell and +not wo-html's `Layout`. + +Two holes, and the difference is the whole escaping story: + +- `{{ expr }}` **HTML-escapes** — it compiles to a call to the `esc` in + scope (wo-html's, unless the app declares its own). Display data goes + here; a typo'd field is a compile error, not a broken page. +- `${ expr }` is **raw** — for markup you built yourself, like the + `${stock}` fragment above or `${content}` in the app shell. + +Nothing is parsed at request time. The literal is a compile-time form: +it produces exactly the string constant and concatenation chain the old +hand-written version did, so there is no template engine to ship, warm +up, or sandbox. + ## The file map (Angular equivalents) | this template | concern | Angular analog | @@ -25,42 +68,53 @@ it and restart: the orders are still there (WAL replay). | `layout/app.wo` | app shell: document, header+footer composition, `ok_html`/`html_error` transport helpers | `app.component.html` | | `layout/header.wo` / `footer.wo` | shared chrome fragments | `header.html` / `footer.html` | | `product_list/view.wo` | VIEW — classes with `fn render() -> Text`, fields = exactly what is displayed | `product-list/view.html` | -| `product_list.controller.wo` | CONTROLLER — query the model, fill the view, answer a `Resp` (one file per feature, root module) | component `.ts` + service | -| `product_page/`, `orders/` | one view module per feature + its root controller file | feature folders | +| `product_list/controller.wo` | CONTROLLER — query the model, fill the view, answer a `Resp`; beside its view in the same module | component `.ts` + service | +| `product_page/`, `orders/` | one module per feature: `view.wo` + `controller.wo` | feature folders | | `static_files/controller.wo` | `/assets/*` from disk, traversal-safe, typed | `angular.json` assets | | `assets/style.css` | ONE real stylesheet, sectioned per feature | the `.scss` files | -| the `render()` bodies | markup-first TODAY: one HTML line per `h = h ..` statement, `${}` holes; story 37 collapses each body to ONE raw literal with `{{ }}` auto-escaped typed holes | Vue `