Layer 2: syntax
Status: planned (decision 182). Lands on the parser port before it freezes (size M; L after).
Layer 2 lets a project, or a package it depends on, change what the parser
recognises, through the syntax table
and nothing else. A syntax module is a sidecar-shaped file (default export,
no top-level await, explicit extensions on relative imports) exporting a
partial table and the four lowering hooks. It is listed in
package.json#mx.syntax.
{ "mx": { "syntax": ["@mesh/syntax"] } }
Scope
Project-scoped. The nearest package.json upward from a file decides, for
every MX file under it of any file kind (.mx, .solid.mx, .ng.mx). The
file kind supplies only the default row. A dependency’s own files parse with
the dependency’s manifest, exactly as mx.tags resolves today. A package
that uses layer 2 ships its syntax module in its own manifest and therefore
needs nothing from its consumers.
Consequences a project accepts by declaring a syntax module:
- its
.mxfiles are outside the Marko parity claim - highlighting does not follow: TextMate and tree-sitter grammars are static, so project-level syntax gets semantic tokens from the language server or nothing; diagnostics and the type projection do follow, because they resolve per file already
- two modules claiming one trigger character with overlapping matchers is a registration error at the manifest
What the table can express
| Field | Example | Lowering |
|---|---|---|
placeholder |
{{ expr }} instead of ${ expr } |
existing Interpolation |
inlineScript |
disable $ lines |
none |
blockTag |
{% for x in xs %} … {% endfor %} |
lowerBlockTag builds For, IfChain through ctx.build |
filter |
::markdown:: … :: |
lowerFilter returns IR |
expressionTriggers |
:draft (atoms), %ui.save, Order Total as one identifier |
lowerTrigger; stand-in keeps the projection one-to-one |
attributeTriggers |
:email sets name, spaced #id and .class |
lowerTrigger with node: "attribute" |
textTriggers |
%ui.save in text |
lowerTrigger; empty on .mx, for languages |
concise |
off | none |
A trigger is a first-character class plus an anchored matcher. The matcher
decides extent; the stand-in is the same-length text the expression parser
sees (:a is 0. today for atoms); node says what core builds. Semantic
checks (“is this name declared”) happen in lowerTrigger or afterLower,
never during lexing. A matcher may capture a terminator and close on it
(closeFrom), or declare balance for nested delimiters; both are data.
The Mesh syntax set
Mesh is the first layer-2 user and the acceptance test. Its module carries what used to be MX core:
- Atoms (
:namein expression position, nested atoms,::namereserved), with the contract vocabulary (type: "atom",values,pattern,ref,declares) checked fromafterLower. - Name sugar
:nameas an attribute trigger:<input :email>setsname="email"in a Mesh project; in a plain MX project it is Marko’s attributevalue:email. - Spaced
#idand.classas attribute triggers whose matcher refuses=:tag #id .classandtag=value #id .classwork,tag #id=123is the trigger’s error. Tag-adjacenttag#id=123is Marko already. The trigger setsterminatesValue, so.classafter an attribute value ends the value rather than continuing a member expression.
Order of work: the table lands, MX’s own atoms and sugars move onto it with no behavior change, then the entries move to the Mesh package. The parser’s atom corpus becomes Mesh’s table tests, run against the MX parser with the Mesh table loaded.
Performance
One array lookup per character in the content and expression states, the
same cost as today’s switch on a character code. Matchers run only on a
trigger hit, anchored, bounded by the RE2 subset. Tables are resolved per
manifest (cached by mtime) and compiled per hash (cached per process).
Not in layer 2
- A trigger on the tag-open character, or any change to attribute syntax beyond the first character of a name.
- A
scan(cursor)callback: the one thing a native lexer cannot run. Extent comes from the matcher,closeFromorbalance. - A text trigger on the
.mxdefault row. - Generated highlighting grammars.