Skip to content

perf: remove the semantic predicate from the NEWLINE lexer rule - #76

Draft
jagalindo wants to merge 7 commits into
developfrom
perf/lexer-newline-predicate
Draft

jagalindo wants to merge 7 commits into
developfrom
perf/lexer-newline-predicate

Conversation

@jagalindo

@jagalindo jagalindo commented Sep 30, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

Parsing UVL is several times slower than it needs to be because of a single semantic predicate
in the NEWLINE lexer rule. This PR removes it from the Python, Java and JavaScript lexers
without changing the token stream. Every model lexes to the same tokens, and indented first lines
are still rejected.

Problem

All three target lexers define

NEWLINE: (
		{atStartOfInput()}? SPACES
		| ( '\r'? '\n' | '\r') SPACES?
	) {handleNewline();};

The {atStartOfInput()}? predicate is reachable from the lexer's start state. In that case
ANTLR's LexerATNSimulator.matchATN never stores the DFA start state (suppressEdge = s0_closure.hasSemanticContext), so every token recomputes the whole ATN closure instead of
reusing the cached DFA.

Measured with the Python target (uvlparser 2.5.0) on a 1.5 MB UVLHub model (BerkeleyDB,
224,698 tokens):

  • computeStartState runs 177,973 times, about once per token (it should run once per mode).
  • With the start state cached, the lexer misses its DFA cache 146 times in total, and parsing is
    about 6x faster: 43.0 s down to 7.5 s on that model, and 0.033 s down to 0.005 s on small
    models. The parse trees are identical.

Java and JavaScript use the same matchATN logic, so they have the same problem.

Change

  • Grammar (uvl/{Python,Java,JavaScript}/UVL*Lexer.g4): NEWLINE loses the predicated
    alternative and becomes ( '\r'? '\n' | '\r') SPACES?.
  • Custom lexers (python/uvl/UVLCustomLexer.py, the Java @members,
    js/src/UVLJavaScriptCustomLexer.js): a new handleLeadingIndentation() runs once, on the
    first nextToken() call.
    • It looks at the leading spaces and tabs and the character after them.
    • When the old predicated alternative would have emitted tokens, it queues the same
      NEWLINE + INDENT pair, with the same offsets, line and column. SKIP_ then consumes the
      spaces as usual.
    • Each target mirrors its own handleNewline() conditions, which already differ slightly
      between targets (Java skips only on //, JavaScript does not check /, Python checks
      / and #), so each target keeps its current behaviour.
  • The now-unused atStartOfInput() helpers are removed.

Behaviour

Unchanged, including the rejection of an indented first line:

  features
      Root

still fails with Line 1:2 - no viable alternative at input ' ', as before.

Tests

  • New faulty model test_models/faulty/indented_first_line.uvl. The Python, JavaScript and Java
    test suites all pick it up and expect a syntax error.
  • New Python test test_lexer_caches_its_dfa_start_state: fails if a predicate reachable from
    the lexer start state is reintroduced.

Verification

  • Token streams (Python target): I simulated the regenerated lexer (the generated 2.5.0 lexer
    with the predicate disabled, plus the new custom lexer) and compared it with the current lexer on
    the 88 test_models and 13 edge cases:
    • The edge cases covered leading spaces, tabs, mixed indentation, //, # and /* */
      comments, blank lines, empty and whitespace-only files, CRLF, and a leading namespace.
    • All 101 inputs gave identical tokens (type, text, line, column, start, stop) and parse
      results
      .
  • UVLHub corpus (1,637 models): I parsed every model with the current lexer, with this PR's lexer
    (Python target, simulated as above), and with a start-state-caching workaround applied in flamapy.
    • Parse results: all 1,637 models identical across the three lexers: parse outcome and error
      text, a structural fingerprint of the parsed model (features, types, attributes, relations,
      constraints), and flamapy's canonical model hash. The one model that fails to parse
      (Truck_gft.uvl) fails with the same error in all three.
    • Operation results: 15,963 flamapy results per lexer, all identical. They cover leaf count,
      depth, branching factor, estimated configurations, atomic sets, language level, SAT
      satisfiable/core/dead features, and exact BDD counts on small models. Three very large models
      exceeded the operation time budget under every lexer and were skipped (KubernetesFM,
      Bausatz_2022_Burgers, embtoolkit); their parse results are identical.
    • Speed (flamapy workaround, which caches the start state like this PR): 153 s to parse the
      whole corpus, against 1,207 s with the current lexer.
  • CI (this PR): the "Run Grammar Tests" workflow regenerated all three parsers and every suite
    passes: Python 86 tests (including test_lexer_caches_its_dfa_start_state and
    faulty/indented_first_line.uvl), JavaScript 84 tests and Java 84 tests. Each suite rejects the new
    faulty model.

🤖 Generated with Claude Code

SundermannC and others added 7 commits July 2, 2026 12:44
* Updates Java maven deploy to react to releases and only to releases

Replace webiny/action-conventional-commits (checks all PR commit messages)
with amannn/action-semantic-pull-request (checks PR title), which validates
the Conventional Commits format on the squash-merge entry point instead of
each individual commit.

* chore: follow yaml styling conventions

Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com>

* Chore: updates version to 0.5.1

---------

Co-authored-by: copilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com>
Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com>
The NEWLINE rule of the Python, Java and JavaScript lexers started with the
predicate {atStartOfInput()}?. ANTLR never caches a lexer's DFA start state
when a semantic predicate is reachable from it (LexerATNSimulator.matchATN
suppresses the s0 edge), so every token recomputed the full ATN closure:
177,973 closure computations for 224,698 tokens on a 1.5 MB model.

Indentation before the first token is now handled once, in the custom lexers'
handleLeadingIndentation(), which queues the same NEWLINE + INDENT tokens the
predicated alternative emitted. The token streams are unchanged, so an
indented first line is still rejected by the parser.

Adds a faulty test model for the indented first line (picked up by the
Python, JavaScript and Java test suites) and a Python test that the lexer's
DFA start state is cached.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants