Repository navigation
Conversation
* Updates Java maven deploy to react to releases and only to releases Replace webiny/action-conventional-commits (checks all PR commit messages) with amannn/action-semantic-pull-request (checks PR title), which validates the Conventional Commits format on the squash-merge entry point instead of each individual commit. * chore: follow yaml styling conventions Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com> * Chore: updates version to 0.5.1 --------- Co-authored-by: copilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com> Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com>
The NEWLINE rule of the Python, Java and JavaScript lexers started with the
predicate {atStartOfInput()}?. ANTLR never caches a lexer's DFA start state
when a semantic predicate is reachable from it (LexerATNSimulator.matchATN
suppresses the s0 edge), so every token recomputed the full ATN closure:
177,973 closure computations for 224,698 tokens on a 1.5 MB model.
Indentation before the first token is now handled once, in the custom lexers'
handleLeadingIndentation(), which queues the same NEWLINE + INDENT tokens the
predicated alternative emitted. The token streams are unchanged, so an
indented first line is still rejected by the parser.
Adds a faulty test model for the indented first line (picked up by the
Python, JavaScript and Java test suites) and a Python test that the lexer's
DFA start state is cached.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Parsing UVL is several times slower than it needs to be because of a single semantic predicate
in the
NEWLINElexer rule. This PR removes it from the Python, Java and JavaScript lexerswithout changing the token stream. Every model lexes to the same tokens, and indented first lines
are still rejected.
Problem
All three target lexers define
The
{atStartOfInput()}?predicate is reachable from the lexer's start state. In that caseANTLR's
LexerATNSimulator.matchATNnever stores the DFA start state (suppressEdge = s0_closure.hasSemanticContext), so every token recomputes the whole ATN closure instead ofreusing the cached DFA.
Measured with the Python target (
uvlparser2.5.0) on a 1.5 MB UVLHub model (BerkeleyDB,224,698 tokens):
computeStartStateruns 177,973 times, about once per token (it should run once per mode).about 6x faster: 43.0 s down to 7.5 s on that model, and 0.033 s down to 0.005 s on small
models. The parse trees are identical.
Java and JavaScript use the same
matchATNlogic, so they have the same problem.Change
uvl/{Python,Java,JavaScript}/UVL*Lexer.g4):NEWLINEloses the predicatedalternative and becomes
( '\r'? '\n' | '\r') SPACES?.python/uvl/UVLCustomLexer.py, the Java@members,js/src/UVLJavaScriptCustomLexer.js): a newhandleLeadingIndentation()runs once, on thefirst
nextToken()call.NEWLINE+INDENTpair, with the same offsets, line and column.SKIP_then consumes thespaces as usual.
handleNewline()conditions, which already differ slightlybetween targets (Java skips only on
//, JavaScript does not check/, Python checks/and#), so each target keeps its current behaviour.atStartOfInput()helpers are removed.Behaviour
Unchanged, including the rejection of an indented first line:
still fails with
Line 1:2 - no viable alternative at input ' ', as before.Tests
test_models/faulty/indented_first_line.uvl. The Python, JavaScript and Javatest suites all pick it up and expect a syntax error.
test_lexer_caches_its_dfa_start_state: fails if a predicate reachable fromthe lexer start state is reintroduced.
Verification
with the predicate disabled, plus the new custom lexer) and compared it with the current lexer on
the 88
test_modelsand 13 edge cases://,#and/* */comments, blank lines, empty and whitespace-only files, CRLF, and a leading
namespace.results.
(Python target, simulated as above), and with a start-state-caching workaround applied in flamapy.
text, a structural fingerprint of the parsed model (features, types, attributes, relations,
constraints), and flamapy's canonical model hash. The one model that fails to parse
(
Truck_gft.uvl) fails with the same error in all three.depth, branching factor, estimated configurations, atomic sets, language level, SAT
satisfiable/core/dead features, and exact BDD counts on small models. Three very large models
exceeded the operation time budget under every lexer and were skipped (
KubernetesFM,Bausatz_2022_Burgers,embtoolkit); their parse results are identical.whole corpus, against 1,207 s with the current lexer.
passes: Python 86 tests (including
test_lexer_caches_its_dfa_start_stateandfaulty/indented_first_line.uvl), JavaScript 84 tests and Java 84 tests. Each suite rejects the newfaulty model.
🤖 Generated with Claude Code