Skip to content

perf(lexer): a word position asks about zsh groups without measuring the file - #162

Open
LESdylan wants to merge 1 commit into
developfrom
perf/source-parse
Open

LESdylan wants to merge 1 commit into
developfrom
perf/source-parse

Conversation

@LESdylan

Copy link
Copy Markdown
Member

What & why

Issue #135: sourcing a 3.6k-line rc cost 181 ms, against 52 ms in bash, and all of it was CPU time.

Under callgrind, 54% of every instruction retired while sourcing the hellishrc_plugins framework (3.5k lines) was strlen, and all of it came from one line:

int	xg_alt_group(const char *at)
{
	return (xg_alt_group_n(at, at + ft_strlen(at), false));
}

At every byte of every word, the lexer asks whether a zsh alternation group (a|b) starts there. at points into the whole input, so each question first measured the rest of the file. That made lexing quadratic in the file's size, in either dialect.

Change (src/execution/case_match_ext3.c): xg_alt_group_n answers 0 unless *at is ( and the zsh dialect is on. Those two tests now come first, and the end of the input is measured only when both pass. The answers are the same by construction.

Instructions retired to source that framework: 457 M → 208 M. Wall time, OPT=1 build, minimum of 7 runs:

bash 5.3.9 develop this PR
-n, 6000 small functions 32 ms 631 ms 28 ms
source the framework 24 ms 100 ms 74 ms
source 6000 trivial lines 16 ms 54 ms 24 ms

Refs #135. It does not close it. What remains is linear, and elsewhere:

  • a chunk is re-lexed each time an open construct grows past a heredoc line;
  • the heredoc pre-scan lexes line by line;
  • the lexer does lookaheads on every byte.

How I verified it

  • A scaling test that fails on develop. tests/parse_scaling_test.py gains a word-heavy case: 2000 lines against 16000, through both -n and source, with no fast pass.
    • On develop's ASan build the ratios are 39.6 and 22.0, which is quadratic.
    • On this branch they are 7.0 and 8.0: linear, and under the new 20.0 bound.
    • The existing cases pass anything under 1.5 s without comparing, and this quadratic stayed under that at their sizes, which is why they never caught it.
  • Local gates on this commit (ASan debug build):
    • golden tests/tester: 5363/5363;
    • tests/run_scripts.sh against bash --posix: 125/125;
    • verify_alloc.sh: identical output on both heaps;
    • alloc_stress.sh: all clean;
    • tests/pty_suite.sh: 111 ok, 6 skipped, 4 failed. None of the four failures is this change:
      • prompt_compat_matrix, prompt_drift_matrix and prompt_jobs_badge expect the non-root %/$ prompt and get #, because the container runs as root. They fail the same way on develop, and CI runs them as a normal user.
      • hxp_framework_test hit the 420 s per-file limit on this 4-core container. develop's binary takes 440 s for it on the same machine, and CI's runners finish it inside the limit.
  • norminette is OK on the touched file.

🤖 Generated with Claude Code

https://claude.ai/code/session_01RAmeHfJNm7XjYbMNrQqvkG


Generated by Claude Code

…the file

    hellish -n 6000-function file:  630 ms  (bash 5.3.9: 32 ms)
    after:                           28 ms

Issue #135: sourcing a 3.6k-line rc cost 181 ms against bash's 52 ms,
"pure CPU". Under callgrind, 54% of every instruction retired while
sourcing the hellishrc_plugins framework (3.5k lines) was strlen,
all of it from one line:

    int xg_alt_group(const char *at)
    {
        return (xg_alt_group_n(at, at + ft_strlen(at), false));
    }

The lexer asks at every byte of every word whether a zsh alternation
group `(a|b)` starts there, with `at` pointing into the whole input,
so each question first measured the rest of the file -- quadratic in
its size, in either dialect. xg_alt_group_n answers 0 unless *at is
'(' and the zsh dialect is on; those two tests now come first and the
end is measured only when they pass. Same answers, by construction.

Instructions to source that framework: 457M -> 208M. OPT=1, min of 7:

                           bash 5.3.9   develop   this
    -n 6000 small funcs        32 ms     631 ms    28 ms
    source the framework       24 ms     100 ms    74 ms
    source 6000 trivial lines  16 ms      54 ms    24 ms

What remains of #135 is linear and elsewhere: re-lexing a chunk each
time an open construct grows past a heredoc line, the heredoc pre-scan
lexing line by line, and the lexer's per-byte lookaheads.

tests/parse_scaling_test.py gains a word-heavy case, 2000 vs 16000
lines via -n and via source, no fast pass: on develop's ASan build the
ratios are 39.6 and 22.0 (quadratic), here 7.0 and 8.0 (linear, under
the new 20.0 bound). The existing cases short-circuit anything under
1.5 s, which this quadratic stayed under at their sizes.

Refs #135

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RAmeHfJNm7XjYbMNrQqvkG

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants