Repository navigation
Conversation
…the file
hellish -n 6000-function file: 630 ms (bash 5.3.9: 32 ms)
after: 28 ms
Issue #135: sourcing a 3.6k-line rc cost 181 ms against bash's 52 ms,
"pure CPU". Under callgrind, 54% of every instruction retired while
sourcing the hellishrc_plugins framework (3.5k lines) was strlen,
all of it from one line:
int xg_alt_group(const char *at)
{
return (xg_alt_group_n(at, at + ft_strlen(at), false));
}
The lexer asks at every byte of every word whether a zsh alternation
group `(a|b)` starts there, with `at` pointing into the whole input,
so each question first measured the rest of the file -- quadratic in
its size, in either dialect. xg_alt_group_n answers 0 unless *at is
'(' and the zsh dialect is on; those two tests now come first and the
end is measured only when they pass. Same answers, by construction.
Instructions to source that framework: 457M -> 208M. OPT=1, min of 7:
bash 5.3.9 develop this
-n 6000 small funcs 32 ms 631 ms 28 ms
source the framework 24 ms 100 ms 74 ms
source 6000 trivial lines 16 ms 54 ms 24 ms
What remains of #135 is linear and elsewhere: re-lexing a chunk each
time an open construct grows past a heredoc line, the heredoc pre-scan
lexing line by line, and the lexer's per-byte lookaheads.
tests/parse_scaling_test.py gains a word-heavy case, 2000 vs 16000
lines via -n and via source, no fast pass: on develop's ASan build the
ratios are 39.6 and 22.0 (quadratic), here 7.0 and 8.0 (linear, under
the new 20.0 bound). The existing cases short-circuit anything under
1.5 s, which this quadratic stayed under at their sizes.
Refs #135
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RAmeHfJNm7XjYbMNrQqvkG
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What & why
Issue #135: sourcing a 3.6k-line rc cost 181 ms, against 52 ms in bash, and all of it was CPU time.
Under callgrind, 54% of every instruction retired while sourcing the hellishrc_plugins framework (3.5k lines) was
strlen, and all of it came from one line:At every byte of every word, the lexer asks whether a zsh alternation group
(a|b)starts there.atpoints into the whole input, so each question first measured the rest of the file. That made lexing quadratic in the file's size, in either dialect.Change (
src/execution/case_match_ext3.c):xg_alt_group_nanswers 0 unless*atis(and the zsh dialect is on. Those two tests now come first, and the end of the input is measured only when both pass. The answers are the same by construction.Instructions retired to source that framework: 457 M → 208 M. Wall time, OPT=1 build, minimum of 7 runs:
-n, 6000 small functionsRefs #135. It does not close it. What remains is linear, and elsewhere:
How I verified it
tests/parse_scaling_test.pygains a word-heavy case: 2000 lines against 16000, through both-nandsource, with no fast pass.tests/tester: 5363/5363;tests/run_scripts.shagainstbash --posix: 125/125;verify_alloc.sh: identical output on both heaps;alloc_stress.sh: all clean;tests/pty_suite.sh: 111 ok, 6 skipped, 4 failed. None of the four failures is this change:prompt_compat_matrix,prompt_drift_matrixandprompt_jobs_badgeexpect the non-root%/$prompt and get#, because the container runs as root. They fail the same way on develop, and CI runs them as a normal user.hxp_framework_testhit the 420 s per-file limit on this 4-core container. develop's binary takes 440 s for it on the same machine, and CI's runners finish it inside the limit.norminetteis OK on the touched file.🤖 Generated with Claude Code
https://claude.ai/code/session_01RAmeHfJNm7XjYbMNrQqvkG
Generated by Claude Code