You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
[PERF] Reduce cold graph extraction work on large Python projects
Problem
The current 5,000-leaf-module comparison corpus contains 5,005 Python files
and 13,738 import edges. The published end-to-end fresh-process median for
one ArchUnitPython boundary rule was 3.59 s (v1.5.0, commit 09c9d199dec095aaa60be227bfd7b5c2b7ffb107, Python 3.13.4, Windows 11).
Exploratory timing of one direct rule.check() found graph extraction taking
roughly 1.8–2.6 s, depending on machine load. Within that phase, _extract_located_imports() accounted for about 1.2–1.7 s over 5,005 files,
and import resolution about 0.37–0.51 s over 13,738 imports. These components
are cumulative instrumentation measurements, not independent benchmark bars.
The source suggests two specific sources of avoidable work:
Each parsed file is walked separately to find TYPE_CHECKING blocks,
conditional import ranges, and actual imports/calls.
Absolute imports are resolved through repeated filesystem existence checks,
even though the scan already has the set of project Python files.
Try a single AST visitor that records contextual import kinds and imports
in one traversal. Compare it with current behavior on nested conditions.
Build a module/file lookup from the discovered project files and use it
for internal absolute-import resolution; retain current semantics for
package __init__.py, relative imports, parent roots, and external modules.
Profile again before considering parallel parsing or a persistent disk
cache. A disk cache needs explicit invalidation semantics and a separate
correctness test matrix.
Acceptance criteria
Existing tests and seeded direct-boundary/cycle correctness gates pass.
Tests cover relative and absolute imports, TYPE_CHECKING, optional imports,
literal dynamic imports, ignore directives, excluded paths, and malformed
Python files.
Report before/after fresh-process p50/p95 at 100/1,000/5,000 modules,
extraction phase time, and peak memory. Preserve zero runtime dependencies.
[PERF] Reduce cold graph extraction work on large Python projects
Problem
The current 5,000-leaf-module comparison corpus contains 5,005 Python files
and 13,738 import edges. The published end-to-end fresh-process median for
one ArchUnitPython boundary rule was 3.59 s (v1.5.0, commit
09c9d199dec095aaa60be227bfd7b5c2b7ffb107, Python 3.13.4, Windows 11).Exploratory timing of one direct
rule.check()found graph extraction takingroughly 1.8–2.6 s, depending on machine load. Within that phase,
_extract_located_imports()accounted for about 1.2–1.7 s over 5,005 files,and import resolution about 0.37–0.51 s over 13,738 imports. These components
are cumulative instrumentation measurements, not independent benchmark bars.
The source suggests two specific sources of avoidable work:
TYPE_CHECKINGblocks,conditional import ranges, and actual imports/calls.
even though the scan already has the set of project Python files.
Proposed investigation and implementation
in one traversal. Compare it with current behavior on nested conditions.
for internal absolute-import resolution; retain current semantics for
package
__init__.py, relative imports, parent roots, and external modules.cache. A disk cache needs explicit invalidation semantics and a separate
correctness test matrix.
Acceptance criteria
TYPE_CHECKING, optional imports,literal dynamic imports, ignore directives, excluded paths, and malformed
Python files.
extraction phase time, and peak memory. Preserve zero runtime dependencies.