fix(verify): make era advisory rule cores-aware - #111
Merged
Merged
Conversation
The era-vs-score check in integrity_check.py used a flat per-chip ceiling (PassMark>1500 before 2006, R23>3000 before 2011) and ignored core/thread count. Cinebench R23 and PassMark run on the physical silicon regardless of launch date, so legitimate pre-2011 high-core enthusiast parts sail past a flat gate: a 6c/12t Gulftown (i7-980X/990X) posts ~6,000-6,500 R23 today and a 4c/8t Bloomfield (i7-920) ~3,000-3,800. Today's CPU advisory re-check (Refs #98) found all 8 era findings were exactly this population of false positives (i7-920/965/870/970/980X/990X, Phenom II X6 1090T/1100T). Scale the ceiling by thread count instead (R23 1000/thread, PassMark 900/thread; threads default to cores, then 1). Pre-2011 microarchitectures top out around ~600 R23 and well under 900 PassMark per thread, so the known-good chips stop flagging, while a genuinely implausible old-chip/modern-score combo still trips it (e.g. a 2c/2009 part claiming 20,000 R23 = 10,000/thread). Extract the rule into an importable, side-effect-free era_score_outliers() helper and guard the scan under __main__ so it can be unit-tested. Add tests/unit/test_integrity_era.py covering the 8 known false positives, the thread-scaling boundary, the thread-count fallback, the pre-2006 PassMark path, and a genuine implausible case that must still flag. Verified against a live TechAPI develop checkout: the era-vs-score section now reports 0 findings (was 8) with the ratio/structural sections unchanged. Refs #98
Seungpyo1007
added a commit
that referenced
this pull request
Sep 29, 2026
Audit follow-up to #111 (cores-aware era rule): swept every rule in integrity_check.py for the same shape of bug — a value compared against a fixed reference that ignores a variable which legitimately shifts what "normal" looks like — and found one more instance. The CPU cross-source ratio detector ran a single global median±MAD over the whole catalog. But ratios like cinebench_r23_multi/geekbench_multi are confounded by core count: R23 multi scales near-linearly with cores while Geekbench multicore compresses, so the ratio climbs monotonically with thread count (measured on live data: ~1.05 at 1-4T up to ~1.48 at 65T+, Pearson corr(threads, log-ratio) +0.52; PassMark/R23 falls, corr -0.59). A global median therefore flagged entire legitimate core-count strata — 90 of 739 R23/GB pairs, median 56 threads vs 16 overall, the whole EPYC/Threadripper/ Xeon many-core cluster with genuine scores — as contamination. Same shape as the flat era ceiling. Fix mirrors #111 (which divided score by threads): regress the confounder out. mad_outliers now accepts an optional per-part covariate, fits a robust Theil-Sen line of log-ratio vs log(covariate), and runs the median±MAD test on the residuals, so each part is judged against the ratio expected for its own core count. CPU pairs pass thread count; the systematic core-count gradient no longer flags while a part anomalous for its own class still does (live R23/GB flags 90 -> 10, survivors are genuine per-class outliers). Callers without a covariate (GPUs) are unchanged. Coarse thread-banding was rejected: it removes the between-band trend but shrinks the within-band envelope, netting more false positives. The GPU cross-source ratios were checked and left as a single population on purpose: they mix a theoretical spec (fp32_tflops) with empirical benchmarks across gaming vs. compute cards and many hardware eras, with no single clean stratifying variable, and stay advisory-only. Adds tests/unit/test_integrity_cross_source.py. Refs #98
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Makes the era-vs-score advisory rule in
integrity_check.pycores/threads-aware, fixing a class of false positives surfaced during today's CPU advisory re-check (Refs #98).Why
The era check flagged an "old chip carrying a modern score" using a flat per-chip ceiling —
passmark_cpu_mark > 1500before 2006 andcinebench_r23_multi > 3000before 2011 — and ignored core/thread count. But Cinebench R23 and PassMark run on the physical silicon regardless of when the part launched, so legitimate pre-2011 high-core-for-their-era chips clear a flat gate:Today's advisory re-check found that all 8 era findings were exactly this population — false positives, not data errors.
What changed
ERA_R23_PER_THREAD = 1000,ERA_PASSMARK_PER_THREAD = 900(threads default to cores, then 1). Pre-2011 microarchitectures (Nehalem/Westmere/K10) top out around ~600 R23 and well under 900 PassMark per thread, so the known-good chips stop flagging — while a genuinely implausible combo still trips it (e.g. a 2c/2009 part claiming 20,000 R23 = 10,000/thread).era_score_outliers(rec)helper and guard the scan body underif __name__ == "__main__":, so the rule can be unit-tested without running the full filesystem scan.tests/unit/test_integrity_era.py: the 8 known false positives (must not flag), a genuine implausible case (must flag), the per-thread boundary, the thread→cores→1 fallback, the pre-2006 PassMark path, missing-score no-ops, and an import-has-no-scan check.Testing
ruff check app tests— all checks passed.mypy app— success, no issues in 111 source files.pytest --cov=app --cov-fail-under=60— 586 passed, total coverage 77.69%.integrity_check.pyagainst a live TechAPI develop checkout: the CPU era-vs-score section now reports 0 findings (was 8), with the cross-source ratio and structural sections unchanged (183 ratio lines, 0 hard anomalies — identical to before).Scope
TechEngine only; no TechAPI data touched. The other advisory tiers (cross-source ratio outliers) are intentionally unchanged — those are the heterogeneous-catalog / benchmark-methodology outliers documented in the #98 re-check, not addressed here.
Refs #98