fix(verify): audit integrity_check.py for threshold rules blind to confounding variables - #112
Merged
Merged
Conversation
Audit follow-up to #111 (cores-aware era rule): swept every rule in integrity_check.py for the same shape of bug — a value compared against a fixed reference that ignores a variable which legitimately shifts what "normal" looks like — and found one more instance. The CPU cross-source ratio detector ran a single global median±MAD over the whole catalog. But ratios like cinebench_r23_multi/geekbench_multi are confounded by core count: R23 multi scales near-linearly with cores while Geekbench multicore compresses, so the ratio climbs monotonically with thread count (measured on live data: ~1.05 at 1-4T up to ~1.48 at 65T+, Pearson corr(threads, log-ratio) +0.52; PassMark/R23 falls, corr -0.59). A global median therefore flagged entire legitimate core-count strata — 90 of 739 R23/GB pairs, median 56 threads vs 16 overall, the whole EPYC/Threadripper/ Xeon many-core cluster with genuine scores — as contamination. Same shape as the flat era ceiling. Fix mirrors #111 (which divided score by threads): regress the confounder out. mad_outliers now accepts an optional per-part covariate, fits a robust Theil-Sen line of log-ratio vs log(covariate), and runs the median±MAD test on the residuals, so each part is judged against the ratio expected for its own core count. CPU pairs pass thread count; the systematic core-count gradient no longer flags while a part anomalous for its own class still does (live R23/GB flags 90 -> 10, survivors are genuine per-class outliers). Callers without a covariate (GPUs) are unchanged. Coarse thread-banding was rejected: it removes the between-band trend but shrinks the within-band envelope, netting more false positives. The GPU cross-source ratios were checked and left as a single population on purpose: they mix a theoretical spec (fp32_tflops) with empirical benchmarks across gaming vs. compute cards and many hardware eras, with no single clean stratifying variable, and stay advisory-only. Adds tests/unit/test_integrity_cross_source.py. Refs #98
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What & why
Follow-up audit to #111 (cores-aware era rule). That fix showed one integrity
rule had a specific shape of bug: it compared a raw value against a fixed
reference while ignoring a variable that legitimately shifts what "normal"
looks like (core count). This PR sweeps every rule in
integrity_check.pyfor the same shape and fixes the one other place it occurs.
Rules audited
threads>1The bug (CPU cross-source ratio outliers)
mad_outliers()computed one global median±MAD over the whole CPU catalogfor ratios like
cinebench_r23_multi / geekbench_multi. That ratio is notscale-free — it's confounded by core count, because the two benchmarks scale
differently with parallelism: Cinebench R23 multi scales near-linearly with
cores while Geekbench multicore compresses at high core counts.
Verified against live TechAPI data (
develop, 4,512 CPUs):1.05 (1-4T) → 1.22 → 1.34 → 1.43 → 1.37 → 1.48 (65T+).corr(threads, log-ratio)= +0.52 for R23/GB, −0.59 forPassMark/R23.
56 vs 16 overall. The high-ratio side (69 parts, thread-median 96)
was the entire EPYC / Threadripper / Xeon many-core cluster; the low-ratio
side (21 parts, thread-median 4) was low-core parts.
(192T) = 2.41, EPYC 9754 155,000 / 62,000 (256T) = 2.50 — real scores that
only look like outliers against a desktop-dominated global median.
Same shape as the flat era ceiling: a fixed reference blind to a variable
(core count) that legitimately shifts normal — so an entire legitimate
population gets flagged.
The fix
Mirrors #111 (which divided the score by thread count): regress the
confounder out.
mad_outliers()now takes an optional per-part covariate,fits a robust Theil–Sen line of
log-ratiovslog(covariate), and runsthe median±MAD test on the residuals, so each part is judged against the
ratio expected for its own core count. CPU pairs pass thread count (falling
back to cores, then 1); callers without a covariate (GPUs) get the original
single-population behaviour unchanged. The check is not deleted — it stays
advisory and still catches a part anomalous for its own class.
Result on live data: R23/GB flags drop 90 → 10, and the survivors are
genuine per-class outliers (Xeon Platinum 8452Y at 5.48, the Skylake-X i9 HEDT
chips, Snapdragon X). The whole legitimate many-core cluster is no longer
flagged. The era / structural / tier / single>multi sections stay clean.
Coarse thread-banding was tried and rejected: it removes the between-band
trend but shrinks the within-band MAD envelope, netting more false positives
(90 → 101). Detrending is the direct analog of #111's per-thread normalisation.
Checked and deliberately left alone
The GPU cross-source ratios (e.g.
passmark_g3d_mark / fp32_tflops) alsoproduce many flags, but the confounder is not a single clean variable: they
pair a theoretical spec (
fp32_tflops) with empirical benchmarks acrossgaming vs. compute cards and many hardware eras. There is no single stratifier
of the era-rule quality, so — per the "don't weaken a rule without showing the
flagged set is a legitimate pattern it fails to account for" bar — this rule is
left as a single population, advisory-only, and documented as such in the code.
Tests & verification
tests/unit/test_integrity_cross_source.py(mirrors the fix(verify): make era advisory rule cores-aware #111 era-ruletests): trend-followers don't flag; the same data without a covariate
reproduces the old false positives; a part anomalous for its own core count
still flags; missing covariate == original global behaviour; <8 points never
flags; zero/None values skipped; reported ratio is the raw
a/b; Theil–Senrecovers a known slope.
pytest tests/unit/test_integrity_cross_source.py tests/unit/test_integrity_era.py→ 16 passed.ruff check app tests→ clean.mypy app→ clean. Fullpytest→ green.Refs #98