Skip to content

fix(verify): audit integrity_check.py for threshold rules blind to confounding variables - #112

Merged
Seungpyo1007 merged 1 commit into
mainfrom
fix/integrity-threshold-audit
Sep 29, 2026
Merged

Seungpyo1007 merged 1 commit into
mainfrom
fix/integrity-threshold-audit

Conversation

@Seungpyo1007

Copy link
Copy Markdown
Member

What & why

Follow-up audit to #111 (cores-aware era rule). That fix showed one integrity
rule had a specific shape of bug: it compared a raw value against a fixed
reference
while ignoring a variable that legitimately shifts what "normal"
looks like
(core count). This PR sweeps every rule in integrity_check.py
for the same shape and fixes the one other place it occurs.

Rules audited

Rule Kind Fixed reference? Verdict
structural (dup slug/name, slug≠file, verified-without-source) hard no — exact identity checks fine
CPU name/tier consistency advisory no — regex on model number fine
CPU single>multi hard no — self-relative invariant, guarded by threads>1 fine
CPU era-vs-score advisory was flat, fixed in #111 already fixed
CPU cross-source ratio outliers advisory yes — single global median±MAD fixed here
GPU cross-source ratio outliers advisory global median±MAD checked, left as-is (see below)

The bug (CPU cross-source ratio outliers)

mad_outliers() computed one global median±MAD over the whole CPU catalog
for ratios like cinebench_r23_multi / geekbench_multi. That ratio is not
scale-free — it's confounded by core count, because the two benchmarks scale
differently with parallelism: Cinebench R23 multi scales near-linearly with
cores while Geekbench multicore compresses at high core counts.

Verified against live TechAPI data (develop, 4,512 CPUs):

  • Per-thread-band median R23/GB ratio climbs monotonically:
    1.05 (1-4T) → 1.22 → 1.34 → 1.43 → 1.37 → 1.48 (65T+).
  • Pearson corr(threads, log-ratio) = +0.52 for R23/GB, −0.59 for
    PassMark/R23.
  • The global median flagged 90 of 739 R23/GB pairs — flagged-thread median
    56 vs 16 overall. The high-ratio side (69 parts, thread-median 96)
    was the entire EPYC / Threadripper / Xeon many-core cluster; the low-ratio
    side (21 parts, thread-median 4) was low-core parts.
  • Spot-checked raw values are genuine, e.g. EPYC 9654 R23 140,000 / GB 58,000
    (192T) = 2.41, EPYC 9754 155,000 / 62,000 (256T) = 2.50 — real scores that
    only look like outliers against a desktop-dominated global median.

Same shape as the flat era ceiling: a fixed reference blind to a variable
(core count) that legitimately shifts normal — so an entire legitimate
population gets flagged.

The fix

Mirrors #111 (which divided the score by thread count): regress the
confounder out
. mad_outliers() now takes an optional per-part covariate,
fits a robust Theil–Sen line of log-ratio vs log(covariate), and runs
the median±MAD test on the residuals, so each part is judged against the
ratio expected for its own core count. CPU pairs pass thread count (falling
back to cores, then 1); callers without a covariate (GPUs) get the original
single-population behaviour unchanged. The check is not deleted — it stays
advisory and still catches a part anomalous for its own class.

Result on live data: R23/GB flags drop 90 → 10, and the survivors are
genuine per-class outliers (Xeon Platinum 8452Y at 5.48, the Skylake-X i9 HEDT
chips, Snapdragon X). The whole legitimate many-core cluster is no longer
flagged. The era / structural / tier / single>multi sections stay clean.

Coarse thread-banding was tried and rejected: it removes the between-band
trend but shrinks the within-band MAD envelope, netting more false positives
(90 → 101). Detrending is the direct analog of #111's per-thread normalisation.

Checked and deliberately left alone

The GPU cross-source ratios (e.g. passmark_g3d_mark / fp32_tflops) also
produce many flags, but the confounder is not a single clean variable: they
pair a theoretical spec (fp32_tflops) with empirical benchmarks across
gaming vs. compute cards and many hardware eras. There is no single stratifier
of the era-rule quality, so — per the "don't weaken a rule without showing the
flagged set is a legitimate pattern it fails to account for" bar — this rule is
left as a single population, advisory-only, and documented as such in the code.

Tests & verification

  • New tests/unit/test_integrity_cross_source.py (mirrors the fix(verify): make era advisory rule cores-aware #111 era-rule
    tests): trend-followers don't flag; the same data without a covariate
    reproduces the old false positives; a part anomalous for its own core count
    still flags; missing covariate == original global behaviour; <8 points never
    flags; zero/None values skipped; reported ratio is the raw a/b; Theil–Sen
    recovers a known slope.
  • pytest tests/unit/test_integrity_cross_source.py tests/unit/test_integrity_era.py → 16 passed.
  • ruff check app tests → clean. mypy app → clean. Full pytest → green.

Refs #98

Audit follow-up to #111 (cores-aware era rule): swept every rule in
integrity_check.py for the same shape of bug — a value compared against a
fixed reference that ignores a variable which legitimately shifts what
"normal" looks like — and found one more instance.

The CPU cross-source ratio detector ran a single global median±MAD over the
whole catalog. But ratios like cinebench_r23_multi/geekbench_multi are
confounded by core count: R23 multi scales near-linearly with cores while
Geekbench multicore compresses, so the ratio climbs monotonically with thread
count (measured on live data: ~1.05 at 1-4T up to ~1.48 at 65T+, Pearson
corr(threads, log-ratio) +0.52; PassMark/R23 falls, corr -0.59). A global
median therefore flagged entire legitimate core-count strata — 90 of 739
R23/GB pairs, median 56 threads vs 16 overall, the whole EPYC/Threadripper/
Xeon many-core cluster with genuine scores — as contamination. Same shape as
the flat era ceiling.

Fix mirrors #111 (which divided score by threads): regress the confounder out.
mad_outliers now accepts an optional per-part covariate, fits a robust
Theil-Sen line of log-ratio vs log(covariate), and runs the median±MAD test on
the residuals, so each part is judged against the ratio expected for its own
core count. CPU pairs pass thread count; the systematic core-count gradient no
longer flags while a part anomalous for its own class still does (live R23/GB
flags 90 -> 10, survivors are genuine per-class outliers). Callers without a
covariate (GPUs) are unchanged. Coarse thread-banding was rejected: it removes
the between-band trend but shrinks the within-band envelope, netting more false
positives.

The GPU cross-source ratios were checked and left as a single population on
purpose: they mix a theoretical spec (fp32_tflops) with empirical benchmarks
across gaming vs. compute cards and many hardware eras, with no single clean
stratifying variable, and stay advisory-only.

Adds tests/unit/test_integrity_cross_source.py.

Refs #98
@Seungpyo1007
Seungpyo1007 merged commit 147d27d into main Sep 29, 2026
1 check passed
@Seungpyo1007
Seungpyo1007 deleted the fix/integrity-threshold-audit branch September 29, 2026 13:54
@Seungpyo1007 Seungpyo1007 added this to the Data & API correctness milestone Sep 29, 2026
@Seungpyo1007 Seungpyo1007 added the bug Something isn't working label Sep 29, 2026
@Seungpyo1007 Seungpyo1007 self-assigned this Sep 29, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

bug Something isn't working

Projects

Status: Done

Development

Successfully merging this pull request may close these issues.

1 participant