fix(data): resolve 8 GPU near-duplicate and spec-accuracy findings - #343
Open
Seungpyo1007 wants to merge 2 commits into
Open
Seungpyo1007 wants to merge 2 commits into
Seungpyo1007 wants to merge 2 commits into
Conversation
Merge 8 confirmed cross-source GPU duplicate pairs where the same physical product was imported twice under different slug conventions (one from the Kaggle bulk dataset, one from TechPowerUp/vendor). Kept the more complete / correctly-sourced side, source-verified every field where the two disagreed, unioned source_urls, and deleted the redundant file (2076 -> 2068). Pairs merged (keep <- delete): - instinct-mi210 <- radeon-instinct-mi210 (same CDNA2 die, 6656 SP / 64 GB HBM2e / 300 W; kept AMD+TPU-sourced 2022-03-22 launch record) - instinct-mi325x <- radeon-instinct-mi325x (kept 256 GB HBM3e / 6 TB/s per AMD official spec; dropped the 288 GB Kaggle value, which is wrong) - geforce-rtx-3060-8-gb <- geforce-rtx-3060-8gb (kept the correct 2022-10-12 launch record; the other listed a wrong 2023-05-25 date) - geforce-rtx-5060-ti-16-gb <- geforce-rtx-5060-ti (kept base_clock 2407 MHz per NVIDIA official 2.41 GHz; dropped the 2250 MHz side) - radeon-rx-9060-xt-16-gb <- radeon-rx-9060-xt (kept the 16 GB record) - tesla-p100-pcie-16-gb <- p100-pcie-16gb (kept the canonical Tesla-named record; folded timespy_score) - v100-pcie-16gb <- tesla-v100-pcie-16-gb (kept TDP 250 W; V100 PCIe is 250 W per NVIDIA, the other side's 300 W is wrong) - v100-sxm2-32gb <- tesla-v100-sxm2-32-gb (kept TDP 300 W; V100 SXM2 is 300 W per NVIDIA; corrected fp32 16.35 -> 15.7 TFLOPS per NVIDIA V100 datasheet) Close-variant families that merely share a name prefix (Tesla K40 d/m/s/t, GRID vGPU profiles, PCI/AGP and MXM/PCIe bus variants, OEM SKUs) were left unmerged: they are distinct real products, not duplicates. Sources: AMD Instinct MI325X datasheet (256 GB HBM3e / 6 TB/s); NVIDIA GeForce RTX 5060 family spec page (base 2.41 GHz / boost 2.57 GHz); NVIDIA Tesla V100 datasheet and product brief (PCIe 250 W, SXM2 300 W, FP32 15.7 TFLOPS). Refs #296
Regenerated the static JSON dump from the updated data via `python -m app.dump` (TechEngine checkout). GPU collection drops 2076 -> 2068; per-record scores shift slightly because the percentile pool changed after removing the 8 duplicate rows. Scope limited to v1/gpus/** plus the top-level manifest. Refs #296
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
GPU near-duplicate and spec-conflict scan of
data/gpu/(2,076 records),following the same method used for CPU (#321/#323), tablet/watch (#324), and
smartphone (#329/#330). Found and merged 8 confirmed cross-source duplicate
pairs — the same physical product imported twice under different slug
conventions (one side from the Kaggle bulk dataset, one from
TechPowerUp/vendor). No configuration variants were merged.
Result: 2,076 → 2,068 GPU records.
Per-pattern breakdown
Pattern 1 — identical scrape under two row-id slugs: none confirmed. The
fingerprint scan surfaced many same-spec groups, but every one is a set of
distinct real products that happen to share core specs — Tesla K40 d/m/s/t,
GRID vGPU profiles, PCI-vs-AGP and MXM-vs-PCIe bus variants, OEM SKUs. Left
unmerged.
Pattern 2 — brand/vendor-prefix path twins: covered by the cross-source
set below (e.g.
radeon-instinct-mi210vsinstinct-mi210).Pattern 3 — cross-source duplicates under different slug conventions (8, all
merged). For each, kept the more complete / correctly-sourced side, verified
every disagreeing field against a cited authority, unioned
source_urls, anddeleted the redundant file:
instinct-mi210radeon-instinct-mi210instinct-mi325xradeon-instinct-mi325xgeforce-rtx-3060-8-gbgeforce-rtx-3060-8gbgeforce-rtx-5060-ti-16-gbgeforce-rtx-5060-tiradeon-rx-9060-xt-16-gbradeon-rx-9060-xttesla-p100-pcie-16-gbp100-pcie-16gbtimespy_scorev100-pcie-16gbtesla-v100-pcie-16-gbv100-sxm2-32gbtesla-v100-sxm2-32-gbPattern 4 — internal spec accuracy spot-check: cross-checked the disputed
fields above plus a sample of surrounding records against manufacturer spec
pages and TechPowerUp. The only genuinely-wrong values sat on the redundant
(deleted) side of a duplicate pair, except the FP32 correction on
v100-sxm2-32gbnoted above. Nothing else was changed on a non-duplicaterecord — anything ambiguous was left as-is.
Sources
Validation
python -m app.validate→ ✅ Data validation passedpython integrity_check.py <data> --strict→ ✅ integrity gate: no hard anomaliespython -m app.dump(last commit); scope limitedto
v1/gpus/**+ the top-level manifest (2076 → 2068, score percentilesshift with the changed pool).
Closes #296