Skip to content

fix(data): resolve 8 GPU near-duplicate and spec-accuracy findings - #343

Open
Seungpyo1007 wants to merge 2 commits into
developfrom
Seungpyo1007/gpu-crosscheck-retry
Open

Seungpyo1007 wants to merge 2 commits into
developfrom
Seungpyo1007/gpu-crosscheck-retry

Conversation

@Seungpyo1007

@Seungpyo1007 Seungpyo1007 commented Sep 29, 2026 •

Copy link
Copy Markdown
Member

Summary

GPU near-duplicate and spec-conflict scan of data/gpu/ (2,076 records),
following the same method used for CPU (#321/#323), tablet/watch (#324), and
smartphone (#329/#330). Found and merged 8 confirmed cross-source duplicate
pairs
— the same physical product imported twice under different slug
conventions (one side from the Kaggle bulk dataset, one from
TechPowerUp/vendor). No configuration variants were merged.

Result: 2,076 → 2,068 GPU records.

Per-pattern breakdown

Pattern 1 — identical scrape under two row-id slugs: none confirmed. The
fingerprint scan surfaced many same-spec groups, but every one is a set of
distinct real products that happen to share core specs — Tesla K40 d/m/s/t,
GRID vGPU profiles, PCI-vs-AGP and MXM-vs-PCIe bus variants, OEM SKUs. Left
unmerged.

Pattern 2 — brand/vendor-prefix path twins: covered by the cross-source
set below (e.g. radeon-instinct-mi210 vs instinct-mi210).

Pattern 3 — cross-source duplicates under different slug conventions (8, all
merged).
For each, kept the more complete / correctly-sourced side, verified
every disagreeing field against a cited authority, unioned source_urls, and
deleted the redundant file:

Keep Delete Resolution (source-verified)
instinct-mi210 radeon-instinct-mi210 same CDNA2 die (6656 SP / 64 GB HBM2e / 300 W); kept AMD+TPU record with correct 2022-03-22 launch
instinct-mi325x radeon-instinct-mi325x kept 256 GB HBM3e / 6 TB/s (AMD official); dropped the 288 GB Kaggle value (wrong)
geforce-rtx-3060-8-gb geforce-rtx-3060-8gb kept correct 2022-10-12 launch; other had wrong 2023-05-25
geforce-rtx-5060-ti-16-gb geforce-rtx-5060-ti kept base 2407 MHz (NVIDIA 2.41 GHz); dropped 2250 MHz side
radeon-rx-9060-xt-16-gb radeon-rx-9060-xt kept the 16 GB record; unioned sources
tesla-p100-pcie-16-gb p100-pcie-16gb kept canonical Tesla-named record; folded timespy_score
v100-pcie-16gb tesla-v100-pcie-16-gb kept TDP 250 W (V100 PCIe is 250 W per NVIDIA); dropped 300 W side
v100-sxm2-32gb tesla-v100-sxm2-32-gb kept TDP 300 W (V100 SXM2 is 300 W per NVIDIA); corrected fp32 16.35 → 15.7 TFLOPS per NVIDIA V100 datasheet

Pattern 4 — internal spec accuracy spot-check: cross-checked the disputed
fields above plus a sample of surrounding records against manufacturer spec
pages and TechPowerUp. The only genuinely-wrong values sat on the redundant
(deleted) side of a duplicate pair, except the FP32 correction on
v100-sxm2-32gb noted above. Nothing else was changed on a non-duplicate
record — anything ambiguous was left as-is.

Sources

  • AMD Instinct MI325X datasheet — 256 GB HBM3e, 6 TB/s
  • NVIDIA GeForce RTX 5060 family spec page — RTX 5060 Ti base 2.41 GHz / boost 2.57 GHz
  • NVIDIA Tesla V100 datasheet + PCIe product brief — PCIe 250 W, SXM2 300 W, FP32 15.7 TFLOPS

Validation

  • python -m app.validate → ✅ Data validation passed
  • python integrity_check.py <data> --strict → ✅ integrity gate: no hard anomalies
  • Static dump refreshed with python -m app.dump (last commit); scope limited
    to v1/gpus/** + the top-level manifest (2076 → 2068, score percentiles
    shift with the changed pool).

Closes #296

Merge 8 confirmed cross-source GPU duplicate pairs where the same physical
product was imported twice under different slug conventions (one from the
Kaggle bulk dataset, one from TechPowerUp/vendor). Kept the more complete /
correctly-sourced side, source-verified every field where the two disagreed,
unioned source_urls, and deleted the redundant file (2076 -> 2068).

Pairs merged (keep <- delete):
- instinct-mi210 <- radeon-instinct-mi210 (same CDNA2 die, 6656 SP / 64 GB
  HBM2e / 300 W; kept AMD+TPU-sourced 2022-03-22 launch record)
- instinct-mi325x <- radeon-instinct-mi325x (kept 256 GB HBM3e / 6 TB/s per
  AMD official spec; dropped the 288 GB Kaggle value, which is wrong)
- geforce-rtx-3060-8-gb <- geforce-rtx-3060-8gb (kept the correct 2022-10-12
  launch record; the other listed a wrong 2023-05-25 date)
- geforce-rtx-5060-ti-16-gb <- geforce-rtx-5060-ti (kept base_clock 2407 MHz
  per NVIDIA official 2.41 GHz; dropped the 2250 MHz side)
- radeon-rx-9060-xt-16-gb <- radeon-rx-9060-xt (kept the 16 GB record)
- tesla-p100-pcie-16-gb <- p100-pcie-16gb (kept the canonical Tesla-named
  record; folded timespy_score)
- v100-pcie-16gb <- tesla-v100-pcie-16-gb (kept TDP 250 W; V100 PCIe is 250 W
  per NVIDIA, the other side's 300 W is wrong)
- v100-sxm2-32gb <- tesla-v100-sxm2-32-gb (kept TDP 300 W; V100 SXM2 is 300 W
  per NVIDIA; corrected fp32 16.35 -> 15.7 TFLOPS per NVIDIA V100 datasheet)

Close-variant families that merely share a name prefix (Tesla K40 d/m/s/t,
GRID vGPU profiles, PCI/AGP and MXM/PCIe bus variants, OEM SKUs) were left
unmerged: they are distinct real products, not duplicates.

Sources: AMD Instinct MI325X datasheet (256 GB HBM3e / 6 TB/s); NVIDIA
GeForce RTX 5060 family spec page (base 2.41 GHz / boost 2.57 GHz); NVIDIA
Tesla V100 datasheet and product brief (PCIe 250 W, SXM2 300 W, FP32 15.7
TFLOPS).

Refs #296
Regenerated the static JSON dump from the updated data via `python -m app.dump`
(TechEngine checkout). GPU collection drops 2076 -> 2068; per-record scores
shift slightly because the percentile pool changed after removing the 8
duplicate rows. Scope limited to v1/gpus/** plus the top-level manifest.

Refs #296

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

bug Something isn't working data Dataset changes enhancement New feature or request

Projects

Status: In Progress

Development

Successfully merging this pull request may close these issues.

Data accuracy: corrections, duplicates and dates (workstream)

1 participant