fix(data): merge 153 confirmed website near-duplicate records - #341
Open
Seungpyo1007 wants to merge 2 commits into
Open
Seungpyo1007 wants to merge 2 commits into
Seungpyo1007 wants to merge 2 commits into
Conversation
Scanned all ~40k data/website records for near-duplicates keyed on normalized homepage_url and display name. Merged 153 groups that denote the same real website imported more than once - QID-suffixed clones, bare-domain stubs vs proper records, protocol/www/trailing-slash/punctuation variants, and same-entity name variants (e.g. Amazon.com/Amazon, ANNO/AustriaN Newspapers Online, translated open-data portal names). Each merge keeps the fuller/verified/better-sourced side, unions source_urls, ORs the verified flag, and fills only missing optional fields. Redundant files deleted (154 files; one 3-way group). Websites have no inbound cross-category references, so deletions are safe. Kept as distinct: different exhibitions/services sharing a museum or portal landing URL, board pages, registration services differing by URL fragment, and borderline company-vs-website / subsite / regional-version pairs. validate.py passes; integrity_check.py --strict reports no hard anomalies. Refs #296
Regenerated the full site/public dump fresh via TechEngine app.dump (engine pinned at 147d27d, scratch DATABASE_URL, TECHAPI_DATA_DIR) against the branch rebased onto develop. Websites collection is 39,930 after removing 154 duplicate files; the remaining non-website deltas are updated_at values that develop's own dump refresh left stale relative to the current data git history. Refs #296
Seungpyo1007
force-pushed
the
Seungpyo1007/website-dedup-scan
branch
from
September 29, 2026 17:40
5e1a416 to
b638563
Compare
Member
Author
|
12/12 categories ? 190,314 records ? diff base develop@37d2f91 🔎 Data verification — Tiers 0–3 (on demand)153 changed data record(s) (showing first 40 for network tiers). Tier 3 is dry-run — no Tier 0 — offline score (changed)
Tier 1 — source-URL liveness (changed)Checked 81 unique URL(s): 80 alive, 1 dead, 0 automation-challenged.
Tier 2 — external cross-reference (changed)
Tier 3 — promotion (dry-run)12 record(s) would promote to
Full-dataset Tier 0 baseline190314 record(s) assessed. %%{init: {"theme":"base","themeVariables":{"pie1":"#3fb950","pie2":"#d29922","pie3":"#f85149","pieStrokeWidth":"0px","pieOpacity":"1"}}}%%
pie showData
title Verification bands — all records
"Green" : 33076
"Yellow" : 156185
"Red" : 1053
Hard violations (forced red):
Requested by @Seungpyo1007 via |
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Scanned all ~40k data/website/ records for near-duplicates (one of the last untouched categories flagged in #296's accuracy tracker). Candidates were grouped by normalized homepage_url and display name; the strong signal was records sharing the exact same normalized URL. After per-group review, 153 groups were confirmed as the same real website imported more than once and merged into a single record; 154 redundant files were deleted (one group was 3-way).
Duplicate patterns merged
Merge rule
For each group I kept the fuller / verified / better-sourced side, unioned source_urls, OR'd the �erified flag, and filled only missing optional fields (never overwriting an existing value). No field values were invented.
Kept as distinct (not merged)
Records that merely share an index/landing URL but denote different real entities were left alone: museum exhibition pages under one archive host, imageboard boards, government registration services differing only by URL fragment (#ul/#ip), and borderline company-vs-website / subsite / regional-version pairs (e.g. itv.com vs ITVX, Daily Mail vs MailOnline, two different municipalities sharing one open-data portal URL).
Verification
Closes #296