Skip to content

fix(dump): prune stale output pages for deleted records - #113

Merged
Seungpyo1007 merged 1 commit into
mainfrom
Seungpyo1007/dump-prune-stale-pages
Sep 29, 2026
Merged

Seungpyo1007 merged 1 commit into
mainfrom
Seungpyo1007/dump-prune-stale-pages

Conversation

@Seungpyo1007

Copy link
Copy Markdown
Member

The bug

python -m app.dump regenerates the static JSON tree by replaying the live
API through an in-process client and writing each record to
site/public/v1/<category>/<slug>/index.json (plus score/index.json for
scored collections). It only ever wrote or updated pages for records that
currently exist — it never compared the output tree against the live data and
removed pages whose backing record was gone.

As a result, every deleted or renamed record in this project's history left an
orphaned page directory sitting under v1/<category>/ forever, and every
future deletion would keep doing the same.

This surfaced during today's (2026-09-29) Atom CPU dedup cleanup: deleting 18
duplicate CPU source records did not remove the corresponding
.../intel-atom-*/index.json and .../intel-atom-*/score/index.json output
pages — they had to be deleted by hand.

The fix

Add _prune_orphaned_pages(collection_dir, valid_slugs), invoked after each
collection is written in generate(). It removes any immediate child
directory of the per-category output dir whose name is not the slug of a
current record, including the nested per-record score/ folder.

It is deliberately narrow and safe:

  • It only ever runs inside a per-category directory (v1/<category>/).
  • It only deletes subdirectories whose name is not a live slug. The
    collection's own index.json list file — and the top-level v1/index.json
    manifest and openapi.json — are non-directory entries, so they are never
    touched.

Pruning runs on every dump rather than behind a flag. The whole point of
this dump is to be a deterministic, accurate mirror of current data, so
"always correct / self-healing" is the safer default. Determinism is
preserved: a re-run with no data changes is byte-identical (verified below).

Tests

  • test_dump_prunes_output_pages_for_deleted_records: seeds fixtures, dumps,
    deletes a record from the DB, re-dumps, and asserts the deleted record's
    output directory is gone while surviving records' pages and the collection
    list file remain.
  • test_prune_orphaned_pages_leaves_files_and_valid_slugs: asserts the helper
    drops only orphan slug dirs and never touches valid-slug dirs or the
    index.json list file.
  • test_prune_orphaned_pages_noop_when_dir_missing: no-op when the category
    dir doesn't exist yet.

Verification

  • ruff check app tests — passes.
  • mypy app — Success: no issues found in 111 source files.
  • pytest — 555 passed.
  • Determinism: two consecutive full dumps of the current dataset (191,837
    records) produced a byte-identical tree (same SHA-256). Planting an orphan
    page dir and re-running removed it and returned the tree to the identical
    baseline hash — pruning introduces no spurious churn.

Refs #100

`python -m app.dump` only ever wrote/updated pages for records that
currently exist; it never compared the output tree against the live data,
so when a source record was deleted or renamed its per-record page
directory (`<category>/<slug>/index.json` and the `score/` subdir) was
left on disk forever. This surfaced during today's Atom CPU dedup, where
18 removed duplicate CPU records left orphaned pages that had to be
deleted by hand.

Add `_prune_orphaned_pages`, invoked after each collection is written, to
remove any immediate child directory of the per-category output dir whose
slug is not backed by a current record (including its nested `score/`
folder). Only per-slug page directories the dump owns are touched: the
collection's own `index.json` list file, the top-level manifest, and
`openapi.json` are non-directory entries and are never at risk. Pruning
runs on every dump so the tree stays a deterministic, accurate mirror of
the data — a no-change re-run remains byte-identical.

Add tests covering a real record deletion (seed, dump, delete, re-dump,
assert the page directory is gone while survivors and the list file
remain) plus the prune helper's file/valid-slug safety guarantees.

Refs #100
@Seungpyo1007
Seungpyo1007 merged commit 39a55bc into main Sep 29, 2026
1 check passed
@Seungpyo1007
Seungpyo1007 deleted the Seungpyo1007/dump-prune-stale-pages branch September 29, 2026 22:37
@Seungpyo1007 Seungpyo1007 added this to the Data & API correctness milestone Sep 29, 2026
@Seungpyo1007 Seungpyo1007 added the bug Something isn't working label Sep 29, 2026
@Seungpyo1007 Seungpyo1007 self-assigned this Sep 29, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

bug Something isn't working

Projects

Status: Done

Development

Successfully merging this pull request may close these issues.

1 participant