fix(dump): prune stale output pages for deleted records - #113
Merged
Merged
Conversation
`python -m app.dump` only ever wrote/updated pages for records that currently exist; it never compared the output tree against the live data, so when a source record was deleted or renamed its per-record page directory (`<category>/<slug>/index.json` and the `score/` subdir) was left on disk forever. This surfaced during today's Atom CPU dedup, where 18 removed duplicate CPU records left orphaned pages that had to be deleted by hand. Add `_prune_orphaned_pages`, invoked after each collection is written, to remove any immediate child directory of the per-category output dir whose slug is not backed by a current record (including its nested `score/` folder). Only per-slug page directories the dump owns are touched: the collection's own `index.json` list file, the top-level manifest, and `openapi.json` are non-directory entries and are never at risk. Pruning runs on every dump so the tree stays a deterministic, accurate mirror of the data — a no-change re-run remains byte-identical. Add tests covering a real record deletion (seed, dump, delete, re-dump, assert the page directory is gone while survivors and the list file remain) plus the prune helper's file/valid-slug safety guarantees. Refs #100
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
The bug
python -m app.dumpregenerates the static JSON tree by replaying the liveAPI through an in-process client and writing each record to
site/public/v1/<category>/<slug>/index.json(plusscore/index.jsonforscored collections). It only ever wrote or updated pages for records that
currently exist — it never compared the output tree against the live data and
removed pages whose backing record was gone.
As a result, every deleted or renamed record in this project's history left an
orphaned page directory sitting under
v1/<category>/forever, and everyfuture deletion would keep doing the same.
This surfaced during today's (2026-09-29) Atom CPU dedup cleanup: deleting 18
duplicate CPU source records did not remove the corresponding
.../intel-atom-*/index.jsonand.../intel-atom-*/score/index.jsonoutputpages — they had to be deleted by hand.
The fix
Add
_prune_orphaned_pages(collection_dir, valid_slugs), invoked after eachcollection is written in
generate(). It removes any immediate childdirectory of the per-category output dir whose name is not the slug of a
current record, including the nested per-record
score/folder.It is deliberately narrow and safe:
v1/<category>/).collection's own
index.jsonlist file — and the top-levelv1/index.jsonmanifest and
openapi.json— are non-directory entries, so they are nevertouched.
Pruning runs on every dump rather than behind a flag. The whole point of
this dump is to be a deterministic, accurate mirror of current data, so
"always correct / self-healing" is the safer default. Determinism is
preserved: a re-run with no data changes is byte-identical (verified below).
Tests
test_dump_prunes_output_pages_for_deleted_records: seeds fixtures, dumps,deletes a record from the DB, re-dumps, and asserts the deleted record's
output directory is gone while surviving records' pages and the collection
list file remain.
test_prune_orphaned_pages_leaves_files_and_valid_slugs: asserts the helperdrops only orphan slug dirs and never touches valid-slug dirs or the
index.jsonlist file.test_prune_orphaned_pages_noop_when_dir_missing: no-op when the categorydir doesn't exist yet.
Verification
ruff check app tests— passes.mypy app— Success: no issues found in 111 source files.pytest— 555 passed.records) produced a byte-identical tree (same SHA-256). Planting an orphan
page dir and re-running removed it and returned the tree to the identical
baseline hash — pruning introduces no spurious churn.
Refs #100