Skip to content

fix: prevent incomplete Search-R1 data cache files - #623

Open
Bluu (Bluuok) wants to merge 2 commits into
microsoft:mainfrom
Bluuok:fix/search-r1-data-cache
Open

Bluu (Bluuok) wants to merge 2 commits into
microsoft:mainfrom
Bluuok:fix/search-r1-data-cache

Conversation

@Bluuok

@Bluuok Bluu (Bluuok) commented Oct 3, 2026 •

Copy link
Copy Markdown

Failed Search-R1 downloads or index assembly can leave non-empty partial cache files that subsequent runs treat as valid. Wildcard shard assembly also includes unrelated part_* files.

Write downloads and the assembled index into same-directory temporary files, clean up failures, and publish each final path only after success. Assemble the index from the two expected shards (part_aa, part_ab). Existing non-empty cache files remain reusable.

Validation in the frozen Linux CPU environment:

python -m pytest tests/examples/test_search_r1_data.py -q -p no:cacheprovider
python -m ruff check tests/examples/test_search_r1_data.py
python -m ruff format --check tests/examples/test_search_r1_data.py
python -m pre_commit run --files examples/search_r1/data_process.sh tests/examples/test_search_r1_data.py
git diff --check

All 4 tests passed. The original download/stale-shard regressions failed on unchanged main. The partial-assembly regression also failed on the previous branch version (1 failed / 3 passed): it emits the first shard and then returns a read error, verifies that no final or temporary index remains, and verifies that a subsequent run recovers the complete index. Scoped Pyright and syntax checks on repository-LF content passed.

Tests execute the full Bash script with isolated Conda/downloader commands and tiny gzip fixtures. No actual dataset download, external service or GPU training was run. Fork CI awaits maintainer approval. AI assistance was used; both the final script and regression tests were reviewed.

@Bluuok
Bluu (Bluuok) marked this pull request as ready for review October 3, 2026 14:30
Copilot AI balanced review requested due to automatic review settings October 3, 2026 14:30

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

Index assembly can still leave a non-empty partial cache file when concatenation fails.

Review effort: Balanced
Findings: 1 High severity

Open (1)
What changed in this PR

Improves Search-R1 cache reliability by publishing successful downloads atomically and restricting index assembly to expected shards.

Changes:

  • Downloads through temporary files with failure cleanup.
  • Uses only part_aa and part_ab for index assembly.
  • Adds end-to-end regression tests for curl/wget failures and stale shards.
File Description
examples/​search_r1/​data_process.sh Adds atomic downloads and explicit shard assembly.
tests/​examples/​test_search_r1_data.py Tests retries, cache preservation, and index contents.

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment thread examples/search_r1/data_process.sh Outdated

if [[ ! -s "$DATA_DIR/e5_Flat.index" ]]; then
cat "$DATA_DIR"/part_* > "$DATA_DIR/e5_Flat.index"
cat "$DATA_DIR/part_aa" "$DATA_DIR/part_ab" > "$DATA_DIR/e5_Flat.index"

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants