Skip to content

test(sampling): seed the sequential vs standard bootstrap uniqueness test (#223) - #229

Merged
Sean-Koval merged 2 commits into
mainfrom
fix/223-seeded-seq-bootstrap-test
Sep 30, 2026
Merged

Sean-Koval merged 2 commits into
mainfrom
fix/223-seeded-seq-bootstrap-test

Conversation

@Sean-Koval

@Sean-Koval Sean-Koval commented Sep 28, 2026 •

Copy link
Copy Markdown
Contributor

Closes #223

Problem

tests/sampling.rs::test_seq_bootstrap_and_ind_matrix ended with a Monte Carlo check (AFML Snippet 4.9): 100 unseeded sequential-bootstrap draws against 100 unseeded uniform draws on the book's three-label example, asserting avg_seq >= avg_std.

Exact enumeration of every draw path (27 uniform samples and 27 weighted sequential paths) gives:

E[avg uniqueness] sd
standard bootstrap 0.64198 0.15635
sequential bootstrap 0.70564 0.13858

The gap is 0.0637. At N = 100 the SE of the difference is 0.0209, so the gap was only about 3.0 SE. Simulating from the exact distributions (200,000 replications) gives a failure rate of 0.114%, about 1 in 880 runs, which matches the one-off failure reported in the issue.

Fix

  • Both samplers now use seeded StdRngs: seq_bootstrap_with_rng for the sequential side and StdRng::random_range for the uniform side. The test is now deterministic.

  • N is raised to 20,000, so the tolerances hold for any seed and survive a change of RNG algorithm. The SE of the difference becomes 0.00148, and the test checks:

    • avg_seq - avg_std > 0.04 (the true gap is 0.0637, a 16 SE margin)
    • each mean within 0.01 of its exact value (more than 9 SE each)

    So the test still checks the AFML claim, and it also pins both expectations to their exact values rather than only their ordering.

  • The reasoning is written up in a comment in the test.

Stability evidence

I ran a temporary test on a scratch branch through CI (workflow_dispatch, run 36374678402, tests job, Linux, debug). The scratch branch has since been deleted.

  • Before (old logic, N = 100, avg_seq >= avg_std): 20 of 20,000 seeds fail (0.10%). A simulation from the exact distributions agrees: 228 failures in 200,000 runs (0.114%).
  • After (new logic, N = 20,000, all three assertions): 0 of 200 seeds fail. Across those seeds the smallest gap was 0.0591 (the threshold is 0.04), and the largest deviations from the exact means were 0.0032 (standard) and 0.0034 (sequential), against a tolerance of 0.01.
  • The committed test itself passes in the same run (test_seq_bootstrap_and_ind_matrix ... ok).

Other tests that assert on unseeded randomness (not changed here)

  • crates/openquant/tests/docs_ef3m_examples.rs::ef3m_page runs M2N::mp_fit (25 runs, and each fit draws its starting p_1 from rand::rng()), then asserts the modal parameters are within 0.02 of the truth. There is no seeded EF3M API, so fixing this needs a *_with_rng variant of fit/single_fit_loop/mp_fit. That should be a follow-up.
  • crates/openquant/tests/bet_sizing.rs::test_bet_size_reserve_fit_and_return_parameters calls bet_size_reserve_full, whose EM starts come from rand::rng(). It only asserts structural properties (sigma > 0, p in (0, 1), sizes in [-1, 1]), so it cannot fail by chance.
  • crates/openquant/tests/ef3m.rs fit(...) tests only assert is_ok() and the parameter lengths.
  • invalid_input.rs calls unseeded seq_bootstrap only for error paths and fully warmed-up samples, so its results are deterministic.
  • No Python test uses unseeded np.random or random.

🤖 Generated with Claude Code

Sean-Koval and others added 2 commits September 27, 2026 23:40
…223)

test_seq_bootstrap_and_ind_matrix compared the mean uniqueness of 100
unseeded sequential-bootstrap draws against 100 unseeded uniform draws.
The exact expectations on the AFML three-label example are 0.70564 vs
0.64198, a gap of only ~3.0 standard errors at N = 100, so the test
failed by chance roughly once in 900 runs.

Both samplers now draw from seeded StdRngs (seq_bootstrap_with_rng and
StdRng::random_range), and N = 20,000 so the checks hold for any seed:
gap > 0.04 (16 SE margin) and each mean within 0.01 of its exact value
(> 9 SE).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…otstrap-test

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@Sean-Koval
Sean-Koval force-pushed the fix/223-seeded-seq-bootstrap-test branch from d15d0f7 to 8cd4614 Compare September 28, 2026 03:46
@Sean-Koval
Sean-Koval merged commit 77f257d into main Sep 30, 2026
12 checks passed
@Sean-Koval
Sean-Koval deleted the fix/223-seeded-seq-bootstrap-test branch October 1, 2026 01:26
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Flaky test: sampling::test_seq_bootstrap_and_ind_matrix uses unseeded randomness

1 participant