test(sampling): seed the sequential vs standard bootstrap uniqueness test (#223) - #229
Merged
Merged
Conversation
…223) test_seq_bootstrap_and_ind_matrix compared the mean uniqueness of 100 unseeded sequential-bootstrap draws against 100 unseeded uniform draws. The exact expectations on the AFML three-label example are 0.70564 vs 0.64198, a gap of only ~3.0 standard errors at N = 100, so the test failed by chance roughly once in 900 runs. Both samplers now draw from seeded StdRngs (seq_bootstrap_with_rng and StdRng::random_range), and N = 20,000 so the checks hold for any seed: gap > 0.04 (16 SE margin) and each mean within 0.01 of its exact value (> 9 SE). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…otstrap-test Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Sean-Koval
force-pushed
the
fix/223-seeded-seq-bootstrap-test
branch
from
September 28, 2026 03:46
d15d0f7 to
8cd4614
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Closes #223
Problem
tests/sampling.rs::test_seq_bootstrap_and_ind_matrixended with a Monte Carlo check (AFML Snippet 4.9): 100 unseeded sequential-bootstrap draws against 100 unseeded uniform draws on the book's three-label example, assertingavg_seq >= avg_std.Exact enumeration of every draw path (27 uniform samples and 27 weighted sequential paths) gives:
The gap is 0.0637. At N = 100 the SE of the difference is 0.0209, so the gap was only about 3.0 SE. Simulating from the exact distributions (200,000 replications) gives a failure rate of 0.114%, about 1 in 880 runs, which matches the one-off failure reported in the issue.
Fix
Both samplers now use seeded
StdRngs:seq_bootstrap_with_rngfor the sequential side andStdRng::random_rangefor the uniform side. The test is now deterministic.N is raised to 20,000, so the tolerances hold for any seed and survive a change of RNG algorithm. The SE of the difference becomes 0.00148, and the test checks:
avg_seq - avg_std > 0.04(the true gap is 0.0637, a 16 SE margin)So the test still checks the AFML claim, and it also pins both expectations to their exact values rather than only their ordering.
The reasoning is written up in a comment in the test.
Stability evidence
I ran a temporary test on a scratch branch through CI (
workflow_dispatch, run 36374678402,testsjob, Linux, debug). The scratch branch has since been deleted.avg_seq >= avg_std): 20 of 20,000 seeds fail (0.10%). A simulation from the exact distributions agrees: 228 failures in 200,000 runs (0.114%).test_seq_bootstrap_and_ind_matrix ... ok).Other tests that assert on unseeded randomness (not changed here)
crates/openquant/tests/docs_ef3m_examples.rs::ef3m_pagerunsM2N::mp_fit(25 runs, and eachfitdraws its startingp_1fromrand::rng()), then asserts the modal parameters are within 0.02 of the truth. There is no seeded EF3M API, so fixing this needs a*_with_rngvariant offit/single_fit_loop/mp_fit. That should be a follow-up.crates/openquant/tests/bet_sizing.rs::test_bet_size_reserve_fit_and_return_parameterscallsbet_size_reserve_full, whose EM starts come fromrand::rng(). It only asserts structural properties (sigma > 0, p in (0, 1), sizes in [-1, 1]), so it cannot fail by chance.crates/openquant/tests/ef3m.rsfit(...)tests only assertis_ok()and the parameter lengths.invalid_input.rscalls unseededseq_bootstraponly for error paths and fully warmed-up samples, so its results are deterministic.np.randomorrandom.🤖 Generated with Claude Code