Repository navigation
cross_validation: walk-forward splits and a no-signal null for features with memory; audit runbooks 11-14 (#217) - #230
Merged
Conversation
…es with memory (#217) Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…s, changelog (#217) Runbook 11 gains a walk-forward Sharpe comparison in its Monte Carlo (re-executed); runbooks 12-14 get audit notes. No published conclusion changes. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Closes #217.
Audit result: no published conclusion or promotion decision changes. Runbooks 11 and 13 fit slowly varying features on purged folds, and their zero-signal nulls through the same splits (plus walk-forward) show no inflation. Runbook 12 fits no model on features. Notebook 08 does no cross-validation. Runbook 14 found the bias and already declined to promote.
Audit, per notebook
vol_ratio(20 vs 250-bar vol),efficiency,ma_gap: slowly varying; purged 5-foldNotebook 11 is the only one whose code changed. Its Monte Carlo gains a walk-forward Sharpe row, and it was re-executed; only that row changed in the outputs. Notebooks 12 to 14 get Markdown-only audit notes and a self-review item. Each runbook's docs page has a short "Audit: features with memory (#217)" section.
Mechanism (what the docs now say)
For a middle fold, the model is also fitted on samples after the test fold. Purging and the embargo remove label-overlap leakage (AFML §7.4). They do not remove what the model learns from those later samples. AFML §12.3 lists "the training set does not trail the testing set" among the pitfalls of CV and points to purging and the embargo as the fix. For features with memory, that fix is not enough.
Concretely, on random walks with the level as the only feature:
Residual walk-forward AUC bias (item 4)
The cause is the metric, not a leak. Pooled or within-fold AUC ranks events at different times against each other, and a later event's score, computed only from past prices, already reflects an earlier event's realised outcome. On random walks, a fixed causal score (minus the deviation from the past-100-bar mean) has a within-fold AUC of 0.54. By the time lag between the two events in a pair:
A fitted walk-forward level model has a per-fold AUC of 0.58 but an accuracy of about 0.50. Recommendation (in the docs and the contract): judge features with memory on per-event measures (accuracy, log loss, net returns) against a null. Do not use pooled AUC. No fix was attempted.
Also found: shuffling returns without replacement is a bad null for level features. It fixes the sum, so every shuffled path is a random-walk bridge pinned to the real end point, and a bridge mean-reverts. A level rule scores 0.517 walk-forward on shuffled paths versus 0.500 on demeaned i.i.d. bootstrap draws. The helper therefore bootstraps.
Added
PurgedKFold::walk_forward_splits(n_samples, min_train_folds) -> Vec<WalkForwardSplit>min_train_foldson, each trained only on the purged earlier samples. The folds are the same, so k-fold and walk-forward compare fold for fold.CrossValidationError::InvalidMinTrainFolds.walk_forward_splits_are_the_kfold_splits_cut_at_the_test_foldwalk_forward_splits_reject_invalid_inputlevel_feature_is_biased_under_purged_kfold_but_not_walk_forward: 400 Gaussian random walks with a fixed SplitMix64 stream, so no dependency bump can change it. Purged k-fold accuracy is 0.524 (t=9.3); walk-forward is 0.498 (t=−0.8).cross_validation.walk_forward_splits.walk_forward_splits,walk_forward_split_with_diagnostics.null_score_distribution(evaluate, splits, make_null, n_null, seed),bootstrap_returns(returns, demean=True)andnull_p_value.Local checks run:
cargo test -p openquant --test cross_validationand the cross_validation doctests-D warnings, and fmt🤖 Generated with Claude Code