Skip to content

cross_validation: walk-forward splits and a no-signal null for features with memory; audit runbooks 11-14 (#217) - #230

Merged
Sean-Koval merged 4 commits into
mainfrom
fix/217-walk-forward-cv
Sep 30, 2026
Merged

Sean-Koval merged 4 commits into
mainfrom
fix/217-walk-forward-cv

Conversation

@Sean-Koval

Copy link
Copy Markdown
Contributor

Closes #217.

Audit result: no published conclusion or promotion decision changes. Runbooks 11 and 13 fit slowly varying features on purged folds, and their zero-signal nulls through the same splits (plus walk-forward) show no inflation. Runbook 12 fits no model on features. Notebook 08 does no cross-validation. Runbook 14 found the bias and already declined to promote.

Audit, per notebook

Notebook Features / CV Null through the same splits (and walk-forward) Conclusions
08 algo wheel (API tour) no CV, in-sample leaderboard n/a unaffected
11 meta-labeling vol_ratio (20 vs 250-bar vol), efficiency, ma_gap: slowly varying; purged 5-fold κ=0 paths are the null. H3 Sharpe meta − primary: k-fold −0.03 (t=−0.7). New: walk-forward −0.01 (t=−0.3). With signal: walk-forward +0.08 (t=2.3) vs k-fold +0.14 (t=2.7). The precision bias under the null runs against meta in both schemes. unchanged; H3 still supported walk-forward
12 CPCV + DSR no fitted features; selects the config with the best training Sharpe; positions from past closes strength-0 null through the same CPCV splits: CPCV mean path SR −0.088, walk-forward −0.077, true −0.091 (planted: 1.506 vs 1.488) unaffected
13 bet sizing runbook 11's features; purged 5-fold κ=0: sized − filter +0.01 (t=0.4); walk-forward +0.03 (t=0.8); 0/30 null paths with best-of-grid DSR ≥ 0.95. With signal walk-forward +0.19 (t=4.3). unchanged
14 fracdiff classifier level-like features; the source of #217 already failed its null and did not promote narrative updated: the residual AUC bias is now explained (below); points to the new tools

Notebook 11 is the only one whose code changed. Its Monte Carlo gains a walk-forward Sharpe row, and it was re-executed; only that row changed in the outputs. Notebooks 12 to 14 get Markdown-only audit notes and a self-review item. Each runbook's docs page has a short "Audit: features with memory (#217)" section.

Mechanism (what the docs now say)

For a middle fold, the model is also fitted on samples after the test fold. Purging and the embargo remove label-overlap leakage (AFML §7.4). They do not remove what the model learns from those later samples. AFML §12.3 lists "the training set does not trail the testing set" among the pitfalls of CV and points to purging and the embargo as the fix. For features with memory, that fix is not enough.

Concretely, on random walks with the level as the only feature:

  • A fitted logistic model fades the level against its training mean. The slope is negative in about 9 of 10 middle folds: a random walk looks mean-reverting inside any window.
  • In k-fold, that training mean includes the levels after the fold. The fixed rule "bet up if below the before+after training mean" is right 54.0% of the time on middle folds. Against the mean of the earlier samples only, it is right 49.9%.

Residual walk-forward AUC bias (item 4)

The cause is the metric, not a leak. Pooled or within-fold AUC ranks events at different times against each other, and a later event's score, computed only from past prices, already reflects an earlier event's realised outcome. On random walks, a fixed causal score (minus the deviation from the past-100-bar mean) has a within-fold AUC of 0.54. By the time lag between the two events in a pair:

Pair lag AUC
≤ 5 bars 0.89
6–40 bars 0.65
200–400 bars 0.50

A fitted walk-forward level model has a per-fold AUC of 0.58 but an accuracy of about 0.50. Recommendation (in the docs and the contract): judge features with memory on per-event measures (accuracy, log loss, net returns) against a null. Do not use pooled AUC. No fix was attempted.

Also found: shuffling returns without replacement is a bad null for level features. It fixes the sum, so every shuffled path is a random-walk bridge pinned to the real end point, and a bridge mean-reverts. A level rule scores 0.517 walk-forward on shuffled paths versus 0.500 on demeaned i.i.d. bootstrap draws. The helper therefore bootstraps.

Added

  • Rust PurgedKFold::walk_forward_splits(n_samples, min_train_folds) -> Vec<WalkForwardSplit>
    • Tests the k-fold folds from min_train_folds on, each trained only on the purged earlier samples. The folds are the same, so k-fold and walk-forward compare fold for fold.
    • The embargo has nothing to remove, which is documented.
    • New typed error CrossValidationError::InvalidMinTrainFolds.
    • Rustdoc with a doctest. The module docs gain a section, "Features with memory: a bias purging does not remove".
  • Rust tests
    • walk_forward_splits_are_the_kfold_splits_cut_at_the_test_fold
    • walk_forward_splits_reject_invalid_input
    • level_feature_is_biased_under_purged_kfold_but_not_walk_forward: 400 Gaussian random walks with a fixed SplitMix64 stream, so no dependency bump can change it. Purged k-fold accuracy is 0.524 (t=9.3); walk-forward is 0.498 (t=−0.8).
  • Python
    • Bindings: cross_validation.walk_forward_splits.
    • Wrappers: walk_forward_splits, walk_forward_split_with_diagnostics.
    • Null helpers: null_score_distribution(evaluate, splits, make_null, n_null, seed), bootstrap_returns(returns, demean=True) and null_p_value.
    • Tests, including the level bias shown through the null helper: the k-fold null is 0.522 (t=5.1), the walk-forward null 0.492.
  • Docs
    • The cross-validation module page gains a section on the mechanism, with Rust and Python examples; the Python example's output was checked locally.
    • The research notebook contract adds a required null-through-the-same-splits control, or walk-forward, for any fitted feature with memory, and a checklist item.
    • Runbook pages gain the audit notes.
    • CHANGELOG updated.
    • Regenerated: stubs, API inventory, Python API reference, runbook gallery, changelog page. Coverage is unchanged.

Local checks run:

  • cargo test -p openquant --test cross_validation and the cross_validation doctests
  • workspace clippy with -D warnings, and fmt
  • the full pytest suite, ruff, mypy and stubtest
  • notebook lint, and notebook 11 re-executed
  • docs: content schema, Rust examples, API drift, Python reference, generated pages, coverage

🤖 Generated with Claude Code

Sean-Koval and others added 4 commits September 30, 2026 14:07
…es with memory (#217)

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…s, changelog (#217)

Runbook 11 gains a walk-forward Sharpe comparison in its Monte Carlo (re-executed);
runbooks 12-14 get audit notes. No published conclusion changes.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@Sean-Koval
Sean-Koval merged commit f4817ac into main Sep 30, 2026
12 checks passed
@Sean-Koval
Sean-Koval deleted the fix/217-walk-forward-cv branch October 1, 2026 01:26
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Purged k-fold CV inflates level-like features even on random walks (runbook 14 null failure)

1 participant