Skip to content
python.financial

Combinatorial purged cross-validation (CPCV) is a resampling method for financial labels whose information intervals can overlap. It divides time into contiguous groups and tests every combination of a fixed number of groups. It purges training labels whose information intervals overlap a test and can apply an embargo after each test interval. The resulting out-of-sample predictions can be arranged into several full-length backtest paths instead of only one.

CPCV gives a distribution of path results instead of one historical path. It does not create new market history, make the paths independent, or prove that a selected strategy will generalize.

Splits and paths are different counts

Let (N) be the number of chronological groups and (k) the number of test groups in each split. The remaining (N-k) groups train the model.

number of splits = C(N, k)
test appearances per group = C(N - 1, k - 1)
number of complete paths = phi(N, k)
                         = k / N * C(N, k)
                         = C(N - 1, k - 1)
training fraction before purge = (N - k) / N

For CPCV(6,2), there are C(6,2) = 15 train/test splits and C(5,1) = 5 complete paths. Each split supplies predictions for two test groups. A path-assembly scheme assigns those repeated group predictions so that each of the five paths contains one out-of-sample prediction for every group.

The original procedure appears in chapter 12 of Marcos Lopez de Prado's Advances in Financial Machine Learning. It requires chronological groups, all test-group combinations, label-aware purging, an optional embargo, model fitting on each cleaned training set, and an explicit mapping from split predictions to paths.

Why purging and embargo matter

Suppose a feature at time (t) predicts a five-day forward return. Its label uses information through (t+5). A training row dated just before a test group can therefore share future prices with a test label. Random or ordinary K-fold splitting does not see that overlap.

  • Purging removes a training observation when its label interval overlaps a test label interval.
  • Embargo removes training observations for a chosen interval after a test interval to reduce leakage through nearby, serially dependent samples.
  • Combinatorial testing repeats that cleaned split for every set of (k) held-out groups.

A fixed gap can be adequate when every label has the same known horizon. Variable event horizons need per-observation prediction and evaluation times. A calendar duration and a row count are not interchangeable when markets close on weekends or observations are irregular.

Check the split counts in Python

This small example checks the split count and group coverage. It does not fit a model or perform label-aware purging.

from collections import Counter
from itertools import combinations
from math import comb


n_groups = 6
n_test_groups = 2
test_sets = list(combinations(range(n_groups), n_test_groups))
coverage = Counter(group for test_set in test_sets for group in test_set)
paths = comb(n_groups - 1, n_test_groups - 1)

print(f"splits: {len(test_sets)}")
print(f"test appearances per group: {sorted(coverage.values())}")
print(f"complete paths: {paths}")
splits: 15
test appearances per group: [5, 5, 5, 5, 5, 5]
complete paths: 5

The identities balance because the splits contain 15 * 2 = 30 test-group slots and the five complete paths contain 5 * 6 = 30 group slots. The arithmetic establishes that five paths can be assembled. It does not specify which split prediction goes into each path.

VectorBT PRO splitter example

VectorBT PRO provides Splitter.from_purged_kfold() for purged and combinatorial train/test masks. The official optimization documentation demonstrates the same API. This example verifies the CPCV(6,2) split count and test coverage on 600 business-day timestamps:

import pandas as pd
import vectorbtpro as vbt


dates = pd.bdate_range("2023-01-01", periods=600)
splitter = vbt.Splitter.from_purged_kfold(
    dates,
    n_folds=6,
    n_test_folds=2,
    purge_td="3D",
    embargo_td="1D",
)
train_masks, test_masks = splitter.get_iter_set_mask_arrs()

print(f"splits: {splitter.n_splits}")
print(f"mask shapes: {train_masks.shape}, {test_masks.shape}")
print(f"test coverage: {sorted(set(test_masks.sum(axis=0)))}")

Executed with VectorBT PRO 2026.9.5, it prints:

splits: 15
mask shapes: (15, 600), (15, 600)
test coverage: [5]

The output proves that the splitter generated all 15 combinations and tested every timestamp five times. It does not itself prove that downstream model predictions were trained without leakage or assembled into five paths. For event labels, pass prediction and evaluation times that describe the real information intervals instead of treating a convenient fixed duration as the label horizon.

How to use CPCV

  1. Define the prediction time and evaluation time for every label.
  2. Partition the ordered sample into (N) contiguous groups without shuffling.
  3. Generate all (C(N,k)) choices of test groups.
  4. For each choice, purge overlapping training-label intervals and apply the embargo.
  5. Fit preprocessing, feature selection, model parameters, and thresholds on that split's training data only.
  6. Store timestamped out-of-sample predictions for every test group.
  7. Assign the repeated predictions to (phi(N,k)) complete paths using a documented path map.
  8. Simulate each path with the same costs, delays, sizing, and risk rules.
  9. Report the path distribution and the entire strategy-selection procedure, not only the best path or parameter set.

Fit scalers and feature selectors inside each split. Preprocessing the complete dataset before CPCV leaks test-distribution information even if the final estimator is refit correctly.

What CPCV can and cannot tell you

CPCV measures sensitivity to different assignments of the same historical groups. Median path performance, dispersion, lower quantiles, loss frequency, drawdown, turnover, and rank stability across parameter choices are more informative than reporting only the mean Sharpe ratio.

The paths remain dependent because they reuse observations and overlapping training sets. Standard confidence intervals that assume independent paths are therefore not automatically valid. CPCV also permits a model tested on an early group to train on later non-overlapping groups. That can be useful for estimating a stationary model's generalization error, but it does not reproduce a historically deployable expanding-window process. Use walk-forward optimization when the question requires training only on information available before each test period.

CPCV does not fix a bad benchmark, survivorship bias, point-in-time feature errors, unrealistic fills, omitted costs, nonstationarity, or multiple-testing bias. If CPCV is repeated across hundreds of strategy variants and only the winner is reported, selection bias remains. Pair path analysis with an untouched final test, parameter-stability checks, and multiplicity-aware statistics such as the deflated Sharpe ratio where appropriate.

When to use CPCV

Use CPCV when labels overlap, enough observations exist to support many fits, and a distribution of path outcomes is worth the computational cost. Choose (N), (k), purge rules, and embargo duration from the label horizon and research design, not by maximizing the reported score. Larger (N) and (k) create more splits and paths but reduce training data per fit and can make the computation grow quickly.

Use a simpler purged K-fold or walk-forward design when data are scarce, the estimator is expensive, or deployment chronology is the main question. PyBroker provides chronological walk-forward analysis, but not CPCV path construction as a built-in workflow.

Choose which optional services may run. You can change these settings at any time.