Purged cross-validation removes a training sample when the interval used to determine its label overlaps a test sample's label interval. An optional embargo excludes training samples for a documented period after a test label ends. These controls address specific forms of temporal leakage. They do not make arbitrary k-fold validation chronological, fix globally fitted features, or prove that a model will generalize.
Marcos Lopez de Prado presents purging and embargo in Chapter 7 of Advances in Financial Machine Learning. The key input is not merely a row timestamp. Each supervised sample needs:
- a prediction time
p_i, when every input required for that row was available - an evaluation time
e_i, when the response or label became fully known - a label interval
I_i = [p_i, e_i], with an explicit closed or half-open boundary convention
For training index i and test index j, purge i when I_i intersects I_j. With closed intervals:
I_i intersects I_j when p_i <= e_j and p_j <= e_i
This interval test handles labels before and after a test block, including variable horizons. A fixed number of omitted rows only matches it when observation spacing and label horizons are both fixed.
Why ordinary k-fold can leak
Suppose each row predicts a return over the next five sessions. The labels at Monday and Wednesday share several future prices. Putting Monday in training and Wednesday in testing does not create independent evidence merely because the rows are different.
Ordinary KFold does not know prediction or evaluation times. Its default need not shuffle, but contiguous folds alone still leave overlapping labels across fold boundaries. Shuffling makes the temporal problem more obvious and can spread overlaps throughout every fold.
Purging acts on the actual response intervals. It is different from deleting an arbitrary number of rows before a test set, and it is different from point-in-time feature construction. Both are required when the research problem has both delayed labels and time-varying source data.
Purging and embargo answer different questions
Purging enforces a checkable rule: no retained training-label interval intersects a test-label interval. If label horizons vary by event, the number of removed rows varies too.
Embargo adds a gap after the latest evaluation time in a test segment before later training predictions may enter. Use it when the information set, market response, data pipeline, or serial dependence creates a defensible contamination window beyond literal label overlap. Embargo duration is a research assumption, not a universal percentage and not a guarantee that autocorrelation has disappeared.
An embargo can throw away large amounts of data. Choose and report it from the information mechanism, sensitivity analysis, and intended deployment cadence. Do not tune it to maximize the validation score.
VectorBT PRO example
The following example has 32 daily samples and variable label horizons of one, three, and five days. It inspects the fold whose test indices are 8 through 15, counts every closed-interval train/test overlap, and compares no filtering, label-aware purging, and a two-day embargo.
import numpy as np
import pandas as pd
import vectorbtpro as vbt
dates = pd.date_range("2026-01-01", periods=32, freq="D")
prediction_time = pd.Series(dates, index=dates)
label_days = np.resize([1, 3, 5], len(dates))
evaluation_time = prediction_time + pd.to_timedelta(label_days, unit="D")
def get_middle_split(embargo):
splitter = vbt.Splitter.from_purged_kfold(
dates,
n_folds=4,
n_test_folds=1,
pred_times=prediction_time,
eval_times=evaluation_time,
embargo_td=embargo,
)
train_masks, test_masks = splitter.get_iter_set_mask_arrs()
for train_mask, test_mask in zip(train_masks, test_masks):
test_idx = np.flatnonzero(test_mask)
if test_idx[0] == 8:
return np.flatnonzero(train_mask), test_idx
raise RuntimeError("target split not found")
train_no_embargo, test_idx = get_middle_split("0D")
train_embargo, _ = get_middle_split("2D")
naive_train = np.setdiff1d(np.arange(len(dates)), test_idx)
def overlap_count(train_idx, test_idx):
return sum(
prediction_time.iloc[i] <= evaluation_time.iloc[j]
and prediction_time.iloc[j] <= evaluation_time.iloc[i]
for i in train_idx
for j in test_idx
)
print("test indices:", test_idx.tolist())
print("naive train:", len(naive_train), "interval overlaps:", overlap_count(naive_train, test_idx))
print("purged train:", len(train_no_embargo), "interval overlaps:", overlap_count(train_no_embargo, test_idx))
print("with 2D embargo:", len(train_embargo), "interval overlaps:", overlap_count(train_embargo, test_idx))
print("removed by embargo:", np.setdiff1d(train_no_embargo, train_embargo).tolist())
Executed with VectorBT PRO 2026.9.5, pandas 2.3.3, and NumPy 2.4.4:
test indices: [8, 9, 10, 11, 12, 13, 14, 15]
naive train: 24 interval overlaps: 13
purged train: 18 interval overlaps: 0
with 2D embargo: 16 interval overlaps: 0
removed by embargo: [20, 21]
The naive training set has 24 rows but violates the label-interval invariant 13 times. Passing the real evaluation times removes six rows and every overlap. The two-day embargo removes indices 20 and 21 as an additional post-test gap. Fewer rows are not automatically better. The meaningful result is zero measured label overlap under the declared boundary rule.
The example uses n_test_folds=1 to show ordinary purged k-fold behavior. Larger values choose combinations of test folds and move toward combinatorial purged cross-validation, which has different split-count and path-assembly questions.
VectorBT PRO details that matter
VectorBT PRO 2026.9.5 exposes both PurgedWalkForwardCV and PurgedKFoldCV through its splitter interface. The current purged k-fold implementation:
- keeps folds unshuffled and accepts explicit
pred_timesandeval_times - drops earlier training rows whose evaluation plus
purge_tdreaches a test prediction time - drops later training predictions through the latest test evaluation time
- extends that post-test exclusion by
embargo_td - supports more than one test fold per split
- returns labeled split masks that can feed parameterized fitting and scoring
Supplying prediction and evaluation times is essential. If both are omitted, the splitter defaults them to the input index. That describes zero-duration labels, so it cannot infer a five-day response horizon from the target values. purge_td can add a conservative duration, but it should not substitute for known per-sample evaluation times.
The VectorBT PRO optimization guide shows how splitters compose with parameterized functions. The splitter creates train/test masks. The user must still fit preprocessing and models only on each training mask, generate test predictions without peeking, assemble scores correctly, and keep an untouched outer evaluation when model selection occurs across folds.
Purged k-fold is not walk-forward deployment
Purged k-fold can train on observations that occur after a held-out test block once overlap and embargo rules are satisfied. That symmetry can be useful for evaluating a candidate family under a stationarity assumption, but it does not reproduce a production system that trains only on the past.
Use walk-forward optimization or a purged walk-forward splitter when the question is, "What would a model fitted only on information available so far have done next?" Use purged k-fold when future training segments are acceptable for the estimand and the goal is broader resampling efficiency. State which question the score answers.
scikit-learn's TimeSeriesSplit is chronological and supports a fixed integer gap before each test set. It does not accept per-sample evaluation times, so a fixed gap needs to cover the longest relevant label horizon. PyBroker similarly supports an explicit lookahead gap in its walk-forward workflow. Neither should be described as label-aware purging when horizons vary.
Leaks that purging cannot remove
Purging and embargo operate on split membership. A valid workflow must separately address:
- features, scalers, imputers, encoders, and selectors fitted on the full dataset
- revised data, current constituents, or other point-in-time failures
- duplicate entities or events appearing in both train and test
- target leakage inside a feature definition
- hyperparameters selected on the folds whose score is later reported
- repeated manual and automated trials
- unrealistic fills, costs, liquidity, borrow, funding, and latency
- regime changes absent from the sample
Purged folds overlap heavily, so fold scores are dependent. Do not treat their count as an equal number of independent experiments. Purging also reduces training size, sometimes unevenly across folds. Report retained counts, class balance, event coverage, and the aggregation method.
Validation checklist
For every split, assert that no retained training interval intersects any test interval. Assert the embargo boundary separately. Fit every learned transformation inside the training subset, preserve chronological causality in features, and record pred_times, eval_times, boundary conventions, purge_td, embargo_td, fold construction, and removed-row counts.
Then inspect sensitivity to plausible label and embargo durations. Pair the result with search-aware controls such as probability of backtest overfitting or the deflated Sharpe ratio when many candidates were tried. Purging removes a defined overlap channel. It does not turn the remaining score into independent proof of an edge.