In-sample data is any data that influenced a strategy choice. That includes fitting model parameters, selecting hyperparameters, choosing features, changing rules, picking a benchmark, choosing a start date, or deciding which result to publish. Out-of-sample data has not influenced any of those choices and is used to evaluate the frozen procedure.
A test set stops being out of sample as soon as its result causes a revision. Renaming it a validation set is then honest, but a new untouched test is needed for a final performance claim.
Training, validation, and test have different jobs
| Partition | Permitted use | What its score means |
|---|---|---|
| Training | Fit model parameters and preprocessing | Fit on data the procedure was allowed to learn from |
| Validation | Select features, rules, hyperparameters, or stopping choices | Selection evidence, not an unbiased final estimate |
| Final test | Evaluate the complete frozen procedure once | One estimate on data unused by selection |
| Paper or live forward period | Observe the frozen implementation under later conditions | Prospective evidence, still subject to regime and execution differences |
For simple rule-based research, training and validation may be combined into one development period. That whole period is in sample. The final test remains separate.
Cawley and Talbot's model-selection study shows that a selection criterion itself can be overfit and that this selection bias can be comparable to reported differences between algorithms. An out-of-sample estimate must evaluate the entire selection procedure, not merely refit the winning model.
Python example with a time-based split
This example generates a 1,000-bar random walk, chooses the best of 1,264 moving-average pairs on the first 600 bars, and evaluates that fixed pair on the remaining 400 bars. Indicators are calculated on the full chronological price series so that the test begins with legitimate past warm-up data, but only training returns choose the parameters.
import numpy as np
import pandas as pd
def strategy_returns(price, fast, slow):
fast_ma = price.rolling(fast).mean()
slow_ma = price.rolling(slow).mean()
position = (fast_ma > slow_ma).astype(float).shift(1).fillna(0)
return position * price.pct_change().fillna(0)
def sharpe(returns, periods_per_year=252):
volatility = returns.std(ddof=1)
if volatility == 0 or np.isnan(volatility):
return np.nan
return np.sqrt(periods_per_year) * returns.mean() / volatility
rng = np.random.default_rng(0)
price = pd.Series(
100 * np.exp(np.cumsum(rng.normal(0.0002, 0.01, 1_000))),
index=pd.bdate_range("2020-01-01", periods=1_000),
)
split = 600
best_score = -np.inf
best_params = None
trials = 0
for fast in range(3, 60, 2):
for slow in range(fast + 3, 250, 5):
trials += 1
returns = strategy_returns(price, fast, slow)
score = sharpe(returns.iloc[:split])
if score > best_score:
best_score = score
best_params = (fast, slow)
selected_returns = strategy_returns(price, *best_params)
test_score = sharpe(selected_returns.iloc[split:])
print(f"trials: {trials}")
print(f"selected parameters: {best_params}")
print(f"in-sample Sharpe: {best_score:.3f}")
print(f"out-of-sample Sharpe: {test_score:.3f}")
trials: 1264
selected parameters: (7, 25)
in-sample Sharpe: 0.661
out-of-sample Sharpe: -0.609
The test result collapses because the apparent training winner was selected from noise. The synthetic drift and seed are fixed only to make the example reproducible. Neither score is evidence about a market strategy.
The one-bar shift(1) is essential. A position derived from today's closing averages earns returns beginning on the next bar. Removing the shift would let the signal capture the same return that completed its inputs.
Warm-up across the boundary
Historical observations before the test start may be used to initialize a causal rolling indicator. That is not leakage because those observations would have existed at deployment. The test score must exclude warm-up-period returns and the warm-up must not be allowed to reach forward into test data.
By contrast, fitting a scaler, imputer, universe filter, feature selector, volatility estimate, or model on the entire series leaks test information even if the final returns are sliced afterward. Scikit-learn's data-leakage guidance recommends fitting transformations only on the appropriate training subset, typically through a pipeline.
How test data becomes contaminated
- Repeated peeking. Inspecting a test result and revising the strategy turns that test into validation data.
- Adaptive split choice. Selecting the boundary, market, or regime after comparing outcomes uses the purported test during design.
- Global preprocessing. Normalizing, imputing, selecting features, or estimating thresholds before splitting transfers distribution information backward.
- Overlapping labels. A training label whose evaluation interval crosses the test boundary shares outcome information with the test period.
- Point-in-time errors. Revised fundamentals, future universe membership, or incorrectly timestamped releases contaminate both partitions.
- Hidden search history. A final rule tested once may already have been chosen from many informal or published alternatives on the same data.
- Implementation mismatch. Using different cost, fill, or sizing rules in test data can make the comparison meaningless even without statistical leakage.
Purged cross-validation removes training labels whose information intervals overlap test labels. An embargo can add a post-test buffer. A fixed chronological split alone does not solve those label-level dependencies.
One split is useful, but not certain
A 60/40 or 80/20 ratio has no universal justification. The development window must cover enough observations and relevant regimes to fit the procedure. The test must be long enough to estimate the metric with useful precision and should reflect the intended deployment horizon.
One test period can be unusually favorable or hostile. Walk-forward optimization evaluates repeated train-then-future-test windows. Nested validation separates inner model selection from outer performance estimation. Neither method creates independent market histories, and both still require a final untouched or prospective check when the research process keeps evolving.
The official scikit-learn TimeSeriesSplit provides expanding chronological splits with optional training-size, test-size, and gap controls for equally spaced samples. A row gap is not automatically a label-aware purge.
A practical process
- State the target deployment period, metric, candidate family, costs, and split policy before testing.
- Reserve the final chronological test and restrict access to its detailed results.
- Perform all feature engineering, preprocessing, fitting, and selection inside development data.
- Preserve only past observations required for causal warm-up at each evaluation boundary.
- Purge overlapping labels and embargo where the information horizon requires it.
- Freeze code, dependencies, data snapshots, and parameters before the final run.
- Report every trial that influenced the selection and uncertainty around the test metric.
- If the result prompts another change, move the boundary forward and treat the old test as development data.
Tool boundaries
Community VectorBT provides rolling and expanding splitter utilities and fast labeled parameter grids. VectorBT PRO adds a unified splitter, purging, embargoing, combinatorial schemes, and split-aware parameterized workflows. Those tools create and apply partitions, but the researcher still decides which outputs influence selection and which remain protected.
PyBroker supports chronological retraining through Strategy.walkforward(). Its plain Strategy.backtest() does not automatically create an out-of-sample evaluation, and walk-forward use is explicit. The previous version of this page incorrectly implied that PyBroker made walk-forward splitting the default.
Treat the reported out-of-sample number as one estimate produced by a documented protocol. Its credibility comes from untouched data, correct timing, consistent simulation assumptions, disclosed search, and later replication, not from the label attached to the slice.