You cannot prove that a trading strategy is free from overfitting. You can make it much harder to fool yourself. Limit what you search, record every choice, repeat the full selection process through time, and use untouched data only after the rules are fixed.
Overfitting in backtesting is a property of the whole research process. A final strategy with two parameters can still be heavily overfit if it survived hundreds of discarded indicators, universes, date ranges, prompts, and manual edits.
Before searching: write down the plan
Before looking at final performance, write down:
- The economic or behavioral hypothesis and a simple benchmark.
- Eligible instruments, point-in-time universe, dates, frequency, and data versions.
- Signal timing, execution delay, order types, sizing, and portfolio constraints.
- Fees, spread, slippage, impact, funding, borrow, and capacity assumptions.
- Candidate families, parameter ranges, constraints, and selection metric.
- Training, validation, and final-test periods, including purge or embargo rules for overlapping labels.
- A rejection rule, not only a rule for declaring success.
This document does not make a hypothesis true. It prevents the search from quietly changing after favorable results appear.
During research: record every attempt
Keep a list of every automated and manual attempt. Record the time, code, data, strategy, settings, score, costs, random seed, train and test periods, and result. Include failed runs and ideas suggested by notebooks, other people, or code-generation tools.
The literal trial count is not always the effective number of independent tests. Nearby moving-average windows are correlated, while undocumented redesigns may add more search freedom than a saved grid reveals. Preserve the literal history and describe dependence instead of inventing a precise adjusted count. The multiple-testing guide explains why judging the winner as a single predeclared test understates false-discovery risk.
Use the simplest family that can express the stated mechanism. Complexity includes more than parameter count:
- Feature definitions, lags, filters, and missing-data choices.
- Entry, exit, stop, sizing, and allocation logic.
- Instruments, sample boundaries, metrics, and benchmarks inspected.
- Alternative cost, fill, and data treatments used to select the result.
Test how you choose the winner, not only the winner
A valid outer test must repeat every operation that learned from data. Fit scalers, features, models, thresholds, parameters, and ensemble weights using only the outer training portion. Then apply the selected result to the next untouched portion.
Walk-forward optimization preserves chronological order, but it is not automatically nested. If you repeatedly change the candidate family after viewing aggregate walk-forward results, those test windows become development data. Overlapping training windows also make fold outcomes dependent. Do not describe them as independent experiments.
For labels whose information intervals overlap, use purging based on label endpoints and an embargo when justified. An arbitrary gap is not equivalent to label-aware purging. Purged cross-validation and combinatorial purged cross-validation address particular leakage and path-analysis problems. They do not correct an undocumented search or create new market regimes.
Worked example: log every candidate and protect a final holdout
This executed example searches nine moving-average pairs on synthetic prices. Eight non-overlapping development test blocks are evaluated by a rolling selection procedure. The 500-observation training windows overlap, so the block results are not independent. The final 300 observations remain untouched until the candidate family and selection rule are frozen.
import numpy as np
import pandas as pd
rng = np.random.default_rng(41)
price = pd.Series(
100 * np.exp(np.cumsum(rng.normal(0.00015, 0.012, 1_600))),
index=pd.bdate_range("2019-01-01", periods=1_600),
)
candidates = [
(fast, slow)
for fast in (5, 10, 20)
for slow in (40, 60, 90)
if fast < slow
]
def strategy_returns(history, fast, slow, fee=0.0005):
fast_ma = history.rolling(fast).mean()
slow_ma = history.rolling(slow).mean()
position = (fast_ma > slow_ma).astype(float).shift(1).fillna(0.0)
turnover = position.diff().abs().fillna(position.abs())
return position * history.pct_change().fillna(0.0) - fee * turnover
def score(returns):
std = returns.std(ddof=1)
return np.sqrt(252) * returns.mean() / std if std > 0 else -np.inf
def select_and_test(prices, train_start, train_end, test_end):
training = prices.iloc[train_start:train_end]
trials = []
for fast, slow in candidates:
value = score(strategy_returns(training, fast, slow))
trials.append({"fast": fast, "slow": slow, "train_sharpe": value})
winner = max(trials, key=lambda row: row["train_sharpe"])
history = prices.iloc[train_start:test_end]
test = strategy_returns(
history, winner["fast"], winner["slow"]
).iloc[train_end - train_start:]
return winner, test, trials
train_len, test_len, holdout_len = 500, 100, 300
dev_end = len(price) - holdout_len
oos_parts, ledger = [], []
for train_end in range(train_len, dev_end, test_len):
winner, test, trials = select_and_test(
price,
train_end - train_len,
train_end,
min(train_end + test_len, dev_end),
)
oos_parts.append(test)
ledger.extend(trials)
dev_oos = pd.concat(oos_parts)
final_winner, final_test, final_trials = select_and_test(
price,
dev_end - train_len,
dev_end,
len(price),
)
print(f"candidates per selection: {len(candidates)}")
print(f"logged development fits: {len(ledger)}")
print(f"development walk-forward Sharpe: {score(dev_oos):.3f}")
print(
f"frozen final choice: fast={final_winner['fast']}, "
f"slow={final_winner['slow']}"
)
print(f"untouched final-holdout Sharpe: {score(final_test):.3f}")
candidates per selection: 9
logged development fits: 72
development walk-forward Sharpe: -0.099
frozen final choice: fast=20, slow=40
untouched final-holdout Sharpe: -0.190
The code retains all 72 development fits rather than only eight winners. It also carries training history into each test block so moving averages have causal warm-up and turnover includes the boundary position. The negative output is not the lesson and is not evidence about moving averages. Synthetic noise was chosen to demonstrate the mechanics without presenting a toy return as a tradable result.
One subtlety remains: if the researcher used the development walk-forward result to redesign the candidate family, that redesign belongs to development. The final 300 observations can still be used once after the revised procedure is frozen. If the researcher changes the strategy after viewing that final result, the period is no longer an untouched test.
Check stability without starting another hidden search
Plot the full parameter surface and compare it across periods, assets, regimes, costs, and defensible data treatments. A broad region is generally less fragile than an isolated peak, but it is not proof of edge. Correlated rules can form smooth surfaces on noise.
Predeclare how robustness affects selection. Choosing the widest plateau, best worst-fold score, or lowest dispersion after inspecting the plots adds another objective. Include that choice in the ledger and place it inside the validation loop. The parameter robustness guide gives a fuller set of diagnostics.
Use statistics that account for the search
No statistic converts a large, adaptive research process into certainty:
- The deflated Sharpe ratio compares a selected Sharpe estimate with an expected-maximum benchmark under assumptions about trials, sample length, skewness, and kurtosis.
- The probability of backtest overfitting estimates how often an in-sample winner ranks poorly out of sample for the supplied candidates, metric, and partitions.
- Family-wise error, false-discovery-rate procedures, White's Reality Check, and the SPA test answer different multiple-testing questions.
- Bootstrap intervals quantify sampling uncertainty only under their resampling design. A confidence interval around the selected winner does not automatically account for the search that selected it.
Report the candidate family, dependence, selection history, sample size, moments, and resampling assumptions beside the statistic. An impressive adjusted number cannot repair leakage, survivorship bias, invalid fills, or omitted costs.
Test trading assumptions separately
Overfitting control and simulation validity are separate gates. Recalculate outcomes under plausible fees, spread, slippage, latency, impact, funding, borrow, and capacity. Hand-check trades around warm-up boundaries, corporate actions, missing data, simultaneous signals, and cash constraints. Use paper trading after historical validation to test current data and operations, not to rescue a weak historical result.
Where tools help, and where they stop
VectorBT PRO can keep the main parts of this work together. It can keep a labeled record of every setting you test. Its rolling, expanding, anchored, grouped, purged, embargoed, and combinatorial splits work with indicators and portfolios. Parallel runs and saved intermediate results help with large tests. DSR, PBO-related methods, heatmaps, and analysis by time split help explain the search results.
That scale increases the need for a research log. VectorBT PRO cannot know about a discarded notebook, an automatically generated variation tried yesterday, or a test period inspected outside the current result. Its splitters implement the validation design you declare. They do not prove that a strategy will generalize.
PyBroker can execute chronological Strategy.walkforward() windows and refit registered models in each training window. Fixed indicators and rules remain fixed unless the user explicitly implements a selection procedure. Its bootstrap output estimates uncertainty under the chosen bootstrap settings, but it does not by itself correct a preceding parameter search.
Questions to answer before paper trading
Freeze the strategy only if you can answer these questions precisely:
- What hypothesis and benchmark were declared?
- How many candidates and discretionary variants were inspected?
- Which data selected features, parameters, costs, and thresholds?
- Were transformations fitted inside the correct training windows?
- How were label overlap, serial dependence, and multiple testing handled?
- Did the result survive parameter, period, asset, data, and execution stress?
- Which period remains genuinely untouched, and what happens after it is opened?
- What trading, performance, or model changes will make you stop the bot?
If the search history is incomplete, say so. Transparent uncertainty is more useful than a precise statistic built on an invented trial count.