Skip to content
python.financial

Multiple-testing bias occurs when researchers test many hypotheses but evaluate the selected result with a single-test threshold. Even when every backtest is calculated correctly, selecting the largest Sharpe ratio, smallest p-value, or best parameter combination changes the probability of a false discovery.

A result from one test is not automatically meaningful. It still depends on the null hypothesis, sampling assumptions, costs, data quality, and uncertainty. Multiple testing adds another requirement: inference must account for the family of alternatives that had a chance to become the reported winner.

The probability behind the bias

Suppose m independent tests have true null hypotheses and each uses significance level alpha. The probability of at least one false rejection, called the family-wise error rate (FWER), is:

FWER = 1 - (1 - alpha)^m

At alpha = 5%, 20 independent null tests produce at least one false rejection about 64.2% of the time. The formula assumes independence. Correlated parameter combinations have a smaller effective number of distinct trials than independent tests, but treating correlation as permission to ignore multiplicity is also wrong.

This simulation uses ideal valid p-values from independent true null hypotheses. It repeats the full 20-test experiment 100,000 times.

import numpy as np

rng = np.random.default_rng(23)
n_experiments, n_trials, alpha = 100_000, 20, 0.05

# Under true continuous null hypotheses, valid p-values are Uniform(0, 1).
p_values = rng.random((n_experiments, n_trials))
minimum_p = p_values.min(axis=1)

raw_fwer = np.mean(minimum_p < alpha)
bonferroni_fwer = np.mean(minimum_p < alpha / n_trials)
independent_theory = 1 - (1 - alpha) ** n_trials

print(f"Single-test alpha: {alpha:.1%}")
print(f"Theoretical FWER for {n_trials} independent tests: {independent_theory:.1%}")
print(f"Simulated FWER without correction: {raw_fwer:.1%}")
print(f"Simulated FWER with Bonferroni: {bonferroni_fwer:.1%}")
Single-test alpha: 5.0%
Theoretical FWER for 20 independent tests: 64.2%
Simulated FWER without correction: 64.3%
Simulated FWER with Bonferroni: 4.9%

Bonferroni compares each p-value with alpha / m. It controls FWER for valid p-values without requiring independent tests, although it can be conservative when trials are numerous or dependent. The simulation is a statistical mechanism check, not a model of market returns or evidence for a strategy.

What belongs in the test family

The relevant count is broader than the final parameter grid. Include every choice influenced by the same evidence and capable of changing the reported result:

  • Strategy families, indicators, features, thresholds, lags, and holding periods.
  • Instruments, universes, bar frequencies, sample start dates, and cost assumptions.
  • Preprocessing, outlier, missing-data, and labeling choices.
  • Metrics and benchmarks inspected during selection.
  • Manual ideas abandoned after looking at disappointing results.
  • Searches run by collaborators, notebooks, scripts, optimizers, or code-generation tools against the same history.
  • Repeated visits to validation or test data.

These trials are rarely independent and may not have explicit p-values. A literal grid size can overstate independent variation, while a published notebook can understate the larger undocumented search. Keep both the literal number of evaluated candidates and evidence about their dependence. When the full history is unknown, report that limitation rather than inventing a precise effective count.

Multiple testing, overfitting, and data snooping

The terms overlap but answer different questions:

  • Multiple testing concerns simultaneous inference across a family of hypotheses.
  • Winner's curse is the upward bias in the selected best estimate.
  • Overfitting adapts a model or rule too closely to the development sample.
  • Data snooping covers the broader reuse of data for discovery and inference.
  • Look-ahead bias uses information unavailable at the historical decision time.

A model can overfit without a formal hypothesis test, and a set of correctly specified tests can still need multiplicity control. Correcting p-values does not repair look-ahead, survivorship bias, poor fills, or omitted costs.

Choose the correction for the claim

No single adjustment answers every research question.

Family-wise error control

Bonferroni and Holm procedures control the probability of one or more false rejections in a declared family. Holm's sequentially rejective procedure is at least as powerful as plain Bonferroni under its standard conditions. These methods require valid candidate-level p-values. Naive p-values from autocorrelated, heteroskedastic, or selected returns are not repaired merely by dividing alpha.

False discovery rate

False discovery rate procedures target the expected proportion of false rejections among the rejected hypotheses. They can be appropriate for screening many signals when some false discoveries are tolerable, but they do not promise that the single selected strategy is genuine. Dependence assumptions and the exact procedure still matter.

Search-aware performance tests

White's Reality Check tests whether the best candidate in a specified search has predictive superiority over a benchmark while using resampling to preserve relevant time-series dependence. Hansen's Superior Predictive Ability test modifies that framework to reduce the influence of poor alternatives. Results depend on the candidate universe, benchmark, loss function, and bootstrap design. Omitted candidates mean omitted search.

The deflated Sharpe ratio compares a selected Sharpe estimate with an expected-maximum benchmark using trial count, sample length, skewness, and kurtosis. It is a conditional probabilistic Sharpe calculation, not the posterior probability that a strategy has real skill. Dependence between trials and incomplete research logs limit its interpretation.

Validation does not erase the research history

An untouched final test can provide evidence not used during selection, but only until researchers inspect it and adapt the strategy. Repeated test-set visits turn the test into development data. Walk-forward and combinatorial purged cross-validation estimate performance across time partitions, but their folds and paths are dependent and do not automatically correct a large candidate search.

A broad parameter plateau can be more credible than an isolated optimum, yet parameter robustness is diagnostic evidence rather than a multiplicity correction. Thousands of smooth, correlated variants can still be selected from noise.

A defensible research workflow

  1. Declare the benchmark, metric, candidate family, and selection rule before the final evaluation.
  2. Log every evaluated candidate and research decision, including failed and manually abandoned ideas.
  3. Fit preprocessing and models only inside each development fold.
  4. Use uncertainty estimates that respect serial dependence and overlapping returns.
  5. Apply FWER, FDR, Reality Check, SPA, DSR, or another method that matches the stated claim.
  6. Inspect parameter and regime stability without treating those inspections as free tests.
  7. Reserve a final chronological test or prospective period and do not recycle it after failure.
  8. Report the full search scope, correction assumptions, costs, and all material selection steps.

What software can and cannot do

VectorBT PRO makes large broadcasted searches, parameter-surface analysis, walk-forward splits, purged splits, CPCV masks, and DSR calculations practical in one workflow. This is valuable because the entire candidate surface can be inspected and carried into validation. It also makes it easy to generate more trials, so the researcher still has to log searches and apply a correction that reflects work performed outside the current object.

Community VectorBT can evaluate array-based grids and calculate DSR inputs and outputs, but it does not know how many ideas were tried in other sessions. PyBroker provides explicit Strategy.walkforward() execution. Plain backtest() does not create a held-out period automatically, and neither method tracks the broader human research history.

Fast backtesting reduces computational cost. It does not reduce the statistical cost of selecting a winner from more attempts.

Choose which optional services may run. You can change these settings at any time.