Skip to content
python.financial

Probability of backtest overfitting (PBO) estimates how often a strategy-selection rule chooses an in-sample winner that ranks below the same candidate set's median out of sample. It is a diagnostic of the supplied search process, performance matrix, metric, and partition scheme. It is not the probability that a strategy has no edge, will lose money, or will fail in live trading.

Bailey, Borwein, Lopez de Prado, and Zhu introduced the framework in The Probability of Backtest Overfitting. Their combinatorially symmetric cross-validation (CSCV) implementation is model-free in the narrow sense that it does not require the candidates' forecasting equations. It still relies on the supplied trials, synchronized observations, performance metric, slice design, and historical sample.

What CSCV measures

Start with a matrix M of shape T x N:

  • T is the number of synchronous return or profit-and-loss observations
  • N is the number of feasible strategy configurations actually compared
  • every column uses the same timestamps, data policy, capital convention, and cost assumptions

Partition the rows into an even number S of equal, contiguous slices. For every one of the C(S, S / 2) choices of half the slices:

  1. Join the chosen slices in their original order as the in-sample set.
  2. Join the complementary slices in their original order as the out-of-sample set.
  3. Calculate the same performance metric for every candidate in both sets.
  4. Select the candidate with the highest in-sample score.
  5. Find that candidate's out-of-sample rank r, where 1 is worst and N is best.
  6. Calculate the relative rank omega = r / (N + 1).
  7. Calculate the logit lambda = ln(omega / (1 - omega)).

The CSCV estimate is:

PBO = count(lambda < 0) / C(S, S / 2)

A negative logit means that the in-sample winner finished below the candidate median out of sample. The complete logit distribution is more informative than the fraction alone because it shows the severity and dispersion of rank degradation.

Python example

This compact implementation uses column-wise unannualized Sharpe ratios because annualization would multiply every candidate by the same positive constant and would not change their ranks. It rejects unequal slices and non-finite inputs instead of silently truncating rows.

from itertools import combinations

import numpy as np


def sharpe_by_column(returns):
    mean = returns.mean(axis=0)
    std = returns.std(axis=0, ddof=1)
    return np.divide(mean, std, out=np.full_like(mean, -np.inf), where=std > 0)


def cscv_pbo(returns, n_slices=8):
    returns = np.asarray(returns, dtype=float)
    n_rows, n_strategies = returns.shape
    if n_slices % 2 or n_rows % n_slices:
        raise ValueError("n_slices must be even and divide the row count")
    if n_strategies < 2 or not np.isfinite(returns).all():
        raise ValueError("need at least two finite, synchronous strategy columns")

    slices = np.split(np.arange(n_rows), n_slices)
    all_slices = set(range(n_slices))
    logits = []

    for train_slices in combinations(range(n_slices), n_slices // 2):
        test_slices = sorted(all_slices.difference(train_slices))
        train_rows = np.concatenate([slices[i] for i in train_slices])
        test_rows = np.concatenate([slices[i] for i in test_slices])

        winner = np.argmax(sharpe_by_column(returns[train_rows]))
        test_scores = sharpe_by_column(returns[test_rows])
        rank = 1 + np.count_nonzero(test_scores < test_scores[winner])
        relative_rank = rank / (n_strategies + 1)
        logits.append(np.log(relative_rank / (1 - relative_rank)))

    logits = np.asarray(logits)
    return np.mean(logits < 0), logits


rng = np.random.default_rng(7)
noise = rng.normal(0, 0.01, size=(800, 50))
stable_column = noise.copy()
stable_column[:, 0] += 0.002

for label, matrix in [("noise", noise), ("injected column", stable_column)]:
    pbo, logits = cscv_pbo(matrix)
    print(f"{label:15s} PBO={pbo:.3f}  splits={len(logits)}")

Executed with Python 3.11 and NumPy 2.4.4:

noise           PBO=0.386  splits=70
injected column PBO=0.029  splits=70

The first matrix contains 50 exchangeable zero-mean columns. Its single-sample estimate is 0.386, not exactly the informationless reference of 0.5. That difference is a useful warning against treating one PBO number as a calibrated universal probability. The second matrix differs only because column 0 has a known, persistent mean shift. Its lower PBO shows that this selection rule identifies the injected column consistently within this constructed sample. It does not establish a market edge.

The synthetic scores are continuous, so exact ties are negligible here. Production code needs an explicit policy for tied in-sample winners and tied out-of-sample ranks, such as deterministic candidate ordering plus average ranks. Report that policy because it can affect a small candidate set.

The example uses S = 8, which creates 70 symmetric splits and keeps the calculation inspectable. S = 16 creates 12,870 splits. More splits give a finer empirical logit distribution but do not create 12,870 independent histories. Every combination reuses observations, so ordinary binomial confidence calculations that assume independent trials are not justified by the combination count alone.

The number is conditional on the candidate set

PBO can change even when the favored strategy and market data do not:

  • hiding failed trials can make the selected candidate's relative rank look better
  • padding the matrix with deliberately poor candidates can also make its rank look better
  • including every intermediate point from an adaptive optimizer can misrepresent the actual final alternatives
  • omitting manual, abandoned, or automatically generated searches leaves the file-drawer problem outside the calculation
  • changing the metric, cost model, time span, slice count, or candidate feasibility rules changes the estimand

Define the candidate family and selection rule before using PBO. For an adaptive search, the paper recommends treating each converged search result as a candidate rather than every guided intermediate step. Keep the complete research ledger so that the reported family matches the real decision process.

What low and high PBO do not mean

A low PBO says that in-sample winners tended to retain an above-median rank among the supplied candidates under these symmetric partitions. It does not prove positive expected return. Every candidate could lose money while one loses less consistently.

A high PBO says that optimizing within this family often selected candidates that ranked poorly in the complementary sample. It does not prove that every candidate lacks skill. The paper notes that a family of similarly strong strategies can have high PBO because selecting one exact winner adds little.

There is no universal interpretation in which 0.49 is safe and 0.51 is doomed. Report the full logit distribution, out-of-sample score distribution, probability of out-of-sample loss, performance degradation, T, N, S, metric, costs, and candidate construction. Predeclare any action threshold rather than choosing it after seeing the estimate.

Limits for time-series tests

CSCV evaluates selection among already generated performance columns. It does not automatically:

  • fit features, scalers, or models separately inside each training partition
  • purge overlapping labels or embargo nearby observations
  • enforce point-in-time data
  • repair look-ahead, survivorship bias, incorrect costs, or optimistic fills
  • model a regime that is absent from the historical sample
  • reproduce a chronological train-then-trade deployment

The distinction matters for machine learning. If candidate predictions depend on fitted preprocessing or models, generate each candidate's out-of-sample predictions without leakage before constructing the matrix, or use a nested procedure suited to that fitting process. Standard CSCV's symmetric, nonchronological recombination answers a different question from walk-forward optimization.

The metric must also behave sensibly on recombined slices. Sharpe-like statistics can use the selected rows directly. Path-dependent measures such as maximum drawdown depend on observation order and portfolio state. Joining disjoint blocks can create artificial transitions, so define state resets and metric construction explicitly.

How the tools fit

The VectorBT family can produce the synchronized T x N candidate matrix needed by the calculation. VectorBT PRO is especially useful for large conditional grids.

PyBroker provides explicit walk-forward analysis. A walk-forward result is not automatically a CSCV result because the training windows, test windows, selection process, and symmetric complement construction differ. PyBroker can help generate chronological out-of-sample evidence, which should complement rather than be relabeled as PBO.

Use deflated Sharpe ratio when the question concerns whether a selected Sharpe clears a search-adjusted benchmark under its stated moment and trial assumptions. Use PBO when the question concerns out-of-sample rank degradation within an explicit candidate family. Neither replaces untouched outer evaluation, realistic simulation, or prospective evidence.

Reporting checklist

Publish the complete candidate definition, all economically feasible trials, return units, alignment policy, costs, metric, T, N, S, tie and missing-value rules, logit distribution, PBO, out-of-sample loss rate, and performance degradation. State which preprocessing and model fitting occurred before the matrix was formed.

Finally, keep PBO outside the optimization objective. Repeatedly changing the strategy until PBO looks favorable overfits the diagnostic itself. Freeze the research procedure, calculate the diagnostic, and use a genuinely untouched outer period or later forward results for the next decision.

Choose which optional services may run. You can change these settings at any time.