The deflated Sharpe ratio (DSR) asks whether an observed Sharpe ratio exceeds a benchmark for the best result expected from a multi-strategy search. It uses sample length, skewness, and kurtosis through the probabilistic Sharpe ratio approximation. The result is an estimated probability under that statistical model.
DSR is a number between zero and one, despite its name. It is not the probability that a strategy has skill, will make money, or will survive out of sample. Its interpretation is conditional on the recorded trial universe, the expected-maximum approximation, return observations, and distributional assumptions.
David Bailey and Marcos Lopez de Prado introduced the method in The Deflated Sharpe Ratio, published in the Journal of Portfolio Management in 2014. The paper targets selection bias under multiple testing and non-normal returns.
The expected maximum Sharpe benchmark
For (N > 1) independent trials whose Sharpe ratios have mean zero and variance (V[SR]), the approximation used here is:
SR0 = sqrt(V[SR]) * [
(1 - gamma) * Phi^-1(1 - 1 / N)
+ gamma * Phi^-1(1 - 1 / (N * e))
]
- (SR_0) is the expected maximum Sharpe ratio under the null search.
- (N) is the number of independent trials represented by the search.
- (V[SR]) is the cross-trial Sharpe-ratio variance.
- (gamma) is the Euler-Mascheroni constant.
- (Phi^{-1}) is the standard-normal quantile function.
The approximation is undefined in this form for one trial because Phi^-1(0) appears. A single prespecified strategy instead uses a benchmark Sharpe such as zero in the ordinary probabilistic Sharpe ratio. Treating N=1 as SR0=0 is a practical convention, not an output of the expected-maximum formula.
The non-normality and sample-length adjustment
Given an observed, per-period Sharpe estimate (\widehat{SR}), the DSR calculation is:
DSR = Phi(
(SR_hat - SR0) * sqrt(T - 1)
/ sqrt(1 - skew * SR_hat + ((kurtosis - 1) / 4) * SR_hat^2)
)
Here (T) is the number of return observations. skew is sample skewness and kurtosis is Pearson kurtosis, where a normal distribution has kurtosis 3. The Sharpe values and benchmark must use the same periodic units. Do not insert an annualized Sharpe into a formula whose variance and sample moments were computed from per-period values.
Python example
This example generates 100 independent, zero-mean strategies. Every strategy is noise. It selects the largest observed Sharpe and then calculates DSR using the complete simulated trial pool.
import numpy as np
from scipy import stats
rng = np.random.default_rng(21)
returns = rng.normal(0, 0.01, size=(500, 100))
sharpes = returns.mean(axis=0) / returns.std(axis=0, ddof=1)
best_idx = int(sharpes.argmax())
best_sr = float(sharpes[best_idx])
n_trials = returns.shape[1]
sr_variance = float(np.var(sharpes, ddof=1))
gamma = np.euler_gamma
sr0 = np.sqrt(sr_variance) * (
(1 - gamma) * stats.norm.ppf(1 - 1 / n_trials)
+ gamma * stats.norm.ppf(1 - 1 / (n_trials * np.e))
)
selected = returns[:, best_idx]
skew = stats.skew(selected, bias=False)
kurtosis = stats.kurtosis(selected, fisher=False, bias=False)
z = (
(best_sr - sr0) * np.sqrt(len(selected) - 1)
/ np.sqrt(1 - skew * best_sr + ((kurtosis - 1) / 4) * best_sr**2)
)
dsr = stats.norm.cdf(z)
print(f"best annualized Sharpe: {best_sr * np.sqrt(252):.3f}")
print(f"expected-max annualized Sharpe: {sr0 * np.sqrt(252):.3f}")
print(f"DSR: {dsr:.3f}")
best annualized Sharpe: 2.117
expected-max annualized Sharpe: 1.940
DSR: 0.598
The selected annualized Sharpe exceeds 2 even though its true expected return is zero. DSR reduces that headline to 0.598 against the recorded 100-trial search. The example is a formula check and selection-bias demonstration, not a market backtest. A DSR of 0.598 does not establish an edge.
What counts as a trial
Count every candidate whose result influenced the selected strategy: parameter combinations, features, signal definitions, assets, sample windows, transformations, benchmarks, and manual revisions. Reporting only the surviving code understates the search.
Correlated variants complicate (N). One hundred neighboring moving-average windows are not equivalent to 100 independent ideas. Using the raw column count with an independent-trial expected-maximum formula can misstate the adjustment. The cross-trial Sharpe variance reflects some similarity, but it does not make an arbitrary candidate set independent or recover experiments that were never logged. Document whether (N) is a literal count or an estimated effective count and explain the dependence treatment.
DSR is also not a substitute for a valid time series. Autocorrelation, overlapping returns, regime changes, stale prices, point-in-time errors, and unrealistic costs can invalidate the inputs before the formula begins.
VectorBT PRO implementation
VectorBT PRO 2026.9.5 exposes deflated_sharpe_ratio() on its returns accessor. For a two-dimensional returns object, the current implementation:
- calculates a non-annualized Sharpe ratio for each column,
- sets (N) to the number of columns,
- estimates Sharpe variance across those columns,
- calculates skewness and Pearson kurtosis per column, and
- returns the normal CDF value for every column.
The following 100-column moving-average grid is intentionally synthetic:
import itertools
import numpy as np
import pandas as pd
import vectorbtpro as vbt
rng = np.random.default_rng(11)
price = pd.Series(
100 * np.exp(np.cumsum(rng.normal(0.0004, 0.012, 500))),
index=pd.bdate_range("2023-01-01", periods=500),
)
combos = list(itertools.product(range(5, 55, 5), range(20, 220, 20)))
fast = vbt.MA.run(price, [x[0] for x in combos], short_name="fast")
slow = vbt.MA.run(price, [x[1] for x in combos], short_name="slow")
portfolio = vbt.PF.from_signals(
price,
fast.ma_crossed_above(slow),
fast.ma_crossed_below(slow),
freq="1D",
year_freq="252 days",
)
dsr = portfolio.returns.vbt.returns.deflated_sharpe_ratio()
best = portfolio.sharpe_ratio.idxmax()
print(
f"best combo: {best}, "
f"Sharpe={portfolio.sharpe_ratio[best]:.4f}, "
f"DSR={dsr[best]:.4f}"
)
best combo: (40, 20), Sharpe=1.9891, DSR=0.9558
The previous version of this page incorrectly reported 0.0850, which is approximately the complement of the actual DSR. The corrected high value is not evidence that this constructed strategy has an edge. The 100 columns share one synthetic price path, nearby moving-average rules are dependent, several parameter pairs describe closely related behavior, and the fixture was designed after its seed and grid were known. The example verifies the API and shows why an automatically inferred column count must be interpreted alongside the actual research history.
Community VectorBT also provides DSR inputs and calculation support. In either product, labeled grids make the tested universe easier to retain, but software cannot count discarded ideas outside the object.
How to report DSR clearly
Report the periodicity, risk-free rate, sample length, skewness convention, Pearson versus excess kurtosis, observed Sharpe, (SR_0), Sharpe variance, literal and effective trial counts, and DSR. A threshold such as 0.95 corresponds to a one-sided 5% convention under the model. It is not a universal certificate.
Use DSR with untouched out-of-sample data, realistic execution costs, parameter robustness, and a complete research ledger. Use probability of backtest overfitting or search-aware bootstrap tests as complementary checks because they ask different questions and rely on different assumptions.