Skip to content
python.financial

The Sharpe ratio is average periodic differential return divided by the standard deviation of that differential return. The usual benchmark is a risk-free return, so the statistic reports average excess return per unit of excess-return variability. A higher historical Sharpe is not automatically a better strategy. The comparison is meaningful only when return definitions, time periods, costs, leverage, and estimation methods are aligned.

William Sharpe introduced the reward-to-variability ratio in his 1966 mutual-fund study. His 1994 treatment of the Sharpe ratio distinguishes an ex ante ratio based on expected returns from an ex post estimate based on observed returns. A backtest reports the latter. Treating that estimate as a forecast requires evidence beyond the ratio itself.

Sharpe ratio formula

For (T) periodic strategy returns (R_t) and benchmark returns (R_{B,t}), define each differential return as:

D_t = R_t - R_B,t

The historical Sharpe estimate at the observation frequency is:

SR_period = mean(D_t) / std(D_t)
  • (R_t) is the strategy's simple return in period (t).
  • (R_{B,t}) is the benchmark return over the same period and in the same currency.
  • (D_t) is the differential return in period (t).
  • mean(D_t) is the arithmetic mean of the periodic differential returns.
  • std(D_t) is their standard deviation. State whether it uses (T) or (T-1) in the denominator.

For the conventional excess-return Sharpe, (R_{B,t}) is a risk-free return. An annual effective rate cannot simply be subtracted from daily returns. Convert it to a matching periodic rate or, preferably, align a time series of observed short-term rates:

daily_rf = (1 + annual_rf)^(1 / 252) - 1

Dividing the annual rate by 252 is an approximation. A leveraged or short strategy may also pay a borrowing, margin, funding, or stock-loan rate that differs from the rate earned on cash. The modeled benchmark should represent the financing economics of the strategy.

When square-root annualization works

With (q) observations per year, the common annualized estimate is:

SR_annual = sqrt(q) * SR_period

This scaling relies on additive differential returns with zero serial correlation and stable first and second moments. It is not a general identity for compounded returns. Daily business-time observations commonly use (q=252), while a continuously traded calendar-day series may use 365. The data frequency alone does not reveal which convention the analyst intended.

Serial correlation changes multi-period variance. If (\gamma_k) is the lag-(k) autocovariance of differential returns, the variance of an additive (q)-period return is:

variance_q = q * gamma_0 + 2 * sum((q - k) * gamma_k, k=1,...,q-1)

Positive autocorrelation increases long-horizon variance and makes naive square-root scaling too optimistic. Negative autocorrelation can have the opposite effect. Andrew Lo's study of Sharpe-ratio statistics derives the stationary-return conversion and shows why square-root scaling is valid only under special conditions.

This example creates five years of synthetic daily returns with positive serial correlation. It compares naive annualization with a heteroskedasticity-and-autocorrelation-consistent (HAC) estimate of long-run variance. The Bartlett window uses 20 lags. Lag choice is an estimation decision, not a universal constant.

import numpy as np


rng = np.random.default_rng(11)
n = 252 * 5
phi = 0.35
innovation = rng.normal(0.0, 0.009, n)
strategy = np.empty(n)
strategy[0] = 0.0005 + innovation[0]
for t in range(1, n):
    strategy[t] = (
        0.0005
        + phi * (strategy[t - 1] - 0.0005)
        + innovation[t]
    )

periods_per_year = 252
annual_rf = 0.04
daily_rf = (1 + annual_rf) ** (1 / periods_per_year) - 1
excess = strategy - daily_rf

naive = (
    np.sqrt(periods_per_year)
    * excess.mean()
    / excess.std(ddof=1)
)


def hac_sharpe(excess_returns, periods_per_year=252, max_lag=20):
    x = np.asarray(excess_returns, dtype=float)
    x = x[np.isfinite(x)]
    centered = x - x.mean()
    gamma0 = centered @ centered / len(x)
    long_run_var = gamma0
    for lag in range(1, max_lag + 1):
        gamma = centered[lag:] @ centered[:-lag] / len(x)
        weight = 1 - lag / (max_lag + 1)
        long_run_var += 2 * weight * gamma
    return (
        np.sqrt(periods_per_year)
        * x.mean()
        / np.sqrt(long_run_var)
    )


adjusted = hac_sharpe(excess)
rho1 = np.corrcoef(excess[1:], excess[:-1])[0, 1]

print(f"Daily risk-free rate: {daily_rf:.8f}")
print(f"Lag-1 autocorrelation: {rho1:.3f}")
print(f"Naive annualized Sharpe: {naive:.3f}")
print(f"HAC-adjusted annualized Sharpe: {adjusted:.3f}")
Daily risk-free rate: 0.00015565
Lag-1 autocorrelation: 0.354
Naive annualized Sharpe: 0.832
HAC-adjusted annualized Sharpe: 0.610

The conventional estimate is higher because it treats the positively correlated observations as independent. The HAC result is an estimator under a stationary-process approximation. It is not ground truth, and its sampling uncertainty remains substantial. The synthetic process demonstrates the calculation only. It is not a strategy or evidence of a tradable edge.

Why there is no universal good Sharpe ratio

Thresholds such as 1, 2, or 3 are descriptive customs, not decision rules. An estimate depends on sample length, observation frequency, asset class, leverage, liquidity, capacity, cost model, and strategy-selection history. A Sharpe of 2 from a short and heavily searched backtest can offer weaker evidence than a lower estimate from a prespecified strategy tested across several regimes.

Compare like with like and report uncertainty. At minimum, disclose:

  • the return series, currency, benchmark, and risk-free-rate source,
  • gross and net results, including fees, spread, slippage, financing, borrow costs, and funding,
  • observation frequency, annualization factor, missing-value policy, and standard-deviation convention,
  • sample start, end, length, and any overlapping returns,
  • serial-dependence treatment and confidence interval or resampling method,
  • every strategy and parameter search that influenced selection, and
  • out-of-sample and live results separately from development results.

For one prespecified independent and identically distributed normal-return model, the Sharpe estimate has a tractable sampling distribution. Real strategy returns often violate those assumptions. Use a dependence-aware bootstrap or appropriate asymptotic method when reporting uncertainty. If a result was selected from many trials, a one-strategy confidence interval is still incomplete. The deflated Sharpe ratio and multiple-testing bias pages explain selection-aware adjustments.

What the Sharpe ratio misses

The ratio compresses a return distribution into its mean and standard deviation. It does not describe the timing or shape of losses. Two strategies can have the same Sharpe and very different skewness, tail exposure, drawdowns, liquidity, and correlation with the rest of a portfolio.

The metric itself does not require normally distributed returns. Normality and independence enter many confidence intervals, significance tests, and simplified interpretations. Short-volatility strategies, smoothed marks, stale prices, and infrequently valued assets can report attractive Sharpes while hiding tail loss or understating measured volatility. Overlapping returns also manufacture dependence.

Constant leverage leaves an idealized excess-return Sharpe unchanged when both mean and volatility scale linearly. Real leverage adds financing, nonlinear transaction costs, margin constraints, liquidation risk, and changing exposures, so realized Sharpe need not remain invariant. A zero-volatility series also needs explicit handling. A positive mean divided by zero is not evidence of a riskless opportunity and often indicates rounded, stale, or incomplete data.

Read Sharpe alongside maximum drawdown, tail and stress measures, turnover, capacity, exposure, and portfolio correlation. The Sortino ratio changes the denominator to downside deviation, but it does not repair bad data, selection bias, or unrealistic execution assumptions.

How Python backtesting tools calculate it

Software can use the same label for different formulas. Check the implementation before comparing output.

Community VectorBT 1.1.0 subtracts a constant per-period risk_free value, computes the arithmetic mean and standard deviation with configurable ddof, and multiplies by the square root of its annualization factor. Its returns-accessor documentation defines risk_free as a constant periodic return. The current global year_freq default is 365 days, so freq="1D" implies 365 observations per year unless year_freq is overridden. For business-day data, pass year_freq="252 days" explicitly.

Backtesting.py 0.6.6 uses a different convention in its statistics report. Its current statistics implementation divides geometrically annualized return minus an annual risk_free_rate by a compounded annualized-volatility estimate. The source notes that this differs from the arithmetic-mean implementation used by Empyrical. The result is internally useful, but it should not be presented as numerically interchangeable with the formula at the top of this page.

Whatever the library, retain the net periodic returns and configuration used to produce the headline. A tear-sheet value without those inputs cannot be audited or compared reliably.

Choose which optional services may run. You can change these settings at any time.