Skip to content
python.financial

A useful Python backtest states exactly which data it uses, when it makes decisions, how orders fill, what trading costs, and how cash and positions change. The code can be short. The assumptions should not be hidden.

This walkthrough implements the same long-only strategy in VectorBT and Backtesting.py. A signal is calculated after bar t closes and may fill at bar t+1 open. Both examples charge trading costs and expose their orders for inspection.

The example uses synthetic data so it is reproducible. Its return is not evidence that the strategy has an edge.

1. Write down the trading rules first

Before coding, answer these questions:

Decision Rule used here
Information time Signal sees completed close bars through time t
Earliest fill Open at t+1
Position Long or flat, one asset
Sizing 95% of available cash on entry
Fees 0.10% of notional on every fill
Spread or slippage Configured explicitly in each framework
Unfilled orders Not modeled because every market order fills in full
Valuation Remaining cash plus shares marked at each close
Boundary state Cash and an open position continue across the split

This is deliberately simple. It is unsuitable for a strategy whose result depends on limit-order queues, partial fills, volume, borrow, funding, market impact, or intrabar stops. The order-types guide and transaction-cost guide explain what a more detailed model must add.

2. Check the data before generating signals

Real data needs stable instrument identifiers, timestamps, timezones, session calendars, missing-data policy, corporate actions, and point-in-time universe membership. Fundamental and alternative data need both observation time and the time the value became knowable. Today's restated value does not belong in a historical decision.

For each column, record:

  • Its source, units, timezone, and adjustment policy.
  • Whether the timestamp marks the start or end of an interval.
  • The publication or arrival delay used by the strategy.
  • How duplicates, gaps, bad ticks, delistings, and symbol changes are handled.
  • A content hash or immutable version for reproducing the run.

Never backfill an unknown historical value from the future. Forward filling can also be wrong when a stale observation should invalidate a signal.

3. Create data and signals once

Both frameworks will use the same synthetic daily prices and the same Bollinger-style rule. The strategy enters when the close crosses below the lower band and exits when it crosses back above the moving average.

import numpy as np
import pandas as pd

rng = np.random.default_rng(17)
n = 1_000
index = pd.bdate_range("2021-01-04", periods=n)
log_returns = rng.normal(0.00015, 0.012, n)
close = pd.Series(100 * np.exp(np.cumsum(log_returns)), index=index)
open_price = close.shift(1) * (1 + rng.normal(0, 0.0015, n))
open_price.iloc[0] = close.iloc[0]

# Backtesting.py expects a complete OHLC table.
wiggle = close * 0.004
prices = pd.DataFrame(
    {
        "Open": open_price,
        "High": np.maximum(open_price, close) + wiggle,
        "Low": np.minimum(open_price, close) - wiggle,
        "Close": close,
        "Volume": 100_000,
    },
    index=index,
)

window, band_width = 20, 2.0
mean = close.rolling(window).mean()
std = close.rolling(window).std(ddof=1)
lower = mean - band_width * std
entries = (close < lower) & (close.shift(1) >= lower.shift(1))
exits = (close > mean) & (close.shift(1) <= mean.shift(1))

Synthetic data keeps the example self-contained. Its return says nothing about whether this strategy has an edge.

4. Run it with VectorBT

VectorBT works directly with the signal arrays. Shifting the Boolean signals by one row makes an event observed at today's close eligible at tomorrow's open. Passing price=open_price tells the portfolio where those orders should fill.

import vectorbt as vbt

portfolio = vbt.Portfolio.from_signals(
    close,
    entries.shift(1, fill_value=False),
    exits.shift(1, fill_value=False),
    price=open_price,
    size=0.95,
    size_type="percent",
    init_cash=10_000,
    fees=0.001,
    slippage=0.0005,
    freq="1D",
)

print(
    portfolio.stats()[
        ["Total Return [%]", "Total Trades", "Max Drawdown [%]"]
    ]
)
print(portfolio.orders.records_readable.head())
Total Return [%]    10.302209
Total Trades                18
Max Drawdown [%]     18.728051

The portfolio records are as important as the summary. Check that the first signal fills on the next date, that fees are present, and that each exit closes the expected position.

5. Run the same idea with Backtesting.py

Backtesting.py expresses the strategy as a class. next() runs after each completed bar. A market order placed there fills at the next bar's open unless trade_on_close=True is set.

from backtesting import Backtest, Strategy


def rolling_mean(values, window):
    return pd.Series(values).rolling(window).mean()


def lower_band(values, window, width):
    values = pd.Series(values)
    return (
        values.rolling(window).mean()
        - width * values.rolling(window).std(ddof=1)
    )


class BollingerCross(Strategy):
    window = 20
    band_width = 2.0

    def init(self):
        self.mean = self.I(rolling_mean, self.data.Close, self.window)
        self.lower = self.I(
            lower_band,
            self.data.Close,
            self.window,
            self.band_width,
        )

    def next(self):
        close = self.data.Close
        crossed_below = (
            close[-1] < self.lower[-1]
            and close[-2] >= self.lower[-2]
        )
        crossed_above = (
            close[-1] > self.mean[-1]
            and close[-2] <= self.mean[-2]
        )

        if not self.position and crossed_below:
            self.buy(size=0.95)
        elif self.position and crossed_above:
            self.position.close()


backtest = Backtest(
    prices,
    BollingerCross,
    cash=10_000,
    commission=0.001,
    spread=0.001,
    exclusive_orders=True,
    finalize_trades=True,
)
stats = backtest.run()
print(stats[["Return [%]", "# Trades", "Max. Drawdown [%]"]])
print(stats["_trades"].head())
backtest.plot()
Return [%]           10.213351
# Trades                    18
Max. Drawdown [%]   -18.672634

The results are close but not identical. VectorBT applies the stated slippage to each fill, while Backtesting.py applies a constant spread. The frameworks also have their own rounding, sizing, and reporting conventions. That difference is useful: it shows why copying a strategy into a second framework is not enough. You must compare its trades.

6. Check the trades before reading the scores

The two framework runs are only a start. Add small examples with answers you can calculate by hand:

  1. A three-bar entry where the next open and exact fee are known.
  2. An exit where sell-side slippage and proceeds can be calculated on paper.
  3. A signal on the last bar, which must remain unfilled.
  4. A position held across the train/test boundary.
  5. Two simultaneous assets competing for insufficient cash if multi-asset support is added.
  6. Missing and duplicated timestamps that must fail validation rather than pass silently.

Reconcile cash, quantity, fees, position, and equity after every fixture. Compare order records between implementations before comparing summary returns. Similar Sharpe ratios can conceal different trades.

The example assumes every market order fills completely. Realistic research may also need bid and ask data, volume participation, latency, partial fills, price limits, minimum notional, borrow, funding, dividends, taxes, FX conversion, and counterfactual market impact. State each omission.

7. Interpret each metric precisely

Total return measures the change in portfolio value over the stated segment. Maximum drawdown measures the worst observed peak-to-trough decline, not the largest possible future loss. The Sharpe ratio in the example annualizes daily arithmetic excess returns with a zero benchmark and an independence-style square-root rule. Serial dependence and irregular exposure can make that convention misleading.

Always report the period, frequency, benchmark or cash rate, number of trades, exposure, turnover, gross and net results, and cost assumptions. Add drawdown duration, tail outcomes, capacity, and per-asset attribution when relevant. Do not select a strategy from one ratio.

8. Keep the final test truly separate

The examples above use the full series to demonstrate the two APIs. For actual research, choose a chronological split before tuning the strategy. Every feature, parameter, threshold, and model selected from data belongs inside the development period. Opening the final period and then revising the strategy turns that period into development data too.

Use rolling or expanding evaluation when the deployment procedure refits through time. Preserve legitimate indicator history, open positions, cash, fees, and turnover across boundaries unless the stated experiment intentionally resets them. Apply purging according to label information intervals when samples overlap. The overfitting workflow covers nested selection, search logs, and final holdouts.

9. Compare framework results before trusting them

Framework defaults are part of the model. Before porting, map every field explicitly:

Setting Questions to answer
Signal delay Does a Boolean at t fill at close t, open t+1, or another configured price?
Costs Are fees and slippage charged per order, per side, per share, or as notional percentages?
Sizing Does 1.0 mean one unit, 100% of cash, or 100% of equity?
Cash sharing Are columns independent or competing inside one portfolio?
Orders Can orders reject, remain open, partially fill, or collide within a bar?
Boundary state Are warm-up, cash, positions, stops, and order IDs carried between chunks?

VectorBT is especially convenient when you want to test many signals, assets, or parameter combinations together. Backtesting.py is easier to read when a strategy makes decisions one bar at a time, and its interactive chart is useful for inspecting individual trades. Neither is automatically more correct. The right choice depends on the strategy, and both require explicit timing and cost settings.

If the strategy later needs order books, partial fills, latency, or a direct path to live trading, consider a more detailed engine such as NautilusTrader. The Python backtesting frameworks guide explains the wider set of choices.

When two engines disagree, do not explain the difference from intuition. Export signals, intended orders, accepted orders, fills, cash, positions, and valuation by timestamp. Find the first divergence and reduce it to a deterministic fixture.

Backtest completion checklist

  • Data and knowledge timestamps are documented and versioned.
  • Signals are causal and the decision-to-fill delay is explicit.
  • Fees, spread, slippage, financing, and capacity match the research question.
  • Cash, positions, orders, and equity pass hand-calculated fixtures.
  • Metrics state their frequency, benchmark, estimator, and sample.
  • The full search is logged and selection occurs inside development data.
  • A final chronological or prospective evaluation remains untouched until freeze.
  • Paper trading and small live tests have their own checks.

A backtest is only as trustworthy as its data, trading rules, accounting, and record of what you tried.

Choose which optional services may run. You can change these settings at any time.