A useful Python backtest states exactly which data it uses, when it makes decisions, how orders fill, what trading costs, and how cash and positions change. The code can be short. The assumptions should not be hidden.
This walkthrough implements the same long-only strategy in VectorBT and Backtesting.py. A signal is calculated after bar t closes and may fill at bar t+1 open. Both examples charge trading costs and expose their orders for inspection.
The example uses synthetic data so it is reproducible. Its return is not evidence that the strategy has an edge.
1. Write down the trading rules first
Before coding, answer these questions:
| Decision | Rule used here |
|---|---|
| Information time | Signal sees completed close bars through time t |
| Earliest fill | Open at t+1 |
| Position | Long or flat, one asset |
| Sizing | 95% of available cash on entry |
| Fees | 0.10% of notional on every fill |
| Spread or slippage | Configured explicitly in each framework |
| Unfilled orders | Not modeled because every market order fills in full |
| Valuation | Remaining cash plus shares marked at each close |
| Boundary state | Cash and an open position continue across the split |
This is deliberately simple. It is unsuitable for a strategy whose result depends on limit-order queues, partial fills, volume, borrow, funding, market impact, or intrabar stops. The order-types guide and transaction-cost guide explain what a more detailed model must add.
2. Check the data before generating signals
Real data needs stable instrument identifiers, timestamps, timezones, session calendars, missing-data policy, corporate actions, and point-in-time universe membership. Fundamental and alternative data need both observation time and the time the value became knowable. Today's restated value does not belong in a historical decision.
For each column, record:
- Its source, units, timezone, and adjustment policy.
- Whether the timestamp marks the start or end of an interval.
- The publication or arrival delay used by the strategy.
- How duplicates, gaps, bad ticks, delistings, and symbol changes are handled.
- A content hash or immutable version for reproducing the run.
Never backfill an unknown historical value from the future. Forward filling can also be wrong when a stale observation should invalidate a signal.
3. Create data and signals once
Both frameworks will use the same synthetic daily prices and the same Bollinger-style rule. The strategy enters when the close crosses below the lower band and exits when it crosses back above the moving average.
import numpy as np
import pandas as pd
rng = np.random.default_rng(17)
n = 1_000
index = pd.bdate_range("2021-01-04", periods=n)
log_returns = rng.normal(0.00015, 0.012, n)
close = pd.Series(100 * np.exp(np.cumsum(log_returns)), index=index)
open_price = close.shift(1) * (1 + rng.normal(0, 0.0015, n))
open_price.iloc[0] = close.iloc[0]
# Backtesting.py expects a complete OHLC table.
wiggle = close * 0.004
prices = pd.DataFrame(
{
"Open": open_price,
"High": np.maximum(open_price, close) + wiggle,
"Low": np.minimum(open_price, close) - wiggle,
"Close": close,
"Volume": 100_000,
},
index=index,
)
window, band_width = 20, 2.0
mean = close.rolling(window).mean()
std = close.rolling(window).std(ddof=1)
lower = mean - band_width * std
entries = (close < lower) & (close.shift(1) >= lower.shift(1))
exits = (close > mean) & (close.shift(1) <= mean.shift(1))
Synthetic data keeps the example self-contained. Its return says nothing about whether this strategy has an edge.
4. Run it with VectorBT
VectorBT works directly with the signal arrays. Shifting the Boolean signals by one row makes an event observed at today's close eligible at tomorrow's open. Passing price=open_price tells the portfolio where those orders should fill.
import vectorbt as vbt
portfolio = vbt.Portfolio.from_signals(
close,
entries.shift(1, fill_value=False),
exits.shift(1, fill_value=False),
price=open_price,
size=0.95,
size_type="percent",
init_cash=10_000,
fees=0.001,
slippage=0.0005,
freq="1D",
)
print(
portfolio.stats()[
["Total Return [%]", "Total Trades", "Max Drawdown [%]"]
]
)
print(portfolio.orders.records_readable.head())
Total Return [%] 10.302209
Total Trades 18
Max Drawdown [%] 18.728051
The portfolio records are as important as the summary. Check that the first signal fills on the next date, that fees are present, and that each exit closes the expected position.
5. Run the same idea with Backtesting.py
Backtesting.py expresses the strategy as a class. next() runs after each completed bar. A market order placed there fills at the next bar's open unless trade_on_close=True is set.
from backtesting import Backtest, Strategy
def rolling_mean(values, window):
return pd.Series(values).rolling(window).mean()
def lower_band(values, window, width):
values = pd.Series(values)
return (
values.rolling(window).mean()
- width * values.rolling(window).std(ddof=1)
)
class BollingerCross(Strategy):
window = 20
band_width = 2.0
def init(self):
self.mean = self.I(rolling_mean, self.data.Close, self.window)
self.lower = self.I(
lower_band,
self.data.Close,
self.window,
self.band_width,
)
def next(self):
close = self.data.Close
crossed_below = (
close[-1] < self.lower[-1]
and close[-2] >= self.lower[-2]
)
crossed_above = (
close[-1] > self.mean[-1]
and close[-2] <= self.mean[-2]
)
if not self.position and crossed_below:
self.buy(size=0.95)
elif self.position and crossed_above:
self.position.close()
backtest = Backtest(
prices,
BollingerCross,
cash=10_000,
commission=0.001,
spread=0.001,
exclusive_orders=True,
finalize_trades=True,
)
stats = backtest.run()
print(stats[["Return [%]", "# Trades", "Max. Drawdown [%]"]])
print(stats["_trades"].head())
backtest.plot()
Return [%] 10.213351
# Trades 18
Max. Drawdown [%] -18.672634
The results are close but not identical. VectorBT applies the stated slippage to each fill, while Backtesting.py applies a constant spread. The frameworks also have their own rounding, sizing, and reporting conventions. That difference is useful: it shows why copying a strategy into a second framework is not enough. You must compare its trades.
6. Check the trades before reading the scores
The two framework runs are only a start. Add small examples with answers you can calculate by hand:
- A three-bar entry where the next open and exact fee are known.
- An exit where sell-side slippage and proceeds can be calculated on paper.
- A signal on the last bar, which must remain unfilled.
- A position held across the train/test boundary.
- Two simultaneous assets competing for insufficient cash if multi-asset support is added.
- Missing and duplicated timestamps that must fail validation rather than pass silently.
Reconcile cash, quantity, fees, position, and equity after every fixture. Compare order records between implementations before comparing summary returns. Similar Sharpe ratios can conceal different trades.
The example assumes every market order fills completely. Realistic research may also need bid and ask data, volume participation, latency, partial fills, price limits, minimum notional, borrow, funding, dividends, taxes, FX conversion, and counterfactual market impact. State each omission.
7. Interpret each metric precisely
Total return measures the change in portfolio value over the stated segment. Maximum drawdown measures the worst observed peak-to-trough decline, not the largest possible future loss. The Sharpe ratio in the example annualizes daily arithmetic excess returns with a zero benchmark and an independence-style square-root rule. Serial dependence and irregular exposure can make that convention misleading.
Always report the period, frequency, benchmark or cash rate, number of trades, exposure, turnover, gross and net results, and cost assumptions. Add drawdown duration, tail outcomes, capacity, and per-asset attribution when relevant. Do not select a strategy from one ratio.
8. Keep the final test truly separate
The examples above use the full series to demonstrate the two APIs. For actual research, choose a chronological split before tuning the strategy. Every feature, parameter, threshold, and model selected from data belongs inside the development period. Opening the final period and then revising the strategy turns that period into development data too.
Use rolling or expanding evaluation when the deployment procedure refits through time. Preserve legitimate indicator history, open positions, cash, fees, and turnover across boundaries unless the stated experiment intentionally resets them. Apply purging according to label information intervals when samples overlap. The overfitting workflow covers nested selection, search logs, and final holdouts.
9. Compare framework results before trusting them
Framework defaults are part of the model. Before porting, map every field explicitly:
| Setting | Questions to answer |
|---|---|
| Signal delay | Does a Boolean at t fill at close t, open t+1, or another configured price? |
| Costs | Are fees and slippage charged per order, per side, per share, or as notional percentages? |
| Sizing | Does 1.0 mean one unit, 100% of cash, or 100% of equity? |
| Cash sharing | Are columns independent or competing inside one portfolio? |
| Orders | Can orders reject, remain open, partially fill, or collide within a bar? |
| Boundary state | Are warm-up, cash, positions, stops, and order IDs carried between chunks? |
VectorBT is especially convenient when you want to test many signals, assets, or parameter combinations together. Backtesting.py is easier to read when a strategy makes decisions one bar at a time, and its interactive chart is useful for inspecting individual trades. Neither is automatically more correct. The right choice depends on the strategy, and both require explicit timing and cost settings.
If the strategy later needs order books, partial fills, latency, or a direct path to live trading, consider a more detailed engine such as NautilusTrader. The Python backtesting frameworks guide explains the wider set of choices.
When two engines disagree, do not explain the difference from intuition. Export signals, intended orders, accepted orders, fills, cash, positions, and valuation by timestamp. Find the first divergence and reduce it to a deterministic fixture.
Backtest completion checklist
- Data and knowledge timestamps are documented and versioned.
- Signals are causal and the decision-to-fill delay is explicit.
- Fees, spread, slippage, financing, and capacity match the research question.
- Cash, positions, orders, and equity pass hand-calculated fixtures.
- Metrics state their frequency, benchmark, estimator, and sample.
- The full search is logged and selection occurs inside development data.
- A final chronological or prospective evaluation remains untouched until freeze.
- Paper trading and small live tests have their own checks.
A backtest is only as trustworthy as its data, trading rules, accounting, and record of what you tried.