Vectorized backtesting represents observations, assets, parameters, and scenarios as array axes, then applies operations across many values or strategy variants at once. This reduces repeated Python interpreter work and keeps parameter dimensions explicit. It does not mean that every loop disappears, every rule is stateless, or simulated fills are realistic.
Many tools use a hybrid design: vectorize independent calculations, run path-dependent cash and position rules in compiled sequential kernels, and use an event scheduler when the exact order of market updates and orders matters.
What vectorization actually changes
Consider an array with shape (time, asset, parameter). A price return with shape (time, asset, 1) can interact with positions for every parameter column. NumPy broadcasting treats compatible length-one dimensions as expanded without copying the smaller input.
The resulting operation still executes loops. Those loops run inside NumPy, pandas, Numba, Rust, or another compiled backend instead of dispatching one Python operation per bar and parameter. The usual gains come from:
- fewer Python function calls and object allocations,
- contiguous or predictable memory access,
- compiled loops and, where applicable, SIMD or native threads,
- batching many variants through one prepared simulation, and
- reusing aligned inputs and metadata.
The Python GIL is only one part of the story. A compiled serial loop can be fast without multiple threads. A vectorized expression can be slow if it allocates several huge temporary arrays or moves more memory than it computes.
Python example: a parameter grid without future data
This example broadcasts one score series against three entry thresholds. A score observed in row (t) controls the position in row (t+1), so the strategy does not earn the return used to create the signal. A five-basis-point cost is charged on every position change.
import numpy as np
asset_returns = np.array(
[0.010, -0.006, 0.012, 0.004, -0.009, 0.015, -0.003, 0.008]
)
score = np.array(
[0.002, 0.011, -0.004, 0.007, 0.015, -0.002, 0.009, 0.003]
)
thresholds = np.array([0.000, 0.005, 0.010])
# Shapes: (8, 1) > (3,) broadcasts to (8, 3).
signal = score[:, None] > thresholds
# A row-t signal becomes a row-(t+1) position.
position = np.vstack([
np.zeros((1, len(thresholds))),
signal[:-1],
]).astype(float)
position_changes = np.abs(
np.diff(
position,
axis=0,
prepend=np.zeros((1, len(thresholds))),
)
)
net_returns = (
position * asset_returns[:, None]
- position_changes * 0.0005
)
total_returns = np.prod(1 + net_returns, axis=0) - 1
# Verify the broadcast result against separate column calculations.
loop_results = []
for threshold in thresholds:
loop_signal = score > threshold
loop_position = np.r_[0.0, loop_signal[:-1]].astype(float)
loop_changes = np.abs(np.diff(loop_position, prepend=0.0))
loop_net = loop_position * asset_returns - loop_changes * 0.0005
loop_results.append(np.prod(1 + loop_net) - 1)
np.testing.assert_allclose(total_returns, loop_results)
print("grid shape:", signal.shape)
for threshold, changes, total in zip(
thresholds,
position_changes.sum(axis=0),
total_returns,
):
print(
f"threshold={threshold:.3f} "
f"position_changes={int(changes)} "
f"net_return={total:.2%}"
)
grid shape: (8, 3)
threshold=0.000 position_changes=5 net_return=1.74%
threshold=0.005 position_changes=5 net_return=2.35%
threshold=0.010 position_changes=4 net_return=2.51%
The vectorized and per-threshold calculations match. That proves implementation parity for this fixture, not a strategy edge. The final positions are not forcibly liquidated, the cost is a fixed assumption, and eight constructed returns provide no performance evidence.
Vectorization does not prevent state
Many time-series operations are stateful but have efficient array or scan implementations:
- cumulative products and running extrema,
- rolling and exponentially weighted statistics,
- position forward-filling,
- drawdown and exposure histories, and
- simple stops expressible from predetermined arrays.
Other rules contain recurrences that depend on simulated decisions: shared cash, partial fills, a trailing reference updated only while a position is open, portfolio-level risk gates, tax lots, or order conflicts. Trying to force these into independent array expressions can create circular logic or a forest of temporary masks.
Use a sequential kernel for such state. Numba can compile a loop that advances cash and positions row by row while columns represent independent strategy variants. Native Rust can do the same. This is still array-oriented research because inputs and outputs remain structured arrays. The internal recurrence is deliberately sequential.
An event scheduler adds another capability: heterogeneous events with explicit priority and timestamps. It is useful when quote, trade, timer, order, fill, cancel, funding, and account events arrive irregularly and affect one another. A regular compiled row scan and an event queue solve different scheduling problems.
You still need to prevent future data leaks
Whole arrays make future values easy to reference accidentally. Every predictor needs a knowledge time, every decision needs a submission time, and every fill needs an earliest executable time. Common failures include:
- using a completed bar's close to earn that same bar's return,
- centered rolling windows or backward-filled features,
- global normalization fitted on the full sample,
- resampling a higher-timeframe final value into earlier rows,
- selecting parameters on the period later labeled out of sample, and
- treating a touched bar price as guaranteed liquidity.
Shifting a signal by one row is correct only when one row matches the required delay. Auction cutoffs, intraday timestamps, weekends, and irregular data need explicit clocks. Use prefix and future-perturbation tests from the look-ahead bias review.
Memory is the limiting resource
Broadcast inputs can be views, but outputs and intermediate expressions often allocate the full result. A float64 array with:
10,000 times * 500 assets * 1,000 parameters * 8 bytes
= 40 GB
already exceeds ordinary workstation memory before signals, positions, records, and metrics are added. More dimensions can make a concise expression impractical.
Control memory by:
- using compatible dtypes and boolean arrays where semantics allow,
- avoiding unnecessary tiled inputs and repeated pandas objects,
- computing in place when safe,
- retaining orders or summary statistics instead of every intermediate,
- chunking independent parameter or asset groups,
- sampling or optimizing rather than materializing an exhaustive grid, and
- moving sequential work into a kernel that emits only required records.
Chunking must preserve dependencies. Parameter columns with independent cash can be split safely. Assets sharing one cash pool, cross-asset ranks, or portfolio constraints may need to remain in the same group. Carry rolling warm-up and simulation state across time chunks rather than restarting each block.
What VectorBT and VectorBT PRO add
Community VectorBT combines pandas-labeled outputs, parameter broadcasting, indicator factories, Numba-compiled portfolio simulation, and record analysis. It is well suited to signals and portfolio variants that fit its public simulation contracts. Array speed does not validate signal timing, universe history, fees, slippage, or fills.
VectorBT PRO 2026.9.5 develops the hybrid model much further:
- parameterization and labeled broadcasting preserve strategy coordinates through indicators, portfolios, records, and metrics,
- flexible selection can reuse scalar or lower-dimensional inputs without tiling every argument,
- signal, order, flexible-order, segment, and post-order callbacks can inspect evolving compiled portfolio state,
- chunking specifications split compatible functions and simulations, while execution engines can schedule independent chunks,
Portfolio.update()continues positions, record identifiers, pending execution settings, and stop state as new rows arrive, and- Numba and Rust accumulators support observation-by-observation computation alongside batch functions.
The official portfolio guide demonstrates compiled callbacks and portfolio continuation. The performance guide documents chunking and execution engines. These mechanisms make broad, stateful research spaces manageable, but they do not turn a bar into an order book or eliminate model-selection bias.
Native Rust without Python
The current VectorBT PRO Rust crate is an independent compute library, not only a hidden Python extension. Its supported algorithms mirror Numba contracts, and typed portfolio strategies can process a full array or step through rows while retaining state. Rust applications can use the crate without installing or embedding Python. The official Python-to-Rust tutorial shows high-level Python, Numba, PyO3-backed Rust, native batch simulation, and native stepping.
This matters when moving supported algorithms to native Rust. Compiler-checked types, builder arguments, and explicit errors make a port easier to inspect. A Rust process can handle data ingestion and simulation, then export records for analysis. An arbitrary Python callback does not automatically become Rust, and matching supported Numba algorithms does not make a strategy correct. Pin matching versions and compare results on controlled tests.
Choose the simplest design that fits
| Requirement | Useful starting point |
|---|---|
| Cross-sectional transforms, indicators, simple positions, broad parameter grids | Vectorized array operations |
| Cash, stops, fills, and other regular path-dependent state over array rows | Compiled sequential simulation |
| Irregular event priority, broker lifecycle, asynchronous feeds, or detailed books | Event-driven simulation |
| Broad discovery followed by execution validation | Array and compiled research, then calibrated event replay |
| Streaming one observation at a time without a Python runtime | Native stateful Rust stepper |
These are not quality rankings. A compiled event engine can process data quickly, and a hybrid vectorized engine can model substantial path dependence. Execution realism comes from timestamps, data, venue rules, and calibrated assumptions, not from the architecture's label. See vectorized versus event-driven backtesting for the direct workflow comparison.