Skip to content
python.financial

Choose an array-centered design when assets, parameters, scenarios, or splits are the main dimensions of the research question. Choose an event-driven design when the order and timing of market, clock, order, fill, and account events matters most. Use a hybrid when broad numerical research and chronological state are both important.

Vectorized and event-driven are implementation descriptions, not quality rankings. Vectorization does not require stateless or unrealistic portfolios. Event scheduling does not guarantee realistic fills or live parity. Modern libraries frequently precompute arrays, run compiled sequential scans, and process selected events in the same workflow.

At a glance

Criterion Array-centered or vectorized design Event-driven design What actually determines the result
Primary organization Dense arrays indexed by time, asset, parameter, scenario, or split Chronological queue of market, timer, order, account, and system events Whether the representation matches the dependencies in the strategy
Typical strength Broad numerical comparison, broadcasting, batch transforms, and population analysis Irregular timing, asynchronous sources, order lifecycle, and explicit causal transitions The exact data, state, and model exposed to the strategy
Sequential state Compiled scans, callbacks, accumulators, and portfolio kernels Strategy, broker, venue, cache, and event-handler objects State ownership and update order, not the architecture label
Parameter scaling Add dimensions, broadcast, compile, chunk, and reduce Repeat jobs with different configuration, often in parallel Array size, per-run setup, retained outputs, dependencies, and hardware
Execution modeling Can model rich order and portfolio rules inside compiled loops Naturally represents commands, acknowledgements, fills, cancels, and timers Data granularity, fill rules, liquidity, latency, and calibration
Data resolution Commonly bars and arrays, but custom quote or trade arrays are possible Bars, ticks, quotes, depth, and external events fit naturally The engine cannot reconstruct information absent from the input
Validation Parameter surfaces and split dimensions can remain labeled Each training or test range is often a separate job or lifecycle Causal preprocessing, purge horizon, trial history, and untouched evidence
Live reuse Batch and streaming state can share kernels, but brokerage may be separate The same event and order abstractions can drive backtest and live modes External reconciliation, network behavior, state recovery, and adapter fidelity
Debugging Inspect whole arrays, labels, records, and compiled callback outputs Trace event order, object state, commands, and responses Deterministic fixtures and observable intermediate state
Common failure Look-ahead through precomputed inputs, memory explosion, accidental independent-column assumptions Hidden defaults, slow repeated jobs, race or ordering mistakes, optimistic venue models Unstated timing, state, cost, and data assumptions

Vectorized code processes many values together

A vectorized expression applies an operation to whole arrays rather than dispatching a Python function for each scalar. A moving average, comparison, return, or statistic can use compiled NumPy loops underneath. Broadcasting aligns smaller inputs across larger shapes without a user-written outer loop.

In quantitative research, columns can represent assets and strategy variants. Additional index levels can preserve window lengths, thresholds, costs, split names, or scenarios. This keeps a research population inspectable as one labeled object instead of producing a directory of unrelated runs.

Vectorization does not mean that every algorithm is parallel across time. Cash, positions, stops, drawdowns, and inventory depend on earlier states. Array-centered engines handle these recurrences with a sequential kernel, often compiled with Numba or Rust. The input and output remain arrays even though the internal calculation advances in order.

VectorBT and VectorBT PRO illustrate this hybrid. Signals and parameters broadcast as arrays, while portfolio constructors process path-dependent state in compiled loops. PRO adds callback hooks, portfolio continuation, streaming accumulators, and native Rust strategy traits without giving up labeled results.

Event-driven code processes one event at a time

An event-driven system places messages on a chronological path. Market data, clock events, order commands, venue responses, fills, account changes, and system events can arrive at different times and trigger different handlers. Strategy code reads current state and emits actions in response.

This maps naturally to live trading because real systems receive asynchronous messages. An order is not merely an array value. It is submitted, accepted or rejected, modified or canceled, partially or fully filled, and reconciled with an external venue. Actors, engines, caches, brokers, and adapters can own different parts of that lifecycle.

NautilusTrader is a detailed example with typed events, simulated venues, L1 through L3 market data, order management, reconciliation, and live adapters. LEAN uses chronological data slices, securities, orders, brokerage models, and live integrations. Backtesting.py is a smaller event-style bar engine whose next() method sees progressively revealed candles.

Event-driven does not imply one slow Python callback per tick. Engines can be written in Rust or C#, batch compatible indicators, use optimized queues, and parallelize independent backtests. The defining property is explicit chronological dispatch and interaction among components.

Both designs can remember earlier events

Consider a cooldown that starts only after a stop order actually fills. The rule cannot be computed correctly from an entry signal alone because rejection, cash, another order, or a different exit can change the path.

An event-driven strategy can update a cooldown counter when it receives the relevant fill event. An array-centered simulator can run a compiled post-order callback, inspect the emitted order record, and update per-column memory. Both are causal when the callback runs after the execution decision and the future cannot influence the state.

The difference is ergonomics and scale. The event engine exposes the order event as a native system object. The array engine can apply the same callback across many assets and parameters while retaining labeled outputs. A pure vector expression is insufficient, but a vectorized research system need not be pure.

Classify each rule before selecting an engine:

  • independent transforms such as returns, rolling statistics, and signal comparisons fit batch arrays,
  • recurrences such as cash, inventory, stops, and cooldowns need a sequential scan,
  • irregular interactions such as acknowledgements, reconnects, external orders, and timers fit event scheduling, and
  • portfolio comparisons across many assets, parameters, and splits benefit from labeled batching.

Realism comes from data and trading rules

An event loop replaying daily close prices is not more realistic than a compiled portfolio loop using the same closes. Neither knows the spread, intrabar path, available size, queue priority, latency, rejection behavior, or market response.

Realism improves when the simulation has relevant information and rules:

  1. causal timestamps and a clear decision-to-order delay,
  2. quotes or trades when spread and sequence matter,
  3. L2 or L3 data when depth and queue assumptions matter,
  4. instrument definitions, calendars, corporate actions, and contract changes,
  5. fill eligibility, partial liquidity, cancellations, and rejections,
  6. fees, spread, slippage, funding, borrow, financing, and nonlinear impact,
  7. account, margin, leverage, liquidation, and portfolio constraints, and
  8. calibration against observed paper or live executions.

An event engine often provides better abstractions for the later items. An array engine can still express many of them in a compiled simulator. In both cases, a historical order book remains immutable and cannot show the counterfactual reaction to the simulated order.

Architecture affects what is convenient to model. It does not prove that the configured model is accurate.

Signal timing and look-ahead

Array code can accidentally use the current close to create a position earning the same bar's return. Precomputing a full indicator can also leak future observations through centered windows, backfills, normalization, or model fitting. Shift signals to the first executable timestamp and fit transformations inside training ranges.

Event scheduling reduces some accidental access because a handler receives the current event frontier. It does not eliminate look-ahead. A custom dataset may contain revised values, a feature can use a future label, a universe can be selected with current constituents, or a simulated fill can use a price that was not yet available.

Write an information ledger for either design. For every input, record when it became knowable, when the strategy consumes it, when an order can reach the market, and which later return the position earns. Test the first few rows by hand.

Speed depends on the job

Array-centered systems often excel when the same operation applies to dense data over many assets or parameter combinations. Broadcasting reduces Python orchestration. Numba and Rust move sequential work into compiled loops. Chunking can trade time for memory and parallelize independent dimensions.

Event-driven systems often excel when events are sparse or irregular and only a small part of state changes at each step. They avoid materializing dense arrays for events that did not occur. A compiled event core can process large chronological streams efficiently.

Neither has constant cost. Array work grows with elements, retained intermediates, records, and memory traffic. Event work grows with events, handlers, objects, queue operations, and repeated job setup. Shared cash or portfolio-wide rules can prevent naive parallelism in either architecture.

Benchmark equal outputs:

  • include data preparation and first-call compilation,
  • report cold and warm runs,
  • measure peak memory and serialized output size,
  • retain the same orders, metrics, and logs,
  • use the same assets, parameters, and execution assumptions, and
  • separate one-run latency from total parameter-search throughput.

A benchmark of a simple moving average cannot decide the architecture for an L3 execution strategy, and one event-driven run cannot decide a 100,000-variant research workload.

Parameter research and validation

Array-centered tools can attach parameter and split labels directly to results. This is valuable for response surfaces, robustness neighborhoods, cross-asset stability, and distributional analysis. VectorBT PRO integrates rolling, purged, embargoed, and combinatorial split workflows with parameterized objects.

Event-driven engines usually execute each configuration as a complete run. Managed or local orchestration can parallelize those jobs. This is appropriate when each candidate needs full engine initialization, complex event state, or granular market replay.

The statistical standard is identical. Record every parameter, feature, market, cost, split, and manual redesign. Keep preprocessing inside training data. Match purging to label horizons. Do not treat overlapping folds as independent. Reserve prospective evidence after selecting the full process.

Faster research increases the number of hypotheses that can be tried. It must be paired with stronger parameter robustness and multiple-testing controls.

Memory, sparsity, and data shape

Dense arrays are effective when most time-asset cells contain meaningful values and the same operations apply. A shape with one million rows, hundreds of assets, and thousands of variants can exceed memory even before simulation records are stored. Use appropriate dtypes, cache shared inputs, chunk independent axes, reduce early, and avoid materializing broadcast views unnecessarily.

Event queues can be more natural for sparse asynchronous data. One instrument may update while another remains unchanged. Corporate actions, news, timers, and broker messages do not fit a rectangular OHLC table cleanly. The engine advances only through present events, though caches and book state still consume memory.

Sparse matrices, columnar event stores, compiled iterators, and hybrid batches blur this distinction. Select the representation that wastes less information and less computation for the actual data.

Backtest-to-live reuse

Event-driven engines can reuse strategy, order, and portfolio abstractions between historical and live nodes. This reduces translation risk. Live operation still differs through wall-clock timing, network latency, external positions, partial history, data-provider differences, exchange behavior, restarts, and human activity.

Array-centered tools can also update more than static signals. Streaming accumulators can update features, and portfolio continuation can carry cash, positions, stops, record identifiers, and custom state across incoming chunks. VBT PRO's standalone Rust crate can run native steppers without Python.

Those capabilities do not supply broker routing or reconciliation. A production service needs an external truth source, idempotent order handling, credential controls, persistence, monitoring, alerts, hard risk limits, and tested recovery.

Do not select architecture on the phrase backtest/live parity. Select it by the exact state and interfaces reused, then test the remaining differences.

Combining both often works best

A research pipeline can combine several execution forms without becoming inconsistent:

  1. batch raw-data cleaning and feature calculation,
  2. broadcast assets and parameters for broad exploration,
  3. compiled sequential scans for positions, stops, and cash,
  4. chronological split application for retraining and validation,
  5. granular event replay for selected execution-sensitive candidates,
  6. streaming accumulators for incoming observations, and
  7. an event-driven live service for broker and venue interaction.

The same product can implement several stages. VBT PRO spans the first six and can move compatible computation into native Rust. NautilusTrader and LEAN span historical event replay and live engine operation. A project can also connect separate tools, but then must reconcile semantics explicitly.

Using two engines is not automatically more rigorous. Porting can change timestamps, sizes, cash sharing, costs, corporate actions, and order rules. Reconcile orders and state on a deterministic fixture before interpreting performance differences.

How to choose

Favor an array-centered research system when most of these are true:

  • the hypothesis is naturally expressed over dense arrays,
  • many assets, parameters, scenarios, or splits need comparison,
  • response surfaces and labeled population analysis matter,
  • bar-level or explicitly modeled portfolio execution is adequate,
  • compiled sequential callbacks cover the path dependence, and
  • research throughput and memory controls are central.

Favor an event-driven engine when most of these are true:

  • heterogeneous events and their exact ordering drive decisions,
  • quotes, trades, depth, timers, or asynchronous feeds matter,
  • order acknowledgements, partial fills, cancels, and reconciliation are first-class,
  • the target is a backtest and live trading system with shared domain objects,
  • the number of research variants is modest or separately orchestrated, and
  • the team can run the required data feeds and adapters.

Favor a hybrid when broad discovery, path-dependent portfolios, granular execution validation, and live operation all matter. Define the boundary between batch arrays, compiled state, historical events, and external live state before choosing products.

A small test before you choose

  1. Freeze one causal dataset, instrument definition, starting cash, fees, and timing convention.
  2. Implement one simple rule and reconcile every signal, order, fill, cash movement, position, and metric.
  3. Add one rule that depends on a prior fill and confirm both engines update state at the same point.
  4. Stress a gap, limit, stop, rejection, insufficient-cash, and final-open-position case.
  5. Scale across the intended assets and parameter grid and measure cold time, warm time, peak memory, and output retention.
  6. Add the most granular available data and identify which execution conclusions actually change.
  7. Run the declared validation process and retain every tested variant.
  8. Continue or restart state, then compare with a one-shot historical reference and any external account state.

The best architecture is the one that represents the project's important dependencies with the fewest hidden assumptions. It may be array-centered, event-driven, or deliberately both.

Choose which optional services may run. You can change these settings at any time.