Skip to content
python.financial

Numba is a just-in-time (JIT) compiler for numerical Python. It specializes a supported Python function for concrete argument types and compiles that specialization to native machine code. This can make explicit loops and stateful backtest kernels fast without moving the surrounding research workflow out of Python.

JIT compilation is not a general accelerator for an entire application. Python object handling, pandas orchestration, I/O, data alignment, plotting, and unsupported functions stay outside the compiled kernel. Performance depends on how much repeated numerical work crosses that boundary.

Lazy specialization by argument type

With no explicit signature, @numba.njit compiles lazily on the first call. A later compatible call reuses the specialization. A materially different signature, such as a different array dtype, can trigger another compilation. Numba's JIT reference documents both lazy and explicit-signature modes.

import numba
import numpy as np


@numba.njit
def moving_average(values, window):
    out = np.full(values.size, np.nan)
    running_sum = 0.0
    for i in range(values.size):
        running_sum += values[i]
        if i >= window:
            running_sum -= values[i - window]
        if i >= window - 1:
            out[i] = running_sum / window
    return out


rng = np.random.default_rng(0)
values64 = rng.normal(size=1_000_000)
print(f"signatures before calls: {len(moving_average.signatures)}")

result64 = moving_average(values64, 20)
print(f"after float64 call: {len(moving_average.signatures)}")

moving_average(values64, 20)
print(f"after repeated float64 call: {len(moving_average.signatures)}")

values32 = values64.astype(np.float32)
moving_average(values32, 20)
print(f"after float32 call: {len(moving_average.signatures)}")

print(f"last mean: {result64[-1]:.6f}")
signatures before calls: 0
after float64 call: 1
after repeated float64 call: 1
after float32 call: 2
last mean: -0.112932

This executed example shows specialization count rather than publishing a machine-specific speedup. On the validation machine, the first float64 call included hundreds of milliseconds of compilation while the warmed call took about two milliseconds. Those timings are not portable across processors, Numba and LLVM versions, power states, or workloads.

An explicit signature can compile eagerly when the decorated function is defined. That moves rather than removes compilation cost and restricts accepted types. Array dimensionality, dtype, layout, and mutability can all affect signature selection.

Nopython mode and supported code

@njit is an alias for JIT compilation in nopython mode. Since Numba 0.59, nopython is also the default for @jit. The compiled function must use types and operations Numba can lower without ordinary Python object execution. Supported numerical scalars, NumPy arrays, loops, branches, selected NumPy operations, tuples, and typed containers work well. Arbitrary Python classes and pandas objects generally do not.

Convert at the boundary and preserve labels outside the kernel:

values = series.to_numpy(dtype=np.float64)
result = moving_average(values, 20)
aligned = series._constructor(result, index=series.index, name="ma_20")

Conversion can copy data. Non-contiguous views, mixed dtypes, object arrays, and repeated pandas-to-NumPy conversion can erase part of the gain. Inspect dtype, shape, strides, contiguity, and allocation behavior before blaming the loop.

Object-mode escapes and objmode blocks exist for integration cases, but crossing back to the interpreter inside a hot loop usually defeats the reason to compile it. Keep kernels narrow, typed, and testable.

Compilation, caching, and reproducible timing

There are three different costs:

  1. Python import and data preparation.
  2. Compilation for a new specialization.
  3. Warm execution of an existing specialization.

Benchmark them separately. Call once before timing steady state, verify output against a simple reference implementation, use representative array sizes and memory layouts, and report software versions and hardware. Do not divide first-call time by warm-call time and present the result as a universal speedup.

cache=True saves eligible compiled functions to disk, which can reduce compilation in later processes. Numba's compiler documentation explains when that saved code may become stale, including changes in imported functions and global constants. File location, permissions, package upgrades, CPU compatibility, and containers also affect reuse. A cache can save time, but it does not prove the result is correct.

Long-running services often warm required signatures during startup. Multiprocessing jobs need deliberate warm-up and cache placement because each process has its own runtime state. A flood of slightly different dtypes and layouts can create compilation latency and memory pressure.

GIL release and parallel execution are explicit choices

Compiling in nopython mode does not by itself promise that surrounding Python threads run concurrently. Numba releases the Python global interpreter lock for a compiled call when nogil=True is requested and compilation supports it, as described in the official nogil documentation. Code then needs the same race and synchronization review as other multithreaded native code.

parallel=True enables supported automatic parallel transformations, and numba.prange expresses parallel loops. More threads can be slower for small arrays, memory-bound work, nested thread pools, or oversubscribed processes. Measure serial and parallel paths under the actual scheduler and control thread counts in shared environments.

Options such as fastmath=True can relax floating-point semantics. That may change NaN, infinity, reassociation, or reproducibility behavior. Bounds checking and error models also affect safety and performance. Trading calculations should test numerical parity and edge cases before enabling aggressive options.

Where Numba fits in backtesting

Numba is most useful when a kernel performs repeated numerical work that ordinary NumPy cannot express cleanly:

  • Stateful order simulation and accounting loops.
  • Rolling or expanding algorithms with custom state.
  • Parameter and asset loops over homogeneous arrays.
  • Signal cleaning, record generation, and reductions.
  • Monte Carlo or bootstrap routines with controlled random streams.

It is less useful for network requests, database access, sparse one-off calls, pandas-heavy joins, or logic dominated by Python objects. Vectorization and JIT compilation are complementary. Arrays define the data model, while compiled loops handle sequential logic inside it.

Community VectorBT follows this pattern: pandas and labeled arrays form the public workflow, while NumPy, Numba, and optional Rust kernels perform supported numerical operations. Callback-based Python logic still has backend limits. PyBroker uses NumPy and Numba for internal calculations while exposing stateful bar callbacks at the strategy layer.

VectorBT PRO: Numba parity with a native Rust path

VectorBT PRO keeps Numba as a complete Python-facing numerical backend and adds vectorbtpro-rust from the same codebase in two forms:

  • A version-matched PyO3 extension that compatible Python calls can select through the jitting registry. Supported calls can prefer Rust and fall back to Numba when Rust is unavailable or unsupported.
  • A regular Rust crate that builds as an rlib and can be used directly with ndarray. Native Rust users do not enable the optional python feature, so neither Python nor PyO3 is required at runtime.

The current crate contract is deliberately mirrored. Public Rust builders have Numba twins and PyO3 wrappers, and the source tree repeats the Python module structure across base operations, data, generic algorithms, indicators, labels, OHLCV, portfolio simulation, records, returns, signals, and utilities. API-contract tests check names, defaults, enums, and record layouts. Python parity tests judge Rust behavior against Numba. The public Rust backend overview describes automatic dispatch, and the native Rust API exposes the crate surface.

"Matches Numba" refers to supported algorithm contracts, not the ability to execute arbitrary Python callbacks in Rust. Fixed kernels can have direct Numba and Rust implementations. Dynamic native simulation uses typed Rust strategies and closures rather than sending a Python callback through an ABI boundary. The Rust crate also provides streaming and Rust-owned callback engines where a direct static-kernel mapping would be the wrong abstraction.

This matters for Rust teams because the same tested algorithms can run inside a native service without embedding Python. Cargo's type checker, builders, generated API documentation, contract tests, and matching module structure also make generated code easier to search and check. Those properties reduce uncertainty during integration, but the result still needs tests for financial meaning, data timing, costs, and version compatibility.

The native API is a low-level kernel API versioned with VectorBT PRO, not an independently stable high-level Rust SDK. The crate is distributed with VectorBT PRO rather than through the public crates.io registry, and callers should pin the matching release and ndarray version.

Production checklist

Before relying on a compiled kernel:

  1. Test it against a readable reference across empty, short, NaN, infinity, dtype, and boundary cases.
  2. Measure compilation and warm execution separately.
  3. Fix and record dtype, layout, versions, thread counts, and random seeds.
  4. Inspect conversions and allocations at every Python-to-native boundary.
  5. Decide whether file caching is valid for the deployment and dependencies.
  6. Verify GIL, threading, process, and nested-parallel behavior under load.
  7. Compare Numba and Rust outputs when changing backends.
  8. Keep financial assumptions and point-in-time tests independent of performance tests.

Native speed can make a research loop more productive. It does not make the underlying strategy, data, or execution assumptions correct.

Choose which optional services may run. You can change these settings at any time.