Skip to content
python.financial

On a standard CPython build, the Global Interpreter Lock (GIL) allows only one thread to execute Python bytecode in an interpreter at a time. This limits CPU scaling for pure-Python loops. It does not prevent useful I/O concurrency, stop native code from running outside the lock, or apply to a standalone Rust or C# process.

The GIL is a CPython implementation mechanism, not a rule of the Python language. It protects interpreter state, but it is not a substitute for application locks. Multi-step updates to orders, positions, caches, and risk state can still race whenever execution switches threads or native code releases the lock.

First, find out what kind of work is slow

Workload What usually dominates Suitable starting point
Broker, database, and market-data I/O Waiting and coordination asyncio or threads with bounded queues
Pure-Python indicator or simulation loop Bytecode and object overhead under the GIL Better algorithm, arrays, Numba, or native code
Independent backtests with large Python objects CPU plus serialization and memory Processes, shared memory, or chunked native kernels
NumPy, BLAS, Numba, or Rust kernel Native implementation, memory bandwidth, internal threads Measure library threading and avoid oversubscription
Live event loop with Python callbacks Callback time, locks, allocation, I/O, logging, and scheduling Profile tail latency and isolate slow work

Threads can overlap blocking I/O because CPython releases the GIL around blocking operations. Native extensions can also detach while performing work that does not access Python objects. The CPython thread-state documentation gives compression and hashing as examples.

A NumPy expression may therefore use native code or an internally threaded BLAS implementation even though the caller is Python. Conversely, a library implemented in C or Rust may keep the GIL around a particular call. Check the actual function or source instead of inferring behavior from the implementation language.

Python and Numba speed example

This benchmark runs four independent, state-dependent integer kernels. The dependency prevents parallelizing iterations inside one task. It then compares sequential and threaded calls for pure Python and for a warmed Numba function compiled with nogil=True.

import os
import platform
import sys
import time
from concurrent.futures import ThreadPoolExecutor

import numba

TASKS = 4
ITERATIONS = 4_000_000


def python_kernel(seed):
    state = seed
    for i in range(ITERATIONS):
        state = (state * 1_664_525 + i + 1_013_904_223) & 0xFFFFFFFF
    return state


@numba.njit(nogil=True)
def numba_kernel(seed):
    state = seed
    for i in range(ITERATIONS):
        state = (state * 1_664_525 + i + 1_013_904_223) & 0xFFFFFFFF
    return state


def measure(function, threaded):
    seeds = list(range(1, TASKS + 1))
    start = time.perf_counter()
    if threaded:
        with ThreadPoolExecutor(max_workers=TASKS) as executor:
            result = list(executor.map(function, seeds))
    else:
        result = [function(seed) for seed in seeds]
    return time.perf_counter() - start, result


numba_kernel(0)
py_serial, expected = measure(python_kernel, threaded=False)
py_threads, threaded_python = measure(python_kernel, threaded=True)
nb_serial, compiled = measure(numba_kernel, threaded=False)
nb_threads, threaded_compiled = measure(numba_kernel, threaded=True)
assert expected == threaded_python == compiled == threaded_compiled

print(f"Python {sys.version_info.major}.{sys.version_info.minor}.{sys.version_info.micro}")
print(f"machine={platform.machine()} logical_cpus={os.cpu_count()}")
print(f"checksum={sum(expected)}")
print(f"pure Python: serial={py_serial:.3f}s threads={py_threads:.3f}s")
print(f"Numba nogil: serial={nb_serial:.3f}s threads={nb_threads:.3f}s")

One observed run with Python 3.11.8, Numba 0.65.0, and an eight-logical-CPU arm64 machine produced:

Python 3.11.8
machine=arm64 logical_cpus=8
checksum=7902759434
pure Python: serial=1.480s threads=1.391s
Numba nogil: serial=0.019s threads=0.005s

The checksum establishes output parity across all four paths. The pure-Python threads did not provide meaningful multicore scaling in this run. The warmed compiled calls were much faster and overlapped across threads because nogil=True released the lock.

These timings are not portable benchmarks. CPU frequency, thermal state, compiler version, thread startup, background work, task size, and optimization all matter. Compilation was deliberately warmed and excluded. Run repeated trials on the target machine, report distributions, and verify results before comparing throughput.

Releasing the GIL and using several cores are different

Numba's nogil option permits other Python threads to run while a supported compiled function executes. It does not parallelize that function by itself. parallel=True and numba.prange are separate mechanisms for parallel work inside a compiled call.

This distinction prevents two common mistakes:

  • a GIL-free serial kernel can overlap with other independent GIL-free calls
  • a parallel kernel can create its own worker pool, which may oversubscribe the machine if outer threads or processes also run it

Control Numba, BLAS, OpenMP, Rayon, process, and executor thread counts as one system. More concurrency can reduce performance through memory-bandwidth saturation, cache contention, context switching, allocation, and nested pools.

Threads, processes, and native code have different costs

Threads share memory and have low handoff cost, but standard CPython serializes Python bytecode. They fit I/O and native calls that explicitly release the GIL. Shared state still needs synchronization.

Processes have separate interpreters and therefore separate GILs. They can run CPU-bound Python concurrently, but arguments and results may need serialization and copying. Startup method, copy-on-write behavior, native thread pools, large arrays, and worker failure recovery matter. Use shared memory or memory-mapped immutable data when measurement shows serialization dominates.

Native extensions can release the GIL around long-running work. Crossing the Python/native boundary, converting arrays, allocating outputs, and invoking Python callbacks can still cost time or reacquire the lock.

A native service or executable avoids the Python runtime on its own hot path. That can simplify latency and concurrency control, but interprocess communication, deployment, observability, and language boundaries remain engineering costs.

What changes in free-threaded CPython

Python 3.13 introduced an optional free-threaded CPython build. Python 3.14 moved the work to the supported phase described by PEP 779, but the ordinary GIL-enabled build remains the default. Free threading is not a flag that turns a standard interpreter into a different binary. The interpreter and extension wheels use a distinct t ABI.

On a free-threaded build, multiple threads can execute Python bytecode concurrently. Check the build with sysconfig.get_config_var("Py_GIL_DISABLED") and the current runtime state with sys._is_gil_enabled(). The GIL can be forced on, and importing an extension that has not declared free-threading support can re-enable it with a warning. The official free-threading guide and extension guide document these boundaries.

Free threading does not make shared mutable strategy state safe. Explicit locks, ownership, immutable messages, and deterministic ordering become more important once bytecode can truly overlap. The Python 3.14 guide reports workload-dependent single-thread overhead averaging roughly 1% on macOS arm64 to 8% on x86-64 Linux in its pyperformance measurements. Measure the actual installed dependencies because extension compatibility and memory use may dominate.

Do not plan a production migration around a guessed year when free threading might become the default. Test the optional build, verify every binary dependency, confirm that the GIL stays disabled after imports, run race detection and stress tests, and compare end-to-end latency rather than isolated Python loops.

How trading libraries use these options

VectorBT PRO 2026.9.5 registers its Numba backend with nogil=True and parallel=False by default. Compatible functions can opt into internal parallelism. Its execution layer also offers serial, thread, process, Pathos, MPIRE, and Dask engines, so independent calls can use a scheduler suited to their data and serialization costs. A thread engine benefits GIL-free native work or I/O, not arbitrary Python callbacks.

The optional VectorBT PRO Rust extension detaches from Python around supported native kernels and can use serial or Rayon paths. The standalone vectorbtpro-rust crate is more fundamental: a Rust application can use it without Python, so there is no Python GIL in that process. Python wrappers, conversions, unsupported callbacks, and analysis around the kernel still have their own costs.

VectorBT gains most of its speed from array operations and Numba-compiled simulation rather than from Python threads. Do not assume every community-package call releases the GIL or uses multiple cores. Inspect and benchmark the selected backend and call.

NautilusTrader places core data, event, order, and execution machinery in Rust. Python strategy and actor callbacks still execute as Python when the Python API is used. A native core reduces Python work and can run native tasks independently, but it does not turn user Python callbacks into GIL-free Rust.

Checklist

Profile first and separate I/O wait, Python bytecode, native kernels, serialization, allocation, garbage collection, and lock contention. Benchmark cold and warm paths with realistic data sizes. Verify whether each extension releases or re-enables the GIL, record all thread-pool sizes, and test nested execution for oversubscription.

For live systems, measure latency percentiles and queue growth under bursts, not only average throughput. Protect shared portfolio and risk state, make event ordering explicit, bound worker queues, propagate failures, and rehearse shutdown and recovery. The GIL is one concurrency constraint. Removing it does not remove the need for a concurrency design.

Choose which optional services may run. You can change these settings at any time.