Skip to content
python.financial

Backtesting, paper trading, and live trading answer different questions. A backtest runs your rules on old data. Paper trading runs them on current data with fake money. Live trading sends real orders and can lose real money. Doing well at one stage does not prove that the next stage will work.

Stage Main question What it tests What it cannot prove
Backtest Did the rules behave acceptably on unseen historical periods? Data alignment, signal logic, sizing, simulated orders, costs, and portfolio accounting Current operations, real routing, queue position, market impact, or future returns
Paper trading Can the system run correctly in real time? Live data, schedules, state, reconnects, simulated orders, alerts, and operator procedures Real fill quality, information leakage, borrow availability, market impact, or capital controls
Controlled live trading Can the complete system operate within defined limits? Real orders, rejects, partial fills, fees, latency, reconciliation, and venue/account behavior Performance at a larger size or in market regimes that have not occurred

The useful progression is not "backtest, then trust." It is a sequence of increasingly expensive tests with explicit reasons to stop.

What a backtest can establish

A backtest replays historical data through a set of rules. It is the cheapest stage for rejecting weak ideas, testing accounting, comparing parameter regions, and measuring sensitivity to fees and execution assumptions. It should include point-in-time inputs, causal signal timing, realistic order rules, and out-of-sample evaluation.

A favorable result means the implementation produced acceptable simulated behavior under the tested data and assumptions. It does not demonstrate a persistent edge. Historical simulations remain exposed to look-ahead bias, survivorship bias, overfitting, and incomplete transaction-cost and slippage models.

Before advancing, retain the configuration, code revision, input-data version, random seed, and trade or order records. Reproducibility is a minimum gate. It lets you distinguish a changed strategy from a changed dataset or simulator later.

Before moving to paper trading

Advance to real-time testing only when all of these statements are true:

  • Signals use only information available at the simulated decision time.
  • The result survives a genuinely untouched period and reasonable cost stress.
  • Performance is not concentrated in one instrument, short interval, or isolated parameter point without an economic explanation.
  • Position sizing, cash, leverage, fees, funding, borrow, and order timing match the intended market closely enough for the decision being made.
  • Individual orders and portfolio totals reconcile, including rejected or unfilled orders when the simulator supports them.

For strategies sensitive to spread, depth, or intrabar sequencing, daily or bar-close data may be insufficient. The event-driven simulation guide explains why a more detailed engine still needs matching and latency assumptions.

Paper trading tests your bot in real time

"Paper trading" covers several different arrangements. Record which one you used because their evidence is not interchangeable.

Paper mode Where orders go Best use Important blind spot
Local dry run The application simulates orders locally Strategy loop, data, state, scheduling, and alerts Broker API and order lifecycle may be bypassed
Broker paper account A broker's separate simulation endpoint Authentication, request validation, broker events, and reconciliation code Fills still follow a simulator's rules
Exchange testnet or sandbox A venue's test environment API integration and venue-specific messages Participants and liquidity may not resemble the live market
Shadow mode Signals run on live data but order submission is blocked Comparing intended orders with the live tape No fill or account-state exercise

Paper trading should expose stale feeds, missed schedules, reconnect bugs, duplicate-order risks, bad symbol mappings, clock errors, state loss, and unusable alerts. It can also reveal a mismatch between backtest signals and signals computed incrementally from live data.

It cannot reproduce execution merely because it uses current quotes. Alpaca's paper-trading specification explicitly excludes effects such as market impact, information leakage, latency slippage, and queue position. Freqtrade likewise warns that exchange sandboxes can have unlike liquidity and trading behavior in its official dry-run guidance.

A concrete dry-run configuration

Freqtrade can run the same strategy definition in backtest, dry-run, and live modes. The surrounding data and execution behavior still differ. Keep dry-run state separate from live state and make the mode obvious in the configuration:

{
  "dry_run": true,
  "dry_run_wallet": 10000,
  "db_url": "sqlite:///tradesv3.dryrun.sqlite",
  "stake_currency": "USDT",
  "stake_amount": 100,
  "max_open_trades": 3,
  "exchange": {
    "name": "kraken",
    "key": "",
    "secret": "",
    "pair_whitelist": ["BTC/USDT"]
  }
}
freqtrade trade --config config.json --strategy MyStrategy --dry-run

This example uses simulated money and a dedicated dry-run database. It does not validate the strategy or predict returns. Check the current Freqtrade command reference and exchange requirements before running it.

Before moving to live trading

Do not promote a system based only on paper profit. First verify operations over enough observations to cover ordinary restarts, quiet periods, bursts, session boundaries, and relevant order types:

  • Every signal can be traced to its inputs, code revision, and intended order.
  • Restarts recover positions, open orders, timers, and strategy state without duplication.
  • Disconnects, stale data, rejects, partial fills, and canceled orders have tested responses.
  • Paper orders reconcile with the application's positions and cash ledger.
  • Alerts reach an operator, and the operator has a tested pause, cancel, and shutdown procedure.
  • The difference between modeled fills and contemporaneous market data is measured rather than ignored.

Live trading tests the complete execution chain

Live trading adds actual order routing, venue rules, account permissions, real fills, and irreversible loss. It is where unexpected rejects, partial fills, rate limits, fee tiers, borrow constraints, funding, queue priority, and market impact become observable. A paper environment cannot certify any of them.

The first live deployment should be a controlled experiment, not a performance target. Use the smallest practical exposure, explicit per-order and aggregate limits, restricted API credentials, independent monitoring, and a tested kill procedure. Compare each intended order with the acknowledged order, every fill with the venue record, and the internal position with the broker or exchange position. Stop on unexplained divergence.

Increasing capital is another model change. Slippage and impact can scale nonlinearly, so a system that behaves correctly at small size has not demonstrated capacity at a larger size. Revisit the backtest and paper assumptions whenever instruments, venues, frequency, leverage, or size changes.

Shared code prevents some mistakes, but the stages still differ

Using the same strategy code at each stage reduces copying mistakes. It does not make the stages equal. Historical and live data differ. Fake fills and exchange fills differ. Live trading also adds networks, saved state, account checks, and activity outside your bot.

NautilusTrader is designed around shared backtest and live components. Its documentation says that the same strategy and execution-algorithm code can run across backtest, sandbox, and live environments, while also stating that live execution adds behavior a simulation may not reproduce. See its backtest and live design boundary. Freqtrade and Lumibot also span simulation and live deployment, with different venue coverage and execution models.

VectorBT PRO serves a different part of the workflow. It can explore large parameter spaces, run walk-forward and purged validation, continue a portfolio simulation as new observations arrive, and turn real fill records into portfolio analytics. Its portfolio continuation preserves cash, positions, order identifiers, and stops across updates. It does not connect to a broker, manage live orders, or reconcile the live account by itself, so actual trading still needs a separate execution layer.

How to investigate disagreement between stages

When results deteriorate, compare the stages at the order level before changing the strategy:

  1. Match signal timestamps and input values. A mismatch usually points to data alignment, warm-up, timezone, or incremental-calculation differences.
  2. Match intended orders. Differences often come from sizing, cash, precision, minimum notional, or state recovery.
  3. Match acknowledgements and fills. Separate rejects, latency, spread, queueing, partial fills, fees, and impact.
  4. Rebuild PnL from fills and cash movements. This catches accounting, funding, borrow, corporate-action, and currency-conversion errors.
  5. Change one assumption at a time and replay the smallest case that reproduces the divergence.

Use the stages as a ladder. Backtests help you reject bad ideas. Paper trading shows whether the bot runs correctly. A small live test shows how real orders behave. None of them guarantees future profits.

Choose which optional services may run. You can change these settings at any time.