Look-ahead bias occurs when a simulated decision uses information or an execution price that was unavailable when that decision would have been made. The test is not whether a value happened on the same calendar date. The test is whether the strategy could have known and acted on that exact value before the assumed order deadline.
Look-ahead bias can make a losing rule look exceptional, but it can also create smaller distortions that survive a casual review. Chronological train and test splits do not repair a feature, universe, or fill that already contains future information.
Decision time and knowledge time
Every simulated action needs a decision timestamp. Every input needs a knowledge timestamp: the earliest time that exact value could have reached the strategy after publication and processing delays. A row labeled 2025-03-31 might represent a quarter ending on that date, a filing released weeks later, or a later revision. Only the release or availability timestamp answers whether the value was usable.
For market data, distinguish at least:
- The interval covered by a bar.
- Whether the bar timestamp denotes its opening or closing boundary.
- The event time at the venue.
- The receipt time at the strategy after feed latency.
- The order cutoff and the earliest executable price after the decision.
For fundamentals and alternative data, also preserve effective dates, publication timestamps, vendor arrival timestamps, and revision versions. A point-in-time database retains what was known then rather than replacing history with the latest restatement.
Common leakage channels
Future shifts and non-causal windows
shift(-1), forward returns, centered rolling windows, future extrema, and backfilled observations explicitly move future values backward. These operations are valid for labels and later evaluation, but not as predictor values at the same timestamp.
Whole-sample preprocessing
Scaling, winsorization, imputation, feature selection, principal components, and model fitting leak when their parameters use later observations. Fit each learned transform only on the relevant training window and then apply it forward. Scikit-learn's data-leakage guidance documents this boundary for preprocessing and pipelines.
Incorrect joins and revisions
Joining an earnings value on its fiscal period end exposes it before release. Joining the latest revised value rewrites the historical record. For irregular releases, a backward pandas.merge_asof can attach the last release at or before a decision time. Its official documentation also supports a tolerance and strict inequality when an equal timestamp was not yet actionable. The input still needs a truthful release timestamp.
Same-bar signals and fills
A daily bar's close is unknown until the bar completes. A rule that consumes that close and fills at the same close needs a real order mechanism with a decision cutoff before the final price is known. Otherwise, use the next executable event. The same issue applies to high, low, volume, VWAP, and finalized indicators derived from the bar.
Even an explicit one-row shift can be wrong when timestamps represent different interval boundaries, markets close at different times, or a slower time frame is aligned before its bar completes.
Future universes and corporate actions
Applying today's index constituents to earlier dates leaks future survival and membership. Split, dividend, symbol, delisting, and merger handling can also reveal or erase later events when a vendor retroactively rewrites data. Survivorship bias overlaps with look-ahead bias when future membership determines the historical sample, but it is useful to audit universe construction separately.
Selection and evaluation leakage
Using the final test period to choose features, thresholds, parameters, or preprocessing is leakage across the research process even if each individual signal is causal. Keep model development separate from genuinely untouched evaluation as described in in-sample versus out-of-sample testing.
Python example: a leak and its fix
This example creates 500 independent synthetic daily returns. The leaked rule decides whether to hold during day t by inspecting day t's completed return. The causal rule can see only the return completed by the previous close. Cash earns zero, trades have no costs, and both rules use the same return path.
import numpy as np
rng = np.random.default_rng(9)
market_returns = rng.normal(0.0002, 0.012, 500)
# At the start of day t, this illegally inspects day t's eventual return.
leaked_position = (market_returns > 0).astype(float)
# This position uses only the return completed by the previous close.
honest_position = np.r_[0.0, (market_returns[:-1] > 0).astype(float)]
def summarize(position):
strategy_returns = position * market_returns
sharpe = np.sqrt(252) * strategy_returns.mean() / strategy_returns.std(ddof=1)
total_return = np.prod(1 + strategy_returns) - 1
return sharpe, total_return
for label, position in [("Leaked", leaked_position), ("Causal", honest_position)]:
sharpe, total_return = summarize(position)
print(f"{label}: Sharpe={sharpe:.2f}, total return={total_return:.1%}")
Leaked: Sharpe=11.39, total return=1135.0%
Causal: Sharpe=-0.47, total return=-14.3%
The leaked result is mechanical: it holds every positive day and skips every negative day. Its performance is proof of the bug built into this example, not a threshold that identifies leakage in general. A biased strategy can have an ordinary Sharpe ratio, and a high Sharpe ratio alone does not locate the cause.
The causal rule is not automatically realistic. It still omits commissions, spread, slippage, latency, and the exact price available after the decision. Removing one future reference fixes only one layer.
Tests that expose future dependence
Use several tests because no single check covers every channel:
- Write an information contract. For each feature, record source time, release time, revision policy, processing delay, decision time, and earliest fill time.
- Run a prefix test. Compute a feature on data ending at time
T, then on the full dataset. Values at or beforeTshould be identical unless the method is intentionally revised. - Perturb the future. Change observations after
T. Earlier predictors and decisions must not change. - Trace one order by hand. Show the raw values known at decision time and the first price eligible for execution.
- Delay signals. A one-event delay is not a universal fix, but a dramatic collapse warrants an alignment investigation.
- Replay vintages. Rebuild several dates from stored raw snapshots or vendor point-in-time queries.
- Audit folds end to end. Fit imputation, scaling, feature selection, models, and thresholds inside each training fold.
- Fail on impossible timestamps. Reject a feature whose knowledge time is later than its decision time rather than silently filling or shifting it.
A framework cannot prevent every leak
An event-driven engine processes the events it receives in an order, but it cannot validate every value placed in those events or stop strategy code from reading a preloaded future array. NautilusTrader documents timestamp-ordered deterministic backtest processing. That ordering helps reproduce an event sequence, while data timestamps, bar boundaries, callbacks, and fill assumptions remain the researcher's responsibility.
Vectorized research is also not inherently biased. Explicitly aligned arrays can be causal, and event-driven code can leak. Community VectorBT makes shifts and order timing visible in array form, which is useful for inspection but leaves the information contract to the user.
VectorBT PRO adds time-aware realignment for opening and closing bar boundaries. Its public safe-resampling example shows realign_closing for multi-time-frame indicators. The current source also marks its future-looking label functions as label-only because using them as predictors can introduce look-ahead bias. These features reduce common alignment mistakes, but they do not certify external release timestamps, revisions, or execution assumptions.
The durable fix is an auditable rule: every predictor must have knowledge_time <= decision_time, and every simulated fill must occur at a price that was still available after the order became actionable.