Walk-forward optimization (WFO) repeatedly selects or fits a strategy on past data, freezes the resulting procedure, and evaluates it on the next chronological segment. The process then advances and repeats. Non-overlapping out-of-sample segments can form one simulated deployment path when cash, positions, and trading costs carry across their boundaries correctly.
The result estimates the performance of the complete re-estimation procedure. It does not prove that the selected parameters are stable, that the strategy has an edge, or that each test fold is statistically independent. Once a researcher uses the combined test result to redesign the procedure, that result has become development information rather than an untouched final test.
Robert Pardo's Walk-Forward Analysis chapter describes evaluating optimization with data outside the optimization window. The same chronological structure is called rolling-origin evaluation in forecasting. Hyndman and Athanasopoulos' time-series cross-validation guide emphasizes that each forecast must use only observations before its origin.
Walk-forward optimization step by step
For each fold:
- Define a training interval that ends before the next test interval.
- Fit every learned transform inside that training interval. This includes feature selection, imputation, scaling, universe selection, model fitting, probability calibration, and parameter selection.
- Apply any required gap or label-aware purge at the boundary.
- Freeze the selected pipeline, including data rules, parameters, risk sizing, and execution assumptions.
- Run it once on the next test segment without using test outcomes to change that fold.
- Advance the origin and repeat according to a predeclared schedule.
- Combine non-overlapping test segments with continuous portfolio accounting.
A six-fold rolling design with 160 training observations, 40 test observations, and a 40-observation step has this shape:
fold 1: train [ 80, 239] -> test [240, 279]
fold 2: train [120, 279] -> test [280, 319]
fold 3: train [160, 319] -> test [320, 359]
...
Each test begins after its training interval. Adjacent test segments do not overlap, so each timestamp contributes once to the deployment path. The training intervals overlap, which makes fold results dependent. Six folds are therefore not six independent experiments.
Rolling, expanding, and anchored windows
The window design encodes assumptions about how quickly information ages.
| Design | Training data in the next fold | Useful when | Main cost |
|---|---|---|---|
| Rolling | Drops the oldest block and adds the newest block | A fixed recent history matches the intended retraining policy | Discards data and can make estimates unstable |
| Expanding | Keeps the original start and adds new observations | All earlier observations remain relevant enough to retain | Old regimes can dominate a changing process |
| Anchored or event-based | Starts or ends at a meaningful market or policy boundary | Deployment is tied to known structural dates | Boundary selection can become another fitted parameter |
Neither rolling nor expanding is automatically better. Training length, test horizon, step size, and retraining frequency are hyperparameters. If they were chosen after inspecting the walk-forward result, their search belongs in the reported experiment count.
Test length should reflect the intended deployment interval and the amount of evidence needed to evaluate it. A one-day test closely imitates daily refitting but can be expensive and creates highly dependent fits. A long test supplies more observations per frozen model but assumes slower adaptation. Calendar windows can be clearer than equal row counts when markets have holidays, missing sessions, or mixed frequencies.
Python example
This example generates 480 synthetic business-day returns with three different autoregressive regimes. It compares momentum lookbacks of 5, 20, and 60 observations. At each fold, it selects the highest net training Sharpe, freezes that lookback for the next 40 observations, and then advances by 40 observations.
The signal is shifted by one row, so the position for date t uses returns only through t - 1. The first training interval starts after an 80-observation history, which gives every lookback a warm-up. Five basis points are charged per unit of turnover. A flip from -1 to +1 is two units and therefore costs 10 basis points.
import numpy as np
import pandas as pd
rng = np.random.default_rng(42)
n = 480
phi = np.repeat([0.65, -0.40, 0.40], [160, 160, 160])
returns = np.zeros(n)
noise = rng.normal(0.0, 0.008, n)
for t in range(1, n):
returns[t] = 0.0002 + phi[t] * returns[t - 1] + noise[t]
returns = pd.Series(
returns,
index=pd.date_range("2023-01-02", periods=n, freq="B"),
name="return",
)
lookbacks = [5, 20, 60]
cost_per_unit_turnover = 0.0005
# Every signal for date t uses returns only through t - 1.
positions = pd.DataFrame(
{
lookback: np.sign(returns.rolling(lookback).sum()).shift(1)
for lookback in lookbacks
}
).fillna(0.0)
def net_returns(position, asset_returns):
turnover = position.diff().abs().fillna(position.abs())
return position * asset_returns - cost_per_unit_turnover * turnover
candidate_returns = pd.DataFrame(
{
lookback: net_returns(positions[lookback], returns)
for lookback in lookbacks
}
)
train_length = 160
test_length = 40
first_train_start = 80 # leaves history for every candidate before fold 1
oos_position = pd.Series(np.nan, index=returns.index, name="position")
fold_rows = []
for train_start in range(
first_train_start,
n - train_length - test_length + 1,
test_length,
):
train = slice(train_start, train_start + train_length)
test = slice(train.stop, train.stop + test_length)
train_slice = candidate_returns.iloc[train]
train_score = (
np.sqrt(252)
* train_slice.mean()
/ train_slice.std(ddof=1)
)
selected = int(train_score.idxmax())
# Freeze the selected lookback throughout this untouched test segment.
oos_position.iloc[test] = positions[selected].iloc[test]
fold_rows.append(
{
"test_start": returns.index[test.start].date(),
"test_end": returns.index[test.stop - 1].date(),
"lookback": selected,
"train_sharpe": train_score[selected],
}
)
oos_position = oos_position.dropna()
oos_returns = net_returns(oos_position, returns.loc[oos_position.index])
for row in fold_rows:
fold_return = oos_returns.loc[
str(row["test_start"]):str(row["test_end"])
]
row["test_return_pct"] = 100 * ((1 + fold_return).prod() - 1)
folds = pd.DataFrame(fold_rows)
print(
folds.round(
{"train_sharpe": 2, "test_return_pct": 2}
).to_string(index=False)
)
print(f"\nOOS observations: {len(oos_returns)}")
print(f"OOS total return: {(1 + oos_returns).prod() - 1:.2%}")
print(
"OOS annualized Sharpe: "
f"{np.sqrt(252) * oos_returns.mean() / oos_returns.std(ddof=1):.2f}"
)
test_start test_end lookback train_sharpe test_return_pct
2023-12-04 2024-01-26 5 1.01 -4.04
2024-01-29 2024-03-22 5 -1.11 -11.99
2024-03-25 2024-05-17 20 -1.57 -1.80
2024-05-20 2024-07-12 60 -1.33 2.86
2024-07-15 2024-09-06 60 -0.87 -3.51
2024-09-09 2024-11-01 5 0.51 12.52
OOS observations: 240
OOS total return: -7.39%
OOS annualized Sharpe: -0.48
The selected lookback changes across folds, but adaptation does not rescue this synthetic strategy. Its stitched test path loses 7.39%. That is a useful outcome because walk-forward optimization is an evaluation procedure, not a mechanism that guarantees profitable out-of-sample returns.
The code computes every candidate's causal position array in advance, which is safe here because the transform has no learned state and every row depends only on earlier returns. A learned scaler, imputer, feature selector, or model could not be fit once on the full array. It would need a new fit inside each training fold.
The example deliberately fixes the candidate set and window schedule. It always selects one candidate even when all training Sharpes are negative. A production rule could instead choose cash, an ensemble, or a stable parameter region, but that rule must be defined before evaluating the test path. The synthetic process, naive Sharpe score, linear cost, full fills, and absence of financing or market impact make this a mechanics check rather than evidence for momentum.
How to stitch the test segments correctly
Concatenating per-fold equity curves can silently invent performance. Build one chronological series of positions or orders first, then run portfolio accounting across it.
Preserve these boundary conditions:
- Non-overlapping tests: If test windows overlap, the same timestamp has several predictions. Do not compound them as separate realized returns. Define an ensemble or evaluate forecasts by origin instead.
- Indicator history: A test indicator may use pre-test observations for warm-up when those observations were already available. Resetting it at every boundary changes the live procedure. Fitting it with future test values leaks.
- Open positions: State whether positions persist, are resized, or are liquidated when parameters change. Forced liquidation needs an order and cost.
- Turnover: Charge the transition from the last position of one fold to the first target of the next. Calculating costs separately inside each fold can miss this rebalance.
- Cash and equity: The next segment should inherit prior cash, equity, margin, and open trade state. Restarting every test with fresh capital distorts sizing and compounding.
- Orders and fills: A signal calculated at a boundary cannot fill at an earlier price. Preserve the same decision-to-fill lag used within folds.
For stateful strategies, a full portfolio rerun using the final dated parameter schedule is often safer than joining isolated fold statistics. Verify that this rerun produces the same orders as the fold-by-fold process.
Gaps, purging, and label horizons
Chronological order does not by itself prevent leakage. Suppose a training label at date t uses the return through t + 5, while the test interval starts at t + 2. That training sample contains outcomes from the test period even though its feature timestamp is earlier.
A fixed gap removes a chosen number of observations between train and test. Scikit-learn's TimeSeriesSplit provides expanding training sets plus gap and max_train_size. Its documentation also requires equally spaced samples for comparable fold durations. A row-count gap is sufficient only when label horizons and observation spacing make it so.
Purging uses each sample's actual prediction and evaluation interval to remove overlaps. It is more precise when labels have variable horizons. An embargo addresses a different boundary and should not be substituted mechanically for a causal test. Neither technique repairs features built from revised or unavailable data.
All input data still need point-in-time correctness. Refit universe membership, corporate-action state, fundamental availability, normalization, and cross-sectional ranks using what was available at each origin. A perfect splitter cannot detect a vendor table that already contains future revisions.
Selection inside each training window
Choosing the single highest training score is simple but unstable. A practical selection rule can require minimum trades, feasible turnover, capacity, drawdown limits, and acceptable performance across subperiods. It can also prefer a stable neighborhood over an isolated optimum or combine several candidates.
Every choice belongs inside training data:
- parameter grid or search distribution,
- objective and tie-breaking rule,
- feature and model family,
- probability threshold and calibration,
- universe, portfolio construction, and risk target,
- transaction costs and capacity constraint, and
- training length, test length, gap, and retraining schedule.
If those choices themselves need tuning, use inner chronological validation within each outer training interval. The next outer test remains untouched. This nested design costs data and computation but separates selection from evaluation. After using outer walk-forward results to choose among several procedures, reserve a later holdout or prospective paper period for the final frozen procedure.
Report all candidates and all procedures that influenced the result. Repeating walk-forward studies until one aggregate curve looks attractive is multiple-testing bias with a more elaborate splitter.
What walk-forward optimization cannot fix
Walk-forward evaluation is useful for strategies that will actually refit or reselect over time. It does not fix:
- a strategy idea or dataset selected after seeing the same history,
- survivorship bias, revised fundamentals, bad timestamps, or incorrect corporate actions,
- optimistic fills, omitted spread, slippage, impact, funding, or borrow costs,
- too few trades or regimes to estimate the objective reliably,
- dependent fold returns caused by overlapping training data,
- a candidate family wide enough to find noise winners repeatedly,
- operational delays between training, approval, and deployable orders, or
- future structural changes absent from every historical fold.
A one-time frozen strategy may be better served by a development period plus one untouched final test. Walk-forward testing is most relevant when periodic retraining is part of the intended live policy. Match the evaluation procedure to deployment rather than adding folds for appearance.
Python framework support
PyBroker provides Strategy.walkforward(). It divides the requested data into equal-sized windows, splits each window according to train_size, trains registered models on the training rows, and runs the strategy on the test rows. Its lookahead excludes target-horizon rows at the boundary, and a dynamic symbol selector receives training data rather than test data.
That method does not make a fixed SMA period an optimized parameter merely because train_size is nonzero. Models, symbol selectors, and explicitly configured optimization logic have something to learn. Fixed rules do not. The former example on this page used fixed 10- and 30-day indicators, so describing its test partitions as parameter optimization was incorrect.
VectorBT PRO provides rolling, expanding, anchored, grouped, scikit-learn, purged walk-forward, and combinatorial split construction through Splitter. The official optimization feature guide shows Splitter.from_rolling, parameterized functions, random subsets, and overlap inspection. Labeled parameter grids can calculate all training candidates, while split-aware application can select and evaluate each fold without manual outer loops.
Those components expose the research design rather than choosing it. The user still defines which set is training, how a winner or ensemble is selected, how chosen parameters reach the next test, and how portfolio state crosses boundaries. Broadcasting thousands of candidates makes complete experiment accounting and an outer evaluation more important, not less.
Walk-forward checklist
Before accepting a result, verify that:
- Every test interval begins after its training information and any required purge or gap.
- All learned preprocessing, universe selection, fitting, calibration, and parameter choice occur inside each training interval.
- Window lengths, step size, objective, candidates, costs, and tie policy were fixed before test inspection.
- Test timestamps contribute once to a deployment curve, or overlapping forecasts have an explicit aggregation rule.
- Indicators receive only causal warm-up history, while cash, positions, orders, and turnover remain continuous across boundaries.
- Labels, features, fundamentals, universes, and corporate actions are point-in-time correct.
- Fees, spread, slippage, impact, financing, borrow, latency, capacity, and operational retraining delay are represented.
- Fold-level selections and results are reported alongside the aggregate path and uncertainty.
- Every alternative walk-forward procedure inspected by the researcher is counted.
- A later untouched or prospective evaluation remains after the process itself is finalized.
Walk-forward optimization is useful when it simulates a clear and repeatable way of retraining a strategy. The most important result is not the best setting in one fold. It is the dated record of training data, frozen decisions, orders, costs, and test results that another researcher could repeat without seeing the future.