Choose PyBroker for rule-based strategies or supervised models whose predictions become bar-based orders, with chronological retraining inside the backtest. Choose FinRL when the question is specifically about reinforcement learning: an agent learns a sequential policy from states, actions, transitions, and rewards.
The tools answer different questions. PyBroker is a general bar-based backtester with integrated model fitting. FinRL provides financial reinforcement-learning environments for research and education. Its maintainers direct newer live-trading work to the separate FinRL-X project. Compare them using the same standards for data timing, costs, validation, reproducibility, and live use.
This page reflects PyBroker 2.0.1 and the original FinRL 0.3.8 repository release, checked on September 19, 2026. FinRL's PyPI package still trails the repository at 0.3.7, so installation source is part of the experiment definition.
PyBroker vs FinRL at a glance
| Criterion | PyBroker | Original FinRL | Decision impact |
|---|---|---|---|
| Primary question | Do declared rules or fitted predictions produce useful bar-based trades? | Which sequential action policy maximizes a declared reward in this environment? | Start from the question, not the ML label |
| Modeling path | Rule functions plus optional classifiers, regressors, and custom models | Gym-style environments plus Stable Baselines 3, ElegantRL, or RLlib agents | PyBroker does not train RL policies. FinRL is not needed for an already specified signal rule |
| Evaluation structure | Backtest or walk-forward windows with model retraining and a configurable label lookahead gap | User-defined train, validation, test, and trade stages around environment rollouts | Neither structure creates an untouched test after the researcher uses its result |
| Trading simulator | Stateful multi-symbol OHLCV bar engine | Environment-specific accounting and transition loop, such as StockTradingEnv |
Both simplify execution and require independent validation for liquidity-sensitive strategies |
| Costs | Fees plus fixed, volatility, volume, random, or custom slippage models | Environment parameters such as buy and sell cost percentages | Configurable cost inputs are not automatically calibrated market models |
| Optimization burden | Optuna parameter search, model choice, feature choice, and window design | Algorithm, network, reward, state, action, seed, and hyperparameter search | RL usually introduces more interacting degrees of freedom, but both require full trial accounting |
| Reproducibility | Data, version, feature pipeline, model, split, cache, seed, and strategy config | All of those plus environment code, reward, action transform, dependency versions, hardware, and multiple seeds | A single best run is inadequate for either workflow |
| Live deployment | Research backtester, with no documented live broker runtime | Original project is positioned for education and research. Basic legacy trading code is not a production control plane | Plan a separate paper and live system, or evaluate another product |
| License | Apache-2.0 with Commons Clause, which restricts selling a substantially derived product or service | MIT | PyBroker is source-available, not plain Apache-2.0. FinRL-X uses a different Apache-2.0 license |
| Current role | Active PyBroker 2 research framework | Preserved original FinRL line. FinRL-X is the separately maintained successor for newer live-trading work | Do not attribute FinRL-X features or license to FinRL |
The modeling question comes first
Supervised learning estimates a target from examples. A PyBroker model might predict next-bar return, event probability, volatility, or a cross-sectional score. The strategy then translates that output into entries, exits, ranks, and sizes. The prediction objective and the trading objective remain separate, so evaluate both forecast quality and portfolio results.
Reinforcement learning estimates a policy or value function through interaction with an environment. A FinRL agent observes state, emits an action, receives a reward, and moves to a new state. This can represent sequential decisions where inventory, cash, and prior actions affect future choices. It also means the researcher defines the world in which the agent learns.
RL is not automatically more adaptive or realistic. If an explicit forecast and sizing rule describe the hypothesis, adding an environment, reward, and policy network creates extra estimation and debugging problems without necessarily adding information. Conversely, PyBroker's prediction-to-order workflow does not directly answer a true policy-learning question.
Both frameworks can run without a meaningful ML edge. PyBroker supports ordinary rule strategies, and a FinRL environment can be stepped with fixed actions. Verify deterministic mechanics before fitting anything.
Data and feature timing
PyBroker consumes pandas data or a supported source, calculates indicators, trains registered models, and calls strategy executions on completed bars. Model inputs can include lags and multiple compressed intervals. A callback can see current portfolio state and schedule an order for a later bar. The researcher still controls feature timestamps, target construction, universe membership, adjustments, and point-in-time availability.
FinRL data processors and applications construct an environment frame containing prices and chosen features. The environment then exposes a state according to its indexing and transition logic. A feature is not causal merely because it appears in an observation. Indicators, turbulence measures, fundamentals, rankings, and normalization must use only data available before the action.
Apply the same audit to both:
- map every raw field to its publication or exchange timestamp,
- fit scalers, imputers, selectors, and models only inside each training interval,
- preserve point-in-time universe membership and security identity,
- keep adjustments, missing bars, calendars, and time zones consistent,
- delay execution until after every signal input is available, and
- make the prediction horizon or reward transition agree with that delay.
One chronological split does not repair a leaky feature pipeline.
Training, validation, and repeated search
PyBroker's Strategy.walkforward() divides time into training and test segments, trains registered models within each window, and withholds the configured lookahead bars at the boundary. Models that use a different bar interval interpret that gap in their own bar units. Fixed strategy rules and constants are not optimized merely because the method is called walk-forward.
FinRL commonly separates training, testing, and trading periods around agent rollouts. A credible study needs a validation period for choosing the algorithm, reward, features, constraints, network, and hyperparameters before a final evaluation. If the final period influences any redesign, it has become part of model development.
The same statistical rule applies to both products: count the complete search. PyBroker trials include feature definitions, estimators, hyperparameters, symbols, windows, slippage settings, and strategy rules. FinRL trials additionally include environment variants, state and action spaces, rewards, algorithms, network architectures, training steps, and random seeds.
RL performance is often noisy because optimization and action sampling are stochastic. Run multiple declared seeds and report their distribution, failed runs, and selection procedure. PyBroker models can also be stochastic, so seed ensembles and stability checks may be needed there too. Do not publish the best seed or best trial as though it were a prespecified strategy.
Walk-forward evaluation can test a repeated retraining process, but its windows are dependent and the process itself can be overfit. Reserve a later prospective period or paper deployment after selecting the full research procedure.
Environment and backtest mechanics
PyBroker is the more complete conventional backtester of the pair. Its stateful bar engine supports multiple symbols and intervals, ranking, rotation, long and short positions, limits, stops, holding periods, fees, margin inputs, and scheduled fills. Strategy callbacks operate per symbol, while portfolio controls coordinate positions. Numba accelerates internal numerical paths.
That breadth does not recover intrabar history. OHLCV bars cannot identify the order of the high and low, queue priority, hidden liquidity, or the market's reaction to an order. Stop precedence and fill timing follow engine rules. A touched limit or stop is only as credible as the bar assumptions and available liquidity model.
FinRL's simulation semantics belong to each environment. In StockTradingEnv, a continuous normalized action is scaled by hmax, converted to an integer share amount, constrained by cash and holdings, and executed against a row price with configured percentage costs. The default reward is based on account-value change times a scaling factor. Those choices define the learned task.
The reward may be the most consequential model component. Account-value change, log return, risk-adjusted reward, drawdown penalty, turnover penalty, or benchmark-relative reward can produce different policies. Reward scaling can affect optimization even when it does not change economic units. State omissions can also violate the Markov assumption that the algorithm relies on.
For either framework, hand-calculate a tiny deterministic fixture. Inspect every order or action, fill, fee, cash change, holding, portfolio value, and terminal condition. Then test gaps, insufficient cash, maximum position, liquidation, missing data, and the final open position.
Transaction costs and capacity
PyBroker 2 provides per-order, per-share, percentage, and custom fees plus built-in slippage models. Fixed basis points, ATR-scaled movement, volume participation with nonlinear impact, randomness, and custom logic cover useful stress cases. They remain abstractions over bar data.
Classic FinRL environments commonly accept buy and sell cost percentages. These can penalize turnover and prevent an agent from learning against a completely frictionless market. They do not by themselves represent spread, nonlinear impact, partial fills, latency, borrow availability, financing, venue rules, or rejected orders.
Use identical calibrated assumptions when comparing results. Measure gross and net performance, turnover, participation, exposure, rejected actions, and break-even costs. If a policy trades more aggressively than a PyBroker strategy, applying the same constant percentage is equal syntax but not equal economic realism. Capacity and impact should respond to order size, liquidity, and horizon.
Neither framework should be the final execution validator for a strategy whose edge depends on tick order, book depth, or latency.
Metrics and baselines
PyBroker reports portfolio, order, trade, and risk metrics and can calculate bootstrap intervals. Its ordinary return bootstrap has assumptions about dependence and stationarity. Treat it as a sensitivity analysis rather than a certificate of generalization.
FinRL examples often plot an agent's account value against a market index or mean-variance portfolio. A benchmark is valid only when it uses the same dates, universe, starting information, costs, rebalancing opportunities, and constraints. Comparing an agent after extensive tuning with an untuned baseline is not a neutral algorithm comparison.
Every experiment should include simple alternatives:
- cash and buy-and-hold where applicable,
- an equal-weight or risk-scaled portfolio,
- a simple declared signal policy,
- the same policy without ML, and
- gross and cost-adjusted variants.
For RL, also include fixed, random, or heuristic policies that exercise the same action space. For supervised learning, compare prediction-based trades with a rule using the same features. Report distributions and drawdowns, not only endpoint return.
Repeating the result and compute needs
PyBroker experiments need the lib-pybroker version, Python environment, data snapshot, feature code, model versions, split definition, strategy configuration, seeds, and cache policy. Cached indicators or models can silently preserve stale results unless keys include every relevant data and code input.
FinRL needs all of that plus the repository tag or commit, agent-library versions, environment source, observation and action definitions, reward, training steps, network configuration, device, deterministic settings, and every seed. PyPI 0.3.7 and repository 0.3.8 are not interchangeable. Pinning only finrl is insufficient when the repository example and installed wheel differ.
Compute cost should be measured per complete research decision, not per backtest episode. A single RL training run may be expensive, while a broad PyBroker Optuna search can consume comparable resources through trial count. Parallelism increases experiment capacity and therefore increases the importance of a registry that records failures and discarded variants.
Live-use limits
PyBroker is a research backtester. It does not document a live broker engine with external-order reconciliation, durable state recovery, monitoring, or operational risk controls. A separate deployment must reproduce its feature availability, prediction timing, sizing, fees, and state transitions.
The original FinRL repository is now described by its maintainers as an education, benchmark, and research-prototyping framework. Its legacy Alpaca-related components do not turn the training environment into a production trading service. The maintainers point users seeking current deployment work to FinRL-X.
FinRL-X is a separate product with a different architecture, strategy interface, live scope, and Apache-2.0 license. Evaluate it separately. Do not use its features to score the original MIT-licensed FinRL project, and do not assume a classic Gym agent migrates unchanged.
For either workflow, a deployment plan needs a paper stage, live data checks, broker reconciliation, durable model and portfolio state, idempotent order handling, hard exposure and loss limits, alerts, restart testing, and a manual shutdown path.
Licensing
PyBroker's repository uses Apache 2.0 with the Commons Clause. The Commons Clause restricts selling a product or service whose value derives substantially from the software. Calling the project simply Apache-2.0 omits a material commercial restriction. Obtain appropriate advice for hosted or commercial use.
Original FinRL uses the permissive MIT License. Models, market data, broker APIs, and dependencies can carry separate terms. FinRL-X's Apache-2.0 license does not replace FinRL's MIT license or change what code is in each repository.
Which should you choose?
Choose PyBroker when most of these are true:
- the hypothesis can be expressed as rules or supervised predictions,
- bars are an adequate research resolution,
- integrated model retraining and prediction inside chronological windows are useful,
- multi-symbol ranking, rotation, or multiple intervals fit the strategy,
- there is a separate plan for final execution validation and deployment, and
- the Commons Clause is acceptable for the intended use.
Choose original FinRL when most of these are true:
- the research contribution is a sequential policy or financial RL environment,
- the state, action, transition, and reward can be stated and audited precisely,
- multiple agents, seeds, and simple policies will be compared fairly,
- educational and research-prototyping scope is acceptable,
- execution will be validated elsewhere, and
- you will pin the repository version and dependencies.
Choose neither merely because a strategy uses machine learning. A separate training pipeline and a conventional backtester can be easier to validate. For newer live-trading work from AI4Finance, compare FinRL-X as a separate project rather than treating it as a direct FinRL upgrade.
A small test before you choose
The models need not be identical, but the evidence standard should be:
- Freeze one point-in-time dataset, universe, calendar, and cost policy.
- Declare the PyBroker target and order mapping, then declare the FinRL state, action transform, transition, and reward with the same economic objective.
- Use chronological training, validation, and final periods without inspecting the final period during design.
- Compare both with the same simple baselines and record every model, environment, hyperparameter, reward, and seed attempted.
- Reconcile a deterministic accounting fixture before training and stress costs, gaps, unavailable actions, and open terminal positions.
- Report the distribution across seeds or retraining windows, not only the selected run.
- Validate the chosen policy in an independent simulator and paper process before any live decision.
This test may show that only one framework matches the hypothesis. That is a useful result. PyBroker and FinRL should not be forced into a winner table when they answer different research questions.