- Platform
- Financial reinforcement-learning framework
- License
- MIT
- Pricing
- Free and open source
- Live trading
- Not provided by the classic FinRL workflow
- Best for
- Learning and prototyping financial RL environments, rewards, and policies rather than conventional signal backtesting or live execution
FinRL is mainly an educational and research framework for financial reinforcement learning (RL). It provides market environments, agent adapters, and a train-test-trade workflow for studying policies that choose actions from market states and rewards. It is not a detailed execution simulator, and its maintainers direct newer live-trading work to the separate FinRL-X project.
How the classic FinRL workflow works
The official architecture has three layers: market environments, deep reinforcement learning (DRL) agents, and financial applications. The repository connects those layers through training, testing, and trading stages. Its agent adapters include Stable Baselines 3, ElegantRL, and RLlib, with documented algorithms such as A2C, DDPG, PPO, SAC, TD3, and DQN.
For the stock environment, the state includes cash, prices, holdings, and selected features. A continuous action is scaled by hmax and converted into share trades. The default reward implementation is the change in total account value multiplied by reward_scaling. Buy and sell cost arrays, a turbulence threshold, and position constraints shape the simulated result. These are model choices, not market facts.
| Design choice | Why it changes the result |
|---|---|
| State and feature timing | Features computed with future or revised data create look-ahead bias |
| Action scaling and share conversion | hmax, integer conversion, cash rules, and shorting constraints define what the agent can trade |
| Reward | Account-value change, risk penalties, or turnover penalties can teach materially different policies |
| Costs and fills | Configured percentages do not reproduce spreads, market impact, partial fills, latency, or borrow constraints |
| Training randomness | Seed, network initialization, sampling, and exploration can produce different policies from the same data |
| Validation split | Reusing a test period for model or reward selection turns that period into training information |
A deterministic environment-mechanics example
Before training an agent, step the environment with fixed actions and check the accounting by hand. The example below targets the official v0.3.8 tag and needs the repository installation described above. It uses one synthetic asset and one feature, so the four-element state is cash, current price, shares held, and momentum.
import numpy as np
import pandas as pd
from finrl.meta.env_stock_trading.env_stocktrading import StockTradingEnv
prices = pd.DataFrame(
{
"date": ["2026-01-02", "2026-01-05", "2026-01-06", "2026-01-07"],
"tic": ["TEST"] * 4,
"close": [100.0, 102.0, 101.0, 104.0],
"momentum": [0.00, 0.02, -0.01, 0.03],
},
index=[0, 1, 2, 3],
)
env = StockTradingEnv(
df=prices,
stock_dim=1,
hmax=10,
initial_amount=1_000,
num_stock_shares=[0],
buy_cost_pct=[0.001],
sell_cost_pct=[0.001],
reward_scaling=1e-4,
state_space=4,
action_space=1,
tech_indicator_list=["momentum"],
print_verbosity=100,
)
observation, _ = env.reset()
scaled_rewards = []
for action in ([0.5], [0.0], [-0.5], [0.0]):
observation, reward, terminated, truncated, _ = env.step(
np.asarray(action, dtype=np.float32)
)
assert not truncated
if terminated:
break
scaled_rewards.append(reward)
print("executed shares:", [int(action[0]) for action in env.actions_memory])
print("account values:", [float(round(value, 3)) for value in env.asset_memory])
print("scaled rewards:", [float(round(value, 7)) for value in scaled_rewards])
print("total cost:", round(env.cost, 3))
print("trades:", env.trades)
print("final cash:", round(observation[0], 3))
The exact displayed script produced:
executed shares: [5, 0, -5]
account values: [1000.0, 1009.5, 1004.5, 1003.995]
scaled rewards: [0.00095, -0.0005, -5.05e-05]
total cost: 1.005
trades: 2
final cash: 1003.995
The normalized actions are multiplied by hmax=10 and converted to integers, so 0.5 buys five shares and -0.5 sells five. Buying at 100 costs 500 plus a 0.5 fee. The environment then advances to the next row and values the position at 102, producing the first 9.5 account-value change and its scaled reward of 0.00095. Selling five shares later adds another 0.505 fee.
This deliberately hand-authored action sequence is not an agent, training run, benchmark, or trading signal. The favorable ending comes from the constructed prices. The run verifies state shape, action scaling, integer share conversion, percentage costs, reward scaling, and Gymnasium termination. It also shows why the environment needs independent execution validation: a buy is priced from the current row with no spread, slippage, partial fill, latency, or volume constraint.
What a credible FinRL experiment should report
A reproducible result needs more than a saved equity curve. Record the FinRL commit or package version, Python and dependency lockfile, data snapshot, universe membership through time, feature timestamps, environment parameters, reward definition, algorithm hyperparameters, random seeds, and hardware. Run multiple seeds and report the distribution, not only the best run.
Keep training, tuning, and final evaluation periods separate. If the policy or reward is changed after inspecting the final period, reserve a new untouched period. Walk-forward tests can better expose regime sensitivity than one fixed split, but they still require point-in-time data and a declared selection procedure. Compare the agent with simple baselines using identical data, costs, and trade constraints.
The 2026 official tutorial trains five Stable Baselines 3 agents and compares them with mean-variance optimization and the DJIA. That tutorial is a workflow demonstration. Its output should not be generalized to other periods, universes, or cost assumptions.
Trading and live-use limits
FinRL environments can charge fixed buy and sell percentages, restrict cash or positions, and use turbulence or volatility controls. Those features do not model an order book, queue priority, latency, venue rules, market impact, financing, or stock-borrow availability. Validate a promising policy in an independent simulator, then paper trade it with monitoring and hard risk limits before considering live capital.
The classic repository includes basic Alpaca-related code, but the maintainers direct newer trading work to FinRL-X. FinRL-X has a different design and an Apache-2.0 license. Its features and versions should not be attributed to this MIT-licensed FinRL package.
Maintenance and license
FinRL uses the MIT License. The original repository is not abandoned, as the March 2026 tutorial release shows, but its stated role is education, benchmarking, and research prototyping. The published package lag and wide dependency surface make a pinned, isolated environment prudent.
When to choose FinRL
Choose FinRL when the research question is genuinely about learning a sequential policy or designing a financial RL environment. Use a conventional backtester when the strategy already has explicit signals and portfolio rules. For live trading, evaluate FinRL-X separately rather than assuming that code will move across unchanged. In either case, treat the best policy from many agents, seeds, rewards, and hyperparameters as a multiple-testing problem.