Reinforcement-learning environments
Median runtime for 1,000 deterministic FinRL StockTradingEnv steps as the asset count grows from 1 to 100.
Tools measured in this benchmark
Open a tool page for its code examples, strengths, and limits. Exact versions are listed in the technical details.
Runtime for 1,000 steps by asset count
Median wall-clock time for a fixed number of deterministic environment steps. Lower is faster. Runtime uses a logarithmic scale.
Environment cost grows with the state and action space
FinRL rises from 65.51 ms at one asset to 2,807.46 ms at 100 assets. The state expands from 4 to 301 values, while the deterministic workload grows from 1,000 to 72,260 executed trades.
See exactly how the result was produced
Exact values
| Assets | State width | FinRL |
|---|---|---|
| 1 | 4 | 65.51 ms |
| 10 | 31 | 369.33 ms |
| 100 | 301 | 2,807.46 ms |
Environment
- Apple M3, 8 logical cores, 24 GB RAM
- macOS 26.5.2, arm64
- CPython 3.11.8
- FinRL 0.3.8, commit cb21549
- NumPy 2.4.6
Procedure
- 1,000 deterministic steps per case
- 1, 10, and 100 assets
- State widths 4, 31, and 301
- 2 warmups and 5 measured repetitions
- Full state, reward checksum, and trade count checked
Scope
This measures environment stepping, not learning. The timing includes environment construction, reset, 1,000 deterministic steps, trading updates, rewards, and state collection. Policy inference, replay buffers, training, and accelerators are excluded.
Peak RSS was not measured consistently. These numbers do not include model inference or training and should not be used as an end-to-end RL throughput claim.
Reproduce and inspect
The downloadable files preserve the measured values, available result detail, environment record, and checksums.
See the editorial process for the publication standard.
Download the data
Download the chart values, detailed results, release details, or file checksums separately.