Scalable Reinforcement Learning via Adaptive Batch Scaling
Jongchan Park
Abstract
Conventional wisdom holds that large-batch training is fundamentally incompatible with Reinforcement Learning (RL) -beyond a modest threshold, increasing batch sizes typically yields diminishing returns or performance degradation due to the inherent non-stationarity of the data distribution. We challenge this view by observing that non-stationarity is not a fixed property of RL, but evolves throughout training: early stages exhibit rapid behavioral shifts that demand small batches for plasticity, whereas late stages approach a quasi-stationary regime where large batches enable precise convergence. Motivated by this observation, we propose Adaptive Batch Scaling (ABS), that dynamically adjusts the effective batch size according to the stability of the learning policy. Central to ABS is Behavioral Divergence, a novel metric that quantifies policy nonstationarity by measuring action-level shifts between consecutive updates, which we use to scale batch size inversely to policy volatility. Integrated with the Parallelised Q-Network (PQN) algorithm and evaluated on the ALE benchmark, ABS seamlessly reconciles early-stage plasticity with latestage stable convergence. Strikingly, contrary to conventional wisdom, our results reveal that the combination of larger networks and larger batch sizes achieves the best performance -a scaling behavior previously thought to be unattainable in RL, now unlocked through adaptive batch control. Our code is available at https:// github.com/daisophila/ABS .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7249dcad-5dba-4945-aa34-496587d86c63Builds on16
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- DAPO: An Open-Source LLM Reinforcement Learning System at ScaleQiying Yu, Zheng Zhang, Ruofei Zhu, Yufeng Yuan et al.NeurIPS 2025 · 2,828 citations
- On the Variance of the Adaptive Learning Rate and BeyondLiyuan Liu, Haoming Jiang, Pengcheng He, Weizhu Chen et al.ICLR 2020 · 2,210 citations
Related papers
- Batch size-invariance for policy optimizationJacob Hilton, Karl Cobbe, John SchulmanNeurIPS 2022 · 41 citations
- The Impact of On-Policy Parallelized Data Collection on Deep Reinforcement Learning NetworksWalter Mayor, Johan S. Obando-Ceron, Aaron C. Courville, Pablo Samuel CastroICML 2025
- Adaptive Q-Network: On-the-fly Target Selection for Deep Reinforcement LearningThéo Vincent, Fabian Wahren, Jan Peters, Boris Belousov et al.ICLR 2025
- Balancing Plasticity and Stability with Fast and Slow Successor FeaturesRaymond Chua, Doina Precup, Blake RichardsICML 2026
- Towards Deeper Deep Reinforcement Learning with Spectral NormalizationJohan Bjorck, Carla P. Gomes, Kilian Q. WeinbergerNeurIPS 2021 · 26 citations
