Batch size-invariance for policy optimization
Jacob Hilton, Karl Cobbe, John Schulman
Abstract
We say an algorithm is batch size-invariant if changes to the batch size can largely be compensated for by changes to other hyperparameters. Stochastic gradient descent is well-known to have this property at small batch sizes, via the learning rate. However, some policy optimization algorithms (such as PPO) do not have this property, because of how they control the size of policy updates. In this work we show how to make these algorithms batch size-invariant. Our key insight is to decouple the proximal policy (used for controlling policy updates) from the behavior policy (used for off-policy corrections). Our experiments help explain why these algorithms work, and additionally show how they can make more efficient use of stale data.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers15
- AREAL: A Large-Scale Asynchronous Reinforcement Learning System for Language ReasoningWei Fu, Jiaxuan Gao, Xujie Shen, Chen Zhu et al.NeurIPS 2025 · 273 citations
- Small batch deep reinforcement learningJohan S. Obando-Ceron, Marc G. Bellemare, Pablo Samuel CastroNeurIPS 2023 · 38 citations
- VER: Scaling On-Policy RL Leads to the Emergence of Navigation in Embodied RearrangementErik Wijmans, Irfan Essa, Dhruv BatraNeurIPS 2022 · 24 citations
- Rethinking the Trust Region in LLM Reinforcement LearningPenghui Qi, Xiangxin Zhou, Zichen Liu, Tianyu Pang et al.ICML 2026 · 22 citations
- Normalization Layer Per-Example Gradients are Sufficient to Predict Gradient Noise Scale in TransformersGavia Gray, Aman Tiwari, Shane Bergsma, Joel HestnessNeurIPS 2024 · 7 citations
Builds on4
- Leveraging Procedural Generation to Benchmark Reinforcement LearningKarl Cobbe, Christopher Hesse, Jacob Hilton, John SchulmanICML 2020 · 685 citations
- Phasic Policy GradientKarl Cobbe, Jacob Hilton, Oleg Klimov, John SchulmanICML 2021 · 191 citations
- Mirror Descent Policy OptimizationManan Tomar, Lior Shani, Yonathan Efroni, Mohammad GhavamzadehICLR 2022 · 111 citations
- Evaluating the Performance of Reinforcement Learning AlgorithmsScott M. Jordan, Yash Chandak, Daniel Cohen, Mengxue Zhang et al.ICML 2020 · 59 citations
Related papers
- Off-Policy Proximal Policy OptimizationWenjia Meng, Qian Zheng, Gang Pan, Yilong YinAAAI 2023 · 27 citations
- Batch Reinforcement Learning with Hyperparameter GradientsByung-Jun Lee, Jongmin Lee, Peter Vrancx, Dongho Kim et al.ICML 2020 · 18 citations
- Staggered Environment Resets Improve Massively Parallel On-Policy Reinforcement LearningSid Bharthulwar, Stone Tao, Hao SuNeurIPS 2025 · 6 citations
- Scalable Reinforcement Learning via Adaptive Batch ScalingJongchan ParkICML 2026 · 1 citation
- Behavior Proximal Policy OptimizationZifeng Zhuang, Kun Lei, Jinxin Liu, Donglin Wang et al.ICLR 2023 · 8 citations
