Batch size-invariance for policy optimization
Jacob Hilton, Karl Cobbe, John Schulman
摘要
We say an algorithm is batch size-invariant if changes to the batch size can largely be compensated for by changes to other hyperparameters. Stochastic gradient descent is well-known to have this property at small batch sizes, via the learning rate. However, some policy optimization algorithms (such as PPO) do not have this property, because of how they control the size of policy updates. In this work we show how to make these algorithms batch size-invariant. Our key insight is to decouple the proximal policy (used for controlling policy updates) from the behavior policy (used for off-policy corrections). Our experiments help explain why these algorithms work, and additionally show how they can make more efficient use of stale data.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper15
- AREAL: A Large-Scale Asynchronous Reinforcement Learning System for Language ReasoningWei Fu, Jiaxuan Gao, Xujie Shen, Chen Zhu 等NeurIPS 2025 · 被引用 273 次
- Small batch deep reinforcement learningJohan S. Obando-Ceron, Marc G. Bellemare, Pablo Samuel CastroNeurIPS 2023 · 被引用 38 次
- VER: Scaling On-Policy RL Leads to the Emergence of Navigation in Embodied RearrangementErik Wijmans, Irfan Essa, Dhruv BatraNeurIPS 2022 · 被引用 24 次
- Rethinking the Trust Region in LLM Reinforcement LearningPenghui Qi, Xiangxin Zhou, Zichen Liu, Tianyu Pang 等ICML 2026 · 被引用 22 次
- Normalization Layer Per-Example Gradients are Sufficient to Predict Gradient Noise Scale in TransformersGavia Gray, Aman Tiwari, Shane Bergsma, Joel HestnessNeurIPS 2024 · 被引用 7 次
它引用的顶会 Paper4
- Leveraging Procedural Generation to Benchmark Reinforcement LearningKarl Cobbe, Christopher Hesse, Jacob Hilton, John SchulmanICML 2020 · 被引用 685 次
- Phasic Policy GradientKarl Cobbe, Jacob Hilton, Oleg Klimov, John SchulmanICML 2021 · 被引用 191 次
- Mirror Descent Policy OptimizationManan Tomar, Lior Shani, Yonathan Efroni, Mohammad GhavamzadehICLR 2022 · 被引用 111 次
- Evaluating the Performance of Reinforcement Learning AlgorithmsScott M. Jordan, Yash Chandak, Daniel Cohen, Mengxue Zhang 等ICML 2020 · 被引用 59 次
相关 Paper
- Off-Policy Proximal Policy OptimizationWenjia Meng, Qian Zheng, Gang Pan, Yilong YinAAAI 2023 · 被引用 27 次
- Batch Reinforcement Learning with Hyperparameter GradientsByung-Jun Lee, Jongmin Lee, Peter Vrancx, Dongho Kim 等ICML 2020 · 被引用 18 次
- Staggered Environment Resets Improve Massively Parallel On-Policy Reinforcement LearningSid Bharthulwar, Stone Tao, Hao SuNeurIPS 2025 · 被引用 6 次
- Scalable Reinforcement Learning via Adaptive Batch ScalingJongchan ParkICML 2026 · 被引用 1 次
- Behavior Proximal Policy OptimizationZifeng Zhuang, Kun Lei, Jinxin Liu, Donglin Wang 等ICLR 2023 · 被引用 8 次
