Value-Based Deep RL Scales Predictably
Oleh Rybkin, Michal Nauman, Preston Fu, Charlie Victor Snell, Pieter Abbeel, Sergey Levine, Aviral Kumar
摘要
Scaling data and compute is critical to the success of modern ML. However, scaling demands predictability: we want methods to not only perform well with more compute or data, but also have their performance be predictable from small-scale runs, without running the large-scale experiment. In this paper, we show that valuebased off-policy RL methods are predictable despite community lore regarding their pathological behavior. First, we show that data and compute requirements to attain a given performance level lie on a Pareto frontier, controlled by the updates-to-data (UTD) ratio. By estimating this frontier, we can predict this data requirement when given more compute, and this compute requirement when given more data. Second, we determine the optimal allocation of a total resource budget across data and compute for a given performance and use it to determine hyperparameters that maximize performance for a given budget. Third, this scaling is enabled by first estimating predictable relationships between hyperparameters, which is used to manage effects of overfitting and plasticity loss unique to RL. We validate our approach using three algorithms: SAC, BRO, and PQL on DeepMind Control, OpenAI gym, and IsaacGym, when extrapolating to higher levels of data, compute, budget, or performance. (I) Compute-Data Pareto frontier (II) Budget extrapolation (III) Fits for multiple J Isaac Gym DMC OpenAI Gym Figure 1: Scaling properties when increasing compute C, data D, budget F, or performance J. Left: Compute versus data requirements Pareto frontier controlled by the UTD ratio σ. We observe that we can trade off data for compute and vice versa, and this relationship is predictable. Middle: Extrapolation from low to high performance. We observe that the optimal resource allocation controlled by σ evolves predictably with increasing budget, and can be used to extrapolate from low to high performance. Right: Pareto frontiers for several performance levels J.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Horizon Reduction Makes RL ScalableSeohong Park, Kevin Frans, Deepinder Mann, Benjamin Eysenbach 等NeurIPS 2025 · 被引用 60 次
- Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task LearnersMichal Nauman, Marek Cygan, Carmelo Sferrazza, Aviral Kumar 等NeurIPS 2025 · 被引用 26 次
- floq: Training Critics via Flow-Matching for Scaling Compute in Value-Based RLBhavya Agrawalla, Michal Nauman, Khush Agrawal, Aviral KumarICLR 2026 · 被引用 22 次
- XQC: Well-conditioned Optimization Accelerates Deep Reinforcement LearningDaniel Palenicek, Florian Vogt, Joe Watson, Ingmar Posner 等ICLR 2026 · 被引用 20 次
- Compute-Optimal Scaling for Value-Based Deep RLPreston Fu, Oleh Rybkin, Zhiyuan Zhou, Michal Nauman 等NeurIPS 2025 · 被引用 7 次
它引用的顶会 Paper21
- Scaling Laws for Reward Model OveroptimizationLeo Gao, John Schulman, Jacob HiltonICML 2023 · 被引用 963 次
- Scaling Data-Constrained Language ModelsNiklas Muennighoff, Alexander M. Rush, Boaz Barak, Teven Le Scao 等NeurIPS 2023 · 被引用 475 次
- The Primacy Bias in Deep Reinforcement LearningEvgenii Nikishin, Max Schwarzer, Pierluca D'Oro, Pierre-Luc Bacon 等ICML 2022 · 被引用 269 次
- Understanding Plasticity in Neural NetworksClare Lyle, Zeyu Zheng, Evgenii Nikishin, Bernardo Ávila Pires 等ICML 2023 · 被引用 162 次
- Implicit Under-Parameterization Inhibits Data-Efficient Deep Reinforcement LearningAviral Kumar, Rishabh Agarwal, Dibya Ghosh, Sergey LevineICLR 2021 · 被引用 155 次
相关 Paper
- IsoCompute Playbook: Optimally Scaling Sampling Compute for LLM RLZhoujun Cheng, Yutao Xie, Yuxiao Qu, Amrith Setlur 等ICML 2026
- Dynamic Update-to-Data Ratio: Minimizing World Model OverfittingNicolai Dorka, Tim Welschehold, Wolfram BurgardICLR 2023 · 被引用 1 次
- Budgeting Counterfactual for Offline RLYao Liu, Pratik Chaudhari, Rasool FakoorNeurIPS 2023 · 被引用 6 次
- Bigger, Regularized, Optimistic: scaling for compute and sample efficient continuous controlMichal Nauman, Mateusz Ostaszewski, Krzysztof Jankowski, Piotr Milos 等NeurIPS 2024 · 被引用 119 次
- The Art of Scaling Reinforcement Learning Compute for LLMsDevvrit Khatri, Lovish Madaan, Rishabh Tiwari, Rachit Bansal 等ICLR 2026 · 被引用 95 次
