CompleteP for RL: Maintaining Feature Learning When Scaling Deep Reinforcement Learning
Adam Lee, M Ganesh Kumar, Blake Bordelon, Cengiz Pehlevan
摘要
The maximal update parameterization () provides an approach to scale parameters while preserving feature learning in deep neural networks. In fixed data distribution settings it has been found to yield more stable learning dynamics, more consistent learned features, and enable optimal hyperparameters such as learning rate to transfer from small models to larger ones. However, it is unclear if these benefits readily transfer to reinforcement learning problems, where learning dynamics are coupled to the non-stationary data distribution induced by an agent's own actions. We empirically study how two regimes, the ''rich'' CompleteP and ''lazy'' Neural Tangent Kernel (NTK) parameterizations affect hyperparameter transfer, feature and policy consistency, and learning behavior as we scale reinforcement learning agents. Ultimately, we show that agents trained using CompleteP consequentially improves compute and reward efficiency compared to the NTK parameterization across multiple control tasks and variants.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper20
- Tuning Large Neural Networks via Zero-Shot Hyperparameter TransferGe Yang, Edward J. Hu, Igor Babuschkin, Szymon Sidor 等NeurIPS 2021 · 被引用 208 次
- The Dormant Neuron Phenomenon in Deep Reinforcement LearningGhada Sokar, Rishabh Agarwal, Pablo Samuel Castro, Utku EvciICML 2023 · 被引用 153 次
- Self-Consistent Dynamical Field Theory of Kernel Evolution in Wide Neural NetworksBlake Bordelon, Cengiz PehlevanNeurIPS 2022 · 被引用 140 次
- Don't be lazy: CompleteP enables compute-efficient deep transformersNolan Dey, Bin Claire Zhang, Lorenzo Noci, Mufan Bill Li 等NeurIPS 2025 · 被引用 77 次
- Tensor Programs VI: Feature Learning in Infinite Depth Neural NetworksGreg Yang, Dingli Yu, Chen Zhu, Soufiane HayouICLR 2024 · 被引用 77 次
相关 Paper
- Weight Decay may matter more than µP for Learning Rate Transfer in PracticeAtli Kosson, Jeremy Welborn, Yang Liu, Martin Jaggi 等ICLR 2026 · 被引用 11 次
- Super Consistency of Neural Network Landscapes and Learning Rate TransferLorenzo Noci, Alexandru Meterez, Thomas Hofmann, Antonio OrvietoNeurIPS 2024 · 被引用 25 次
- On the Provable Separation of Scales in Maximal Update ParameterizationLetong Hong, Zhangyang WangICML 2025
- Understanding the Mechanisms of Fast Hyperparameter TransferNikhil Ghosh, Denny Wu, Alberto BiettiICLR 2026 · 被引用 8 次
- Deep Linear Network Training Dynamics from Random Initialization: Data, Width, Depth, and Hyperparameter TransferBlake Bordelon, Cengiz PehlevanICML 2025
