Identifying Policy Gradient Subspaces
Jan Schneider, Pierre Schumacher, Simon Guist, Le Chen, Daniel F. B. Haeufle, Bernhard Schölkopf, Dieter Büchler
摘要
Policy gradient methods hold great potential for solving complex continuous control tasks. Still, their training efficiency can be improved by exploiting structure within the optimization problem. Recent work indicates that supervised learning can be accelerated by leveraging the fact that gradients lie in a low-dimensional and slowly-changing subspace. In this paper, we conduct a thorough evaluation of this phenomenon for two popular deep policy gradient methods on various simulated benchmark tasks. Our results demonstrate the existence of such gradient subspaces despite the continuously changing data distribution inherent to reinforcement learning. These findings reveal promising directions for future work on more efficient reinforcement learning, e.g., through improving parameter-space exploration or enabling second-order optimization.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- SubTrack++ : Gradient Subspace Tracking for Scalable LLM TrainingSahar Rajabi, Nayeema Nonta, Sirisha RambhatlaNeurIPS 2025 · 被引用 19 次
- The Ladder in Chaos: Improving Policy Learning by Harnessing the Parameter Evolving Path in A Low-dimensional SpaceHongyao Tang, Min Zhang, Chen Chen, Jianye HaoNeurIPS 2024 · 被引用 2 次
- Mitigating Plasticity Loss in Continual Reinforcement Learning by Reducing ChurnHongyao Tang, Johan S. Obando-Ceron, Pablo Samuel Castro, Aaron C. Courville 等ICML 2025
- Does SGD really happen in tiny subspaces?Minhak Song, Kwangjun Ahn, Chulhee YunICLR 2025
它引用的顶会 Paper12
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Bypassing the Ambient Dimension: Private SGD with Gradient Subspace IdentificationYingxue Zhou, Steven Wu, Arindam BanerjeeICLR 2021 · 被引用 118 次
- A Closer Look at Deep Policy GradientsAndrew Ilyas, Logan Engstrom, Shibani Santurkar, Dimitris Tsipras 等ICLR 2020 · 被引用 107 次
- Subspace Adversarial TrainingTao Li, Yingwen Wu, Sizhe Chen, Kun Fang 等CVPR 2022 · 被引用 59 次
- Is High Variance Unavoidable in RL? A Case Study in Continuous ControlJohan Bjorck, Carla P. Gomes, Kilian Q. WeinbergerICLR 2022 · 被引用 36 次
相关 Paper
- From Parameters to Behaviors: Unsupervised Compression of the Policy SpaceDavide Tenedini, Riccardo Zamboni, Mirco Mutti, Marcello RestelliICLR 2026 · 被引用 4 次
- Mollification Effects of Policy Gradient MethodsTao Wang, Sylvia L. Herbert, Sicun GaoICML 2024 · 被引用 2 次
- A Parametric Class of Approximate Gradient Updates for Policy OptimizationRamki Gummadi, Saurabh Kumar, Junfeng Wen, Dale SchuurmansICML 2022
- Deep Bayesian Quadrature Policy OptimizationRavi Tej Akella, Kamyar Azizzadenesheli, Mohammad Ghavamzadeh, Animashree Anandkumar 等AAAI 2021 · 被引用 5 次
- Deterministic Value-Policy GradientsQingpeng Cai, Ling Pan, Pingzhong TangAAAI 2020 · 被引用 1 次
