Understanding and Improving Hyperbolic Deep Reinforcement Learning
Timo Klein, Thomas Lang, Andrii Shkabrii, Alexander Sturm, Kevin Sidak, Lukas Miklautz, Claudia Plant, Yllka Velaj, Sebastian Tschiatschek
摘要
The exponential volume growth of hyperbolic geometry can embed the hierarchical relationships between states in reinforcement learning (RL) with far less distortion than Euclidean space. However, hyperbolic deep RL faces severe optimization challenges, and formal analysis of why optimization fails is lacking. We identify key factors that determine the success and failure of training hyperbolic deep RL agents. By analyzing the gradients of core operations in the Poincaré Ball and Hyperboloid models of hyperbolic geometry, we show that large-norm embeddings destabilize gradient-based training, leading to trust-region violations in proximal policy optimization (PPO). Based on these insights, we introduce HYPER++, a new hyperbolic deep RL agent that consists of three components: (i) feature regularization guaranteeing bounded norms while avoiding the curse of dimensionality from clipping; (ii) a categorical value loss for stable critic training; and (iii) a more optimization-friendly formulation of hyperbolic network layers. On ProcGen, we show that HYPER++ guarantees stable learning, outperforms prior hyperbolic agents, and reduces wall-clock time by approximately 30%. On Atari-5 with Double DQN, HYPER++ strongly outperforms Euclidean and hyperbolic baselines. We release our code at https://github.com/Probabilistic-and-Interactive-ML/hyper-rl .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper20
- Deep Reinforcement Learning at the Edge of the Statistical PrecipiceRishabh Agarwal, Max Schwarzer, Pablo Samuel Castro, Aaron C. Courville 等NeurIPS 2021 · 被引用 1,067 次
- Hyperbolic Neural Networks++Ryohei Shimizu, Yusuke Mukuta, Tatsuya HaradaICLR 2021 · 被引用 791 次
- Leveraging Procedural Generation to Benchmark Reinforcement LearningKarl Cobbe, Christopher Hesse, Jacob Hilton, John SchulmanICML 2020 · 被引用 685 次
- Phasic Policy GradientKarl Cobbe, Jacob Hilton, Oleg Klimov, John SchulmanICML 2021 · 被引用 191 次
- Understanding Plasticity in Neural NetworksClare Lyle, Zeyu Zheng, Evgenii Nikishin, Bernardo Ávila Pires 等ICML 2023 · 被引用 162 次
相关 Paper
- Hyperbolic Deep Reinforcement LearningEdoardo Cetin, Benjamin Paul Chamberlain, Michael M. Bronstein, Jonathan J. HuntICLR 2023 · 被引用 4 次
- Exploiting Geometric Structures for Modeling Multi-Agent Behaviors: A New ThinkingBohao Qu, Xiaofeng Cao, Bing Li, Menglin Zhang 等AAAI 2026
- Improving Robustness of Hyperbolic Neural Networks by Lipschitz AnalysisYuekang Li, Yidan Mao, Yifei Yang, Dongmian ZouKDD 2024 · 被引用 1 次
- No Representation, No Trust: Connecting Representation, Collapse, and Trust Issues in PPOSkander Moalla, Andrea Miele, Daniil Pyatko, Razvan Pascanu 等NeurIPS 2024 · 被引用 35 次
- Robust Deep Reinforcement Learning against Adversarial Perturbations on State ObservationsHuan Zhang, Hongge Chen, Chaowei Xiao, Bo Li 等NeurIPS 2020 · 被引用 437 次
