XQC: Well-conditioned Optimization Accelerates Deep Reinforcement Learning
Daniel Palenicek, Florian Vogt, Joe Watson, Ingmar Posner, Jan Peters
摘要
Sample efficiency is a central property of effective deep reinforcement learning algorithms. Recent work has improved this through added complexity, such as larger models, exotic network architectures, and more complex algorithms, which are typically motivated purely by empirical performance. We take a more principled approach by focusing on the optimization landscape of the critic network. Using the eigenspectrum and condition number of the critic's Hessian, we systematically investigate the impact of common architectural design decisions on training dynamics. Our analysis reveals that a novel combination of batch normalization (BN), weight normalization (WN), and a distributional crossentropy (CE) loss produces condition numbers orders of magnitude smaller than baselines. This combination also naturally bounds gradient norms, a property critical for maintaining a stable effective learning rate under non-stationary targets and bootstrapping. Based on these insights, we introduce XQC: a well-motivated, sample-efficient deep actor-critic algorithm built upon soft actor-critic that embodies these optimization-aware principles. We achieve state-of-the-art sample efficiency across 55 proprioception and 15 vision-based continuous control tasks, all while using significantly fewer parameters than competing methods. Our code is available at danielpalenicek.github.io/projects/xqc. Figure 1 : Well-conditioned network architectures yield state-of-the-art RL performance. Our algorithm, XQC with a BN and WN-based architecture and a CE loss, achieves competitive performance against state-of-the-art baselines across 55 proprioceptive continuous control tasks from four different benchmarks with a single set of hyperparameters. Notably, with ∼ 4.5× fewer parameters and ∼ 5× less compute in terms of FLOP/S than SIMBA-V2, the closest competitor. XQC's efficiency carries over to RL from pixels on 15 vision-based DMC tasks, significantly improving on DRQ-V2.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Relative Entropy Pathwise Policy OptimizationClaas Voelcker, Axel Brunnbauer, Marcel Hussing, Michal Nauman 等ICLR 2026 · 被引用 6 次
- Stable Deep Reinforcement Learning via Isotropic Gaussian RepresentationsAli Saheb pasand, Johan Obando-Ceron, Aaron Courville, Pouya Bashivan 等ICML 2026 · 被引用 5 次
- Use the Online Network If You Can: Towards Fast and Stable Reinforcement LearningAhmed Hendawy, Henrik Metternich, Théo Vincent, Mahdi Kallel 等ICLR 2026 · 被引用 4 次
- What Does Flow-Matching Bring to TD-Learning?Bhavya Agrawalla, Michal Nauman, Aviral KumarICML 2026 · 被引用 1 次
- Trust Region Inverse Reinforcement Learning: Explicit Dual Ascent using Local Policy UpdatesAnish Abhijit Diwan, Davide Tateo, Christopher Mower, Haitham Bou Ammar 等ICML 2026
它引用的顶会 Paper27
- Dream to Control: Learning Behaviors by Latent ImaginationDanijar Hafner, Timothy P. Lillicrap, Jimmy Ba, Mohammad NorouziICLR 2020 · 被引用 1,852 次
- Deep Reinforcement Learning at the Edge of the Statistical PrecipiceRishabh Agarwal, Max Schwarzer, Pablo Samuel Castro, Aaron C. Courville 等NeurIPS 2021 · 被引用 1,067 次
- TD-MPC2: Scalable, Robust World Models for Continuous ControlNicklas Hansen, Hao Su, Xiaolong WangICLR 2024 · 被引用 388 次
- Efficient Online Reinforcement Learning with Offline DataPhilip J. Ball, Laura Smith, Ilya Kostrikov, Sergey LevineICML 2023 · 被引用 326 次
- The Primacy Bias in Deep Reinforcement LearningEvgenii Nikishin, Max Schwarzer, Pierluca D'Oro, Pierre-Luc Bacon 等ICML 2022 · 被引用 269 次
相关 Paper
- CrossQ: Batch Normalization in Deep Reinforcement Learning for Greater Sample Efficiency and SimplicityAditya Bhatt, Daniel Palenicek, Boris Belousov, Max Argus 等ICLR 2024 · 被引用 106 次
- Dropout Q-Functions for Doubly Efficient Reinforcement LearningTakuya Hiraoka, Takahisa Imagawa, Taisei Hashimoto, Takashi Onishi 等ICLR 2022 · 被引用 157 次
- Striving for Simplicity and Performance in Off-Policy DRL: Output Normalization and Non-Uniform SamplingChe Wang, Yanqiu Wu, Quan Vuong, Keith W. RossICML 2020 · 被引用 38 次
- Scaling Off-Policy Reinforcement Learning with Batch and Weight NormalizationDaniel Palenicek, Florian Vogt, Joe Watson, Jan PetersNeurIPS 2025 · 被引用 22 次
- SimBa: Simplicity Bias for Scaling Up Parameters in Deep Reinforcement LearningHojoon Lee, Dongyoon Hwang, Donghu Kim, Hyunseung Kim 等ICLR 2025
