XQC: Well-conditioned Optimization Accelerates Deep Reinforcement Learning
Daniel Palenicek, Florian Vogt, Joe Watson, Ingmar Posner, Jan Peters
Abstract
Sample efficiency is a central property of effective deep reinforcement learning algorithms. Recent work has improved this through added complexity, such as larger models, exotic network architectures, and more complex algorithms, which are typically motivated purely by empirical performance. We take a more principled approach by focusing on the optimization landscape of the critic network. Using the eigenspectrum and condition number of the critic's Hessian, we systematically investigate the impact of common architectural design decisions on training dynamics. Our analysis reveals that a novel combination of batch normalization (BN), weight normalization (WN), and a distributional crossentropy (CE) loss produces condition numbers orders of magnitude smaller than baselines. This combination also naturally bounds gradient norms, a property critical for maintaining a stable effective learning rate under non-stationary targets and bootstrapping. Based on these insights, we introduce XQC: a well-motivated, sample-efficient deep actor-critic algorithm built upon soft actor-critic that embodies these optimization-aware principles. We achieve state-of-the-art sample efficiency across 55 proprioception and 15 vision-based continuous control tasks, all while using significantly fewer parameters than competing methods. Our code is available at danielpalenicek.github.io/projects/xqc. Figure 1 : Well-conditioned network architectures yield state-of-the-art RL performance. Our algorithm, XQC with a BN and WN-based architecture and a CE loss, achieves competitive performance against state-of-the-art baselines across 55 proprioceptive continuous control tasks from four different benchmarks with a single set of hyperparameters. Notably, with ∼ 4.5× fewer parameters and ∼ 5× less compute in terms of FLOP/S than SIMBA-V2, the closest competitor. XQC's efficiency carries over to RL from pixels on 15 vision-based DMC tasks, significantly improving on DRQ-V2.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers5
- Relative Entropy Pathwise Policy OptimizationClaas Voelcker, Axel Brunnbauer, Marcel Hussing, Michal Nauman et al.ICLR 2026 · 6 citations
- Stable Deep Reinforcement Learning via Isotropic Gaussian RepresentationsAli Saheb pasand, Johan Obando-Ceron, Aaron Courville, Pouya Bashivan et al.ICML 2026 · 5 citations
- Use the Online Network If You Can: Towards Fast and Stable Reinforcement LearningAhmed Hendawy, Henrik Metternich, Théo Vincent, Mahdi Kallel et al.ICLR 2026 · 4 citations
- What Does Flow-Matching Bring to TD-Learning?Bhavya Agrawalla, Michal Nauman, Aviral KumarICML 2026 · 1 citation
- Trust Region Inverse Reinforcement Learning: Explicit Dual Ascent using Local Policy UpdatesAnish Abhijit Diwan, Davide Tateo, Christopher Mower, Haitham Bou Ammar et al.ICML 2026
Builds on27
- Dream to Control: Learning Behaviors by Latent ImaginationDanijar Hafner, Timothy P. Lillicrap, Jimmy Ba, Mohammad NorouziICLR 2020 · 1,852 citations
- Deep Reinforcement Learning at the Edge of the Statistical PrecipiceRishabh Agarwal, Max Schwarzer, Pablo Samuel Castro, Aaron C. Courville et al.NeurIPS 2021 · 1,067 citations
- TD-MPC2: Scalable, Robust World Models for Continuous ControlNicklas Hansen, Hao Su, Xiaolong WangICLR 2024 · 388 citations
- Efficient Online Reinforcement Learning with Offline DataPhilip J. Ball, Laura Smith, Ilya Kostrikov, Sergey LevineICML 2023 · 326 citations
- The Primacy Bias in Deep Reinforcement LearningEvgenii Nikishin, Max Schwarzer, Pierluca D'Oro, Pierre-Luc Bacon et al.ICML 2022 · 269 citations
Related papers
- CrossQ: Batch Normalization in Deep Reinforcement Learning for Greater Sample Efficiency and SimplicityAditya Bhatt, Daniel Palenicek, Boris Belousov, Max Argus et al.ICLR 2024 · 106 citations
- Dropout Q-Functions for Doubly Efficient Reinforcement LearningTakuya Hiraoka, Takahisa Imagawa, Taisei Hashimoto, Takashi Onishi et al.ICLR 2022 · 157 citations
- Striving for Simplicity and Performance in Off-Policy DRL: Output Normalization and Non-Uniform SamplingChe Wang, Yanqiu Wu, Quan Vuong, Keith W. RossICML 2020 · 38 citations
- Scaling Off-Policy Reinforcement Learning with Batch and Weight NormalizationDaniel Palenicek, Florian Vogt, Joe Watson, Jan PetersNeurIPS 2025 · 22 citations
- SimBa: Simplicity Bias for Scaling Up Parameters in Deep Reinforcement LearningHojoon Lee, Dongyoon Hwang, Donghu Kim, Hyunseung Kim et al.ICLR 2025
