Extreme Q-Learning: MaxEnt RL without Entropy
Divyansh Garg, Joey Hejna, Matthieu Geist, Stefano Ermon
Abstract
Modern Deep Reinforcement Learning (RL) algorithms require estimates of the maximal Q-value, which are difficult to compute in continuous domains with an infinite number of possible actions. In this work, we introduce a new update rule for online and offline RL which directly models the maximal value using Extreme Value Theory (EVT), drawing inspiration from economics. By doing so, we avoid computing Q-values using out-of-distribution actions which is often a substantial source of error. Our key insight is to introduce an objective that directly estimates the optimal soft-value functions (LogSumExp) in the maximum entropy RL setting without needing to sample from a policy. Using EVT, we derive our Extreme Q-Learning framework and consequently online and, for the first time, offline MaxEnt Q-learning algorithms, that do not explicitly require access to a policy or its entropy. Our method obtains consistently strong performance in the D4RL benchmark, outperforming prior works by 10+ points on the challenging Franka Kitchen tasks while offering moderate improvements over SAC and TD3 on online DM Control tasks. Visualizations and code can be found on our website at https://div99.github.io/XQL/.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 59d96d00-6edb-4ad8-bce4-207ab0a450afCited by top-tier papers81
- HIQL: Offline Goal-Conditioned RL with Latent States as ActionsSeohong Park, Dibya Ghosh, Benjamin Eysenbach, Sergey LevineNeurIPS 2023 · 173 citations
- For SALE: State-Action Representation Learning for Deep Reinforcement LearningScott Fujimoto, Wei-Di Chang, Edward J. Smith, Shixiang Gu et al.NeurIPS 2023 · 128 citations
- Is Value Learning Really the Main Bottleneck in Offline RL?Seohong Park, Kevin Frans, Sergey Levine, Aviral KumarNeurIPS 2024 · 99 citations
- Inverse Preference Learning: Preference-based RL without a Reward FunctionJoey Hejna, Dorsa SadighNeurIPS 2023 · 92 citations
- Consistency Models as a Rich and Efficient Policy Class for Reinforcement LearningZihan Ding, Chi JinICLR 2024 · 73 citations
Builds on7
- Conservative Q-Learning for Offline Reinforcement LearningAviral Kumar, Aurick Zhou, George Tucker, Sergey LevineNeurIPS 2020 · 2,881 citations
- Decision Transformer: Reinforcement Learning via Sequence ModelingLili Chen, Kevin Lu, Aravind Rajeswaran, Kimin Lee et al.NeurIPS 2021 · 2,557 citations
- Offline Reinforcement Learning with Implicit Q-LearningIlya Kostrikov, Ashvin Nair, Sergey LevineICLR 2022 · 1,402 citations
- A Minimalist Approach to Offline Reinforcement LearningScott Fujimoto, Shixiang Shane GuNeurIPS 2021 · 1,292 citations
- Uncertainty-Based Offline Reinforcement Learning with Diversified Q-EnsembleGaon An, Seungyong Moon, Jang-Hyun Kim, Hyun Oh SongNeurIPS 2021 · 430 citations
Related papers
- Maximum Entropy Reinforcement Learning via Energy-Based Normalizing FlowChen-Hao Chao, Chien Feng, Wei-Fang Sun, Cheng-Kuang Lee et al.NeurIPS 2024 · 29 citations
- DR-SAC: Distributionally Robust Soft Actor-Critic for Reinforcement Learning under UncertaintyMingxuan Cui, Duo Zhou, Yuxuan Han, Grani A. Hanasusanto et al.ICLR 2026 · 6 citations
- Mind Your Entropy: From Maximum Entropy to Trajectory Entropy-Constrained RLGuojian Zhan, Likun Wang, Pengcheng Wang, Feihong Zhang et al.ICML 2026
- Striving for Simplicity and Performance in Off-Policy DRL: Output Normalization and Non-Uniform SamplingChe Wang, Yanqiu Wu, Quan Vuong, Keith W. RossICML 2020 · 38 citations
- Maximum Entropy Reinforcement Learning with Diffusion PolicyXiaoyi Dong, Jian Cheng, Xi Sheryl ZhangICML 2025
