Stop Regressing: Training Value Functions via Classification for Scalable Deep RL
Jesse Farebrother, Jordi Orbay, Quan Vuong, Adrien Ali Taïga, Yevgen Chebotar, Ted Xiao, Alex Irpan, Sergey Levine, Pablo Samuel Castro, Aleksandra Faust, Aviral Kumar, Rishabh Agarwal
Abstract
Value functions are a central component of deep reinforcement learning (RL). These functions, parameterized by neural networks, are trained using a mean squared error regression objective to match bootstrapped target values. However, scaling value-based RL methods that use regression to large networks, such as high-capacity Transformers, has proven challenging. This difficulty is in stark contrast to supervised learning: by leveraging a cross-entropy classification loss, supervised methods have scaled reliably to massive networks. Observing this discrepancy, in this paper, we investigate whether the scalability of deep RL can also be improved simply by using classification in place of regression for training value functions. We demonstrate that value functions trained with categorical cross-entropy significantly improves performance and scalability in a variety of domains. These include: single-task RL on Atari 2600 games with SoftMoEs, multi-task RL on Atari with large-scale ResNets, robotic manipulation with Q-transformers, playing Chess without search, and a language-agent Wordle task with high-capacity Transformers, achieving state-of-the-art results on these domains. Through careful analysis, we show that the benefits of categorical cross-entropy primarily stem from its ability to mitigate issues inherent to value-based RL, such as noisy targets and non-stationarity. Overall, we argue that a simple shift to training value functions with categorical cross-entropy can yield substantial improvements in the scalability of deep RL at little-to-no cost.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c6797050-a481-4700-a9d3-a14f5ffc7671Cited by top-tier papers63
- DigiRL: Training In-The-Wild Device-Control Agents with Autonomous Reinforcement LearningHao Bai, Yifei Zhou, Jiayi Pan, Mert Cemri et al.NeurIPS 2024 · 239 citations
- Is Behavior Cloning All You Need? Understanding Horizon in Imitation LearningDylan J. Foster, Adam Block, Dipendra MisraNeurIPS 2024 · 112 citations
- REBEL: Reinforcement Learning via Regressing Relative RewardsZhaolin Gao, Jonathan D. Chang, Wenhao Zhan, Owen Oertell et al.NeurIPS 2024 · 82 citations
- Mixtures of Experts Unlock Parameter Scaling for Deep RLJohan S. Obando-Ceron, Ghada Sokar, Timon Willi, Clare Lyle et al.ICML 2024 · 74 citations
- Horizon Reduction Makes RL ScalableSeohong Park, Kevin Frans, Deepinder Mann, Benjamin Eysenbach et al.NeurIPS 2025 · 60 citations
Builds on23
- Conservative Q-Learning for Offline Reinforcement LearningAviral Kumar, Aurick Zhou, George Tucker, Sergey LevineNeurIPS 2020 · 2,881 citations
- Deep Reinforcement Learning at the Edge of the Statistical PrecipiceRishabh Agarwal, Max Schwarzer, Pablo Samuel Castro, Aaron C. Courville et al.NeurIPS 2021 · 1,067 citations
- An Optimistic Perspective on Offline Reinforcement LearningRishabh Agarwal, Dale Schuurmans, Mohammad NorouziICML 2020 · 568 citations
- TD-MPC2: Scalable, Robust World Models for Continuous ControlNicklas Hansen, Hao Su, Xiaolong WangICLR 2024 · 388 citations
- Multi-Game Decision TransformersKuang-Huei Lee, Ofir Nachum, Mengjiao Yang, Lisa Lee et al.NeurIPS 2022 · 279 citations
Related papers
- TQL: Scaling Q-Functions with Transformers by Preventing Attention CollapsePerry Dong, Kuo-Han Hung, Alexander Swerdlow, Dorsa Sadigh et al.ICML 2026 · 7 citations
- AMAGO-2: Breaking the Multi-Task Barrier in Meta-Reinforcement Learning with TransformersJake Grigsby, Justin Sasek, Samyak Parajuli, Daniel Adebi et al.NeurIPS 2024 · 19 citations
- Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task LearnersMichal Nauman, Marek Cygan, Carmelo Sferrazza, Aviral Kumar et al.NeurIPS 2025 · 26 citations
- Offline Actor-Critic Reinforcement Learning Scales to Large ModelsJost Tobias Springenberg, Abbas Abdolmaleki, Jingwei Zhang, Oliver Groth et al.ICML 2024 · 37 citations
- Label Encoding for Regression NetworksDeval Shah, Zi Yu Xue, Tor M. AamodtICLR 2022 · 23 citations
