Why Should I Trust You, Bellman? The Bellman Error is a Poor Replacement for Value Error
Scott Fujimoto, David Meger, Doina Precup, Ofir Nachum, Shixiang Shane Gu
Abstract
In this work, we study the use of the Bellman equation as a surrogate objective for value prediction accuracy. While the Bellman equation is uniquely solved by the true value function over all state-action pairs, we find that the Bellman error (the difference between both sides of the equation) is a poor proxy for the accuracy of the value function. In particular, we show that (1) due to cancellations from both sides of the Bellman equation, the magnitude of the Bellman error is only weakly related to the distance to the true value function, even when considering all state-action pairs, and (2) in the finite data regime, the Bellman equation can be satisfied exactly by infinitely many suboptimal solutions. This means that the Bellman error can be minimized without improving the accuracy of the value function. We demonstrate these phenomena through a series of propositions, illustrative toy examples, and empirical analysis in standard benchmark domains.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f8b7a4a9-1ff6-4e61-a243-e5ede0412a83Cited by top-tier papers17
- For SALE: State-Action Representation Learning for Deep Reinforcement LearningScott Fujimoto, Wei-Di Chang, Edward J. Smith, Shixiang Gu et al.NeurIPS 2023 · 128 citations
- Is Value Learning Really the Main Bottleneck in Offline RL?Seohong Park, Kevin Frans, Sergey Levine, Aviral KumarNeurIPS 2024 · 99 citations
- Disentangling Transfer in Continual Reinforcement LearningMaciej Wolczyk, Michal Zajac, Razvan Pascanu, Lukasz Kucinski et al.NeurIPS 2022 · 46 citations
- Learning to Mitigate AI Collusion on Economic PlatformsGianluca Brero, Eric Mibuari, Nicolas Lepore, David C. ParkesNeurIPS 2022 · 22 citations
- Offline Reinforcement Learning with Closed-Form Policy Improvement OperatorsJiachen Li, Edwin Zhang, Ming Yin, Qinxun Bai et al.ICML 2023 · 18 citations
Builds on7
- Agent57: Outperforming the Atari Human BenchmarkAdrià Puigdomènech Badia, Bilal Piot, Steven Kapturowski, Pablo Sprechmann et al.ICML 2020 · 584 citations
- Implementation Matters in Deep RL: A Case Study on PPO and TRPOLogan Engstrom, Andrew Ilyas, Shibani Santurkar, Dimitris Tsipras et al.ICLR 2020 · 305 citations
- What are the Statistical Limits of Offline RL with Linear Function Approximation?Ruosong Wang, Dean P. Foster, Sham M. KakadeICLR 2021 · 172 citations
- Representation Matters: Offline Pretraining for Sequential Decision MakingMengjiao Yang, Ofir NachumICML 2021 · 126 citations
- Benchmarks for Deep Off-Policy EvaluationJustin Fu, Mohammad Norouzi, Ofir Nachum, George Tucker et al.ICLR 2021 · 112 citations
Related papers
- Solving hidden monotone variational inequalities with surrogate lossesRyan D'Orazio, Danilo Vucetic, Zichu Liu, Junhyung Lyle Kim et al.ICLR 2025
- Defining and Characterizing Reward GamingJoar Skalse, Nikolaus H. R. Howe, Dmitrii Krasheninnikov, David KruegerNeurIPS 2022 · 466 citations
- On the Statistical Benefits of Temporal Difference LearningDavid Cheikhi, Daniel RussoICML 2023 · 6 citations
- Reinforcement Learning in Linear MDPs: Constant Regret and Representation SelectionMatteo Papini, Andrea Tirinzoni, Aldo Pacchiano, Marcello Restelli et al.NeurIPS 2021 · 26 citations
- Bellman-consistent Pessimism for Offline Reinforcement LearningTengyang Xie, Ching-An Cheng, Nan Jiang, Paul Mineiro et al.NeurIPS 2021 · 339 citations
