Revisiting Bellman Errors for Offline Model Selection
Joshua P. Zitovsky, Daniel de Marchi, Rishabh Agarwal, Michael Rene Kosorok
Abstract
Offline model selection (OMS), that is, choosing the best policy from a set of many policies given only logged data, is crucial for applying offline RL in real-world settings. One idea that has been extensively explored is to select policies based on the mean squared Bellman error (MSBE) of the associated Q-functions. However, previous work has struggled to obtain adequate OMS performance with Bellman errors, leading many researchers to abandon the idea. To this end, we elucidate why previous work has seen pessimistic results with Bellman errors and identify conditions under which OMS algorithms based on Bellman errors will perform well. Moreover, we develop a new estimator of the MSBE that is more accurate than prior methods. Our estimator obtains impressive OMS performance on diverse discrete control tasks, including Atari games.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6c7d0c60-9b34-4a2d-acd6-e8a44a3322d4Cited by top-tier papers2
- Towards Robust Multi-Modal Reasoning via Model SelectionXiangyan Liu, Rongxue Li, Wei Ji, Tao LinICLR 2024 · 9 citations
- Model Selection for Off-policy Evaluation: New Algorithms and Experimental ProtocolPai Liu, Lingfeng Zhao, Shivangi Agarwal, Jinghan Liu et al.NeurIPS 2025 · 6 citations
Builds on11
- An Optimistic Perspective on Offline Reinforcement LearningRishabh Agarwal, Dale Schuurmans, Mohammad NorouziICML 2020 · 568 citations
- Implicit Under-Parameterization Inhibits Data-Efficient Deep Reinforcement LearningAviral Kumar, Rishabh Agarwal, Dibya Ghosh, Sergey LevineICLR 2021 · 155 citations
- Online and Offline Reinforcement Learning by Planning with a Learned ModelJulian Schrittwieser, Thomas Hubert, Amol Mandhane, Mohammadamin Barekatain et al.NeurIPS 2021 · 149 citations
- Batch Value-function Approximation with Only RealizabilityTengyang Xie, Nan JiangICML 2021 · 131 citations
- Benchmarks for Deep Off-Policy EvaluationJustin Fu, Mohammad Norouzi, Ofir Nachum, George Tucker et al.ICLR 2021 · 112 citations
Related papers
- Model-Bellman Inconsistency for Model-based Offline Reinforcement LearningYihao Sun, Jiaji Zhang, Chengxing Jia, Haoxin Lin et al.ICML 2023 · 61 citations
- Oracle Inequalities for Model Selection in Offline Reinforcement LearningJonathan N. Lee, George Tucker, Ofir Nachum, Bo Dai et al.NeurIPS 2022 · 14 citations
- Model-Based Offline Reinforcement Learning with Local MisspecificationKefan Dong, Yannis Flet-Berliac, Allen Nie, Emma BrunskillAAAI 2023 · 6 citations
- OPERA: Automatic Offline Policy Evaluation with Re-weighted Aggregates of Multiple EstimatorsAllen Nie, Yash Chandak, Christina J. Yuan, Anirudhan Badrinath et al.NeurIPS 2024 · 7 citations
- Policy-Adaptive Estimator Selection for Off-Policy EvaluationTakuma Udagawa, Haruka Kiyohara, Yusuke Narita, Yuta Saito et al.AAAI 2023 · 29 citations
