Bellman Residual Orthogonalization for Offline Reinforcement Learning
Andrea Zanette, Martin J. Wainwright
Abstract
We propose and analyze a reinforcement learning principle that approximates the Bellman equations by enforcing their validity only along an user-defined space of test functions. Focusing on applications to model-free offline RL with function approximation, we exploit this principle to derive confidence intervals for off-policy evaluation, as well as to optimize over policies within a prescribed policy class. We prove an oracle inequality on our policy optimization procedure in terms of a trade-off between the value and uncertainty of an arbitrary comparator policy. Different choices of test function spaces allow us to tackle different problems within a common framework. We characterize the loss of efficiency in moving from on-policy to off-policy data using our procedures, and establish connections to concentrability coefficients studied in past work. We examine in depth the implementation of our methods with linear function approximation, and provide theoretical guarantees with polynomial-time implementations even when Bellman closure does not hold.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 053f289b-6924-47fc-b9a4-285f15cbc502Cited by top-tier papers6
- Future-Dependent Value-Based Off-Policy Evaluation in POMDPsMasatoshi Uehara, Haruka Kiyohara, Andrew Bennett, Victor Chernozhukov et al.NeurIPS 2023 · 31 citations
- On the Statistical Efficiency of Reward-Free Exploration in Non-Linear RLJinglin Chen, Aditya Modi, Akshay Krishnamurthy, Nan Jiang et al.NeurIPS 2022 · 31 citations
- When is Realizability Sufficient for Off-Policy Reinforcement Learning?Andrea ZanetteICML 2023 · 16 citations
- Offline Minimax Soft-Q-learning Under Realizability and Partial CoverageMasatoshi Uehara, Nathan Kallus, Jason D. Lee, Wen SunNeurIPS 2023 · 10 citations
- OMPO: A Unified Framework for RL under Policy and Dynamics ShiftsYu Luo, Tianying Ji, Fuchun Sun, Jianwei Zhang et al.ICML 2024 · 5 citations
Builds on26
- MOPO: Model-based Offline Policy OptimizationTianhe Yu, Garrett Thomas, Lantao Yu, Stefano Ermon et al.NeurIPS 2020 · 989 citations
- An Optimistic Perspective on Offline Reinforcement LearningRishabh Agarwal, Dale Schuurmans, Mohammad NorouziICML 2020 · 568 citations
- Is Pessimism Provably Efficient for Offline RL?Ying Jin, Zhuoran Yang, Zhaoran WangICML 2021 · 419 citations
- Bridging Offline Reinforcement Learning and Imitation Learning: A Tale of PessimismParia Rashidinejad, Banghua Zhu, Cong Ma, Jiantao Jiao et al.NeurIPS 2021 · 373 citations
- Bellman-consistent Pessimism for Offline Reinforcement LearningTengyang Xie, Ching-An Cheng, Nan Jiang, Paul Mineiro et al.NeurIPS 2021 · 339 citations
Related papers
- Offline Learning in Markov Games with General Function ApproximationYuheng Zhang, Yu Bai, Nan JiangICML 2023 · 17 citations
- Non-asymptotic Confidence Intervals of Off-policy Evaluation: Primal and Dual BoundsYihao Feng, Ziyang Tang, Na Zhang, Qiang LiuICLR 2021 · 13 citations
- Provable Benefits of Actor-Critic Methods for Offline Reinforcement LearningAndrea Zanette, Martin J. Wainwright, Emma BrunskillNeurIPS 2021 · 140 citations
- Model-based Reinforcement Learning for Confounded POMDPsMao Hong, Zhengling Qi, Yanxun XuICML 2024 · 5 citations
- Oracle Inequalities for Model Selection in Offline Reinforcement LearningJonathan N. Lee, George Tucker, Ofir Nachum, Bo Dai et al.NeurIPS 2022 · 14 citations
