The Optimal Approximation Factors in Misspecified Off-Policy Value Function Estimation
Philip Amortila, Nan Jiang, Csaba Szepesvári
Abstract
Theoretical guarantees in reinforcement learning (RL) are known to suffer multiplicative blow-up factors with respect to the misspecification error of function approximation. Yet, the nature of such approximation factors -- especially their optimal form in a given learning problem -- is poorly understood. In this paper we study this question in linear off-policy value function estimation, where many open questions remain. We study the approximation factor in a broad spectrum of settings, such as with the weighted -norm (where the weighting is the offline state distribution), the norm, the presence vs. absence of state aliasing, and full vs. partial coverage of the state space. We establish the optimal asymptotic approximation factors (up to constants) for all of these settings. In particular, our bounds identify two instance-dependent factors for the norm and only one for the norm, which are shown to dictate the hardness of off-policy evaluation under misspecification.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 41a147d9-eba8-4501-aee4-1b563f9f9667Cited by top-tier papers3
- Model Selection for Off-policy Evaluation: New Algorithms and Experimental ProtocolPai Liu, Lingfeng Zhao, Shivangi Agarwal, Jinghan Liu et al.NeurIPS 2025 · 6 citations
- Beyond Least Squares: Uniform Approximation and the Hidden Cost of MisspecificationDavide Maran, Csaba SzepesváriNeurIPS 2025 · 2 citations
- A Unifying View of Coverage in Linear Off-policy EvaluationPhilip Amortila, Audrey Huang, Akshay Krishnamurthy, Nan JiangICLR 2026 · 2 citations
Builds on8
- Minimax Weight and Q-Function Learning for Off-Policy EvaluationMasatoshi Uehara, Jiawei Huang, Nan JiangICML 2020 · 199 citations
- Learning with Good Feature Representations in Bandits and in RL with a Generative ModelTor Lattimore, Csaba Szepesvári, Gellért WeiszICML 2020 · 181 citations
- What are the Statistical Limits of Offline RL with Linear Function Approximation?Ruosong Wang, Dean P. Foster, Sham M. KakadeICLR 2021 · 172 citations
- Minimax-Optimal Off-Policy Evaluation with Linear Function ApproximationYaqi Duan, Zeyu Jia, Mengdi WangICML 2020 · 161 citations
- Representation Learning for Online and Offline RL in Low-rank MDPsMasatoshi Uehara, Xuezhou Zhang, Wen SunICLR 2022 · 138 citations
Related papers
- Beyond the Return: Off-policy Function Estimation under User-specified Error-measuring DistributionsAudrey Huang, Nan JiangNeurIPS 2022 · 9 citations
- Variance-Aware Off-Policy Evaluation with Linear Function ApproximationYifei Min, Tianhao Wang, Dongruo Zhou, Quanquan GuNeurIPS 2021 · 43 citations
- Minimax Optimal and Computationally Efficient Algorithms for Distributionally Robust Offline Reinforcement LearningZhishuai Liu, Pan XuNeurIPS 2024 · 21 citations
- Offline Actor-Critic for Average Reward MDPsWilliam G. Powell, Jeongyeol Kwon, Qiaomin Xie, Hanbaek LyuNeurIPS 2025
- On the Sample Complexity of Vanilla Model-Based Offline Reinforcement Learning with Dependent SamplesMustafa O. Karabag, Ufuk TopcuAAAI 2023 · 6 citations
