Minimax Weight and Q-Function Learning for Off-Policy Evaluation
Masatoshi Uehara, Jiawei Huang, Nan Jiang
摘要
We provide theoretical investigations into off-policy evaluation in reinforcement learning using function approximators for (marginalized) importance weights and value functions. Our contributions include: (1) A new estimator, MWL, that directly estimates importance ratios over the state-action distributions, removing the reliance on knowledge of the behavior policy as in prior work (Liu et al., 2018). (2) Another new estimator, MQL, obtained by swapping the roles of importance weights and value-functions in MWL. MQL has an intuitive interpretation of minimizing average Bellman errors and can be combined with MWL in a doubly robust manner. (3) Several additional results that offer further insights into these methods, including the sample complexity analyses of MWL and MQL, their asymptotic optimality in the tabular setting, how the learned importance weights depend the choice of the discriminator class, and how our methods provide a unified view of some old and new algorithms in RL.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper91
- Is Pessimism Provably Efficient for Offline RL?Ying Jin, Zhuoran Yang, Zhaoran WangICML 2021 · 被引用 419 次
- Bridging Offline Reinforcement Learning and Imitation Learning: A Tale of PessimismParia Rashidinejad, Banghua Zhu, Cong Ma, Jiantao Jiao 等NeurIPS 2021 · 被引用 373 次
- Pessimistic Model-based Offline Reinforcement Learning under Partial CoverageMasatoshi Uehara, Wen SunICLR 2022 · 被引用 176 次
- What are the Statistical Limits of Offline RL with Linear Function Approximation?Ruosong Wang, Dean P. Foster, Sham M. KakadeICLR 2021 · 被引用 172 次
- Provable Benefits of Actor-Critic Methods for Offline Reinforcement LearningAndrea Zanette, Martin J. Wainwright, Emma BrunskillNeurIPS 2021 · 被引用 140 次
相关 Paper
- Minimax Value Interval for Off-Policy Evaluation and Policy OptimizationNan Jiang, Jiawei HuangNeurIPS 2020 · 被引用 68 次
- Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy EvaluationHongyi Zhou, Josiah P. Hanna, Jin Zhu, Ying Yang 等ICML 2025
- Beyond the Return: Off-policy Function Estimation under User-specified Error-measuring DistributionsAudrey Huang, Nan JiangNeurIPS 2022 · 被引用 9 次
- Double Reinforcement Learning for Efficient and Robust Off-Policy EvaluationNathan Kallus, Masatoshi UeharaICML 2020 · 被引用 6 次
- Doubly Robust Off-Policy Value and Gradient Estimation for Deterministic PoliciesNathan Kallus, Masatoshi UeharaNeurIPS 2020 · 被引用 16 次
