A Maximum-Entropy Approach to Off-Policy Evaluation in Average-Reward MDPs
Nevena Lazic, Dong Yin, Mehrdad Farajtabar, Nir Levine, Dilan Görür, Chris Harris, Dale Schuurmans
摘要
This work focuses on off-policy evaluation (OPE) with function approximation in infinite-horizon undiscounted Markov decision processes (MDPs). For MDPs that are ergodic and linear (i.e. where rewards and dynamics are linear in some known features), we provide the first finite-sample OPE error bound, extending existing results beyond the episodic and discounted cases. In a more general setting, when the feature dynamics are approximately linear and for arbitrary rewards, we propose a new approach for estimating stationary distributions with function approximation. We formulate this problem as finding the maximum-entropy distribution subject to matching feature expectations under empirical dynamics. We show that this results in an exponential-family distribution whose sufficient statistics are the features, paralleling maximum-entropy approaches in supervised learning. We demonstrate the effectiveness of the proposed OPE approaches in multiple environments. Introduction Recently, there have been considerable advances in reinforcement learning (RL), with algorithms achieving impressive performance on game playing and simple robotic tasks. Successful approaches typically learn through direct (online) interaction with the environment. However, in many real applications, access to the environment is limited to a fixed dataset, due to considerations of cost, safety, or time. One key challenge in this setting is off-policy evaluation (OPE): the task of evaluating the performance of a target policy given samples collected by a behavior policy. The focus of our work is OPE in infinite-horizon undiscounted MDPs, which capture long-horizon tasks such as game playing, routing, and the control of physical systems. Most recent state-of-the-art OPE methods for this setting estimate the ratios of stationary distributions of the target and behavior policy [
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Average-Reward Off-Policy Policy Evaluation with Function ApproximationShangtong Zhang, Yi Wan, Richard S. Sutton, Shimon WhitesonICML 2021 · 被引用 39 次
- Sparse Feature Selection Makes Batch Reinforcement Learning More Sample EfficientBotao Hao, Yaqi Duan, Tor Lattimore, Csaba Szepesvári 等ICML 2021 · 被引用 29 次
- Proximal Point Imitation LearningLuca Viano, Angeliki Kamoutsi, Gergely Neu, Igor Krawczuk 等NeurIPS 2022 · 被引用 27 次
- Non-asymptotic Confidence Intervals of Off-policy Evaluation: Primal and Dual BoundsYihao Feng, Ziyang Tang, Na Zhang, Qiang LiuICLR 2021 · 被引用 13 次
- Imitation Learning in Discounted Linear MDPs without exploration assumptionsLuca Viano, Stratis Skoulakis, Volkan CevherICML 2024 · 被引用 10 次
它引用的顶会 Paper2
相关 Paper
- A Unifying View of Coverage in Linear Off-policy EvaluationPhilip Amortila, Audrey Huang, Akshay Krishnamurthy, Nan JiangICLR 2026 · 被引用 2 次
- Distributional Offline Policy Evaluation with Predictive Error GuaranteesRunzhe Wu, Masatoshi Uehara, Wen SunICML 2023 · 被引用 19 次
- Black-box Off-policy Estimation for Infinite-Horizon Reinforcement LearningAli Mousavi, Lihong Li, Qiang Liu, Denny ZhouICLR 2020 · 被引用 33 次
- Asymptotically Exact Error Characterization of Offline Policy Evaluation with Misspecified Linear ModelsKohei MiyaguchiNeurIPS 2021 · 被引用 3 次
- A Principled Path to Fitted Distributional EvaluationSungee Hong, Jiayi Wang, Zhengling Qi, Raymond K. W. WongNeurIPS 2025
