Black-box Off-policy Estimation for Infinite-Horizon Reinforcement Learning
Ali Mousavi, Lihong Li, Qiang Liu, Denny Zhou
摘要
Off-policy estimation for long-horizon problems is important in many real-life applications such as healthcare and robotics, where high-fidelity simulators may not be available and on-policy evaluation is expensive or impossible. Recently, proposed an approach that avoids the curse of horizon suffered by typical importance-sampling-based methods. While showing promising results, this approach is limited in practice as it requires data being collected by a known behavior policy. In this work, we propose a novel approach that eliminates such limitations. In particular, we formulate the problem as solving for the fixed point of a "backward flow" operator and show that the fixed point solution gives the desired importance ratios of stationary distributions between the target and behavior policies. We analyze its asymptotic consistency and finite-sample generalization. Experiments on benchmarks verify the effectiveness of our proposed approach.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- GradientDICE: Rethinking Generalized Offline Estimation of Stationary ValuesShangtong Zhang, Bo Liu, Shimon WhitesonICML 2020 · 被引用 107 次
- Learning and Planning in Average-Reward Markov Decision ProcessesYi Wan, Abhishek Naik, Richard S. SuttonICML 2021 · 被引用 82 次
- Mean-Variance Policy Iteration for Risk-Averse Reinforcement LearningShangtong Zhang, Bo Liu, Shimon WhitesonAAAI 2021 · 被引用 44 次
- Average-Reward Off-Policy Policy Evaluation with Function ApproximationShangtong Zhang, Yi Wan, Richard S. Sutton, Shimon WhitesonICML 2021 · 被引用 39 次
- A Unified Framework for Alternating Offline Model Training and Policy LearningShentao Yang, Shujian Zhang, Yihao Feng, Mingyuan ZhouNeurIPS 2022 · 被引用 18 次
相关 Paper
- A Maximum-Entropy Approach to Off-Policy Evaluation in Average-Reward MDPsNevena Lazic, Dong Yin, Mehrdad Farajtabar, Nir Levine 等NeurIPS 2020 · 被引用 13 次
- Zero-Shot Off-Policy LearningArip Asadulaev, Maksim Bobrin, Salem Lahlou, Dmitry V. Dylov 等ICML 2026 · 被引用 1 次
- Doubly Robust Bias Reduction in Infinite Horizon Off-Policy EstimationZiyang Tang, Yihao Feng, Lihong Li, Dengyong Zhou 等ICLR 2020 · 被引用 72 次
- Doubly Robust Off-Policy Value and Gradient Estimation for Deterministic PoliciesNathan Kallus, Masatoshi UeharaNeurIPS 2020 · 被引用 16 次
- GenDICE: Generalized Offline Estimation of Stationary ValuesRuiyi Zhang, Bo Dai, Lihong Li, Dale SchuurmansICLR 2020 · 被引用 184 次
