Black-box Off-policy Estimation for Infinite-Horizon Reinforcement Learning
Ali Mousavi, Lihong Li, Qiang Liu, Denny Zhou
Abstract
Off-policy estimation for long-horizon problems is important in many real-life applications such as healthcare and robotics, where high-fidelity simulators may not be available and on-policy evaluation is expensive or impossible. Recently, proposed an approach that avoids the curse of horizon suffered by typical importance-sampling-based methods. While showing promising results, this approach is limited in practice as it requires data being collected by a known behavior policy. In this work, we propose a novel approach that eliminates such limitations. In particular, we formulate the problem as solving for the fixed point of a "backward flow" operator and show that the fixed point solution gives the desired importance ratios of stationary distributions between the target and behavior policies. We analyze its asymptotic consistency and finite-sample generalization. Experiments on benchmarks verify the effectiveness of our proposed approach.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 75ea5a15-68f9-4ff7-bfaf-d25ac8292501Cited by top-tier papers10
- GradientDICE: Rethinking Generalized Offline Estimation of Stationary ValuesShangtong Zhang, Bo Liu, Shimon WhitesonICML 2020 · 107 citations
- Learning and Planning in Average-Reward Markov Decision ProcessesYi Wan, Abhishek Naik, Richard S. SuttonICML 2021 · 82 citations
- Mean-Variance Policy Iteration for Risk-Averse Reinforcement LearningShangtong Zhang, Bo Liu, Shimon WhitesonAAAI 2021 · 44 citations
- Average-Reward Off-Policy Policy Evaluation with Function ApproximationShangtong Zhang, Yi Wan, Richard S. Sutton, Shimon WhitesonICML 2021 · 39 citations
- A Unified Framework for Alternating Offline Model Training and Policy LearningShentao Yang, Shujian Zhang, Yihao Feng, Mingyuan ZhouNeurIPS 2022 · 18 citations
Related papers
- A Maximum-Entropy Approach to Off-Policy Evaluation in Average-Reward MDPsNevena Lazic, Dong Yin, Mehrdad Farajtabar, Nir Levine et al.NeurIPS 2020 · 13 citations
- Zero-Shot Off-Policy LearningArip Asadulaev, Maksim Bobrin, Salem Lahlou, Dmitry V. Dylov et al.ICML 2026 · 1 citation
- Doubly Robust Bias Reduction in Infinite Horizon Off-Policy EstimationZiyang Tang, Yihao Feng, Lihong Li, Dengyong Zhou et al.ICLR 2020 · 72 citations
- Doubly Robust Off-Policy Value and Gradient Estimation for Deterministic PoliciesNathan Kallus, Masatoshi UeharaNeurIPS 2020 · 16 citations
- GenDICE: Generalized Offline Estimation of Stationary ValuesRuiyi Zhang, Bo Dai, Lihong Li, Dale SchuurmansICLR 2020 · 184 citations
