Deeply-Debiased Off-Policy Interval Estimation
Chengchun Shi, Runzhe Wan, Victor Chernozhukov, Rui Song
Abstract
Off-policy evaluation learns a target policy's value with a historical dataset generated by a different behavior policy. In addition to a point estimate, many applications would benefit significantly from having a confidence interval (CI) that quantifies the uncertainty of the point estimate. In this paper, we propose a novel deeply-debiasing procedure to construct an efficient, robust, and flexible CI on a target policy's value. Our method is justified by theoretical results and numerical experiments. A Python implementation of the proposed procedure is available at https://github.com/RunzheStat/ D2OPE .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 14c19637-b756-444d-80a7-ee9ffcdbbc2dCited by top-tier papers15
- On Well-posedness and Minimax Optimal Rates of Nonparametric Q-function Estimation in Off-policy EvaluationXiaohong Chen, Zhengling QiICML 2022 · 36 citations
- A Minimax Learning Approach to Off-Policy Evaluation in Confounded Partially Observable Markov Decision ProcessesChengchun Shi, Masatoshi Uehara, Jiawei Huang, Nan JiangICML 2022 · 31 citations
- Optimal Treatment Allocation for Efficient Policy Evaluation in Sequential Decision MakingTing Li, Chengchun Shi, Jianing Wang, Fan Zhou et al.NeurIPS 2023 · 21 citations
- Safe Exploration for Efficient Policy Evaluation and ComparisonRunzhe Wan, Branislav Kveton, Rui SongICML 2022 · 16 citations
- Doubly Robust Alignment for Large Language ModelsErhan Xu, Kai Ye, Hongyi Zhou, Luhan Zhu et al.NeurIPS 2025 · 14 citations
Builds on8
- Minimax Weight and Q-Function Learning for Off-Policy EvaluationMasatoshi Uehara, Jiawei Huang, Nan JiangICML 2020 · 199 citations
- Inference for Batched BanditsKelly W. Zhang, Lucas Janson, Susan A. MurphyNeurIPS 2020 · 115 citations
- CoinDICE: Off-Policy Confidence Interval EstimationBo Dai, Ofir Nachum, Yinlam Chow, Lihong Li et al.NeurIPS 2020 · 96 citations
- Doubly Robust Bias Reduction in Infinite Horizon Off-Policy EstimationZiyang Tang, Yihao Feng, Lihong Li, Dengyong Zhou et al.ICLR 2020 · 72 citations
- Minimax Value Interval for Off-Policy Evaluation and Policy OptimizationNan Jiang, Jiawei HuangNeurIPS 2020 · 68 citations
Related papers
- An Instrumental Variable Approach to Confounded Off-Policy EvaluationYang Xu, Jin Zhu, Chengchun Shi, Shikai Luo et al.ICML 2023 · 24 citations
- Accountable Off-Policy Evaluation With Kernel Bellman StatisticsYihao Feng, Tongzheng Ren, Ziyang Tang, Qiang LiuICML 2020 · 45 citations
- High-Confidence Off-Policy (or Counterfactual) Variance EstimationYash Chandak, Shiv Shankar, Philip S. ThomasAAAI 2021 · 10 citations
- Non-asymptotic Confidence Intervals of Off-policy Evaluation: Primal and Dual BoundsYihao Feng, Ziyang Tang, Na Zhang, Qiang LiuICLR 2021 · 13 citations
- Simultaneous Statistical Inference for Off-Policy Evaluation in Reinforcement LearningTianpai Luo, Xinyuan Fan, Weichi WuNeurIPS 2025 · 1 citation
