High-Confidence Off-Policy (or Counterfactual) Variance Estimation
Yash Chandak, Shiv Shankar, Philip S. Thomas
Abstract
Many sequential decision-making systems leverage data collected using prior policies to propose a new policy. For critical applications, it is important that high-confidence guarantees on the new policy’s behavior are provided before deployment, to ensure that the policy will behave as desired. Prior works have studied high-confidence off-policy estimation of the expected return, however, high-confidence off-policy estimation of the variance of returns can be equally critical for high-risk applications. In this paper we tackle the previously open problem of estimating and bounding, with high confidence, the variance of returns from off-policy data.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4f2e4898-7613-4053-98f3-d6f0c2873d99Cited by top-tier papers4
- Universal Off-Policy EvaluationYash Chandak, Scott Niekum, Bruno C. da Silva, Erik G. Learned-Miller et al.NeurIPS 2021 · 64 citations
- Off-Policy Risk Assessment in Contextual BanditsAudrey Huang, Liu Leqi, Zachary C. Lipton, Kamyar AzizzadenesheliNeurIPS 2021 · 44 citations
- Near-Minimax-Optimal Distributional Reinforcement Learning with a Generative ModelMark Rowland, Kevin Kevin Li, Rémi Munos, Clare Lyle et al.NeurIPS 2024 · 9 citations
- Supervised Learning with General Risk FunctionalsLiu Leqi, Audrey Huang, Zachary C. Lipton, Kamyar AzizzadenesheliICML 2022 · 7 citations
Related papers
- Empirical Likelihood for Contextual BanditsNikos Karampatziakis, John Langford, Paul MineiroNeurIPS 2020 · 11 citations
- Accountable Off-Policy Evaluation With Kernel Bellman StatisticsYihao Feng, Tongzheng Ren, Ziyang Tang, Qiang LiuICML 2020 · 45 citations
- Deeply-Debiased Off-Policy Interval EstimationChengchun Shi, Runzhe Wan, Victor Chernozhukov, Rui SongICML 2021 · 43 citations
- Off-Policy Interval Estimation with Lipschitz Value IterationZiyang Tang, Yihao Feng, Na Zhang, Jian Peng et al.NeurIPS 2020 · 6 citations
- Variance-Aware Off-Policy Evaluation with Linear Function ApproximationYifei Min, Tianhao Wang, Dongruo Zhou, Quanquan GuNeurIPS 2021 · 43 citations
