Policy-Adaptive Estimator Selection for Off-Policy Evaluation
Takuma Udagawa, Haruka Kiyohara, Yusuke Narita, Yuta Saito, Kei Tateno
摘要
Off-policy evaluation (OPE) aims to accurately evaluate the performance of counterfactual policies using only offline logged data. Although many estimators have been developed, there is no single estimator that dominates the others, because the estimators' accuracy can vary greatly depending on a given OPE task such as the evaluation policy, number of actions, and noise level. Thus, the data-driven estimator selection problem is becoming increasingly important and can have a significant impact on the accuracy of OPE. However, identifying the most accurate estimator using only the logged data is quite challenging because the ground-truth estimation accuracy of estimators is generally unavailable. This paper thus studies this challenging problem of estimator selection for OPE for the first time. In particular, we enable an estimator selection that is adaptive to a given OPE task, by appropriately subsampling available logged data and constructing pseudo policies useful for the underlying estimator selection task. Comprehensive experiments on both synthetic and real-world company data demonstrate that the proposed procedure substantially improves the estimator selection compared to a non-adaptive heuristic. Note that complete version with technical appendix is available on arXiv: http://arxiv.org/abs/2211.13904.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper14
- Off-Policy Evaluation for Large Action Spaces via Conjunct Effect ModelingYuta Saito, Qingyang Ren, Thorsten JoachimsICML 2023 · 被引用 34 次
- Off-Policy Evaluation of Slate Bandit Policies via Optimizing AbstractionHaruka Kiyohara, Masahiro Nomura, Yuta SaitoWWW 2024 · 被引用 18 次
- On (Normalised) Discounted Cumulative Gain as an Off-Policy Evaluation Metric for Top-n RecommendationOlivier Jeunen, Ivan Potapov, Aleksei UstimenkoKDD 2024 · 被引用 16 次
- Towards Assessing and Benchmarking Risk-Return Tradeoff of Off-Policy EvaluationHaruka Kiyohara, Ren Kishimoto, Kosuke Kawakami, Ken Kobayashi 等ICLR 2024 · 被引用 15 次
- Off-Policy Evaluation of Ranking Policies under Diverse User BehaviorHaruka Kiyohara, Masatoshi Uehara, Yusuke Narita, Nobuyuki Shimizu 等KDD 2023 · 被引用 8 次
它引用的顶会 Paper7
- Doubly robust off-policy evaluation with shrinkageYi Su, Maria Dimakopoulou, Akshay Krishnamurthy, Miroslav DudíkICML 2020 · 被引用 128 次
- Off-Policy Evaluation for Large Action Spaces via EmbeddingsYuta Saito, Thorsten JoachimsICML 2022 · 被引用 62 次
- Subgaussian and Differentiable Importance Sampling for Off-Policy Evaluation and LearningAlberto Maria Metelli, Alessio Russo, Marcello RestelliNeurIPS 2021 · 被引用 55 次
- Adaptive Estimator Selection for Off-Policy EvaluationYi Su, Pavithra Srinath, Akshay KrishnamurthyICML 2020 · 被引用 55 次
- Towards Hyperparameter-free Policy Selection for Offline Reinforcement LearningSiyuan Zhang, Nan JiangNeurIPS 2021 · 被引用 47 次
相关 Paper
- OPERA: Automatic Offline Policy Evaluation with Re-weighted Aggregates of Multiple EstimatorsAllen Nie, Yash Chandak, Christina J. Yuan, Anirudhan Badrinath 等NeurIPS 2024 · 被引用 7 次
- Off-Policy Evaluation under Nonignorable Missing DataHan Wang, Yang Xu, Wenbin Lu, Rui SongICML 2025
- Counterfactual Learning with General Data-Generating PoliciesYusuke Narita, Kyohei Okumura, Akihiro Shimizu, Kohei YataAAAI 2023 · 被引用 2 次
- Cross-Domain Off-Policy Evaluation and Learning for Contextual BanditsYuta Natsubori, Masataka Ushiku, Yuta SaitoICLR 2025
- Active Offline Policy SelectionKsenia Konyushkova, Yutian Chen, Thomas Paine, Çaglar Gülçehre 等NeurIPS 2021 · 被引用 35 次
