Off-Policy Evaluation via Adaptive Weighting with Data from Contextual Bandits
Ruohan Zhan, Vitor Hadad, David A. Hirshberg, Susan Athey
摘要
It has become increasingly common for data to be collected adaptively, for example using contextual bandits. Historical data of this type can be used to evaluate other treatment assignment policies to guide future innovation or experiments. However, policy evaluation is challenging if the target policy differs from the one used to collect data, and popular estimators, including doubly robust (DR) estimators, can be plagued by bias, excessive variance, or both. In particular, when the pattern of treatment assignment in the collected data looks little like the pattern generated by the policy to be evaluated, the importance weights used in DR estimators explode, leading to excessive variance. In this paper, we improve the DR estimator by adaptively weighting observations to control its variance. We show that a t-statistic based on our improved estimator is asymptotically normal under certain conditions, allowing us to form confidence intervals and test hypotheses. Using synthetic data and public benchmarks, we provide empirical evidence for our estimator's improved accuracy and inferential properties relative to existing alternatives.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper17
- Statistical Inference with M-Estimators on Adaptively Collected DataKelly W. Zhang, Lucas Janson, Susan A. MurphyNeurIPS 2021 · 被引用 66 次
- Double/Debiased Machine Learning for Dynamic Treatment EffectsGreg Lewis, Vasilis SyrgkanisNeurIPS 2021 · 被引用 50 次
- Offline Neural Contextual Bandits: Pessimism, Optimization and GeneralizationThanh Nguyen-Tang, Sunil Gupta, A. Tuan Nguyen, Svetha VenkateshICLR 2022 · 被引用 35 次
- On Instance-Dependent Bounds for Offline Reinforcement Learning with Linear Function ApproximationThanh Nguyen-Tang, Ming Yin, Sunil Gupta, Svetha Venkatesh 等AAAI 2023 · 被引用 24 次
- The Adaptive Doubly Robust Estimator and a Paradox Concerning Logging PolicyMasahiro Kato, Kenichiro McAlinn, Shota YasuiNeurIPS 2021 · 被引用 23 次
它引用的顶会 Paper3
- Doubly robust off-policy evaluation with shrinkageYi Su, Maria Dimakopoulou, Akshay Krishnamurthy, Miroslav DudíkICML 2020 · 被引用 128 次
- Optimal Off-Policy Evaluation from Multiple Logging PoliciesNathan Kallus, Yuta Saito, Masatoshi UeharaICML 2021 · 被引用 44 次
- On Conditional Versus Marginal Bias in Multi-Armed BanditsJaehyeok Shin, Aaditya Ramdas, Alessandro RinaldoICML 2020 · 被引用 13 次
相关 Paper
- Post-Contextual-Bandit InferenceAurélien Bibaut, Maria Dimakopoulou, Nathan Kallus, Antoine Chambaz 等NeurIPS 2021 · 被引用 58 次
- Marginal Density Ratio for Off-Policy Evaluation in Contextual BanditsMuhammad Faaiz Taufiq, Arnaud Doucet, Rob Cornish, Jean-Francois TonNeurIPS 2023 · 被引用 14 次
- Statistical Inference on Multi-armed Bandits with Delayed FeedbackLei Shi, Jingshen Wang, Tianhao WuICML 2023 · 被引用 7 次
- Online Multi-Armed Bandits with Adaptive InferenceMaria Dimakopoulou, Zhimei Ren, Zhengyuan ZhouNeurIPS 2021 · 被引用 47 次
- Distributionally Robust Policy Evaluation and Learning for Continuous Treatment with Observational DataCheuk Hang Leung, Yiyan Huang, Yijun Li, Qi WuAAAI 2025 · 被引用 1 次
