Statistical Inference with M-Estimators on Adaptively Collected Data
Kelly W. Zhang, Lucas Janson, Susan A. Murphy
摘要
Bandit algorithms are increasingly used in real-world sequential decision-making problems. Associated with this is an increased desire to be able to use the resulting datasets to answer scientific questions like: Did one type of ad lead to more purchases? In which contexts is a mobile health intervention effective? However, classical statistical approaches fail to provide valid confidence intervals when used with data collected with bandit algorithms. Alternative methods have recently been developed for simple models (e.g., comparison of means). Yet there is a lack of general methods for conducting statistical inference using more complex models on data collected with (contextual) bandit algorithms; for example, current methods cannot be used for valid inference on parameters in a logistic regression model for a binary reward. In this work, we develop theory justifying the use of M-estimators-which includes estimators based on empirical risk minimization as well as maximum likelihood-on data collected with adaptive algorithms, including (contextual) bandit algorithms. Specifically, we show that M-estimators, modified with particular adaptive weights, can be used to construct asymptotically valid confidence regions for a variety of inferential targets.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper14
- Active Statistical InferenceTijana Zrnic, Emmanuel J. CandèsICML 2024 · 被引用 34 次
- On Instance-Dependent Bounds for Offline Reinforcement Learning with Linear Function ApproximationThanh Nguyen-Tang, Ming Yin, Sunil Gupta, Svetha Venkatesh 等AAAI 2023 · 被引用 24 次
- The Adaptive Doubly Robust Estimator and a Paradox Concerning Logging PolicyMasahiro Kato, Kenichiro McAlinn, Shota YasuiNeurIPS 2021 · 被引用 23 次
- CLIP-OGD: An Experimental Design for Adaptive Neyman Allocation in Sequential ExperimentsJessica Dai, Paula Gradu, Christopher HarshawNeurIPS 2023 · 被引用 21 次
- Constrained Stochastic Nonconvex Optimization with State-dependent Markov DataAbhishek Roy, Krishnakumar Balasubramanian, Saeed GhadimiNeurIPS 2022 · 被引用 14 次
它引用的顶会 Paper3
- Personalized HeartSteps: A Reinforcement Learning Algorithm for Optimizing Physical ActivityPeng Liao, Kristjan H. Greenewald, Predrag V. Klasnja, Susan A. MurphyUbiComp 2020 · 被引用 163 次
- Inference for Batched BanditsKelly W. Zhang, Lucas Janson, Susan A. MurphyNeurIPS 2020 · 被引用 115 次
- Off-Policy Evaluation via Adaptive Weighting with Data from Contextual BanditsRuohan Zhan, Vitor Hadad, David A. Hirshberg, Susan AtheyKDD 2021 · 被引用 22 次
相关 Paper
- Post-Contextual-Bandit InferenceAurélien Bibaut, Maria Dimakopoulou, Nathan Kallus, Antoine Chambaz 等NeurIPS 2021 · 被引用 58 次
- Replicable BanditsHossein Esfandiari, Alkis Kalavasis, Amin Karbasi, Andreas Krause 等ICLR 2023 · 被引用 1 次
- Robust Tests in Online Decision-MakingGi-Soo Kim, Jane P. Kim, Hyun-Joon YangAAAI 2022
- RoME: A Robust Mixed-Effects Bandit Algorithm for Optimizing Mobile Health InterventionsEaston K. Huch, Jieru Shi, Madeline R. Abbott, Jessica R. Golbus 等NeurIPS 2024 · 被引用 6 次
- Likelihood Ratio Confidence Sets for Sequential Decision MakingNicolas Emmenegger, Mojmir Mutny, Andreas KrauseNeurIPS 2023 · 被引用 13 次
