Simulation-Based Inference for Adaptive Experiments
Brian Cho, Aurélien Bibaut, Nathan Kallus
Abstract
Multi-arm bandit experimental designs are increasingly being adopted over standard randomized trials due to their potential to improve outcomes for study participants, enable faster identification of the best-performing options, and/or enhance the precision of estimating key parameters. Current approaches for inference after adaptive sampling either rely on asymptotic normality under restricted experiment designs or underpowered martingale concentration inequalities that lead to weak power in practice. To bypass these limitations, we propose a simulation-based approach for conducting hypothesis tests and constructing confidence intervals for arm specific means and their differences. Our simulation-based approach uses positively biased nuisances to generate additional trajectories of the experiment, which we call simulation with optimism. Using these simulations, we characterize the distribution potentially non-normal sample mean test statistic to conduct inference. We provide guarantees for (i) asymptotic type I error control, (ii) convergence of our confidence intervals, and (iii) asymptotic strong consistency of our estimator over a wide variety of common bandit designs. Our empirical results show that our approach achieves the desired coverage while reducing confidence interval widths by up to 50%, with drastic improvements for arms not targeted by the design.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ca1ce0b0-034a-4e2c-bde1-37c515c9c54fCited by top-tier papers1
Ask how each one uses itBuilds on5
- Statistical Inference with M-Estimators on Adaptively Collected DataKelly W. Zhang, Lucas Janson, Susan A. MurphyNeurIPS 2021 · 66 citations
- Post-Contextual-Bandit InferenceAurélien Bibaut, Maria Dimakopoulou, Nathan Kallus, Antoine Chambaz et al.NeurIPS 2021 · 58 citations
- Batched Thompson SamplingCem Kalkanli, Ayfer ÖzgürNeurIPS 2021 · 29 citations
- Off-Policy Evaluation via Adaptive Weighting with Data from Contextual BanditsRuohan Zhan, Vitor Hadad, David A. Hirshberg, Susan AtheyKDD 2021 · 22 citations
- Peeking with PEAK: Sequential, Nonparametric Composite Hypothesis Tests for Means of Multiple Data StreamsBrian Cho, Kyra Gan, Nathan KallusICML 2024 · 14 citations
Related papers
- Inference for Batched BanditsKelly W. Zhang, Lucas Janson, Susan A. MurphyNeurIPS 2020 · 115 citations
- On Conditional Versus Marginal Bias in Multi-Armed BanditsJaehyeok Shin, Aaditya Ramdas, Alessandro RinaldoICML 2020 · 13 citations
- Optimistic Algorithms for Adaptive Estimation of the Average Treatment EffectOjash Neopane, Aaditya Ramdas, Aarti SinghICML 2025
- Optimal Estimation of the Best Mean in Multi-Armed BanditsTakayuki Osogami, Junya Honda, Junpei KomiyamaNeurIPS 2025
- On the Robustness of Bandit Multiple TestingZhengyu Zhou, Weiwei LiuAAAI 2026
