Exploiting Similarities in A/B Testing with Off-Policy Estimation
Otmane Sakhi, Alexandre Gilotte, David Rohde
Abstract
We study A/B testing, the standard protocol for measuring the performance gain of a new decision system relative to a baseline. Traditional A/B testing treats both systems as black boxes, ignoring potential similarities between them. In practice, however, new and baseline systems are rarely radically different and often share significant structure, which can be captured by their propensities to make similar decisions. We show that in such cases, the commonly used difference-in-means estimator, though unbiased, is statistically suboptimal. Leveraging off-policy estimation, we introduce a family of A/B testing estimators that exploit the propensities of the tested systems to achieve improved concentration properties. This family is flexible enough to be tailored to practical decision-making. The resulting estimators are simple, robust to propensities misspecification, substantially more accurate when the tested systems exhibit similarities, and gracefully fall back to the difference-in-means estimator when such similarities are absent. Our theoretical analysis and empirical studies confirm their efficiency and practicality.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9d106055-4e02-4dee-9a4f-2874daf9ac8eCited by top-tier papers1
Ask how each one uses itBuilds on13
- Doubly robust off-policy evaluation with shrinkageYi Su, Maria Dimakopoulou, Akshay Krishnamurthy, Miroslav DudíkICML 2020 · 128 citations
- Off-Policy Evaluation for Large Action Spaces via EmbeddingsYuta Saito, Thorsten JoachimsICML 2022 · 62 citations
- Subgaussian and Differentiable Importance Sampling for Off-Policy Evaluation and LearningAlberto Maria Metelli, Alessio Russo, Marcello RestelliNeurIPS 2021 · 55 citations
- Markovian Interference in ExperimentsVivek F. Farias, Andrew A. Li, Tianyi Peng, Andrew ZhengNeurIPS 2022 · 52 citations
- Understanding the Curse of Horizon in Off-Policy Evaluation via Conditional Importance SamplingYao Liu, Pierre-Luc Bacon, Emma BrunskillICML 2020 · 49 citations
Related papers
- Unraveling the Interplay between Carryover Effects and Reward Autocorrelations in Switchback ExperimentsQianglin Wen, Chengchun Shi, Ying Yang, Niansheng Tang et al.ICML 2025
- Optimal Treatment Allocation for Efficient Policy Evaluation in Sequential Decision MakingTing Li, Chengchun Shi, Jianing Wang, Fan Zhou et al.NeurIPS 2023 · 21 citations
- Pessimistic Data Integration for Policy EvaluationXiangkun Wu, Ting Li, Gholamali Aminian, Armin Behnamnia et al.NeurIPS 2025 · 2 citations
- Strategic A/B testing via Maximum Probability-driven Two-armed BanditYu Zhang, Shanshan Zhao, Bokui Wan, Jinjuan Wang et al.ICML 2025
- Doubly-Robust Estimation of Counterfactual Policy Mean EmbeddingsHoussam Zenati, Bariscan Bozkurt, Arthur GrettonNeurIPS 2025 · 3 citations
