Short-lived High-volume Bandits
Su Jia, Nishant Oli, Ian Anderson, Paul Duff, Andrew A. Li, R. Ravi
Abstract
We study how to efficiently perform A/B/n testing for a high-volume of short-lived treatments. We formulate the problem as a multiple-play bandits model. In each round a set of k actions arrive. Each action is available for w rounds and has an unknown reward rate. In each round, the learner selects a multiset of n actions and immediately observes the realized rewards. We aim to minimize the average loss under a random input model where the instance is randomly drawn from a known prior distribution D. We show that if k = O(n ρ ) for some ρ > 0, our policy achieves Õ(n -minρ, 1 2 (1+ 1 w ) -1 ) average loss on a sufficiently large class of prior distributions. We also complement this result by showing that every policy suffers Ω(n -minρ, 1 2 ) average loss on the same class of distributions. We further validate the effectiveness of our policy through a large-scale field experiment on Glance, a content card-serving platform.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- Online Experimental Design With Estimation-Regret Trade-off Under Network InterferenceZhiheng Zhang, Zichen WangNeurIPS 2025 · 12 citations
- Design-Based Bandits Under Network Interference: Trade-Off Between Regret and Statistical InferenceZichen Wang, Haoyang Hong, Chuanhao Li, Haoxuan Li et al.NeurIPS 2025 · 3 citations
- Multi-Armed Bandits with Interference: Bridging Causal Inference and Adversarial BanditsSu Jia, Peter I. Frazier, Nathan KallusICML 2025
Builds on1
Related papers
- Asymptotically Optimal and Computationally Efficient Average Treatment Effect Estimation in A/B testingVikas Deep, Achal Bassamboo, Sandeep K. JunejaICML 2024 · 1 citation
- A/B/n Testing with Control in the Presence of SubpopulationsYoan Russac, Christina Katsimerou, Dennis Bohle, Olivier Cappé et al.NeurIPS 2021 · 34 citations
- Empirical Bayes Selection for Value MaximizationDominic Coey, Kenneth HungKDD 2025 · 1 citation
- Choice BanditsArpit Agarwal, Nicholas Johnson, Shivani AgarwalNeurIPS 2020 · 19 citations
- Interference, Bias, and Variance in Two-Sided Marketplace Experimentation: Guidance for PlatformsHannah Li, Geng Zhao, Ramesh Johari, Gabriel Y. WeintraubWWW 2022 · 46 citations
