Model Selection for Production System via Automated Online Experiments
Zhenwen Dai, Praveen Chandar, Ghazal Fazelnia, Benjamin A. Carterette, Mounia Lalmas
Abstract
A challenge that machine learning practitioners in the industry face is the task of selecting the best model to deploy in production. As a model is often an intermediate component of a production system, online controlled experiments such as A/B tests yield the most reliable estimation of the effectiveness of the whole system, but can only compare two or a few models due to budget constraints. We propose an automated online experimentation mechanism that can efficiently perform model selection from a large pool of models with a small number of online experiments. We derive the probability distribution of the metric of interest that contains the model uncertainty from our Bayesian surrogate model trained using historical logs. Our method efficiently identifies the best model by sequentially selecting and deploying a list of models from the candidate set that balance exploration-exploitation. Using simulations based on real data, we demonstrate the effectiveness of our method on two different tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e28a799c-1ff2-4379-b4c8-4a76d5c6aa26Cited by top-tier papers1
Ask how each one uses itRelated papers
- LABO: LLM-Accelerated Bayesian Optimization through Broad Exploration and Selective ExperimentationZhuo Chen, Xinzhe Yuan, Jianshu Zhang, Jinzong Dong et al.ICML 2026 · 2 citations
- Optimal Treatment Allocation for Efficient Policy Evaluation in Sequential Decision MakingTing Li, Chengchun Shi, Jianing Wang, Fan Zhou et al.NeurIPS 2023 · 21 citations
- Effect Size Estimation for Duration Recommendation in Online Experiments: Leveraging Hierarchical Models and Objective Utility ApproachesYu Liu, Runzhe Wan, James McQueen, Doug Hains et al.AAAI 2024 · 1 citation
- inRAN: Interpretable Online Bayesian Learning for Network Automation in Open Radio Access NetworksMing Zhao, Yuru Zhang, Qiang Liu, Ahan Kak et al.INFOCOM 2026 · 1 citation
- Balancing Risk and Reward: A Batched-Bandit Strategy for Automated Phased ReleaseYufan Li, Jialiang Mao, Iavor BojinovNeurIPS 2023 · 1 citation
