How Did the Model Change? Efficiently Assessing Machine Learning API Shifts
Lingjiao Chen, Matei Zaharia, James Zou
Abstract
ML prediction APIs from providers like Amazon and Google have made it simple to use ML in applications. A challenge for users is that such APIs continuously change over time as the providers update models, and changes can happen silently without users knowing. It is thus important to monitor when and how much the ML APIs' performance shifts. To provide detailed change assessment, we model ML API shifts as confusion matrix differences, and propose a principled algorithmic framework, MASA, to provably assess these shifts efficiently given a sample budget constraint. MASA employs an upper-confidence bound based approach to adaptively determine on which data point to query the ML API to estimate shifts. Empirically, we observe significant ML API shifts from 2020 to 2021 among 12 out of 36 applications using commercial APIs from Google, Microsoft, Amazon, and other providers. These real-world shifts include both improvements and reductions in accuracy. Extensive experiments show that MASA can estimate such API shifts more accurately than standard approaches given the same budget.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers5
- Estimating and Explaining Model Performance When Both Covariates and Labels ShiftLingjiao Chen, Matei Zaharia, James Y. ZouNeurIPS 2022 · 34 citations
- Ecosystem-level Analysis of Deployed Machine Learning Reveals Homogeneous OutcomesConnor Toups, Rishi Bommasani, Kathleen Creel, Sarah H. Bana et al.NeurIPS 2023 · 25 citations
- Efficient Online ML API Selection for Multi-Label Classification TasksLingjiao Chen, Matei Zaharia, James ZouICML 2022 · 22 citations
- Log Probability Tracking of LLM APIsTimothee Chauvin, Erwan Le Merrer, Francois Taiani, Gilles TredanICLR 2026 · 12 citations
- Token-Efficient Change Detection in LLM APIsTimothee Chauvin, Clément Lalanne, Erwan Le Merrer, Jean-Michel Loubes et al.ICML 2026 · 4 citations
Builds on3
- WILDS: A Benchmark of in-the-Wild Distribution ShiftsPang Wei Koh, Shiori Sagawa, Henrik Marklund, Sang Michael Xie et al.ICML 2021 · 1,773 citations
- From ImageNet to Image Classification: Contextualizing Progress on BenchmarksDimitris Tsipras, Shibani Santurkar, Logan Engstrom, Andrew Ilyas et al.ICML 2020 · 146 citations
- Beyond Accuracy: Behavioral Testing of NLP Models with CheckListMarco Túlio Ribeiro, Tongshuang Wu, Carlos Guestrin, Sameer SinghACL 2020 · 51 citations
Related papers
- FrugalML: How to use ML Prediction APIs more accurately and cheaplyLingjiao Chen, Matei Zaharia, James Y. ZouNeurIPS 2020 · 57 citations
- Are Machine Learning Cloud APIs Used Correctly?Chengcheng Wan, Shicheng Liu, Henry Hoffmann, Michael Maire et al.ICSE 2021 · 37 citations
- Model Equality Testing: Which Model is this API Serving?Irena Gao, Percy Liang, Carlos GuestrinICLR 2025
- ChameleonAPI: Automatic and Efficient Customization of Neural Networks for ML ApplicationsYuhan Liu, Chengcheng Wan, Kuntai Du, Henry Hoffmann et al.OSDI 2024 · 1 citation
- Run-Time Prevention of Software Integration Failures of Machine Learning APIsChengcheng Wan, Yuhan Liu, Kuntai Du, Henry Hoffmann et al.OOPSLA 2023 · 5 citations
