Biathlon: Harnessing Model Resilience for Accelerating ML Inference Pipelines
Chaokun Chang, Eric Lo, Chunxiao Ye
Abstract
Machine learning inference pipelines commonly encountered in data science and industries often require real-time responsiveness due to their user-facing nature. However, meeting this requirement becomes particularly challenging when certain input features require aggregating a large volume of data online. Recent literature on interpretable machine learning reveals that most machine learning models exhibit a notable degree of resilience to variations in input. This suggests that machine learning models can effectively accommodate approximate input features with minimal discernible impact on accuracy. In this paper, we introduce Biathlon, a novel ML serving system that leverages the inherent resilience of models and determines the optimal degree of approximation for each aggregation feature. This approach enables maximum speedup while ensuring a guaranteed bound on accuracy loss. We evaluate Biathlon on real pipelines from both industry applications and data science competitions, demonstrating its ability to meet real-time latency requirements by achieving 5.3× to 16.6× speedup with almost no accuracy loss.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext cad819c5-ff46-4df5-b787-daf4d8c752eaCited by top-tier papers2
- Apt-Serve: Adaptive Request Scheduling on Hybrid Cache for Scalable LLM Inference ServingShihong Gao, Xin Zhang, Yanyan Shen, Lei ChenSIGMOD 2025 · 7 citations
- Ken: An Execution Engine for Unstructured Database SystemsFerdinand Kossmann, Ziniu Wu, Alex Turk, Nesime Tatbul et al.VLDB 2026
Builds on9
- Feature Importance-aware Transferable Adversarial AttacksZhibo Wang, Hengchang Guo, Zhifei Zhang, Wenxin Liu et al.ICCV 2021 · 306 citations
- A Tensor Compiler for Unified Machine Learning Prediction ServingSupun Nakandala, Karla Saur, Gyeong-In Yu, Konstantinos Karanasos et al.OSDI 2020 · 60 citations
- Query Processing on Tensor Computation RuntimesDong He, Supun Chathuranga Nakandala, Dalitso Banda, Rathijit Sen et al.VLDB 2022 · 54 citations
- End-to-end Optimization of Machine Learning Prediction QueriesKwanghyun Park, Karla Saur, Dalitso Banda, Rathijit Sen et al.SIGMOD 2022 · 50 citations
- JoinBoost: Grow Trees Over Normalized Data Using Only SQLZezhou Huang, Rathijit Sen, Jiaxiang Liu, Eugene WuVLDB 2023 · 23 citations
Related papers
- Proteus: A High-Throughput Inference-Serving System with Accuracy ScalingSohaib Ahmad, Hui Guan, Brian D. Friedman, Thomas Williams et al.ASPLOS 2024 · 31 citations
- Loki: A System for Serving ML Inference Pipelines with Hardware and Accuracy ScalingSohaib Ahmad, Hui Guan, Ramesh K. SitaramanHPDC 2024 · 8 citations
- SHEPHERD: Serving DNNs in the WildHong Zhang, Yupeng Tang, Anurag Khandelwal, Ion StoicaNSDI 2023 · 161 citations
- Serving DNNs like Clockwork: Performance Predictability from the Bottom UpArpan Gujarati, Reza Karimi, Safya Alzayat, Wei Hao et al.OSDI 2020 · 392 citations
- JITServe: SLO-aware LLM Serving with Imprecise Request InformationWei Zhang, Zhiyu Wu, Yi Mu, Rui Ning et al.NSDI 2026 · 29 citations
