Proteus: A High-Throughput Inference-Serving System with Accuracy Scaling
Sohaib Ahmad, Hui Guan, Brian D. Friedman, Thomas Williams, Ramesh K. Sitaraman, Thomas Y. C. Woo
Abstract
Existing machine learning inference-serving systems largely rely on hardware scaling by adding more devices or using more powerful accelerators to handle increasing query demands. However, hardware scaling might not be feasible for fixed-size edge clusters or private clouds due to their limited hardware resources. A viable alternate solution is accuracy scaling, which adapts the accuracy of ML models instead of hardware resources to handle varying query demands. This work studies the design of a high-throughput inference-serving system with accuracy scaling that can meet throughput requirements while maximizing accuracy. To achieve the goal, this work proposes to identify the right amount of accuracy scaling by jointly optimizing three sub-problems: how to select model variants, how to place them on heterogeneous devices, and how to assign query workloads to each device. It also proposes a new adaptive batching algorithm to handle variations in query arrival times and minimize SLO violations. Based on the proposed techniques, we build an inference-serving system called Proteus and empirically evaluate it on real-world and synthetic traces. We show that Proteus reduces accuracy drop by up to 3× and latency timeouts by 2--10× with respect to baseline schemes, while meeting throughput requirements.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers7
- Cache-Craft: Managing Chunk-Caches for Efficient Retrieval-Augmented GenerationShubham Agarwal, Sai Sundaresan, Subrata Mitra, Debabrata Mahapatra et al.SIGMOD 2025 · 20 citations
- Katz: Efficient Workflow Serving for Diffusion Models with Many AdaptersSuyi Li, Lingyun Yang, Xiaoxiao Jiang, Hanfeng Lu et al.USENIX ATC 2025 · 14 citations
- Loki: A System for Serving ML Inference Pipelines with Hardware and Accuracy ScalingSohaib Ahmad, Hui Guan, Ramesh K. SitaramanHPDC 2024 · 8 citations
- MaverIQ: Fingerprint-Guided Extrapolation and Fragmentation-Aware Layering for Intent-Based LLM ServingDimitrios Liakopoulos, Prasoon Sinha, Tianrui Hu, Myungjin Lee et al.SC 2025 · 2 citations
- MOSEL: Inference Serving Using Dynamic Modality SelectionBodun Hu, Le Xu, Jeongyoon Moon, Neeraja J. Yadwadkar et al.EMNLP 2024 · 2 citations
Builds on10
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel et al.ICLR 2020 · 7,418 citations
- Serving DNNs like Clockwork: Performance Predictability from the Bottom UpArpan Gujarati, Reza Karimi, Safya Alzayat, Wei Hao et al.OSDI 2020 · 392 citations
- INFaaS: Automated Model-less Inference ServingFrancisco Romero, Qian Li, Neeraja J. Yadwadkar, Christos KozyrakisUSENIX ATC 2021 · 325 citations
- INFless: a native serverless system for low-latency, high-throughput inferenceYanan Yang, Laiping Zhao, Yiming Li, Huanyu Zhang et al.ASPLOS 2022 · 145 citations
- RecSSD: near data processing for solid state drive based recommendation inferenceMark Wilkening, Udit Gupta, Samuel Hsia, Caroline Trippel et al.ASPLOS 2021 · 100 citations
Related papers
- SHEPHERD: Serving DNNs in the WildHong Zhang, Yupeng Tang, Anurag Khandelwal, Ion StoicaNSDI 2023 · 161 citations
- SuperServe: Fine-Grained Inference Serving for Unpredictable WorkloadsAlind Khare, Dhruv Garg, Sukrit Kalra, Snigdha Grandhi et al.NSDI 2025
- InfScaler: Enabling Efficient ML Inference Serving on Multi-Accelerator Edge Devices via Asymmetric Auto-ScalingBorui Li, Tiange Xia, Shuai Wang, Shuai WangDAC 2025 · 2 citations
- Serving Heterogeneous Machine Learning Models on Multi-GPU Servers with Spatio-Temporal SharingSeungbeom Choi, Sunho Lee, Yeonjae Kim, Jongse Park et al.USENIX ATC 2022 · 200 citations
- Optimizing Inference Serving on Serverless PlatformsAhsan Ali, Riccardo Pinciroli, Feng Yan, Evgenia SmirniVLDB 2022 · 76 citations
