SuperServe: Fine-Grained Inference Serving for Unpredictable Workloads
Alind Khare, Dhruv Garg, Sukrit Kalra, Snigdha Grandhi, Ion Stoica, Alexey Tumanov
摘要
The increasing deployment of ML models on the critical path of production applications in both datacenter and the edge requires ML inference serving systems to serve these models under unpredictable and bursty request arrival rates. Serving models under such conditions requires these systems to strike a careful balance between the latency and accuracy requirements of the application and the overall efficiency of utilization of scarce resources. State-of-the-art systems resolve this tension by either choosing a static point in the latency-accuracy tradeoff space to serve all requests or load specific models on the critical path of request serving.
In this work, we instead resolve this tension by simultaneously serving the entire-range of models spanning the latencyaccuracy tradeoff space. Our novel mechanism, SubNetAct, achieves this by carefully inserting specialized operators in weight-shared SuperNetworks. These operators enable Sub-NetAct to dynamically route requests through the network to meet a latency and accuracy target. SubNetAct requires upto 2.6× lower memory to serve a vastly-higher number of models than prior state-of-the-art. In addition, SubNetAct's near-instantaneous actuation of models unlocks the design space of fine-grained, reactive scheduling policies. We explore the design of one such extremely effective policy, SlackFit and instantiate both SubNetAct and SlackFit in a real system, SuperServe. SuperServe achieves 4.67% higher accuracy for the same SLO attainment and 2.85× higher SLO attainment for the same accuracy on a trace derived from the real-world Microsoft Azure Functions workload and yields the best tradeoffs on a wide range of extremely-bursty synthetic traces automatically.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- PowerInfer: Fast Large Language Model Serving with a Consumer-grade GPUYixin Song, Zeyu Mi, Haotong Xie, Haibo ChenSOSP 2024 · 被引用 86 次
- TetriServe: Efficiently Serving Mixed DiT WorkloadsRunyu Lu, Shiqi He, Wenxuan Tan, Shenggui Li 等ASPLOS 2026 · 被引用 1 次
- VillainNet: Targeted Poisoning Attacks Against SuperNets Along the Accuracy-Latency Pareto FrontierDavid Oygenblik, Abhinav Vemulapalli, Animesh Agrawal, Debopam Sanyal 等CCS 2025
- CloserToMe: A Unified Framework for Accurate and Transferable Latency Prediction Across Heterogeneous DevicesCheng Tang, Guochong Sui, Wenqi Lou, Zihan Wang 等AAAI 2026
- Ken: An Execution Engine for Unstructured Database SystemsFerdinand Kossmann, Ziniu Wu, Alex Turk, Nesime Tatbul 等VLDB 2026
它引用的顶会 Paper16
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Searching for MobileNetV3Andrew Howard, Ruoming Pang, Hartwig Adam, Quoc V. Le 等ICCV 2019 · 被引用 9,163 次
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa 等ICML 2021 · 被引用 8,974 次
- A ConvNet for the 2020sZhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer 等CVPR 2022 · 被引用 6,782 次
- Once-for-All: Train One Network and Specialize it for Efficient DeploymentHan Cai, Chuang Gan, Tianzhe Wang, Zhekai Zhang 等ICLR 2020 · 被引用 1,522 次
相关 Paper
- Proteus: A High-Throughput Inference-Serving System with Accuracy ScalingSohaib Ahmad, Hui Guan, Brian D. Friedman, Thomas Williams 等ASPLOS 2024 · 被引用 31 次
- Jellyfish: Timely Inference Serving for Dynamic Edge NetworksVinod Nigade, Pablo Bauszat, Henri E. Bal, Lin WangRTSS 2022 · 被引用 40 次
- QoServe: Breaking the Silos of LLM Inference ServingKanishk Goel, Jayashree Mohan, Nipun Kwatra, Ravi Shreyas Anupindi 等ASPLOS 2026 · 被引用 3 次
- AugServe: Adaptive Request Scheduling for Augmented Large Language Model Inference ServingYing Wang, Zhen Jin, Zhenqian Chen, Jiexiong Xu 等ICML 2026 · 被引用 4 次
- Model Selection for Latency-Critical Inference ServingDaniel Mendoza, Francisco Romero, Caroline TrippelEuroSys 2024 · 被引用 16 次
