INFaaS: Automated Model-less Inference Serving
Francisco Romero, Qian Li, Neeraja J. Yadwadkar, Christos Kozyrakis
Abstract
Despite existing work in machine learning inference serving, ease-of-use and cost efficiency remain challenges at large scales. Developers must manually search through thousands of model-variants -versions of already-trained models that differ in hardware, resource footprints, latencies, costs, and accuracies -to meet the diverse application requirements. Since requirements, query load, and applications themselves evolve over time, these decisions need to be made dynamically for each inference query to avoid excessive costs through naive autoscaling. To avoid navigating through the large and complex trade-off space of model-variants, developers often fix a variant across queries, and replicate it when load increases. However, given the diversity across variants and hardware platforms in the cloud, a lack of understanding of the trade-off space can incur significant costs to developers.
This paper introduces INFaaS, an automated model-less system for distributed inference serving, where developers simply specify the performance and accuracy requirements for their applications without needing to specify a specific model-variant for each query. INFaaS generates model-variants from already trained models, and efficiently navigates the large trade-off space of model-variants on behalf of developers to meet application-specific objectives: (a) for each query, it selects a model, hardware architecture, and model optimizations, (b) it combines VM-level horizontal autoscaling with model-level autoscaling, where multiple, different modelvariants are used to serve queries within each machine. By leveraging diverse variants and sharing hardware resources across models, INFaaS achieves 1.3× higher throughput, violates latency objectives 1.6× less often, and saves up to 21.6× in cost (8.5× on average) compared to state-of-the-art inference serving systems on AWS EC2.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ededca78-6686-4274-a548-0ef53e30d79fCited by top-tier papers59
- Splitwise: Efficient Generative LLM Inference Using Phase SplittingPratyush Patel, Esha Choukse, Chaojie Zhang, Aashaka Shah et al.ISCA 2024 · 282 citations
- AlpaServe: Statistical Multiplexing with Model Parallelism for Deep Learning ServingZhuohan Li, Lianmin Zheng, Yinmin Zhong, Vincent Liu et al.OSDI 2023 · 211 citations
- Serving Heterogeneous Machine Learning Models on Multi-GPU Servers with Spatio-Temporal SharingSeungbeom Choi, Sunho Lee, Yeonjae Kim, Jongse Park et al.USENIX ATC 2022 · 200 citations
- SHEPHERD: Serving DNNs in the WildHong Zhang, Yupeng Tang, Anurag Khandelwal, Ion StoicaNSDI 2023 · 161 citations
- Microsecond-scale Preemption for Concurrent GPU-accelerated DNN InferencesMingcong Han, Hanze Zhang, Rong Chen, Haibo ChenOSDI 2022 · 153 citations
Builds on6
- Serverless in the Wild: Characterizing and Optimizing the Serverless Workload at a Large Cloud ProviderMohammad Shahrad, Rodrigo Fonseca, Iñigo Goiri, Gohar Irfan Chaudhry et al.USENIX ATC 2020 · 946 citations
- MLPerf Inference BenchmarkVijay Janapa Reddi, Christine Cheng, David Kanter, Peter Mattson et al.ISCA 2020 · 517 citations
- Serving DNNs like Clockwork: Performance Predictability from the Bottom UpArpan Gujarati, Reza Karimi, Safya Alzayat, Wei Hao et al.OSDI 2020 · 392 citations
- DeepRecSys: A System for Optimizing End-To-End At-Scale Neural Recommendation InferenceUdit Gupta, Samuel Hsia, Vikram Saraph, Xiaodong Wang et al.ISCA 2020 · 149 citations
- Shredder: Learning Noise Distributions to Protect Inference PrivacyFatemehsadat Mireshghallah, Mohammadkazem Taram, Prakash Ramrakhyani, Ali Jalali et al.ASPLOS 2020 · 80 citations
Related papers
- Proteus: A High-Throughput Inference-Serving System with Accuracy ScalingSohaib Ahmad, Hui Guan, Brian D. Friedman, Thomas Williams et al.ASPLOS 2024 · 31 citations
- SMIless: Serving DAG-based Inference with Dynamic Invocations under Serverless ComputingChengzhi Lu, Huanle Xu, Yudan Li, Wenyan Chen et al.SC 2024 · 8 citations
- MOSEL: Inference Serving Using Dynamic Modality SelectionBodun Hu, Le Xu, Jeongyoon Moon, Neeraja J. Yadwadkar et al.EMNLP 2024 · 2 citations
- Automated End-to-End Model Serving with Cooperative Compilation and SchedulingYikang Zhang, Junlong Chen, Wei Wang, Jia Liu et al.EuroSys 2026
- HexGen-3: A Fully Disaggregated LLM Serving Framework with Fine-Grained Heterogeneous Resource AutoscalingYouhe Jiang, Wenshuang Li, You Peng, Jintao Zhang et al.ICML 2026
