Sommelier: Curating DNN Models for the Masses
Peizhen Guo, Bo Hu, Wenjun Hu
Abstract
Deep learning model repositories are indispensable in machine learning ecosystems today to facilitate model reuse. However, existing model repositories provide a bare-bone interface for model retrieval. The onus is on the user to profile and select from potentially hundreds of choices, barely relieving an average user of the expertise required to design the model in the first place. In this paper, we present Sommelier, an indexing and query system above typical DNN model repositories to interface directly with inference serving or other use cases. Given a desirable accuracy target and resource budget for an inference task category, Sommelier automatically searches through the repository for the most suitable model, without requiring manual profiling from the user. Motivated by manual iterative model search processes and typical model design strategies that generate model variants or models with common segments, Sommelier organizes DNN models based on their semantic correlation, defined as the probability of models producing the same results. This is further combined with a resource index based on relative resource consumption. Sommelier is implemented as a standalone query engine that can interface with an existing repository such as TF-Hub. A case study of 163 models in TF-Hub highlights the extent of model correlation across different model series, suggesting the best candidate model can easily evade manual profiling. Extensive evaluation shows that Sommelier returns the ideal model for over 95% of the queries; When interfaced with an inference server, Sommelier can reduce the 90th percentile tail latency of inference tasks by a factor of 6 via automatic model switching, far more than typical scale-out system optimizations.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 130e85b1-7b98-4481-a080-64423dee92b4Cited by top-tier papers5
- Gemel: Model Merging for Memory-Efficient, Real-Time Video Analytics at the EdgeArthi Padmanabhan, Neil Agarwal, Anand P. Iyer, Ganesh Ananthanarayanan et al.NSDI 2023 · 94 citations
- Tabi: An Efficient Multi-Level Inference System for Large Language ModelsYiding Wang, Kai Chen, Haisheng Tan, Kun GuoEuroSys 2023 · 64 citations
- Proteus: A High-Throughput Inference-Serving System with Accuracy ScalingSohaib Ahmad, Hui Guan, Brian D. Friedman, Thomas Williams et al.ASPLOS 2024 · 31 citations
- Loki: A System for Serving ML Inference Pipelines with Hardware and Accuracy ScalingSohaib Ahmad, Hui Guan, Ramesh K. SitaramanHPDC 2024 · 8 citations
- Alsatian: Optimizing Model Search for Deep Transfer LearningNils Strassenburg, Boris Glavic, Tilmann RablSIGMOD 2025 · 2 citations
Related papers
- SHiFT: An Efficient, Flexible Search Engine for Transfer LearningCédric Renggli, Xiaozhe Yao, Luka Kolar, Luka Rimanic et al.VLDB 2023 · 8 citations
- Which Model to Transfer? Finding the Needle in the Growing HaystackCédric Renggli, André Susano Pinto, Luka Rimanic, Joan Puigcerver et al.CVPR 2022 · 13 citations
- SilverTorch: A Unified Model-based System to Democratize Large-Scale Recommendation on GPUsBi Xue, Hong Wu, Lei Chen, Chao Yang et al.SIGIR 2026
- Searching the Deployable Convolution Neural Networks for GPUsLinnan Wang, Chenhan Yu, Satish Salian, Slawomir Kierat et al.CVPR 2022 · 8 citations
- An Empirical Study of Pre-Trained Model Reuse in the Hugging Face Deep Learning Model RegistryWenxin Jiang, Nicholas Synovic, Matt Hyatt, Taylor R. Schorlemmer et al.ICSE 2023 · 62 citations
