Model Slicing for Supporting Complex Analytics with Elastic Inference Cost and Resource Constraints
Shaofeng Cai, Gang Chen, Beng Chin Ooi, Jinyang Gao
摘要
Deep learning models have been used to support analytics beyond simple aggregation, where deeper and wider models have been shown to yield great results. These models consume a huge amount of memory and computational operations. However, most of the large-scale industrial applications are often computational budget constrained. In practice, the peak workload of inference service could be 10x higher than the average cases, with the presence of unpredictable extreme cases. Lots of computational resources could be wasted during off-peak hours and the system may crash when the workload exceeds system capacity. How to support deep learning services with dynamic workload cost-efficiently remains a challenging problem. In this paper, we address the challenge with a general and novel training scheme called model slicing , which enables deep learning models to provide predictions within the prescribed computational resource budget dynamically. Model slicing could be viewed as an elastic computation solution without requiring more computational resources. Succinctly, each layer in the model is divided into groups of contiguous block of basic components (i.e. neurons in dense layers and channels in convolutional layers), and then partially ordered relation is introduced to these groups by enforcing that groups participated in each forward pass always starts from the first group to the dynamically-determined rightmost group. Trained by dynamically indexing the rightmost group with a single parameter slice rate , the network is engendered to build up group-wise and residual representation. Then during inference, a sub-model with fewer groups can be readily deployed for efficiency whose computation is roughly quadratic to the width controlled by the slice rate. Extensive experiments show that models trained with model slicing can effectively support on-demand workload with elastic inference cost.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- Dynamic slicing for deep neural networksZiqi Zhang, Yuanchun Li, Yao Guo, Xiangqun Chen 等FSE 2020 · 被引用 34 次
- Optimizing Machine Learning Inference Queries with Correlative Proxy ModelsZhihui Yang, Zuozhi Wang, Yicong Huang, Yao Lu 等VLDB 2022 · 被引用 33 次
- MLCask: Efficient Management of Component Evolution in Collaborative Data Analytics PipelinesZhaojing Luo, Sai Ho Yeung, Meihui Zhang, Kaiping Zheng 等ICDE 2021 · 被引用 31 次
- Database Native Model Selection: Harnessing Deep Neural Networks in Database SystemsNaili Xing, Shaofeng Cai, Gang Chen, Zhaojing Luo 等VLDB 2024 · 被引用 14 次
- Powering In-Database Dynamic Model Slicing for Structured Data AnalyticsLingze Zeng, Naili Xing, Shaofeng Cai, Gang Chen 等VLDB 2024 · 被引用 7 次
相关 Paper
- Compile-Time QoS Scheme for Deep Learning InferencesSungin Hong, Hyunjun Kim, Hwansoo HanSC 2025 · 被引用 1 次
- Navigating Scaling Laws: Compute Optimality in Adaptive Model TrainingSotiris Anagnostidis, Gregor Bachmann, Imanol Schlag, Thomas HofmannICML 2024 · 被引用 2 次
- Microscope: mobile service traffic decomposition for network slicing as a serviceChaoyun Zhang, Marco Fiore, Cezary Ziemlicki, Paul PatrasMobiCom 2020 · 被引用 37 次
- USHER: Holistic Interference Avoidance for Resource Optimized ML InferenceSudipta Saha Shubha, Haiying Shen, Anand P. IyerOSDI 2024 · 被引用 35 次
- Geryon: Accelerating Distributed CNN Training by Network-Level Flow SchedulingShuai Wang, Dan Li, Jinkun GengINFOCOM 2020 · 被引用 59 次
