LegoDNN: block-grained scaling of deep neural networks for mobile vision
Rui Han, Qinglong Zhang, Chi Harold Liu, Guoren Wang, Jian Tang, Lydia Y. Chen
Abstract
Deep neural networks (DNNs) have become ubiquitous techniques in mobile and embedded systems for applications such as image/object recognition and classification. The trend of executing multiple DNNs simultaneously exacerbate the existing limitations of meeting stringent latency/accuracy requirements on resource constrained mobile devices. The prior art sheds light on exploring the accuracy-resource tradeoff by scaling the model sizes in accordance to resource dynamics. However, such model scaling approaches face to imminent challenges: (i) large space exploration of model sizes, and (ii) prohibitively long training time for different model combinations. In this paper, we present LegoDNN, a lightweight, block-grained scaling solution for running multi-DNN workloads in mobile vision systems. LegoDNN guarantees short model training times by only extracting and training a small number of common blocks (e.g. 5 in VGG and 8 in ResNet) in a DNN. At run-time, LegoDNN optimally combines the descendant models of these blocks to maximize accuracy under specific resources and latency constraints, while reducing switching overhead via smart block-level scaling of the DNN. We implement LegoDNN in TensorFlow Lite and extensively evaluate it against state-of-the-art techniques (FLOP scaling, knowledge distillation and model compression) using a set of 12 popular DNN models. Evaluation results show that LegoDNN provides 1,296x to 279,936x more options in model sizes without increasing training time, thus achieving as much as 31.74% improvement in inference accuracy and 71.07% reduction in scaling energy consumptions.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2d149c36-3007-4a35-b0b3-be0f4658e8b5Cited by top-tier papers12
- AdaptiveNet: Post-deployment Neural Architecture Adaptation for Diverse Edge EnvironmentsHao Wen, Yuanchun Li, Zunshuai Zhang, Shiqi Jiang et al.MobiCom 2023 · 55 citations
- A Workload-Aware DVFS Robust to Concurrent Tasks for Mobile DevicesChengdong Lin, Kun Wang, Zhenjiang Li, Yu PuMobiCom 2023 · 52 citations
- Mandheling: mixed-precision on-device DNN training with DSP offloadingDaliang Xu, Mengwei Xu, Qipeng Wang, Shangguang Wang et al.MobiCom 2022 · 43 citations
- MergeSFL: Split Federated Learning with Feature Merging and Batch Size RegulationYunming Liao, Yang Xu, Hongli Xu, Lun Wang et al.ICDE 2024 · 41 citations
- FlexNN: Efficient and Adaptive DNN Inference on Memory-Constrained Edge DevicesXiangyu Li, Yuanchun Li, Yuanzhe Li, Ting Cao et al.MobiCom 2024 · 40 citations
Builds on5
- Contrastive Representation DistillationYonglong Tian, Dilip Krishnan, Phillip IsolaICLR 2020 · 1,305 citations
- Provable Filter Pruning for Efficient Neural NetworksLucas Liebenwein, Cenk Baykal, Harry Lang, Dan Feldman et al.ICLR 2020 · 161 citations
- And the Bit Goes Down: Revisiting the Quantization of Neural NetworksPierre Stock, Armand Joulin, Rémi Gribonval, Benjamin Graham et al.ICLR 2020 · 157 citations
- NeuOS: A Latency-Predictable Multi-Dimensional Optimization Framework for DNN-driven Autonomous SystemsSoroush Bateni, Cong LiuUSENIX ATC 2020 · 49 citations
- HRank: Filter Pruning Using High-Rank Feature MapMingbao Lin, Rongrong Ji, Yan Wang, Yichen Zhang et al.CVPR 2020
Related papers
- Resource-aware Deployment of Dynamic DNNs over Multi-tiered Interconnected SystemsChetna Singhal, Yashuo Wu, Francesco Malandrino, Marco Levorato et al.INFOCOM 2024 · 15 citations
- FlashMem: Supporting Modern DNN Workloads on Mobile with GPU Memory Hierarchy OptimizationsZhihao Shu, Md. Musfiqur Rahman Sanim, Hangyu Zheng, Kunxiong Zhu et al.ASPLOS 2026
- Convolutional Neural Network Compression through Generalized Kronecker Product DecompositionMarawan Gamal Abdel Hameed, Marzieh S. Tahaei, Ali Mosleh, Vahid Partovi NiaAAAI 2022 · 33 citations
- Non-uniform DNN Structured Subnets Sampling for Dynamic InferenceLi Yang, Zhezhi He, Yu Cao, Deliang FanDAC 2020 · 12 citations
- Distilling Optimal Neural Networks: Rapid Search in Diverse SpacesBert Moons, Parham Noorzad, Andrii Skliar, Giovanni Mariani et al.ICCV 2021 · 46 citations
