NeuLens: spatial-based dynamic acceleration of convolutional neural networks on edge
Xueyu Hou, Yongjie Guan, Tao Han
Abstract
Convolutional neural networks (CNNs) play an important role in today's mobile and edge computing systems for vision-based tasks like object classification and detection. However, state-of-the-art methods on CNN acceleration are trapped in either limited practical latency speed-up on general computing platforms or latency speed-up with severe accuracy loss. In this paper, we propose a spatial-based dynamic CNN acceleration framework, NeuLens, for mobile and edge platforms. Specially, we design a novel dynamic inference mechanism, assemble region-aware convolution (ARAC) supernet, that peels off redundant operations inside CNN models as many as possible based on spatial redundancy and channel slicing. In ARAC supernet, the CNN inference flow is split into multiple independent micro-flows, and the computational cost of each can be autonomously adjusted based on its tiled-input content and application requirements. These micro-flows can be loaded into hardware like GPUs as single models. Consequently, its operation reduction can be well translated into latency speed-up and is compatible with hardware-level accelerations. Moreover, the inference accuracy can be well preserved by identifying critical regions on images and processing them in the original resolution with large micro-flow. Based on our evaluation, NeuLens outperforms baseline methods by up to 58% latency reduction with the same accuracy and by up to 67.9% accuracy improvement under the same latency/memory constraints.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 44dc8422-022b-415c-a33b-39a7f92602b3Cited by top-tier papers3
- MetaStream: Live Volumetric Content Capture, Creation, Delivery, and Rendering in Real TimeYongjie Guan, Xueyu Hou, Nan Wu, Bo Han et al.MobiCom 2023 · 47 citations
- SwapMoE: Serving Off-the-shelf MoE-based Large Language Models with Tunable Memory BudgetRui Kong, Yuanchun Li, Qingtian Feng, Weijun Wang et al.ACL 2024 · 12 citations
- Panopticus: Omnidirectional 3D Object Detection on Resource-constrained Edge DevicesJeho Lee, Chanyoung Jung, Jiwon Kim, Hojung ChaMobiCom 2024 · 4 citations
Related papers
- Latency-aware Spatial-wise Dynamic NetworksYizeng Han, Zhihang Yuan, Yifan Pu, Chenhao Xue et al.NeurIPS 2022 · 30 citations
- Anticipating and eliminating redundant computations in accelerated sparse trainingJonathan S. Lew, Yunpeng Liu, Wenyi Gong, Negar Goli et al.ISCA 2022 · 10 citations
- Dynamic Resolution NetworkMingjian Zhu, Kai Han, Enhua Wu, Qiulin Zhang et al.NeurIPS 2021 · 71 citations
- Flexible high-resolution object detection on edge devices with tunable latencyShiqi Jiang, Zhiqi Lin, Yuanchun Li, Yuanchao Shu et al.MobiCom 2021 · 103 citations
- Elf: accelerate high-resolution mobile deep vision with content-aware parallel offloadingWuyang Zhang, Zhezhi He, Luyang Liu, Zhenhua Jia et al.MobiCom 2021 · 171 citations
