Non-uniform DNN Structured Subnets Sampling for Dynamic Inference
Li Yang, Zhezhi He, Yu Cao, Deliang Fan
摘要
With the success of Deep Neural Networks (DNN), many recent works have been focusing on developing hardware accelerator for power and resource-limited system via model compression techniques, such as quantization, pruning, low-rank approximation and etc. However, almost all existing compressed DNNs are fixed after deployment, which lacks run-time adaptive structure to adapt to its dynamic hardware resource allocation, power budget, throughput requirement, as well as dynamic workload. As the countermeasure, to construct a novel run-time dynamic DNN structure, we propose a novel DNN sub-network sampling method via non-uniform channel selection for subnets generation. Thus, user can trade off between power, speed, computing load and accuracy on-the-fly after the deployment, depending on the dynamic requirements or specifications of the given system. We verify the proposed model on both CIFAR-10 and ImageNet dataset using ResNets, which outperforms the same sub-nets trained individually and other related works. It shows that, our method can achieve latency trade-off among 13.4, 24.6, 41.3, 62.1(ms) and 30.5, 38.7, 51, 65.4(ms) for GPU with 128 batch-size and CPU respectively on ImageNet using ResNet18.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- RepNet: Efficient On-Device Learning via Feature ReprogrammingLi Yang, Adnan Siraj Rakin, Deliang FanCVPR 2022 · 被引用 18 次
- Dancing along Battery: Enabling Transformer with Run-time Reconfigurability on Mobile DevicesYuhong Song, Weiwen Jiang, Bingbing Li, Panjie Qi 等DAC 2021 · 被引用 16 次
它引用的顶会 Paper1
相关 Paper
- PIM-Prune: Fine-Grain DCNN Pruning for Crossbar-Based Process-In-Memory ArchitectureChaoqun Chu, Yanzhi Wang, Yilong Zhao, Xiaolong Ma 等DAC 2020 · 被引用 64 次
- Dynamic Structure Pruning for Compressing CNNsJun-Hyung Park, Yeachan Kim, Junho Kim, Joon-Young Choi 等AAAI 2023 · 被引用 24 次
- SmartExchange: Trading Higher-cost Memory Storage/Access for Lower-cost ComputationYang Zhao, Xiaohan Chen, Yue Wang, Chaojian Li 等ISCA 2020 · 被引用 44 次
- OPQ: Compressing Deep Neural Networks with One-shot Pruning-QuantizationPeng Hu, Xi Peng, Hongyuan Zhu, Mohamed M. Sabry Aly 等AAAI 2021 · 被引用 79 次
- AOWS: Adaptive and Optimal Network Width Search With Latency ConstraintsMaxim Berman, Leonid Pishchulin, Ning Xu, Matthew B. Blaschko 等CVPR 2020
