Latency-aware Spatial-wise Dynamic Networks
Yizeng Han, Zhihang Yuan, Yifan Pu, Chenhao Xue, Shiji Song, Guangyu Sun, Gao Huang
摘要
Spatial-wise dynamic convolution has become a promising approach to improving the inference efficiency of deep networks. By allocating more computation to the most informative pixels, such an adaptive inference paradigm reduces the spatial redundancy in image features and saves a considerable amount of unnecessary computation. However, the theoretical efficiency achieved by previous methods can hardly translate into a realistic speedup, especially on the multi-core processors (e.g. GPUs). The key challenge is that the existing literature has only focused on designing algorithms with minimal computation, ignoring the fact that the practical latency can also be influenced by scheduling strategies and hardware properties. To bridge the gap between theoretical computation and practical efficiency, we propose a latency-aware spatial-wise dynamic network (LASNet), which performs coarse-grained spatially adaptive inference under the guidance of a novel latency prediction model. The latency prediction model can efficiently estimate the inference latency of dynamic networks by simultaneously considering algorithms, scheduling strategies, and hardware properties. We use the latency predictor to guide both the algorithm design and the scheduling optimization on various hardware platforms. Experiments on image classification, object detection and instance segmentation demonstrate that the proposed framework significantly improves the practical inference efficiency of deep networks. For example, the average latency of a ResNet-101 on the ImageNet validation set could be reduced by 36% and 46% on a server GPU (Nvidia Tesla-V100) and an edge device (Nvidia Jetson TX2 GPU) respectively without sacrificing the accuracy. Code is available at https://github.com/LeapLabTHU/LASNet .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper16
- FLatten Transformer: Vision Transformer using Focused Linear AttentionDongchen Han, Xuran Pan, Yizeng Han, Shiji Song 等ICCV 2023 · 被引用 358 次
- Adaptive Rotated Convolution for Rotated Object DetectionYifan Pu, Yiru Wang, Zhuofan Xia, Yizeng Han 等ICCV 2023 · 被引用 154 次
- Rank-DETR for High Quality Object DetectionYifan Pu, Weicong Liang, Yiduo Hao, Yuhui Yuan 等NeurIPS 2023 · 被引用 138 次
- CF-ViT: A General Coarse-to-Fine Method for Vision TransformerMengzhao Chen, Mingbao Lin, Ke Li, Yunhang Shen 等AAAI 2023 · 被引用 105 次
- Dynamic Perceiver for Efficient Visual RecognitionYizeng Han, Dongchen Han, Zeyu Liu, Yulin Wang 等ICCV 2023 · 被引用 45 次
它引用的顶会 Paper5
- Glance and Focus: a Dynamic Approach to Reducing Spatial Redundancy in Image ClassificationYulin Wang, Kangchen Lv, Rui Huang, Shiji Song 等NeurIPS 2020 · 被引用 179 次
- Batch-shaping for learning conditional channel gated networksBabak Ehteshami Bejnordi, Tijmen Blankevoort, Max WellingICLR 2020 · 被引用 82 次
- Resolution Adaptive Networks for Efficient InferenceLe Yang, Yizeng Han, Xi Chen, Shiji Song 等CVPR 2020
- Designing Network Design SpacesIlija Radosavovic, Raj Prateek Kosaraju, Ross B. Girshick, Kaiming He 等CVPR 2020
- Dynamic Convolutions: Exploiting Spatial Sparsity for Faster InferenceThomas Verelst, Tinne TuytelaarsCVPR 2020
相关 Paper
- NeuLens: spatial-based dynamic acceleration of convolutional neural networks on edgeXueyu Hou, Yongjie Guan, Tao HanMobiCom 2022 · 被引用 12 次
- SpeedDETR: Speed-aware Transformers for End-to-end Object DetectionPeiyan Dong, Zhenglun Kong, Xin Meng, Peng Zhang 等ICML 2023 · 被引用 21 次
- ZeroBN: Learning Compact Neural Networks For Latency-Critical Edge SystemsShuo Huai, Lei Zhang, Di Liu, Weichen Liu 等DAC 2021 · 被引用 16 次
- BRP-NAS: Prediction-based NAS using GCNsLukasz Dudziak, Thomas Chau, Mohamed S. Abdelfattah, Royson Lee 等NeurIPS 2020 · 被引用 233 次
- Dynamic Resolution NetworkMingjian Zhu, Kai Han, Enhua Wu, Qiulin Zhang 等NeurIPS 2021 · 被引用 71 次
