Fractional Skipping: Towards Finer-Grained Dynamic CNN Inference
Jianghao Shen, Yue Wang, Pengfei Xu, Yonggan Fu, Zhangyang Wang, Yingyan Lin
摘要
While increasingly deep networks are still in general desired for achieving state-of-the-art performance, for many specific inputs a simpler network might already suffice. Existing works exploited this observation by learning to skip convolutional layers in an input-dependent manner. However, we argue their binary decision scheme, i.e., either fully executing or completely bypassing one layer for a specific input, can be enhanced by introducing finer-grained, “softer” decisions. We therefore propose a Dynamic Fractional Skipping (DFS) framework. The core idea of DFS is to hypothesize layer-wise quantization (to different bitwidths) as intermediate “soft” choices to be made between fully utilizing and skipping a layer. For each input, DFS dynamically assigns a bitwidth to both weights and activations of each layer, where fully executing and skipping could be viewed as two “extremes” (i.e., full bitwidth and zero bitwidth). In this way, DFS can “fractionally” exploit a layer's expressive power during input-adaptive inference, enabling finer-grained accuracy-computational cost trade-offs. It presents a unified view to link input-adaptive layer skipping and input-adaptive hybrid quantization. Extensive experimental results demonstrate the superior tradeoff between computational cost and model expressive power (accuracy) achieved by DFS. More visualizations also indicate a smooth and consistent transition in the DFS behaviors, especially the learned choices between layer skipping and different quantizations when the total computational budgets vary, validating our hypothesis that layer quantization could be viewed as intermediate variants of layer skipping. Our source code and supplementary material are available at https://github.com/Torment123/DFS.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper12
- HW-NAS-Bench: Hardware-Aware Neural Architecture Search BenchmarkChaojian Li, Zhongzhi Yu, Yonggan Fu, Yongan Zhang 等ICLR 2021 · 被引用 128 次
- DAdaQuant: Doubly-adaptive quantization for communication-efficient Federated LearningRobert Hönig, Yiren Zhao, Robert MullinsICML 2022 · 被引用 87 次
- IntraQ: Learning Synthetic Images with Intra-Class Heterogeneity for Zero-Shot Network QuantizationYunshan Zhong, Mingbao Lin, Gongrui Nan, Jianzhuang Liu 等CVPR 2022 · 被引用 79 次
- FracTrain: Fractionally Squeezing Bit Savings Both Temporally and Spatially for Efficient DNN TrainingYonggan Fu, Haoran You, Yang Zhao, Yue Wang 等NeurIPS 2020 · 被引用 36 次
- CPT: Efficient Deep Neural Network Training via Cyclic PrecisionYonggan Fu, Han Guo, Meng Li, Xin Yang 等ICLR 2021 · 被引用 36 次
相关 Paper
- Arbitrary Bit-width Network: A Joint Layer-Wise Quantization and Adaptive Inference ApproachChen Tang, Haoyu Zhai, Kai Ouyang, Zhi Wang 等ACM MM 2022 · 被引用 15 次
- AdaBits: Neural Network Quantization With Adaptive Bit-WidthsQing Jin, Linjie Yang, Zhenyu LiaoCVPR 2020
- CSQ: Growing Mixed-Precision Quantization Scheme with Bi-level Continuous SparsificationLirui Xiao, Huanrui Yang, Zhen Dong, Kurt Keutzer 等DAC 2023 · 被引用 9 次
- Mixed Precision DNNs: All you need is a good parametrizationStefan Uhlich, Lukas Mauch, Fabien Cardinaux, Kazuki Yoshiyama 等ICLR 2020 · 被引用 159 次
- Any-Precision Deep Neural NetworksHaichao Yu, Haoxiang Li, Humphrey Shi, Thomas S. Huang 等AAAI 2021 · 被引用 79 次
