Generalizing Reuse Patterns for Efficient DNN on Microcontrollers
Jiesong Liu, Bin Ren, Xipeng Shen
摘要
Deep Neural Networks (DNNs) face challenges in deployment on resource-constrained devices due to their high computational demands. Leveraging redundancy in input data and activation maps for computation reuse is an effective way to accelerate DNN inference, especially for microcontrollers where the computing power is very limited. This work points out an important limitation in current reuse-based DNN optimizations, the narrow definition of reuse patterns in data. It proposes the concept of generalized reuse and uncovers the relations between generalized reuse patterns and row/column reorder of a matrix view of the input or activation map of a DNN. It revolutionizes the conventional view of explorable reuse patterns, drastically expanding the reuse space. It further develops two novel analytical models for analyzing the impacts of reuse patterns on the accuracy and latency of DNNs, enabling efficient selection of appropriate reuse patterns. Experiments show that generalized reuse consistently brings significant benefits, regardless of the differences among DNNs or microcontroller hardware. It delivers 1.03-2.2x speedups or 1-8% accuracy improvement over conventional reuse.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper12
- MCUNet: Tiny Deep Learning on IoT DevicesJi Lin, Wei-Ming Chen, Yujun Lin, John Cohn 等NeurIPS 2020 · 被引用 827 次
- Memory-efficient Patch-based Inference for Tiny Deep LearningJi Lin, Wei-Ming Chen, Han Cai, Chuang Gan 等NeurIPS 2021 · 被引用 190 次
- ModelDiff: testing-based DNN similarity comparison for model reuse detectionYuanchun Li, Ziqi Zhang, Bingyan Liu, Ziyue Yang 等ISSTA 2021 · 被引用 44 次
- G-TADOC: Enabling Efficient GPU-Based Text Analytics without DecompressionFeng Zhang, Zaifeng Pan, Yanliang Zhou, Jidong Zhai 等ICDE 2021 · 被引用 31 次
- Cascading structured pruning: enabling high data reuse for sparse DNN acceleratorsEdward Hanson, Shiyu Li, Hai Helen Li, Yiran ChenISCA 2022 · 被引用 30 次
相关 Paper
- Space-Efficient TREC for Enabling Deep Learning on MicrocontrollersJiesong Liu, Feng Zhang, Jiawei Guan, Hsin-Hsuan Sung 等ASPLOS 2023 · 被引用 8 次
- TREC: Transient Redundancy Elimination-based ConvolutionJiawei Guan, Feng Zhang, Jiesong Liu, Hsin-Hsuan Sung 等NeurIPS 2022 · 被引用 6 次
- RePIM: Joint Exploitation of Activation and Weight Repetitions for In-ReRAM DNN AccelerationChen-Yang Tsai, Chin-Fu Nien, Tz-Ching Yu, Hung-Yu Yeh 等DAC 2021 · 被引用 22 次
- PattPIM: A Practical ReRAM-Based DNN Accelerator by Reusing Weight Pattern RepetitionsYuhao Zhang, Zhiping Jia, Yungang Pan, Hongchao Du 等DAC 2020 · 被引用 27 次
- ConfuciuX: Autonomous Hardware Resource Assignment for DNN Accelerators using Reinforcement LearningSheng-Chun Kao, Geonhwa Jeong, Tushar KrishnaMICRO 2020 · 被引用 115 次
