MEET: Towards Memory-Efficient Temporal Sparse Deep Neural Networks
Zeqi Zhu, Ibrahim Batuhan Akkaya, Luc Waeijen, Egor Bondarev, Arash Pourtaherian, Orlando Moreira
Abstract
Deep Neural Networks (DNNs) are accurate but computeintensive, leading to substantial energy consumption during inference. Exploiting temporal redundancy through ∆-Σ convolution [26] in video processing has proven to greatly enhance computation efciency. However, temporal ∆-Σ DNNs typically require substantial memory for storing neuron states to compute inter-frame differences, hindering their on-chip deployment. To mitigate this memory cost, directly compressing the states can disrupt the linearity of temporal ∆-Σ convolution, causing accumulated errors in long-term ∆-Σ processing. Thus, we propose MEET, an optimization framework for MEmory-Efcient Temporal ∆-Σ DNNs. MEET transfers the state compression challenge to a well-established weight compression problem by trading fewer activations for more weights and introduces a co-design of network architecture and suppression method to optimize for mixed spatial-temporal execution. Evaluations on three vision applications demonstrate a reduction of 5.1∼13.3 × in total memory compared to the most computation-efcient temporal DNNs, while preserving the computation efciency and model accuracy in long-term ∆-Σ processing. MEET facilitates the deployment of temporal ∆-Σ DNNs within on-chip memory of embedded eventdriven platforms, empowering low-power edge processing.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext cef5d3b7-d2ee-4598-aaad-7573493adbc1Builds on8
- A ConvNet for the 2020sZhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer et al.CVPR 2022 · 6,782 citations
- Scaling Up Your Kernels to 31×31: Revisiting Large Kernel Design in CNNsXiaohan Ding, Xiangyu Zhang, Jungong Han, Guiguang DingCVPR 2022 · 1,298 citations
- FairNAS: Rethinking Evaluation Fairness of Weight Sharing Neural Architecture SearchXiangxiang Chu, Bo Zhang, Ruijun XuICCV 2021 · 362 citations
- Pruning vs Quantization: Which is Better?Andrey Kuzmin, Markus Nagel, Mart van Baalen, Arash Behboodi et al.NeurIPS 2023 · 152 citations
- More ConvNets in the 2020s: Scaling up Kernels Beyond 51x51 using SparsityShiwei Liu, Tianlong Chen, Xiaohan Chen, Xuxi Chen et al.ICLR 2023 · 87 citations
Related papers
- Skip-Convolutions for Efficient Video ProcessingAmirhossein Habibian, Davide Abati, Taco S. Cohen, Babak Ehteshami BejnordiCVPR 2021
- Deep Learning in Latent Space for Video Prediction and CompressionBowen Liu, Yu Chen, Shiyu Liu, Hun-Seok KimCVPR 2021
- Shortcut-V2V: Compression Framework for Video-to-Video Translation based on Temporal Redundancy ReductionChaeyeon Chung, Yeojeong Park, Seunghwan Choi, Munkhsoyol Ganbat et al.ICCV 2023 · 4 citations
- DeltaCNN: End-to-End CNN Inference of Sparse Frame Differences in VideosMathias Parger, Chengcheng Tang, Christopher D. Twigg, Cem Keskin et al.CVPR 2022 · 31 citations
- MotionDeltaCNN: Sparse CNN Inference of Frame Differences in Moving Camera Videos with Spherical Buffers and Padded ConvolutionsMathias Parger, Chengcheng Tang, Thomas Neff, Christopher D. Twigg et al.ICCV 2023 · 11 citations
