Cascading structured pruning: enabling high data reuse for sparse DNN accelerators
Edward Hanson, Shiyu Li, Hai Helen Li, Yiran Chen
摘要
Performance and efficiency of running modern Deep Neural Networks (DNNs) are heavily bounded by data movement. To mitigate the data movement bottlenecks, recent DNN inference accelerator designs widely adopt aggressive compression techniques and sparse-skipping mechanisms. These mechanisms avoid transferring or computing with zero-valued weights or activations to save time and energy. However, such sparse-skipping logic involves large input buffers and irregular data access patterns, thus precluding many energy-efficient data reuse opportunities and dataflows. In this work, we propose Cascading Structured Pruning (CSP), a technique that preserves significantly more data reuse opportunities for higher energy efficiency while maintaining comparable performance relative to recent sparse architectures such as SparTen. CSP includes the following two components: At algorithm level, CSP-A induces a predictable sparsity pattern that allows for low-overhead compression of weight data and sequential access to both activation and weight data. At architecture level, CSP-H leverages CSP-A's induced sparsity pattern with a novel dataflow to access unique activation data only once, thus removing the demand for large input buffers. Each CSP-H processing element (PE) employs a novel accumulation buffer design and a counter-based sparse-skipping mechanism to support the dataflow with minimum controller overhead. We verify our approach on several representative models. Our simulated results show that CSP achieves on average 15× energy efficiency improvement over SparTen with comparable or superior speedup under most evaluations.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper4
- SOFA: A Compute-Memory Optimized Sparsity Accelerator via Cross-Stage Coordinated TilingHuizheng Wang, Jiahao Fang, Xinru Tang, Zhiheng Yue 等MICRO 2024 · 被引用 31 次
- MCBP: A Memory-Compute Efficient LLM Inference Accelerator Leveraging Bit-Slice-enabled Sparsity and RepetitivenessHuizheng Wang, Zichuan Wang, Zhiheng Yue, Yousheng Long 等MICRO 2025 · 被引用 10 次
- PADE: A Predictor-Free Sparse Attention Accelerator via Unified Execution and Stage FusionHuizheng Wang, Hongbin Wang, Zichuan Wang, Zhiheng Yue 等HPCA 2026 · 被引用 2 次
- Generalizing Reuse Patterns for Efficient DNN on MicrocontrollersJiesong Liu, Bin Ren, Xipeng ShenASPLOS 2025 · 被引用 2 次
相关 Paper
- CSCNN: Algorithm-hardware Co-design for CNN Accelerators using Centrosymmetric FiltersJiajun Li, Ahmed Louri, Avinash Karanth, Razvan C. BunescuHPCA 2021 · 被引用 9 次
- SmartExchange: Trading Higher-cost Memory Storage/Access for Lower-cost ComputationYang Zhao, Xiaohan Chen, Yue Wang, Chaojian Li 等ISCA 2020 · 被引用 44 次
- HBP: Hierarchically Balanced Pruning and Accelerator Co-Design for Efficient DNN InferenceAo Ren, Yuhao Wang, Tao Zhang, Jiaxing Shi 等DAC 2023 · 被引用 4 次
- CANDLES: Channel-Aware Novel Dataflow-Microarchitecture Co-Design for Low Energy Sparse Neural Network AccelerationSumanth Gudaparthi, Sarabjeet Singh, Surya Narayanan, Rajeev Balasubramonian 等HPCA 2022 · 被引用 23 次
- SparseTrain: Exploiting Dataflow Sparsity for Efficient Convolutional Neural Networks TrainingPengcheng Dai, Jianlei Yang, Xucheng Ye, Xingzhou Cheng 等DAC 2020 · 被引用 27 次
