SPARK: Scalable and Precision-Aware Acceleration of Neural Networks via Efficient Encoding
Fangxin Liu, Ning Yang, Haomin Li, Zongwu Wang, Zhuoran Song, Songwen Pei, Li Jiang
摘要
Deep Neural Networks (DNNs) have demonstrated remarkable success; however, their increasing model size poses a challenge due to the widening gap between model size and hardware capacity. To address this, model compression techniques have been proposed, but existing compression methods struggle to effectively handle the significant parameter variations (activations and weights) within the model. Moreover, current variance-aware encoding solutions for compression introduce complex logic, leading to limited compression benefits and hardware efficiency. In this context, we present SPARK, a novel algorithm/architecture co-designed solution that utilizes variable-length data representation for local parameter value processing, offering low hardware overhead and high-performance gains. Our key insight is that the high-order part in quantized values are often sparse, allowing us to employ an identity bit to assign the appropriate encoding length, thereby eliminating redundant bit-length footprints. This reduction in data representation based on data characteristics enables a serialized structured data encoding scheme that seamlessly integrates with existing hardware accelerators, such as systolic arrays. We evaluate SPARK-based accelerators against some existing encoding-based accelerator, and our results demonstrate significant improvements. The SPARK-based accelerator achieves up to 4.65 × speedup and 74.7% energy reduction, while maintaining superior model accuracy.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper2
- MCBP: A Memory-Compute Efficient LLM Inference Accelerator Leveraging Bit-Slice-enabled Sparsity and RepetitivenessHuizheng Wang, Zichuan Wang, Zhiheng Yue, Yousheng Long 等MICRO 2025 · 被引用 10 次
- -LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical FormatsYuzong Chen, Chao Fang, Xilai Dai, Yuheng Wu 等ISCA 2026 · 被引用 4 次
相关 Paper
- Tight Compression: Compressing CNN Model Tightly Through Unstructured Pruning and Simulated Annealing Based PermutationXizi Chen, Jingyang Zhu, Jingbo Jiang, Chi-Ying TsuiDAC 2020 · 被引用 30 次
- BLOOM: Bit-Slice Framework for DNN Acceleration with Mixed-PrecisionFangxin Liu, Ning Yang, Zongwu Wang, Xuanpeng Zhu 等DAC 2025 · 被引用 1 次
- SmartExchange: Trading Higher-cost Memory Storage/Access for Lower-cost ComputationYang Zhao, Xiaohan Chen, Yue Wang, Chaojian Li 等ISCA 2020 · 被引用 44 次
- BBS: Bi-Directional Bit-Level Sparsity for Deep Learning AccelerationYuzong Chen, Jian Meng, Jae-sun Seo, Mohamed S. AbdelfattahMICRO 2024 · 被引用 25 次
- CANDLES: Channel-Aware Novel Dataflow-Microarchitecture Co-Design for Low Energy Sparse Neural Network AccelerationSumanth Gudaparthi, Sarabjeet Singh, Surya Narayanan, Rajeev Balasubramonian 等HPCA 2022 · 被引用 23 次
