SmartExchange: Trading Higher-cost Memory Storage/Access for Lower-cost Computation
Yang Zhao, Xiaohan Chen, Yue Wang, Chaojian Li, Haoran You, Yonggan Fu, Yuan Xie, Zhangyang Wang, Yingyan Lin
摘要
We present SmartExchange, an algorithm-hardware co-design framework to trade higher-cost memory storage/access for lower-cost computation, for energy-efficient inference of deep neural networks (DNNs). We develop a novel algorithm to enforce a specially favorable DNN weight structure, where each layerwise weight matrix can be stored as the product of a small basis matrix and a large sparse coefficient matrix whose non-zero elements are all power-of-2. To our best knowledge, this algorithm is the first formulation that integrates three mainstream model compression ideas: sparsification or pruning, decomposition, and quantization, into one unified framework. The resulting sparse and readily-quantized DNN thus enjoys greatly reduced energy consumption in data movement as well as weight storage. On top of that, we further design a dedicated accelerator to fully utilize the SmartExchange-enforced weights to improve both energy efficiency and latency performance. Extensive experiments show that 1) on the algorithm level, SmartExchange outperforms stateof-the-art compression techniques, including merely sparsification or pruning, decomposition, and quantization, in various ablation studies based on nine models and four datasets; and 2) on the hardware level, SmartExchange can boost the energy efficiency by up to 6.7× and reduce the latency by up to 19.2× over four state-of-the-art DNN accelerators, when benchmarked on seven DNN models (including four standard DNNs, two compact DNN models, and one segmentation model) and three datasets.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper13
- SpAtten: Efficient Sparse Attention Architecture with Cascade Token and Head PruningHanrui Wang, Zhekai Zhang, Song HanHPCA 2021 · 被引用 412 次
- HW-NAS-Bench: Hardware-Aware Neural Architecture Search BenchmarkChaojian Li, Zhongzhi Yu, Yonggan Fu, Yongan Zhang 等ICLR 2021 · 被引用 128 次
- Unified Visual Transformer CompressionShixing Yu, Tianlong Chen, Jiayi Shen, Huan Yuan 等ICLR 2022 · 被引用 118 次
- ShiftAddNet: A Hardware-Inspired Deep NetworkHaoran You, Xiaohan Chen, Yongan Zhang, Chaojian Li 等NeurIPS 2020 · 被引用 99 次
- Instant-3D: Instant Neural Radiance Field Training Towards On-Device AR/VR 3D ReconstructionSixu Li, Chaojian Li, Wenbo Zhu, Boyang Tony Yu 等ISCA 2023 · 被引用 79 次
它引用的顶会 Paper2
- Differentiable Soft Quantization: Bridging Full-Precision and Low-Bit Neural NetworksRuihao Gong, Xianglong Liu, Shenghu Jiang, Tianxiang Li 等ICCV 2019 · 被引用 540 次
- Interstellar: Using Halide's Scheduling Language to Analyze DNN AcceleratorsXuan Yang, Mingyu Gao, Qiaoyi Liu, Jeff Setter 等ASPLOS 2020 · 被引用 237 次
相关 Paper
- Cascading structured pruning: enabling high data reuse for sparse DNN acceleratorsEdward Hanson, Shiyu Li, Hai Helen Li, Yiran ChenISCA 2022 · 被引用 30 次
- STC: Significance-aware Transform-based Codec Framework for External Memory Access ReductionFeng Xiong, Fengbin Tu, Man Shi, Yang Wang 等DAC 2020 · 被引用 23 次
- CSCNN: Algorithm-hardware Co-design for CNN Accelerators using Centrosymmetric FiltersJiajun Li, Ahmed Louri, Avinash Karanth, Razvan C. BunescuHPCA 2021 · 被引用 9 次
- BitPattern: Enabling Efficient Bit-Serial Acceleration of Deep Neural Networks through Bit-Pattern PruningGang Wang, Siqi Cai, Zhenyu Li, Wenjie Li 等DAC 2025
- ETTE: Efficient Tensor-Train-based Computing Engine for Deep Neural NetworksYu Gong, Miao Yin, Lingyi Huang, Jinqi Xiao 等ISCA 2023 · 被引用 12 次
