NN-AdderNet: Nonnegative and Sparse Weight Optimization Towards Ultra-Low Bitwidth AdderNet Quantization and Compression
Yunxiang Zhang, Gengchen Sun, Lizhi Fang, Biao Sun, Wenfeng Zhao
摘要
Emerging efficient deep neural network (DNN) models, such as AdderNet, have shown great promise in significantly improving hardware efficiency compared to traditional convolutional neural networks (CNNs). However, achieving low bitwidth quantization and effective model compression remains a major challenge. To this end, we introduce Nonnegative AdderNet (NNAdderNet), a quantization- and compression-friendly AdderNet variant that enables model compression down to 4 bits or even lower. We begin by proposing an equivalent transformation of the sum-of-absolute-difference (SAD) kernel in AdderNet, which allows for the formulation of nonnegative weights. This transformation effectively eliminates the need for a sign bit, thus saving 1 bit per weight. Next, we propose to exploit the dual-sparsity pattern in the weights of the activation-oriented NN-AdderNet quantized model. This inherent sparsity enhances the lossless compression performance over NN-AdderNet. Experimental results show that NN-AdderNet can achieve an average compressed weight bitwidth down to 4 bits or even lower, while achieving negligible accuracy loss as compared to full-precision AdderNet models. Such benefits are further illustrated with hardware-level energy and latency improvements in designing DNN inference accelerators. Consequently, the NN-AdderNet model exhibits both algorithmic and hardware efficiency, thus making it a promising candidate for resource-limited applications.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- AdderNet 2.0: Optimal FPGA Acceleration of AdderNet with Activation-Oriented Quantization and Fused Bias Removal based Memory OptimizationYunxiang Zhang, Omar Al Kailani, Wenfeng ZhaoDAC 2024 · 被引用 2 次
- Redistribution of Weights and Activations for AdderNet QuantizationYing Nie, Kai Han, Haikang Diao, Chuanjian Liu 等NeurIPS 2022 · 被引用 14 次
- ShiftAddNet: A Hardware-Inspired Deep NetworkHaoran You, Xiaohan Chen, Yongan Zhang, Chaojian Li 等NeurIPS 2020 · 被引用 99 次
- A2Q: Accumulator-Aware Quantization with Guaranteed Overflow AvoidanceIan Colbert, Alessandro Pappalardo, Jakoba Petri-KoenigICCV 2023 · 被引用 19 次
- Mix and Match: A Novel FPGA-Centric Deep Neural Network Quantization FrameworkSung-En Chang, Yanyu Li, Mengshu Sun, Runbin Shi 等HPCA 2021 · 被引用 125 次
