PsumQuant: In-line Post-training Partial Sum Quantizer for Energy Efficient NPU Inference
Sangwoo Hwang, Yeeun Hong, Jaeha Kung
摘要
The rapid growth of deep neural networks (DNNs) has intensified the demand for efficient hardware acceleration under quantization. While prior research has successfully reduced weight and activation precision, partial sums generated during accumulation often retain high precision, resulting in significant energy overhead. In this work, we analyze psum distributions in tiled architectures and reveal that within-tile outliers are input-dependent. We propose PsumQuant, a post-training, input-aware quantization that predicts psum scales on-the-fly. By leveraging the crest factor of input activations, our learnable scale predictor effectively bounds the psum bit-width while handling the extreme outliers in DNNs. Experimental results on a systolic array demonstrate that PsumQuant compresses psum precision down to 8-bit within only a 1% accuracy drop on ResNet-18 and a marginal 0.04 perplexity increase on Llama-3.1. Furthermore, bit-width reduction with PsumQuant results in a 45% reduction in total energy with minimal accuracy loss, demonstrating that PsumQuant provides a highly efficient solution for actual NPU architectures.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper8
- WinoGrande: An Adversarial Winograd Schema Challenge at ScaleKeisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, Yejin ChoiAAAI 2020 · 被引用 3,037 次
- Learned Step Size quantizationSteven K. Esser, Jeffrey L. McKinstry, Deepika Bablani, Rathinakumar Appuswamy 等ICLR 2020 · 被引用 1,037 次
- Up or Down? Adaptive Rounding for Post-Training QuantizationMarkus Nagel, Rana Ali Amjad, Mart van Baalen, Christos Louizos 等ICML 2020 · 被引用 816 次
- HAWQ: Hessian AWare Quantization of Neural Networks With Mixed-PrecisionZhen Dong, Zhewei Yao, Amir Gholami, Michael W. Mahoney 等ICCV 2019 · 被引用 645 次
- SpikedAttention: Training-Free and Fully Spike-Driven Transformer-to-SNN Conversion with Winner-Oriented Spike Shift for Softmax OperationSangwoo Hwang, Seunghyun Lee, Dahoon Park, Donghun Lee 等NeurIPS 2024 · 被引用 23 次
相关 Paper
- APSQ: Additive Partial Sum Quantization with Algorithm-Hardware Co-DesignYonghao Tan, Pingcheng Dong, Yongkun Wu, Yu Liu 等DAC 2025 · 被引用 1 次
- QPP: Real-Time Quantization Parameter Prediction for Deep Neural NetworksVladimir Kryzhanovskiy, Gleb Balitskiy, Nikolay Kozyrskiy, Aleksandr ZuruevCVPR 2021
- A2Q: Accumulator-Aware Quantization with Guaranteed Overflow AvoidanceIan Colbert, Alessandro Pappalardo, Jakoba Petri-KoenigICCV 2023 · 被引用 19 次
- PD-Quant: Post-Training Quantization Based on Prediction Difference MetricJiawei Liu, Lin Niu, Zhihang Yuan, Dawei Yang 等CVPR 2023
- SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language ModelsGuangxuan Xiao, Ji Lin, Mickaël Seznec, Hao Wu 等ICML 2023 · 被引用 1,493 次
