INSPIRE: Accelerating Deep Neural Networks via Hardware-friendly Index-Pair Encoding
Fangxin Liu, Ning Yang, Zhiyan Song, Zongwu Wang, Haomin Li, Shiyuan Huang, Zhuoran Song, Songwen Pei, Li Jiang
摘要
Deep Neural Network (DNN) inference consumes significant computing resources and development efforts due to the growing model size. Quantization is a promising technique to reduce the computation and memory cost of DNNs. Most existing quantization methods rely on fixed-point integers or floating-point types, which require more bits to maintain model accuracy. In contrast, variable-length quantization, which combines high precision for values with significant magnitudes (i.e., outliers) and low precision for normal values, offers algorithmic advantages but introduces significant hardware overhead due to variable-length encoding and decoding. Also, existing quantization methods are less effective for both (dynamic) activations and (static) weights due to the presence of outliers.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- ANT: Exploiting Adaptive Numerical Data Type for Low-bit Deep Neural Network QuantizationCong Guo, Chen Zhang, Jingwen Leng, Zihan Liu 等MICRO 2022 · 被引用 109 次
- Drift: Leveraging Distribution-based Dynamic Precision Quantization for Efficient Deep Neural Network AccelerationLian Liu, Zhaohui Xu, Yintao He, Ying Wang 等DAC 2024 · 被引用 5 次
- BLOOM: Bit-Slice Framework for DNN Acceleration with Mixed-PrecisionFangxin Liu, Ning Yang, Zongwu Wang, Xuanpeng Zhu 等DAC 2025 · 被引用 1 次
- Instance-Aware Dynamic Neural Network QuantizationZhenhua Liu, Yunhe Wang, Kai Han, Siwei Ma 等CVPR 2022 · 被引用 38 次
- REx: Data-Free Residual Quantization Error ExpansionEdouard Yvinec, Arnaud Dapogny, Matthieu Cord, Kevin BaillyNeurIPS 2023 · 被引用 11 次
