Block and Subword-Scaling Floating-Point (BSFP) : An Efficient Non-Uniform Quantization For Low Precision Inference
Yun-Chen Lo, Tse-Kuang Lee, Ren-Shuo Liu
摘要
In this paper, we propose Block and Subword-Scaling Floating-Point (BSFP), a non-uniform quantization scheme for the skewed and non-uniform distribution of weight vectors in neural networks. By quantizing each weight vector as the superposition of multiple subword vectors (in two's complement) with scaling factors (in Low-bit Floating-Point, LBFP), BSFP can effectively fit the distribution of weight vectors while maintaining high computation efficiency. Furthermore, we present a grid search-based MSE-optimal quantization flow and a scaled serial processing engine to complete the quantization pipeline and the infrastructure. The experimental results on the ImageNet classification task show that our proposed method outperforms state-of-the-art Microsoft Floating Point (MSFP) by up to 20.56% top-1 accuracy at the same weight precision and reduces up to 10.3% model size. Furthermore, BSFP outperforms MSFP by up to 2.0 computing throughput and up to 5.3 energy efficiency under the same silicon area budget.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper3
- MX+: Pushing the Limits of Microscaling Formats for Efficient Large Language Model ServingJungi Lee, Junyong Park, Soohyun Cha, Jaehoon Cho 等MICRO 2025 · 被引用 7 次
- M2XFP: A Metadata-Augmented Microscaling Data Format for Efficient Low-bit QuantizationWeiming Hu, Zihan Zhang, Haoyan Zhang, Chen Zhang 等ASPLOS 2026 · 被引用 2 次
- Pushing the Limits of BFP on Narrow Precision LLM InferenceHui Wang, Yuan Cheng, Xiaomeng Han, Zhengpeng Zhao 等AAAI 2025 · 被引用 1 次
相关 Paper
- Additive Powers-of-Two Quantization: An Efficient Non-uniform Discretization for Neural NetworksYuhang Li, Xin Dong, Wei WangICLR 2020 · 被引用 315 次
- RMSMP: A Novel Deep Neural Network Quantization Framework with Row-wise Mixed Schemes and Multiple PrecisionsSung-En Chang, Yanyu Li, Mengshu Sun, Weiwen Jiang 等ICCV 2021 · 被引用 14 次
- BiE: Bi-Exponent Block Floating-Point for Large Language Models QuantizationLancheng Zou, Wenqian Zhao, Shuo Yin, Chen Bai 等ICML 2024 · 被引用 13 次
- BiQGEMM: matrix multiplication with lookup table for binary-coding-based quantized DNNsYongkweon Jeon, Baeseong Park, Se Jung Kwon, Byeongwook Kim 等SC 2020 · 被引用 31 次
- Improving Low-Precision Network Quantization via Bin RegularizationTiantian Han, Dong Li, Ji Liu, Lu Tian 等ICCV 2021 · 被引用 45 次
