SC2020Top-tier venue
BiQGEMM: matrix multiplication with lookup table for binary-coding-based quantized DNNs
Yongkweon Jeon, Baeseong Park, Se Jung Kwon, Byeongwook Kim, Jeongin Yun, Dongsoo Lee
Abstract
The number of parameters in deep neural networks (DNNs) is rapidly increasing to support complicated tasks and to improve model accuracy. Correspondingly, the amount of computations and required memory footprint increase as well. Quantization is an efficient method to address such concerns by compressing DNNs such that computations can be simplified while required storage footprint is significantly reduced. Unfortunately, commercial CPUs and GPUs do not fully support quantization because only fixed data transfers (such as 32 bits) are allowed. As a result, even if weights are quantized (by a non-uniform quantization scheme) into a few bits, CPUs and GPUs may not access multiple quantized weights without memory bandwidth waste. Success of quantization in practice, hence, relies on an efficient computation engine design, especially for matrix multiplication that is a basic computation engine in most DNNs. In this paper, we propose a novel matrix multiplication method, called BiQGEMM, dedicated to quantized DNNs. BiQGEMM can access multiple quantized weights simultaneously in one instruction. In addition, BiQGEMM pre-computes intermediate results that are highly redundant when quantization leads to limited available computation space. Since pre-computed values are stored in lookup tables and reused, BiQGEMM achieves lower amount of overall computations. Our extensive experimental results show that BiQGEMM presents higher performance than conventional schemes when DNNs are quantized.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3992fc4c-6588-474e-94bf-593746bb27b5Cited by top-tier papers14
- LUT-GEMM: Quantized Matrix Multiplication based on LUTs for Efficient Inference in Large-Scale Generative Language ModelsGunho Park, Baeseong Park, Minsub Kim, Sungjae Lee et al.ICLR 2024 · 134 citations
- Nonuniform-to-Uniform Quantization: Towards Accurate Quantization via Generalized Straight-Through EstimationZechun Liu, Kwang-Ting Cheng, Dong Huang, Eric P. Xing et al.CVPR 2022 · 108 citations
- ShiftAddLLM: Accelerating Pretrained LLMs via Post-Training Multiplication-Less ReparameterizationHaoran You, Yipin Guo, Yichao Fu, Wei Zhou et al.NeurIPS 2024 · 47 citations
- Mr.BiQ: Post-Training Non-Uniform Quantization based on Minimizing the Reconstruction ErrorYongkweon Jeon, Chungman Lee, Eulrang Cho, Yeonju RoCVPR 2022 · 28 citations
- LUT Tensor Core: A Software-Hardware Co-Design for LUT-Based Low-Bit LLM InferenceZhiwen Mo, Lei Wang, Jianyu Wei, Zhichen Zeng et al.ISCA 2025 · 17 citations
Builds on3
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel et al.ICLR 2020 · 7,418 citations
- ERNIE 2.0: A Continual Pre-Training Framework for Language UnderstandingYu Sun, Shuohuan Wang, Yu-Kun Li, Shikun Feng et al.AAAI 2020 · 885 citations
- And the Bit Goes Down: Revisiting the Quantization of Neural NetworksPierre Stock, Armand Joulin, Rémi Gribonval, Benjamin Graham et al.ICLR 2020 · 157 citations
Related papers
- Mix and Match: A Novel FPGA-Centric Deep Neural Network Quantization FrameworkSung-En Chang, Yanyu Li, Mengshu Sun, Runbin Shi et al.HPCA 2021 · 125 citations
- Winning Both the Accuracy of Floating Point Activation and the Simplicity of Integer ArithmeticYulhwa Kim, Jaeyong Jang, Jehun Lee, Jihoon Park et al.ICLR 2023
- Term quantization: furthering quantization at run timeHsiang-Tsung Kung, Bradley McDanel, Sai Qian ZhangSC 2020 · 10 citations
- BBS: Bi-Directional Bit-Level Sparsity for Deep Learning AccelerationYuzong Chen, Jian Meng, Jae-sun Seo, Mohamed S. AbdelfattahMICRO 2024 · 25 citations
- LiquidGEMM: Hardware-Efficient W4A8 GEMM Kernel for High-Performance LLM ServingHuanqi Hu, Bowen Xiao, Shixuan Sun, Jianian Yin et al.SC 2025 · 2 citations
