QUIK: Towards End-to-end 4-Bit Inference on Generative Large Language Models
Saleh Ashkboos, Ilia Markov, Elias Frantar, Tingxuan Zhong, Xincheng Wang, Jie Ren, Torsten Hoefler, Dan Alistarh
Abstract
Large Language Models (LLMs) from the GPT family have become extremely popular, leading to a race towards reducing their inference costs to allow for efficient local computation. However, the vast majority of existing work focuses on weight-only quantization, which can reduce runtime costs in the memory-bound onetoken-at-a-time generative setting, but does not address costs in compute-bound scenarios, such as batched inference or prompt processing. In this paper, we address the general quantization problem, where both weights and activations should be quantized, which leads to computational improvements in general. We show that the majority of inference computations for large generative models can be performed with both weights and activations being cast to 4 bits, while at the same time maintaining good accuracy. We achieve this via a hybrid quantization strategy called QUIK that compresses most of the weights and activations to 4-bit, while keeping a small fraction of "outlier" weights and activations in higher-precision. QUIK is that it is designed with computational efficiency in mind: we provide GPU kernels matching the QUIK format with highly-efficient layer-wise runtimes, which lead to practical end-to-end throughput improvements of up to 3.4x relative to FP16 execution. We provide detailed studies for models from the OPT, LLaMA-2 and Falcon families, as well as a first instance of accurate inference using quantization plus 2:4 sparsity. Anonymized code is available here.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d65d3c1a-6e24-405f-8cc8-0ac911440ab5Cited by top-tier papers5
- WUSH: Near-Optimal Adaptive Transforms for LLM QuantizationJiale Chen, Vage Egiazarian, Roberto Castro, Torsten Hoefler et al.ICML 2026 · 7 citations
- ARCQuant: Boosting NVFP4 Quantization with Augmented Residual Channels for LLMsHaoqian Meng, Yilun Luo, Yafei Zhao, Wenyuan Liu et al.ACL 2026 · 5 citations
- TACO: Efficient Communication Compression of Intermediate Tensors for Scalable Tensor-Parallel LLM TrainingMan Liu, Xingchen Liu, Xingjian Tian, Bing Lu et al.HPDC 2026 · 1 citation
- LATMiX: Learnable Affine Transformations for Microscaling Quantization of LLMsOfir Gordon, Lior Dikstein, Arnon Netzer, Idan Achituve et al.ICML 2026
- ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank ResidualsUtkarsh Saxena, Sayeh Sharify, Kaushik Roy, Xin WangICML 2025
Builds on9
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra et al.NeurIPS 2022 · 5,493 citations
- SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language ModelsGuangxuan Xiao, Ji Lin, Mickaël Seznec, Hao Wu et al.ICML 2023 · 1,493 citations
- Fast Inference from Transformers via Speculative DecodingYaniv Leviathan, Matan Kalman, Yossi MatiasICML 2023 · 1,472 citations
- SparseGPT: Massive Language Models Can be Accurately Pruned in One-ShotElias Frantar, Dan AlistarhICML 2023 · 1,240 citations
- FlexGen: High-Throughput Generative Inference of Large Language Models with a Single GPUYing Sheng, Lianmin Zheng, Binhang Yuan, Zhuohan Li et al.ICML 2023 · 683 citations
Related papers
- COMET: Towards Practical W4A4KV4 LLMs ServingLian Liu, Long Cheng, Haimeng Ren, Zhaohui Xu et al.ASPLOS 2025 · 5 citations
- Agile-Quant: Activation-Guided Quantization for Faster Inference of LLMs on the EdgeXuan Shen, Peiyan Dong, Lei Lu, Zhenglun Kong et al.AAAI 2024 · 59 citations
- TwinQuant: Learnable Subspace Decomposition for 4-Bit LLM QuantizationHaodong WANG, Junjie Liu, Zicong Hong, Qianli Liu et al.ICML 2026
- MARLIN: Mixed-Precision Auto-Regressive Parallel Inference on Large Language ModelsElias Frantar, Roberto L. Castro, Jiale Chen, Torsten Hoefler et al.PPoPP 2025 · 24 citations
- Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data FormatChao Fang, Man Shi, Robin Geens, Arne Symons et al.HPCA 2025 · 15 citations
