GPT3.int8(): 8-bit Matrix Multiplication for Transformers at Scale
Tim Dettmers, Mike Lewis, Younes Belkada, Luke Zettlemoyer
Abstract
Large language models have been widely adopted but require significant GPU memory for inference and finetuning. We develop methods for Int8 matrix multiplication for transformer multi-layer perceptron (MLP) and attention projection layers, which cut the required memory for inference by half while retaining full precision performance. With our method, a 16/32-bit checkpoint can be loaded, converted to Int8, and used immediately without performance degradation -no post-quantization training is required. The key challenge, which we empirically show for the first time, is that existing quantization methods perform poorly at scale due to emergent outlier feature dimensions. We find that standard quantization techniques for matrix multiplication fail beyond 1.3B parameters. To overcome this barrier, we develop vector-wise quantization, which keeps separate normalization constants for each inner product in the matrix multiplication. Additionally, we identify layer and input invariant feature dimensions in the hidden states, which heavily influence attention and disrupt quantization methods starting at 13B parameters. To scale to 13B, we develop a new mixed-precision matrix decomposition scheme, which allows scaling without performance degradation to at least 13B parameters. This result makes large transformers more accessible, for example, by enabling inference with GPT-J and T5-11B on a single free cloud GPU, GPT-NeoX-20B on a single gaming-grade GPU, and OPT-30B on a single data-center-grade GPU. We open source our software.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext fbb2459b-4b2f-4dd9-9808-99444688ff0fCited by top-tier papers5
- ReLoRA: High-Rank Training Through Low-Rank UpdatesVladislav Lialin, Sherin Muckatira, Namrata Shivagunde, Anna RumshiskyICLR 2024 · 214 citations
- Treasures in Discarded Weights for LLM QuantizationHao Yu, Yang Zhou, Bohua Chen, Zelan Yang et al.AAAI 2025 · 1 citation
- IM-Unpack: Training and Inference with Arbitrarily Low Precision IntegersZhanpeng Zeng, Karthikeyan Sankaralingam, Vikas SinghICML 2024 · 1 citation
- No outlier channels but with outlier blocksShanwen Mao, Hao Zhang, Jiasheng Li, Haoyu Qiao et al.ICLR 2026
- One Leak Away: How Pretrained Model Exposure Amplifies Jailbreak Risks in Finetuned LLMsYixin Tan, Yu Zhe, Rui Wen, Jun SakumaCCS 2026
Builds on8
- Differentiable Soft Quantization: Bridging Full-Precision and Low-Bit Neural NetworksRuihao Gong, Xianglong Liu, Shenghu Jiang, Tianxiang Li et al.ICCV 2019 · 540 citations
- Training with Quantization Noise for Extreme Model CompressionPierre Stock, Angela Fan, Benjamin Graham, Edouard Grave et al.ICLR 2021 · 262 citations
- TernaryBERT: Distillation-aware Ultra-low Bit BERTWei Zhang, Lu Hou, Yichun Yin, Lifeng Shang et al.EMNLP 2020 · 147 citations
- A Statistical Framework for Low-bitwidth Training of Deep Neural NetworksJianfei Chen, Yu Gai, Zhewei Yao, Michael W. Mahoney et al.NeurIPS 2020 · 75 citations
- Efficient Large Scale Language Modeling with Mixtures of ExpertsMikel Artetxe, Shruti Bhosale, Naman Goyal, Todor Mihaylov et al.EMNLP 2022 · 71 citations
Related papers
- ZeroQuant: Efficient and Affordable Post-Training Quantization for Large-Scale TransformersZhewei Yao, Reza Yazdani Aminabadi, Minjia Zhang, Xiaoxia Wu et al.NeurIPS 2022 · 816 citations
- LUT-GEMM: Quantized Matrix Multiplication based on LUTs for Efficient Inference in Large-Scale Generative Language ModelsGunho Park, Baeseong Park, Minsub Kim, Sungjae Lee et al.ICLR 2024 · 134 citations
- ABQ-LLM: Arbitrary-Bit Quantized Inference Acceleration for Large Language ModelsChao Zeng, Songwei Liu, Yusheng Xie, Hong Liu et al.AAAI 2025 · 24 citations
- SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language ModelsGuangxuan Xiao, Ji Lin, Mickaël Seznec, Hao Wu et al.ICML 2023 · 1,493 citations
- GANQ: GPU-Adaptive Non-Uniform Quantization for Large Language ModelsPengxiang Zhao, Xiaoming YuanICML 2025
