GPT3.int8(): 8-bit Matrix Multiplication for Transformers at Scale
Tim Dettmers, Mike Lewis, Younes Belkada, Luke Zettlemoyer
摘要
Large language models have been widely adopted but require significant GPU memory for inference and finetuning. We develop methods for Int8 matrix multiplication for transformer multi-layer perceptron (MLP) and attention projection layers, which cut the required memory for inference by half while retaining full precision performance. With our method, a 16/32-bit checkpoint can be loaded, converted to Int8, and used immediately without performance degradation -no post-quantization training is required. The key challenge, which we empirically show for the first time, is that existing quantization methods perform poorly at scale due to emergent outlier feature dimensions. We find that standard quantization techniques for matrix multiplication fail beyond 1.3B parameters. To overcome this barrier, we develop vector-wise quantization, which keeps separate normalization constants for each inner product in the matrix multiplication. Additionally, we identify layer and input invariant feature dimensions in the hidden states, which heavily influence attention and disrupt quantization methods starting at 13B parameters. To scale to 13B, we develop a new mixed-precision matrix decomposition scheme, which allows scaling without performance degradation to at least 13B parameters. This result makes large transformers more accessible, for example, by enabling inference with GPT-J and T5-11B on a single free cloud GPU, GPT-NeoX-20B on a single gaming-grade GPU, and OPT-30B on a single data-center-grade GPU. We open source our software.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- ReLoRA: High-Rank Training Through Low-Rank UpdatesVladislav Lialin, Sherin Muckatira, Namrata Shivagunde, Anna RumshiskyICLR 2024 · 被引用 214 次
- Treasures in Discarded Weights for LLM QuantizationHao Yu, Yang Zhou, Bohua Chen, Zelan Yang 等AAAI 2025 · 被引用 1 次
- IM-Unpack: Training and Inference with Arbitrarily Low Precision IntegersZhanpeng Zeng, Karthikeyan Sankaralingam, Vikas SinghICML 2024 · 被引用 1 次
- No outlier channels but with outlier blocksShanwen Mao, Hao Zhang, Jiasheng Li, Haoyu Qiao 等ICLR 2026
- One Leak Away: How Pretrained Model Exposure Amplifies Jailbreak Risks in Finetuned LLMsYixin Tan, Yu Zhe, Rui Wen, Jun SakumaCCS 2026
它引用的顶会 Paper8
- Differentiable Soft Quantization: Bridging Full-Precision and Low-Bit Neural NetworksRuihao Gong, Xianglong Liu, Shenghu Jiang, Tianxiang Li 等ICCV 2019 · 被引用 540 次
- Training with Quantization Noise for Extreme Model CompressionPierre Stock, Angela Fan, Benjamin Graham, Edouard Grave 等ICLR 2021 · 被引用 262 次
- TernaryBERT: Distillation-aware Ultra-low Bit BERTWei Zhang, Lu Hou, Yichun Yin, Lifeng Shang 等EMNLP 2020 · 被引用 147 次
- A Statistical Framework for Low-bitwidth Training of Deep Neural NetworksJianfei Chen, Yu Gai, Zhewei Yao, Michael W. Mahoney 等NeurIPS 2020 · 被引用 75 次
- Efficient Large Scale Language Modeling with Mixtures of ExpertsMikel Artetxe, Shruti Bhosale, Naman Goyal, Todor Mihaylov 等EMNLP 2022 · 被引用 71 次
相关 Paper
- ZeroQuant: Efficient and Affordable Post-Training Quantization for Large-Scale TransformersZhewei Yao, Reza Yazdani Aminabadi, Minjia Zhang, Xiaoxia Wu 等NeurIPS 2022 · 被引用 816 次
- LUT-GEMM: Quantized Matrix Multiplication based on LUTs for Efficient Inference in Large-Scale Generative Language ModelsGunho Park, Baeseong Park, Minsub Kim, Sungjae Lee 等ICLR 2024 · 被引用 134 次
- ABQ-LLM: Arbitrary-Bit Quantized Inference Acceleration for Large Language ModelsChao Zeng, Songwei Liu, Yusheng Xie, Hong Liu 等AAAI 2025 · 被引用 24 次
- SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language ModelsGuangxuan Xiao, Ji Lin, Mickaël Seznec, Hao Wu 等ICML 2023 · 被引用 1,493 次
- GANQ: GPU-Adaptive Non-Uniform Quantization for Large Language ModelsPengxiang Zhao, Xiaoming YuanICML 2025
