Training Transformers with 4-bit Integers
Haocheng Xi, Changhao Li, Jianfei Chen, Jun Zhu
Abstract
Quantizing the activation, weight, and gradient to 4-bit is promising to accelerate neural network training. However, existing 4-bit training methods require custom numerical formats which are not supported by contemporary hardware. In this work, we propose a training method for transformers with all matrix multiplications implemented with the INT4 arithmetic. Training with an ultra-low INT4 precision is challenging. To achieve this, we carefully analyze the specific structures of activation and gradients in transformers to propose dedicated quantizers for them. For forward propagation, we identify the challenge of outliers and propose a Hadamard quantizer to suppress the outliers. For backpropagation, we leverage the structural sparsity of gradients by proposing bit splitting and leverage score sampling techniques to quantize gradients accurately. Our algorithm achieves competitive accuracy on a wide range of tasks including natural language understanding, machine translation, and image classification. Unlike previous 4-bit training methods, our algorithm can be implemented on the current generation of GPUs. Our prototypical linear operator implementation is up to 2.2 times faster than the FP16 counterparts and speeds up the training by up to 35.1%. Our code is available at https://github.com/xijiu9/Train_Transformers_with_INT4 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 73c4bdfe-9bd5-46fb-9a08-ad85f9fc6b7cCited by top-tier papers29
- QuaRot: Outlier-Free 4-Bit Inference in Rotated LLMsSaleh Ashkboos, Amirkeivan Mohtashami, Maximilian L. Croci, Bo Li et al.NeurIPS 2024 · 723 citations
- QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice CodebooksAlbert Tseng, Jerry Chee, Qingyao Sun, Volodymyr Kuleshov et al.ICML 2024 · 295 citations
- DuQuant: Distributing Outliers via Dual Transformation Makes Stronger Quantized LLMsHaokun Lin, Haobo Xu, Yichen Wu, Jingzhi Cui et al.NeurIPS 2024 · 206 citations
- Federated Full-Parameter Tuning of Billion-Sized Language Models with Communication Cost under 18 KilobytesZhen Qin, Daoyuan Chen, Bingchen Qian, Bolin Ding et al.ICML 2024 · 73 citations
- Quartet: Native FP4 Training Can Be Optimal for Large Language ModelsRoberto L. Castro, Andrei Panferov, Rush Tabesh, Oliver Sieberling et al.NeurIPS 2025 · 38 citations
Builds on22
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
- MLP-Mixer: An all-MLP Architecture for VisionIlya O. Tolstikhin, Neil Houlsby, Alexander Kolesnikov, Lucas Beyer et al.NeurIPS 2021 · 3,862 citations
- Reformer: The Efficient TransformerNikita Kitaev, Lukasz Kaiser, Anselm LevskayaICLR 2020 · 2,878 citations
- Sharpness-aware Minimization for Efficiently Improving GeneralizationPierre Foret, Ariel Kleiner, Hossein Mobahi, Behnam NeyshaburICLR 2021 · 1,861 citations
Related papers
- Ultra-Low Precision 4-bit Training of Deep Neural NetworksXiao Sun, Naigang Wang, Chia-Yu Chen, Jiamin Ni et al.NeurIPS 2020 · 227 citations
- FP4 All the Way: Fully Quantized Training of Large Language ModelsBrian Chmiel, Maxim Fishman, Ron Banner, Daniel SoudryNeurIPS 2025 · 9 citations
- Attn-QAT: 4-Bit Attention With Quantization-Aware TrainingPeiyuan Zhang, Matthew Noto, Wenxuan Tan, Chengquan Jiang et al.ICML 2026 · 2 citations
- Bit-by-Bit: Progressive QAT Strategy with Outlier Channel Splitting for Stable Low-Bit LLMsBinxing Xu, Hao Gu, Lujun Li, Hao Wang et al.ACL 2026 · 2 citations
- Bridging the Gap Between Promise and Performance for Microscaling FP4 QuantizationVage Egiazarian, Roberto L. Castro, Denis Kuznedelev, Andrei Panferov et al.ICLR 2026 · 38 citations
