Quartet: Native FP4 Training Can Be Optimal for Large Language Models
Roberto L. Castro, Andrei Panferov, Rush Tabesh, Oliver Sieberling, Jiale Chen, Mahdi Nikdan, Saleh Ashkboos, Dan Alistarh
摘要
Training large language models (LLMs) models directly in low-precision offers a way to address computational costs by improving both throughput and energy efficiency. For those purposes, NVIDIA's recent Blackwell architecture facilitates very low-precision operations using FP4 variants. Yet, current algorithms for training LLMs in FP4 precision face significant accuracy degradation and often rely on mixed-precision fallbacks. In this paper, we investigate hardware-supported FP4 training and introduce a new approach for accurate, end-to-end FP4 training with all the major computations (i.e., linear layers) in low precision. Through extensive evaluations on Llama-type models, we reveal a new low-precision scaling law that quantifies performance trade-offs across bit-widths and training setups. Guided by this investigation, we design an "optimal" technique in terms of accuracyvs-computation, called Quartet. We implement Quartet using optimized CUDA kernels tailored for Blackwell, demonstrating that fully FP4-based training is a competitive alternative to FP16 half-precision and to FP8 training. Our code is available at https://github.com/IST-DASLab/Quartet.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper12
- INT vs. FP: A Comprehensive Study of Fine-Grained Low-bit Quantization FormatsMengzhao Chen, Meng Wu, Hui Jin, Zhihang Yuan 等ICML 2026 · 被引用 21 次
- QeRL: Beyond Efficiency - Quantization-enhanced Reinforcement Learning for LLMsWei Huang, Yi Ge, Shuai Yang, Yicheng Xiao 等ICLR 2026 · 被引用 19 次
- Quartet II: Accurate LLM Pre-Training in NVFP4 by Improved Unbiased Gradient EstimationAndrei Panferov, Erik Schultheis, Soroush Tabesh, Dan AlistarhICML 2026 · 被引用 11 次
- TetraJet-v2: Accurate NVFP4 Training for Large Language Models with Oscillation Suppression and Outlier ControlYuxiang Chen, Yifan Liu, Xiaoming Xu, Pengle Zhang 等ICML 2026 · 被引用 11 次
- Unveiling the Potential of Quantization with MXFP4: Strategies for Quantization Error ReductionJatin Chhugani, Geonhwa Jeong, Bor-Yiing Su, Yunjie Pan 等ICML 2026 · 被引用 6 次
它引用的顶会 Paper20
- WinoGrande: An Adversarial Winograd Schema Challenge at ScaleKeisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, Yejin ChoiAAAI 2020 · 被引用 3,037 次
- Learned Step Size quantizationSteven K. Esser, Jeffrey L. McKinstry, Deepika Bablani, Rathinakumar Appuswamy 等ICLR 2020 · 被引用 1,037 次
- QuaRot: Outlier-Free 4-Bit Inference in Rotated LLMsSaleh Ashkboos, Amirkeivan Mohtashami, Maximilian L. Croci, Bo Li 等NeurIPS 2024 · 被引用 723 次
- QuIP: 2-Bit Quantization of Large Language Models With GuaranteesJerry Chee, Yaohui Cai, Volodymyr Kuleshov, Christopher De SaNeurIPS 2023 · 被引用 503 次
- Ultra-Low Precision 4-bit Training of Deep Neural NetworksXiao Sun, Naigang Wang, Chia-Yu Chen, Jiamin Ni 等NeurIPS 2020 · 被引用 227 次
相关 Paper
- MicroMix: Efficient Mixed-Precision Quantization with Microscaling Formats for Large Language ModelsWenyuan Liu, Haoqian Meng, Yilun Luo, Peng Zhang 等ICLR 2026 · 被引用 12 次
- Optimizing Large Language Model Training Using FP4 QuantizationRuizhe Wang, Yeyun Gong, Xiao Liu, Guoshuai Zhao 等ICML 2025
- COMET: Towards Practical W4A4KV4 LLMs ServingLian Liu, Long Cheng, Haimeng Ren, Zhaohui Xu 等ASPLOS 2025 · 被引用 5 次
- FP4 All the Way: Fully Quantized Training of Large Language ModelsBrian Chmiel, Maxim Fishman, Ron Banner, Daniel SoudryNeurIPS 2025 · 被引用 9 次
- any4: Learned 4-bit Numeric Representation for LLMsMostafa Elhoushi, Jeff JohnsonICML 2025
