TetraJet-v2: Accurate NVFP4 Training for Large Language Models with Oscillation Suppression and Outlier Control
Yuxiang Chen, Yifan Liu, Xiaoming Xu, Pengle Zhang, Michael Beyer, Martin Rapp, Jun Zhu, Jianfei Chen
Abstract
Large Language Models (LLMs) training is prohibitively expensive, driving interest in lowprecision fully-quantized training (FQT). While novel 4-bit formats like NVFP4 offer substantial efficiency gains, achieving near-lossless training at such low precision remains challenging. We introduce TetraJet-v2, an end-to-end 4-bit FQT method that leverages NVFP4 for activations, weights and gradients in all linear layers. We identify two critical issues hindering low-precision LLM training: weight oscillation and outliers. To address them, we propose: 1) an unbiased double-block quantization method for NVFP4 linear layers with practically optimal convergence in LLM training, 2) OsciReset, the first effective algorithm to suppress LLMs' weight oscillation bottleneck, and 3) OutControl, a mix-precision algorithm to retain outlier accuracy. TetraJet-v2 outperforms prior methods on FP4 pre-training for LLMs across models up to 370M parameters trained up to 212B tokens, reducing the performance gap to BF16 by an average of 51.3% while enabling 1.67× end-to-end speedup over FP8. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3d2b65bf-7245-420d-b104-516c52eaa2f1Cited by top-tier papers2
- Quartet: Native FP4 Training Can Be Optimal for Large Language ModelsRoberto L. Castro, Andrei Panferov, Rush Tabesh, Oliver Sieberling et al.NeurIPS 2025 · 38 citations
- Quartet II: Accurate LLM Pre-Training in NVFP4 by Improved Unbiased Gradient EstimationAndrei Panferov, Erik Schultheis, Soroush Tabesh, Dan AlistarhICML 2026 · 11 citations
Builds on19
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- WinoGrande: An Adversarial Winograd Schema Challenge at ScaleKeisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, Yejin ChoiAAAI 2020 · 3,037 citations
- PIQA: Reasoning about Physical Commonsense in Natural LanguageYonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao et al.AAAI 2020 · 2,916 citations
- FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningTri DaoICLR 2024 · 2,600 citations
- Efficient Streaming Language Models with Attention SinksGuangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han et al.ICLR 2024 · 1,714 citations
Related papers
- Optimizing Large Language Model Training Using FP4 QuantizationRuizhe Wang, Yeyun Gong, Xiao Liu, Guoshuai Zhao et al.ICML 2025
- Oscillation-Reduced MXFP4 Training for Vision TransformersYuxiang Chen, Haocheng Xi, Jun Zhu, Jianfei ChenICML 2025
- FP4 All the Way: Fully Quantized Training of Large Language ModelsBrian Chmiel, Maxim Fishman, Ron Banner, Daniel SoudryNeurIPS 2025 · 9 citations
- PTQ1.61: Push the Real Limit of Extremely Low-Bit Post-Training Quantization Methods for Large Language ModelsJiaqi Zhao, Miao Zhang, Ming Wang, Yuzhang Shang et al.ACL 2025 · 7 citations
- Metis: Training LLMs with FP4 QuantizationHengjie Cao, Mengyi Chen, Yifeng Yang, Fang Dong et al.ICLR 2026 · 10 citations
