Oscillation-Reduced MXFP4 Training for Vision Transformers
Yuxiang Chen, Haocheng Xi, Jun Zhu, Jianfei Chen
Abstract
Pre-training Transformers in FP4 precision is becoming a promising approach to gain substantial speedup, but it comes with a considerable loss of accuracy. Microscaling (MX) data format provides a fine-grained per-group quantization method to improve the representation ability of the FP4 format and is supported by the nextgeneration Blackwell GPU architecture. However, training with MXFP4 data format still results in significant degradation and there is a lack of systematic research on the reason. In this work, we propose a novel training method TetraJet for a more accurate FP4 training. We comprehensively evaluate all of the quantizers involved in the training, and identify the weight oscillation problem in the forward pass as the main source of the degradation in MXFP4 training. Therefore, we introduce two novel methods, EMA Quantizer (Q-EMA) and Adaptive Ramping Optimizer (Q-Ramping), to resolve the oscillation problem. Extensive experiments on Vision Transformers demonstrate that TetraJet consistently outperforms the existing 4-bit training methods, and Q-EMA & Q-Ramping can provide additional enhancement by effectively reducing oscillation. We decreased the accuracy degradation by more than 50% compared to the baseline, and can even achieve competitive performance compared to full precision training. The codes are available at https://github.com/thu-ml/ TetraJet-MXFP4Training .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 714829f9-ba5d-439c-84a9-4b2c18e96924Cited by top-tier papers4
- INT vs. FP: A Comprehensive Study of Fine-Grained Low-bit Quantization FormatsMengzhao Chen, Meng Wu, Hui Jin, Zhihang Yuan et al.ICML 2026 · 21 citations
- MicroMix: Efficient Mixed-Precision Quantization with Microscaling Formats for Large Language ModelsWenyuan Liu, Haoqian Meng, Yilun Luo, Peng Zhang et al.ICLR 2026 · 12 citations
- TetraJet-v2: Accurate NVFP4 Training for Large Language Models with Oscillation Suppression and Outlier ControlYuxiang Chen, Yifan Liu, Xiaoming Xu, Pengle Zhang et al.ICML 2026 · 11 citations
- Benchmarking Post-Training Quantization of Large Language Models under Microscaling Floating Point FormatsManyi Zhang, Ji-Fu Li, Zhongao Sun, Haoli Bai et al.ACL 2026 · 2 citations
Builds on10
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
- Ultra-Low Precision 4-bit Training of Deep Neural NetworksXiao Sun, Naigang Wang, Chia-Yu Chen, Jiamin Ni et al.NeurIPS 2020 · 227 citations
- Overcoming Oscillations in Quantization-Aware TrainingMarkus Nagel, Marios Fournarakis, Yelysei Bondarenko, Tijmen BlankevoortICML 2022 · 163 citations
- Stable and low-precision training for large-scale vision-language modelsMitchell Wortsman, Tim Dettmers, Luke Zettlemoyer, Ari Morcos et al.NeurIPS 2023 · 101 citations
Related papers
- Bridging the Gap Between Promise and Performance for Microscaling FP4 QuantizationVage Egiazarian, Roberto L. Castro, Denis Kuznedelev, Andrei Panferov et al.ICLR 2026 · 38 citations
- Quartet II: Accurate LLM Pre-Training in NVFP4 by Improved Unbiased Gradient EstimationAndrei Panferov, Erik Schultheis, Soroush Tabesh, Dan AlistarhICML 2026 · 11 citations
- Block Rotation is All You Need for MXFP4 QuantizationYuantian Shao, Peisong Wang, Yuanteng Chen, Chang Xu et al.ICML 2026 · 16 citations
- Quartet: Native FP4 Training Can Be Optimal for Large Language ModelsRoberto L. Castro, Andrei Panferov, Rush Tabesh, Oliver Sieberling et al.NeurIPS 2025 · 38 citations
- Is Finer Better? The Limits of Microscaling Formats in Large Language ModelsAndrea Fasoli, Monodeep Kar, Chi-Chun (Charlie) Liu, Swagath Venkataramani et al.ICLR 2026 · 7 citations
