Scaling FP8 training to trillion-token LLMs
Maxim Fishman, Brian Chmiel, Ron Banner, Daniel Soudry
Abstract
We train, for the first time, large language models using FP8 precision on datasets up to 2 trillion tokens -- a 20-fold increase over previous limits. Through these extended training runs, we uncover critical instabilities in FP8 training that were not observable in earlier works with shorter durations. We trace these instabilities to outlier amplification by the SwiGLU activation function. Interestingly, we show, both analytically and empirically, that this amplification happens only over prolonged training periods, and link it to a SwiGLU weight alignment process. To address this newly identified issue, we introduce Smooth-SwiGLU, a novel modification that ensures stable FP8 training without altering function behavior. We also demonstrate, for the first time, FP8 quantization of both Adam optimizer moments. Combining these innovations, we successfully train a 7B parameter model using FP8 precision on 256 Intel Gaudi2 accelerators, achieving on-par results with the BF16 baseline while delivering up to a throughput improvement. A reference implementation is supplied in https://github.com/Anonymous1252022/Megatron-DeepSpeed.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers30
- Quartet: Native FP4 Training Can Be Optimal for Large Language ModelsRoberto L. Castro, Andrei Panferov, Rush Tabesh, Oliver Sieberling et al.NeurIPS 2025 · 38 citations
- Understanding and Minimising Outlier Features in Transformer TrainingBobby He, Lorenzo Noci, Daniele Paliotta, Imanol Schlag et al.NeurIPS 2024 · 27 citations
- FP64 is All You Need: Rethinking Failure Modes in Physics-Informed Neural NetworksChenhui Xu, Dancheng Liu, Amir Nassereldine, Jinjun XiongNeurIPS 2025 · 18 citations
- Scaling Law for Quantization-Aware TrainingMengzhao Chen, Chaoyi Zhang, Jing Liu, Zeng et al.ICML 2026 · 16 citations
- HALO: Hadamard-Assisted Lower-Precision Optimization for LLMsSaleh Ashkboos, Mahdi Nikdan, Rush Tabesh, Roberto L. Castro et al.NeurIPS 2025 · 16 citations
Builds on3
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Quantizable Transformers: Removing Outliers by Helping Attention Heads Do NothingYelysei Bondarenko, Markus Nagel, Tijmen BlankevoortNeurIPS 2023 · 196 citations
- SpinQuant: LLM Quantization with Learned RotationsZechun Liu, Changsheng Zhao, Igor Fedorov, Bilge Soran et al.ICLR 2025
Related papers
- Towards Fully FP8 GEMM LLM Training at ScaleAlejandro Hernández-Cano, Dhia Garbaya, Imanol Schlag, Martin JaggiNeurIPS 2025 · 13 citations
- SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language ModelsGuangxuan Xiao, Ji Lin, Mickaël Seznec, Hao Wu et al.ICML 2023 · 1,493 citations
- TWEO: Transformers Without Extreme Outliers Enables FP8 Training And Quantization For DummiesGuang Liang, Jie Shao, Ningyuan Tang, Xinyao Liu et al.CVPR 2026 · 5 citations
- Stable and low-precision training for large-scale vision-language modelsMitchell Wortsman, Tim Dettmers, Luke Zettlemoyer, Ari Morcos et al.NeurIPS 2023 · 101 citations
- MOSS: Efficient and Accurate FP8 LLM Training with Microscaling and Automatic ScalingYu Zhang, Huiling Zhen, Mingxuan Yuan, Bei YuICLR 2026 · 3 citations
