Scaling FP8 training to trillion-token LLMs
Maxim Fishman, Brian Chmiel, Ron Banner, Daniel Soudry
摘要
We train, for the first time, large language models using FP8 precision on datasets up to 2 trillion tokens -- a 20-fold increase over previous limits. Through these extended training runs, we uncover critical instabilities in FP8 training that were not observable in earlier works with shorter durations. We trace these instabilities to outlier amplification by the SwiGLU activation function. Interestingly, we show, both analytically and empirically, that this amplification happens only over prolonged training periods, and link it to a SwiGLU weight alignment process. To address this newly identified issue, we introduce Smooth-SwiGLU, a novel modification that ensures stable FP8 training without altering function behavior. We also demonstrate, for the first time, FP8 quantization of both Adam optimizer moments. Combining these innovations, we successfully train a 7B parameter model using FP8 precision on 256 Intel Gaudi2 accelerators, achieving on-par results with the BF16 baseline while delivering up to a throughput improvement. A reference implementation is supplied in https://github.com/Anonymous1252022/Megatron-DeepSpeed.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper30
- Quartet: Native FP4 Training Can Be Optimal for Large Language ModelsRoberto L. Castro, Andrei Panferov, Rush Tabesh, Oliver Sieberling 等NeurIPS 2025 · 被引用 38 次
- Understanding and Minimising Outlier Features in Transformer TrainingBobby He, Lorenzo Noci, Daniele Paliotta, Imanol Schlag 等NeurIPS 2024 · 被引用 27 次
- FP64 is All You Need: Rethinking Failure Modes in Physics-Informed Neural NetworksChenhui Xu, Dancheng Liu, Amir Nassereldine, Jinjun XiongNeurIPS 2025 · 被引用 18 次
- Scaling Law for Quantization-Aware TrainingMengzhao Chen, Chaoyi Zhang, Jing Liu, Zeng 等ICML 2026 · 被引用 16 次
- HALO: Hadamard-Assisted Lower-Precision Optimization for LLMsSaleh Ashkboos, Mahdi Nikdan, Rush Tabesh, Roberto L. Castro 等NeurIPS 2025 · 被引用 16 次
它引用的顶会 Paper3
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Quantizable Transformers: Removing Outliers by Helping Attention Heads Do NothingYelysei Bondarenko, Markus Nagel, Tijmen BlankevoortNeurIPS 2023 · 被引用 196 次
- SpinQuant: LLM Quantization with Learned RotationsZechun Liu, Changsheng Zhao, Igor Fedorov, Bilge Soran 等ICLR 2025
相关 Paper
- Towards Fully FP8 GEMM LLM Training at ScaleAlejandro Hernández-Cano, Dhia Garbaya, Imanol Schlag, Martin JaggiNeurIPS 2025 · 被引用 13 次
- SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language ModelsGuangxuan Xiao, Ji Lin, Mickaël Seznec, Hao Wu 等ICML 2023 · 被引用 1,493 次
- TWEO: Transformers Without Extreme Outliers Enables FP8 Training And Quantization For DummiesGuang Liang, Jie Shao, Ningyuan Tang, Xinyao Liu 等CVPR 2026 · 被引用 5 次
- Stable and low-precision training for large-scale vision-language modelsMitchell Wortsman, Tim Dettmers, Luke Zettlemoyer, Ari Morcos 等NeurIPS 2023 · 被引用 101 次
- MOSS: Efficient and Accurate FP8 LLM Training with Microscaling and Automatic ScalingYu Zhang, Huiling Zhen, Mingxuan Yuan, Bei YuICLR 2026 · 被引用 3 次
