Towards Fully FP8 GEMM LLM Training at Scale
Alejandro Hernández-Cano, Dhia Garbaya, Imanol Schlag, Martin Jaggi
Abstract
Despite the significant potential of FP8 data formats for large language model (LLM) pre-training, their adoption has been limited due to challenges in maintaining stability at scale. Existing approaches often rely on suboptimal fine-grained FP8 kernels or fall back to higher-precision matrix multiplications (GEMMs) in sensitive components, such as attention projections, compromising potential throughput gains. We introduce a new class of LLM architectures that, for the first time, support FP8 computation for all GEMMs within transformer blocks during both forward and backward passes. This enables unprecedented throughput gains, particularly at scale, while matching the downstream performance of standard BF16 training. Our architecture design reduces large outlier activations, promoting stable long-term FP8 training. In addition, we identify key metrics to monitor low-precision training and predict potential future divergences.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers6
- Apertus: Democratizing Open and Compliant LLMs for Global Language EnvironmentsAlejandro Hernández-Cano, Alexander Hägele, Allen Hao Huang, Angelika Romanou et al.ACL 2026 · 51 citations
- Quartet II: Accurate LLM Pre-Training in NVFP4 by Improved Unbiased Gradient EstimationAndrei Panferov, Erik Schultheis, Soroush Tabesh, Dan AlistarhICML 2026 · 11 citations
- TWEO: Transformers Without Extreme Outliers Enables FP8 Training And Quantization For DummiesGuang Liang, Jie Shao, Ningyuan Tang, Xinyao Liu et al.CVPR 2026 · 5 citations
- Attn-QAT: 4-Bit Attention With Quantization-Aware TrainingPeiyuan Zhang, Matthew Noto, Wenxuan Tan, Chengquan Jiang et al.ICML 2026 · 2 citations
- Rank-Aware Spectral Bounds on Attention Logits for Stable Low-Precision TrainingSeyed Morteza EmadiICML 2026 · 1 citation
Builds on12
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Swin Transformer V2: Scaling Up Capacity and ResolutionZe Liu, Han Hu, Yutong Lin, Zhuliang Yao et al.CVPR 2022 · 2,138 citations
- GShard: Scaling Giant Models with Conditional Computation and Automatic ShardingDmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen et al.ICLR 2021 · 1,954 citations
- Going deeper with Image TransformersHugo Touvron, Matthieu Cord, Alexandre Sablayrolles, Gabriel Synnaeve et al.ICCV 2021 · 1,279 citations
- ZeRO: memory optimizations toward training trillion parameter modelsSamyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, Yuxiong HeSC 2020 · 852 citations
Related papers
- Scaling FP8 training to trillion-token LLMsMaxim Fishman, Brian Chmiel, Ron Banner, Daniel SoudryICLR 2025 · 1 citation
- HALO: Hadamard-Assisted Lower-Precision Optimization for LLMsSaleh Ashkboos, Mahdi Nikdan, Rush Tabesh, Roberto L. Castro et al.NeurIPS 2025 · 16 citations
- µnit Scaling: Simple and Scalable FP8 LLM TrainingSaaketh Narayan, Abhay Gupta, Mansheej Paul, Davis W. BlalockICML 2025
- Optimizing Large Language Model Training Using FP4 QuantizationRuizhe Wang, Yeyun Gong, Xiao Liu, Guoshuai Zhao et al.ICML 2025
- NestedFP: High-Performance, Memory-Efficient Dual-Precision Floating Point Support for LLMsHaeun Lee, Omin Kwon, Yeonhong Park, Jae W. LeeNeurIPS 2025 · 5 citations
