MOSS: Efficient and Accurate FP8 LLM Training with Microscaling and Automatic Scaling
Yu Zhang, Huiling Zhen, Mingxuan Yuan, Bei Yu
Abstract
Training large language models with FP8 formats offers significant efficiency gains. However, the reduced numerical precision of FP8 poses challenges for stable and accurate training. Current frameworks preserve training performance using mixed-granularity quantization, i.e., applying per-group quantization for activations and per-tensor/block quantization for weights. While effective, per-group quantization requires scaling along the inner dimension of matrix multiplication, introducing additional dequantization overhead. Moreover, these frameworks often rely on just-in-time scaling to dynamically adjust scaling factors based on the current data distribution. However, this online quantization is inefficient for FP8 training, as it involves multiple memory reads and writes that negate the performance benefits of FP8. To overcome these limitations, we propose MOSS, a novel FP8 training framework that ensures both efficiency and numerical stability. MOSS introduces two key innovations: (1) a two-level microscaling strategy for quantizing sensitive activations, which balances precision and dequantization cost by combining a high-precision global scale with compact, power-of-two local scales; and (2) automatic scaling for weights in linear layers, which eliminates the need for costly max-reduction operations by predicting and adjusting scaling factors during training. Leveraging these techniques, MOSS enables efficient FP8 training of a 7B parameter model, achieving performance comparable to the BF16 baseline while achieving up to 34% higher training throughput.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ffceca51-8708-4515-a54b-81bf37b5b8d8Builds on4
- MAmmoTH: Building Math Generalist Models through Hybrid Instruction TuningXiang Yue, Xingwei Qu, Ge Zhang, Yao Fu et al.ICLR 2024 · 558 citations
- NumGLUE: A Suite of Fundamental yet Challenging Mathematical Reasoning TasksSwaroop Mishra, Arindam Mitra, Neeraj Varshney, Bhavdeep Singh Sachdeva et al.ACL 2022 · 138 citations
- Unit Scaling: Out-of-the-Box Low-Precision TrainingCharlie Blake, Douglas Orr, Carlo LuschiICML 2023 · 15 citations
- Scaling FP8 training to trillion-token LLMsMaxim Fishman, Brian Chmiel, Ron Banner, Daniel SoudryICLR 2025 · 1 citation
Related papers
- COAT: Compressing Optimizer states and Activations for Memory-Efficient FP8 TrainingHaocheng Xi, Han Cai, Ligeng Zhu, Yao Lu et al.ICLR 2025
- SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language ModelsGuangxuan Xiao, Ji Lin, Mickaël Seznec, Hao Wu et al.ICML 2023 · 1,493 citations
- Towards Fully FP8 GEMM LLM Training at ScaleAlejandro Hernández-Cano, Dhia Garbaya, Imanol Schlag, Martin JaggiNeurIPS 2025 · 13 citations
- Optimizing Large Language Model Training Using FP4 QuantizationRuizhe Wang, Yeyun Gong, Xiao Liu, Guoshuai Zhao et al.ICML 2025
- µnit Scaling: Simple and Scalable FP8 LLM TrainingSaaketh Narayan, Abhay Gupta, Mansheej Paul, Davis W. BlalockICML 2025
