Unit Scaling: Out-of-the-Box Low-Precision Training
Charlie Blake, Douglas Orr, Carlo Luschi
Abstract
We present unit scaling, a paradigm for designing deep learning models that simplifies the use of low-precision number formats. Training in FP16 or the recently proposed FP8 formats offers substantial efficiency gains, but can lack sufficient range for out-of-the-box training. Unit scaling addresses this by introducing a principled approach to model numerics: seeking unit variance of all weights, activations and gradients at initialisation. Unlike alternative methods, this approach neither requires multiple training runs to find a suitable scale nor has significant computational overhead. We demonstrate the efficacy of unit scaling across a range of models and optimisers. We further show that existing models can be adapted to be unit-scaled, training BERT-Large in FP16 and then FP8 with no degradation in accuracy.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6c165952-a50d-464e-8a78-d6746267746dCited by top-tier papers9
- Stable and low-precision training for large-scale vision-language modelsMitchell Wortsman, Tim Dettmers, Luke Zettlemoyer, Ari Morcos et al.NeurIPS 2023 · 101 citations
- Scaling Exponents Across Parameterizations and OptimizersKatie E. Everett, Lechao Xiao, Mitchell Wortsman, Alexander A. Alemi et al.ICML 2024 · 59 citations
- MOSS: Efficient and Accurate FP8 LLM Training with Microscaling and Automatic ScalingYu Zhang, Huiling Zhen, Mingxuan Yuan, Bei YuICLR 2026 · 3 citations
- Rank-Aware Spectral Bounds on Attention Logits for Stable Low-Precision TrainingSeyed Morteza EmadiICML 2026 · 1 citation
- FALQON: Accelerating LoRA Fine-tuning with Low-Bit Floating-Point ArithmeticKanghyun Choi, Hyeyoon Lee, Sunjong Park, Dain Kwon et al.NeurIPS 2025 · 1 citation
Builds on7
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language ModelsGuangxuan Xiao, Ji Lin, Mickaël Seznec, Hao Wu et al.ICML 2023 · 1,493 citations
- Large Batch Optimization for Deep Learning: Training BERT in 76 minutesYang You, Jing Li, Sashank J. Reddi, Jonathan Hseu et al.ICLR 2020 · 1,170 citations
- High-Performance Large-Scale Image Recognition Without NormalizationAndy Brock, Soham De, Samuel L. Smith, Karen SimonyanICML 2021 · 613 citations
- Understanding and Overcoming the Challenges of Efficient Transformer QuantizationYelysei Bondarenko, Markus Nagel, Tijmen BlankevoortEMNLP 2021 · 74 citations
Related papers
- u-μP: The Unit-Scaled Maximal Update ParametrizationCharlie Blake, Constantin Eichenberg, Josef Dean, Lukas Balles et al.ICLR 2025
- µnit Scaling: Simple and Scalable FP8 LLM TrainingSaaketh Narayan, Abhay Gupta, Mansheej Paul, Davis W. BlalockICML 2025
- FP8 Quantization: The Power of the ExponentAndrey Kuzmin, Mart van Baalen, Yuwei Ren, Markus Nagel et al.NeurIPS 2022 · 154 citations
- Shifted and Squeezed 8-bit Floating Point format for Low-Precision Training of Deep Neural NetworksLéopold Cambier, Anahita Bhiwandiwalla, Ting Gong, Oguz H. Elibol et al.ICLR 2020 · 53 citations
- Towards Fully FP8 GEMM LLM Training at ScaleAlejandro Hernández-Cano, Dhia Garbaya, Imanol Schlag, Martin JaggiNeurIPS 2025 · 13 citations
