Collage: Light-Weight Low-Precision Strategy for LLM Training
Tao Yu, Gaurav Gupta, Karthick Gopalswamy, Amith R. Mamidala, Hao Zhou, Jeffrey Huynh, Youngsuk Park, Ron Diamant, Anoop Deoras, Luke Huan
摘要
Large models training is plagued by the intense compute cost and limited hardware memory. A practical solution is low-precision representation but is troubled by loss in numerical accuracy and unstable training rendering the model less useful. We argue that low-precision floating points can perform well provided the error is properly compensated at the critical locations in the training process. We propose Collage which utilizes multi-component float representation in low-precision to accurately perform operations with numerical errors accounted. To understand the impact of imprecision to training, we propose a simple and novel metric which tracks the lost information during training as well as differentiates various precision strategies. Our method works with commonly used low-precision such as half-precision (-bit floating points) and can be naturally extended to work with even lower precision such as -bit. Experimental results show that pre-training using Collage removes the requirement of using -bit floating-point copies of the model and attains similar/better training performance compared to -bit mixed-precision strategy, with up to speedup and to less memory usage in practice.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- A Convergence Analysis of Adaptive Optimizers under Floating-point QuantizationXuan Tang, Jichu Li, Difan ZouICLR 2026 · 被引用 8 次
- Low Precision Streaming PCASanjoy Dasgupta, Syamantak Kumar, Shourya Pandey, Purnamrita SarkarNeurIPS 2025 · 被引用 3 次
- Log-Normal Multiplicative Dynamics for Stable Low-Precision Deep LearningKeigo Nishida, Eren Mehmet KIRAL, Kenichi Bannai, Mohammad Emtiyaz Khan 等ICML 2026 · 被引用 2 次
- ECO: Quantized Training without Full-Precision Master WeightsMahdi Nikdan, Amir Zandieh, Dan Alistarh, Vahab MirrokniICML 2026 · 被引用 2 次
- ELMO : Efficiency via Low-precision and Peak Memory Optimization in Large Output SpacesJinbin Zhang, Nasib Ullah, Erik Schultheis, Rohit BabbarICML 2025
它引用的顶会 Paper11
- The case for 4-bit precision: k-bit Inference Scaling LawsTim Dettmers, Luke ZettlemoyerICML 2023 · 被引用 315 次
- Learning N: M Fine-grained Structured Sparse Neural Networks From ScratchAojun Zhou, Yukun Ma, Junnan Zhu, Jianbo Liu 等ICLR 2021 · 被引用 301 次
- Structured Pruning Learns Compact and Accurate ModelsMengzhou Xia, Zexuan Zhong, Danqi ChenACL 2022 · 被引用 236 次
- Width & Depth Pruning for Vision TransformersFang Yu, Kun Huang, Meng Wang, Yuan Cheng 等AAAI 2022 · 被引用 159 次
- Stable and low-precision training for large-scale vision-language modelsMitchell Wortsman, Tim Dettmers, Luke Zettlemoyer, Ari Morcos 等NeurIPS 2023 · 被引用 101 次
相关 Paper
- Shifted and Squeezed 8-bit Floating Point format for Low-Precision Training of Deep Neural NetworksLéopold Cambier, Anahita Bhiwandiwalla, Ting Gong, Oguz H. Elibol 等ICLR 2020 · 被引用 53 次
- Multi-Precision Policy Enforced Training (MuPPET) : A Precision-Switching Strategy for Quantised Fixed-Point Training of CNNsAditya Rajagopal, Diederik Adriaan Vink, Stylianos I. Venieris, Christos-Savvas BouganisICML 2020 · 被引用 17 次
- M+Adam: Low-Precision Training via Additive–Multiplicative OptimizationXiaoyuan Liang, Sebastian Loeschcke, Mads Toftrup, Anima AnandkumarICML 2026 · 被引用 1 次
- Post-training Quantization with Multiple Points: Mixed Precision without Mixed PrecisionXingchao Liu, Mao Ye, Dengyong Zhou, Qiang LiuAAAI 2021 · 被引用 54 次
- Training Quantized Neural Networks With a Full-Precision Auxiliary ModuleBohan Zhuang, Lingqiao Liu, Mingkui Tan, Chunhua Shen 等CVPR 2020
