1-bit Adam: Communication Efficient Large-Scale Training with Adam's Convergence Speed
Hanlin Tang, Shaoduo Gan, Ammar Ahmad Awan, Samyam Rajbhandari, Conglong Li, Xiangru Lian, Ji Liu, Ce Zhang, Yuxiong He
Abstract
Scalable training of large models (like BERT and GPT-3) requires careful optimization rooted in model design, architecture, and system capabilities. From a system standpoint, communication has become a major bottleneck, especially on commodity systems with standard TCP interconnects that offer limited network bandwidth. Communication compression is an important technique to reduce training time on such systems. One of the most effective methods is error-compensated compression, which offers robust convergence speed even under 1-bit compression. However, state-of-the-art error compensation techniques only work with basic optimizers like SGD and Momentum SGD, which are linearly dependent on the gradients. They do not work with non-linear gradient-based optimizers like Adam, which offer state-of-the-art convergence efficiency and accuracy for models like BERT. In this paper, we propose 1-bit Adam that reduces the communication volume by up to 5×, offers much better scalability, and provides the same sample-wise convergence speed as uncompressed Adam. Our key finding is that Adam's variance (non-linear term) becomes stable during training, hence we can run Adam in the beginning (warmup phase) and use it as a precondition for Momentum SGD during the rest of the training (compression phase). Experiments on up to 256 GPUs show that 1-bit Adam enables up to 3.3× higher throughput for BERT-Large pre-training and up to 2.9× higher throughput for SQuAD fine-tuning. In addition, we provide theoretical analysis for our proposed work.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8cb80a59-cd4d-4055-89ec-2fb4507b7902Cited by top-tier papers13
- Communication-Efficient Adaptive Federated LearningYujia Wang, Lu Lin, Jinghui ChenICML 2022 · 101 citations
- CocktailSGD: Fine-tuning Foundation Models over 500Mbps NetworksJue Wang, Yucheng Lu, Binhang Yuan, Beidi Chen et al.ICML 2023 · 60 citations
- Smart-Infinity: Fast Large Language Model Training using Near-Storage Processing on a Real SystemHongsun Jang, Jaeyong Song, Jaewon Jung, Jaeyoung Park et al.HPCA 2024 · 26 citations
- CO2: Efficient Distributed Training with Full Communication-Computation OverlapWeigao Sun, Zhen Qin, Weixuan Sun, Shidi Li et al.ICLR 2024 · 17 citations
- LLM.265: Video Codecs are Secretly Tensor CodecsCeyu Xu, Yongji Wu, Xinyu Yang, Beidi Chen et al.MICRO 2025 · 13 citations
Builds on5
- Decentralized Deep Learning with Arbitrary Communication CompressionAnastasia Koloskova, Tao Lin, Sebastian U. Stich, Martin JaggiICLR 2020 · 263 citations
- Linearly Converging Error Compensated SGDEduard Gorbunov, Dmitry Kovalev, Dmitry Makarenko, Peter RichtárikNeurIPS 2020 · 90 citations
- A new regret analysis for Adam-type algorithmsAhmet Alacaoglu, Yura Malitsky, Panayotis Mertikopoulos, Volkan CevherICML 2020 · 50 citations
- CSER: Communication-efficient SGD with Error ResetCong Xie, Shuai Zheng, Oluwasanmi Koyejo, Indranil Gupta et al.NeurIPS 2020 · 50 citations
- Adam with Bandit Sampling for Deep LearningRui Liu, Tianyi Wu, Barzan MozafariNeurIPS 2020 · 30 citations
Related papers
- Maximizing Communication Efficiency for Large-scale Training via 0/1 AdamYucheng Lu, Conglong Li, Minjia Zhang, Christopher De Sa et al.ICLR 2023 · 4 citations
- SLAMB: Accelerated Large Batch Training with Sparse CommunicationHang Xu, Wenxuan Zhang, Jiawei Fei, Yuzhe Wu et al.ICML 2023 · 7 citations
- Birder: Communication-Efficient 1-bit Adaptive Optimizer for Practical Distributed DNN TrainingHanyang Peng, Shuang Qin, Yue Yu, Jin Wang et al.NeurIPS 2023 · 5 citations
- MicroAdam: Accurate Adaptive Optimization with Low Space Overhead and Provable ConvergenceIonut-Vlad Modoranu, Mher Safaryan, Grigory Malinovsky, Eldar Kurtic et al.NeurIPS 2024 · 32 citations
- SDP4Bit: Toward 4-bit Communication Quantization in Sharded Data Parallelism for LLM TrainingJinda Jia, Cong Xie, Hanlin Lu, Daoce Wang et al.NeurIPS 2024 · 23 citations
