Large Batch Optimization for Deep Learning Using New Complete Layer-Wise Adaptive Rate Scaling
Zhouyuan Huo, Bin Gu, Heng Huang
Abstract
Training deep neural networks using a large batch size has shown promising results and benefits many real-world applications. Warmup is one of nontrivial techniques to stabilize the convergence of large batch training. However, warmup is an empirical method and it is still unknown whether there is a better algorithm with theoretical underpinnings. In this paper, we propose a novel Complete Layer-wise Adaptive Rate Scaling (CLARS) algorithm for large-batch training. We prove the convergence of our algorithm by introducing a new fine-grained analysis of gradient-based methods. Furthermore, the new analysis also helps to understand two other empirical tricks, layer-wise adaptive rate scaling and linear learning rate scaling. We conduct extensive experiments and demonstrate that the proposed algorithm outperforms gradual warmup technique by a large margin and defeats the convergence of the state-of-the-art large-batch optimizer in training advanced deep neural networks (ResNet, DenseNet, Mo-bileNet) on ImageNet dataset.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b49ce3f4-2e80-4f35-96d8-5d37e5902d5aCited by top-tier papers5
- How to Scale Your EMADan Busbridge, Jason Ramapuram, Pierre Ablin, Tatiana Likhomanenko et al.NeurIPS 2023 · 33 citations
- CAME: Confidence-guided Adaptive Memory Efficient OptimizationYang Luo, Xiaozhe Ren, Zangwei Zheng, Zhuo Jiang et al.ACL 2023 · 7 citations
- Asynchronous Distributed Bilevel OptimizationYang Jiao, Kai Yang, Tiancheng Wu, Dongjin Song et al.ICLR 2023 · 6 citations
- Integrally Pre-Trained Transformer Pyramid NetworksYunjie Tian, Lingxi Xie, Zhaozhi Wang, Longhui Wei et al.CVPR 2023
- AdaFisher: Adaptive Second Order Optimization via Fisher InformationDamien Martins Gomes, Yanlei Zhang, Eugene Belilovsky, Guy Wolf et al.ICLR 2025
Related papers
- Not All Layers Are Equal: A Layer-Wise Adaptive Approach Toward Large-Scale DNN TrainingYun-Yong Ko, Dongwon Lee, Sang-Wook KimWWW 2022 · 11 citations
- Large Batch Optimization for Deep Learning: Training BERT in 76 minutesYang You, Jing Li, Sashank J. Reddi, Jonathan Hseu et al.ICLR 2020 · 1,170 citations
- AdaScale SGD: A User-Friendly Algorithm for Distributed TrainingTyler B. Johnson, Pulkit Agrawal, Haijie Gu, Carlos GuestrinICML 2020 · 41 citations
- Concurrent Adversarial Learning for Large-Batch TrainingYong Liu, Xiangning Chen, Minhao Cheng, Cho-Jui Hsieh et al.ICLR 2022 · 14 citations
- High-Performance Large-Scale Image Recognition Without NormalizationAndy Brock, Soham De, Samuel L. Smith, Karen SimonyanICML 2021 · 613 citations
