Large Batch Optimization for Deep Learning Using New Complete Layer-Wise Adaptive Rate Scaling
Zhouyuan Huo, Bin Gu, Heng Huang
摘要
Training deep neural networks using a large batch size has shown promising results and benefits many real-world applications. Warmup is one of nontrivial techniques to stabilize the convergence of large batch training. However, warmup is an empirical method and it is still unknown whether there is a better algorithm with theoretical underpinnings. In this paper, we propose a novel Complete Layer-wise Adaptive Rate Scaling (CLARS) algorithm for large-batch training. We prove the convergence of our algorithm by introducing a new fine-grained analysis of gradient-based methods. Furthermore, the new analysis also helps to understand two other empirical tricks, layer-wise adaptive rate scaling and linear learning rate scaling. We conduct extensive experiments and demonstrate that the proposed algorithm outperforms gradual warmup technique by a large margin and defeats the convergence of the state-of-the-art large-batch optimizer in training advanced deep neural networks (ResNet, DenseNet, Mo-bileNet) on ImageNet dataset.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- How to Scale Your EMADan Busbridge, Jason Ramapuram, Pierre Ablin, Tatiana Likhomanenko 等NeurIPS 2023 · 被引用 33 次
- CAME: Confidence-guided Adaptive Memory Efficient OptimizationYang Luo, Xiaozhe Ren, Zangwei Zheng, Zhuo Jiang 等ACL 2023 · 被引用 7 次
- Asynchronous Distributed Bilevel OptimizationYang Jiao, Kai Yang, Tiancheng Wu, Dongjin Song 等ICLR 2023 · 被引用 6 次
- Integrally Pre-Trained Transformer Pyramid NetworksYunjie Tian, Lingxi Xie, Zhaozhi Wang, Longhui Wei 等CVPR 2023
- AdaFisher: Adaptive Second Order Optimization via Fisher InformationDamien Martins Gomes, Yanlei Zhang, Eugene Belilovsky, Guy Wolf 等ICLR 2025
相关 Paper
- Not All Layers Are Equal: A Layer-Wise Adaptive Approach Toward Large-Scale DNN TrainingYun-Yong Ko, Dongwon Lee, Sang-Wook KimWWW 2022 · 被引用 11 次
- Large Batch Optimization for Deep Learning: Training BERT in 76 minutesYang You, Jing Li, Sashank J. Reddi, Jonathan Hseu 等ICLR 2020 · 被引用 1,170 次
- AdaScale SGD: A User-Friendly Algorithm for Distributed TrainingTyler B. Johnson, Pulkit Agrawal, Haijie Gu, Carlos GuestrinICML 2020 · 被引用 41 次
- Concurrent Adversarial Learning for Large-Batch TrainingYong Liu, Xiangning Chen, Minhao Cheng, Cho-Jui Hsieh 等ICLR 2022 · 被引用 14 次
- High-Performance Large-Scale Image Recognition Without NormalizationAndy Brock, Soham De, Samuel L. Smith, Karen SimonyanICML 2021 · 被引用 613 次
