Not All Layers Are Equal: A Layer-Wise Adaptive Approach Toward Large-Scale DNN Training
Yun-Yong Ko, Dongwon Lee, Sang-Wook Kim
Abstract
A large-batch training with data parallelism is a widely adopted approach to efficiently train a large deep neural network (DNN) model. Large-batch training, however, often suffers from the problem of the model quality degradation because of its fewer iterations. To alleviate this problem, in general, learning rate (lr) scaling methods have been applied, which increases the learning rate to make an update larger at each iteration. Unfortunately, however, we observe that large-batch training with state-of-the-art lr scaling methods still often degrade the model quality when a batch size crosses a specific limit, rendering such lr methods less useful. To this phenomenon, we hypothesize that existing lr scaling methods overlook the subtle but important differences across "layers" in training, which results in the degradation of the overall model quality. From this hypothesis, we propose a novel approach (LENA) toward the learning rate scaling for large-scale DNN training, employing: (1) a layer-wise adaptive lr scaling to adjust lr for each layer individually, and (2) a layer-wise state-aware warm-up to track the state of the training for each layer and finish its warm-up automatically. The comprehensive evaluation with variations of batch sizes demonstrates that LENA achieves the target accuracy (i.e., the accuracy of single-worker training): (1) within the fewest iterations across different batch sizes (up to 45.2% fewer iterations and 44.7% shorter time than the existing state-of-the-art method), and (2) for training very large-batch sizes, surpassing the limits of all baselines. CCS CONCEPTS • Theory of computation → Distributed algorithms.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers5
- KHAN: Knowledge-Aware Hierarchical Attention Networks for Accurate Political Stance PredictionYun-Yong Ko, Seongeun Ryu, Soeun Han, Youngseung Jeon et al.WWW 2023 · 21 citations
- Leveraging the two-timescale regime to demonstrate convergence of neural networksPierre Marion, Raphaël BerthierNeurIPS 2023 · 19 citations
- Stealthy Backdoor Attack in Federated Learning via Adaptive Layer-Wise Gradient AlignmentQingqian Yang, Peishen Yan, Xiaoyu Wu, Jiaru Zhang et al.ICCV 2025 · 2 citations
- PRO-VPT: Distribution-Adaptive Visual Prompt Tuning via Prompt RelocationChikai Shang, Mengke Li, Yiqun Zhang, Zhen Chen et al.ICCV 2025 · 1 citation
- Beyond What's Shared: Recovering Lost Unique Information from Intermediate Layers to Boost Multimodal Geo-Foundation ModelsJangHyeon Lee, Philipe Ambrozio Dias, Yao-Yi Chiang, Dalton LungaCVPR 2026
Builds on3
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- ECLARE: Extreme Classification with Label Graph CorrelationsAnshul Mittal, Noveen Sachdeva, Sheshansh Agrawal, Sumeet Agarwal et al.WWW 2021 · 71 citations
- AdaScale SGD: A User-Friendly Algorithm for Distributed TrainingTyler B. Johnson, Pulkit Agrawal, Haijie Gu, Carlos GuestrinICML 2020 · 41 citations
Related papers
- Large Batch Optimization for Deep Learning Using New Complete Layer-Wise Adaptive Rate ScalingZhouyuan Huo, Bin Gu, Heng HuangAAAI 2021 · 35 citations
- Large Batch Optimization for Deep Learning: Training BERT in 76 minutesYang You, Jing Li, Sashank J. Reddi, Jonathan Hseu et al.ICLR 2020 · 1,170 citations
- SLAMB: Accelerated Large Batch Training with Sparse CommunicationHang Xu, Wenxuan Zhang, Jiawei Fei, Yuzhe Wu et al.ICML 2023 · 7 citations
- JABAS: Joint Adaptive Batching and Automatic Scaling for DNN Training on Heterogeneous GPUsGyeongchan Yun, Junesoo Kang, Hyunjoon Jeong, Sanghyeon Eom et al.EuroSys 2025 · 2 citations
- Augment Your Batch: Improving Generalization Through Instance RepetitionElad Hoffer, Tal Ben-Nun, Itay Hubara, Niv Giladi et al.CVPR 2020
