No Parameters Left Behind: Sensitivity Guided Adaptive Learning Rate for Training Large Transformer Models
Chen Liang, Haoming Jiang, Simiao Zuo, Pengcheng He, Xiaodong Liu, Jianfeng Gao, Weizhu Chen, Tuo Zhao
Abstract
Recent research has shown the existence of significant redundancy in large Transformer models. One can prune the redundant parameters without significantly sacrificing the generalization performance. However, we question whether the redundant parameters could have contributed more if they were properly trained. To answer this question, we propose a novel training strategy that encourages all parameters to be trained sufficiently. Specifically, we adaptively adjust the learning rate for each parameter according to its sensitivity, a robust gradient-based measure reflecting this parameter's contribution to the model performance. A parameter with low sensitivity is redundant, and we improve its fitting by increasing its learning rate. In contrast, a parameter with high sensitivity is well-trained, and we regularize it by decreasing its learning rate to prevent further overfitting. We conduct extensive experiments on natural language understanding, neural machine translation, and image classification to demonstrate the effectiveness of the proposed schedule. Analysis shows that the proposed schedule indeed reduces the redundancy and improves generalization performance.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers9
- Learn more, but bother less: parameter efficient continual learningFuli Qiao, Mehrdad MahdaviNeurIPS 2024 · 36 citations
- Seeking Neural Nuggets: Knowledge Transfer in Large Language Models from a Parametric PerspectiveMing Zhong, Chenxin An, Weizhu Chen, Jiawei Han et al.ICLR 2024 · 16 citations
- Learning List-Level Domain-Invariant Representations for RankingRuicheng Xian, Honglei Zhuang, Zhen Qin, Hamed Zamani et al.NeurIPS 2023 · 11 citations
- Fine-tuning Happens in Tiny Subspaces: Exploring Intrinsic Task-specific Subspaces of Pre-trained Language ModelsZhong Zhang, Bang Liu, Junming ShaoACL 2023 · 8 citations
- PALM: Pushing Adaptive Learning Rate Mechanisms for Continual Test-Time AdaptationSarthak Kumar Maharana, Baoming Zhang, Yunhui GuoAAAI 2025 · 7 citations
Builds on12
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- On the Variance of the Adaptive Learning Rate and BeyondLiyuan Liu, Haoming Jiang, Pengcheng He, Weizhu Chen et al.ICLR 2020 · 2,210 citations
- Sharpness-aware Minimization for Efficiently Improving GeneralizationPierre Foret, Ariel Kleiner, Hossein Mobahi, Behnam NeyshaburICLR 2021 · 1,861 citations
- Reducing Transformer Depth on Demand with Structured DropoutAngela Fan, Edouard Grave, Armand JoulinICLR 2020 · 695 citations
Related papers
- The Importance of Being Parameters: An Intra-Distillation Method for Serious GainsHaoran Xu, Philipp Koehn, Kenton MurrayEMNLP 2022 · 2 citations
- Enhancing Large Language Model Performance with Gradient-Based Parameter SelectionHaoling Li, Xin Zhang, Xiao Liu, Yeyun Gong et al.AAAI 2025
- Gradient-based Gradual Pruning for Language-Specific Multilingual Neural Machine TranslationDan He, Minh-Quang Pham, Thanh-Le Ha, Marco TurchiEMNLP 2023 · 2 citations
- A Better Start: Sensitivity-Aware Warm-Up for Robust and Efficient Fine-TuningYile Chen, Zeyi Wen, Jian Chen, Jin HuangAAAI 2026
- Acceleration of Large Transformer Model Training by Sensitivity-Based Layer DroppingYujie Zeng, Wenlong He, Ihor V. Vasyltsov, Jiali Pang et al.AAAI 2023 · 2 citations
