AutoLR: Layer-wise Pruning and Auto-tuning of Learning Rates in Fine-tuning of Deep Networks
Youngmin Ro, Jin Young Choi
Abstract
Existing fine-tuning methods use a single learning rate over all layers. In this paper, first, we discuss that trends of layer-wise weight variations by fine-tuning using a single learning rate do not match the well-known notion that lower-level layers extract general features and higher-level layers extract specific features. Based on our discussion, we propose an algorithm that improves fine-tuning performance and reduces network complexity through layer-wise pruning and auto-tuning of layer-wise learning rates. The proposed algorithm has verified the effectiveness by achieving state-of-the-art performance on the image retrieval benchmark datasets (CUB-200, Cars-196, Stanford online product, and Inshop). Code is available at https://github.com/youngminPIL/AutoLR.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6635f6b0-b559-48e7-8fd2-87f96f50518aCited by top-tier papers10
- Efficient Federated Learning for Modern NLPDongqi Cai, Yaozong Wu, Shangguang Wang, Felix Xiaozhu Lin et al.MobiCom 2023 · 63 citations
- Surgical Fine-Tuning Improves Adaptation to Distribution ShiftsYoonho Lee, Annie S. Chen, Fahim Tajwar, Ananya Kumar et al.ICLR 2023 · 47 citations
- Adaptive Test-Time Personalization for Federated LearningWenxuan Bao, Tianxin Wei, Haohan Wang, Jingrui HeNeurIPS 2023 · 41 citations
- Temperature Balancing, Layer-wise Weight Analysis, and Neural Network TrainingYefan Zhou, Tianyu Pang, Keqin Liu, Charles H. Martin et al.NeurIPS 2023 · 29 citations
- Leveraging the two-timescale regime to demonstrate convergence of neural networksPierre Marion, Raphaël BerthierNeurIPS 2023 · 19 citations
Related papers
- Comparing Rewinding and Fine-tuning in Neural Network PruningAlex Renda, Jonathan Frankle, Michael CarbinICLR 2020 · 437 citations
- One LR Doesn’t Fit All: Heavy-Tail Guided Layerwise Learning Rates for LLMsDi He, Songjun Tu, Keyu Wang, Lu Yin et al.ICML 2026 · 3 citations
- One Step Learning, One Step ReviewXiaolong Huang, Qiankun Li, Xueran Li, Xuesong GaoAAAI 2024 · 3 citations
- Lottery Ticket Preserves Weight Correlation: Is It Desirable or Not?Ning Liu, Geng Yuan, Zhengping Che, Xuan Shen et al.ICML 2021 · 34 citations
- LoRA+: Efficient Low Rank Adaptation of Large ModelsSoufiane Hayou, Nikhil Ghosh, Bin YuICML 2024 · 388 citations
