AutoLR: Layer-wise Pruning and Auto-tuning of Learning Rates in Fine-tuning of Deep Networks
Youngmin Ro, Jin Young Choi
摘要
Existing fine-tuning methods use a single learning rate over all layers. In this paper, first, we discuss that trends of layer-wise weight variations by fine-tuning using a single learning rate do not match the well-known notion that lower-level layers extract general features and higher-level layers extract specific features. Based on our discussion, we propose an algorithm that improves fine-tuning performance and reduces network complexity through layer-wise pruning and auto-tuning of layer-wise learning rates. The proposed algorithm has verified the effectiveness by achieving state-of-the-art performance on the image retrieval benchmark datasets (CUB-200, Cars-196, Stanford online product, and Inshop). Code is available at https://github.com/youngminPIL/AutoLR.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- Efficient Federated Learning for Modern NLPDongqi Cai, Yaozong Wu, Shangguang Wang, Felix Xiaozhu Lin 等MobiCom 2023 · 被引用 63 次
- Surgical Fine-Tuning Improves Adaptation to Distribution ShiftsYoonho Lee, Annie S. Chen, Fahim Tajwar, Ananya Kumar 等ICLR 2023 · 被引用 47 次
- Adaptive Test-Time Personalization for Federated LearningWenxuan Bao, Tianxin Wei, Haohan Wang, Jingrui HeNeurIPS 2023 · 被引用 41 次
- Temperature Balancing, Layer-wise Weight Analysis, and Neural Network TrainingYefan Zhou, Tianyu Pang, Keqin Liu, Charles H. Martin 等NeurIPS 2023 · 被引用 29 次
- Leveraging the two-timescale regime to demonstrate convergence of neural networksPierre Marion, Raphaël BerthierNeurIPS 2023 · 被引用 19 次
相关 Paper
- Comparing Rewinding and Fine-tuning in Neural Network PruningAlex Renda, Jonathan Frankle, Michael CarbinICLR 2020 · 被引用 437 次
- One LR Doesn’t Fit All: Heavy-Tail Guided Layerwise Learning Rates for LLMsDi He, Songjun Tu, Keyu Wang, Lu Yin 等ICML 2026 · 被引用 3 次
- One Step Learning, One Step ReviewXiaolong Huang, Qiankun Li, Xueran Li, Xuesong GaoAAAI 2024 · 被引用 3 次
- Lottery Ticket Preserves Weight Correlation: Is It Desirable or Not?Ning Liu, Geng Yuan, Zhengping Che, Xuan Shen 等ICML 2021 · 被引用 34 次
- LoRA+: Efficient Low Rank Adaptation of Large ModelsSoufiane Hayou, Nikhil Ghosh, Bin YuICML 2024 · 被引用 388 次
