Pay Attention to Small Weights
Chao Zhou, Tom Jacobs, Advait Gadhikar, Rebekka Burkholz
摘要
Finetuning large pretrained neural networks is known to be resource-intensive, both in terms of memory and computational cost. To mitigate this, a common approach is to restrict training to a subset of the model parameters. By analyzing the relationship between gradients and weights during finetuning, we observe a notable pattern: large gradients are often associated with small-magnitude weights. This correlation is more pronounced in finetuning settings than in training from scratch. Motivated by this observation, we propose NANOADAM, which dynamically updates only the small-magnitude weights during finetuning and offers several practical advantages: first, this criterion is gradient-free -- the parameter subset can be determined without gradient computation; second, it preserves large-magnitude weights, which are likely to encode critical features learned during pretraining, thereby reducing the risk of catastrophic forgetting; thirdly, it permits the use of larger learning rates and consistently leads to better generalization performance in experiments. We demonstrate this for both NLP and vision tasks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper13
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- DoRA: Weight-Decomposed Low-Rank AdaptationShih-Yang Liu, Chien-Yi Wang, Hongxu Yin, Pavlo Molchanov 等ICML 2024 · 被引用 820 次
- 8-bit Optimizers via Block-wise QuantizationTim Dettmers, Mike Lewis, Sam Shleifer, Luke ZettlemoyerICLR 2022 · 被引用 457 次
- GaLore: Memory-Efficient LLM Training by Gradient Low-Rank ProjectionJiawei Zhao, Zhenyu Zhang, Beidi Chen, Zhangyang Wang 等ICML 2024 · 被引用 433 次
相关 Paper
- Powerpropagation: A sparsity inducing weight reparameterisationJonathan Schwarz, Siddhant M. Jayakumar, Razvan Pascanu, Peter E. Latham 等NeurIPS 2021 · 被引用 63 次
- Adaptive Budget Allocation for Parameter-Efficient Fine-TuningQingru Zhang, Minshuo Chen, Alexander Bukharin, Pengcheng He 等ICLR 2023 · 被引用 32 次
- NeuroAda: Activating Each Neuron's Potential for Parameter-Efficient Fine-TuningZhi Zhang, Yixian Shen, Congfeng Cao, Ekaterina ShutovaEMNLP 2025
- Model-Dowser: Data-Free Importance Probing to Mitigate Catastrophic Forgetting in Multimodal Large Language ModelsHyeontaek Hwang, DINH SON NGUYEN, Daeyoung KimICML 2026
- HFT: Half Fine-Tuning for Large Language ModelsTingfeng Hui, Zhenyu Zhang, Shuohuan Wang, Weiran Xu 等ACL 2025
