WeightLoRA: Keep Only Necessary Adapters
Andrey Veprikov, Vladimir Solodkin, Alexander Zyl, Andrey V. Savchenko, Aleksandr Beznosikov
摘要
The widespread utilization of language models in modern applications is inconceivable without Parameter-Efficient Fine-Tuning techniques, such as low-rank adaptation (), which adds trainable adapters to selected layers. Although may obtain accurate solutions, it requires significant memory to train large models and intuition on which layers to add adapters. In this paper, we propose a novel method, , which overcomes this issue by adaptive selection of the most critical heads throughout the optimization process. As a result, we can significantly reduce the number of trainable parameters while maintaining the capability to obtain consistent or even superior metric values. We conduct experiments for a series of competitive benchmarks and DeBERTa, BART, and Llama models, comparing our method with different adaptive approaches. The experimental results demonstrate the efficacy of and the superior performance of in almost all cases.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper6
- Deberta: decoding-Enhanced Bert with Disentangled AttentionPengcheng He, Xiaodong Liu, Jianfeng Gao, Weizhu ChenICLR 2021 · 被引用 3,729 次
- DeBERTaV3: Improving DeBERTa using ELECTRA-Style Pre-Training with Gradient-Disentangled Embedding SharingPengcheng He, Jianfeng Gao, Weizhu ChenICLR 2023 · 被引用 394 次
- Flexora: Flexible Low-Rank Adaptation for Large Language ModelsChenxing Wei, Yao Shu, Ying Tiffany He, Fei YuACL 2025 · 被引用 12 次
- The Iterative Optimal Brain Surgeon: Faster Sparse Recovery by Leveraging Second-Order InformationDiyuan Wu, Ionut-Vlad Modoranu, Mher Safaryan, Denis Kuznedelev 等NeurIPS 2024 · 被引用 8 次
- Zeroth-Order Hard-Thresholding: Gradient Error vs. ExpansivityWilliam de Vazelhes, Hualin Zhang, Huimin Wu, Xiaotong Yuan 等NeurIPS 2022 · 被引用 4 次
相关 Paper
- DenseLoRA: Dense Low-Rank Adaptation of Large Language ModelsLin Mu, Xiaoyu Wang, Li Ni, Yang Li 等ACL 2025 · 被引用 3 次
- MELoRA: Mini-Ensemble Low-Rank Adapters for Parameter-Efficient Fine-TuningPengjie Ren, Chengshun Shi, Shiguang Wu, Mengqi Zhang 等ACL 2024
- BSLoRA: Enhancing the Parameter Efficiency of LoRA with Intra-Layer and Inter-Layer SharingYuhua Zhou, Ruifeng Li, Changhai Zhou, Fei Yang 等ICML 2025
- Beyond Higher Rank: Token-wise Input-Output Projections for Efficient Low-Rank AdaptationShiwei Li, Xiandi Luo, Haozhao Wang, Xing Tang 等NeurIPS 2025 · 被引用 10 次
- Stable-LoRA: Stabilizing Feature Learning of Low-Rank AdaptationYize Wu, Ke Gao, Ling Li, Yanjun WuICLR 2026 · 被引用 1 次
