LoRA-E2: Effective and Efficient Low-rank Adaptation
Shengkun Zhu, Jinshan Zeng, Yiming Wang, Sheng Wang, Yuan Sun, Shangfeng Chen, Yuan Yao, Qiang Yang
摘要
Low-rank adaptation (LoRA) has emerged as an efficient fine-tuning technique for large language models, enabling parameter-efficient updates while maintaining task performance. However, LoRA suffers from two key issues: 1) inefficient feature learning when the width n (embedding dimension) is large, and 2) ineffective updates to the adapter matrix A due to the initialization of B as zero. We propose LoRA-E2, which utilizes a Gaussian initialization with variance Θ(n-3/4) for A, and employs the Gauss-Seidel iteration to train B and A. We theoretically show that LoRA-E2 enables more stable and efficient feature learning with effective parameter updates over standard LoRA. Empirically, LoRA-E2 achieves consistent gains in both natural language understanding and generation tasks. On the GLUE benchmark with T5-base, it improves performance by 1-10% over LoRA. When fine-tuning LLaMA 2-7B on MetaMathQA with GSM8K as validation, LoRA-E2 surpasses LoRA by 1-2% and converges up to ∼1/43× faster. Code is available at https://github.com/whu-totemdb/LoRA-E2.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- LoRA-GA: Low-Rank Adaptation with Gradient ApproximationShaowen Wang, Linxi Yu, Jian LiNeurIPS 2024 · 被引用 194 次
- LoRA Done RITE: Robust Invariant Transformation Equilibration for LoRA OptimizationJui-Nan Yen, Si Si, Zhao Meng, Felix X. Yu 等ICLR 2025
- LoRA+: Efficient Low Rank Adaptation of Large ModelsSoufiane Hayou, Nikhil Ghosh, Bin YuICML 2024 · 被引用 388 次
- GoRA: Gradient-driven Adaptive Low Rank AdaptationHaonan He, Peng Ye, Yuchen Ren, Yuan Yuan 等NeurIPS 2025 · 被引用 17 次
- PiSSA: Principal Singular Values and Singular Vectors Adaptation of Large Language ModelsFanxu Meng, Zhaohui Wang, Muhan ZhangNeurIPS 2024 · 被引用 374 次
