Make LoRA Great Again: Boosting LoRA with Adaptive Singular Values and Mixture-of-Experts Optimization Alignment
Chenghao Fan, Zhenyi Lu, Sichen Liu, Chengfeng Gu, Xiaoye Qu, Wei Wei, Yu Cheng
摘要
While Low-Rank Adaptation (LoRA) enables parameter-efficient fine-tuning for Large Language Models (LLMs), its performance often falls short of Full Fine-Tuning (Full FT). Current methods optimize LoRA by initializing with static singular value decomposition (SVD) subsets, leading to suboptimal leveraging of pre-trained knowledge. Another path for improving LoRA is incorporating a Mixture-of-Experts (MoE) architecture. However, weight misalignment and complex gradient dynamics make it challenging to adopt SVD prior to the LoRA MoE architecture. To mitigate these issues, we propose Great LoRA Mixture-of-Expert (GOAT), a framework that (1) adaptively integrates relevant priors using an SVD-structured MoE, and (2) aligns optimization with full finetuned MoE by deriving a theoretical scaling factor. We demonstrate that proper scaling, without modifying the architecture or training algorithms, boosts LoRA MoE's efficiency and performance. Experiments across 25 datasets, including natural language understanding, commonsense reasoning, image classification, and natural language generation, demonstrate GOAT's state-of-the-art performance, closing the gap with Full FT.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- Correlated Low-Rank Adaptation for ConvNetsWu Ran, Weijia Zhang, Shuyang Pang, Qi Zhu 等NeurIPS 2025 · 被引用 5 次
- Sparse Spectral LoRA: Routed Experts for Medical VLMsOmid Nejati Manzari, Hojat Asgariandehkordi, Taha Koleilat, Yiming Xiao 等CVPR 2026 · 被引用 3 次
- Beyond Adapter Retrieval: Latent Geometry-Preserving Composition via Sparse Task ProjectionPengfei Jin, Peng Shu, Sifan Song, Sekeun Kim 等AAAI 2026 · 被引用 2 次
- CLIP-MoE: Towards Building Mixture of Experts for CLIP with Diversified Multiplet UpcyclingJihai Zhang, Xiaoye Qu, Tong Zhu, Yu ChengEMNLP 2025 · 被引用 1 次
- Decomposing the Basic Abilities of Large Language Models: Mitigating Cross-Task Interference in Multi-Task Instruct-TuningBing Wang, Ximing Li, Changchun Li, Jinjin Chi 等ICML 2026 · 被引用 1 次
它引用的顶会 Paper16
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- PIQA: Reasoning about Physical Commonsense in Natural LanguageYonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao 等AAAI 2020 · 被引用 2,916 次
- GaLore: Memory-Efficient LLM Training by Gradient Low-Rank ProjectionJiawei Zhao, Zhenyu Zhang, Beidi Chen, Zhangyang Wang 等ICML 2024 · 被引用 433 次
- PiSSA: Principal Singular Values and Singular Vectors Adaptation of Large Language ModelsFanxu Meng, Zhaohui Wang, Muhan ZhangNeurIPS 2024 · 被引用 374 次
相关 Paper
- S'MoRE: Structural Mixture of Residual Experts for Parameter-Efficient LLM Fine-tuningHanqing Zeng, Yinglong Xia, Zhuokai Zhao, Chuan Jiang 等NeurIPS 2025 · 被引用 3 次
- MoA: Heterogeneous Mixture of Adapters for Parameter-Efficient Fine-Tuning of Large Language ModelsJie Cao, Tianwei Lin, Bo Yuan, Rolan Yan 等ACL 2026 · 被引用 2 次
- TalkLoRA: Communication-Aware Mixture of Low-Rank Adaptation for Large Language ModelsLin Mu, Haiyang Wang, Li Ni, Lei Sang 等ACL 2026 · 被引用 1 次
- LoRACoE: Improving Large Language Model via Composition-based LoRA ExpertGuanyu Li, Zhiheng Xi, Zhihao Zhang, Boyang Hong 等EMNLP 2025
- Each Rank Could be an Expert: Single-Ranked Mixture of Experts LoRA for Multi-task LearningZiyu Zhao, Yixiao Zhou, Xin Yu, Zhi Zhang 等KDD 2026 · 被引用 13 次
