Mixture of LoRA Experts
Xun Wu, Shaohan Huang, Furu Wei
Abstract
Low-Rank Adaptation (LoRA) (Hu et al., 2021) has emerged as a pivotal technique for fine-tuning large pre-trained models, renowned for its efficacy across a wide array of tasks. The modular architecture of LoRA has catalyzed further research into the synergistic composition of multiple trained LoRAs, aiming to amplify performance across various tasks. However, the effective composition of these trained LoRAs presents a formidable challenge: (1) Linear arithmetic composition can lead to the diminution of the generative capabilities inherent in the original pre-trained models or the distinctive attributes of the individually trained LoRAs, potentially resulting in suboptimal outcomes. (2) Reference tuning-based composition exhibits limitations in adaptability and incurs significant computational costs due to the requirements to retrain a large model. In response to these challenges, we propose Mixture of LoRA Experts (MOLE). MOLE treats each layer of trained LoRAs as a distinct expert and implements hierarchical weight control by integrating a learnable gating function within each layer to learn optimal composition weights tailored specifically to the objectives of a given domain. MOLE not only demonstrates enhanced performance in LoRA composition but also preserves the essential flexibility necessary for effective composition of trained LoRAs with minimal computational overhead. Extensive experiments conducted in both Natural Language Processing (NLP) and Vision & Language (V&L) domains validate the effects of MOLE. Our code are available at https://github.com/yushuiwx/MoLE.git .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 88d226cf-cec2-4164-9f77-34dab7162f0eCited by top-tier papers97
- Cartridges: Lightweight and general-purpose long context representations via self-studySabri Eyuboglu, Ryan Ehrlich, Simran Arora, Neel Guha et al.ICLR 2026 · 74 citations
- Learning to Route Among Specialized Experts for Zero-Shot GeneralizationMohammed Muqeeth, Haokun Liu, Yufan Liu, Colin RaffelICML 2024 · 63 citations
- How to Continually Adapt Text-to-Image Diffusion Models for Flexible Customization?Jiahua Dong, Wenqi Liang, Hongliu Li, Duzhen Zhang et al.NeurIPS 2024 · 42 citations
- FlyLoRA: Boosting Task Decoupling and Parameter Efficiency via Implicit Rank-Wise Mixture-of-ExpertsHeming Zou, Yunliang Zang, Wutong Xu, Yao Zhu et al.NeurIPS 2025 · 38 citations
- AuroRA: Breaking Low-Rank Bottleneck of LoRA with Nonlinear MappingHaonan Dong, Wenhao Zhu, Guojie Song, Liang WangNeurIPS 2025 · 31 citations
Builds on11
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Adversarial NLI: A New Benchmark for Natural Language UnderstandingYixin Nie, Adina Williams, Emily Dinan, Mohit Bansal et al.ACL 2020 · 602 citations
- An Image is Worth One Word: Personalizing Text-to-Image Generation using Textual InversionRinon Gal, Yuval Alaluf, Yuval Atzmon, Or Patashnik et al.ICLR 2023 · 464 citations
- SVDiff: Compact Parameter Space for Diffusion Fine-TuningLigong Han, Yinxiao Li, Han Zhang, Peyman Milanfar et al.ICCV 2023 · 384 citations
Related papers
- LoRACoE: Improving Large Language Model via Composition-based LoRA ExpertGuanyu Li, Zhiheng Xi, Zhihao Zhang, Boyang Hong et al.EMNLP 2025
- S'MoRE: Structural Mixture of Residual Experts for Parameter-Efficient LLM Fine-tuningHanqing Zeng, Yinglong Xia, Zhuokai Zhao, Chuan Jiang et al.NeurIPS 2025 · 3 citations
- Each Rank Could be an Expert: Single-Ranked Mixture of Experts LoRA for Multi-task LearningZiyu Zhao, Yixiao Zhou, Xin Yu, Zhi Zhang et al.KDD 2026 · 13 citations
- LD-MoLE: Learnable Dynamic Routing for Mixture of LoRA ExpertsYuan Zhuang, Yi Shen, Yuexin Bian, Qing Su et al.ICLR 2026 · 15 citations
- TalkLoRA: Communication-Aware Mixture of Low-Rank Adaptation for Large Language ModelsLin Mu, Haiyang Wang, Li Ni, Lei Sang et al.ACL 2026 · 1 citation
