Exploring Training on Heterogeneous Data with Mixture of Low-rank Adapters
Yuhang Zhou, Zihua Zhao, Siyuan Du, Haolin Li, Jiangchao Yao, Ya Zhang, Yanfeng Wang
摘要
Training a unified model to take multiple targets into account is a trend towards artificial general intelligence. However, how to efficiently mitigate the training conflicts among heterogeneous data collected from different domains or tasks remains under-explored. In this study, we explore to leverage Mixture of Low-rank Adapters (MoLA) to mitigate conflicts in heterogeneous data training, which requires to jointly train the multiple lowrank adapters and their shared backbone. Specifically, we introduce two variants of MoLA, namely, MoLA-Grad and MoLA-Router, to respectively handle the target-aware and target-agnostic scenarios during inference. The former uses task identifiers to assign personalized low-rank adapters to each task, disentangling task-specific knowledge towards their adapters, thereby mitigating heterogeneity conflicts. The latter uses a novel Task-wise Decorrelation (TwD) loss to intervene the router to learn oriented weight combinations of adapters to homogeneous tasks, achieving similar effects. We conduct comprehensive experiments to verify the superiority of MoLA over previous state-of-the-art methods and present indepth analysis on its working mechanism. Source code is available at: https://github.com/ MediaBrain-SJTU/MoLA Exploring Training on Heterogeneous Data with Mixture of Low-rank Adapters Domain heterogeneity Multi-input task heterogeneity Single-input task heterogeneity Art Clipart Product Real World Chair
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Dynamic Modeling of Patients, Modalities and Tasks via Multi-modal Multi-task Mixture of ExpertsChenwei Wu, Zitao Shuai, Zhengxu Tang, Luning Wang 等ICLR 2025
- ELoRA: Low-Rank Adaptation for Equivariant GNNsChen Wang, Siyu Hu, Guangming Tan, Weile JiaICML 2025
- Ensembles of Low-Rank Expert AdaptersYinghao Li, Vianne R. Gao, Chao Zhang, MohamadAli TorkamaniICLR 2025
- Solving Token Gradient Conflict in Mixture-of-Experts for Large Vision-Language ModelLongrong Yang, Dong Shen, Chaoxiang Cai, Fan Yang 等ICLR 2025
它引用的顶会 Paper17
- Gradient Surgery for Multi-Task LearningTianhe Yu, Saurabh Kumar, Abhishek Gupta, Sergey Levine 等NeurIPS 2020 · 被引用 2,261 次
- Conflict-Averse Gradient Descent for Multi-task learningBo Liu, Xingchao Liu, Xiaojie Jin, Peter Stone 等NeurIPS 2021 · 被引用 686 次
- Just Pick a Sign: Optimizing Deep Multitask Models with Gradient Sign DropoutZhao Chen, Jiquan Ngiam, Yanping Huang, Thang Luong 等NeurIPS 2020 · 被引用 313 次
- Multi-Task Learning as a Bargaining GameAviv Navon, Aviv Shamsian, Idan Achituve, Haggai Maron 等ICML 2022 · 被引用 243 次
- Gradient Vaccine: Investigating and Improving Multi-task Optimization in Massively Multilingual ModelsZirui Wang, Yulia Tsvetkov, Orhan Firat, Yuan CaoICLR 2021 · 被引用 241 次
相关 Paper
- Each Rank Could be an Expert: Single-Ranked Mixture of Experts LoRA for Multi-task LearningZiyu Zhao, Yixiao Zhou, Xin Yu, Zhi Zhang 等KDD 2026 · 被引用 13 次
- LoRA-Mixer: Coordinate Modular LoRA Experts Through Serial Attention RoutingWenbing Li, Zikai Song, Hang Zhou, Junqing Yu 等ICLR 2026 · 被引用 20 次
- FlyLoRA: Boosting Task Decoupling and Parameter Efficiency via Implicit Rank-Wise Mixture-of-ExpertsHeming Zou, Yunliang Zang, Wutong Xu, Yao Zhu 等NeurIPS 2025 · 被引用 38 次
- A Stronger Mixture of Low-Rank Experts for Fine-Tuning Foundation ModelsMengyang Sun, Yihao Wang, Tao Feng, Dan Zhang 等ICML 2025
- Batched Low-Rank Adaptation of Foundation ModelsYeming Wen, Swarat ChaudhuriICLR 2024 · 被引用 32 次
