Q-MoE: Connector for MLLMs with Text-Driven Routing
Hanzi Wang, Jiamin Ren, Yifeng Ding, Lei Ren, Huixing Jiang, Wei Chen, Fangxiang Feng, Xiaojie Wang
摘要
Multimodal Large Language Models (MLLMs) have showcased remarkable advances in handling various vision-language tasks. These models typically consist of a Large Language Model (LLM), a vision encoder and a connector structure, which is used to bridge the modality gap between vision and language. It is challenging for the connector to filter the right visual information for LLM according to the task in hand. Most of previous connectors, such as light-weight projection and Q-former, treat visual information for diverse tasks uniformly, therefore lacking task-specific visual information extraction capabilities. To address the issue, this paper proposes Q-MoE, a query-based connector with Mixture-of-Experts (MoE) to extract task-specific information with text-driven routing. Furthermore, an optimal path based training strategy is proposed to find an optimal expert combination. Extensive experiments on two popular open-source LLMs and several different visual-language tasks demonstrate the effectiveness of the Q-MoE connecter.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper1
问问它们各自怎么用它相关 Paper
- MoME: Mixture of Multimodal Experts for Generalist Multimodal Large Language ModelsLeyang Shen, Gongwei Chen, Rui Shao, Weili Guan 等NeurIPS 2024 · 被引用 55 次
- Uni-Med: A Unified Medical Generalist Foundation Model For Multi-Task Learning Via Connector-MoEXun Zhu, Ying Hu, Fanbin Mo, Miao Li 等NeurIPS 2024 · 被引用 29 次
- Visual Anchors Are Strong Information Aggregators For Multimodal Large Language ModelHaogeng Liu, Quanzeng You, Xiaotian Han, Yongfei Liu 等NeurIPS 2024 · 被引用 5 次
- ParGo: Bridging Vision-Language with Partial and Global ViewsAn-Lan Wang, Bin Shan, Wei Shi, Kun-Yu Lin 等AAAI 2025 · 被引用 42 次
- MoVA: Adapting Mixture of Vision Experts to Multimodal ContextZhuofan Zong, Bingqi Ma, Dazhong Shen, Guanglu Song 等NeurIPS 2024 · 被引用 110 次
