Wings: Learning Multimodal LLMs without Text-only Forgetting
Yi-Kai Zhang, Shiyin Lu, Yang Li, Yanqing Ma, Qingguo Chen, Zhao Xu, Weihua Luo, Kaifu Zhang, De-Chuan Zhan, Han-Jia Ye
摘要
Multimodal large language models (MLLMs), initiated with a trained LLM, first align images with text and then fine-tune on multimodal mixed inputs. However, the MLLM catastrophically forgets the text-only instructions, which do not include images and can be addressed within the initial LLM. In this paper, we present Wings, a novel MLLM that excels in both text-only dialogues and multimodal comprehension. Analyzing MLLM attention in multimodal instructions reveals that text-only forgetting is related to the attention shifts from pre-image to post-image text. From that, we construct extra modules that act as the boosted learner to compensate for the attention shift. The complementary visual and textual learners, like"wings"on either side, are connected in parallel within each layer's attention block. Initially, image and text inputs are aligned with visual learners operating alongside the main attention, balancing focus on visual elements. Textual learners are later collaboratively integrated with attention-based routing to blend the outputs of the visual and textual learners. We design the Low-Rank Residual Attention (LoRRA) to guarantee high efficiency for learners. Our experimental results demonstrate that Wings outperforms equally-scaled MLLMs in both text-only and visual question-answering tasks. On a newly constructed Interleaved Image-Text (IIT) benchmark, Wings exhibits superior performance from text-only-rich to multimodal-rich question-answering tasks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- Initializing Variable-sized Vision Transformers from Learngene with Learnable TransformationShiyu Xia, Yuankun Zu, Xu Yang, Xin GengNeurIPS 2024 · 被引用 9 次
- Mitigating Intra- and Inter-modal Forgetting in Continual Learning of Unified Multimodal ModelsXiwen Wei, Mustafa Munir, Radu MarculescuNeurIPS 2025 · 被引用 9 次
- Linearly Decomposing and Recomposing Vision Transformers for Diverse-Scale ModelsShuxia Lin, Miaosen Zhang, Ruiming Chen, Xu Yang 等NeurIPS 2024 · 被引用 7 次
- Decouple to Generalize: Context-First Self-Evolving Learning for Data-Scarce Vision-Language ReasoningTingyu Li, Zheng Sun, Jingxuan Wei, Conghui He 等CVPR 2026 · 被引用 2 次
- Data Selection Matters: Towards Robust Instruction Tuning of Large Multimodal ModelsXu Yang, Chen Liu, Ying WeiNeurIPS 2025 · 被引用 2 次
它引用的顶会 Paper48
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
相关 Paper
- Libra: Building Decoupled Vision System on Large Language ModelsYifan Xu, Xiaoshan Yang, Yaguang Song, Changsheng XuICML 2024 · 被引用 11 次
- LLaDA-V: Large Language Diffusion Models with Visual Instruction TuningZebin You, Shen Nie, Xiaolu Zhang, JUN ZHOU 等CVPR 2026 · 被引用 154 次
- SMoLoRa: Exploring and Defying Dual Catastrophic Forgetting in Continual Visual Instruction TuningZiqi Wang, Chang Che, Qi Wang, Yangyang Li 等ICCV 2025 · 被引用 4 次
- LRM-LLaVA: Overcoming the Modality Gap of Multilingual Large Language-Vision Model for Low-Resource LanguagesJunchen Li, Qing Yang, Bojian Jiang, Shaolin Zhu 等AAAI 2025 · 被引用 3 次
- VFA: Empowering Multilingual MLLMs via Vision-Free AdaptationYixia Li, Yaqing Shi, Zhiwen Ruan, Dongdong Zhang 等ACL 2026
