ModalPrompt: Towards Efficient Multimodal Continual Instruction Tuning with Dual-Modality Guided Prompt
Fanhu Zeng, Fei Zhu, Haiyang Guo, Xu-Yao Zhang, Cheng-Lin Liu
Abstract
Large Multimodal Models (LMMs) exhibit remarkable multi-tasking ability by learning mixed instruction datasets. However, novel tasks would be encountered sequentially in dynamic world, which urges for equipping LMMs with multimodal continual instruction learning (MCIT) ability especially for diverse and challenging generative tasks. Existing MCIT methods do not fully exploit the unique attribute of LMMs and often gain performance at the expense of efficiency. In this paper, we propose a novel prompt learning framework for MCIT to effectively alleviate forgetting of previous knowledge while managing computational complexity with natural image-text supervision. Concretely, we learn prompts for each task and exploit efficient prompt fusion for knowledge transfer and prompt selection for complexity management with dual-modality guidance. Extensive experiments demonstrate that our approach achieves substantial +14.26% performance gain on MCIT benchmarks with remarkable ×1.42 inference speed free from growing computation. Code is available at https:// github.com/AuroraZengfh/ModalPrompt .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2e8670a7-229a-4e6e-a332-107309e53469Cited by top-tier papers2
- Reinforcement Fine-Tuning Naturally Mitigates Forgetting in Continual Post-TrainingSong Lai, Haohan Zhao, Rong Feng, Changyi Ma et al.ICML 2026 · 46 citations
- FOREVER: Forgetting Curve-Inspired Memory Replay for Language Model Continual LearningYujie Feng, Hao Wang, Jian Li, Xu Chu et al.ACL 2026 · 3 citations
Builds on18
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech et al.NeurIPS 2022 · 6,707 citations
- InstructBLIP: Towards General-purpose Vision-Language Models with Instruction TuningWenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong et al.NeurIPS 2023 · 4,013 citations
- MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual ContextsPan Lu, Hritik Bansal, Tony Xia, Jiacheng Liu et al.ICLR 2024 · 1,472 citations
Related papers
- M²PT: Multimodal Prompt Tuning for Zero-shot Instruction LearningTaowen Wang, Yiyang Liu, James Liang, Junhan Zhao et al.EMNLP 2024 · 31 citations
- Leveraging Large Vision-Language Model as User Intent-Aware Encoder for Composed Image RetrievalZelong Sun, Dong Jing, Guoxing Yang, Nanyi Fei et al.AAAI 2025 · 13 citations
- Federated Continual Instruction TuningHaiyang Guo, Fanhu Zeng, Fei Zhu, Wenzhuo Liu et al.ICCV 2025 · 2 citations
- Mitigating the Evolving Semantic Entanglement in Continual Learning of Vision-Language ModelsYiliang Zhu, Dayan Wu, Qinghang Su, Zexian Yang et al.ACM MM 2025
- COMMA: Co-articulated Multi-Modal LearningLianyu Hu, Liqing Gao, Zekang Liu, Chi-Man Pun et al.AAAI 2024 · 7 citations
