CoMMIT: Coordinated Multimodal Instruction Tuning
Xintong Li, Junda Wu, Tong Yu, Rui Wang, Yu Wang, Xiang Chen, Jiuxiang Gu, Lina Yao, Julian J. McAuley, Jingbo Shang
Abstract
Instruction tuning in multimodal large language models (MLLMs) generally involves cooperative learning between a backbone LLM and a feature encoder of non-text input modalities. The major challenge is how to efficiently find the synergy between the two modules so that LLMs can adapt their reasoning abilities to downstream tasks while feature encoders can adjust to provide more task-specific information about its modality. In this paper, we analyze the MLLM instruction tuning from both theoretical and empirical perspectives, where we find the unbalanced learning between the feature encoder and the LLM can cause problems of oscillation and biased learning that lead to sub-optimal convergence. Inspired by our findings, we propose a Multimodal Balance Coefficient that enables quantitative measurement of the balance of learning. Based on this, we further design a dynamic learning scheduler that better coordinates the learning between the LLM and feature encoder, alleviating the problems of oscillation and biased learning. In addition, we introduce an auxiliary regularization on the gradient to promote updating with larger step sizes, which potentially allows for a more accurate estimation of the proposed Mul-tiModal Balance Coefficient and further improves the training sufficiency. Our proposed approach is agnostic to the architecture of LLM and feature encoder, so it can be generically integrated with various MLLMs. We conduct experiments on multiple downstream tasks with various MLLMs, demonstrating the proposed method is more effective than the baselines in MLLM instruction tuning.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f29dbc74-78e2-4782-b9f5-e0dcd2c5b508Cited by top-tier papers2
- When the Prompt Becomes Visual: Vision-Centric Jailbreak Attacks for Large Image Editing ModelsJiacheng Hou, Yining Sun, Ruochong Jin, Haochen Han et al.ICML 2026 · 3 citations
- CapeLLM: Support-Free Category-Agnostic Pose Estimation with Multimodal Large Language ModelsJunho Kim, Hyungjin Chung, Byung-Hoon KimICCV 2025 · 1 citation
Builds on15
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- InstructBLIP: Towards General-purpose Vision-Language Models with Instruction TuningWenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong et al.NeurIPS 2023 · 4,013 citations
- Cross-Entropy Loss Functions: Theoretical Analysis and ApplicationsAnqi Mao, Mehryar Mohri, Yutao ZhongICML 2023 · 790 citations
Related papers
- Beyond Fixed Biases: Decoding the Role of Reasoning Uncertainty in MLLM Modality ConflictsZhuoran Zhang, Tengyue Wang, Xilin Gong, Yang Shi et al.ICML 2026
- Dynamic Mixture of Curriculum LoRA Experts for Continual Multimodal Instruction TuningChendi Ge, Xin Wang, Zeyang Zhang, Hong Chen et al.ICML 2025
- LLaVA Steering: Visual Instruction Tuning with 500x Fewer Parameters through Modality Linear Representation-SteeringJinhe Bi, Yujun Wang, Haokun Chen, Xun Xiao et al.ACL 2025
- Adaptive Task Balancing for Visual Instruction Tuning via Inter-Task Contribution and Intra-Task DifficultyYanqi Dai, Yong Wang, Zebin You, Dong Jing et al.WWW 2026 · 4 citations
- M²PT: Multimodal Prompt Tuning for Zero-shot Instruction LearningTaowen Wang, Yiyang Liu, James Liang, Junhan Zhao et al.EMNLP 2024 · 31 citations
