Self-Improving Teacher Cultivates Better Student: Distillation Calibration for Multimodal Large Language Models
Xinwei Li, Li Lin, Shuai Wang, Chen Qian
Abstract
Multimodal content generation, which leverages visual information to enhance the comprehension of cross-modal understanding, plays a critical role in Multimodal Information Retrieval. With the development of large language models (LLMs), recent research has adopted visual instruction tuning to inject the knowledge of LLMs into downstream multimodal tasks. The high complexity and great demand for resources urge researchers to study efficient distillation solutions to transfer the knowledge from pre-trained multimodal models.(teachers) to more compact student models. However, the instruction tuning for knowledge distillation in multimodal LLMs is resource-intensive and capability-restricted. The comprehension of students is highly reliant on the teacher models. To address this issue, we propose a novel Multimodal Distillation Calibration framework (MmDC). The main idea is to generate high-quality training instances that challenge student models to comprehend and prompt the teacher to calibrate the knowledge transferred to students, ultimately cultivating a better student model in downstream tasks. This framework comprises two stages: (1) multimodal alignment and (2) knowledge distillation calibration. In the first stage, parameter-efficient fine-tuning is used to enhance feature alignment between different modalities. In the second stage, we develop a calibration strategy to assess the student model's capability and generate high-quality instances to calibrate knowledge distillation from teacher to student. The experiments on diverse datasets show that our framework efficiently improves the student model's capabilities. Our 7B-size student model, after three iterations of distillation calibration, outperforms the current state-of-the-art LLaVA-13B model on the ScienceQA and LLaVA Test datasets and also exceeds other strong baselines in a zero-shot setting.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 63fe0fd0-19a6-46a6-83ee-e28a1b1aba3bCited by top-tier papers1
Ask how each one uses itRelated papers
- LLaVA-KD: A Framework of Distilling Multimodal Large Language ModelsYuxuan Cai, Jiangning Zhang, Haoyang He, Xinwei He et al.ICCV 2025 · 9 citations
- Beyond Next-Token Alignment: Distilling Multimodal Large Language Models via Token InteractionsLin Chen, zhaoxiaoke, Kun Ding, Weiwei Feng et al.ICML 2026 · 4 citations
- LLaVA-MoD: Making LLaVA Tiny via MoE-Knowledge DistillationFangxun Shu, Yue Liao, Lei Zhang, Le Zhuo et al.ICLR 2025
- CompoDistill: Attention Distillation for Compositional Reasoning in Multimodal LLMsJiwan Kim, Kibum Kim, Sangwoo Seo, Chanyoung ParkICLR 2026 · 13 citations
- HieRD: Hierarchical Relational Distillation for Vision-Language Embedding ModelsVinh Le, Nguyen Dang, Tu Vu, Linh Van et al.ICML 2026
