LLaVA-MoD: Making LLaVA Tiny via MoE-Knowledge Distillation
Fangxun Shu, Yue Liao, Lei Zhang, Le Zhuo, Chenning Xu, Guanghao Zhang, Haonan Shi, Long Chan, Tao Zhong, Zhelun Yu, Wanggui He, Siming Fu
Abstract
We introduce LLaVA-MoD, a novel framework designed to enable the efficient training of small-scale Multimodal Language Models (s-MLLM) distilling knowledge from large-scale MLLM (l-MLLM). Our approach tackles two fundamental challenges in MLLM distillation. First, we optimize the network structure of s-MLLM by integrating a sparse Mixture of Experts (MoE) architecture into the language model, striking a balance between computational efficiency and model expressiveness. Second, we propose a progressive knowledge transfer strategy for comprehensive knowledge transfer. This strategy begins with mimic distillation, where we minimize the Kullback-Leibler (KL) divergence between output distributions to enable s-MLLM to emulate l-MLLM's understanding. Following this, we introduce preference distillation via Preference Optimization (PO), where the key lies in treating l-MLLM as the reference model. During this phase, the s-MLLM's ability to discriminate between superior and inferior examples is significantly enhanced beyond l-MLLM, leading to a better s-MLLM that surpasses l-MLLM, particularly in hallucination benchmarks. Extensive experiments demonstrate that LLaVA-MoD surpasses existing works across various benchmarks while maintaining minimal activated parameters and low computational costs. Remarkably, LLaVA-MoD-2B surpasses Qwen-VL-Chat-7B with an average gain of 8.8%, using merely 0.3% of the training data and 23% trainable parameters. The results underscore LLaVA-MoD's ability to effectively distill comprehensive knowledge from its teacher model, paving the way for developing efficient MLLMs.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers20
- DeepEyes: Incentivizing "Thinking with Images" via Reinforcement LearningZiwei Zheng, Michael Yang, Jack Hong, Chenxiao Zhao et al.ICLR 2026 · 321 citations
- MoESD: Unveil Speculative Decoding's Potential for Accelerating Sparse MoEZongle Huang, Lei Zhu, Zongyuan Zhan, Ting Hu et al.NeurIPS 2025 · 23 citations
- PrefixKV: Adaptive Prefix KV Cache is What Vision Instruction-Following Models Need for Efficient GenerationAo Wang, Hui Chen, Jianchao Tan, Kefeng Zhang et al.NeurIPS 2025 · 16 citations
- CompoDistill: Attention Distillation for Compositional Reasoning in Multimodal LLMsJiwan Kim, Kibum Kim, Sangwoo Seo, Chanyoung ParkICLR 2026 · 13 citations
- Discovering Important Experts for Mixture-of-Experts Models Pruning Through a Theoretical PerspectiveWeizhong Huang, Yuxin Zhang, Xiawu Zheng, Fei Chao et al.NeurIPS 2025 · 12 citations
Builds on20
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech et al.NeurIPS 2022 · 6,707 citations
- MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language ModelsDeyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li et al.ICLR 2024 · 3,079 citations
Related papers
- LLaVA-KD: A Framework of Distilling Multimodal Large Language ModelsYuxuan Cai, Jiangning Zhang, Haoyang He, Xinwei He et al.ICCV 2025 · 9 citations
- Routing Experts: Learning to Route Dynamic Experts in Existing Multi-modal Large Language ModelsQiong Wu, Zhaoxi Ke, Yiyi Zhou, Xiaoshuai Sun et al.ICLR 2025
- Beyond Next-Token Alignment: Distilling Multimodal Large Language Models via Token InteractionsLin Chen, zhaoxiaoke, Kun Ding, Weiwei Feng et al.ICML 2026 · 4 citations
- Self-Improving Teacher Cultivates Better Student: Distillation Calibration for Multimodal Large Language ModelsXinwei Li, Li Lin, Shuai Wang, Chen QianSIGIR 2024 · 4 citations
- MoVE-KD: Knowledge Distillation for VLMs with Mixture of Visual EncodersJiajun Cao, Yuan Zhang, Tao Huang, Ming Lu et al.CVPR 2025
