MoVE-KD: Knowledge Distillation for VLMs with Mixture of Visual Encoders
Jiajun Cao, Yuan Zhang, Tao Huang, Ming Lu, Qizhe Zhang, Ruichuan An, Ningning Ma, Shanghang Zhang
Abstract
Visual encoders are fundamental components in visionlanguage models (VLMs), each showcasing unique strengths derived from various pre-trained visual foundation models. To leverage the various capabilities of these encoders, recent studies incorporate multiple encoders within a single VLM, leading to a considerable increase in computational cost. In this paper, we present Mixture-of-Visual-Encoder Knowledge Distillation (MoVE-KD), a novel framework that distills the unique proficiencies of multiple vision encoders into a single, efficient encoder model. Specifically, to mitigate conflicts and retain the unique characteristics of each teacher encoder, we employ low-rank adaptation (LoRA) and mixture-of-experts (MoEs) to selectively activate specialized knowledge based on input features, enhancing both adaptability and efficiency. To regularize the KD process and enhance performance, we propose an attention-based distillation strategy that adaptively weighs the different encoders and emphasizes valuable visual tokens, reducing the burden of replicating comprehensive but distinct features from multiple teachers. Comprehensive experiments on popular VLMs, such as LLaVA and LLaVA-NeXT, validate the effectiveness of our method. Our code is available at: https://github.com/hey-cjj/MoVE-KD .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers16
- MoLe-VLA: Dynamic Layer-skipping Vision Language Action Model via Mixture-of-Layers for Efficient Robot ManipulationRongyu Zhang, Menghang Dong, Yuan Zhang, Liang Heng et al.AAAI 2026 · 56 citations
- FastDriveVLA: Efficient End-to-End Driving via Plug-and-Play Reconstruction-based Token PruningJiajun Cao, Qizhe Zhang, Peidong Jia, Xuhui Zhao et al.AAAI 2026 · 18 citations
- Masking Teacher and Reinforcing Student for Distilling Vision-Language ModelsByung-Kwan Lee, Yu-Chiang Frank Wang, Ryo HachiumaCVPR 2026 · 7 citations
- Hawaii: Hierarchical Visual Knowledge Transfer for Efficient Vision-Language ModelsYimu Wang, Mozhgan Nasr Azadani, Sean Sedwards, Krzysztof CzarneckiNeurIPS 2025 · 6 citations
- Gated Relational Alignment via Confidence-based Distillation for Efficient VLMsYanlong Chen, Amir Habibian, Luca Benini, Yawei LiICML 2026 · 5 citations
Builds on14
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- A ConvNet for the 2020sZhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer et al.CVPR 2022 · 6,782 citations
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech et al.NeurIPS 2022 · 6,707 citations
- Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question AnsweringPan Lu, Swaroop Mishra, Tanglin Xia, Liang Qiu et al.NeurIPS 2022 · 2,727 citations
Related papers
- LLaVA-KD: A Framework of Distilling Multimodal Large Language ModelsYuxuan Cai, Jiangning Zhang, Haoyang He, Xinwei He et al.ICCV 2025 · 9 citations
- D2MoRA: Diversity-Regulated Asymmetric MoE-LoRA Decomposition for Efficient Multi-Task AdaptationJianhui Zuo, Xuemeng Song, Haokun Wen, Meng Liu et al.AAAI 2026
- Each Rank Could be an Expert: Single-Ranked Mixture of Experts LoRA for Multi-task LearningZiyu Zhao, Yixiao Zhou, Xin Yu, Zhi Zhang et al.KDD 2026 · 13 citations
- Beyond Next-Token Alignment: Distilling Multimodal Large Language Models via Token InteractionsLin Chen, zhaoxiaoke, Kun Ding, Weiwei Feng et al.ICML 2026 · 4 citations
- Visual Perception by Large Language Model's WeightsFeipeng Ma, Hongwei Xue, Yizhou Zhou, Guangting Wang et al.NeurIPS 2024 · 24 citations
