Collaborative Training of Tiny-Large Vision Language Models
Shichen Lu, Longteng Guo, Wenxuan Wang, Zijia Zhao, Tongtian Yue, Jing Liu, Si Liu
Abstract
Recently, large vision language models (LVLMs) have advanced AI by integrating visual and linguistic data for tasks like visual conversation, image captioning, and visual question answering. Current LVLM research either scales up model size for performance or reduces parameters for limited computational resources. We believe both large and tiny models have unique strengths and that collaborative training yields better results than independent training. We propose Collaborative Training of Tiny-Large Vision Language Models (CTVLMs), a framework connecting large and tiny models via a projection layer and leveraging a synergistic training strategy. Our framework improves training efficiency by strengthening the interconnection between large and tiny models. Using the parameter efficiency of tiny models, we effectively align image-text features, then apply knowledge distillation to help large models better align cross-modal information. During fine-tuning, the large model's extensive knowledge enhances tiny model's performance. This collaborative approach allows models to adapt to various computational resources and outperforms existing methods in vision-language tasks.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Related papers
- Large-Small Model Synergy with Multimodal Fine-Grained Heuristics for Knowledge-Based Visual Question AnsweringZhongfan Sun, Kan Guo, Yongli Hu, Daxin Tian et al.ACM MM 2025
- VL2Lite: Task-Specific Knowledge Distillation from Large Vision-Language Models to Lightweight NetworksJinseong Jang, Chunfei Ma, Byeongwon LeeCVPR 2025
- A-VL: Adaptive Attention for Large Vision-Language ModelsJunyang Zhang, Mu Yuan, Ruiguang Zhong, Puhan Luo et al.AAAI 2025 · 6 citations
- Align-KD: Distilling Cross-Modal Alignment Knowledge for Mobile Vision-Language Large Model EnhancementQianhan Feng, Wenshuo Li, Tong Lin, Xinghao ChenCVPR 2025
- LLaVA-KD: A Framework of Distilling Multimodal Large Language ModelsYuxuan Cai, Jiangning Zhang, Haoyang He, Xinwei He et al.ICCV 2025 · 9 citations
