Eve: Efficient Multimodal Vision Language Models with Elastic Visual Experts
Miao Rang, Zhenni Bi, Chuanjian Liu, Yehui Tang, Kai Han, Yunhe Wang
Abstract
Multimodal vision language models (VLMs) have made significant progress with the support of continuously increasing model sizes and data volumes. Running VLMs on edge devices has become a challenge for their widespread application. There are several efficient VLM efforts, but they often sacrifice linguistic capabilities to enhance multimodal abilities, or require extensive training. To address this quandary, we introduce the innovative framework of Efficient Vision Language Models with Elastic Visual Experts (Eve). By strategically incorporating adaptable visual expertise at multiple stages of training, Eve strikes a balance between preserving linguistic abilities and augmenting multimodal capabilities. This balanced approach results in a versatile model with only 1.8B parameters that delivers significant improvements in both multimodal and linguistic tasks. Notably, in configurations below 3B parameters, Eve distinctly outperforms in language benchmarks and achieves state-of-the-art results 68.87% in VLM Benchmarks. Additionally, its multimodal accuracy outstrips that of the larger 7B LLaVA-1.5 model. Our code is available at https://github.com/rangmiao/Eve .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4b07bde3-b15b-41cb-8d5e-d82eedebfe50Cited by top-tier papers3
- Circle-RoPE: Cone-like Decoupled Rotary Positional Embedding for Vision-Language ModelsChengcheng Wang, Jianyuan Guo, Hongguang Li, Yuchuan Tian et al.ICML 2026 · 14 citations
- EvoMoE: Expert Evolution in Mixture of Experts for Multimodal Large Language ModelsLinglin Jing, Yuting Gao, Zhigang Wang, Wang Lan et al.AAAI 2026 · 4 citations
- MedFG-VQA: Low-Frequency Memory and Graph Attention for Lightweight Medical VQAHaowen Gu, Gensheng Pei, Zeren Sun, Mingwu Ren et al.CVPR 2026 · 2 citations
Builds on16
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 6,549 citations
Related papers
- Unveiling Encoder-Free Vision-Language ModelsHaiwen Diao, Yufeng Cui, Xiaotong Li, Yueze Wang et al.NeurIPS 2024 · 107 citations
- DUET-VLM: Dual stage Unified Efficient Token reduction for VLM Training and InferenceAditya Kumar Singh, Hitesh Kandala, Pratik Prabhanjan Brahma, Zicheng Liu et al.CVPR 2026
- A Comprehensive Overhaul of Multimodal Assistant with Small Language ModelsMinjie Zhu, Yichen Zhu, Ning Liu, Xin Liu et al.AAAI 2025 · 30 citations
- Improved Baselines with Visual Instruction TuningHaotian Liu, Chunyuan Li, Yuheng Li, Yong Jae LeeCVPR 2024
- Feast Your Eyes: Mixture-of-Resolution Adaptation for Multimodal Large Language ModelsGen Luo, Yiyi Zhou, Yuxin Zhang, Xiawu Zheng et al.ICLR 2025
