Investigating Redundancy in Multimodal Large Language Models with Multiple Vision Encoders
Yizhou Wang, Song Mao, Yang Chen, Yufan Shen, Pinlong Cai, Ding Wang, Guohang Yan, Zhi Yu, Yinqiao Yan, Xuming Hu, Botian Shi
摘要
Recent multimodal large language models (MLLMs) increasingly integrate multiple vision encoders to improve performance on various benchmarks, assuming that diverse pretraining objectives yield complementary visual signals. However, we show this assumption often fails in practice. Through systematic encoder masking across representative multi-encoder MLLMs, we find that performance typically degrades gracefully-and sometimes even improves-when selected encoders are masked, revealing pervasive encoder redundancy. To quantify this effect, we introduce two principled metrics: the Conditional Utilization Rate (CUR), which measures an encoder's marginal contribution in the presence of others, and the Information Gap (IG), which captures heterogeneity in encoder utility within a model. Using these tools, we observe: (i) strong specialization on tasks like OCR & Chart, where a single encoder can dominate with a CUR > 90%, (ii) high redundancy on general VQA and knowledge-based tasks, where encoders are largely interchangeable, (iii) instances of detrimental encoders with negative CUR. Notably, masking specific encoders can yield up to 16% higher accuracy on a specific task category and 3.6% overall performance boost compared to the full model. Furthermore, single-and dual-encoder variants recover over 90% of baseline on most non-OCR tasks with substantially lower training resources and inference latency. Our analysis challenges the "more encoders are better" heuristic in MLLMs and provides actionable diagnostics for developing more efficient and effective multimodal architectures. The project website is available at https://github.com/MaoSong2022/Encoder-Redundancy .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper20
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao 等ICCV 2023 · 被引用 13,211 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa 等ICML 2021 · 被引用 8,974 次
- A ConvNet for the 2020sZhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer 等CVPR 2022 · 被引用 6,782 次
相关 Paper
- Deciphering Cross-Modal Alignment in Large Vision-Language Models Via Modality Integration RateQidong Huang, Xiaoyi Dong, Pan Zhang, Yuhang Zang 等ICCV 2025 · 被引用 8 次
- METEOR: Multi-Encoder Collaborative Token Pruning for Efficient Vision Language ModelsYuchen Liu, Yaoming Wang, Bowen Shi, Xiaopeng Zhang 等ICCV 2025 · 被引用 2 次
- Skip-It? Theoretical Conditions for Layer Skipping in Vision–Language ModelsMax Hartman, Vidhata Jayaraman, Moulik Choraria, Akhil Bhimaraju 等ICML 2026 · 被引用 1 次
- MoDES: Accelerating Mixture-of-Experts Multimodal Large Language Models via Dynamic Expert SkippingYushi Huang, Zining Wang, Zhihang Yuan, Yifu Ding 等CVPR 2026 · 被引用 15 次
- Multi-modal Data Spectrum: Multi-modal Datasets are Multi-dimensionalDivyam Madaan, Varshan Muhunthan, Kyunghyun Cho, Sumit ChopraICLR 2026 · 被引用 3 次
