Unifying Specialized Visual Encoders for Video Language Models
Jihoon Chung, Tyler Zhu, Max Gonzalez Saez-Diez, Juan Carlos Niebles, Honglu Zhou, Olga Russakovsky
摘要
Recent advances in vision backbones have yielded powerful and diverse visual and video encoders. Yet, current Video Large Language Models encode visual inputs using an encoder from a single backbone family, limiting the amount and type of visual information they can process. We propose MERV, a Multi-Encoder Video Representation, which utilizes multiple encoders for a comprehensive video representation. To optimize heterogeneous features from a broad spectrum of encoders and ensure efficient and coherent feature integration, MERV first aligns encoder features spatio-temporally, then projects them into a unified structure, and finally fuses them through cross-attention. Under fair comparison, MERV achieves up to 4.62% higher accuracy than its base model, while introducing minimal extra parameters and training faster than equivalent singleencoder methods after parallelizing visual processing. Qualitative analysis shows MERV successfully captures and integrates domain knowledge from each encoder, opening new possibilities for scaling enhanced video understanding.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Causality Matters: How Temporal Information Emerges in Video Language ModelsYumeng Shi, Quanyu Long, Yin Wu, Wenya WangAAAI 2026 · 被引用 4 次
- Unleashing the Intrinsic Visual Representation Capability of Multimodal Large Language ModelsHengzhuang Li, Xinsong Zhang, QIMING PENG, Bin Luo 等CVPR 2026 · 被引用 2 次
- State-Space Hierarchical Compression with Gated Attention and Learnable Sampling for Hour-Long Video Understanding in Large Multimodal ModelsGeewook Kim, Minjoon SeoAAAI 2026 · 被引用 1 次
- Exploring the Visual Feature Space for Multimodal Neural DecodingWeihao Xia, A. Cengiz ÖztireliICCV 2025 · 被引用 1 次
- SegMo: Co-Designing Content-Aware Sparsity and Locally-Cohesive Segment Parallelism for Efficient VLM InferenceHaojuan Li, Ruohan Tang, Dongzhou Cheng, Zongpu Zhang 等CVPR 2026
它引用的顶会 Paper29
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech 等NeurIPS 2022 · 被引用 6,707 次
- ViViT: A Video Vision TransformerAnurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun 等ICCV 2021 · 被引用 2,947 次
相关 Paper
- Mavors: Multi-granularity Video Representation for Multimodal Large Language ModelYang Shi, Jiaheng Liu, Yushuo Guan, Zhenhua Wu 等ACM MM 2025 · 被引用 1 次
- Multiview Transformers for Video RecognitionShen Yan, Xuehan Xiong, Anurag Arnab, Zhichao Lu 等CVPR 2022 · 被引用 279 次
- Video-LLaVA: Learning United Visual Representation by Alignment Before ProjectionBin Lin, Yang Ye, Bin Zhu, Jiaxi Cui 等EMNLP 2024 · 被引用 231 次
- MoVE-KD: Knowledge Distillation for VLMs with Mixture of Visual EncodersJiajun Cao, Yuan Zhang, Tao Huang, Ming Lu 等CVPR 2025
- FlashVID: Efficient Video Large Language Models via Training-free Tree-based Spatiotemporal Token MergingZiyang Fan, Keyu Chen, Ruilong Xing, Yulin Li 等ICLR 2026 · 被引用 15 次
