Breaking the Encoder Barrier for Seamless Video-Language Understanding
Handong Li, Yiyuan Zhang, Longteng Guo, Xiangyu Yue, Jing Liu
Abstract
Most Video-Large Language Models (Video-LLMs) adopt an encoder-decoder framework, where a vision encoder extracts frame-wise features for processing by a language model. However, this approach incurs high computational costs, introduces resolution biases, and struggles to capture fine-grained multimodal interactions. To overcome these limitations, we propose ELVA, an encoder-free Video-LLM that directly models nuanced video-language interactions without relying on a vision encoder. ELVA employs token merging to construct a bottom-up hierarchical representation and incorporates a video guidance supervisor for direct spatiotemporal representation learning. Additionally, a hybrid-resolution mechanism strategically integrates high- and low-resolution frames as inputs to achieve an optimal balance between performance and efficiency. With only 7M publicly available video-text pairs, ELVA achieves performance on par with encoder-based Video-LLMs while reducing FLOPs by up to 95% and inference latency by 92%, offering a scalable and efficient solution for real-time video understanding.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- Exploring the Potential of Encoder-free Architectures in 3D LMMsYiwen Tang, Ziyu Guo, Zhuhao Wang, Renrui Zhang et al.ICLR 2026 · 19 citations
- See&Trek: Training-Free Spatial Prompting for Multimodal Large Language ModelPengteng Li, Pinhao Song, Wuyang Li, Huizai Yao et al.NeurIPS 2025 · 13 citations
- VideoNSA: Native Sparse Attention Scales Video UnderstandingEnxin Song, Wenhao Chai, Shusheng Yang, Ethan Armand et al.ICLR 2026 · 11 citations
- AdaSpark: Adaptive Sparsity for Efficient Long-Video UnderstandingHandong Li, Zikang Liu, Longteng Guo, Tongtian Yue et al.CVPR 2026
Builds on15
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- InstructBLIP: Towards General-purpose Vision-Language Models with Instruction TuningWenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong et al.NeurIPS 2023 · 4,013 citations
- Sigmoid Loss for Language Image Pre-TrainingXiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, Lucas BeyerICCV 2023 · 2,932 citations
- Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language ModelsMuhammad Maaz, Hanoona Abdul Rasheed, Salman Khan, Fahad KhanACL 2024 · 279 citations
- Video-LLaVA: Learning United Visual Representation by Alignment Before ProjectionBin Lin, Yang Ye, Bin Zhu, Jiaxi Cui et al.EMNLP 2024 · 231 citations
Related papers
- FlashVID: Efficient Video Large Language Models via Training-free Tree-based Spatiotemporal Token MergingZiyang Fan, Keyu Chen, Ruilong Xing, Yulin Li et al.ICLR 2026 · 15 citations
- EarlyTom: Early Token Compression Completes Fast Video UnderstandingHesong Wang, Xin Jin, Lu Lu, Chenhaowen Li et al.CVPR 2026 · 7 citations
- iLLaVA: An Image is Worth Fewer Than 1/3 Input Tokens in Large Multimodal ModelsLianyu Hu, Liqing Gao, Fanhua Shang, Liang Wan et al.ICLR 2026 · 9 citations
- Video-Panda: Parameter-efficient Alignment for Encoder-free Video-Language ModelsJinhui Yi, Syed Talal Wasim, Yanan Luo, Muzammal Naseer et al.CVPR 2025
- AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and PruningYiwu Zhong, Zhuoming Liu, Yin Li, Liwei WangICCV 2025 · 1 citation
