Understanding Long Videos with Multimodal Language Models
Kanchana Ranasinghe, Xiang Li, Kumara Kahatapitiya, Michael S. Ryoo
Abstract
Large Language Models (LLMs) have allowed recent LLM-based approaches to achieve excellent performance on long-video understanding benchmarks. We investigate how extensive world knowledge and strong reasoning skills of underlying LLMs influence this strong performance. Surprisingly, we discover that LLMbased approaches can yield surprisingly good accuracy on long-video tasks with limited video information, sometimes even with no video-specific information. Building on this, we explore injecting video-specific information into an LLMbased framework. We utilize off-the-shelf vision tools to extract three objectcentric information modalities from videos, and then leverage natural language as a medium for fusing this information. Our resulting Multimodal Video Understanding (MVU) framework demonstrates state-of-the-art performance across multiple video understanding benchmarks. Strong performance also on robotics domain tasks establishes its strong generality. Code: github.com/kahnchana/mvu 🖼 Selected Frames 💬 Ques/on 💬 Candidates 🖼 Center Frame LLM VLM 💬 Ques/on 💬 Candidates Just LLM Single Frame VLM 💬 Ques/on 💬 Candidates 🎞 Video Mul3modal Video Understanding (MVU)
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers6
- Logic-in-Frames: Dynamic Keyframe Search via Visual Semantic-Logical Verification for Long Video UnderstandingWeiyu Guo, Ziyang Chen, Shaoguang Wang, JianXiang He et al.NeurIPS 2025 · 35 citations
- Scaling up Memory for Robotic Control via Experience RetrievalAjay Sridhar, Jennifer Pan, Satvik Sharma, Chelsea FinnICLR 2026 · 21 citations
- Compass: SLO-aware Query Planner for Compound AI Serving at ScaleBanruo Liu, Wei-Yu Lin, Minghao Fang, Yihan Jiang et al.VLDB 2026 · 5 citations
- On Discriminative vs. Generative classifiers: Rethinking MLLMs for Action UnderstandingZhanzhong Pang, Dibyadip Chatterjee, Fadime Sener, Angela YaoICLR 2026 · 1 citation
- Re-thinking Temporal Search for Long-Form Video UnderstandingJinhui Ye, Zihan Wang, Haosen Sun, Keshigeyan Chandrasegaran et al.CVPR 2025
Builds on35
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- ViperGPT: Visual Inference via Python Execution for ReasoningDídac Surís, Sachit Menon, Carl VondrickICCV 2023 · 732 citations
- Large Language Models as Commonsense Knowledge for Large-Scale Task PlanningZirui Zhao, Wee Sun Lee, David HsuNeurIPS 2023 · 423 citations
Related papers
- Visual Context Window Extension: A New Perspective for Long Video UnderstandingHongchen Wei, Zhenzhong ChenACM MM 2025
- DrVideo: Document Retrieval Based Long Video UnderstandingZiyu Ma, Chenhui Gou, Hengcan Shi, Bin Sun et al.CVPR 2025
- MA-LMM: Memory-Augmented Large Multimodal Model for Long-Term Video UnderstandingBo He, Hengduo Li, Young Kyun Jang, Menglin Jia et al.CVPR 2024
- Video Summarization with Large Language ModelsMin Jung Lee, Dayoung Gong, Minsu ChoCVPR 2025
- OmAgent: A Multi-modal Agent Framework for Complex Video Understanding with Task Divide-and-ConquerLu Zhang, Tiancheng Zhao, Heting Ying, Yibo Ma et al.EMNLP 2024 · 9 citations
