Apollo: An Exploration of Video Understanding in Large Multimodal Models
Orr Zohar, Xiaohan Wang, Yann Dubois, Nikhil Mehta, Tong Xiao, Philippe Hansen-Estruch, Licheng Yu, Xiaofang Wang, Felix Juefei-Xu, Ning Zhang, Serena Yeung-Levy, Xide Xia
摘要
Despite the rapid integration of video perception capabilities into Large Multimodal Models (LMMs), what drives their video perception remains poorly understood. Consequently, many design decisions in this domain are made without proper justification or analysis. The high computational cost of training and evaluating such models and limited open research hinder the development of video-LMMs. To address this, we present a comprehensive study that helps uncover what effectively drives video understanding in LMMs. We begin by critically examining the primary contributors to the high computational requirements associated with video-LMM research and discover Scaling Consistency, wherein design and training decisions made on smaller models and datasets (up to a critical size) effectively transfer to larger models. Leveraging these insights, we explored many video-specific aspects of video-LMMs, including video sampling, architectures, data composition, training schedules, and more. Guided by these findings, we introduce Apollo, a state-of-the-art family of LMMs that achieve superior performance across different model sizes. Our models process over 1-hour videos efficiently, with the 3B parameter variant outperforming most existing 7B models. Apollo-7B is state-of-the-art compared to 7B LMMs with a 70.9 on MLVU, and 63.3 on Video-MME.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper44
- Molmo2: Open Weights and Data for Vision-Language Models with Video Understanding and GroundingChristopher Clark, Jieyu Zhang, Zixian Ma, Jae Sung Park 等CVPR 2026 · 被引用 144 次
- Cambrian-S: Towards Spatial Supersensing in VideoShusheng Yang, Jihan Yang, Pinzhi Huang, Ellis Brown 等ICLR 2026 · 被引用 139 次
- PerceptionLM: Open-Access Data and Models for Detailed Visual UnderstandingJang Hyun Cho, Andrea Madotto, Effrosyni Mavroudi, Triantafyllos Afouras 等NeurIPS 2025 · 被引用 97 次
- Scaling RL to Long VideosYukang Chen, Wei Huang, Baifeng Shi, Qinghao Hu 等NeurIPS 2025 · 被引用 91 次
- SpatialLM: Training Large Language Models for Structured Indoor ModelingYongsen Mao, Junhao Zhong, Chuan Fang, Jia Zheng 等NeurIPS 2025 · 被引用 89 次
它引用的顶会 Paper21
- MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language ModelsDeyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li 等ICLR 2024 · 被引用 3,079 次
- Sigmoid Loss for Language Image Pre-TrainingXiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, Lucas BeyerICCV 2023 · 被引用 2,932 次
- VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-TrainingZhan Tong, Yibing Song, Jue Wang, Limin WangNeurIPS 2022 · 被引用 2,336 次
- Perceiver: General Perception with Iterative AttentionAndrew Jaegle, Felix Gimeno, Andy Brock, Oriol Vinyals 等ICML 2021 · 被引用 1,399 次
- LanguageBind: Extending Video-Language Pretraining to N-modality by Language-based Semantic AlignmentBin Zhu, Bin Lin, Munan Ning, Yang Yan 等ICLR 2024 · 被引用 403 次
相关 Paper
- Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video AnalysisChaoyou Fu, Yuhan Dai, Yongdong Luo, Lei Li 等CVPR 2025
- ReMoRa: Multimodal Large Language Model based on Refined Motion Representation for Long-Video UnderstandingDaichi Yashima, Shuhei Kurita, Yusuke Oda, Komei SugiuraCVPR 2026 · 被引用 6 次
- Improving LLM Video Understanding with 16 Frames Per SecondYixuan Li, Changli Tang, Jimin Zhuang, Yudong Yang 等ICML 2025
- HERMES: KV Cache as Hierarchical Memory for Efficient Streaming Video UnderstandingHaowei Zhang, Shudong Yang, Jinlan Fu, See-Kiong Ng 等ACL 2026 · 被引用 17 次
- Free-Moref: Instantly Multiplexing Context Perception Capabilities of Video-Mllms Within Single InferenceKuo Wang, Quanlong Zheng, Junlin Xie, Yanhao Zhang 等ICCV 2025
