Unleashing Hour-Scale Video Training for Long Video-Language Understanding
Jingyang Lin, Jialian Wu, Ximeng Sun, Ze Wang, Jiang Liu, Yusheng Su, Xiaodong Yu, Hao Chen, Jiebo Luo, Zicheng Liu, Emad Barsoum
Abstract
Recent long-form video-language understanding benchmarks have driven progress in video large multimodal models (Video-LMMs). However, the scarcity of well-annotated long videos has left the training of hour-long Video-LMMs underexplored. To close this gap, we present VideoMarathon, a large-scale hour-long video instruction-following dataset. This dataset includes around 9,700 hours of long videos sourced from diverse domains, ranging from 3 to 60 minutes per video. Specifically, it contains 3.3M high-quality QA pairs, spanning six fundamental topics: temporality, spatiality, object, action, scene, and event. Compared to existing video instruction datasets, VideoMarathon significantly extends training video durations up to 1 hour, and supports 22 diverse tasks requiring both short- and long-term video comprehension. Building on VideoMarathon, we propose Hour-LLaVA, a powerful and efficient Video-LMM for hour-scale video-language modeling. It enables hour-long video training and inference at 1-FPS sampling by leveraging a memory augmentation module, which adaptively integrates question-relevant and spatiotemporally informative semantics from the cached full video context. In our experiments, Hour-LLaVA achieves the best performance on multiple representative long video-language benchmarks, demonstrating the high quality of the VideoMarathon dataset and the superiority of the Hour-LLaVA model.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c113deeb-373d-4d53-9aa8-626ad598ea61Cited by top-tier papers5
- ReWatch-R1: Boosting Complex Video Reasoning in Large Vision-Language Models through Agentic Data SynthesisCongzhi Zhang, Zhibin Wang, Yinchao Ma, Jiawei Peng et al.ICLR 2026 · 24 citations
- TimeSearch-R: Adaptive Temporal Search for Long-Form Video Understanding via Self-Verification Reinforcement LearningJunwen Pan, Qizhe Zhang, Rui Zhang, Ming Lu et al.ICLR 2026 · 20 citations
- VideoSeek: Long-Horizon Video Agent with Tool-Guided SeekingJingyang Lin, Jialian Wu, Jiang Liu, Ximeng Sun et al.CVPR 2026 · 15 citations
- REVISOR: Beyond Textual Reflection, Towards Multimodal Introspective Reasoning in Long-Form Video UnderstandingJiaze Li, Hao Yin, Wenhui Tan, Jingyang Chen et al.CVPR 2026 · 14 citations
- Select Less, Reason More: Prioritizing Evidence Purity for Video ReasoningXuchen Li, Xuzhao Li, Shiyu Hu, Kaiqi HuangCVPR 2026 · 5 citations
Builds on28
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 4,104 citations
- Sigmoid Loss for Language Image Pre-TrainingXiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, Lucas BeyerICCV 2023 · 2,932 citations
- Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question AnsweringPan Lu, Swaroop Mishra, Tanglin Xia, Liang Qiu et al.NeurIPS 2022 · 2,727 citations
- HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video ClipsAntoine Miech, Dimitri Zhukov, Jean-Baptiste Alayrac, Makarand Tapaswi et al.ICCV 2019 · 1,437 citations
Related papers
- VISTA: Enhancing Long-Duration and High-Resolution Video Understanding by Video Spatiotemporal AugmentationWeiming Ren, Huan Yang, Jie Min, Cong Wei et al.CVPR 2025
- ProLongVid: A Simple but Strong Baseline for Long-context Video Instruction TuningRui Wang, Bohao Li, Xiyang Dai, Jianwei Yang et al.EMNLP 2025
- SALOVA: Segment-Augmented Long Video Assistant for Targeted Retrieval and Routing in Long-Form Video AnalysisJunho Kim, Hyunjun Kim, Hosu Lee, Yong Man RoCVPR 2025
- PPLLaVA: Varied Video Sequence Understanding With Prompt GuidanceShangkun Sun, Ruyang Liu, Haoran Tang, Yixiao Ge et al.ICLR 2026 · 21 citations
- VideoRefer Suite: Advancing Spatial-Temporal Object Understanding with Video LLMYuqian Yuan, Hang Zhang, Wentong Li, Zesen Cheng et al.CVPR 2025
