ScaleLong: A Multi-Timescale Benchmark for Long Video Understanding
David Ma, Huaqing Yuan, Xingjian Wang, Qianbo Zang, Tianci Liu, Xinyang He, Yanbin Wei, Jiawei Guo, nijiahui, Zhenzhu Yang, Meng Cao, Shanghaoran Quan
Abstract
Although long-video understanding demands that models capture hierarchical temporal information—from clip and shot to event and story—existing benchmarks either neglect this multi-scale design or scatter scale-specific questions across different videos, preventing direct comparison of model performance across timescales on the same content. To address this, we introduce ScaleLong, the first benchmark to disentangle these factors by embedding questions targeting four hierarchical timescalesclip, shot, event, and storyall within the same video content. This within-content multi-timescale questioning design enables direct comparison of model performance across timescales on identical videos. ScaleLong features 269 long videos (avg. 86 min) from 5 main categories and 36 sub-categories, with 4–8 carefully designed questions, with at least one question targeting each timescale. Evaluating 23 MLLMs reveals a distinct U-shaped performance trend: higher accuracy at the shortest (clip) and longest (story) timescales, with a dip at intermediate levels. Furthermore, ablation studies demonstrate that increased visual token capacity consistently enhances reasoning across all timescales. ScaleLong offers a crucial fine-grained, multi-timescale benchmark for advancing MLLM capabilities in long-video understanding. The code and dataset are available at https://github.com/multimodal-art-projection/ScaleLong
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get cb0d4604-cf53-4587-b30d-740f7e1bb0acCited by top-tier papers4
- Thinking with Drafts: Speculative Temporal Reasoning for Efficient Long Video UnderstandingPengfei Hu, Meng Cao, Yingyao Wang, Yi Wang et al.CVPR 2026 · 3 citations
- VideoTrace-R1: Long Video-based Retrieval-Augmented Generation via Reinforcement LearningZongsheng Cao, Anran Liu, Jun Xie, Feng Chen et al.ICML 2026
- Think in Cloud, Look at Edges: Semantic-Driven Query Decomposition for Efficient Video ReasoningWenhao Zou, Zhijie Cai, Minchen Yu, Zongshuai Zhang et al.ICML 2026
- Seeing the Scene Matters: Revealing Forgetting in Video Understanding Models with a Scene-Aware Long-Video BenchmarkSeng Nam Chen, Hao Chen, Chenglam Ho, Xinyu Mao et al.CVPR 2026
Related papers
- ALLVB: All-in-One Long Video Understanding BenchmarkXichen Tan, Yuanjing Luo, Yunfan Ye, Fang Liu et al.AAAI 2025 · 13 citations
- CG-Bench: Clue-grounded Question Answering Benchmark for Long Video UnderstandingGuo Chen, Yicheng Liu, Yifei Huang, Baoqi Pei et al.ICLR 2025
- Unleashing Hour-Scale Video Training for Long Video-Language UnderstandingJingyang Lin, Jialian Wu, Ximeng Sun, Ze Wang et al.NeurIPS 2025 · 25 citations
- MLVU: Benchmarking Multi-task Long Video UnderstandingJunjie Zhou, Yan Shu, Bo Zhao, Boya Wu et al.CVPR 2025
- Universal Video Temporal Grounding with Generative Multi-modal Large Language ModelsZeqian Li, Shangzhe Di, Zhonghua Zhai, Weilin Huang et al.NeurIPS 2025 · 30 citations
