ICLR2026

ScaleLong: A Multi-Timescale Benchmark for Long Video Understanding

David Ma, Huaqing Yuan, Xingjian Wang, Qianbo Zang, Tianci Liu, Xinyang He, Yanbin Wei, Jiawei Guo, nijiahui, Zhenzhu Yang, Meng Cao, Shanghaoran Quan, Yizhi LI, Wangchunshu Zhou, Jiaheng Liu, Wenhao Huang, Ge Zhang, Shiwen Ni, Xiaojie Jin

5 citations

DOI arXiv Publisher

Abstract

Although long-video understanding demands that models capture hierarchical temporal information-from clip (seconds) and shot (tens of seconds) to event (minutes) and story (hours)-existing benchmarks either neglect this multi-scale design or scatter scale-specific questions across different videos, preventing direct comparison of model performance across timescales on the same content. To address this, we introduce ScaleLong, the first benchmark to disentangle these factors by embedding questions targeting four hierarchical timescales-clip (seconds), shot (tens of seconds), event (minutes), and story (hours)-all within the same video content. This within-content multi-timescale questioning design enables direct comparison of model performance across timescales on identical videos. ScaleLong features 269 long videos (avg. 86 min) from 5 main categories and 36 sub-categories, with 4-8 carefully designed questions, with at least one question targeting each timescale. Evaluating 23 MLLMs reveals a distinct U-shaped performance trend: higher accuracy at the shortest (clip) and longest (story) timescales, with a dip at intermediate levels. Furthermore, ablation studies demonstrate that increased visual token capacity consistently enhances reasoning across all timescales. ScaleLong offers a crucial fine-grained, multi-timescale benchmark for advancing MLLM capabilities in long-video understanding. The code and dataset are available at https://github.com/multimodal-art-projection/ScaleLong (b)Video categories (a)Task Distribution Figure 1. (a) Task distribution in LongVideoBench. LongVideoBench consists of a total of 5 tasks, ensuring comprehensive evaluation of the model's capabilities. (b) Video Categories. LongVideoBench includes videos spanning 5 major categories and 37 subcategories, ensuring diverse topical coverage.