Movie101: A New Movie Understanding Benchmark
Zihao Yue, Qi Zhang, Anwen Hu, Liang Zhang, Ziheng Wang, Qin Jin
Abstract
To help the visually impaired enjoy movies, automatic movie narrating systems are expected to narrate accurate, coherent, and role-aware plots when there are no speaking lines of actors. Existing works benchmark this challenge as a normal video captioning task via some simplifications, such as removing role names and evaluating narrations with ngram-based metrics, which makes it difficult for automatic systems to meet the needs of real application scenarios. To narrow this gap, we construct a large-scale Chinese movie benchmark, named Movie101. Closer to real scenarios, the Movie Clip Narrating (MCN) task in our benchmark asks models to generate role-aware narration paragraphs for complete movie clips where no actors are speaking. External knowledge, such as role information and movie genres, is also provided for better movie understanding. Besides, we propose a new metric called Movie Narration Score (MNScore) for movie narrating evaluation, which achieves the best correlation with human evaluation. Our benchmark also supports the Temporal Narration Grounding (TNG) task to investigate clip localization given text descriptions. For both two tasks, our proposed methods well leverage external knowledge and outperform carefully designed baselines. The dataset and codes are released at https://github.com/yuezih/Movie101 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7cb5243d-f4e7-4f45-973a-e0d40abc0716Cited by top-tier papers5
- TimeChat: A Time-sensitive Multimodal Large Language Model for Long Video UnderstandingShuhuai Ren, Linli Yao, Shicheng Li, Xu Sun et al.CVPR 2024 · 83 citations
- TimeViper: A Hybrid Mamba-Transformer Vision-Language Model for Efficient Long Video UnderstandingBoshen Xu, Zihan Xiao, Jiaze Li, Jianzhong Ju et al.CVPR 2026 · 5 citations
- Movie101v2: Improved Movie Narration BenchmarkZihao Yue, Yepeng Zhang, Ziheng Wang, Qin JinACL 2025 · 5 citations
- VideoVista-CulturalLingo: 360° Horizons-Bridging Cultures, Languages, and Domains in Video ComprehensionXinyu Chen, Yunxin Li, Haoyuan Shi, Baotian Hu et al.ACL 2025
- MLVU: Benchmarking Multi-task Long Video UnderstandingJunjie Zhou, Yan Shu, Bo Zhao, Boya Wu et al.CVPR 2025
Builds on10
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video ClipsAntoine Miech, Dimitri Zhukov, Jean-Baptiste Alayrac, Makarand Tapaswi et al.ICCV 2019 · 1,437 citations
- CLIPScore: A Reference-free Evaluation Metric for Image CaptioningJack Hessel, Ari Holtzman, Maxwell Forbes, Ronan Le Bras et al.EMNLP 2021 · 937 citations
- Learning 2D Temporal Adjacent Networks for Moment Localization with Natural LanguageSongyang Zhang, Houwen Peng, Jianlong Fu, Jiebo LuoAAAI 2020 · 579 citations
- Detecting Moments and Highlights in Videos via Natural Language QueriesJie Lei, Tamara L. Berg, Mohit BansalNeurIPS 2021 · 425 citations
Related papers
- What You See is What You Ask: Evaluating Audio DescriptionsDivy Kala, Eshika Khandelwal, Makarand TapaswiEMNLP 2025
- Fine-grained Audible Video DescriptionXuyang Shen, Dong Li, Jinxing Zhou, Zhen Qin et al.CVPR 2023
- HowToNarrate: A General-Domain Benchmark for Synchronized Video Narration with External KnowledgeXueyan Wang, Dingyi Yang, Qin JinACL 2026
- AutoAD II: The Sequel - Who, When, and What in Movie Audio DescriptionTengda Han, Max Bain, Arsha Nagrani, Gül Varol et al.ICCV 2023 · 55 citations
- Synchronized Video Storytelling: Generating Video Narrations with Structured StorylineDingyi Yang, Chunru Zhan, Ziheng Wang, Biao Wang et al.ACL 2024
