MovieCORE: COgnitive REasoning in Movies
Gueter Josmy Faure, Min-Hung Chen, Jia-Fong Yeh, Ying Cheng, Hung-Ting Su, Yung-Hao Tang, Shang-Hong Lai, Winston H. Hsu
Abstract
Contrasts wisdom and warmth (old age) with energy and curiosity (youth). (Q: How are intergenerational themes demonstrated through specific scenes in the video?) Captures shift in emotional state due to external factors. (Q: How do changes in settings impact the elderly character's emotions and sense of identity?) Emotional/Psychological States Character Contrasts Cause-Effect Relationships Figure 1: Beyond Shallow Video Understanding: The proposed benchmark, MovieCORE, challenges visionlanguage models (VLMs) to understand the subtle interplay between emotions (Top, Middle), character dynamics and causality (Middle, Bottom), and psychological complexity (Top, Middle). From empathy to introspection, from wisdom to curiosity MovieCORE tests VLMs' ability to comprehend the deeper elements of movies.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on8
- InstructBLIP: Towards General-purpose Vision-Language Models with Instruction TuningWenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong et al.NeurIPS 2023 · 4,013 citations
- Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language ModelsMuhammad Maaz, Hanoona Abdul Rasheed, Salman Khan, Fahad KhanACL 2024 · 279 citations
- MAD: A Scalable Dataset for Language Grounding in Videos from Movie Audio DescriptionsMattia Soldan, Alejandro Pardo, Juan León Alcázar, Fabian Caba Heilbron et al.CVPR 2022 · 84 citations
- HERMES: Temporal-Coherent Long-form Understanding with Episodes and SemanticsGueter Josmy Faure, Jia-Fong Yeh, Min-Hung Chen, Hung-Ting Su et al.ICCV 2025 · 2 citations
- Towards Long-Form Video UnderstandingChao-Yuan Wu, Philipp KrähenbühlCVPR 2021
Related papers
- HumanVBench: Probing Human-Centric Video Understanding in MLLMs with Automatically Synthesized BenchmarksTing Zhou, Daoyuan Chen, Qirui Jiao, Bolin Ding et al.CVPR 2026
- IDA-VLM: Towards Movie Understanding via ID-Aware Large Vision-Language ModelYatai Ji, Shilong Zhang, Jie Wu, Peize Sun et al.ICLR 2025
- ShotBench: Expert-Level Cinematic Understanding in Vision-Language ModelsHongbo Liu, Jingwen He, Yi Jin, Dian Zheng et al.NeurIPS 2025 · 24 citations
- VidComposition: Can MLLMs Analyze Compositions in Compiled Videos?Yunlong Tang, Junjia Guo, Hang Hua, Susan Liang et al.CVPR 2025
- MotionBench: Benchmarking and Improving Fine-grained Video Motion Understanding for Vision Language ModelsWenyi Hong, Yean Cheng, Zhuoyi Yang, Weihan Wang et al.CVPR 2025
