ShotBench: Expert-Level Cinematic Understanding in Vision-Language Models
Hongbo Liu, Jingwen He, Yi Jin, Dian Zheng, Yuhao Dong, Fan Zhang, Ziqi Huang, Yinan He, Weichao Chen, Yu Qiao, Wanli Ouyang, Shengjie Zhao, Ziwei Liu
Abstract
Cinematography, the fundamental visual language of film, is essential for conveying narrative, emotion, and aesthetic quality. While recent Vision-Language Models (VLMs) demonstrate strong general visual understanding, their proficiency in comprehending the nuanced cinematic grammar embedded within individual shots remains largely unexplored and lacks robust evaluation. This critical gap limits both fine-grained visual comprehension and the precision of AI-assisted video generation. To address this, we introduce ShotBench, a comprehensive benchmark specifically designed for cinematic language understanding. It features over 3.5k expert-annotated QA pairs from images and video clips, meticulously curated from over 200 acclaimed (predominantly Oscar-nominated) films and spanning eight key cinematography dimensions. Our evaluation of 24 leading VLMs on ShotBench reveals their substantial limitations: even the top-performing model achieves less than 60% average accuracy, particularly struggling with fine-grained visual cues and complex spatial reasoning. To catalyze advancement in this domain, we construct ShotQA, a large-scale multimodal dataset comprising approximately 70k cinematic QA pairs. Leveraging ShotQA, we develop ShotVL through supervised fine-tuning and Group Relative Policy Optimization. ShotVL significantly outperforms all existing open-source and proprietary models on ShotBench, establishing new state-of-the-art performance. We open-source our models, data, and code to foster rapid progress in this crucial area of AI-driven cinematic understanding and generation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers5
- ViStoryBench: Comprehensive Benchmark Suite for Story VisualizationCailin Zhuang, Ailin Huang, Hu Yaoqi, Jingwei Wu et al.CVPR 2026 · 37 citations
- UniVBench: Towards Unified Evaluation for Video Foundation ModelsJianhui Wei, Xiaotian Zhang, Yichen Li, Yuan Wang et al.CVPR 2026 · 12 citations
- Seeing without Pixels: Perception from Camera TrajectoriesZihui Xue, Kristen Grauman, Dima Damen, Andrew Zisserman et al.CVPR 2026 · 4 citations
- LiViBench: An Omnimodal Benchmark for Interactive Livestream Video UnderstandingXiaodong Wang, Langling Huang, Zhirong Wu, Xu Zhao et al.AAAI 2026 · 1 citation
- JoPPO: Hierarchical Photography Assessment via Contrastive Joint Conditional Probabilistic Reinforcement LearningYifan Yang, Juntuo Wang, Yuming Qiao, Xudong Zhang et al.CVPR 2026
Builds on19
- InstructBLIP: Towards General-purpose Vision-Language Models with Instruction TuningWenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong et al.NeurIPS 2023 · 4,013 citations
- FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningTri DaoICLR 2024 · 2,600 citations
- Visual-RFT: Visual Reinforcement Fine-TuningZiyu Liu, Zeyi Sun, Yuhang Zang, Xiaoyi Dong et al.ICCV 2025 · 563 citations
- VL-Rethinker: Incentivizing Self-Reflection of Vision-Language Models with Reinforcement LearningHaozhe Wang, Chao Qu, Zuming Huang, Wei Chu et al.NeurIPS 2025 · 356 citations
- DigiRL: Training In-The-Wild Device-Control Agents with Autonomous Reinforcement LearningHao Bai, Yifei Zhou, Jiayi Pan, Mert Cemri et al.NeurIPS 2024 · 239 citations
Related papers
- VCapsBench: A Large-scale Fine-grained Benchmark for Video Caption Quality EvaluationShi-Xue Zhang, Hongfa Wang, Duojun Huang, Xin Li et al.AAAI 2026 · 5 citations
- VidComposition: Can MLLMs Analyze Compositions in Compiled Videos?Yunlong Tang, Junjia Guo, Hang Hua, Susan Liang et al.CVPR 2025
- VL-RewardBench: A Challenging Benchmark for Vision-Language Generative Reward ModelsLei Li, Yuancheng Wei, Zhihui Xie, Xuqing Yang et al.CVPR 2025
- V2P-Bench: Evaluating Video-Language Understanding with Visual Prompts for Better Human-Model InteractionYiming Zhao, Yu Zeng, Yukun Qi, YaoYang Liu et al.ICLR 2026 · 8 citations
- Seeing Across Views: Benchmarking Spatial Reasoning of Vision-Language Models in Robotic ScenesZhiYuan Feng, Zhaolu Kang, Qijie Wang, Zhiying Du et al.ICLR 2026 · 23 citations
