VEU-Bench: Towards Comprehensive Understanding of Video Editing
Bozheng Li, Yongliang Wu, Yi Lu, Jiashuo Yu, Licheng Tang, Jiawang Cao, Wenqing Zhu, Yuyang Sun, Jay Wu, Wenbo Zhu
Abstract
Level Q: What Shot type does the given clip belong to? A: Over-the-shoulder shot. Recognition Q: Clarify shot motion of given video and explain grounds for the judgment. A: The shot-motion of the given video is handheld. The camera exhibits slight vertical and horizontal jitters as it follows the girl walking and then moves to focus on the two boys conversing. Q: what is the cut type used in the given video, and what is its function? A: The video uses emphasis/deemphasis cut. By cutting from medium shot of reporter, Gayle Hailstorm, to close-up of Heather in car, the emphasis cut draws attention to Heather. have made great progress in general video understanding tasks, their capabilities in video editing understanding (VEU) tasks remain unexplored. To address this gap, in this paper, we introduce VEU-Bench (Video Editing Understanding Benchmark), a comprehensive benchmark that categorizes video editing components across various dimensions, from intra-frame features like shot size to intershot attributes such as cut types and transitions. Unlike previous video editing understanding benchmarks that focus mainly on editing element classification, VEU-Bench encompasses 19 fine-grained tasks across three stages: recognition, reasoning, and judging. To enhance the annotation of VEU automatically, we built an annotation pipeline inte-grated with an ontology-based knowledge base. Through extensive experiments with 11 state-of-the-art Vid-LLMs, our findings reveal that current Vid-LLMs face significant challenges in VEU tasks, with some performing worse than random choice. To alleviate this issue, we develop Oscars 1 , a VEU expert model fine-tuned on the curated VEU-Bench dataset. It outperforms existing open-source Vid-LLMs on VEU-Bench by over 28.3% in accuracy and achieves performance comparable to commercial models like GPT-4o. We also demonstrate that incorporating VEU data significantly enhances the performance of Vid-LLMs on general video understanding benchmarks, with an average improvement of 8.3% across nine reasoning tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9982208f-c4f7-481c-b79f-85f1e671ea8bCited by top-tier papers4
- RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-ThoughtYi Lu, Jiawang Cao, Yongliang Wu, Bozheng Li et al.ACL 2025 · 15 citations
- VideoNSA: Native Sparse Attention Scales Video UnderstandingEnxin Song, Wenhao Chai, Shusheng Yang, Ethan Armand et al.ICLR 2026 · 11 citations
- AutoCut: End-to-end advertisement video editing based on multimodal discretization and controllable generationMilton Zhou, Sizhong Qin, Yongzhi Li, Quan Chen et al.CVPR 2026 · 4 citations
- VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative VideosJiashuo Yu, Yue Wu, Meng Chu, Zhifei Ren et al.ICCV 2025 · 3 citations
Builds on11
- Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language ModelsMuhammad Maaz, Hanoona Abdul Rasheed, Salman Khan, Fahad KhanACL 2024 · 279 citations
- Video-LLaVA: Learning United Visual Representation by Alignment Before ProjectionBin Lin, Yang Ye, Bin Zhu, Jiaxi Cui et al.EMNLP 2024 · 231 citations
- LVBench: An Extreme Long Video Understanding BenchmarkWeihan Wang, Zehai He, Wenyi Hong, Yean Cheng et al.ICCV 2025 · 28 citations
- Learning to Cut by Watching MoviesAlejandro Pardo, Fabian Caba Heilbron, Juan León Alcázar, Ali K. Thabet et al.ICCV 2021 · 26 citations
- Frame Order Matters: A Temporal Sequence-Aware Model for Few-Shot Action RecognitionBozheng Li, Mushui Liu, Gaoang Wang, Yunlong YuAAAI 2025 · 14 citations
Related papers
- ShotBench: Expert-Level Cinematic Understanding in Vision-Language ModelsHongbo Liu, Jingwen He, Yi Jin, Dian Zheng et al.NeurIPS 2025 · 24 citations
- GLIMPSE: Do Large Vision-Language Models Truly Think With Videos or Just Glimpse at Them?Yiyang Zhou, Linjie Li, Shi Qiu, Zhengyuan Yang et al.EMNLP 2025 · 2 citations
- V2P-Bench: Evaluating Video-Language Understanding with Visual Prompts for Better Human-Model InteractionYiming Zhao, Yu Zeng, Yukun Qi, YaoYang Liu et al.ICLR 2026 · 8 citations
- VidComposition: Can MLLMs Analyze Compositions in Compiled Videos?Yunlong Tang, Junjia Guo, Hang Hua, Susan Liang et al.CVPR 2025
- MVBench: A Comprehensive Multi-modal Video Understanding BenchmarkKunchang Li, Yali Wang, Yinan He, Yizhuo Li et al.CVPR 2024
