VidEvent: A Large Dataset for Understanding Dynamic Evolution of Events in Videos
Baoyu Liang, Qile Su, Shoutai Zhu, Yuchen Liang, Chao Tong
Abstract
Despite the significant impact of visual events on human cognition, understanding events in videos remains a challenging task for AI due to their complex structures, semantic hierarchies, and dynamic evolution. To address this, we propose the task of video event understanding that extracts event scripts and makes predictions with these scripts from videos. To support this task, we introduce VidEvent, a large-scale dataset containing over 23,000 well-labeled events, featuring detailed event structures, broad hierarchies, and logical relations extracted from movie recap videos. The dataset was created through a meticulous annotation process, ensuring highquality and reliable event data. We also provide comprehensive baseline models offering detailed descriptions of their architecture and performance metrics. These models serve as benchmarks for future research, facilitating comparisons and improvements. Our analysis of VidEvent and the baseline models highlights the dataset's potential to advance video event understanding and encourages the exploration of innovative algorithms and models. The dataset and related resources are publicly available at www.videvent.top.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4fd856f5-1498-4f82-84b1-910376fe0ab8Cited by top-tier papers2
- Video-CoE: Reinforcing Video Event Prediction via Chain of EventsQile Su, Jing Tang, Rui Chen, Lei Sun et al.CVPR 2026 · 2 citations
- EventFormer: A Node-graph Hierarchical Attention Transformer for Action-centric Video Event PredictionQile Su, Shoutai Zhu, Shuai Zhang, Baoyu Liang et al.ACM MM 2025
Builds on14
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 4,104 citations
- Is Space-Time Attention All You Need for Video Understanding?Gedas Bertasius, Heng Wang, Lorenzo TorresaniICML 2021 · 2,927 citations
- HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video ClipsAntoine Miech, Dimitri Zhukov, Jean-Baptiste Alayrac, Makarand Tapaswi et al.ICCV 2019 · 1,437 citations
- Revisiting Skeleton-based Action RecognitionHaodong Duan, Yue Zhao, Kai Chen, Dahua Lin et al.CVPR 2022 · 752 citations
Related papers
- MovieRecapsQA: A Multimodal Open-Ended Video Question-Answering BenchmarkShaden Shaar, Bradon Thymes, Sirawut Chaixanien, Claire Cardie et al.CVPR 2026 · 4 citations
- Video ReCap: Recursive Captioning of Hour-Long VideosMd Mohaiminul Islam, Ngan Ho, Xitong Yang, Tushar Nagarajan et al.CVPR 2024
- A Very Big Video Reasoning SuiteMaijunxian Wang, Ruisi Wang, Juyi Lin, Ran Ji et al.ICML 2026 · 20 citations
- A Local-to-Global Approach to Multi-Modal Movie Scene SegmentationAnyi Rao, Linning Xu, Yu Xiong, Guodong Xu et al.CVPR 2020
- A Graph-Based Framework to Bridge Movies and SynopsesYu Xiong, Qingqiu Huang, Lingfeng Guo, Hang Zhou et al.ICCV 2019 · 71 citations
