VidEvent: A Large Dataset for Understanding Dynamic Evolution of Events in Videos
Baoyu Liang, Qile Su, Shoutai Zhu, Yuchen Liang, Chao Tong
摘要
Despite the significant impact of visual events on human cognition, understanding events in videos remains a challenging task for AI due to their complex structures, semantic hierarchies, and dynamic evolution. To address this, we propose the task of video event understanding that extracts event scripts and makes predictions with these scripts from videos. To support this task, we introduce VidEvent, a large-scale dataset containing over 23,000 well-labeled events, featuring detailed event structures, broad hierarchies, and logical relations extracted from movie recap videos. The dataset was created through a meticulous annotation process, ensuring highquality and reliable event data. We also provide comprehensive baseline models offering detailed descriptions of their architecture and performance metrics. These models serve as benchmarks for future research, facilitating comparisons and improvements. Our analysis of VidEvent and the baseline models highlights the dataset's potential to advance video event understanding and encourages the exploration of innovative algorithms and models. The dataset and related resources are publicly available at www.videvent.top.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Video-CoE: Reinforcing Video Event Prediction via Chain of EventsQile Su, Jing Tang, Rui Chen, Lei Sun 等CVPR 2026 · 被引用 2 次
- EventFormer: A Node-graph Hierarchical Attention Transformer for Action-centric Video Event PredictionQile Su, Shoutai Zhu, Shuai Zhang, Baoyu Liang 等ACM MM 2025
它引用的顶会 Paper14
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 被引用 4,104 次
- Is Space-Time Attention All You Need for Video Understanding?Gedas Bertasius, Heng Wang, Lorenzo TorresaniICML 2021 · 被引用 2,927 次
- HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video ClipsAntoine Miech, Dimitri Zhukov, Jean-Baptiste Alayrac, Makarand Tapaswi 等ICCV 2019 · 被引用 1,437 次
- Revisiting Skeleton-based Action RecognitionHaodong Duan, Yue Zhao, Kai Chen, Dahua Lin 等CVPR 2022 · 被引用 752 次
相关 Paper
- MovieRecapsQA: A Multimodal Open-Ended Video Question-Answering BenchmarkShaden Shaar, Bradon Thymes, Sirawut Chaixanien, Claire Cardie 等CVPR 2026 · 被引用 4 次
- Video ReCap: Recursive Captioning of Hour-Long VideosMd Mohaiminul Islam, Ngan Ho, Xitong Yang, Tushar Nagarajan 等CVPR 2024
- A Very Big Video Reasoning SuiteMaijunxian Wang, Ruisi Wang, Juyi Lin, Ran Ji 等ICML 2026 · 被引用 20 次
- A Local-to-Global Approach to Multi-Modal Movie Scene SegmentationAnyi Rao, Linning Xu, Yu Xiong, Guodong Xu 等CVPR 2020
- A Graph-Based Framework to Bridge Movies and SynopsesYu Xiong, Qingqiu Huang, Lingfeng Guo, Hang Zhou 等ICCV 2019 · 被引用 71 次
