EventFormer: A Node-graph Hierarchical Attention Transformer for Action-centric Video Event Prediction
Qile Su, Shoutai Zhu, Shuai Zhang, Baoyu Liang, Chao Tong
Abstract
Script event induction, which aims to predict the subsequent event based on the context, is a challenging task in NLP, achieving remarkable success in practical applications. However, human events are mostly recorded and presented in the form of videos rather than scripts, yet there is a lack of related research in the realm of vision. To address this problem, we introduce AVEP (Action-centric Video Event Prediction), a task that distinguishes itself from existing video prediction tasks through its incorporation of more complex logic and richer semantic information. We present a large structured dataset, which consists of about 35K annotated videos and more than 178K video clips of event, built upon existing video event datasets to support this task. The dataset offers more fine-grained annotations, where the atomic unit is represented as a multimodal event argument node, providing better structured representations of video events. Due to the complexity of event structures, traditional visual models that take patches or frames as input are not well-suited for AVEP. We propose EventFormer, a node-graph hierarchical attention based video event prediction model, which can capture both the relationships between events and their arguments and the coreferencial relationships between arguments. We conducted experiments using several SOTA video prediction models as well as LVLMs on AVEP, demonstrating both the complexity of the task and the value of the dataset. Our approach outperforms all these video prediction models. We will release the dataset and code for replicating the experiments and annotations.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0fcbbef6-3e6a-48b8-abe5-c88f2c45b4fcCited by top-tier papers2
- Video-as-Answer: Predict and Generate Next Video Event with Joint-GRPOJunhao Cheng, Liang Hou, Xin Tao, Jing LiaoCVPR 2026 · 6 citations
- Video-CoE: Reinforcing Video Event Prediction via Chain of EventsQile Su, Jing Tang, Rui Chen, Lei Sun et al.CVPR 2026 · 2 citations
Builds on15
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- VaTeX: A Large-Scale, High-Quality Multilingual Dataset for Video-and-Language ResearchXin Wang, Jiawei Wu, Jun-Kun Chen, Lei Li et al.ICCV 2019 · 688 citations
- Ego4D: Around the World in 3, 000 Hours of Egocentric VideoKristen Grauman, Andrew Westbury, Eugene Byrne, Zachary Chavis et al.CVPR 2022 · 525 citations
- HACS: Human Action Clips and Segments Dataset for Recognition and Temporal LocalizationHang Zhao, Antonio Torralba, Lorenzo Torresani, Zhicheng YanICCV 2019 · 298 citations
- Anticipative Video TransformerRohit Girdhar, Kristen GraumanICCV 2021 · 270 citations
Related papers
- Rich Event Modeling for Script Event PredictionLong Bai, Saiping Guan, Zixuan Li, Jiafeng Guo et al.AAAI 2023 · 11 citations
- StepFormer: Self-Supervised Step Discovery and Localization in Instructional VideosNikita Dvornik, Isma Hadji, Ran Zhang, Konstantinos G. Derpanis et al.CVPR 2023
- VidEvent: A Large Dataset for Understanding Dynamic Evolution of Events in VideosBaoyu Liang, Qile Su, Shoutai Zhu, Yuchen Liang et al.AAAI 2025 · 5 citations
- Towards Long-Form Video UnderstandingChao-Yuan Wu, Philipp KrähenbühlCVPR 2021
- Fostering Video Reasoning via Next-Event PredictionHaonan Wang, Hongfu Liu, Xiangyan Liu, Chao Du et al.ICLR 2026 · 14 citations
