Video Event Extraction via Tracking Visual States of Arguments
Guang Yang, Manling Li, Jiajie Zhang, Xudong Lin, Heng Ji, Shih-Fu Chang
Abstract
Video event extraction aims to detect salient events from a video and identify the arguments for each event as well as their semantic roles. Existing methods focus on capturing the overall visual scene of each frame, ignoring fine-grained argument-level information. Inspired by the definition of events as changes of states, we propose a novel framework to detect video events by tracking the changes in the visual states of all involved arguments, which are expected to provide the most informative evidence for the extraction of video events. In order to capture the visual state changes of arguments, we decompose them into changes in pixels within objects, displacements of objects, and interactions among multiple arguments. We further propose Object State Embedding, Object Motion-aware Embedding and Argument Interaction Embedding to encode and track these changes respectively. Experiments on various video event extraction tasks demonstrate significant improvements compared to state-of-the-art models. In particular, on verb classification, we achieve 3.49% absolute gains (19.53% relative gains) in F1@5 on Video Situation Recognition. Our Code is publicly available at https://github.com/Shinetism/VStates for research purposes.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext aab65ff1-a0cd-4337-9b11-7b20a0a46a90Cited by top-tier papers3
- Constructing Holistic Spatio-Temporal Scene Graph for Video Semantic Role LabelingYu Zhao, Hao Fei, Yixin Cao, Bobo Li et al.ACM MM 2023 · 31 citations
- Localizing Active Objects from Egocentric Vision with Symbolic World KnowledgeTe-Lin Wu, Yu Zhou, Nanyun PengEMNLP 2023 · 5 citations
- Video Event Extraction with Multi-View Interaction Knowledge DistillationKaiwen Wei, Runyan Du, Li Jin, Jian Liu et al.AAAI 2024 · 5 citations
Builds on8
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 4,104 citations
- Is Space-Time Attention All You Need for Video Understanding?Gedas Bertasius, Heng Wang, Lorenzo TorresaniICML 2021 · 2,927 citations
- Event Extraction as Machine Reading ComprehensionJian Liu, Yubo Chen, Kang Liu, Wei Bi et al.EMNLP 2020 · 300 citations
- CLIP-Event: Connecting Text and Images with Event StructuresManling Li, Ruochen Xu, Shuohang Wang, Luowei Zhou et al.CVPR 2022 · 103 citations
- Cross-media Structured Common Space for Multimedia Event ExtractionManling Li, Alireza Zareian, Qi Zeng, Spencer Whitehead et al.ACL 2020 · 87 citations
Related papers
- Object-Centric Framework for Video Moment RetrievalZongyao Li, Yongkang Wong, Satoshi Yamazaki, Jianquan Liu et al.AAAI 2026
- SAGE: A Unified Framework for Generalizable Object State Recognition with State-Action Graph EmbeddingYuan Zang, Zitian Tang, Junho Cho, Jaewook Yoo et al.NeurIPS 2025
- Symbiotic Attention with Privileged Information for Egocentric Action RecognitionXiaohan Wang, Yu Wu, Linchao Zhu, Yi YangAAAI 2020 · 66 citations
- Syntax-Aware Action Targeting for Video CaptioningQi Zheng, Chaoyue Wang, Dacheng TaoCVPR 2020
- Attention-Based Context Aware Reasoning for Situation RecognitionThilini Cooray, Ngai-Man Cheung, Wei LuCVPR 2020
