Video Event Extraction via Tracking Visual States of Arguments
Guang Yang, Manling Li, Jiajie Zhang, Xudong Lin, Heng Ji, Shih-Fu Chang
摘要
Video event extraction aims to detect salient events from a video and identify the arguments for each event as well as their semantic roles. Existing methods focus on capturing the overall visual scene of each frame, ignoring fine-grained argument-level information. Inspired by the definition of events as changes of states, we propose a novel framework to detect video events by tracking the changes in the visual states of all involved arguments, which are expected to provide the most informative evidence for the extraction of video events. In order to capture the visual state changes of arguments, we decompose them into changes in pixels within objects, displacements of objects, and interactions among multiple arguments. We further propose Object State Embedding, Object Motion-aware Embedding and Argument Interaction Embedding to encode and track these changes respectively. Experiments on various video event extraction tasks demonstrate significant improvements compared to state-of-the-art models. In particular, on verb classification, we achieve 3.49% absolute gains (19.53% relative gains) in F1@5 on Video Situation Recognition. Our Code is publicly available at https://github.com/Shinetism/VStates for research purposes.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Constructing Holistic Spatio-Temporal Scene Graph for Video Semantic Role LabelingYu Zhao, Hao Fei, Yixin Cao, Bobo Li 等ACM MM 2023 · 被引用 31 次
- Localizing Active Objects from Egocentric Vision with Symbolic World KnowledgeTe-Lin Wu, Yu Zhou, Nanyun PengEMNLP 2023 · 被引用 5 次
- Video Event Extraction with Multi-View Interaction Knowledge DistillationKaiwen Wei, Runyan Du, Li Jin, Jian Liu 等AAAI 2024 · 被引用 5 次
它引用的顶会 Paper8
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 被引用 4,104 次
- Is Space-Time Attention All You Need for Video Understanding?Gedas Bertasius, Heng Wang, Lorenzo TorresaniICML 2021 · 被引用 2,927 次
- Event Extraction as Machine Reading ComprehensionJian Liu, Yubo Chen, Kang Liu, Wei Bi 等EMNLP 2020 · 被引用 300 次
- CLIP-Event: Connecting Text and Images with Event StructuresManling Li, Ruochen Xu, Shuohang Wang, Luowei Zhou 等CVPR 2022 · 被引用 103 次
- Cross-media Structured Common Space for Multimedia Event ExtractionManling Li, Alireza Zareian, Qi Zeng, Spencer Whitehead 等ACL 2020 · 被引用 87 次
相关 Paper
- Object-Centric Framework for Video Moment RetrievalZongyao Li, Yongkang Wong, Satoshi Yamazaki, Jianquan Liu 等AAAI 2026
- SAGE: A Unified Framework for Generalizable Object State Recognition with State-Action Graph EmbeddingYuan Zang, Zitian Tang, Junho Cho, Jaewook Yoo 等NeurIPS 2025
- Symbiotic Attention with Privileged Information for Egocentric Action RecognitionXiaohan Wang, Yu Wu, Linchao Zhu, Yi YangAAAI 2020 · 被引用 66 次
- Syntax-Aware Action Targeting for Video CaptioningQi Zheng, Chaoyue Wang, Dacheng TaoCVPR 2020
- Attention-Based Context Aware Reasoning for Situation RecognitionThilini Cooray, Ngai-Man Cheung, Wei LuCVPR 2020
