Action Genome: Actions As Compositions of Spatio-Temporal Scene Graphs
Jingwei Ji, Ranjay Krishna, Li Fei-Fei, Juan Carlos Niebles
摘要
Action recognition has typically treated actions and activities as monolithic events that occur in videos. However, there is evidence from Cognitive Science and Neuroscience that people actively encode activities into consistent hierarchical part structures. However, in Computer Vision, few explorations on representations that encode event partonomies have been made. Inspired by evidence that the prototypical unit of an event is an action-object interaction, we introduce Action Genome, a representation that decomposes actions into spatio-temporal scene graphs. Action Genome captures changes between objects and their pairwise relationships while an action occurs. It contains 10K videos with 0.4M objects and 1.7M visual relationships annotated. With Action Genome, we extend an existing action recognition model by incorporating scene graphs as spatiotemporal feature banks to achieve better performance on the Charades dataset. Next, by decomposing and learning the temporal changes in visual relationships that result in an action, we demonstrate the utility of a hierarchical event decomposition by enabling few-shot action recognition, achieving 42.7% mAP using as few as 10 examples. Finally, we benchmark existing scene graph models on the new task of spatio-temporal scene graph prediction.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper111
- Video-of-Thought: Step-by-Step Video Reasoning from Perception to CognitionHao Fei, Shengqiong Wu, Wei Ji, Hanwang Zhang 等ICML 2024 · 被引用 182 次
- Spatial-Temporal Transformer for Dynamic Scene Graph GenerationYuren Cong, Wentong Liao, Hanno Ackermann, Bodo Rosenhahn 等ICCV 2021 · 被引用 163 次
- RLIP: Relational Language-Image Pre-training for Human-Object Interaction DetectionHangjie Yuan, Jianwen Jiang, Samuel Albanie, Tao Feng 等NeurIPS 2022 · 被引用 88 次
- Learning to Generate Scene Graph from Natural Language SupervisionYiwu Zhong, Jing Shi, Jianwei Yang, Chenliang Xu 等ICCV 2021 · 被引用 88 次
- Target Adaptive Context Aggregation for Video Scene Graph GenerationYao Teng, Limin Wang, Zhifeng Li, Gangshan WuICCV 2021 · 被引用 80 次
它引用的顶会 Paper5
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 被引用 4,104 次
- TSM: Temporal Shift Module for Efficient Video UnderstandingJi Lin, Chuang Gan, Song HanICCV 2019 · 被引用 2,049 次
- HACS: Human Action Clips and Segments Dataset for Recognition and Temporal LocalizationHang Zhao, Antonio Torralba, Lorenzo Torresani, Zhicheng YanICCV 2019 · 被引用 298 次
- Specifying Object Attributes and Relations in Interactive Scene GenerationOron Ashual, Lior WolfICCV 2019 · 被引用 190 次
- Scene Graph Prediction With Limited LabelsRanjay Krishna, Vincent S. Chen, Paroma Varma, Michael S. Bernstein 等ICCV 2019 · 被引用 5 次
相关 Paper
- Home Action Genome: Cooperative Compositional Action UnderstandingNishant Rai, Haofeng Chen, Jingwei Ji, Rishi Desai 等CVPR 2021
- Detecting Human-Object Relationships in VideosJingwei Ji, Rishi Desai, Juan Carlos NieblesICCV 2021 · 被引用 47 次
- Prompt-guided Disentangled Representation for Action RecognitionTianci Wu, Guangming Zhu, Jiang Lu, Siyuan Wang 等NeurIPS 2025 · 被引用 1 次
- Dynamic Scene Graph Generation via Anticipatory Pre-trainingYiming Li, Xiaoshan Yang, Changsheng XuCVPR 2022 · 被引用 38 次
- Visual Knowledge Graph for Human Action Reasoning in VideosYue Ma, Yali Wang, Yue Wu, Ziyu Lyu 等ACM MM 2022 · 被引用 29 次
