Action Genome: Actions As Compositions of Spatio-Temporal Scene Graphs
Jingwei Ji, Ranjay Krishna, Li Fei-Fei, Juan Carlos Niebles
Abstract
Action recognition has typically treated actions and activities as monolithic events that occur in videos. However, there is evidence from Cognitive Science and Neuroscience that people actively encode activities into consistent hierarchical part structures. However, in Computer Vision, few explorations on representations that encode event partonomies have been made. Inspired by evidence that the prototypical unit of an event is an action-object interaction, we introduce Action Genome, a representation that decomposes actions into spatio-temporal scene graphs. Action Genome captures changes between objects and their pairwise relationships while an action occurs. It contains 10K videos with 0.4M objects and 1.7M visual relationships annotated. With Action Genome, we extend an existing action recognition model by incorporating scene graphs as spatiotemporal feature banks to achieve better performance on the Charades dataset. Next, by decomposing and learning the temporal changes in visual relationships that result in an action, we demonstrate the utility of a hierarchical event decomposition by enabling few-shot action recognition, achieving 42.7% mAP using as few as 10 examples. Finally, we benchmark existing scene graph models on the new task of spatio-temporal scene graph prediction.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 03d662db-985e-4eb4-9ff1-da8945263933Cited by top-tier papers111
- Video-of-Thought: Step-by-Step Video Reasoning from Perception to CognitionHao Fei, Shengqiong Wu, Wei Ji, Hanwang Zhang et al.ICML 2024 · 182 citations
- Spatial-Temporal Transformer for Dynamic Scene Graph GenerationYuren Cong, Wentong Liao, Hanno Ackermann, Bodo Rosenhahn et al.ICCV 2021 · 163 citations
- RLIP: Relational Language-Image Pre-training for Human-Object Interaction DetectionHangjie Yuan, Jianwen Jiang, Samuel Albanie, Tao Feng et al.NeurIPS 2022 · 88 citations
- Learning to Generate Scene Graph from Natural Language SupervisionYiwu Zhong, Jing Shi, Jianwei Yang, Chenliang Xu et al.ICCV 2021 · 88 citations
- Target Adaptive Context Aggregation for Video Scene Graph GenerationYao Teng, Limin Wang, Zhifeng Li, Gangshan WuICCV 2021 · 80 citations
Builds on5
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 4,104 citations
- TSM: Temporal Shift Module for Efficient Video UnderstandingJi Lin, Chuang Gan, Song HanICCV 2019 · 2,049 citations
- HACS: Human Action Clips and Segments Dataset for Recognition and Temporal LocalizationHang Zhao, Antonio Torralba, Lorenzo Torresani, Zhicheng YanICCV 2019 · 298 citations
- Specifying Object Attributes and Relations in Interactive Scene GenerationOron Ashual, Lior WolfICCV 2019 · 190 citations
- Scene Graph Prediction With Limited LabelsRanjay Krishna, Vincent S. Chen, Paroma Varma, Michael S. Bernstein et al.ICCV 2019 · 5 citations
Related papers
- Home Action Genome: Cooperative Compositional Action UnderstandingNishant Rai, Haofeng Chen, Jingwei Ji, Rishi Desai et al.CVPR 2021
- Detecting Human-Object Relationships in VideosJingwei Ji, Rishi Desai, Juan Carlos NieblesICCV 2021 · 47 citations
- Prompt-guided Disentangled Representation for Action RecognitionTianci Wu, Guangming Zhu, Jiang Lu, Siyuan Wang et al.NeurIPS 2025 · 1 citation
- Dynamic Scene Graph Generation via Anticipatory Pre-trainingYiming Li, Xiaoshan Yang, Changsheng XuCVPR 2022 · 38 citations
- Visual Knowledge Graph for Human Action Reasoning in VideosYue Ma, Yali Wang, Yue Wu, Ziyu Lyu et al.ACM MM 2022 · 29 citations
