Unified Graph Structured Models for Video Understanding
Anurag Arnab, Chen Sun, Cordelia Schmid
Abstract
Accurate video understanding involves reasoning about the relationships between actors, objects and their environment, often over long temporal intervals. In this paper, we propose a message passing graph neural network that explicitly models these spatio-temporal relations and can use explicit representations of objects, when supervision is available, and implicit representations otherwise. Our formulation generalises previous structured models for video understanding, and allows us to study how different design choices in graph structure and representation affect the model’s performance. We demonstrate our method on two different tasks requiring relational reasoning in videos – spatio-temporal action detection on AVA and UCF101-24, and video scene graph classification on the recent Action Genome dataset – and achieve state-of-the-art results on all three datasets. Furthermore, we show quantitatively and qualitatively how our method is able to more effectively model relationships between relevant entities in the scene.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 025ede07-ffec-42a0-9fc3-96b6d6359e6dCited by top-tier papers13
- ViViT: A Video Vision TransformerAnurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun et al.ICCV 2021 · 2,947 citations
- Multiview Transformers for Video RecognitionShen Yan, Xuehan Xiong, Anurag Arnab, Zhichao Lu et al.CVPR 2022 · 279 citations
- Dense and Aligned Captions (DAC) Promote Compositional Reasoning in VL ModelsSivan Doveh, Assaf Arbelle, Sivan Harary, Roei Herzig et al.NeurIPS 2023 · 93 citations
- Object-Region Video TransformersRoei Herzig, Elad Ben-Avraham, Karttikeya Mangalam, Amir Bar et al.CVPR 2022 · 75 citations
- Helping Hands: An Object-Aware Ego-Centric Video Recognition ModelChuhan Zhang, Ankush Gupta, Andrew ZissermanICCV 2023 · 39 citations
Builds on5
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 4,104 citations
- Action Genome: Actions As Compositions of Spatio-Temporal Scene GraphsJingwei Ji, Ranjay Krishna, Li Fei-Fei, Juan Carlos NieblesCVPR 2020
- Something-Else: Compositional Action Recognition With Spatial-Temporal Interaction NetworksJoanna Materzynska, Tete Xiao, Roei Herzig, Huijuan Xu et al.CVPR 2020
- Dynamic Graph Message Passing NetworksLi Zhang, Dan Xu, Anurag Arnab, Philip H. S. TorrCVPR 2020
- In Defense of Grid Features for Visual Question AnsweringHuaizu Jiang, Ishan Misra, Marcus Rohrbach, Erik G. Learned-Miller et al.CVPR 2020
Related papers
- Spatial-Temporal Transformer for Dynamic Scene Graph GenerationYuren Cong, Wentong Liao, Hanno Ackermann, Bodo Rosenhahn et al.ICCV 2021 · 163 citations
- Detecting Human-Object Relationships in VideosJingwei Ji, Rishi Desai, Juan Carlos NieblesICCV 2021 · 47 citations
- End-to-End Entity-Predicate Association Reasoning for Dynamic Scene Graph GenerationLiwei Wang, Yanduo Zhang, Tao Lu, Fang Liu et al.ICCV 2025 · 1 citation
- Object-Relation Reasoning Graph for Action RecognitionYangjun Ou, Li Mi, Zhenzhong ChenCVPR 2022 · 25 citations
- Dynamic Scene Graph Generation via Anticipatory Pre-trainingYiming Li, Xiaoshan Yang, Changsheng XuCVPR 2022 · 38 citations
