Spatial-Temporal Transformer for Dynamic Scene Graph Generation
Yuren Cong, Wentong Liao, Hanno Ackermann, Bodo Rosenhahn, Michael Ying Yang
Abstract
Dynamic scene graph generation aims at generating a scene graph of the given video. Compared to the task of scene graph generation from images, it is more challenging because of the dynamic relationships between objects and the temporal dependencies between frames allowing for a richer semantic interpretation. In this paper, we propose Spatial-temporal Transformer (STTran), a neural network that consists of two core modules: (1) a spatial encoder that takes an input frame to extract spatial context and reason about the visual relationships within a frame, and (2) a temporal decoder which takes the output of the spatial encoder as input in order to capture the temporal dependencies between frames and infer the dynamic relationships. Furthermore, STTran is flexible to take varying lengths of videos as input without clipping, which is especially important for long videos. Our method is validated on the benchmark dataset Action Genome (AG). The experimental results demonstrate the superior performance of our method in terms of dynamic scene graphs. Moreover, a set of ablative studies is conducted and the effect of each proposed module is justified. Code available at: https://github.com/yrcong/STTran .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ad3a6cf0-9f46-46bd-b2b4-c009fd7aba40Cited by top-tier papers47
- Video-of-Thought: Step-by-Step Video Reasoning from Perception to CognitionHao Fei, Shengqiong Wu, Wei Ji, Hanwang Zhang et al.ICML 2024 · 182 citations
- Delving into Sequential Patches for Deepfake DetectionJiazhi Guan, Hang Zhou, Zhibin Hong, Errui Ding et al.NeurIPS 2022 · 84 citations
- MASTER: Market-Guided Stock Transformer for Stock Price ForecastingTong Li, Zhaoyang Liu, Yanyan Shen, Xue Wang et al.AAAI 2024 · 80 citations
- RNTrajRec: Road Network Enhanced Trajectory Recovery with Spatial-Temporal TransformerYuqi Chen, Hanyuan Zhang, Weiwei Sun, Baihua ZhengICDE 2023 · 70 citations
- (2.5+1)D Spatio-Temporal Scene Graphs for Video Question AnsweringAnoop Cherian, Chiori Hori, Tim K. Marks, Jonathan Le RouxAAAI 2022 · 48 citations
Builds on11
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
- TSM: Temporal Shift Module for Efficient Video UnderstandingJi Lin, Chuang Gan, Song HanICCV 2019 · 2,049 citations
- VL-BERT: Pre-training of Generic Visual-Linguistic RepresentationsWeijie Su, Xizhou Zhu, Yue Cao, Bin Li et al.ICLR 2020 · 1,825 citations
- Specifying Object Attributes and Relations in Interactive Scene GenerationOron Ashual, Lior WolfICCV 2019 · 190 citations
Related papers
- Dynamic Scene Graph Generation via Anticipatory Pre-trainingYiming Li, Xiaoshan Yang, Changsheng XuCVPR 2022 · 38 citations
- OED: Towards One-stage End-to-End Dynamic Scene Graph GenerationGuan Wang, Zhimin Li, Qingchao Chen, Yang LiuCVPR 2024 · 12 citations
- End-to-End Entity-Predicate Association Reasoning for Dynamic Scene Graph GenerationLiwei Wang, Yanduo Zhang, Tao Lu, Fang Liu et al.ICCV 2025 · 1 citation
- Unified Graph Structured Models for Video UnderstandingAnurag Arnab, Chen Sun, Cordelia SchmidICCV 2021 · 57 citations
- Detecting Human-Object Relationships in VideosJingwei Ji, Rishi Desai, Juan Carlos NieblesICCV 2021 · 47 citations
