Unbiased Scene Graph Generation in Videos
Sayak Nag, Kyle Min, Subarna Tripathi, Amit K. Roy-Chowdhury
Abstract
The task of dynamic scene graph generation (SGG) from videos is complicated and challenging due to the inherent dynamics of a scene, temporal fluctuation of model predictions, and the long-tailed distribution of the visual relationships in addition to the already existing challenges in image-based SGG. Existing methods for dynamic SGG have primarily focused on capturing spatio-temporal context using complex architectures without addressing the challenges mentioned above, especially the long-tailed distribution of relationships. This often leads to the generation of biased scene graphs. To address these challenges, we introduce a new framework called TEMPURA: TEmporal consistency and Memory Prototype guided UnceRtainty Attenuation for unbiased dynamic SGG. TEMPURA employs object-level temporal consistencies via transformerbased sequence modeling, learns to synthesize unbiased relationship representations using memory-guided training, and attenuates the predictive uncertainty of visual relations using a Gaussian Mixture Model (GMM). Extensive experiments demonstrate that our method achieves significant (up to 10% in some cases) performance gain over existing methods highlighting its superiority in generating more unbiased scene graphs.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b1eb5fe2-e9c1-4672-bfa4-ad20079fd409Cited by top-tier papers17
- Constructing Holistic Spatio-Temporal Scene Graph for Video Semantic Role LabelingYu Zhao, Hao Fei, Yixin Cao, Bobo Li et al.ACM MM 2023 · 31 citations
- Action Scene Graphs for Long-Form Understanding of Egocentric VideosIvan Rodin, Antonino Furnari, Kyle Min, Subarna Tripathi et al.CVPR 2024 · 14 citations
- CYCLO: Cyclic Graph Transformer Approach to Multi-Object Relationship Modeling in Aerial VideosTrong-Thuan Nguyen, Pha A. Nguyen, Xin Li, Jackson David Cothren et al.NeurIPS 2024 · 13 citations
- OED: Towards One-stage End-to-End Dynamic Scene Graph GenerationGuan Wang, Zhimin Li, Qingchao Chen, Yang LiuCVPR 2024 · 12 citations
- Motion-aware Contrastive Learning for Temporal Panoptic Scene Graph GenerationThong Thanh Nguyen, Xiaobao Wu, Yi Bin, Cong-Duy T. Nguyen et al.AAAI 2025 · 8 citations
Builds on19
- ViViT: A Video Vision TransformerAnurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun et al.ICCV 2021 · 2,947 citations
- VideoBERT: A Joint Model for Video and Language Representation LearningChen Sun, Austin Myers, Carl Vondrick, Kevin Murphy et al.ICCV 2019 · 1,396 citations
- Spatial-Temporal Transformer for Dynamic Scene Graph GenerationYuren Cong, Wentong Liao, Hanno Ackermann, Bodo Rosenhahn et al.ICCV 2021 · 163 citations
- Estimating and Exploiting the Aleatoric Uncertainty in Surface Normal EstimationGwangbin Bae, Ignas Budvytis, Roberto CipollaICCV 2021 · 154 citations
- Active Learning for Deep Object Detection via Probabilistic ModelingJiwoong Choi, Ismail Elezi, Hyuk-Jae Lee, Clément Farabet et al.ICCV 2021 · 144 citations
Related papers
- Unbiased Video Scene Graph Generation via Visual and Semantic Dual DebiasingYanjun Li, Zhaoyang Li, Honghui Chen, Lizhi XuCVPR 2025
- TD²-Net: Toward Denoising and Debiasing for Video Scene Graph GenerationXin Lin, Chong Shi, Yibing Zhan, Zuopeng Yang et al.AAAI 2024 · 8 citations
- Towards Unbiased and Robust Spatio-Temporal Scene Graph Generation and AnticipationRohith Peddi, Saurabh, Ayush Abhay Shrivastava, Parag Singla et al.CVPR 2025
- Iterative Learning with Extra and Inner Knowledge for Long-tail Dynamic Scene Graph GenerationYiming Li, Xiaoshan Yang, Changsheng XuACM MM 2023 · 1 citation
- Resistance Training Using Prior Bias: Toward Unbiased Scene Graph GenerationChao Chen, Yibing Zhan, Baosheng Yu, Liu Liu et al.AAAI 2022 · 52 citations
