Unbiased Scene Graph Generation in Videos
Sayak Nag, Kyle Min, Subarna Tripathi, Amit K. Roy-Chowdhury
摘要
The task of dynamic scene graph generation (SGG) from videos is complicated and challenging due to the inherent dynamics of a scene, temporal fluctuation of model predictions, and the long-tailed distribution of the visual relationships in addition to the already existing challenges in image-based SGG. Existing methods for dynamic SGG have primarily focused on capturing spatio-temporal context using complex architectures without addressing the challenges mentioned above, especially the long-tailed distribution of relationships. This often leads to the generation of biased scene graphs. To address these challenges, we introduce a new framework called TEMPURA: TEmporal consistency and Memory Prototype guided UnceRtainty Attenuation for unbiased dynamic SGG. TEMPURA employs object-level temporal consistencies via transformerbased sequence modeling, learns to synthesize unbiased relationship representations using memory-guided training, and attenuates the predictive uncertainty of visual relations using a Gaussian Mixture Model (GMM). Extensive experiments demonstrate that our method achieves significant (up to 10% in some cases) performance gain over existing methods highlighting its superiority in generating more unbiased scene graphs.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper17
- Constructing Holistic Spatio-Temporal Scene Graph for Video Semantic Role LabelingYu Zhao, Hao Fei, Yixin Cao, Bobo Li 等ACM MM 2023 · 被引用 31 次
- Action Scene Graphs for Long-Form Understanding of Egocentric VideosIvan Rodin, Antonino Furnari, Kyle Min, Subarna Tripathi 等CVPR 2024 · 被引用 14 次
- CYCLO: Cyclic Graph Transformer Approach to Multi-Object Relationship Modeling in Aerial VideosTrong-Thuan Nguyen, Pha A. Nguyen, Xin Li, Jackson David Cothren 等NeurIPS 2024 · 被引用 13 次
- OED: Towards One-stage End-to-End Dynamic Scene Graph GenerationGuan Wang, Zhimin Li, Qingchao Chen, Yang LiuCVPR 2024 · 被引用 12 次
- Motion-aware Contrastive Learning for Temporal Panoptic Scene Graph GenerationThong Thanh Nguyen, Xiaobao Wu, Yi Bin, Cong-Duy T. Nguyen 等AAAI 2025 · 被引用 8 次
它引用的顶会 Paper19
- ViViT: A Video Vision TransformerAnurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun 等ICCV 2021 · 被引用 2,947 次
- VideoBERT: A Joint Model for Video and Language Representation LearningChen Sun, Austin Myers, Carl Vondrick, Kevin Murphy 等ICCV 2019 · 被引用 1,396 次
- Spatial-Temporal Transformer for Dynamic Scene Graph GenerationYuren Cong, Wentong Liao, Hanno Ackermann, Bodo Rosenhahn 等ICCV 2021 · 被引用 163 次
- Estimating and Exploiting the Aleatoric Uncertainty in Surface Normal EstimationGwangbin Bae, Ignas Budvytis, Roberto CipollaICCV 2021 · 被引用 154 次
- Active Learning for Deep Object Detection via Probabilistic ModelingJiwoong Choi, Ismail Elezi, Hyuk-Jae Lee, Clément Farabet 等ICCV 2021 · 被引用 144 次
相关 Paper
- Unbiased Video Scene Graph Generation via Visual and Semantic Dual DebiasingYanjun Li, Zhaoyang Li, Honghui Chen, Lizhi XuCVPR 2025
- TD²-Net: Toward Denoising and Debiasing for Video Scene Graph GenerationXin Lin, Chong Shi, Yibing Zhan, Zuopeng Yang 等AAAI 2024 · 被引用 8 次
- Towards Unbiased and Robust Spatio-Temporal Scene Graph Generation and AnticipationRohith Peddi, Saurabh, Ayush Abhay Shrivastava, Parag Singla 等CVPR 2025
- Iterative Learning with Extra and Inner Knowledge for Long-tail Dynamic Scene Graph GenerationYiming Li, Xiaoshan Yang, Changsheng XuACM MM 2023 · 被引用 1 次
- Resistance Training Using Prior Bias: Toward Unbiased Scene Graph GenerationChao Chen, Yibing Zhan, Baosheng Yu, Liu Liu 等AAAI 2022 · 被引用 52 次
