Towards Unbiased and Robust Spatio-Temporal Scene Graph Generation and Anticipation
Rohith Peddi, Saurabh, Ayush Abhay Shrivastava, Parag Singla, Vibhav Gogate
Abstract
Spatio-Temporal Scene Graphs (STSGs) provide a concise and expressive representation of dynamic scenes by modeling objects and their evolving relationships over time. However, real-world visual relationships often exhibit a longtailed distribution, causing existing methods for tasks like Video Scene Graph Generation (VidSGG) and Scene Graph Anticipation (SGA) to produce biased scene graphs. To this end, we propose IMPARTAIL, a novel training framework that leverages loss masking and curriculum learning to mitigate bias in the generation and anticipation of spatiotemporal scene graphs. Unlike prior methods that add extra architectural components to learn unbiased estimators, we propose an impartial training objective that reduces the dominance of head classes during learning and focuses on underrepresented tail relationships. Our curriculum-driven mask generation strategy further empowers the model to adaptively adjust its bias mitigation strategy over time, enabling more balanced and robust estimations. To thoroughly assess performance under various distribution shifts, we also introduce two new tasks-Robust Spatio-Temporal Scene Graph Generation and Robust Scene Graph Anticipation-offering a challenging benchmark for evaluating the resilience of STSG models. Extensive experiments on the Action Genome dataset demonstrate the superior unbiased performance and robustness of our method compared to existing baselines.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 33a5fceb-68db-4e1c-9aad-b335cb1a1ef1Cited by top-tier papers1
Ask how each one uses itBuilds on9
- Decoupling Representation and Classifier for Long-Tailed RecognitionBingyi Kang, Saining Xie, Marcus Rohrbach, Zhicheng Yan et al.ICLR 2020 · 1,496 citations
- Asymmetric Loss For Multi-Label ClassificationTal Ridnik, Emanuel Ben Baruch, Nadav Zamir, Asaf Noy et al.ICCV 2021 · 778 citations
- Long-tailed Recognition by Routing Diverse Distribution-Aware ExpertsXudong Wang, Long Lian, Zhongqi Miao, Ziwei Liu et al.ICLR 2021 · 481 citations
- Assembly101: A Large-Scale Multi-View Video Dataset for Understanding Procedural ActivitiesFadime Sener, Dibyadip Chatterjee, Daniel Shelepov, Kun He et al.CVPR 2022 · 168 citations
- Spatial-Temporal Transformer for Dynamic Scene Graph GenerationYuren Cong, Wentong Liao, Hanno Ackermann, Bodo Rosenhahn et al.ICCV 2021 · 163 citations
Related papers
- PPDL: Predicate Probability Distribution based Loss for Unbiased Scene Graph GenerationWei Li, Haiwei Zhang, Qijie Bai, Guoqing Zhao et al.CVPR 2022 · 64 citations
- Dynamic Scene Graph Generation via Anticipatory Pre-trainingYiming Li, Xiaoshan Yang, Changsheng XuCVPR 2022 · 38 citations
- Unbiased Scene Graph Generation in VideosSayak Nag, Kyle Min, Subarna Tripathi, Amit K. Roy-ChowdhuryCVPR 2023
- Environment-Invariant Curriculum Relation Learning for Fine-Grained Scene Graph GenerationYukuan Min, Aming Wu, Cheng DengICCV 2023 · 16 citations
- Prior Knowledge-driven Dynamic Scene Graph Generation with Causal InferenceJiale Lu, Lianggangxu Chen, Youqi Song, Shaohui Lin et al.ACM MM 2023 · 7 citations
