Classification-Then-Grounding: Reformulating Video Scene Graphs as Temporal Bipartite Graphs
Kaifeng Gao, Long Chen, Yulei Niu, Jian Shao, Jun Xiao
Abstract
Today's VidSGG models are all proposal-based methods, i.e., they first generate numerous paired subject-object snippets as proposals, and then conduct predicate classification for each proposal. In this paper, we argue that this prevalent proposal-based framework has three inherent drawbacks: 1) The ground-truth predicate labels for proposals are partially correct. 2) They break the high-order relations among different predicate instances of a same subject-object pair. 3) VidSGG performance is upper-bounded by the quality of the proposals. To this end, we propose a new classification-then-grounding framework for VidSGG, which can avoid all the three overlooked drawbacks. Meanwhile, under this framework, we reformulate the video scene graphs as temporal bipartite graphs, where the entities and predicates are two types of nodes with time slots, and the edges denote different semantic roles between these nodes. This formulation takes full advantage of our new framework. Accordingly, we further propose a novel BIpartite Graph based SGG model: BIG. It consists of a classification stage and a grounding stage, where the former aims to classify the categories of all the nodes and the edges, and the latter tries to localize the temporal location of each relation instance. Extensive ablations on two VidSGG datasets have attested to the effectiveness of our framework and BIG. Code is available at https://github.com/Dawn-LX/VidSGG-BIG.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 43084bb3-82ed-48ea-a209-923efcf3b497Cited by top-tier papers10
- The Devil is in the Labels: Noisy Label Correction for Robust Scene Graph GenerationLin Li, Long Chen, Yifeng Huang, Zhimeng Zhang et al.CVPR 2022 · 103 citations
- Integrating Object-aware and Interaction-aware Knowledge for Weakly Supervised Scene Graph GenerationXingchen Li, Long Chen, Wenbo Ma, Yi Yang et al.ACM MM 2022 · 22 citations
- Compositional Prompt Tuning with Motion Cues for Open-vocabulary Video Relation DetectionKaifeng Gao, Long Chen, Hanwang Zhang, Jun Xiao et al.ICLR 2023 · 9 citations
- TD²-Net: Toward Denoising and Debiasing for Video Scene Graph GenerationXin Lin, Chong Shi, Yibing Zhan, Zuopeng Yang et al.AAAI 2024 · 8 citations
- Rethinking Multi-Modal Alignment in Multi-Choice VideoQA from Feature and Sample PerspectivesShaoning Xiao, Long Chen, Kaifeng Gao, Zhao Wang et al.EMNLP 2022 · 5 citations
Builds on22
- CrossFormer: A Versatile Vision Transformer Hinging on Cross-scale AttentionWenxiao Wang, Lu Yao, Long Chen, Binbin Lin et al.ICLR 2022 · 367 citations
- Deconfounded Video Moment Retrieval with Causal InterventionXun Yang, Fuli Feng, Wei Ji, Meng Wang et al.SIGIR 2021 · 198 citations
- Boundary Proposal Network for Two-stage Natural Language Video LocalizationShaoning Xiao, Long Chen, Songyang Zhang, Wei Ji et al.AAAI 2021 · 186 citations
- Rethinking the Bottom-Up Framework for Query-Based Video LocalizationLong Chen, Chujie Lu, Siliang Tang, Jun Xiao et al.AAAI 2020 · 182 citations
- Spatial-Temporal Transformer for Dynamic Scene Graph GenerationYuren Cong, Wentong Liao, Hanno Ackermann, Bodo Rosenhahn et al.ICCV 2021 · 163 citations
Related papers
- Triple Correlations-Guided Label Supplementation for Unbiased Video Scene Graph GenerationWenqing Wang, Kaifeng Gao, Yawei Luo, Tao Jiang et al.ACM MM 2023 · 7 citations
- SGTR: End-to-end Scene Graph Generation with TransformerRongjie Li, Songyang Zhang, Xuming HeCVPR 2022 · 108 citations
- Constructing Holistic Spatio-Temporal Scene Graph for Video Semantic Role LabelingYu Zhao, Hao Fei, Yixin Cao, Bobo Li et al.ACM MM 2023 · 31 citations
- Unbiased Video Scene Graph Generation via Visual and Semantic Dual DebiasingYanjun Li, Zhaoyang Li, Honghui Chen, Lizhi XuCVPR 2025
- End-to-End Entity-Predicate Association Reasoning for Dynamic Scene Graph GenerationLiwei Wang, Yanduo Zhang, Tao Lu, Fang Liu et al.ICCV 2025 · 1 citation
