Grounding Consistency: Distilling Spatial Common Sense for Precise Visual Relationship Detection
Markos Diomataris, Nikolaos Gkanatsios, Vassilis Pitsikalis, Petros Maragos
Abstract
Scene Graph Generators (SGGs) are models that, given an image, build a directed graph where each edge represents a predicted subject predicate object triplet. Most SGGs silently exploit datasets' bias on relationships' context, i.e. its subject and object, to improve recall and neglect spatial and visual evidence, e.g. having seen a glut of data for person wearing shirt, they are overconfident that every person is wearing every shirt. Such imprecise predictions are mainly ascribed to the lack of negative examples for most relationships, which obstructs models from meaningfully learning predicates, even those that have ample positive examples. We first present an indepth investigation of the context bias issue to showcase that all examined state-of-the-art SGGs share the above vulnerabilities. In response, we propose a semi-supervised scheme that forces predicted triplets to be grounded consistently back to the image, in a closed-loop manner. The developed spatial common sense can be then distilled to a student SGG and substantially enhance its spatial reasoning ability. This Grounding Consistency Distillation (GCD) approach is model-agnostic and benefits from the superfluous unlabeled samples to retain the valuable context information and avert memorization of annotations. Furthermore, we demonstrate that current metrics disregard unlabeled samples, rendering themselves incapable of reflecting context bias, then we mine and incorporate during evaluation hard-negatives to reformulate precision as a reliable metric. Extensive experimental comparisons exhibit large quantitative - up to 70% relative precision boost on VG200 dataset - and qualitative improvements to prove the significance of our GCD method and our metrics towards refocusing graph generation as a core aspect of scene understanding. Code available at https://github.com/deeplab-ai/grounding-consistent-vrd.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b3be1025-e95a-4bb7-982d-66846ebfb4cdCited by top-tier papers1
Ask how each one uses itBuilds on12
- S4L: Self-Supervised Semi-Supervised LearningLucas Beyer, Xiaohua Zhai, Avital Oliver, Alexander KolesnikovICCV 2019 · 854 citations
- NLNL: Negative Learning for Noisy LabelsYoungdong Kim, Junho Yim, Juseung Yun, Junmo KimICCV 2019 · 338 citations
- Detecting Unseen Visual Relations Using AnalogiesJulia Peyre, Josef Sivic, Ivan Laptev, Cordelia SchmidICCV 2019 · 135 citations
- One-Shot Learning for Long-Tail Visual Relation DetectionWeitao Wang, Meng Wang, Sen Wang, Guodong Long et al.AAAI 2020 · 20 citations
- Visual Relationship Detection with Low Rank Non-Negative Tensor DecompositionMohammed Haroon Dupty, Zhen Zhang, Wee Sun LeeAAAI 2020 · 9 citations
Related papers
- Dark Knowledge Balance Learning for Unbiased Scene Graph GenerationZhiqing Chen, Yawei Luo, Jian Shao, Yi Yang et al.ACM MM 2023 · 9 citations
- Semi-Supervised Clustering Framework for Fine-grained Scene Graph GenerationJiarui Yang, Chuan Wang, Jun Zhang, Shuyi Wu et al.AAAI 2025 · 2 citations
- Unbiased Video Scene Graph Generation via Visual and Semantic Dual DebiasingYanjun Li, Zhaoyang Li, Honghui Chen, Lizhi XuCVPR 2025
- DSGG: Dense Relation Transformer for an End-to-End Scene Graph GenerationZeeshan Hayder, Xuming HeCVPR 2024
- Visual Distant Supervision for Scene Graph GenerationYuan Yao, Ao Zhang, Xu Han, Mengdi Li et al.ICCV 2021 · 41 citations
