Visual Distant Supervision for Scene Graph Generation
Yuan Yao, Ao Zhang, Xu Han, Mengdi Li, Cornelius Weber, Zhiyuan Liu, Stefan Wermter, Maosong Sun
Abstract
Scene graph generation aims to identify objects and their relations in images, providing structured image representations that can facilitate numerous applications in computer vision. However, scene graph models usually require supervised learning on large quantities of labeled data with intensive human annotation. In this work, we propose visual distant supervision, a novel paradigm of visual relation learning, which can train scene graph models without any human-labeled data. The intuition is that by aligning commonsense knowledge bases and images, we can automatically create large-scale labeled data to provide distant supervision for visual relation learning. To alleviate the noise in distantly labeled data, we further propose a framework that iteratively estimates the probabilistic relation labels and eliminates the noisy ones. Comprehensive experimental results show that our distantly supervised model outperforms strong weakly supervised and semi-supervised baselines. By further incorporating human-labeled data in a semi-supervised fashion, our model outperforms state-of-the-art fully supervised models by a large margin (e.g., 8.3 micro- and 7.8 macro-recall@50 improvements for predicate classification in Visual Genome evaluation). We make the data and code for this paper publicly available at https://github.com/thunlp/VisualDS.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers18
- SGTR: End-to-end Scene Graph Generation with TransformerRongjie Li, Songyang Zhang, Xuming HeCVPR 2022 · 108 citations
- Visually-Prompted Language Model for Fine-Grained Scene Graph Generation in an Open WorldQifan Yu, Juncheng Li, Yu Wu, Siliang Tang et al.ICCV 2023 · 51 citations
- DP-SSL: Towards Robust Semi-supervised Learning with A Few Labeled SamplesYi Xu, Jiandong Ding, Lu Zhang, Shuigeng ZhouNeurIPS 2021 · 34 citations
- Scene Graph Generation with Role-Playing Large Language ModelsGuikun Chen, Jin Li, Wenguan WangNeurIPS 2024 · 33 citations
- Multi-Prototype Space Learning for Commonsense-Based Scene Graph GenerationLianggangxu Chen, Youqi Song, Yiqing Cai, Jiale Lu et al.AAAI 2024 · 11 citations
Builds on8
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Unpaired Image Captioning via Scene Graph AlignmentsJiuxiang Gu, Shafiq R. Joty, Jianfei Cai, Handong Zhao et al.ICCV 2019 · 191 citations
- NOTE-RCNN: NOise Tolerant Ensemble RCNN for Semi-Supervised Object DetectionJiyang Gao, Jiang Wang, Shengyang Dai, Li-Jia Li et al.ICCV 2019 · 99 citations
- Scene Graph Prediction With Limited LabelsRanjay Krishna, Vincent S. Chen, Paroma Varma, Michael S. Bernstein et al.ICCV 2019 · 5 citations
- Learning From Noisy Anchors for One-Stage Object DetectionHengduo Li, Zuxuan Wu, Chen Zhu, Caiming Xiong et al.CVPR 2020
Related papers
- Learning to Generate Language-Supervised and Open-Vocabulary Scene Graph Using Pre-Trained Visual-Semantic SpaceYong Zhang, Yingwei Pan, Ting Yao, Rui Huang et al.CVPR 2023
- A Simple Baseline for Weakly-Supervised Scene Graph GenerationJing Shi, Yiwu Zhong, Ning Xu, Yin Li et al.ICCV 2021 · 34 citations
- Not All Relations are Equal: Mining Informative Labels for Scene Graph GenerationArushi Goel, Basura Fernando, Frank Keller, Hakan BilenCVPR 2022 · 30 citations
- Weakly Supervised Visual Semantic ParsingAlireza Zareian, Svebor Karaman, Shih-Fu ChangCVPR 2020
- Dynamic Scene Graph Generation via Anticipatory Pre-trainingYiming Li, Xiaoshan Yang, Changsheng XuCVPR 2022 · 38 citations
