Learning to Generate Language-Supervised and Open-Vocabulary Scene Graph Using Pre-Trained Visual-Semantic Space
Yong Zhang, Yingwei Pan, Ting Yao, Rui Huang, Tao Mei, Chang Wen Chen
Abstract
Scene graph generation (SGG) aims to abstract an image into a graph structure, by representing objects as graph nodes and their relations as labeled edges. However, two knotty obstacles limit the practicability of current SGG methods in real-world scenarios: 1) training SGG models requires time-consuming ground-truth annotations, and 2) the closed-set object categories make the SGG models limited in their ability to recognize novel objects outside of training corpora. To address these issues, we novelly exploit a powerful pre-trained visual-semantic space (VSS) to trigger language-supervised and open-vocabulary SGG in a simple yet effective manner. Specifically, cheap scene graph supervision data can be easily obtained by parsing image language descriptions into semantic graphs. Next, the noun phrases on such semantic graphs are directly grounded over image regions through region-word alignment in the pre-trained VSS. In this way, we enable open-vocabulary object detection by performing object category name grounding with a text prompt in this VSS. On the basis of visually-grounded objects, the relation representations are naturally built for relation recognition, pursuing open-vocabulary SGG. We validate our proposed approach with extensive experiments on the Visual Genome benchmark across various SGG scenarios (i.e., supervised / language-supervised, closed-set / open-vocabulary). Consistent superior performances are achieved compared with existing methods, demonstrating the potential of exploiting pre-trained VSS for SGG in more practical scenarios.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a0cebdc6-7fa1-4519-90b0-a4544ecbd2e7Cited by top-tier papers18
- Scene Graph Generation with Role-Playing Large Language ModelsGuikun Chen, Jin Li, Wenguan WangNeurIPS 2024 · 33 citations
- Multi-Prototype Space Learning for Commonsense-Based Scene Graph GenerationLianggangxu Chen, Youqi Song, Yiqing Cai, Jiale Lu et al.AAAI 2024 · 11 citations
- Improving Scene Graph Generation with Superpixel-Based Interaction LearningJingyi Wang, Can Zhang, Jinfa Huang, Botao Ren et al.ACM MM 2023 · 8 citations
- Vision-Language Interactive Relation Mining for Open-Vocabulary Scene Graph GenerationYukuan Min, Muli Yang, Jinhao Zhang, Yuxuan Wang et al.ICCV 2025 · 2 citations
- SPADE: Spatial-Aware Denoising Network for Open-Vocabulary Panoptic Scene Graph Generation with Long- and Local-Range Context ReasoningXin Hu, Ke Qin, Guiduo Duan, Ming Li et al.ICCV 2025 · 1 citation
Builds on33
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen et al.ICML 2021 · 5,401 citations
- Align before Fuse: Vision and Language Representation Learning with Momentum DistillationJunnan Li, Ramprasaath R. Selvaraju, Akhilesh Gotmare, Shafiq R. Joty et al.NeurIPS 2021 · 2,985 citations
- VL-BERT: Pre-training of Generic Visual-Linguistic RepresentationsWeijie Su, Xizhou Zhu, Yue Cao, Bin Li et al.ICLR 2020 · 1,825 citations
Related papers
- From Pixels to Graphs: Open-Vocabulary Scene Graph Generation with Vision-Language ModelsRongjie Li, Songyang Zhang, Dahua Lin, Kai Chen et al.CVPR 2024
- Visual Distant Supervision for Scene Graph GenerationYuan Yao, Ao Zhang, Xu Han, Mengdi Li et al.ICCV 2021 · 41 citations
- IS-GGT: Iterative Scene Graph Generation with Generative TransformersSanjoy Kundu, Sathyanarayanan N. AakurCVPR 2023
- Mixture-of-Experts based Feature Decoupling for Open Vocabulary Scene Graph GenerationYiming Li, Sisi You, Bing-Kun BaoCVPR 2026
- Open-Vocabulary Video Scene Graph Generation via Union-aware Semantic AlignmentZiyue Wu, Junyu Gao, Changsheng XuACM MM 2024 · 8 citations
