LLM4SGG: Large Language Models for Weakly Supervised Scene Graph Generation
Kibum Kim, Kanghoon Yoon, Jaehyeong Jeon, Yeonjun In, Jinyoung Moon, Donghyun Kim, Chanyoung Park
Abstract
Weakly-Supervised Scene Graph Generation (WSSGG) research has recently emerged as an alternative to the fullysupervised approach that heavily relies on costly annotations. In this regard, studies on WSSGG have utilized image captions to obtain unlocalized triplets while primarily focusing on grounding the unlocalized triplets over image regions. However, they have overlooked the two issues involved in the triplet formation process from the captions: 1) Semantic over-simplification issue arises when extracting triplets from captions, where fine-grained predicates in captions are undesirably converted into coarse-grained predicates, resulting in a long-tailed predicate distribution, and 2) Low-density scene graph issue arises when aligning the triplets in the caption with entity/predicate classes of interest, where many triplets are discarded and not used in training, leading to insufficient supervision. To tackle the two issues, we propose a new approach, i.e., Large Language Model for weakly-supervised SGG (LLM4SGG), where we mitigate the two issues by leveraging the LLM's in-depth understanding of language and reasoning ability during the extraction of triplets from captions and alignment of entity/predicate classes with target data. To further engage the LLM in these processes, we adopt the idea of Chainof-Thought and the in-context few-shot learning strategy. To validate the effectiveness of LLM4SGG, we conduct extensive experiments on Visual Genome and GQA datasets, showing significant improvements in both Recall@K and mean Recall@K compared to the state-of-the-art WSSGG methods. A further appeal is that LLM4SGG is dataefficient, enabling effective model training with a small amount of training images. Our code is available on https://github.com/rlqja1107/torch-LLM4SGG
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ccfc3899-a363-4b35-ab30-9c04a48a8138Cited by top-tier papers15
- RA-SGG: Retrieval-Augmented Scene Graph Generation Framework via Multi-Prototype LearningKanghoon Yoon, Kibum Kim, Jaehyeong Jeon, Yeonjun In et al.AAAI 2025 · 8 citations
- LLM Meets Scene Graph: Can Large Language Models Understand and Generate Scene Graphs? A Benchmark and Empirical StudyDongil Yang, Minjin Kim, Sunghwan Kim, Beong-woo Kwak et al.ACL 2025 · 8 citations
- Robo-SGG: Exploiting Layout-Oriented Normalization and Restitution Can Improve Robust Scene Graph GenerationChangsheng Lv, Zijian Fu, Mengshi QiCVPR 2026 · 4 citations
- Knowledge Image Matters: Improving Knowledge-Based Visual Reasoning with Multi-Image Large Language ModelsGuanghui Ye, Huan Zhao, Zhixue Zhao, Xupeng Zha et al.ACL 2025 · 2 citations
- Token-Efficient Item Representation via Images for LLM Recommender SystemsKibum Kim, Sein Kim, Hongseok Kang, Jiwan Kim et al.ICLR 2026 · 2 citations
Builds on26
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language ModelsDeyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li et al.ICLR 2024 · 3,079 citations
- Align before Fuse: Vision and Language Representation Learning with Momentum DistillationJunnan Li, Ramprasaath R. Selvaraju, Akhilesh Gotmare, Shafiq R. Joty et al.NeurIPS 2021 · 2,985 citations
Related papers
- Weakly Supervised Video Scene Graph Generation via Natural Language SupervisionKibum Kim, Kanghoon Yoon, Yeonjun In, Jaehyeong Jeon et al.ICLR 2025
- Linguistic Structures As Weak Supervision for Visual Scene Graph GenerationKeren Ye, Adriana KovashkaCVPR 2021
- Learning to Generate Language-Supervised and Open-Vocabulary Scene Graph Using Pre-Trained Visual-Semantic SpaceYong Zhang, Yingwei Pan, Ting Yao, Rui Huang et al.CVPR 2023
- Leveraging Large Language Models for Node Generation in Few-Shot Learning on Text-Attributed GraphsJianxiang Yu, Yuxiang Ren, Chenghua Gong, Jiaqi Tan et al.AAAI 2025 · 32 citations
- Not All Relations are Equal: Mining Informative Labels for Scene Graph GenerationArushi Goel, Basura Fernando, Frank Keller, Hakan BilenCVPR 2022 · 30 citations
