Scene Graph Generation with Role-Playing Large Language Models
Guikun Chen, Jin Li, Wenguan Wang
Abstract
Current approaches for open-vocabulary scene graph generation (OVSGG) use vision-language models such as CLIP and follow a standard zero-shot pipeline -- computing similarity between the query image and the text embeddings for each category (i.e., text classifiers). In this work, we argue that the text classifiers adopted by existing OVSGG methods, i.e., category-/part-level prompts, are scene-agnostic as they remain unchanged across contexts. Using such fixed text classifiers not only struggles to model visual relations with high variance, but also falls short in adapting to distinct contexts. To plug these intrinsic shortcomings, we devise SDSGG, a scene-specific description based OVSGG framework where the weights of text classifiers are adaptively adjusted according to the visual content. In particular, to generate comprehensive and diverse descriptions oriented to the scene, an LLM is asked to play different roles (e.g., biologist and engineer) to analyze and discuss the descriptive features of a given scene from different views. Unlike previous efforts simply treating the generated descriptions as mutually equivalent text classifiers, SDSGG is equipped with an advanced renormalization mechanism to adjust the influence of each text classifier based on its relevance to the presented scene (this is what the term"specific"means). Furthermore, to capture the complicated interplay between subjects and objects, we propose a new lightweight module called mutual visual adapter. It refines CLIP's ability to recognize relations by learning an interaction-aware semantic space. Extensive experiments on prevalent benchmarks show that SDSGG outperforms top-leading methods by a clear margin.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3f0c9c4d-5b06-4e16-a59a-47ba70cdd0f9Cited by top-tier papers13
- Human-Object Interaction Detection Collaborated with Large Relation-driven Diffusion ModelsLiulei Li, Wenguan Wang, Yi YangNeurIPS 2024 · 29 citations
- OmniGaze: Reward-inspired Generalizable Gaze Estimation in the WildHongyu Qu, Jianan Wei, Xiangbo Shu, Yazhou Yao et al.NeurIPS 2025 · 15 citations
- Learning Human-Object Interaction as GroupsJiajun Hong, Jianan Wei, Wenguan WangNeurIPS 2025 · 6 citations
- Relation-R1: Progressively Cognitive Chain-of-Thought Guided Reinforcement Learning for Unified Relation ComprehensionLin Li, Wei Chen, Jiahui Li, Kwang-Ting Cheng et al.AAAI 2026 · 6 citations
- SinkTrack: Attention Sink based Context Anchoring for Large Language ModelsXu Liu, Guikun Chen, Wenguan WangICLR 2026 · 4 citations
Builds on60
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 6,549 citations
- Big Self-Supervised Models are Strong Semi-Supervised LearnersTing Chen, Simon Kornblith, Kevin Swersky, Mohammad Norouzi et al.NeurIPS 2020 · 2,611 citations
Related papers
- Relation-aware Hierarchical Prompt for Open-vocabulary Scene Graph GenerationTao Liu, Rongjie Li, Chongyu Wang, Xuming HeAAAI 2025
- Mixture-of-Experts based Feature Decoupling for Open Vocabulary Scene Graph GenerationYiming Li, Sisi You, Bing-Kun BaoCVPR 2026
- CLIP-Adapted Region-to-Text Learning for Generative Open-Vocabulary Semantic SegmentationJiannan Ge, Lingxi Xie, Hongtao Xie, Pandeng Li et al.ICCV 2025 · 3 citations
- Vision-Language Interactive Relation Mining for Open-Vocabulary Scene Graph GenerationYukuan Min, Muli Yang, Jinhao Zhang, Yuxuan Wang et al.ICCV 2025 · 2 citations
- From Pixels to Graphs: Open-Vocabulary Scene Graph Generation with Vision-Language ModelsRongjie Li, Songyang Zhang, Dahua Lin, Kai Chen et al.CVPR 2024
