APT: Towards Universal Scene Graph Generation via Plug-in Adaptive Prompt Tuning
Ruikun Luo, Changwei Gu, Jing Yang, Yuan Gao, Jieming Yang, Song Wu, Hai Jin, Xiaoyu Xia
Abstract
Scene Graph Generation (SGG) is pivotal for structured visual understanding, yet it remains hindered by a fundamental limitation: the reliance on fixed, frozen semantic representations from pre-trained language models. These semantic priors, while beneficial in other domains, are inherently misaligned with the dynamic, context-sensitive nature of visual relationships, leading to biased and suboptimal performance. In this paper, we transcend the traditional one-stage v.s. two-stage architectural debate and identify this representational bottleneck as the core issue. We introduce Adaptive Prompt Tuning (APT), a universal paradigm that converts frozen semantic features into dynamic, context-aware representations through lightweight, learnable prompts. APT acts as a plug-in module that can be seamlessly integrated into existing SGG frameworks. Extensive experiments demonstrate that APT achieves +2.7 improvement in mR@100 on PredCls, +3.6 gain in F@100 and up to +6.0 gain in mR@50 in open-vocabulary novel splits. Notably, it achieves this with less than 0.5M additonal parameters (<1.5% overhead) and reduced 7.8%-25% training time, establishing a new state-of-the-art while offering a unified, efficient, and scalable solution for future SGG research. The source code of APT is available at https://github.com/CGCL-codes/APT.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 51f6b934-a1fd-4fbb-8291-22e2a8131017Builds on22
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Conditional Prompt Learning for Vision-Language ModelsKaiyang Zhou, Jingkang Yang, Chen Change Loy, Ziwei LiuCVPR 2022 · 1,438 citations
- InfoGCN: Representation Learning for Human Skeleton-based Action RecognitionHyung-Gun Chi, Myoung Hoon Ha, Seung-geun Chi, Sang Wan Lee et al.CVPR 2022 · 383 citations
- FedFed: Feature Distillation against Data Heterogeneity in Federated LearningZhiqin Yang, Yonggang Zhang, Yu Zheng, Xinmei Tian et al.NeurIPS 2023 · 166 citations
Related papers
- Scene Graph Generation with Role-Playing Large Language ModelsGuikun Chen, Jin Li, Wenguan WangNeurIPS 2024 · 33 citations
- From Pixels to Graphs: Open-Vocabulary Scene Graph Generation with Vision-Language ModelsRongjie Li, Songyang Zhang, Dahua Lin, Kai Chen et al.CVPR 2024
- Learning Hierarchical Prompt with Structured Linguistic Knowledge for Vision-Language ModelsYubin Wang, Xinyang Jiang, De Cheng, Dongsheng Li et al.AAAI 2024 · 50 citations
- Relation-aware Hierarchical Prompt for Open-vocabulary Scene Graph GenerationTao Liu, Rongjie Li, Chongyu Wang, Xuming HeAAAI 2025
- Predicate Debiasing in Vision-Language Models Integration for Scene Graph Generation EnhancementYuxuan Wang, Xiaoyuan LiuEMNLP 2024 · 1 citation
