APT: Towards Universal Scene Graph Generation via Plug-in Adaptive Prompt Tuning
Ruikun Luo, Changwei Gu, Jing Yang, Yuan Gao, Jieming Yang, Song Wu, Hai Jin, Xiaoyu Xia
摘要
Scene Graph Generation (SGG) is pivotal for structured visual understanding, yet it remains hindered by a fundamental limitation: the reliance on fixed, frozen semantic representations from pre-trained language models. These semantic priors, while beneficial in other domains, are inherently misaligned with the dynamic, context-sensitive nature of visual relationships, leading to biased and suboptimal performance. In this paper, we transcend the traditional one-stage v.s. two-stage architectural debate and identify this representational bottleneck as the core issue. We introduce Adaptive Prompt Tuning (APT), a universal paradigm that converts frozen semantic features into dynamic, context-aware representations through lightweight, learnable prompts. APT acts as a plug-in module that can be seamlessly integrated into existing SGG frameworks. Extensive experiments demonstrate that APT achieves +2.7 improvement in mR@100 on PredCls, +3.6 gain in F@100 and up to +6.0 gain in mR@50 in open-vocabulary novel splits. Notably, it achieves this with less than 0.5M additonal parameters (<1.5% overhead) and reduced 7.8%-25% training time, establishing a new state-of-the-art while offering a unified, efficient, and scalable solution for future SGG research. The source code of APT is available at https://github.com/CGCL-codes/APT.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper22
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Conditional Prompt Learning for Vision-Language ModelsKaiyang Zhou, Jingkang Yang, Chen Change Loy, Ziwei LiuCVPR 2022 · 被引用 1,438 次
- InfoGCN: Representation Learning for Human Skeleton-based Action RecognitionHyung-Gun Chi, Myoung Hoon Ha, Seung-geun Chi, Sang Wan Lee 等CVPR 2022 · 被引用 383 次
- FedFed: Feature Distillation against Data Heterogeneity in Federated LearningZhiqin Yang, Yonggang Zhang, Yu Zheng, Xinmei Tian 等NeurIPS 2023 · 被引用 166 次
相关 Paper
- Scene Graph Generation with Role-Playing Large Language ModelsGuikun Chen, Jin Li, Wenguan WangNeurIPS 2024 · 被引用 33 次
- From Pixels to Graphs: Open-Vocabulary Scene Graph Generation with Vision-Language ModelsRongjie Li, Songyang Zhang, Dahua Lin, Kai Chen 等CVPR 2024
- Learning Hierarchical Prompt with Structured Linguistic Knowledge for Vision-Language ModelsYubin Wang, Xinyang Jiang, De Cheng, Dongsheng Li 等AAAI 2024 · 被引用 50 次
- Relation-aware Hierarchical Prompt for Open-vocabulary Scene Graph GenerationTao Liu, Rongjie Li, Chongyu Wang, Xuming HeAAAI 2025
- Predicate Debiasing in Vision-Language Models Integration for Scene Graph Generation EnhancementYuxuan Wang, Xiaoyuan LiuEMNLP 2024 · 被引用 1 次
