Visually-Prompted Language Model for Fine-Grained Scene Graph Generation in an Open World
Qifan Yu, Juncheng Li, Yu Wu, Siliang Tang, Wei Ji, Yueting Zhuang
摘要
Scene Graph Generation (SGG) aims to extract <subject, predicate, object> relationships in images for vision understanding. Although recent works have made steady progress on SGG, they still suffer long-tail distribution issues that tail-predicates are more costly to train and hard to distinguish due to a small amount of annotated data compared to frequent predicates. Existing re-balancing strategies try to handle it via prior rules but are still confined to pre-defined conditions, which are not scalable for various models and datasets. In this paper, we propose a Crossmodal prediCate boosting (CaCao) framework, where a visually-prompted language model is learned to generate diverse fine-grained predicates in a low-resource way. The proposed CaCao can be applied in a plug-and-play fashion and automatically strengthen existing SGG to tackle the long-tailed problem. Based on that, we further introduce a novel Entangled cross-modal prompt approach for open-world predicate scene graph generation (Epic), where models can generalize to unseen predicates in a zeroshot manner. Comprehensive experiments on three benchmark datasets show that CaCao consistently boosts the performance of multiple scene graph generation models in a model-agnostic way. Moreover, our Epic achieves competitive performance on open-world predicate prediction. The data and code for this paper are publicly available. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper27
- Fine-tuning Multimodal LLMs to Follow Zero-shot Demonstrative InstructionsJuncheng Li, Kaihang Pan, Zhiqi Ge, Minghe Gao 等ICLR 2024 · 被引用 95 次
- Zero-shot Visual Relation Detection via Composite Visual Cues from Large Language ModelsLin Li, Jun Xiao, Guikun Chen, Jian Shao 等NeurIPS 2023 · 被引用 52 次
- Compositional Feature Augmentation for Unbiased Scene Graph GenerationLin Li, Guikun Chen, Jun Xiao, Yi Yang 等ICCV 2023 · 被引用 36 次
- Gradient-Regulated Meta-Prompt Learning for Generalizable Vision-Language ModelsJuncheng Li, Minghe Gao, Longhui Wei, Siliang Tang 等ICCV 2023 · 被引用 34 次
- Scene Graph Generation with Role-Playing Large Language ModelsGuikun Chen, Jin Li, Wenguan WangNeurIPS 2024 · 被引用 33 次
它引用的顶会 Paper33
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen 等ICML 2021 · 被引用 5,401 次
- ViLT: Vision-and-Language Transformer Without Convolution or Region SupervisionWonjae Kim, Bokyung Son, Ildoo KimICML 2021 · 被引用 2,258 次
相关 Paper
- Learning to Generate an Unbiased Scene Graph by Using Attribute-Guided Predicate FeaturesLei Wang, Zejian Yuan, Badong ChenAAAI 2023 · 被引用 8 次
- PPDL: Predicate Probability Distribution based Loss for Unbiased Scene Graph GenerationWei Li, Haiwei Zhang, Qijie Bai, Guoqing Zhao 等CVPR 2022 · 被引用 64 次
- Fast Contextual Scene Graph Generation with Unbiased Context AugmentationTianlei Jin, Fangtai Guo, Qiwei Meng, Shiqiang Zhu 等CVPR 2023
- From General to Specific: Informative Scene Graph Generation via Balance AdjustmentYuyu Guo, Lianli Gao, Xuanhan Wang, Yuxuan Hu 等ICCV 2021 · 被引用 96 次
- Fine-Grained Predicates Learning for Scene Graph GenerationXinyu Lyu, Lianli Gao, Yuyu Guo, Zhou Zhao 等CVPR 2022 · 被引用 48 次
