Structure-CLIP: Towards Scene Graph Knowledge to Enhance Multi-Modal Structured Representations
Yufeng Huang, Jiji Tang, Zhuo Chen, Rongsheng Zhang, Xinfeng Zhang, Weijie Chen, Zeng Zhao, Zhou Zhao, Tangjie Lv, Zhipeng Hu, Wen Zhang
摘要
Large-scale vision-language pre-training has achieved significant performance in multi-modal understanding and generation tasks. However, existing methods often perform poorly on image-text matching tasks that require structured representations, i.e., representations of objects, attributes, and relations. As illustrated in Fig. 1 (a), the models cannot make a distinction between "An astronaut rides a horse" and "A horse rides an astronaut". This is because they fail to fully leverage structured knowledge when learning representations in multi-modal scenarios. In this paper, we present an end-to-end framework Structure-CLIP, which integrates Scene Graph Knowledge (SGK) to enhance multimodal structured representations. Firstly, we use scene graphs to guide the construction of semantic negative examples, which results in an increased emphasis on learning structured representations. Moreover, a Knowledge-Enhance Encoder (KEE) is proposed to leverage SGK as input to further enhance structured representations. To verify the effectiveness of the proposed framework, we pre-train our model with the aforementioned approaches and conduct experiments on downstream tasks. Experimental results demonstrate that Structure-CLIP achieves state-of-the-art (SOTA) performance on VG-Attribution and VG-Relation datasets, with 12.5% and 4.1% ahead of the multi-modal SOTA model respectively. Meanwhile, the results on MSCOCO indicate that Structure-CLIP significantly enhances the structured representations while maintaining the ability of general representations. Our code is available at https://github.com/zjukg/ Structure-CLIP .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper16
- CF-VLM: CounterFactual Vision-Language Fine-tuningJusheng Zhang, Kaitong Cai, Yijia Fan, Jian Wang 等NeurIPS 2025 · 被引用 71 次
- Beyond Accuracy: Ensuring Correct Predictions With Correct RationalesTang Li, Mengmeng Ma, Xi PengNeurIPS 2024 · 被引用 6 次
- EvdCLIP: Improving Vision-Language Retrieval with Entity Visual Descriptions from Large Language ModelsGuanghao Meng, Sunan He, Jinpeng Wang, Tao Dai 等AAAI 2025 · 被引用 5 次
- Enhancing Compositional Reasoning in CLIP via Reconstruction and Alignment of Text DescriptionsJihoon Kwon, Kyle Min, Jy-yong SohnNeurIPS 2025 · 被引用 4 次
- Decoupled Global-Local Alignment for Improving Compositional UnderstandingXiaoxing Hu, Kaicheng Yang, Jun Wang, Haoran Xu 等ACM MM 2025 · 被引用 4 次
它引用的顶会 Paper16
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 被引用 6,549 次
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen 等ICML 2021 · 被引用 5,401 次
- ViLT: Vision-and-Language Transformer Without Convolution or Region SupervisionWonjae Kim, Bokyung Son, Ildoo KimICML 2021 · 被引用 2,258 次
- BARTScore: Evaluating Generated Text as Text GenerationWeizhe Yuan, Graham Neubig, Pengfei LiuNeurIPS 2021 · 被引用 1,143 次
相关 Paper
- Contrastive Language-Image Pre-Training with Knowledge GraphsXuran Pan, Tianzhu Ye, Dongchen Han, Shiji Song 等NeurIPS 2022 · 被引用 81 次
- ERNIE-ViL: Knowledge Enhanced Vision-Language Representations through Scene GraphsFei Yu, Jiji Tang, Weichong Yin, Yu Sun 等AAAI 2021 · 被引用 414 次
- Coarse-to-Fine Contrastive Learning in Image-Text-Graph Space for Improved Vision-Language CompositionalityHarman Singh, Pengchuan Zhang, Qifan Wang, Mengjiao Wang 等EMNLP 2023 · 被引用 10 次
- Object-centric binding in Contrastive Language-Image PretrainingRim Assouel, Pietro Astolfi, Florian Bordes, Michal Drozdzal 等NeurIPS 2025 · 被引用 14 次
- LBMKGC: Large Model-Driven Balanced Multimodal Knowledge Graph CompletionYuan Guo, Qian Ma, Hui Li, Qiao Ning 等NeurIPS 2025 · 被引用 3 次
