Structure-CLIP: Towards Scene Graph Knowledge to Enhance Multi-Modal Structured Representations
Yufeng Huang, Jiji Tang, Zhuo Chen, Rongsheng Zhang, Xinfeng Zhang, Weijie Chen, Zeng Zhao, Zhou Zhao, Tangjie Lv, Zhipeng Hu, Wen Zhang
Abstract
Large-scale vision-language pre-training has achieved significant performance in multi-modal understanding and generation tasks. However, existing methods often perform poorly on image-text matching tasks that require structured representations, i.e., representations of objects, attributes, and relations. As illustrated in Fig. 1 (a), the models cannot make a distinction between "An astronaut rides a horse" and "A horse rides an astronaut". This is because they fail to fully leverage structured knowledge when learning representations in multi-modal scenarios. In this paper, we present an end-to-end framework Structure-CLIP, which integrates Scene Graph Knowledge (SGK) to enhance multimodal structured representations. Firstly, we use scene graphs to guide the construction of semantic negative examples, which results in an increased emphasis on learning structured representations. Moreover, a Knowledge-Enhance Encoder (KEE) is proposed to leverage SGK as input to further enhance structured representations. To verify the effectiveness of the proposed framework, we pre-train our model with the aforementioned approaches and conduct experiments on downstream tasks. Experimental results demonstrate that Structure-CLIP achieves state-of-the-art (SOTA) performance on VG-Attribution and VG-Relation datasets, with 12.5% and 4.1% ahead of the multi-modal SOTA model respectively. Meanwhile, the results on MSCOCO indicate that Structure-CLIP significantly enhances the structured representations while maintaining the ability of general representations. Our code is available at https://github.com/zjukg/ Structure-CLIP .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers16
- CF-VLM: CounterFactual Vision-Language Fine-tuningJusheng Zhang, Kaitong Cai, Yijia Fan, Jian Wang et al.NeurIPS 2025 · 71 citations
- Beyond Accuracy: Ensuring Correct Predictions With Correct RationalesTang Li, Mengmeng Ma, Xi PengNeurIPS 2024 · 6 citations
- EvdCLIP: Improving Vision-Language Retrieval with Entity Visual Descriptions from Large Language ModelsGuanghao Meng, Sunan He, Jinpeng Wang, Tao Dai et al.AAAI 2025 · 5 citations
- Enhancing Compositional Reasoning in CLIP via Reconstruction and Alignment of Text DescriptionsJihoon Kwon, Kyle Min, Jy-yong SohnNeurIPS 2025 · 4 citations
- Decoupled Global-Local Alignment for Improving Compositional UnderstandingXiaoxing Hu, Kaicheng Yang, Jun Wang, Haoran Xu et al.ACM MM 2025 · 4 citations
Builds on16
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 6,549 citations
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen et al.ICML 2021 · 5,401 citations
- ViLT: Vision-and-Language Transformer Without Convolution or Region SupervisionWonjae Kim, Bokyung Son, Ildoo KimICML 2021 · 2,258 citations
- BARTScore: Evaluating Generated Text as Text GenerationWeizhe Yuan, Graham Neubig, Pengfei LiuNeurIPS 2021 · 1,143 citations
Related papers
- Contrastive Language-Image Pre-Training with Knowledge GraphsXuran Pan, Tianzhu Ye, Dongchen Han, Shiji Song et al.NeurIPS 2022 · 81 citations
- ERNIE-ViL: Knowledge Enhanced Vision-Language Representations through Scene GraphsFei Yu, Jiji Tang, Weichong Yin, Yu Sun et al.AAAI 2021 · 414 citations
- Coarse-to-Fine Contrastive Learning in Image-Text-Graph Space for Improved Vision-Language CompositionalityHarman Singh, Pengchuan Zhang, Qifan Wang, Mengjiao Wang et al.EMNLP 2023 · 10 citations
- Object-centric binding in Contrastive Language-Image PretrainingRim Assouel, Pietro Astolfi, Florian Bordes, Michal Drozdzal et al.NeurIPS 2025 · 14 citations
- LBMKGC: Large Model-Driven Balanced Multimodal Knowledge Graph CompletionYuan Guo, Qian Ma, Hui Li, Qiao Ning et al.NeurIPS 2025 · 3 citations
