Contrastive Language-Image Pre-Training with Knowledge Graphs
Xuran Pan, Tianzhu Ye, Dongchen Han, Shiji Song, Gao Huang
摘要
Recent years have witnessed the fast development of large-scale pre-training frameworks that can extract multi-modal representations in a unified form and achieve promising performances when transferred to downstream tasks. Nevertheless, existing approaches mainly focus on pre-training with simple image-text pairs, while neglecting the semantic connections between concepts from different modalities. In this paper, we propose a knowledge-based pre-training framework, dubbed Knowledge-CLIP, which injects semantic information into the widely used CLIP model [38] . Through introducing knowledge-based objectives in the pre-training process and utilizing different types of knowledge graphs as training data, our model can semantically align the representations in vision and language with higher quality, and enhance the reasoning ability across scenarios and modalities. Extensive experiments on various vision-language downstream tasks demonstrate the effectiveness of Knowledge-CLIP compared with the original CLIP and competitive baselines.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper14
- FLatten Transformer: Vision Transformer using Focused Linear AttentionDongchen Han, Xuran Pan, Yizeng Han, Shiji Song 等ICCV 2023 · 被引用 358 次
- Structural Information Guided Multimodal Pre-training for Vehicle-Centric PerceptionXiao Wang, Wentao Wu, Chenglong Li, Zhicheng Zhao 等AAAI 2024 · 被引用 10 次
- Language Semantic Graph Guided Data-Efficient LearningWenxuan Ma, Shuang Li, Lincan Cai, Jingxuan KangNeurIPS 2023 · 被引用 6 次
- Does Progress On Object Recognition Benchmarks Improve Generalization on Crowdsourced, Global Data?Megan Richards, Polina Kirichenko, Diane Bouchacourt, Mark IbrahimICLR 2024 · 被引用 2 次
- VL-KGE: Vision-Language Models Meet Knowledge Graph EmbeddingsAthanasios Efthymiou, Stevan Rudinac, Monika Kackovic, Nachoem Wijnberg 等WWW 2026 · 被引用 2 次
它引用的顶会 Paper22
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa 等ICML 2021 · 被引用 8,974 次
- Align before Fuse: Vision and Language Representation Learning with Momentum DistillationJunnan Li, Ramprasaath R. Selvaraju, Akhilesh Gotmare, Shafiq R. Joty 等NeurIPS 2021 · 被引用 2,985 次
- ViLT: Vision-and-Language Transformer Without Convolution or Region SupervisionWonjae Kim, Bokyung Son, Ildoo KimICML 2021 · 被引用 2,258 次
相关 Paper
- Structure-CLIP: Towards Scene Graph Knowledge to Enhance Multi-Modal Structured RepresentationsYufeng Huang, Jiji Tang, Zhuo Chen, Rongsheng Zhang 等AAAI 2024 · 被引用 65 次
- Retrieval-based Knowledge Augmented Vision Language Pre-trainingJiahua Rao, Zifei Shan, Longpo Liu, Yao Zhou 等ACM MM 2023 · 被引用 13 次
- RankCLIP: Ranking-Consistent Language-Image PretrainingYiming Zhang, Zhuokai Zhao, Zhaorun Chen, Zhili Feng 等ICCV 2025 · 被引用 1 次
- Hierarchical Cross-Modal Prompt Learning for Vision-Language ModelsHao Zheng, Shunzhi Yang, Zhuoxin He, Jinfeng Yang 等ICCV 2025 · 被引用 5 次
- DenseCLIP: Language-Guided Dense Prediction with Context-Aware PromptingYongming Rao, Wenliang Zhao, Guangyi Chen, Yansong Tang 等CVPR 2022 · 被引用 527 次
