RLEG: Vision-Language Representation Learning with Diffusion-based Embedding Generation
Liming Zhao, Kecheng Zheng, Yun Zheng, Deli Zhao, Jingren Zhou
Abstract
Vision-language representation learning models (e.g., CLIP) have achieved state-of-the-art performance on various downstream tasks, which usually need large-scale training data to learn discriminative representation. Recent progress on generative diffusion models (e.g., DALL-E 2) has demonstrated that diverse high-quality samples can be synthesized by randomly sampling from generative distribution. By virtue of generative capability in this paper, we propose a novel visionlanguage Representation Learning method with diffusion-based Embedding Generation (RLEG), which exploits diffusion models to generate feature embedding online for learning effective vision-language representation. Specifically, we first adopt image and text encoders to extract the corresponding embeddings. Secondly, pretrained diffusion-based embedding generators are harnessed to transfer the embedding modality online between vision and language domains. The embeddings generated from the generators are then served as augmented embedding-level samples, which are applied to contrastive learning with the variant of the CLIP framework. Experimental results show that the proposed method could learn effective representation and achieve state-of-theart performance on various tasks including image classification, image-text retrieval, object detection, semantic segmentation, and text-conditional image generation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 80f6012e-7f99-491f-af99-6084cb7611acCited by top-tier papers2
- MomentDiff: Generative Video Moment Retrieval from Random to RealPandeng Li, Chen-Wei Xie, Hongtao Xie, Liming Zhao et al.NeurIPS 2023 · 113 citations
- LoTLIP: Improving Language-Image Pre-training for Long Text UnderstandingWei Wu, Kecheng Zheng, Shuailei Ma, Fan Lu et al.NeurIPS 2024 · 35 citations
Builds on17
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
Related papers
- RWKV-CLIP: A Robust Vision-Language Representation LearnerTiancheng Gu, Kaicheng Yang, Xiang An, Ziyong Feng et al.EMNLP 2024 · 11 citations
- Shifted Diffusion for Text-to-image GenerationYufan Zhou, Bingchen Liu, Yizhe Zhu, Xiao Yang et al.CVPR 2023
- DiffDis: Empowering Generative Diffusion Model with Cross-Modal Discrimination CapabilityRunhui Huang, Jianhua Han, Guansong Lu, Xiaodan Liang et al.ICCV 2023 · 10 citations
- Pre-trained Text-to-Image Diffusion Models Are Versatile Representation Learners for ControlGunshi Gupta, Karmesh Yadav, Yarin Gal, Dhruv Batra et al.NeurIPS 2024 · 17 citations
- RelCLIP: Adapting Language-Image Pretraining for Visual Relationship Detection via Relational Contrastive LearningYi Zhu, Zhaoqing Zhu, Bingqian Lin, Xiaodan Liang et al.EMNLP 2022 · 8 citations
