Zero-Shot Contrastive Loss for Text-Guided Diffusion Image Style Transfer
Serin Yang, Hyunmin Hwang, Jong Chul Ye
Abstract
Diffusion models have shown great promise in text-guided image style transfer, but there is a trade-off between style transformation and content preservation due to their stochastic nature. Existing methods require computationally expensive fine-tuning of diffusion models or additional neural network. To address this, here we propose a zero-shot contrastive loss for diffusion models that doesn’t require additional fine-tuning or auxiliary networks. By leveraging patch-wise contrastive loss between generated samples and original image embeddings in the pre-trained diffusion model, our method can generate images with the same semantic content as the source image in a zero-shot manner. Our approach outperforms existing methods while preserving content and requiring no additional training, not only for image style transfer but also for image-to-image translation and manipulation. Our experimental results validate the effectiveness of our proposed method. Code is available at https://github.com/YSerin/ZeCon.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 05842e84-32a2-4b17-8fd5-abb4256b684dCited by top-tier papers28
- ControlStyle: Text-Driven Stylized Image Generation Using Diffusion PriorsJingwen Chen, Yingwei Pan, Ting Yao, Tao MeiACM MM 2023 · 45 citations
- Contrastive Denoising Score for Text-Guided Latent Diffusion Image EditingHyelin Nam, Gihyun Kwon, Geon Yeong Park, Jong Chul YeCVPR 2024 · 27 citations
- Infrared and Visible Image Fusion with Language-Driven Loss in CLIP Embedding SpaceYuhao Wang, Lingjuan Miao, Zhiqiang Zhou, Lei Zhang et al.ACM MM 2025 · 18 citations
- GDA: Generalized Diffusion for Robust Test-Time AdaptationYun-Yun Tsai, Fu-Chen Chen, Albert Y. C. Chen, Junfeng Yang et al.CVPR 2024 · 11 citations
- Localize, Understand, Collaborate: Semantic-Aware Dragging via Intention ReasonerXing Cui, Peipei Li, Zekun Li, Xuannan Liu et al.NeurIPS 2024 · 11 citations
Builds on21
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
Related papers
- Zero-shot Image-to-Image TranslationGaurav Parmar, Krishna Kumar Singh, Richard Zhang, Yijun Li et al.SIGGRAPH 2023 · 355 citations
- Diffusion-based Image Translation using disentangled style and content representationGihyun Kwon, Jong Chul YeICLR 2023 · 46 citations
- Steered Diffusion: A Generalized Framework for Plug-and-Play Conditional Image SynthesisNithin Gopalakrishnan Nair, Anoop Cherian, Suhas Lohit, Ye Wang et al.ICCV 2023 · 22 citations
- UniversalBooth: Model-Agnostic Personalized Text-To-Image GenerationSonghua Liu, Ruonan Yu, Xinchao WangICCV 2025 · 2 citations
- Less is More: Masking Elements in Image Condition Features Avoids Content Leakages in Style Transfer Diffusion ModelsLin Zhu, Xinbing Wang, Chenghu Zhou, Qinying Gu et al.ICLR 2025
