Text-to-Image Diffusion Models are Great Sketch-Photo Matchmakers
Subhadeep Koley, Ayan Kumar Bhunia, Aneeshan Sain, Pinaki Nath Chowdhury, Tao Xiang, Yi-Zhe Song
Abstract
Figure 1. Sketch-based image retrieval frameworks [12, 102, 103] usually employ ImageNet pre-trained CNNs [3, 15, 103], JFT-trained vision transformers (ViT) [67, 78], or visual encoders of vision-language models like CLIP [77] as backbone feature extractors. Rich knowledge from large-scale pre-training offers a good initialisation, which when further fine-tuned on sketch-photo datasets, performs way better than training from random initialisation [62]. While one can extract features either by discarding the classification head for ImageNet pre-trained models, auxiliary task head for self-supervised models, or by using CLIP's visual encoder, text-to-image diffusion models (e.g., stable diffusion) lack any specific feature embedding space. However, we find that its intermediate representations implicitly hold robust cross-modal features at multiple granularities. Unlike prior SBIR backbones, pre-trained with discriminative tasks, we propose to leverage denoising diffusion models pre-trained with text-to-image generative tasks to bridge the sketch-photo domain gap. Being a text-to-image generation model trained on a large corpus of text-image pairs (LAION [82, 83]), it holds both semantic and shape prior [89]. PCA representation [89] (right) of intermediate UNet features (sketch/photo) from different upsampling blocks (details in Sec. 5) depict that they share significant semantic similarity (denoted by similar colours).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 255b7a12-fa42-4d97-a02f-562d34a45dffCited by top-tier papers12
- VQ-SGen: A Vector Quantized Stroke Representation for Creative Sketch GenerationJiawei Wang, Zhiming Cui, Changjian LiICCV 2025 · 3 citations
- Leveraging Prior Knowledge of Diffusion Model for Person SearchGiyeol Kim, Sooyoung Yang, Jihyong Oh, Myungjoo Kang et al.ICCV 2025 · 2 citations
- Modeling the Visual Ambiguity of Human SketchesYang Zhou, Ping Ni, Jin Wang, Senyun Jia et al.CVPR 2026
- You'll Never Walk Alone: A Sketch and Text Duet for Fine-Grained Image RetrievalSubhadeep Koley, Ayan Kumar Bhunia, Aneeshan Sain, Pinaki Nath Chowdhury et al.CVPR 2024
- Intermediate Domain Alignment and Morphology Analogy for Patent-Product Image RetrievalHaifan Gong, Xuanye Zhang, Ruifei Zhang, Yun Su et al.NeurIPS 2025
Builds on53
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
Related papers
- SketchFusion: Learning Universal Sketch Features through Fusing Foundation ModelsSubhadeep Koley, Tapas Kumar Dutta, Aneeshan Sain, Pinaki Nath Chowdhury et al.CVPR 2025
- TVT: Three-Way Vision Transformer through Multi-Modal Hypersphere Learning for Zero-Shot Sketch-Based Image RetrievalJialin Tian, Xing Xu, Fumin Shen, Yang Yang et al.AAAI 2022 · 54 citations
- Photo Pre-Training, But for SketchKe Li, Kaiyue Pang, Yi-Zhe SongCVPR 2023
- Semantic-Aware Knowledge Preservation for Zero-Shot Sketch-Based Image RetrievalQing Liu, Lingxi Xie, Huiyu Wang, Alan L. YuilleICCV 2019 · 126 citations
- More Photos Are All You Need: Semi-Supervised Learning for Fine-Grained Sketch Based Image RetrievalAyan Kumar Bhunia, Pinaki Nath Chowdhury, Aneeshan Sain, Yongxin Yang et al.CVPR 2021
