Text-to-Image Diffusion Models are Great Sketch-Photo Matchmakers
Subhadeep Koley, Ayan Kumar Bhunia, Aneeshan Sain, Pinaki Nath Chowdhury, Tao Xiang, Yi-Zhe Song
摘要
Figure 1. Sketch-based image retrieval frameworks [12, 102, 103] usually employ ImageNet pre-trained CNNs [3, 15, 103], JFT-trained vision transformers (ViT) [67, 78], or visual encoders of vision-language models like CLIP [77] as backbone feature extractors. Rich knowledge from large-scale pre-training offers a good initialisation, which when further fine-tuned on sketch-photo datasets, performs way better than training from random initialisation [62]. While one can extract features either by discarding the classification head for ImageNet pre-trained models, auxiliary task head for self-supervised models, or by using CLIP's visual encoder, text-to-image diffusion models (e.g., stable diffusion) lack any specific feature embedding space. However, we find that its intermediate representations implicitly hold robust cross-modal features at multiple granularities. Unlike prior SBIR backbones, pre-trained with discriminative tasks, we propose to leverage denoising diffusion models pre-trained with text-to-image generative tasks to bridge the sketch-photo domain gap. Being a text-to-image generation model trained on a large corpus of text-image pairs (LAION [82, 83]), it holds both semantic and shape prior [89]. PCA representation [89] (right) of intermediate UNet features (sketch/photo) from different upsampling blocks (details in Sec. 5) depict that they share significant semantic similarity (denoted by similar colours).
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper12
- VQ-SGen: A Vector Quantized Stroke Representation for Creative Sketch GenerationJiawei Wang, Zhiming Cui, Changjian LiICCV 2025 · 被引用 3 次
- Leveraging Prior Knowledge of Diffusion Model for Person SearchGiyeol Kim, Sooyoung Yang, Jihyong Oh, Myungjoo Kang 等ICCV 2025 · 被引用 2 次
- Modeling the Visual Ambiguity of Human SketchesYang Zhou, Ping Ni, Jin Wang, Senyun Jia 等CVPR 2026
- You'll Never Walk Alone: A Sketch and Text Duet for Fine-Grained Image RetrievalSubhadeep Koley, Ayan Kumar Bhunia, Aneeshan Sain, Pinaki Nath Chowdhury 等CVPR 2024
- Intermediate Domain Alignment and Morphology Analogy for Patent-Product Image RetrievalHaifan Gong, Xuanye Zhang, Ruifei Zhang, Yun Su 等NeurIPS 2025
它引用的顶会 Paper53
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 被引用 13,211 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
相关 Paper
- SketchFusion: Learning Universal Sketch Features through Fusing Foundation ModelsSubhadeep Koley, Tapas Kumar Dutta, Aneeshan Sain, Pinaki Nath Chowdhury 等CVPR 2025
- TVT: Three-Way Vision Transformer through Multi-Modal Hypersphere Learning for Zero-Shot Sketch-Based Image RetrievalJialin Tian, Xing Xu, Fumin Shen, Yang Yang 等AAAI 2022 · 被引用 54 次
- Photo Pre-Training, But for SketchKe Li, Kaiyue Pang, Yi-Zhe SongCVPR 2023
- Semantic-Aware Knowledge Preservation for Zero-Shot Sketch-Based Image RetrievalQing Liu, Lingxi Xie, Huiyu Wang, Alan L. YuilleICCV 2019 · 被引用 126 次
- More Photos Are All You Need: Semi-Supervised Learning for Fine-Grained Sketch Based Image RetrievalAyan Kumar Bhunia, Pinaki Nath Chowdhury, Aneeshan Sain, Yongxin Yang 等CVPR 2021
