Lune

CVPR2024顶会

Text-to-Image Diffusion Models are Great Sketch-Photo Matchmakers

Subhadeep Koley, Ayan Kumar Bhunia, Aneeshan Sain, Pinaki Nath Chowdhury, Tao Xiang, Yi-Zhe Song

2024年份
12顶会引用

摘要

Figure 1. Sketch-based image retrieval frameworks [12, 102, 103] usually employ ImageNet pre-trained CNNs [3, 15, 103], JFT-trained vision transformers (ViT) [67, 78], or visual encoders of vision-language models like CLIP [77] as backbone feature extractors. Rich knowledge from large-scale pre-training offers a good initialisation, which when further fine-tuned on sketch-photo datasets, performs way better than training from random initialisation [62]. While one can extract features either by discarding the classification head for ImageNet pre-trained models, auxiliary task head for self-supervised models, or by using CLIP's visual encoder, text-to-image diffusion models (e.g., stable diffusion) lack any specific feature embedding space. However, we find that its intermediate representations implicitly hold robust cross-modal features at multiple granularities. Unlike prior SBIR backbones, pre-trained with discriminative tasks, we propose to leverage denoising diffusion models pre-trained with text-to-image generative tasks to bridge the sketch-photo domain gap. Being a text-to-image generation model trained on a large corpus of text-image pairs (LAION [82, 83]), it holds both semantic and shape prior [89]. PCA representation [89] (right) of intermediate UNet features (sketch/photo) from different upsampling blocks (details in Sec. 5) depict that they share significant semantic similarity (denoted by similar colours).

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

引用它的顶会 Paper12

问问它们各自怎么用它

它引用的顶会 Paper53

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖