SketchFusion: Learning Universal Sketch Features through Fusing Foundation Models
Subhadeep Koley, Tapas Kumar Dutta, Aneeshan Sain, Pinaki Nath Chowdhury, Ayan Kumar Bhunia, Yi-Zhe Song
Abstract
Figure 1. (Left): Apart from high-resolution image generation, text-to-image diffusion models (e.g., Stable diffusion (SD) [71]) with their innate object understanding capability [84, 107], have shown remarkable performance across a wide range of image-based vision tasks (e.g., segmentation [94], depth estimation [110], etc. ). However, upon analysing the PCA representation of SD's intermediate UNet features, we observe that it struggles to achieve similar results when working with freehand abstract sketches (detail in Sec. 4). Unlike pixel-perfect photos, highly abstract freehand sketches are sparse and lack detailed textures and colours [26] , making it harder for the SD model to extract meaningful features. Furthermore, investigating the SD denoising process in the frequency domain (via Fourier Transform), we observe the predominance of high-frequency (HF) components, rather than their low-frequency (LF) counterpart -crucial for capturing comprehensive semantic context. To mitigate this inherent bias within SD, we reinforce the diffusion process with another pretrained model (i.e., CLIP [67]) whose bias is complementary (i.e., focuses on LF) to SD. Consequently, the proposed extractor can extract semantically meaningful and accurate features from both sketches and photos, that encapsulate a broader frequency spectrum (i.e., HF and LF). (Right:) Testing the proposed method with different sketch-based discriminative and dense prediction tasks (requiring knowledge of both sketch and image), we find a marked improvement over baseline SD+CLIP hybrid feature extractor.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- Modeling the Visual Ambiguity of Human SketchesYang Zhou, Ping Ni, Jin Wang, Senyun Jia et al.CVPR 2026
- O3SLM: Open Weight, Open Data, and Open Vocabulary Sketch-Language ModelRishi Gupta, Mukilan Karuppasamy, Shyam Marjit, Aditay Tripathi et al.AAAI 2026
- When Lines Meet Textures: Spatial-Frequency Aligned Diffusion Features for Cross-Sparsity CorrespondenceMingrui Zhu, Fengzhi Wang, Xin Wei, Jue Wang et al.CVPR 2026
Builds on64
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
Related papers
- Text-to-Image Diffusion Models are Great Sketch-Photo MatchmakersSubhadeep Koley, Ayan Kumar Bhunia, Aneeshan Sain, Pinaki Nath Chowdhury et al.CVPR 2024
- Text-to-Image Diffusion Models are Zero Shot ClassifiersKevin Clark, Priyank JainiNeurIPS 2023 · 192 citations
- One-Shot Reference-based Structure-Aware Image to Sketch SynthesisRui Yang, Honghong Yang, Li Zhao, Qin Lei et al.AAAI 2025 · 2 citations
- StableRep: Synthetic Images from Text-to-Image Models Make Strong Visual Representation LearnersYonglong Tian, Lijie Fan, Phillip Isola, Huiwen Chang et al.NeurIPS 2023 · 251 citations
- Enhancing Compositional Text-to-Image Generation with Reliable Random SeedsShuangqi Li, Hieu Le, Jingyi Xu, Mathieu SalzmannICLR 2025
