SketchFusion: Learning Universal Sketch Features through Fusing Foundation Models
Subhadeep Koley, Tapas Kumar Dutta, Aneeshan Sain, Pinaki Nath Chowdhury, Ayan Kumar Bhunia, Yi-Zhe Song
摘要
Figure 1. (Left): Apart from high-resolution image generation, text-to-image diffusion models (e.g., Stable diffusion (SD) [71]) with their innate object understanding capability [84, 107], have shown remarkable performance across a wide range of image-based vision tasks (e.g., segmentation [94], depth estimation [110], etc. ). However, upon analysing the PCA representation of SD's intermediate UNet features, we observe that it struggles to achieve similar results when working with freehand abstract sketches (detail in Sec. 4). Unlike pixel-perfect photos, highly abstract freehand sketches are sparse and lack detailed textures and colours [26] , making it harder for the SD model to extract meaningful features. Furthermore, investigating the SD denoising process in the frequency domain (via Fourier Transform), we observe the predominance of high-frequency (HF) components, rather than their low-frequency (LF) counterpart -crucial for capturing comprehensive semantic context. To mitigate this inherent bias within SD, we reinforce the diffusion process with another pretrained model (i.e., CLIP [67]) whose bias is complementary (i.e., focuses on LF) to SD. Consequently, the proposed extractor can extract semantically meaningful and accurate features from both sketches and photos, that encapsulate a broader frequency spectrum (i.e., HF and LF). (Right:) Testing the proposed method with different sketch-based discriminative and dense prediction tasks (requiring knowledge of both sketch and image), we find a marked improvement over baseline SD+CLIP hybrid feature extractor.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Modeling the Visual Ambiguity of Human SketchesYang Zhou, Ping Ni, Jin Wang, Senyun Jia 等CVPR 2026
- O3SLM: Open Weight, Open Data, and Open Vocabulary Sketch-Language ModelRishi Gupta, Mukilan Karuppasamy, Shyam Marjit, Aditay Tripathi 等AAAI 2026
- When Lines Meet Textures: Spatial-Frequency Aligned Diffusion Features for Cross-Sparsity CorrespondenceMingrui Zhu, Fengzhi Wang, Xin Wei, Jue Wang 等CVPR 2026
它引用的顶会 Paper64
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 被引用 13,211 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
相关 Paper
- Text-to-Image Diffusion Models are Great Sketch-Photo MatchmakersSubhadeep Koley, Ayan Kumar Bhunia, Aneeshan Sain, Pinaki Nath Chowdhury 等CVPR 2024
- Text-to-Image Diffusion Models are Zero Shot ClassifiersKevin Clark, Priyank JainiNeurIPS 2023 · 被引用 192 次
- One-Shot Reference-based Structure-Aware Image to Sketch SynthesisRui Yang, Honghong Yang, Li Zhao, Qin Lei 等AAAI 2025 · 被引用 2 次
- StableRep: Synthetic Images from Text-to-Image Models Make Strong Visual Representation LearnersYonglong Tian, Lijie Fan, Phillip Isola, Huiwen Chang 等NeurIPS 2023 · 被引用 251 次
- Enhancing Compositional Text-to-Image Generation with Reliable Random SeedsShuangqi Li, Hieu Le, Jingyi Xu, Mathieu SalzmannICLR 2025
