When Lines Meet Textures: Spatial-Frequency Aligned Diffusion Features for Cross-Sparsity Correspondence
Mingrui Zhu, Fengzhi Wang, Xin Wei, Jue Wang, Nannan Wang, Xinbo Gao
摘要
Establishing accurate correspondence between sparse line representations and rich textured imagery remains a formidable challenge. While diffusion features excel in semantic correspondence, they struggle to bridge the fundamental gap between abstract sketches and texture-rich photographs. We identify two critical disparities: spatial domain misalignment from structural abstraction differences, and frequency domain inconsistencies from texture density variations. Based on this analysis, we propose SFA-DIFT, a novel approach that learns spatial-frequency aligned diffusion features for robust cross-modal correspondence. Unlike previous methods focusing solely on spatial alignment, our key innovation performs dual-domain alignment by learning unified clean diffusion features while strategically aggregating low-frequency components in the frequency domain. This comprehensive spatial-frequency alignment enables equitable understanding between sparse abstractions and rich textures. To validate our approach, we extend the existing sketch-photo correspondence dataset (PSC6K) by generating multi-style textured imagery, creating MS-PSC6K, a comprehensive correspondence benchmark. Extensive experiments demonstrate that SFA-DIFT achieves state-of-the-art performance, delivering substantial improvements with an average of 0.87% on PCK@1, 2.20% on PCK@5, and 0.95% on PCK@10 over previous best methods, validating the effectiveness and robustness of our dual-domain alignment approach. Codes are available on https://github.com/Mofr77/SFA-DIFT.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper18
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Emergent Correspondence from Image DiffusionLuming Tang, Menglin Jia, Qianqian Wang, Cheng Perng Phoo 等NeurIPS 2023 · 被引用 555 次
- A Tale of Two Features: Stable Diffusion Complements DINO for Zero-Shot Semantic CorrespondenceJunyi Zhang, Charles Herrmann, Junhwa Hur, Luisa Polania Cabrera 等NeurIPS 2023 · 被引用 371 次
- CLIPasso: semantically-aware object sketchingYael Vinker, Ehsan Pajouheshgar, Jessica Y. Bo, Roman Christian Bachmann 等SIGGRAPH 2022 · 被引用 219 次
相关 Paper
- Learning Dense Correspondences between Photos and SketchesXuanchen Lu, Xiaolong Wang, Judith E. FanICML 2023 · 被引用 6 次
- SpecXNet: A Dual-Domain Convolutional Network for Robust Deepfake DetectionInzamamul Alam, Md Tanvir Islam, Simon S. WooACM MM 2025 · 被引用 6 次
- Diffusion Hyperfeatures: Searching Through Time and Space for Semantic CorrespondenceGrace Luo, Lisa Dunlap, Dong Huk Park, Aleksander Holynski 等NeurIPS 2023 · 被引用 261 次
- Unsupervised Semantic Correspondence Using Stable DiffusionEric Hedlin, Gopal Sharma, Shweta Mahajan, Hossam Isack 等NeurIPS 2023 · 被引用 152 次
- Towards Transformer-Based Aligned Generation with Self-Coherence GuidanceShulei Wang, Wang Lin, Hai Huang, Hanting Wang 等CVPR 2025
