Learning Dense Correspondences between Photos and Sketches
Xuanchen Lu, Xiaolong Wang, Judith E. Fan
摘要
Humans effortlessly grasp the connection between sketches and real-world objects, even when these sketches are far from realistic. Moreover, human sketch understanding goes beyond categorization -- critically, it also entails understanding how individual elements within a sketch correspond to parts of the physical world it represents. What are the computational ingredients needed to support this ability? Towards answering this question, we make two contributions: first, we introduce a new sketch-photo correspondence benchmark, , containing 150K annotations of 6250 sketch-photo pairs across 125 object categories, augmenting the existing Sketchy dataset with fine-grained correspondence metadata. Second, we propose a self-supervised method for learning dense correspondences between sketch-photo pairs, building upon recent advances in correspondence learning for pairs of photos. Our model uses a spatial transformer network to estimate the warp flow between latent representations of a sketch and photo extracted by a contrastive learning-based ConvNet backbone. We found that this approach outperformed several strong baselines and produced predictions that were quantitatively consistent with other warp-based methods. However, our benchmark also revealed systematic differences between predictions of the suite of models we tested and those of humans. Taken together, our work suggests a promising path towards developing artificial systems that achieve more human-like understanding of visual images at different levels of abstraction. Project page: https://photo-sketch-correspondence.github.io
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Latent Trajectory Learning for Limited Timestamps under Distribution Shift over TimeQiuhao Zeng, Changjian Shui, Long-Kai Huang, Peng Liu 等ICLR 2024 · 被引用 15 次
- Sparse-to-dense Multimodal Image Registration via Multi-Task LearningKaining Zhang, Jiayi MaICML 2024 · 被引用 6 次
- Fuse2Match: Training-Free Fusion of Flow, Diffusion, and Contrastive Models for Zero-Shot Semantic MatchingJing Zuo, Jiaqi Wang, Yonggang Qi, Yi-Zhe SongNeurIPS 2025
- Doodle Your Keypoints: Sketch-Based Few-Shot Keypoint DetectionSubhajit Maity, Ayan Kumar Bhunia, Subhadeep Koley, Pinaki Nath Chowdhury 等ICCV 2025
- When Lines Meet Textures: Spatial-Frequency Aligned Diffusion Features for Cross-Sparsity CorrespondenceMingrui Zhu, Fengzhi Wang, Xin Wei, Jue Wang 等CVPR 2026
它引用的顶会 Paper24
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec 等NeurIPS 2020 · 被引用 9,171 次
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen 等ICML 2021 · 被引用 5,401 次
- Big Self-Supervised Models are Strong Semi-Supervised LearnersTing Chen, Simon Kornblith, Kevin Swersky, Mohammad Norouzi 等NeurIPS 2020 · 被引用 2,611 次
相关 Paper
- Warp Consistency for Unsupervised Learning of Dense CorrespondencesPrune Truong, Martin Danelljan, Fisher Yu, Luc Van GoolICCV 2021 · 被引用 60 次
- Learning Transformation-Predictive Representations for Detection and Description of Local FeaturesZihao Wang, Chunxu Wu, Yifei Yang, Zhen LiCVPR 2023
- Semi-Supervised Learning of Semantic Correspondence with Pseudo-LabelsJiwon Kim, Kwangrok Ryoo, Junyoung Seo, Gyuseong Lee 等CVPR 2022 · 被引用 23 次
- Jamais Vu: Exposing the Generalization Gap in Supervised Semantic CorrespondenceOctave Mariotti, Zhipeng Du, Yash Bhalgat, Oisin Mac Aodha 等NeurIPS 2025 · 被引用 8 次
- Improving Semantic Correspondence with Viewpoint-Guided Spherical MapsOctave Mariotti, Oisin Mac Aodha, Hakan BilenCVPR 2024 · 被引用 12 次
