Jamais Vu: Exposing the Generalization Gap in Supervised Semantic Correspondence
Octave Mariotti, Zhipeng Du, Yash Bhalgat, Oisin Mac Aodha, Hakan Bilen
Abstract
The goal of semantic correspondence (SC) estimation is to establish semantically meaningful matches across different instances of an object category. In this work, we illustrate how recent supervised SC methods generalize poorly beyond the annotated keypoints seen during training, thus effectively acting as keypoint detectors.
To address this, we propose a new approach for learning dense correspondences by lifting 2D keypoints into a canonical 3D space using monocular depth estimation.
Our method constructs a continuous canonical manifold that captures object geometry without requiring explicit 3D supervision or camera annotations. Additionally, we introduce SPair-U, an extension of SPair-71k with novel keypoint annotations, to better assess generalization. Experiments not only demonstrate that our model significantly outperforms supervised baselines on unseen keypoints, highlighting its effectiveness in learning robust correspondences, but that unsupervised baselines outperform supervised counterparts when evaluated across different datasets.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 80a02ca5-a98a-405b-a778-52e6384757b0Cited by top-tier papers1
Ask how each one uses itBuilds on30
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- Depth Anything V2Lihe Yang, Bingyi Kang, Zilong Huang, Zhen Zhao et al.NeurIPS 2024 · 2,305 citations
- Emergent Correspondence from Image DiffusionLuming Tang, Menglin Jia, Qianqian Wang, Cheng Perng Phoo et al.NeurIPS 2023 · 555 citations
Related papers
- SemAlign3D: Semantic Correspondence between RGB-Images through Aligning 3D Object-Class RepresentationsKrispin Wandel, Hesheng WangCVPR 2025
- Improving Semantic Correspondence with Viewpoint-Guided Spherical MapsOctave Mariotti, Oisin Mac Aodha, Hakan BilenCVPR 2024 · 12 citations
- HumanGPS: Geodesic PreServing Feature for Dense Human CorrespondencesFeitong Tan, Danhang Tang, Mingsong Dou, Kaiwen Guo et al.CVPR 2021
- MonoRUn: Monocular 3D Object Detection by Reconstruction and Uncertainty PropagationHansheng Chen, Yuyao Huang, Wei Tian, Zhong Gao et al.CVPR 2021
- Seeing Behind Objects for 3D Multi-Object Tracking in RGB-D SequencesNorman Müller, Yu-Shiang Wong, Niloy J. Mitra, Angela Dai et al.CVPR 2021
