Doppelgangers: Learning to Disambiguate Images of Similar Structures
Ruojin Cai, Joseph Tung, Qianqian Wang, Hadar Averbuch-Elor, Bharath Hariharan, Noah Snavely
Abstract
We consider the visual disambiguation task of determining whether a pair of visually similar images depict the same or distinct 3D surfaces (e.g., the same or opposite sides of a symmetric building). Illusory image matches, where two images observe distinct but visually similar 3D surfaces, can be challenging for humans to differentiate, and can also lead 3D reconstruction algorithms to produce erroneous results. We propose a learning-based approach to visual disambiguation, formulating it as a binary classification task on image pairs. To that end, we introduce a new dataset for this problem, Doppelgangers, which includes image pairs of similar structures with ground truth labels. We also design a network architecture that takes the spatial distribution of local keypoints and matches as input, allowing for better reasoning about both local and global cues. Our evaluation shows that our method can distinguish illusory matches in difficult cases, and can be integrated into SfM pipelines to produce correct, disambiguated 3D reconstructions. See our project page for our code, datasets, and more results: doppelgangers-3d.github.io.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext cfa503a6-4af6-48bd-a2f8-0248a24b5679Cited by top-tier papers14
- Dense Optical Tracking: Connecting the DotsGuillaume Le Moing, Jean Ponce, Cordelia SchmidCVPR 2024 · 24 citations
- Global Structure-from-Motion Meets Feedforward ReconstructionLinfei Pan, Johannes L. Schönberger, Marc PollefeysCVPR 2026 · 5 citations
- Same or Not? Enhancing Visual Perception in Vision-Language ModelsDamiano Marsili, Aditya Mehta, Ryan Y. Lin, Georgia GkioxariCVPR 2026 · 5 citations
- Long-Tail Internet Photo ReconstructionYuan Li, Yuanbo Xiangli, Hadar Averbuch-Elor, Noah Snavely et al.CVPR 2026 · 4 citations
- Leveraging Camera Triplets for Efficient and Accurate Structure-from-MotionLalit Manam, Venu Madhav GovinduCVPR 2024 · 4 citations
Builds on5
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- Towers of Babel: Combining Images, Language, and 3D Geometry for Learning Multimodal VisionXiaoshi Wu, Hadar Averbuch-Elor, Jin Sun, Noah SnavelyICCV 2021 · 26 citations
- SuperGlue: Learning Feature Matching With Graph Neural NetworksPaul-Edouard Sarlin, Daniel DeTone, Tomasz Malisiewicz, Andrew RabinovichCVPR 2020
- Google Landmarks Dataset v2 - A Large-Scale Benchmark for Instance-Level Recognition and RetrievalTobias Weyand, André Araújo, Bingyi Cao, Jack SimCVPR 2020
- Visual ChiralityZhiqiu Lin, Jin Sun, Abe Davis, Noah SnavelyCVPR 2020
Related papers
- ArchSym: Detecting 3D-Grounded Architectural Symmetries in the WildHanyu Chen, Ruojin Cai, Steve Marschner, Noah SnavelyCVPR 2026
- D2IM-Net: Learning Detail Disentangled Implicit Fields From Single ImagesManyi Li, Hao ZhangCVPR 2021
- Deep Unsupervised 3D SfM Face Reconstruction Based on Massive Landmark Bundle AdjustmentYuxing Wang, Yawen Lu, Zhihua Xie, Guoyu LuACM MM 2021 · 15 citations
- Structure from Duplicates: Neural Inverse Graphics from a Pile of ObjectsTianhang Cheng, Wei-Chiu Ma, Kaiyu Guan, Antonio Torralba et al.NeurIPS 2023 · 5 citations
- 3D Visual Illusion Depth EstimationChengtang Yao, Zhidan Liu, Jiaxi Zeng, Lidong Yu et al.NeurIPS 2025 · 3 citations
