Co-Attention for Conditioned Image Matching
Olivia Wiles, Sébastien Ehrhardt, Andrew Zisserman
Abstract
We propose a new approach to determine correspondences between image pairs in the wild under large changes in illumination, viewpoint, context, and material. While other approaches find correspondences between pairs of images by treating the images independently, we instead condition on both images to implicitly take account of the differences between them. To achieve this, we introduce (i) a spatial attention mechanism (a co-attention module, CoAM) for conditioning the learned features on both images, and (ii) a distinctiveness score used to choose the best matches at test time. CoAM can be added to standard architectures and trained using self-supervision or supervised data, and achieves a significant performance improvement under hard conditions, e.g. large viewpoint changes. We demonstrate that models using CoAM achieve state of the art or competitive results on a wide range of tasks: local matching, camera localization, 3D reconstruction, and image stylization.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers10
- Stacked Hybrid-Attention and Group Collaborative Learning for Unbiased Scene Graph GenerationXingning Dong, Tian Gan, Xuemeng Song, Jianlong Wu et al.CVPR 2022 · 116 citations
- Decoupling Makes Weakly Supervised Local Feature BetterKunhong Li, Longguang Wang, Li Liu, Qing Ran et al.CVPR 2022 · 58 citations
- Guide Local Feature Matching by Overlap EstimationYing Chen, Dihe Huang, Shang Xu, Jianlin Liu et al.AAAI 2022 · 35 citations
- Continuous Parametric Optical FlowJianqin Luo, Zhexiong Wan, Yuxin Mao, Bo Li et al.NeurIPS 2023 · 6 citations
- TAB: Transformer Attention Bottlenecks Enable User Intervention and Debugging in Vision-Language ModelsPooyan Rahmanzadehgervi, Hung Huy Nguyen, Rosanne Liu, Long Mai et al.ICCV 2025 · 3 citations
Builds on3
- SuperGlue: Learning Feature Matching With Graph Neural NetworksPaul-Edouard Sarlin, Daniel DeTone, Tomasz Malisiewicz, Andrew RabinovichCVPR 2020
- MAST: A Memory-Augmented Self-Supervised TrackerZihang Lai, Erika Lu, Weidi XieCVPR 2020
- Correspondence Networks With Adaptive Neighbourhood ConsensusShuda Li, Kai Han, Theo W. Costain, Henry Howard-Jenkins et al.CVPR 2020
Related papers
- TransforMatcher: Match-to-Match Attention for Semantic CorrespondenceSeungwook Kim, Juhong Min, Minsu ChoCVPR 2022 · 26 citations
- COTR: Correspondence Transformer for Matching Across ImagesWei Jiang, Eduard Trulls, Jan Hosang, Andrea Tagliasacchi et al.ICCV 2021 · 318 citations
- Geometry-Free View Synthesis: Transformers and no 3D PriorsRobin Rombach, Patrick Esser, Björn OmmerICCV 2021 · 115 citations
- Self-Supervised Spatial Correspondence Across ModalitiesAyush Shrivastava, Andrew OwensCVPR 2025
- LoFTR: Detector-Free Local Feature Matching With TransformersJiaming Sun, Zehong Shen, Yuang Wang, Hujun Bao et al.CVPR 2021
