TransforMatcher: Match-to-Match Attention for Semantic Correspondence
Seungwook Kim, Juhong Min, Minsu Cho
Abstract
Establishing correspondences between images remains a challenging task, especially under large appearance changes due to different viewpoints or intra-class variations. In this work, we introduce a strong semantic image matching learner, dubbed TransforMatcher, which builds on the success of transformer networks in vision domains. Un-like existing convolution- or attention-based schemes for correspondence, TransforMatcher performs global match-to-match attention for precise match localization and dynamic refinement. To handle a large number of matches in a dense correlation map, we develop a light-weight attention architecture to consider the global match-to-match interactions. We also propose to utilize a multi-channel correlation map for refinement, treating the multi-level scores as features instead of a single score to fully exploit the richer layer-wise semantics. In experiments, TransforMatcher sets a new state of the art on SPair-71k while performing on par with existing SOTA methods on the PF-PASCAL dataset.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers18
- Emergent Correspondence from Image DiffusionLuming Tang, Menglin Jia, Qianqian Wang, Cheng Perng Phoo et al.NeurIPS 2023 · 555 citations
- Unleashing Diffusion Transformers for Visual Correspondence by Modulating Massive ActivationsChaofan Gan, Yuanpeng Tu, Xi Chen, Tieyuan Chen et al.NeurIPS 2025 · 22 citations
- Unifying Feature and Cost Aggregation with Transformers for Semantic and Visual CorrespondenceSunghwan Hong, Seokju Cho, Seungryong Kim, Stephen LinICLR 2024 · 16 citations
- 3D Geometric Shape Assembly via Efficient Point Cloud MatchingNahyuk Lee, Juhong Min, Junha Lee, Seungwook Kim et al.ICML 2024 · 12 citations
- Jamais Vu: Exposing the Generalization Gap in Supervised Semantic CorrespondenceOctave Mariotti, Zhipeng Du, Yash Bhalgat, Oisin Mac Aodha et al.NeurIPS 2025 · 8 citations
Builds on22
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
- Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without ConvolutionsWenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan et al.ICCV 2021 · 4,909 citations
- Reformer: The Efficient TransformerNikita Kitaev, Lukasz Kaiser, Anselm LevskayaICLR 2020 · 2,878 citations
Related papers
- Correspondence Transformers with Asymmetric Feature Learning and Matching Flow Super-ResolutionYixuan Sun, Dongyang Zhao, Zhangyue Yin, Yiwen Huang et al.CVPR 2023
- LoFTR: Detector-Free Local Feature Matching With TransformersJiaming Sun, Zehong Shen, Yuang Wang, Hujun Bao et al.CVPR 2021
- COTR: Correspondence Transformer for Matching Across ImagesWei Jiang, Eduard Trulls, Jan Hosang, Andrea Tagliasacchi et al.ICCV 2021 · 318 citations
- Dynamic Context Correspondence Network for Semantic AlignmentShuaiyi Huang, Qiuyue Wang, Songyang Zhang, Shipeng Yan et al.ICCV 2019 · 97 citations
- Multi-scale Matching Networks for Semantic CorrespondenceDongyang Zhao, Ziyang Song, Zhenghao Ji, Gangming Zhao et al.ICCV 2021 · 56 citations
