Improving Transformer-based Image Matching by Cascaded Capturing Spatially Informative Keypoints
Chenjie Cao, Yanwei Fu
Abstract
Learning robust local image feature matching is a fundamental low-level vision task, which has been widely explored in the past few years. Recently, detector-free local feature matchers based on transformers have shown promising results, which largely outperform pure Convolutional Neural Network (CNN) based ones. But correlations produced by transformer-based methods are spatially limited to the center of source views' coarse patches, because of the costly attention learning. In this work, we rethink this issue and find that such matching formulation degrades pose estimation, especially for low-resolution images. So we propose a transformer-based cascade matching model -Cascade feature Matching TRansformer (CasMTR) § , to efficiently learn dense feature correlations, which allows us to choose more reliable matching pairs for the relative pose estimation. Instead of re-training a new detector, we use a simple yet effective Non-Maximum Suppression (NMS) post-process to filter keypoints through the confidence map, and largely improve the matching precision. CasMTR achieves state-of-the-art performance in indoor and outdoor pose estimation as well as visual localization. Moreover, thorough ablations show the efficacy of the proposed components and techniques.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers12
- MVSFormer++: Revealing the Devil in Transformer's Details for Multi-View StereoChenjie Cao, Xinlin Ren, Yanwei FuICLR 2024 · 68 citations
- MESA: Matching Everything by Segmenting AnythingYesheng Zhang, Xu ZhaoCVPR 2024 · 19 citations
- Semantic-aware Representation Learning for Homography EstimationYuhan Liu, Qianxin Huang, Siqi Hui, Jingwen Fu et al.ACM MM 2024 · 6 citations
- Multi-scale Consistency for Robust 3D Registration via Hierarchical Sinkhorn TreeChengwei Ren, Yifan Feng, Weixiang Zhang, Xiao-Ping (Steven) Zhang et al.NeurIPS 2024 · 6 citations
- PRISM: PRogressive dependency maxImization for Scale-invariant image MatchingXudong Cai, Yongcai Wang, Lun Luo, Minhang Wang et al.ACM MM 2024 · 3 citations
Builds on21
- Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without ConvolutionsWenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan et al.ICCV 2021 · 4,909 citations
- Twins: Revisiting the Design of Spatial Attention in Vision TransformersXiangxiang Chu, Zhi Tian, Yuqing Wang, Bo Zhang et al.NeurIPS 2021 · 1,388 citations
- DISK: Learning local features with policy gradientMichal J. Tyszkiewicz, Pascal Fua, Eduard TrullsNeurIPS 2020 · 652 citations
- LST: Ladder Side-Tuning for Parameter and Memory Efficient Transfer LearningYi-Lin Sung, Jaemin Cho, Mohit BansalNeurIPS 2022 · 347 citations
- COTR: Correspondence Transformer for Matching Across ImagesWei Jiang, Eduard Trulls, Jan Hosang, Andrea Tagliasacchi et al.ICCV 2021 · 318 citations
Related papers
- LoFTR: Detector-Free Local Feature Matching With TransformersJiaming Sun, Zehong Shen, Yuang Wang, Hujun Bao et al.CVPR 2021
- Towards Accurate Facial Landmark Detection via Cascaded TransformersHui Li, Zidong Guo, Seon-Min Rhee, Seungju Han et al.CVPR 2022 · 45 citations
- Collaborative Feature Matching with Progressive Correspondence LearningXin Liu, Yanbing Han, Rong Qin, Bing Wang et al.AAAI 2026
- Correspondence Transformers with Asymmetric Feature Learning and Matching Flow Super-ResolutionYixuan Sun, Dongyang Zhao, Zhangyue Yin, Yiwen Huang et al.CVPR 2023
- Geometrized Transformer for Self-Supervised Homography EstimationJiazhen Liu, Xirong LiICCV 2023 · 27 citations
