Improving Transformer-based Image Matching by Cascaded Capturing Spatially Informative Keypoints
Chenjie Cao, Yanwei Fu
摘要
Learning robust local image feature matching is a fundamental low-level vision task, which has been widely explored in the past few years. Recently, detector-free local feature matchers based on transformers have shown promising results, which largely outperform pure Convolutional Neural Network (CNN) based ones. But correlations produced by transformer-based methods are spatially limited to the center of source views' coarse patches, because of the costly attention learning. In this work, we rethink this issue and find that such matching formulation degrades pose estimation, especially for low-resolution images. So we propose a transformer-based cascade matching model -Cascade feature Matching TRansformer (CasMTR) § , to efficiently learn dense feature correlations, which allows us to choose more reliable matching pairs for the relative pose estimation. Instead of re-training a new detector, we use a simple yet effective Non-Maximum Suppression (NMS) post-process to filter keypoints through the confidence map, and largely improve the matching precision. CasMTR achieves state-of-the-art performance in indoor and outdoor pose estimation as well as visual localization. Moreover, thorough ablations show the efficacy of the proposed components and techniques.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper12
- MVSFormer++: Revealing the Devil in Transformer's Details for Multi-View StereoChenjie Cao, Xinlin Ren, Yanwei FuICLR 2024 · 被引用 68 次
- MESA: Matching Everything by Segmenting AnythingYesheng Zhang, Xu ZhaoCVPR 2024 · 被引用 19 次
- Semantic-aware Representation Learning for Homography EstimationYuhan Liu, Qianxin Huang, Siqi Hui, Jingwen Fu 等ACM MM 2024 · 被引用 6 次
- Multi-scale Consistency for Robust 3D Registration via Hierarchical Sinkhorn TreeChengwei Ren, Yifan Feng, Weixiang Zhang, Xiao-Ping (Steven) Zhang 等NeurIPS 2024 · 被引用 6 次
- PRISM: PRogressive dependency maxImization for Scale-invariant image MatchingXudong Cai, Yongcai Wang, Lun Luo, Minhang Wang 等ACM MM 2024 · 被引用 3 次
它引用的顶会 Paper21
- Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without ConvolutionsWenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan 等ICCV 2021 · 被引用 4,909 次
- Twins: Revisiting the Design of Spatial Attention in Vision TransformersXiangxiang Chu, Zhi Tian, Yuqing Wang, Bo Zhang 等NeurIPS 2021 · 被引用 1,388 次
- DISK: Learning local features with policy gradientMichal J. Tyszkiewicz, Pascal Fua, Eduard TrullsNeurIPS 2020 · 被引用 652 次
- LST: Ladder Side-Tuning for Parameter and Memory Efficient Transfer LearningYi-Lin Sung, Jaemin Cho, Mohit BansalNeurIPS 2022 · 被引用 347 次
- COTR: Correspondence Transformer for Matching Across ImagesWei Jiang, Eduard Trulls, Jan Hosang, Andrea Tagliasacchi 等ICCV 2021 · 被引用 318 次
相关 Paper
- LoFTR: Detector-Free Local Feature Matching With TransformersJiaming Sun, Zehong Shen, Yuang Wang, Hujun Bao 等CVPR 2021
- Towards Accurate Facial Landmark Detection via Cascaded TransformersHui Li, Zidong Guo, Seon-Min Rhee, Seungju Han 等CVPR 2022 · 被引用 45 次
- Collaborative Feature Matching with Progressive Correspondence LearningXin Liu, Yanbing Han, Rong Qin, Bing Wang 等AAAI 2026
- Correspondence Transformers with Asymmetric Feature Learning and Matching Flow Super-ResolutionYixuan Sun, Dongyang Zhao, Zhangyue Yin, Yiwen Huang 等CVPR 2023
- Geometrized Transformer for Self-Supervised Homography EstimationJiazhen Liu, Xirong LiICCV 2023 · 被引用 27 次
