Scene-Aware Feature Matching
Xiaoyong Lu, Yaping Yan, Tong Wei, Songlin Du
Abstract
Current feature matching methods focus on point-level matching, pursuing better representation learning of individual features, but lacking further understanding of the scene. This results in significant performance degradation when handling challenging scenes such as scenes with large viewpoint and illumination changes. To tackle this problem, we propose a novel model named SAM, which applies attentional grouping to guide Scene-Aware feature Matching. SAM handles multi-level features, i.e., image tokens and group tokens, with attention layers, and groups the image tokens with the proposed token grouping module. Our model can be trained by ground-truth matches only and produce reasonable grouping results. With the sense-aware grouping guidance, SAM is not only more accurate and robust but also more interpretable than conventional feature matching models. Sufficient experiments on various applications, including homography estimation, pose estimation, and image matching, demonstrate that our model achieves state-of-the-art performance.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on7
- GroupViT: Semantic Segmentation Emerges from Text SupervisionJiarui Xu, Shalini De Mello, Sifei Liu, Wonmin Byeon et al.CVPR 2022 · 398 citations
- Learning Two-View Correspondences and Geometry Using Order-Aware NetworkJiahui Zhang, Dawei Sun, Zixin Luo, Anbang Yao et al.ICCV 2019 · 362 citations
- Dual-Resolution Correspondence NetworksXinghui Li, Kai Han, Shuda Li, Victor PrisacariuNeurIPS 2020 · 207 citations
- Guide Local Feature Matching by Overlap EstimationYing Chen, Dihe Huang, Shang Xu, Jianlin Liu et al.AAAI 2022 · 35 citations
- ParaFormer: Parallel Attention Transformer for Efficient Feature MatchingXiaoyong Lu, Yaping Yan, Bin Kang, Songlin DuAAAI 2023 · 21 citations
Related papers
- MESA: Matching Everything by Segmenting AnythingYesheng Zhang, Xu ZhaoCVPR 2024 · 19 citations
- Learning Compact 3D Representations from Feed-Forward Novel View SynthesisHonggyu An, Jaewoo Jung, Mungyeom Kim, Chaehyun Kim et al.CVPR 2026
- End2End Multi-View Feature Matching with Differentiable Pose OptimizationBarbara Roessle, Matthias NießnerICCV 2023 · 34 citations
- TopicFM: Robust and Interpretable Topic-Assisted Feature MatchingKhang Truong Giang, Soohwan Song, Sungho JoAAAI 2023 · 73 citations
- Improving Transformer-based Image Matching by Cascaded Capturing Spatially Informative KeypointsChenjie Cao, Yanwei FuICCV 2023 · 23 citations
