ParaFormer: Parallel Attention Transformer for Efficient Feature Matching
Xiaoyong Lu, Yaping Yan, Bin Kang, Songlin Du
Abstract
Heavy computation is a bottleneck limiting deep-learning-based feature matching algorithms to be applied in many real-time applications. However, existing lightweight networks optimized for Euclidean data cannot address classical feature matching tasks, since sparse keypoint based descriptors are expected to be matched. This paper tackles this problem and proposes two concepts: 1) a novel parallel attention model entitled ParaFormer and 2) a graph based U-Net architecture with attentional pooling. First, ParaFormer fuses features and keypoint positions through the concept of amplitude and phase, and integrates self- and cross-attention in a parallel manner which achieves a win-win performance in terms of accuracy and efficiency. Second, with U-Net architecture and proposed attentional pooling, the ParaFormer-U variant significantly reduces computational complexity, and minimize performance loss caused by downsampling. Sufficient experiments on various applications, including homography estimation, pose estimation, and image matching, demonstrate that ParaFormer achieves state-of-the-art performance while maintaining high efficiency. The efficient ParaFormer-U variant achieves comparable performance with less than 50% FLOPs of the existing attention-based models.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9b733d6a-05de-491e-8a95-458b2ec6898fCited by top-tier papers3
- Scene-Aware Feature MatchingXiaoyong Lu, Yaping Yan, Tong Wei, Songlin DuICCV 2023 · 9 citations
- Matching While Perceiving: Enhance Image Feature Matching with Applicable Semantic AmalgamationShihua Zhang, Zhenjie Zhu, Zizhuo Li, Tao Lu et al.AAAI 2025 · 6 citations
- JamMa: Ultra-lightweight Local Feature Matching with Joint MambaXiaoyong Lu, Songlin DuCVPR 2025
Builds on9
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without ConvolutionsWenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan et al.ICCV 2021 · 4,909 citations
- Conditional Positional Encodings for Vision TransformersXiangxiang Chu, Zhi Tian, Bo Zhang, Xinlong Wang et al.ICLR 2023 · 406 citations
- Learning Two-View Correspondences and Geometry Using Order-Aware NetworkJiahui Zhang, Dawei Sun, Zixin Luo, Anbang Yao et al.ICCV 2019 · 362 citations
- An Image Patch is a Wave: Phase-Aware Vision MLPYehui Tang, Kai Han, Jianyuan Guo, Chang Xu et al.CVPR 2022 · 137 citations
Related papers
- XFeat: Accelerated Features for Lightweight Image MatchingGuilherme A. Potje, Felipe Cadar, André Araújo, Renato Martins et al.CVPR 2024 · 128 citations
- Learning to Match Features with Seeded Graph Matching NetworkHongkai Chen, Zixin Luo, Jiahui Zhang, Lei Zhou et al.ICCV 2021 · 165 citations
- ResMatch: Residual Attention Learning for Feature MatchingYuxin Deng, Kaining Zhang, Shihua Zhang, Yansheng Li et al.AAAI 2024 · 15 citations
- Geometrized Transformer for Self-Supervised Homography EstimationJiazhen Liu, Xirong LiICCV 2023 · 27 citations
- ETO: Efficient Transformer-based Local Feature Matching by Organizing Multiple Homography HypothesesJunjie Ni, Guofeng Zhang, Guanglin Li, Yijin Li et al.NeurIPS 2024 · 14 citations
