ParaFormer: Parallel Attention Transformer for Efficient Feature Matching
Xiaoyong Lu, Yaping Yan, Bin Kang, Songlin Du
摘要
Heavy computation is a bottleneck limiting deep-learning-based feature matching algorithms to be applied in many real-time applications. However, existing lightweight networks optimized for Euclidean data cannot address classical feature matching tasks, since sparse keypoint based descriptors are expected to be matched. This paper tackles this problem and proposes two concepts: 1) a novel parallel attention model entitled ParaFormer and 2) a graph based U-Net architecture with attentional pooling. First, ParaFormer fuses features and keypoint positions through the concept of amplitude and phase, and integrates self- and cross-attention in a parallel manner which achieves a win-win performance in terms of accuracy and efficiency. Second, with U-Net architecture and proposed attentional pooling, the ParaFormer-U variant significantly reduces computational complexity, and minimize performance loss caused by downsampling. Sufficient experiments on various applications, including homography estimation, pose estimation, and image matching, demonstrate that ParaFormer achieves state-of-the-art performance while maintaining high efficiency. The efficient ParaFormer-U variant achieves comparable performance with less than 50% FLOPs of the existing attention-based models.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Scene-Aware Feature MatchingXiaoyong Lu, Yaping Yan, Tong Wei, Songlin DuICCV 2023 · 被引用 9 次
- Matching While Perceiving: Enhance Image Feature Matching with Applicable Semantic AmalgamationShihua Zhang, Zhenjie Zhu, Zizhuo Li, Tao Lu 等AAAI 2025 · 被引用 6 次
- JamMa: Ultra-lightweight Local Feature Matching with Joint MambaXiaoyong Lu, Songlin DuCVPR 2025
它引用的顶会 Paper9
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without ConvolutionsWenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan 等ICCV 2021 · 被引用 4,909 次
- Conditional Positional Encodings for Vision TransformersXiangxiang Chu, Zhi Tian, Bo Zhang, Xinlong Wang 等ICLR 2023 · 被引用 406 次
- Learning Two-View Correspondences and Geometry Using Order-Aware NetworkJiahui Zhang, Dawei Sun, Zixin Luo, Anbang Yao 等ICCV 2019 · 被引用 362 次
- An Image Patch is a Wave: Phase-Aware Vision MLPYehui Tang, Kai Han, Jianyuan Guo, Chang Xu 等CVPR 2022 · 被引用 137 次
相关 Paper
- XFeat: Accelerated Features for Lightweight Image MatchingGuilherme A. Potje, Felipe Cadar, André Araújo, Renato Martins 等CVPR 2024 · 被引用 128 次
- Learning to Match Features with Seeded Graph Matching NetworkHongkai Chen, Zixin Luo, Jiahui Zhang, Lei Zhou 等ICCV 2021 · 被引用 165 次
- ResMatch: Residual Attention Learning for Feature MatchingYuxin Deng, Kaining Zhang, Shihua Zhang, Yansheng Li 等AAAI 2024 · 被引用 15 次
- Geometrized Transformer for Self-Supervised Homography EstimationJiazhen Liu, Xirong LiICCV 2023 · 被引用 27 次
- ETO: Efficient Transformer-based Local Feature Matching by Organizing Multiple Homography HypothesesJunjie Ni, Guofeng Zhang, Guanglin Li, Yijin Li 等NeurIPS 2024 · 被引用 14 次
