ETO: Efficient Transformer-based Local Feature Matching by Organizing Multiple Homography Hypotheses
Junjie Ni, Guofeng Zhang, Guanglin Li, Yijin Li, Xinyang Liu, Zhaoyang Huang, Hujun Bao
Abstract
Recent developments have led to the emergence of transformer-based approaches for local feature matching, resulting in enhanced accuracy of matches. However, the time required for transformer-based feature enhancement is excessively long, which limits their practical application. In this paper, we propose methods to reduce the computational load of transformers during both the coarse matching and refinement stages. During the coarse matching phase, we organize multiple homography hypotheses to approximate continuous matches. Each hypothesis encompasses several features to be matched, significantly reducing the number of features that require enhancement via transformers. In the refinement stage, we reduce the bidirectional self-attention and cross-attention mechanisms to unidirectional cross-attention, thereby substantially decreasing the cost of computation. Overall, our method demonstrates at least 4 times faster compared to other transformerbased feature matching algorithms. Comprehensive evaluations on other open datasets such as Megadepth, YFCC100M, ScanNet, and HPatches demonstrate our method's efficacy, highlighting its potential to significantly enhance a wide array of downstream applications.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- A Global Depth-Range-Free Multi-View Stereo Transformer Network with Pose EmbeddingYitong Dong, Yijin Li, Zhaoyang Huang, Weikang Bian et al.NeurIPS 2024 · 7 citations
- EDM: Efficient Deep Feature MatchingXi Li, Tong Rao, Cihui PanICCV 2025 · 6 citations
- No Calibration, No Depth, No Problem: Cross-Sensor View Synthesis with 3D ConsistencyCho-Ying Wu, Zixun Huang, Xinyu Huang, Liu RenCVPR 2026
Builds on21
- Transformers are RNNs: Fast Autoregressive Transformers with Linear AttentionAngelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, François FleuretICML 2020 · 2,665 citations
- LightGlue: Local Feature Matching at Light SpeedPhilipp Lindenberger, Paul-Edouard Sarlin, Marc PollefeysICCV 2023 · 936 citations
- COTR: Correspondence Transformer for Matching Across ImagesWei Jiang, Eduard Trulls, Jan Hosang, Andrea Tagliasacchi et al.ICCV 2021 · 318 citations
- Depth-Adaptive TransformerMaha Elbayad, Jiatao Gu, Edouard Grave, Michael AuliICLR 2020 · 264 citations
- Quadtree Attention for Vision TransformersShitao Tang, Jiahui Zhang, Siyu Zhu, Ping TanICLR 2022 · 194 citations
Related papers
- Geometrized Transformer for Self-Supervised Homography EstimationJiazhen Liu, Xirong LiICCV 2023 · 27 citations
- LoFTR: Detector-Free Local Feature Matching With TransformersJiaming Sun, Zehong Shen, Yuang Wang, Hujun Bao et al.CVPR 2021
- Unifying Feature and Cost Aggregation with Transformers for Semantic and Visual CorrespondenceSunghwan Hong, Seokju Cho, Seungryong Kim, Stephen LinICLR 2024 · 16 citations
- Improving Transformer-based Image Matching by Cascaded Capturing Spatially Informative KeypointsChenjie Cao, Yanwei FuICCV 2023 · 23 citations
- LocalTrans: A Multiscale Local Transformer Network for Cross-Resolution Homography EstimationRuizhi Shao, Gaochang Wu, Yuemei Zhou, Ying Fu et al.ICCV 2021 · 57 citations
