Geometrized Transformer for Self-Supervised Homography Estimation
Jiazhen Liu, Xirong Li
Abstract
For homography estimation, we propose Geometrized Transformer (GeoFormer), a new detector-free feature matching method. Current detector-free methods, e.g. LoFTR, lack an effective mean to accurately localize small and thus computationally feasible regions for cross-attention diffusion. We resolve the challenge with an extremely simple idea: using the classical RANSAC geometry for attentive region search. Given coarse matches by LoFTR, a homography is obtained with ease. Such a homography allows us to compute cross-attention in a focused manner, where key/value sets required by Transformers can be reduced to small fix-sized regions rather than an entire image. Local features can thus be enhanced by standard Transformers. We integrate GeoFormer into the LoFTR framework. By minimizing a multi-scale cross-entropy based matching loss on auto-generated training data, the network is trained in a fully self-supervised manner. Extensive experiments are conducted on multiple real-world datasets covering natural images, heavily manipulated pictures and retinal images. The proposed method compares favorably against the state-of-the-art.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b9adccea-1267-46d8-a219-57b1f5d5f7f9Cited by top-tier papers3
- Discrete Latent Perspective Learning for Segmentation and DetectionDeyi Ji, Feng Zhao, Lanyun Zhu, Wenwei Jin et al.ICML 2024 · 21 citations
- Semantic-aware Representation Learning for Homography EstimationYuhan Liu, Qianxin Huang, Siqi Hui, Jingwen Fu et al.ACM MM 2024 · 6 citations
- Auto-Regressive Transformation for Image AlignmentKanggeon Lee, Soochahn Lee, Kyoung Mu LeeICCV 2025 · 1 citation
Builds on6
- Quadtree Attention for Vision TransformersShitao Tang, Jiahui Zhang, Siyu Zhu, Ping TanICLR 2022 · 194 citations
- GLAMpoints: Greedily Learned Accurate Match PointsPrune Truong, Stefanos Apostolopoulos, Agata Mosinska, Samuel Stucky et al.ICCV 2019 · 77 citations
- Motion Basis Learning for Unsupervised Deep Homography Estimation with Subspace ProjectionNianjin Ye, Chuan Wang, Haoqiang Fan, Shuaicheng LiuICCV 2021 · 69 citations
- Unsupervised Homography Estimation with Coplanarity-Aware GANMingbo Hong, Yuhang Lu, Nianjin Ye, Chunyu Lin et al.CVPR 2022 · 62 citations
- SuperGlue: Learning Feature Matching With Graph Neural NetworksPaul-Edouard Sarlin, Daniel DeTone, Tomasz Malisiewicz, Andrew RabinovichCVPR 2020
Related papers
- LoFTR: Detector-Free Local Feature Matching With TransformersJiaming Sun, Zehong Shen, Yuang Wang, Hujun Bao et al.CVPR 2021
- Efficient LoFTR: Semi-Dense Local Feature Matching with Sparse-Like SpeedYifan Wang, Xingyi He, Sida Peng, Dongli Tan et al.CVPR 2024 · 126 citations
- Improving Transformer-based Image Matching by Cascaded Capturing Spatially Informative KeypointsChenjie Cao, Yanwei FuICCV 2023 · 23 citations
- ETO: Efficient Transformer-based Local Feature Matching by Organizing Multiple Homography HypothesesJunjie Ni, Guofeng Zhang, Guanglin Li, Yijin Li et al.NeurIPS 2024 · 14 citations
- Multi-Frame Self-Supervised Depth with TransformersVitor Guizilini, Rares Ambrus, Dian Chen, Sergey Zakharov et al.CVPR 2022 · 95 citations
