D2Former: Jointly Learning Hierarchical Detectors and Contextual Descriptors via Agent-Based Transformers
Jianfeng He, Yuan Gao, Tianzhu Zhang, Zhe Zhang, Feng Wu
Abstract
Establishing pixel-level matches between image pairs is vital for a variety of computer vision applications. However, achieving robust image matching remains challenging because CNN extracted descriptors usually lack discriminative ability in texture-less regions and keypoint detectors are only good at identifying keypoints with a specific level of structure. To deal with these issues, a novel image matching method is proposed by Jointly Learning Hierarchical Detectors and Contextual Descriptors via Agentbased Transformers (D 2 Former), including a contextual feature descriptor learning (CFDL) module and a hierarchical keypoint detector learning (HKDL) module. The proposed D 2 Former enjoys several merits. First, the proposed CFDL module can model long-range contexts efficiently and effectively with the aid of designed descriptor agents. Second, the HKDL module can generate keypoint detectors in a hierarchical way, which is helpful for detecting keypoints with diverse levels of structures. Extensive experimental results on four challenging benchmarks show that our proposed method significantly outperforms stateof-the-art image matching methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 11b614ad-fd84-489f-b7d0-86b09cbffe73Builds on9
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- DISK: Learning local features with policy gradientMichal J. Tyszkiewicz, Pascal Fua, Eduard TrullsNeurIPS 2020 · 652 citations
- Learning Two-View Correspondences and Geometry Using Order-Aware NetworkJiahui Zhang, Dawei Sun, Zixin Luo, Anbang Yao et al.ICCV 2019 · 362 citations
- Dual-Resolution Correspondence NetworksXinghui Li, Kai Han, Shuda Li, Victor PrisacariuNeurIPS 2020 · 207 citations
- GLAMpoints: Greedily Learned Accurate Match PointsPrune Truong, Stefanos Apostolopoulos, Agata Mosinska, Samuel Stucky et al.ICCV 2019 · 77 citations
Related papers
- SD2Event: Self-Supervised Learning of Dynamic Detectors and Contextual Descriptors for Event CamerasYuan Gao, Yuqing Zhu, Xinjun Li, Yimin Du et al.CVPR 2024
- Collaborative Feature Matching with Progressive Correspondence LearningXin Liu, Yanbing Han, Rong Qin, Bing Wang et al.AAAI 2026
- LoFTR: Detector-Free Local Feature Matching With TransformersJiaming Sun, Zehong Shen, Yuang Wang, Hujun Bao et al.CVPR 2021
- Improving Transformer-based Image Matching by Cascaded Capturing Spatially Informative KeypointsChenjie Cao, Yanwei FuICCV 2023 · 23 citations
- 2D3D-MATR: 2D-3D Matching Transformer for Detection-free Registration between Images and Point CloudsMinhao Li, Zheng Qin, Zhirui Gao, Renjiao Yi et al.ICCV 2023 · 30 citations
