RDD: Robust Feature Detector and Descriptor using Deformable Transformer
Gonglin Chen, Tianwen Fu, Haiwei Chen, Wenbin Teng, Hanyuan Xiao, Yajie Zhao
Abstract
As a core step in structure-from-motion and SLAM, robust feature detection and description under challenging scenarios such as significant viewpoint changes remain unresolved despite their ubiquity. While recent works have identified the importance of local features in modeling geometric transformations, these methods fail to learn the visual cues present in long-range relationships. We present Robust Deformable Detector (RDD), a novel and robust keypoint detector/descriptor leveraging the deformable transformer, which captures global context and geometric invariance through deformable self-attention mechanisms. Specifically, we observed that deformable attention focuses on key locations, effectively reducing the search space complexity and modeling the geometric invariance. Furthermore, we collected an Air-to-Ground dataset for training in addition to the standard MegaDepth dataset. Our proposed method outperforms all state-of-the-art keypoint detection/description methods in sparse matching tasks and is also capable of semi-dense matching. To ensure comprehensive evaluation, we introduce two challenging benchmarks: one emphasizing large viewpoint and scale variations, and the other being an Air-to-Ground benchmark -an evaluation setting that has recently gaining popularity for 3D reconstruction across different altitudes. Project
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers6
- AsymLoc: Towards Asymmetric Feature Matching for Efficient Visual LocalizationMohammad Omama, Gabriele Berton, Eric Foxlin, Yelin KimCVPR 2026
- SGAT: Learning Feature Matching with Singularity-enhanced Graph Attention NetworkYizhuo Zhang, Kun Sun, Chang Tang, Yuanyuan Liu et al.AAAI 2026
- From Pairs to Sequences: Track-Aware Policy Gradients for Keypoint DetectionYepeng Liu, Hao Li, Liwen Yang, Fangzhen Li et al.CVPR 2026
- AerialFusion: Co-Motion-Driven Unified Registration and Fusion on Multi-modal Data Streams from Aerial ViewJunhui Qiu, Xiang Xiang, Hongyun Wang, Jiaqi GuiAAAI 2026
- MMVIP: A Visible-infrared Paired Dataset for Multi-weather Marine VisionYunpeng Yin, Lihan Wang, Zhaoshen He, Xinqiang He et al.CVPR 2026
Builds on14
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li et al.ICLR 2021 · 7,353 citations
- 3D Gaussian Splatting for Real-Time Radiance Field RenderingBernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, George DrettakisSIGGRAPH 2023 · 5,687 citations
- LightGlue: Local Feature Matching at Light SpeedPhilipp Lindenberger, Paul-Edouard Sarlin, Marc PollefeysICCV 2023 · 936 citations
- DISK: Learning local features with policy gradientMichal J. Tyszkiewicz, Pascal Fua, Eduard TrullsNeurIPS 2020 · 652 citations
- Local Relation Networks for Image RecognitionHan Hu, Zheng Zhang, Zhenda Xie, Stephen LinICCV 2019 · 555 citations
Related papers
- Key-Grid: Unsupervised 3D Keypoints Detection using Grid Heatmap FeaturesChengkai Hou, Zhengrong Xue, Bingyang Zhou, Jinghan Ke et al.NeurIPS 2024 · 9 citations
- UniMapping: Unified SLAM Framework for Map-Centric Embodied PerceptionXiaze Zhang, Ziheng Ding, Yuejie Zhang, lifeng chen et al.ICML 2026
- Rotation-Invariant Transformer for Point Cloud MatchingHao Yu, Zheng Qin, Ji Hou, Mahdi Saleh et al.CVPR 2023
- Context and Geometry Aware Voxel Transformer for Semantic Scene CompletionZhu Yu, Runmin Zhang, Jiacheng Ying, Junchen Yu et al.NeurIPS 2024 · 73 citations
- RGGT: A Generative-Prior-Guided Transformer for Unified Rigid and Non-Rigid Point Cloud RegistrationChengyu Zheng, Songlin Yang, Jin Huang, Honghua Chen et al.ICML 2026
