RIPE: Reinforcement Learning on Unlabeled Image Pairs for Robust Keypoint Extraction
Johannes Künzel, Anna Hilsmann, Peter Eisert
Abstract
We introduce RIPE, an innovative reinforcement learning-based framework for weakly-supervised training of a keypoint extractor that excels in both detection and description tasks. In contrast to conventional training regimes that depend heavily on artificial transformations, pre-generated models, or 3D data, RIPE requires only a binary label indicating whether paired images represent the same scene. This minimal supervision significantly expands the pool of training data, enabling the creation of a highly generalized and robust keypoint extractor. RIPE utilizes the encoder's intermediate layers for the description of the keypoints with a hyper-column approach to integrate information from different scales. Additionally, we propose an auxiliary loss to enhance the discriminative capability of the learned descriptors. Comprehensive evaluations on standard benchmarks demonstrate that RIPE simplifies data preparation while achieving competitive performance compared to state-of-the-art techniques, marking a significant advancement in robust keypoint extraction and description. To support further research, we have made our code publicly available at https://github.com/fraunhoferhhi/RIPE.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e53f03fc-446e-4350-bf84-625ba0bd83d6Cited by top-tier papers2
- Restore Text First, Enhance Image Later: Two-Stage Scene Text Image Super-Resolution with Glyph Structure GuidanceMinxing Luo, Linlong Fan, Qiushi Wang, Ge Wu et al.CVPR 2026 · 2 citations
- From Pairs to Sequences: Track-Aware Policy Gradients for Keypoint DetectionYepeng Liu, Hao Li, Liwen Yang, Fangzhen Li et al.CVPR 2026
Builds on10
- ACDC: The Adverse Conditions Dataset with Correspondences for Semantic Driving Scene UnderstandingChristos Sakaridis, Dengxin Dai, Luc Van GoolICCV 2021 · 655 citations
- DISK: Learning local features with policy gradientMichal J. Tyszkiewicz, Pascal Fua, Eduard TrullsNeurIPS 2020 · 652 citations
- DUSt3R: Geometric 3D Vision Made EasyShuzhe Wang, Vincent Leroy, Yohann Cabon, Boris Chidlovskii et al.CVPR 2024 · 302 citations
- Efficient LoFTR: Semi-Dense Local Feature Matching with Sparse-Like SpeedYifan Wang, Xingyi He, Sida Peng, Dongli Tan et al.CVPR 2024 · 126 citations
- Hyperpixel Flow: Semantic Correspondence With Multi-Layer Neural FeaturesJuhong Min, Jongmin Lee, Jean Ponce, Minsu ChoICCV 2019 · 120 citations
Related papers
- S-TREK: Sequential Translation and Rotation Equivariant Keypoints for local feature extractionEmanuele Santellani, Christian Sormann, Mattia Rossi, Andreas Kuhn et al.ICCV 2023 · 17 citations
- Reinforced Feature Points: Optimizing Feature Detection and Description for a High-Level TaskAritra Bhowmik, Stefan Gumhold, Carsten Rother, Eric BrachmannCVPR 2020
- Unsupervised Learning of Object Landmarks via Self-Training CorrespondenceDimitrios Mallis, Enrique Sanchez, Matthew Bell, Georgios TzimiropoulosNeurIPS 2020 · 20 citations
- Unsupervised Learning of Visual 3D Keypoints for ControlBoyuan Chen, Pieter Abbeel, Deepak PathakICML 2021 · 46 citations
- Learning Transformation-Predictive Representations for Detection and Description of Local FeaturesZihao Wang, Chunxu Wu, Yifei Yang, Zhen LiCVPR 2023
