Self-Supervised Geometric Perception
Heng Yang, Wei Dong, Luca Carlone, Vladlen Koltun
Abstract
We present self-supervised geometric perception (SGP), the first general framework to learn a feature descriptor for correspondence matching without any ground-truth geometric model labels (e.g., camera poses, rigid transformations). Our first contribution is to formulate geometric perception as an optimization problem that jointly optimizes the feature descriptor and the geometric models given a large corpus of visual measurements (e.g., images, point clouds). Under this optimization formulation, we show that two important streams of research in vision, namely robust model fitting and deep feature learning, correspond to optimizing one block of the unknown variables while fixing the other block. This analysis naturally leads to our second contribution -the SGP algorithm that performs alternating minimization to solve the joint optimization. SGP iteratively executes two meta-algorithms: a teacher that performs robust model fitting given learned features to generate geometric pseudo-labels, and a student that performs deep feature learning under noisy supervision of the pseudo-labels. As a third contribution, we apply SGP to two perception problems on large-scale real datasets, namely relative camera pose estimation on MegaDepth and point cloud registration on 3DMatch. We demonstrate that SGP achieves stateof-the-art performance that is on-par or superior to the supervised oracles trained using ground-truth labels. 1 * Equal contribution. Work performed during internship at Intel Labs.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 48d9830b-4907-4bea-8e1e-ca8bc5821932Cited by top-tier papers9
- Lepard: Learning partial point cloud matching in rigid and deformable scenesYang Li, Tatsuya HaradaCVPR 2022 · 163 citations
- GIM: Learning Generalizable Image Matcher From Internet VideosXuelun Shen, Zhipeng Cai, Wei Yin, Matthias Müller et al.ICLR 2024 · 80 citations
- PUMP: Pyramidal and Uniqueness Matching Priors for Unsupervised Learning of Local DescriptorsJérôme Revaud, Vincent Leroy, Philippe Weinzaepfel, Boris ChidlovskiiCVPR 2022 · 16 citations
- Dynamical Pose EstimationHeng Yang, Chris Doran, Jean-Jacques E. SlotineICCV 2021 · 10 citations
- TUSK: Task-Agnostic Unsupervised KeypointsYuhe Jin, Weiwei Sun, Jan Hosang, Eduard Trulls et al.NeurIPS 2022 · 6 citations
Builds on16
- Digging Into Self-Supervised Monocular Depth EstimationClément Godard, Oisin Mac Aodha, Michael Firman, Gabriel J. BrostowICCV 2019 · 2,416 citations
- Deep Closest Point: Learning Representations for Point Cloud RegistrationYue Wang, Justin SolomonICCV 2019 · 1,026 citations
- Fully Convolutional Geometric FeaturesChristopher B. Choy, Jaesik Park, Vladlen KoltunICCV 2019 · 807 citations
- Rethinking Pre-training and Self-trainingBarret Zoph, Golnaz Ghiasi, Tsung-Yi Lin, Yin Cui et al.NeurIPS 2020 · 755 citations
- DPOD: 6D Pose Object Detector and RefinerSergey Zakharov, Ivan Shugurov, Slobodan IlicICCV 2019 · 486 citations
Related papers
- DKM: Dense Kernelized Feature Matching for Geometry EstimationJohan Edstedt, Ioannis Athanasiadis, Mårten Wadenbäck, Michael FelsbergCVPR 2023
- Warp Consistency for Unsupervised Learning of Dense CorrespondencesPrune Truong, Martin Danelljan, Fisher Yu, Luc Van GoolICCV 2021 · 60 citations
- CorrNet3D: Unsupervised End-to-End Learning of Dense Correspondence for 3D Point CloudsYiming Zeng, Yue Qian, Zhiyu Zhu, Junhui Hou et al.CVPR 2021
- Bootstrap Your Own CorrespondencesMohamed El Banani, Justin JohnsonICCV 2021 · 45 citations
- Leveraging SE(3) Equivariance for Self-supervised Category-Level Object Pose Estimation from Point CloudsXiaolong Li, Yijia Weng, Li Yi, Leonidas J. Guibas et al.NeurIPS 2021 · 61 citations
