Sim2Real Object-Centric Keypoint Detection and Description
Chengliang Zhong, Chao Yang, Fuchun Sun, Jinshan Qi, Xiaodong Mu, Huaping Liu, Wenbing Huang
Abstract
Keypoint detection and description play a central role in computer vision. Most existing methods are in the form of scene-level prediction, without returning the object classes of different keypoints. In this paper, we propose the objectcentric formulation, which, beyond the conventional setting, requires further identifying which object each interest point belongs to. With such fine-grained information, our framework enables more downstream potentials, such as objectlevel matching and pose estimation in a clustered environment. To get around the difficulty of label collection in the real world, we develop a sim2real contrastive learning mechanism that can generalize the model trained in simulation to real-world applications. The novelties of our training method are three-fold: (i) we integrate the uncertainty into the learning framework to improve feature description of hard cases, e.g., less-textured or symmetric patches; (ii) we decouple the object descriptor into two output branches-intra-object salience and inter-object distinctness, resulting in a better pixel-wise description; (iii) we enforce cross-view semantic consistency for enhanced robustness in representation learning. Comprehensive experiments on image matching and 6D pose estimation verify the encouraging generalization ability of our method from simulation to reality. Particularly for 6D pose estimation, our method significantly outperforms typical unsupervised/sim2real methods, achieving a closer gap with the fully supervised counterpart. Additional results and videos can be found at https://zhongcl-thu.github.io/rock/ .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- OnePose++: Keypoint-Free One-Shot Object Pose Estimation without CAD ModelsXingyi He, Jiaming Sun, Yuang Wang, Di Huang et al.NeurIPS 2022 · 190 citations
- 3D Implicit Transporter for Temporally Consistent Keypoint DiscoveryChengliang Zhong, Yuhang Zheng, Yupeng Zheng, Hao Zhao et al.ICCV 2023 · 23 citations
- SNAKE: Shape-aware Neural 3D Keypoint FieldChengliang Zhong, Peixing You, Xiaoxue Chen, Hao Zhao et al.NeurIPS 2022 · 17 citations
Builds on8
- Digging Into Self-Supervised Monocular Depth EstimationClément Godard, Oisin Mac Aodha, Michael Firman, Gabriel J. BrostowICCV 2019 · 2,416 citations
- DISK: Learning local features with policy gradientMichal J. Tyszkiewicz, Pascal Fua, Eduard TrullsNeurIPS 2020 · 652 citations
- Key.Net: Keypoint Detection by Handcrafted and Learned CNN FiltersAxel Barroso Laguna, Edgar Riba, Daniel Ponsa, Krystian MikolajczykICCV 2019 · 323 citations
- D3VO: Deep Depth, Deep Pose and Deep Uncertainty for Monocular Visual OdometryNan Yang, Lukas von Stumberg, Rui Wang, Daniel CremersCVPR 2020
- Dense Contrastive Learning for Self-Supervised Visual Pre-TrainingXinlong Wang, Rufeng Zhang, Chunhua Shen, Tao Kong et al.CVPR 2021
Related papers
- Self-Supervised Visual Representation Learning with Semantic GroupingXin Wen, Bingchen Zhao, Anlin Zheng, Xiangyu Zhang et al.NeurIPS 2022 · 104 citations
- Instance-Adaptive and Geometric-Aware Keypoint Learning for Category-Level 6D Object Pose EstimationXiao Lin, Wenfei Yang, Yuan Gao, Tianzhu ZhangCVPR 2024
- Symmetry and Uncertainty-Aware Object SLAM for 6DoF Object Pose EstimationNathaniel Merrill, Yuliang Guo, Xingxing Zuo, Xinyu Huang et al.CVPR 2022 · 42 citations
- Point-GCC: Universal Self-supervised 3D Scene Pre-training via Geometry-Color ContrastGuofan Fan, Zekun Qi, Wenkai Shi, Kaisheng MaACM MM 2024 · 12 citations
- Learning Symmetry-Aware Geometry Correspondences for 6D Object Pose EstimationHeng Zhao, Shenxing Wei, Dahu Shi, Wenming Tan et al.ICCV 2023 · 33 citations
