Unsupervised Learning of Visual 3D Keypoints for Control
Boyuan Chen, Pieter Abbeel, Deepak Pathak
Abstract
Learning sensorimotor control policies from highdimensional images crucially relies on the quality of the underlying visual representations. Prior works show that structured latent space such as visual keypoints often outperforms unstructured representations for robotic control. However, most of these representations, whether structured or unstructured are learned in a 2D space even though the control tasks are usually performed in a 3D environment. In this work, we propose a framework to learn such a 3D geometric structure directly from images in an end-toend unsupervised manner. The input images are embedded into latent 3D keypoints via a differentiable encoder which is trained to optimize both a multi-view consistency loss and downstream task objective. These discovered 3D keypoints tend to meaningfully capture robot joints as well as object movements in a consistent manner across both time and 3D space. The proposed approach outperforms prior state-of-art methods across a variety of reinforcement learning benchmarks. Code and videos at https://buoyancy99.github. io/unsup-3d-keypoints/ .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers11
- Multi-View Masked World Models for Visual Robotic ManipulationYounggyo Seo, Junsu Kim, Stephen James, Kimin Lee et al.ICML 2023 · 99 citations
- DynaMo: In-Domain Dynamics Pretraining for Visuo-Motor ControlZichen Jeff Cui, Hengkai Pan, Aadhithya Iyer, Siddhant Haldar et al.NeurIPS 2024 · 61 citations
- MoVie: Visual Model-Based Policy Adaptation for View GeneralizationSizhe Yang, Yanjie Ze, Huazhe XuNeurIPS 2023 · 29 citations
- 3D Implicit Transporter for Temporally Consistent Keypoint DiscoveryChengliang Zhong, Yuhang Zheng, Yupeng Zheng, Hao Zhao et al.ICCV 2023 · 23 citations
- Neural Marionette: Unsupervised Learning of Motion Skeleton and Latent Dynamics from Volumetric VideoJinseok Bae, Hojun Jang, Cheol-Hui Min, Hyungun Choi et al.AAAI 2022 · 6 citations
Builds on4
- CURL: Contrastive Unsupervised Representations for Reinforcement LearningMichael Laskin, Aravind Srinivas, Pieter AbbeelICML 2020 · 1,261 citations
- Image Augmentation Is All You Need: Regularizing Deep Reinforcement Learning from PixelsDenis Yarats, Ilya Kostrikov, Rob FergusICLR 2021 · 911 citations
- Reinforcement Learning with Augmented DataMichael Laskin, Kimin Lee, Adam Stooke, Lerrel Pinto et al.NeurIPS 2020 · 833 citations
- Learning Deep Network for Detecting 3D Object Keypoints and 6D PosesWanqing Zhao, Shaobo Zhang, Ziyu Guan, Wei Zhao et al.CVPR 2020
Related papers
- Unsupervised 3D Structure Inference from Category-Specific Image CollectionsWeikang Wang, Dongliang Cao, Florian BernardCVPR 2024
- Key-Grid: Unsupervised 3D Keypoints Detection using Grid Heatmap FeaturesChengkai Hou, Zhengrong Xue, Bingyang Zhou, Jinghan Ke et al.NeurIPS 2024 · 9 citations
- Towards Unsupervised Learning of Generative Models for 3D Controllable Image SynthesisYiyi Liao, Katja Schwarz, Lars M. Mescheder, Andreas GeigerCVPR 2020
- 3D-DLP: Self-supervised 3D Object-centric Scene Representation LearningEllina Zhang, Madhavan Iyengar, Amir Zadeh, Chuan Li et al.ICML 2026
- GeoMoLa: Geometry-Aware Motion Latents for Learning Robust Manipulation PoliciesYunchao Zhang, Yijia Weng, Ruizhe Liu, Ming Hu et al.ICML 2026
