RoboPEPP: Vision-Based Robot Pose and Joint Angle Estimation through Embedding Predictive Pre-Training
Raktim Gautam Goswami, Prashanth Krishnamurthy, Yann LeCun, Farshad Khorrami
Abstract
Vision-based pose estimation of articulated robots with unknown joint angles has applications in collaborative robotics and human-robot interaction tasks. Current frameworks use neural network encoders to extract image features and downstream layers to predict joint angles and robot pose. While images of robots inherently contain rich information about the robot's physical structures, existing methods often fail to leverage it fully; therefore, limiting performance under occlusions and truncations. To address this, we introduce RoboPEPP, a method that fuses information about the robot's physical model into the encoder using a masking-based self-supervised embedding-predictive architecture. Specifically, we mask the robot's joints and pre-train an encoder-predictor model to infer the joints' embeddings from surrounding unmasked regions, enhancing the encoder's understanding of the robot's physical model. The pre-trained encoder-predictor pair, along with joint angle and keypoint prediction networks, is then fine-tuned for pose and joint angle estimation. Random masking of input during fine-tuning and keypoint filtering during evaluation further improves robustness. Our method, evaluated on several datasets, achieves the best results in robot pose and joint angle estimation while being the least sensitive to occlusions and requiring the lowest execution time. The code is available at https://github.com/raktimgg/RoboPEPP .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 80d75c99-8c17-4965-b283-54ddeba682bfCited by top-tier papers2
- RoboTAG: End-to-end Robot Pose Estimation via Topological Alignment GraphYifan Liu, Fangneng Zhan, Wanhua Li, Haowen Sun et al.CVPR 2026
- Visual-RRT: Finding Paths toward Visual-Goals via Differentiable RenderingSebin Lee, Jumin Lee, Taeyeon Kim, Youngju Na et al.CVPR 2026
Builds on9
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- VICReg: Variance-Invariance-Covariance Regularization for Self-Supervised LearningAdrien Bardes, Jean Ponce, Yann LeCunICLR 2022 · 1,226 citations
- Self-labelling via simultaneous clustering and representation learningYuki Markus Asano, Christian Rupprecht, Andrea VedaldiICLR 2020 · 873 citations
- End-to-End Learnable Geometric Vision by Backpropagating PnP OptimizationBo Chen, Álvaro Parra, Jiewei Cao, Nan Li et al.CVPR 2020
- Learning Human-to-Robot Handovers from Point CloudsSammy Joe Christen, Wei Yang, Claudia Pérez-D'Arpino, Otmar Hilliges et al.CVPR 2023
Related papers
- Single-View Robot Pose and Joint Angle Estimation via Render & CompareYann Labbé, Justin Carpentier, Mathieu Aubry, Josef SivicCVPR 2021
- Robot Structure Prior Guided Temporal Attention for Camera-to-Robot Pose Estimation from Image SequenceYang Tian, Jiyao Zhang, Zekai Yin, Hao DongCVPR 2023
- Markerless Camera-to-Robot Pose Estimation via Self-Supervised Sim-to-Real TransferJingpei Lu, Florian Richter, Michael C. YipCVPR 2023
- SO-Pose: Exploiting Self-Occlusion for Direct 6D Pose EstimationYan Di, Fabian Manhardt, Gu Wang, Xiangyang Ji et al.ICCV 2021 · 163 citations
- Category-Level Articulated Object Pose EstimationXiaolong Li, He Wang, Li Yi, Leonidas J. Guibas et al.CVPR 2020
