Markerless Camera-to-Robot Pose Estimation via Self-Supervised Sim-to-Real Transfer
Jingpei Lu, Florian Richter, Michael C. Yip
Abstract
Solving the camera-to-robot pose is a fundamental requirement for vision-based robot control, and is a process that takes considerable effort and cares to make accurate. Traditional approaches require modification of the robot via markers, and subsequent deep learning approaches enabled markerless feature extraction. Mainstream deep learning methods only use synthetic data and rely on Domain Randomization to fill the sim-to-real gap, because acquiring the 3D annotation is labor-intensive. In this work, we go beyond the limitation of 3D annotations for real-world data. We propose an end-to-end pose estimation framework that is capable of online camera-to-robot calibration and a self-supervised training method to scale the training to unlabeled real-world data. Our framework combines deep learning and geometric vision for solving the robot pose, and the pipeline is fully differentiable. To train the Camerato-Robot Pose Estimation Network (CtRNet), we leverage foreground segmentation and differentiable rendering for image-level self-supervision. The pose prediction is visualized through a renderer and the image loss with the input image is back-propagated to train the neural network. Our experimental results on two public real datasets confirm the effectiveness of our approach over existing works. We also integrate our framework into a visual servoing system to demonstrate the promise of real-time precise robot pose estimation for automation tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 919ac562-367f-4d0e-a471-c84697411596Cited by top-tier papers4
- Towards Long-Horizon Vision-Language-Action System: Reasoning, Acting and MemoryDaixun Li, Yusi Zhang, Mingxiang Cao, Donglai Liu et al.ICCV 2025 · 2 citations
- EgoRoC: Towards Egocentric Robotic Control via Task-Agnostic Visual AlignmentWei Feng, Chi Zhang, Nan Li, Qian Zhang et al.CVPR 2026
- RoboPEPP: Vision-Based Robot Pose and Joint Angle Estimation through Embedding Predictive Pre-TrainingRaktim Gautam Goswami, Prashanth Krishnamurthy, Yann LeCun, Farshad KhorramiCVPR 2025
- RoboTAG: End-to-end Robot Pose Estimation via Topological Alignment GraphYifan Liu, Fangneng Zhan, Wanhua Li, Haowen Sun et al.CVPR 2026
Builds on3
- Soft Rasterizer: A Differentiable Renderer for Image-Based 3D ReasoningShichen Liu, Weikai Chen, Tianye Li, Hao LiICCV 2019 · 789 citations
- End-to-End Learnable Geometric Vision by Backpropagating PnP OptimizationBo Chen, Álvaro Parra, Jiewei Cao, Nan Li et al.CVPR 2020
- Single-View Robot Pose and Joint Angle Estimation via Render & CompareYann Labbé, Justin Carpentier, Mathieu Aubry, Josef SivicCVPR 2021
Related papers
- SMOC-Net: Leveraging Camera Pose for Self-Supervised Monocular Object Pose EstimationTao Tan, Qiulei DongCVPR 2023
- Generalizing to the Open World: Deep Visual Odometry With Online AdaptationShunkai Li, Xin Wu, Yingdian Cao, Hongbin ZhaCVPR 2021
- CrossLoc: Scalable Aerial Localization Assisted by Multimodal Synthetic DataQi Yan, Jianhao Zheng, Simon Reding, Shanci Li et al.CVPR 2022 · 23 citations
- GDR-Net: Geometry-Guided Direct Regression Network for Monocular 6D Object Pose EstimationGu Wang, Fabian Manhardt, Federico Tombari, Xiangyang JiCVPR 2021
- FlowCam: Training Generalizable 3D Radiance Fields without Camera Poses via Pixel-Aligned Scene FlowCameron Smith, Yilun Du, Ayush Tewari, Vincent SitzmannNeurIPS 2023 · 43 citations
