Robot Structure Prior Guided Temporal Attention for Camera-to-Robot Pose Estimation from Image Sequence
Yang Tian, Jiyao Zhang, Zekai Yin, Hao Dong
Abstract
In this work, we tackle the problem of online camera-torobot pose estimation from single-view successive frames of an image sequence, a crucial task for robots to interact with the world. The primary obstacles of this task are the robot's self-occlusions and the ambiguity of single-view images. This work demonstrates, for the first time, the effectiveness of temporal information and the robot structure prior in addressing these challenges. Given the successive frames and the robot joint configuration, our method learns to accurately regress the 2D coordinates of the predefined robot's keypoints (e.g. joints). With the camera intrinsic and robotic joints status known, we get the camerato-robot pose using a Perspective-n-point (PnP) solver. We further improve the camera-to-robot pose iteratively using the robot structure prior. To train the whole pipeline, we build a large-scale synthetic dataset generated with domain randomisation to bridge the sim-to-real gap. The extensive experiments on synthetic and real-world datasets and the downstream robotic grasping task demonstrate that our method achieves new state-of-the-art performances and outperforms traditional hand-eye calibration algorithms in real-time (36 FPS).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ce34010c-20cc-465c-a1b8-c2a62cacf41fCited by top-tier papers2
- RoboPEPP: Vision-Based Robot Pose and Joint Angle Estimation through Embedding Predictive Pre-TrainingRaktim Gautam Goswami, Prashanth Krishnamurthy, Yann LeCun, Farshad KhorramiCVPR 2025
- RoboTAG: End-to-end Robot Pose Estimation via Topological Alignment GraphYifan Liu, Fangneng Zhan, Wanhua Li, Haowen Sun et al.CVPR 2026
Builds on12
- CenterNet: Keypoint Triplets for Object DetectionKaiwen Duan, Song Bai, Lingxi Xie, Honggang Qi et al.ICCV 2019 · 3,348 citations
- TrackFormer: Multi-Object Tracking with TransformersTim Meinhardt, Alexander Kirillov, Laura Leal-Taixé, Christoph FeichtenhoferCVPR 2022 · 927 citations
- EPro-PnP: Generalized End-to-End Probabilistic Perspective-n-Points for Monocular Object Pose EstimationHansheng Chen, Pichao Wang, Fan Wang, Wei Tian et al.CVPR 2022 · 175 citations
- PnP-DETR: Towards Efficient Visual Analysis with TransformersTao Wang, Li Yuan, Yunpeng Chen, Jiashi Feng et al.ICCV 2021 · 125 citations
- RePOSE: Fast 6D Object Pose Refinement via Deep Texture RenderingShun Iwase, Xingyu Liu, Rawal Khirodkar, Rio Yokota et al.ICCV 2021 · 103 citations
Related papers
- Single-View Robot Pose and Joint Angle Estimation via Render & CompareYann Labbé, Justin Carpentier, Mathieu Aubry, Josef SivicCVPR 2021
- Markerless Camera-to-Robot Pose Estimation via Self-Supervised Sim-to-Real TransferJingpei Lu, Florian Richter, Michael C. YipCVPR 2023
- MoreFusion: Multi-object Reasoning for 6D Pose Estimation from Volumetric FusionKentaro Wada, Edgar Sucar, Stephen James, Daniel Lenton et al.CVPR 2020
- Online Unsupervised Learning of the 3D Kinematic Structure of Arbitrary Rigid BodiesUrbano Miguel Nunes, Yiannis DemirisICCV 2019 · 4 citations
- SO-Pose: Exploiting Self-Occlusion for Direct 6D Pose EstimationYan Di, Fabian Manhardt, Gu Wang, Xiangyang Ji et al.ICCV 2021 · 163 citations
