RoboTAG: End-to-end Robot Pose Estimation via Topological Alignment Graph
Yifan Liu, Fangneng Zhan, Wanhua Li, Haowen Sun, Katerina Fragkiadaki, Hanspeter Pfister
Abstract
Estimating robot pose from a monocular RGB image is a challenge in robotics and computer vision. Existing methods typically build networks on top of 2D visual backbones and depend heavily on labeled data for training, which is often scarce in real-world scenarios, causing a sim-to-real gap. Moreover, these approaches reduce the 3D-based problem to 2D domain, neglecting the 3D priors. To address these, we propose Robot Topological Alignment Graph (RoboTAG), which incorporates a 3D branch to inject 3D priors while enabling co-evolution of the 2D and 3D representations, alleviating the reliance on labels. Specifically, the RoboTAG consists of a 3D branch and a 2D branch, where nodes represent the states of the camera and robot system, and edges capture the dependencies between these variables or denote alignments between them. Closed loops are then defined in the graph, on which a consistency supervision across branches can be applied. Experimental results demonstrate that our method is effective across robot types, suggesting new possibilities of alleviating the data bottleneck in robotics.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on11
- Video Diffusion ModelsJonathan Ho, Tim Salimans, Alexey A. Gritsenko, William Chan et al.NeurIPS 2022 · 2,948 citations
- Depth Anything V2Lihe Yang, Bingyi Kang, Zilong Huang, Zhen Zhao et al.NeurIPS 2024 · 2,305 citations
- ScoreHOI: Physically Plausible Reconstruction of Human-Object Interaction via Score-Guided DiffusionAo Li, Jinpeng Liu, Yixuan Zhu, Yansong TangICCV 2025 · 1 citation
- End-to-End Learnable Geometric Vision by Backpropagating PnP OptimizationBo Chen, Álvaro Parra, Jiewei Cao, Nan Li et al.CVPR 2020
- Learning Human-to-Robot Handovers from Point CloudsSammy Joe Christen, Wei Yang, Claudia Pérez-D'Arpino, Otmar Hilliges et al.CVPR 2023
Related papers
- Markerless Camera-to-Robot Pose Estimation via Self-Supervised Sim-to-Real TransferJingpei Lu, Florian Richter, Michael C. YipCVPR 2023
- Robot Structure Prior Guided Temporal Attention for Camera-to-Robot Pose Estimation from Image SequenceYang Tian, Jiyao Zhang, Zekai Yin, Hao DongCVPR 2023
- CrossLoc: Scalable Aerial Localization Assisted by Multimodal Synthetic DataQi Yan, Jianhao Zheng, Simon Reding, Shanci Li et al.CVPR 2022 · 23 citations
- SE(3)-Equivariance with Geometric and Topological Guidance for Category-Level Object Pose EstimationSheng Yu, Di-Hua Zhai, Yuanqing XiaCVPR 2026
- Beyond Human Perception: Understanding Multi-Object World from Monocular ViewKeyu Guo, Yongle Huang, Shijie Sun, Xiangyu Song et al.CVPR 2025
