RoboTAG: End-to-end Robot Pose Estimation via Topological Alignment Graph
Yifan Liu, Fangneng Zhan, Wanhua Li, Haowen Sun, Katerina Fragkiadaki, Hanspeter Pfister
摘要
Estimating robot pose from a monocular RGB image is a challenge in robotics and computer vision. Existing methods typically build networks on top of 2D visual backbones and depend heavily on labeled data for training, which is often scarce in real-world scenarios, causing a sim-to-real gap. Moreover, these approaches reduce the 3D-based problem to 2D domain, neglecting the 3D priors. To address these, we propose Robot Topological Alignment Graph (RoboTAG), which incorporates a 3D branch to inject 3D priors while enabling co-evolution of the 2D and 3D representations, alleviating the reliance on labels. Specifically, the RoboTAG consists of a 3D branch and a 2D branch, where nodes represent the states of the camera and robot system, and edges capture the dependencies between these variables or denote alignments between them. Closed loops are then defined in the graph, on which a consistency supervision across branches can be applied. Experimental results demonstrate that our method is effective across robot types, suggesting new possibilities of alleviating the data bottleneck in robotics.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper11
- Video Diffusion ModelsJonathan Ho, Tim Salimans, Alexey A. Gritsenko, William Chan 等NeurIPS 2022 · 被引用 2,948 次
- Depth Anything V2Lihe Yang, Bingyi Kang, Zilong Huang, Zhen Zhao 等NeurIPS 2024 · 被引用 2,305 次
- ScoreHOI: Physically Plausible Reconstruction of Human-Object Interaction via Score-Guided DiffusionAo Li, Jinpeng Liu, Yixuan Zhu, Yansong TangICCV 2025 · 被引用 1 次
- End-to-End Learnable Geometric Vision by Backpropagating PnP OptimizationBo Chen, Álvaro Parra, Jiewei Cao, Nan Li 等CVPR 2020
- Learning Human-to-Robot Handovers from Point CloudsSammy Joe Christen, Wei Yang, Claudia Pérez-D'Arpino, Otmar Hilliges 等CVPR 2023
相关 Paper
- Markerless Camera-to-Robot Pose Estimation via Self-Supervised Sim-to-Real TransferJingpei Lu, Florian Richter, Michael C. YipCVPR 2023
- Robot Structure Prior Guided Temporal Attention for Camera-to-Robot Pose Estimation from Image SequenceYang Tian, Jiyao Zhang, Zekai Yin, Hao DongCVPR 2023
- CrossLoc: Scalable Aerial Localization Assisted by Multimodal Synthetic DataQi Yan, Jianhao Zheng, Simon Reding, Shanci Li 等CVPR 2022 · 被引用 23 次
- SE(3)-Equivariance with Geometric and Topological Guidance for Category-Level Object Pose EstimationSheng Yu, Di-Hua Zhai, Yuanqing XiaCVPR 2026
- Beyond Human Perception: Understanding Multi-Object World from Monocular ViewKeyu Guo, Yongle Huang, Shijie Sun, Xiangyu Song 等CVPR 2025
