CoordAR: One-Reference 6D Pose Estimation of Novel Objects via Autoregressive Coordinate Map Generation
Dexin Zuo, Ang Li, Wei Wang, Wenxian Yu, Danping Zou
Abstract
Object 6D pose estimation, a crucial task for robotics and augmented reality applications, becomes particularly challenging when dealing with novel objects whose 3D models are not readily available. To reduce dependency on 3D models, recent studies have explored one-reference-based pose estimation, which requires only a single reference view instead of a complete 3D model. However, existing methods that rely on real-valued coordinate regression suffer from limited global consistency due to the local nature of convolutional architectures and face challenges in symmetric or occluded scenarios owing to a lack of uncertainty modeling. We present CoordAR, a novel autoregressive framework for one-reference 6D pose estimation of unseen objects. CoordAR formulates 3D-3D correspondences between the reference and query views as a map of discrete tokens, which is obtained in an autoregressive and probabilistic manner. To enable accurate correspondence regression, CoordAR introduces 1) a novel coordinate map tokenization that enables probabilistic prediction over discretized 3D space; 2) a modality-decoupled encoding strategy that separately encodes RGB appearance and coordinate cues; and 3) an autoregressive transformer decoder conditioned on both position-aligned query features and the partially generated token sequence. With these novel mechanisms, CoordAR significantly outperforms existing methods on multiple benchmarks and demonstrates strong robustness to symmetry, occlusion, and other challenges in real-world tests.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 039716b6-e331-414d-b286-702b4452c0e4Builds on23
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Visual Autoregressive Modeling: Scalable Image Generation via Next-Scale PredictionKeyu Tian, Yi Jiang, Zehuan Yuan, Bingyue Peng et al.NeurIPS 2024 · 1,199 citations
- Autoregressive Image Generation without Vector QuantizationTianhong Li, Yonglong Tian, He Li, Mingyang Deng et al.NeurIPS 2024 · 758 citations
- CDPN: Coordinates-Based Disentangled Pose Network for Real-Time RGB-Based 6-DoF Object Pose EstimationZhigang Li, Gu Wang, Xiangyang JiICCV 2019 · 482 citations
- FoundationPose: Unified 6D Pose Estimation and Tracking of Novel ObjectsBowen Wen, Wei Yang, Jan Kautz, Stan BirchfieldCVPR 2024 · 215 citations
Related papers
- One2Any: One-Reference 6D Pose Estimation for Any ObjectMengya Liu, Siyuan Li, Ajad Chhatkuli, Prune Truong et al.CVPR 2025
- Pix2Pose: Pixel-Wise Coordinate Regression of Objects for 6D Pose EstimationKiru Park, Timothy Patten, Markus VinczeICCV 2019 · 527 citations
- 3D-Aware Hypothesis & Verification for Generalizable Relative Object Pose EstimationChen Zhao, Tong Zhang, Mathieu SalzmannICLR 2024 · 13 citations
- RefPose: Leveraging Reference Geometric Correspondences for Accurate 6D Pose Estimation of Unseen ObjectsJaeguk Kim, Jaewoo Park, Keuntek Lee, Nam Ik ChoCVPR 2025
- AutoRF: Learning 3D Object Radiance Fields from Single View ObservationsNorman Müller, Andrea Simonelli, Lorenzo Porzi, Samuel Rota Bulò et al.CVPR 2022 · 45 citations
