Uni6D: A Unified CNN Framework without Projection Breakdown for 6D Pose Estimation
Xiaoke Jiang, Donghai Li, Hao Chen, Ye Zheng, Rui Zhao, Liwei Wu
Abstract
As RGB-D sensors become more affordable, using RGB- D images to obtain high-accuracy 6D pose estimation results becomes a better option. State-of-the-art approaches typically use different backbones to extract features for RGB and depth images. They use a 2D CNN for RGB images and a perpixel point cloud network for depth data, as well as a fusion network for feature fusion. We find that the essential reason for using two independent backbones is the “projection breakdown” problem. In the depth image plane, the projected 3D structure of the physical world is preserved by the 1D depth value and its built-in 2D pixel coordinate (UV). Any spatial transformation that modifies UV, such as resize, flip, crop, or pooling operations in the CNN pipeline, breaks the binding between the pixel value and UV coordinate. As a consequence, the 3D structure is no longer preserved by a modified depth image or feature. To address this issue, we propose a simple yet effective method denoted as Uni6D that explicitly takes the extra UV data along with RGB-D images as input. Our method has a Unified CNN framework for 6D pose estimation with a single CNN backbone. In particular, the architecture of our method is based on Mask R-CNN with two extra heads, one named RT head for directly predicting 6D pose and the other named abc head for guiding the network to map the visible points to their coordinates in the 3D model as an auxiliary module. This end-to-end approach balances simplicity and accuracy, achieving comparable accuracy with state of the arts and 7.2x faster inference speed on the YCB-Video dataset.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 49845b84-3f97-4eb0-9b83-07778a92949eCited by top-tier papers5
- Deep Fusion Transformer Network with Weighted Vector-Wise Keypoints Voting for Robust 6D Object Pose EstimationJun Zhou, Kai Chen, Linlin Xu, Qi Dou et al.ICCV 2023 · 42 citations
- Learning Symmetry-Aware Geometry Correspondences for 6D Object Pose EstimationHeng Zhao, Shenxing Wei, Dahu Shi, Wenming Tan et al.ICCV 2023 · 33 citations
- RFMPose: Generative Category-level Object Pose Estimation via Riemannian Flow MatchingWenzhe Ouyang, Qi Ye, Jinghua Wang, Zenglin Xu et al.NeurIPS 2025 · 5 citations
- HS-Pose: Hybrid Scope Feature Extraction for Category-level Object Pose EstimationLinfang Zheng, Chen Wang, Yinghan Sun, Esha Dasgupta et al.CVPR 2023
- GIVEPose: Gradual Intra-class Variation Elimination for RGB-based Category-Level Object Pose EstimationZiqin Huang, Gu Wang, Chenyangguang Zhang, Ruida Zhang et al.CVPR 2025
Builds on9
- CDPN: Coordinates-Based Disentangled Pose Network for Real-Time RGB-Based 6-DoF Object Pose EstimationZhigang Li, Gu Wang, Xiangyang JiICCV 2019 · 482 citations
- MoreFusion: Multi-object Reasoning for 6D Pose Estimation from Volumetric FusionKentaro Wada, Edgar Sucar, Stephen James, Daniel Lenton et al.CVPR 2020
- PVN3D: A Deep Point-Wise 3D Keypoints Voting Network for 6DoF Pose EstimationYisheng He, Wei Sun, Haibin Huang, Jianran Liu et al.CVPR 2020
- Learning Canonical Shape Space for Category-Level 6D Object Pose and Size EstimationDengsheng Chen, Jun Li, Zheng Wang, Kai XuCVPR 2020
- FFB6D: A Full Flow Bidirectional Fusion Network for 6D Pose EstimationYisheng He, Haibin Huang, Haoqiang Fan, Qifeng Chen et al.CVPR 2021
Related papers
- DCNet: Dense Correspondence Neural Network for 6DoF Object Pose Estimation in Occluded ScenesZhi Chen, Wei Yang, Zhenbo Xu, Xike Xie et al.ACM MM 2020 · 3 citations
- PR-GCN: A Deep Graph Convolutional Network with Point Refinement for 6D Pose EstimationGuangyuan Zhou, Huiqun Wang, Jiaxin Chen, Di HuangICCV 2021 · 45 citations
- ES6D: A Computation Efficient and Symmetry-Aware 6D Pose Regression FrameworkNingkai Mo, Wanshui Gan, Naoto Yokoya, Shifeng ChenCVPR 2022 · 31 citations
- One2Any: One-Reference 6D Pose Estimation for Any ObjectMengya Liu, Siyuan Li, Ajad Chhatkuli, Prune Truong et al.CVPR 2025
- SO-Pose: Exploiting Self-Occlusion for Direct 6D Pose EstimationYan Di, Fabian Manhardt, Gu Wang, Xiangyang Ji et al.ICCV 2021 · 163 citations
