G2L-Net: Global to Local Network for Real-Time 6D Pose Estimation With Embedding Vector Features
Wei Chen, Xi Jia, Hyung Jin Chang, Jinming Duan, Ales Leonardis
Abstract
In this paper, we propose a novel real-time 6D object pose estimation framework, named G2L-Net. Our network operates on point clouds from RGB-D detection in a divideand-conquer fashion. Specifically, our network consists of three steps. First, we extract the coarse object point cloud from the RGB-D image by 2D detection. Second, we feed the coarse object point cloud to a translation localization network to perform 3D segmentation and object translation prediction. Third, via the predicted segmentation and translation, we transfer the fine object point cloud into a local canonical coordinate, in which we train a rotation localization network to estimate initial object rotation. In the third step, we define point-wise embedding vector features to capture viewpoint-aware information. To calculate more accurate rotation, we adopt a rotation residual estimator to estimate the residual between initial rotation and ground truth, which can boost initial pose estimation performance. Our proposed G2L-Net is real-time despite the fact multiple steps are stacked via the proposed coarse-to-fine framework. Extensive experiments on two benchmark datasets show that G2L-Net achieves state-of-the-art performance in terms of both accuracy and speed. 1 Net.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 556dc043-b4c9-4966-871b-fb2ae49b53c1Cited by top-tier papers5
- OVE6D: Object Viewpoint Encoding for Depth-based 6D Object Pose EstimationDingding Cai, Janne Heikkilä, Esa RahtuCVPR 2022 · 62 citations
- CatFormer: Category-Level 6D Object Pose Estimation with TransformerSheng Yu, Di-Hua Zhai, Yuanqing XiaAAAI 2024 · 14 citations
- VinT-6D: A Large-Scale Object-in-hand Dataset from Vision, Touch and ProprioceptionZhaoliang Wan, Yonggen Ling, Senlin Yi, Lu Qi et al.ICML 2024 · 11 citations
- KeyPose: Category-Level 6D Object Pose Estimation with Self-Adaptive KeypointsSheng Yu, Di-Hua Zhai, Yuanqing XiaAAAI 2025 · 2 citations
- FS-Net: Fast Shape-Based Network for Category-Level 6D Object Pose Estimation With Decoupled Rotation MechanismWei Chen, Xi Jia, Hyung Jin Chang, Jinming Duan et al.CVPR 2021
Builds on1
Related papers
- DCNet: Dense Correspondence Neural Network for 6DoF Object Pose Estimation in Occluded ScenesZhi Chen, Wei Yang, Zhenbo Xu, Xike Xie et al.ACM MM 2020 · 3 citations
- DSC-PoseNet: Learning 6DoF Object Pose Estimation via Dual-Scale ConsistencyZongxin Yang, Xin Yu, Yi YangCVPR 2021
- CDPN: Coordinates-Based Disentangled Pose Network for Real-Time RGB-Based 6-DoF Object Pose EstimationZhigang Li, Gu Wang, Xiangyang JiICCV 2019 · 482 citations
- PVN3D: A Deep Point-Wise 3D Keypoints Voting Network for 6DoF Pose EstimationYisheng He, Wei Sun, Haibin Huang, Jianran Liu et al.CVPR 2020
- GPV-Pose: Category-level Object Pose Estimation via Geometry-guided Point-wise VotingYan Di, Ruida Zhang, Zhiqiang Lou, Fabian Manhardt et al.CVPR 2022 · 141 citations
