CatFormer: Category-Level 6D Object Pose Estimation with Transformer
Sheng Yu, Di-Hua Zhai, Yuanqing Xia
Abstract
Although there has been significant progress in category-level object pose estimation in recent years, there is still considerable room for improvement. In this paper, we propose a novel transformer-based category-level 6D pose estimation method called CatFormer to enhance the accuracy pose estimation. CatFormer comprises three main parts: a coarse deformation part, a fine deformation part, and a recurrent refinement part. In the coarse and fine deformation sections, we introduce a transformer-based deformation module that performs point cloud deformation and completion in the feature space. Additionally, after each deformation, we incorporate a transformer-based graph module to adjust fused features and establish geometric and topological relationships between points based on these features. Furthermore, we present an end-to-end recurrent refinement module that enables the prior point cloud to deform multiple times according to real scene features. We evaluate CatFormer's performance by training and testing it on CAMERA25 and REAL275 datasets. Experimental results demonstrate that CatFormer surpasses state-of-the-art methods. Moreover, we extend the usage of CatFormer to instance-level object pose estimation on the LINEMOD dataset, as well as object pose estimation in real-world scenarios. The experimental results validate the effectiveness and generalization capabilities of Cat-Former. Our code and the supplemental materials are avaliable at https://github.com/BIT-robot-group/CatFormer.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 275dc527-be0d-486a-a901-07dcac01c24aCited by top-tier papers2
- KeyPose: Category-Level 6D Object Pose Estimation with Self-Adaptive KeypointsSheng Yu, Di-Hua Zhai, Yuanqing XiaAAAI 2025 · 2 citations
- SE(3)-Equivariance with Geometric and Topological Guidance for Category-Level Object Pose EstimationSheng Yu, Di-Hua Zhai, Yuanqing XiaCVPR 2026
Builds on19
- Pix2Pose: Pixel-Wise Coordinate Regression of Objects for 6D Pose EstimationKiru Park, Timothy Patten, Markus VinczeICCV 2019 · 527 citations
- DPOD: 6D Pose Object Detector and RefinerSergey Zakharov, Ivan Shugurov, Slobodan IlicICCV 2019 · 486 citations
- CDPN: Coordinates-Based Disentangled Pose Network for Real-Time RGB-Based 6-DoF Object Pose EstimationZhigang Li, Gu Wang, Xiangyang JiICCV 2019 · 482 citations
- SGPA: Structure-Guided Prior Adaptation for Category-Level 6D Object Pose EstimationKai Chen, Qi DouICCV 2021 · 183 citations
- DualPoseNet: Category-level 6D Object Pose and Size Estimation Using Dual Pose Network with Refined Learning of Pose ConsistencyJiehong Lin, Zewei Wei, Zhihao Li, Songcen Xu et al.ICCV 2021 · 169 citations
Related papers
- 3D Object Detection With PointformerXuran Pan, Zhuofan Xia, Shiji Song, Li Erran Li et al.CVPR 2021
- GeoReF: Geometric Alignment Across Shape Variation for Category-level Object Pose RefinementLinfang Zheng, Tze Ho Elden Tse, Chen Wang, Yinghan Sun et al.CVPR 2024 · 6 citations
- Leveraging SE(3) Equivariance for Self-supervised Category-Level Object Pose Estimation from Point CloudsXiaolong Li, Yijia Weng, Li Yi, Leonidas J. Guibas et al.NeurIPS 2021 · 61 citations
- IST-Net: Prior-free Category-level Pose Estimation with Implicit Space TransformationJianhui Liu, Yukang Chen, Xiaoqing Ye, Xiaojuan QiICCV 2023 · 64 citations
- Query6DoF: Learning Sparse Queries as Implicit Shape Prior for Category-Level 6DoF Pose EstimationRuiqi Wang, Xinggang Wang, Te Li, Rong Yang et al.ICCV 2023 · 31 citations
