Generative Category-level Object Pose Estimation via Diffusion Models
Jiyao Zhang, Mingdong Wu, Hao Dong
Abstract
Object pose estimation plays a vital role in embodied AI and computer vision, enabling intelligent agents to comprehend and interact with their surroundings. Despite the practicality of category-level pose estimation, current approaches encounter challenges with partially observed point clouds, known as the multihypothesis issue. In this study, we propose a novel solution by reframing categorylevel object pose estimation as conditional generative modeling, departing from traditional point-to-point regression. Leveraging score-based diffusion models, we estimate object poses by sampling candidates from the diffusion model and aggregating them through a two-step process: filtering out outliers via likelihood estimation and subsequently mean-pooling the remaining candidates. To avoid the costly integration process when estimating the likelihood, we introduce an alternative method that trains an energy-based model from the original scorebased model, enabling end-to-end likelihood estimation. Our approach achieves state-of-the-art performance on the REAL275 dataset and demonstrates promising generalizability to novel categories sharing similar symmetric properties without fine-tuning. Furthermore, it can readily adapt to object pose tracking tasks, yielding comparable results to the current state-of-the-art baselines. Our checkpoints and demonstrations can be found at https://sites.google.com/view/genpose .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 08d85f0d-40ee-412c-8925-49536ebdcf12Cited by top-tier papers18
- DiffCAD: Weakly-Supervised Probabilistic CAD Model Retrieval and Alignment from an RGB ImageDaoyi Gao, Dávid Rozenberszki, Stefan Leutenegger, Angela DaiSIGGRAPH 2024 · 28 citations
- SceneMaker: Open-set 3D Scene Generation with Decoupled De-occlusion and Pose Estimation ModelYukai Shi, Weiyu Li, Zihao Wang, Hongyang Li et al.CVPR 2026 · 17 citations
- Rectified Point Flow: Generic Point Cloud Pose EstimationTao Sun, Liyuan Zhu, Shengyu Huang, Shuran Song et al.NeurIPS 2025 · 14 citations
- Learning 3D Object Spatial Relationships From Pre-Trained 2D Diffusion ModelsSangwon Baik, Hyeonwoo Kim, Hanbyul JooICCV 2025 · 7 citations
- CADGrasp: Learning Contact and Collision Aware General Dexterous Grasping in Cluttered ScenesJiyao Zhang, Zhiyuan Ma, Tianhao Wu, Zeyuan Chen et al.NeurIPS 2025 · 7 citations
Builds on17
- Elucidating the Design Space of Diffusion-Based Generative ModelsTero Karras, Miika Aittala, Timo Aila, Samuli LaineNeurIPS 2022 · 3,959 citations
- Improved Techniques for Training Score-Based Generative ModelsYang Song, Stefano ErmonNeurIPS 2020 · 1,527 citations
- Score-Based Generative Modeling through Stochastic Differential EquationsYang Song, Jascha Sohl-Dickstein, Diederik P. Kingma, Abhishek Kumar et al.ICLR 2021 · 1,270 citations
- Maximum Likelihood Training of Score-Based Diffusion ModelsYang Song, Conor Durkan, Iain Murray, Stefano ErmonNeurIPS 2021 · 958 citations
- Solving Inverse Problems in Medical Imaging with Score-Based Generative ModelsYang Song, Liyue Shen, Lei Xing, Stefano ErmonICLR 2022 · 721 citations
Related papers
- Joint Learning of Pose Regression and Denoising Diffusion with Score Scaling Sampling for Category-Level 6D Pose EstimationSeunghyun Lee, Tae-Kyun KimICCV 2025 · 2 citations
- DICArt: Advancing Category-level Articulated Object Pose Estimation in Discrete State-SpacesLi Zhang, Mingyu Mei, Ailing Wang, Xianhui Meng et al.CVPR 2026 · 2 citations
- SE(3)-Equivariance with Geometric and Topological Guidance for Category-Level Object Pose EstimationSheng Yu, Di-Hua Zhai, Yuanqing XiaCVPR 2026
- DiffPose: Multi-hypothesis Human Pose Estimation using Diffusion ModelsKarl Holmquist, Bastian WandtICCV 2023 · 92 citations
- Leveraging SE(3) Equivariance for Self-supervised Category-Level Object Pose Estimation from Point CloudsXiaolong Li, Yijia Weng, Li Yi, Leonidas J. Guibas et al.NeurIPS 2021 · 61 citations
