DiffusionRegPose: Enhancing Multi-Person Pose Estimation Using a Diffusion-Based End-to-End Regression Approach
Dayi Tan, Hansheng Chen, Wei Tian, Lu Xiong
Abstract
This paper presents the DiffusionRegPose, a novel approach to multi-person pose estimation that converts a one-stage, end-to-end keypoint regression model into a diffusion-based sampling process. Existing one-stage deterministic re-gression methods, though efficient, are often prone to missed or false detections in crowded or occluded scenes, due to their inability to reason pose ambiguity. To address these challenges, we handle ambiguous poses in a generative fashion, i.e., sampling from the image-conditioned pose distributions characterized by a diffusion probabilistic model. Specifically, with initial pose tokens extracted from the image, noisy pose candidates are progressively refined by inter-acting with the initial tokens via attention layers. Extensive evaluations on the COCO and CrowdPose datasets show that DiffusionRegPose clearly improves the pose accuracy in crowded scenarios, as evidenced by a notable 4. 0 AP in-crease in the AP<inf xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">H</inf> metric on the CrowdPose dataset. This demonstrates the model's potential for robust and precise human pose estimation in real-world applications. Code will be available at https://github.com/cici203IDiffusionRegPose.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 55073a90-9e56-4d77-98f0-b846cf861421Cited by top-tier papers3
- Indoor Multi-View Radar Object Detection via 3D Bounding Box DiffusionRyoma Yataka, Pu Perry Wang, Petros Boufounos, Ryuhei TakahashiAAAI 2026 · 1 citation
- UDAPose: Unsupervised Domain Adaptation for Low-Light Human Pose EstimationHaopeng Chen, Yihao Ai, Kabeen Kim, Robby T. Tan et al.CVPR 2026 · 1 citation
- Attentive Keypoint Identification: Progressive Spatiotemporal Refinement for Video-based Human Pose EstimationSifan Wu, Haipeng Chen, Yingda Lyu, Shaojing Fan et al.AAAI 2026
Builds on32
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 6,759 citations
- Video Diffusion ModelsJonathan Ho, Tim Salimans, Alexey A. Gritsenko, William Chan et al.NeurIPS 2022 · 2,948 citations
- Score-Based Generative Modeling through Stochastic Differential EquationsYang Song, Jascha Sohl-Dickstein, Diederik P. Kingma, Abhishek Kumar et al.ICLR 2021 · 1,270 citations
Related papers
- Semantic-aware Transfer with Instance-adaptive Parsing for Crowded Scenes Pose EstimationXuanhan Wang, Lianli Gao, Yan Dai, Yixuan Zhou et al.ACM MM 2021 · 14 citations
- Bottom-Up Human Pose Estimation via Disentangled Keypoint RegressionZigang Geng, Ke Sun, Bin Xiao, Zhaoxiang Zhang et al.CVPR 2021
- DiffPose: Multi-hypothesis Human Pose Estimation using Diffusion ModelsKarl Holmquist, Bastian WandtICCV 2023 · 92 citations
- DiffPose: Toward More Reliable 3D Pose EstimationJia Gong, Lin Geng Foo, Zhipeng Fan, Qiuhong Ke et al.CVPR 2023
- QueryPose: Sparse Multi-Person Pose Regression via Spatial-Aware Part-Level QueryYabo Xiao, Kai Su, Xiaojuan Wang, Dongdong Yu et al.NeurIPS 2022 · 32 citations
