DiffPose: SpatioTemporal Diffusion Model for Video-Based Human Pose Estimation
Runyang Feng, Yixing Gao, Tze Ho Elden Tse, Xueqing Ma, Hyung Jin Chang
Abstract
Denoising diffusion probabilistic models that were initially proposed for realistic image generation have recently shown success in various perception tasks (e.g., object detection and image segmentation) and are increasingly gaining attention in computer vision. However, extending such models to multi-frame human pose estimation is non-trivial due to the presence of the additional temporal dimension in videos. More importantly, learning representations that focus on keypoint regions is crucial for accurate localization of human joints. Nevertheless, the adaptation of the diffusion-based methods remains unclear on how to achieve such objective. In this paper, we present DiffPose, a novel diffusion architecture that formulates video-based human pose estimation as a conditional heatmap generation problem. First, to better leverage temporal information, we propose SpatioTemporal Representation Learner which aggregates visual evidences across frames and uses the resulting features in each denoising step as a condition. In addition, we present a mechanism called Lookup-based Multi-Scale Feature Interaction that determines the correlations between local joints and global contexts across multiple scales. This mechanism generates delicate representations that focus on keypoint regions. Altogether, by extending diffusion models, we show two unique characteristics from DiffPose on pose estimation task: (i) the ability to combine multiple sets of pose estimates to improve prediction accuracy, particularly for challenging joints, and (ii) the ability to adjust the number of iterative steps for feature refinement without retraining the model. DiffPose sets new state-of-the-art results on three benchmarks: PoseTrack2017, PoseTrack2018, and PoseTrack21.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext cee9a4c8-7277-4742-9cd7-d6d323fc43d0Cited by top-tier papers28
- MonoDiff: Monocular 3D Object Detection and Pose Estimation with Diffusion ModelsYasiru Ranasinghe, Deepti Hegde, Vishal M. PatelCVPR 2024 · 21 citations
- DPMesh: Exploiting Diffusion Prior for Occluded Human Mesh RecoveryYixuan Zhu, Ao Li, Yansong Tang, Wenliang Zhao et al.CVPR 2024 · 10 citations
- SynFER: Towards Boosting Facial Expression Recognition With Synthetic DataXilin He, Cheng Luo, Xiaole Xian, Bing Li et al.ICCV 2025 · 6 citations
- FIP: Endowing Robust Motion Capture on Daily Garment by Fusing Flex and Inertial SensorsRuonan Zheng, Jiawei Fang, Yuan Yao, Xiaoxia Gao et al.CHI 2025 · 5 citations
- SpatioTemporal Learning for Human Pose Estimation in Sparsely-Labeled VideosYingying Jiao, Zhigang Wang, Sifan Wu, Shaojing Fan et al.AAAI 2025 · 5 citations
Builds on32
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
Related papers
- DiffusionPose: Markov-Optimized Diffusion Model for Human Pose EstimationZhigang Wang, Zhenguang Liu, Shaojing Fan, Sifan Wu et al.AAAI 2026
- Attentive Keypoint Identification: Progressive Spatiotemporal Refinement for Video-based Human Pose EstimationSifan Wu, Haipeng Chen, Yingda Lyu, Shaojing Fan et al.AAAI 2026
- DiffPose: Multi-hypothesis Human Pose Estimation using Diffusion ModelsKarl Holmquist, Bastian WandtICCV 2023 · 92 citations
- Deep Dual Consecutive Network for Human Pose EstimationZhenguang Liu, Haoming Chen, Runyang Feng, Shuang Wu et al.CVPR 2021
- Video-Based Human Pose Regression via Decoupled Space-Time AggregationJijie He, Wenwu YangCVPR 2024
