Geometry-Guided Diffusion Model with Masked Transformer for Robust Multi-View 3D Human Pose Estimation
Xinyi Zhang, Qinpeng Cui, Qiqi Bao, Wenming Yang, Qingmin Liao
Abstract
Recent research on Diffusion Models and Transformers has brought significant advancements to 3D Human Pose Estimation (HPE). Nonetheless, existing methods often fail to concurrently address the issues of accuracy and generalization. In this paper, we propose a Geometry-guided Dif fusion Model with Masked Transformer (Masked Gifformer) for robust multi-view 3D HPE. Within the framework of the diffusion model, a hierarchical multi-view trans-former-based denoiser is exploited to fit the 3D pose distribution by systematically integrating joint and view information. To address the long-standing problem of poor generalization, we introduce a fully random mask mechanism without any additional learnable modules or parameters. Furthermore, we incorporate geometric guidance into the diffusion model to enhance the accuracy of the model. This is achieved by optimizing the sampling process to minimize reprojection errors through modeling a conditional guidance distribution. Extensive experiments on two benchmarks demonstrate that Masked Gifformer effectively achieves a trade-off between accuracy and generalization. Specifically, our method outperforms other probabilistic methods by > 40% and achieves comparable results with state-of-the-art deterministic methods. In addition, our method exhibits robustness to varying camera numbers, spatial arrangements, and datasets.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 8a3b2b7b-8b98-430b-a7b7-ff036fe11d31Cited by top-tier papers1
Ask how each one uses itRelated papers
- Multiple View Geometry Transformers for 3D Human Pose EstimationZiwei Liao, Jialiang Zhu, Chunyu Wang, Han Hu et al.CVPR 2024
- Efficient Hierarchical Multi-view Fusion Transformer for 3D Human Pose EstimationKangkang Zhou, Lijun Zhang, Feng Lu, Xiang-Dong Zhou et al.ACM MM 2023 · 17 citations
- Diffusion-Based 3D Human Pose Estimation with Multi-Hypothesis AggregationWenkang Shan, Zhenhua Liu, Xinfeng Zhang, Zhao Wang et al.ICCV 2023 · 148 citations
- SVTformer: Spatial-View-Temporal Transformer for Multi-View 3D Human Pose EstimationWanruo Zhang, Mengyuan Liu, Hong Liu, Wenhao LiAAAI 2025 · 4 citations
- Deep Semantic Graph Transformer for Multi-View 3D Human Pose EstimationLijun Zhang, Kangkang Zhou, Feng Lu, Xiang-Dong Zhou et al.AAAI 2024 · 14 citations
