Efficient Hierarchical Multi-view Fusion Transformer for 3D Human Pose Estimation
Kangkang Zhou, Lijun Zhang, Feng Lu, Xiang-Dong Zhou, Yu Shi
Abstract
In multi-view 3D human pose estimation (HPE), information from different viewpoints is highly variable due to complex factors such as background and occlusion, making cross-view feature extrac tion and fusion difficult. Most existing methods have problems of over-reliance on camera parameters or insufficient semantic feature extraction. To address these issues, this paper proposes a hierar chical multi-view fusion transformer (HMVformer) framework for 3D HPE, incorporating cross-view feature fusion methods into the spatial and temporal feature extraction process in a coarse-to-fine manner. To begin, global to local attention graph features are ex tracted and incorporated with the original pose features to better preserve the spatial structure semantic knowledge. Then, various cross-view feature fusion modules are built and embedded into the pose feature extraction for consistent and distinctive information fusion across multiple viewpoints. Furthermore, sequential tem poral information is extracted and fused with spatial knowledge for feature refinement and depth uncertainty reduction. Extensive experiments on three popular 3D HPE benchmarks show that HMV former achieves state-of-the-art results without relying on complex loss functions or providing camera parameters, simple but effective in mitigating depth ambiguity and improving 3D pose prediction accuracy. Codes and models are available1.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get b0fae4ff-eaf8-452f-87e3-1cdb23c13843Cited by top-tier papers5
- SuperVLAD: Compact and Robust Image Descriptors for Visual Place RecognitionFeng Lu, Xinyao Zhang, Canming Ye, Shuting Dong et al.NeurIPS 2024 · 24 citations
- Deep Semantic Graph Transformer for Multi-View 3D Human Pose EstimationLijun Zhang, Kangkang Zhou, Feng Lu, Xiang-Dong Zhou et al.AAAI 2024 · 14 citations
- SVTformer: Spatial-View-Temporal Transformer for Multi-View 3D Human Pose EstimationWanruo Zhang, Mengyuan Liu, Hong Liu, Wenhao LiAAAI 2025 · 4 citations
- Towards Practical Human Motion Prediction with LiDAR Point CloudsXiao Han, Yiming Ren, Yichen Yao, Yujing Sun et al.ACM MM 2024 · 2 citations
- From Sparse to Dense: Spatio-Temporal Fusion for Multi-View 3D Human Pose Estimation with DenseWarperLing Li, Changjie Chen, Yuyan Wang, Jiaqing Lyu et al.ICLR 2026
Related papers
- FusionFormer: A Concise Unified Feature Fusion Transformer for 3D Pose EstimationYanlu Cai, Weizhong Zhang, Yuan Wu, Cheng JinAAAI 2024 · 23 citations
- Multiple View Geometry Transformers for 3D Human Pose EstimationZiwei Liao, Jialiang Zhu, Chunyu Wang, Han Hu et al.CVPR 2024
- MHFormer: Multi-Hypothesis Transformer for 3D Human Pose EstimationWenhao Li, Hong Liu, Hao Tang, Pichao Wang et al.CVPR 2022 · 403 citations
- Geometry-Guided Diffusion Model with Masked Transformer for Robust Multi-View 3D Human Pose EstimationXinyi Zhang, Qinpeng Cui, Qiqi Bao, Wenming Yang et al.ACM MM 2024 · 3 citations
- 3D Human Pose Estimation with Spatial and Temporal TransformersCe Zheng, Sijie Zhu, Matías Mendieta, Taojiannan Yang et al.ICCV 2021 · 648 citations
