APP: Adaptive Pose Pooling for 3D Human Pose Estimation from Videos
Jinyan Zhang, Mengyuan Liu, Hong Liu, Guoquan Wang, Wenhao Li
Abstract
Current advancements in 3D human pose estimation have attained notable success by converting 2D poses into their 3D counterparts. However, this approach is inherently influenced by the errors introduced by 2D pose detectors and overlooks the intrinsic spatial information embedded within RGB images. To address these challenges, we introduce a versatile module called Adaptive Pose Pooling (APP), which is compatible with many existing 2D-to-3D lifting models. The APP module includes three novel sub-modules: Pose-Aware Offsets Generation (PAOG), Pose-Aware Sampling (PAS), and Spatial Temporal Information Fusion (STIF). First, we extract latent features of the multi-frame lifting model. Then, a 2D pose detector is utilized to extract multi-level feature maps from the image. After that, PAOG generates offsets according to featuremaps. PAS uses offsets to sample featuremaps. Then, STIF can fuse PAS sampling features and latent features. This innovative design allows the APP module to simultaneously capture spatial and temporal information. We conduct comprehensive experiments on two widely used datasets: Human3.6M and MPI-INF-3DHP. Meanwhile, we employ various lifting models to demonstrate the efficacy of the APP module. Our results show that the proposed APP module consistently enhances the performance of lifting models, achieving state-of-the-art results. Significantly, our module achieves these performance boosts without necessitating alterations to the architecture of the lifting model. Our code is available at https://github.com/jinyanzhang/APP.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 2eef3a79-7a03-48e8-80a4-f62397eca600Related papers
- Lifting by Image - Leveraging Image Cues for Accurate 3D Human Pose EstimationFeng Zhou, Jianqin Yin, Peiyang LiAAAI 2024 · 18 citations
- GLA-GCN: Global-local Adaptive Graph Convolutional Network for 3D Human Pose Estimation from Monocular VideoBruce X. B. Yu, Zhi Zhang, Yongxu Liu, Sheng-Hua Zhong et al.ICCV 2023 · 131 citations
- Towards Alleviating the Modeling Ambiguity of Unsupervised Monocular 3D Human Pose EstimationZhenbo Yu, Bingbing Ni, Jingwei Xu, Junjie Wang et al.ICCV 2021 · 39 citations
- PandaPose: 3D Human Pose Lifting from a Single Image via Propagating 2D Pose Prior to 3D Anchor SpaceJinghong Zheng, Changlong Jiang, Yang Xiao, Jiaqi Li et al.NeurIPS 2025 · 1 citation
- Geometry-Driven Self-Supervised Method for 3D Human Pose EstimationYang Li, Kan Li, Shuai Jiang, Ziyue Zhang et al.AAAI 2020 · 40 citations
