Improving Robustness and Accuracy via Relative Information Encoding in 3D Human Pose Estimation
Wenkang Shan, Haopeng Lu, Shanshe Wang, Xinfeng Zhang, Wen Gao
Abstract
Most of the existing 3D human pose estimation approaches mainly focus on predicting 3D positional relationships between the root joint and other human joints (local motion) instead of the overall trajectory of the human body (global motion). Despite the great progress achieved by these approaches, they are not robust to global motion, and lack the ability to accurately predict local motion with a small movement range. To alleviate these two problems, we propose a relative information encoding method that yields positional and temporal enhanced representations. Firstly, we encode positional information by utilizing relative coordinates of 2D poses to enhance the consistency between the input and output distribution. The same posture with different absolute 2D positions can be mapped to a common representation. It is beneficial to resist the interference of global motion on the prediction results. Second, we encode temporal information by establishing the connection between the current pose and other poses of the same person within a period of time. More attention will be paid to the movement changes before and after the current pose, resulting in better prediction performance on local motion with a small movement range. The ablation studies validate the effectiveness of the proposed relative information encoding method. Besides, we introduce a multi-stage optimization method to the whole framework to further exploit the positional and temporal enhanced representations. Our method outperforms state-of-the-art methods on two public datasets. Code is available at https://github.com/paTRICK-swk/Pose3D-RIE.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6455b824-4c15-4fde-b751-196db8d1e492Cited by top-tier papers11
- Diffusion-Based 3D Human Pose Estimation with Multi-Hypothesis AggregationWenkang Shan, Zhenhua Liu, Xinfeng Zhang, Zhao Wang et al.ICCV 2023 · 148 citations
- GLA-GCN: Global-local Adaptive Graph Convolutional Network for 3D Human Pose Estimation from Monocular VideoBruce X. B. Yu, Zhi Zhang, Yongxu Liu, Sheng-Hua Zhong et al.ICCV 2023 · 131 citations
- Ray3D: ray-based 3D human pose estimation for monocular absolute 3D localizationYu Zhan, Fenghai Li, Renliang Weng, Wongun ChoiCVPR 2022 · 62 citations
- CEE-Net: Complementary End-to-End Network for 3D Human Pose Generation and EstimationHaolun Li, Chi-Man PunAAAI 2023 · 46 citations
- FinePOSE: Fine-Grained Prompt-Driven 3D Human Pose Estimation via Diffusion ModelsJinglin Xu, Yijie Guo, Yuxin PengCVPR 2024 · 39 citations
Builds on9
- Exploiting Spatial-Temporal Relationships for 3D Pose Estimation via Graph Convolutional NetworksYujun Cai, Liuhao Ge, Jun Liu, Jianfei Cai et al.ICCV 2019 · 504 citations
- Optimizing Network Structure for 3D Human Pose EstimationHai Ci, Chunyu Wang, Xiaoxuan Ma, Yizhou WangICCV 2019 · 267 citations
- Occlusion-Aware Networks for 3D Human Pose Estimation in VideoYu Cheng, Bo Yang, Bo Wang, Wending Yan et al.ICCV 2019 · 223 citations
- Not All Parts Are Created Equal: 3D Pose Estimation by Modeling Bi-Directional Dependencies of Body PartsJue Wang, Shaoli Huang, Xinchao Wang, Dacheng TaoICCV 2019 · 66 citations
- Weakly-Supervised 3D Human Pose Learning via Multi-View Images in the WildUmar Iqbal, Pavlo Molchanov, Jan KautzCVPR 2020
Related papers
- Encoder-decoder with Multi-level Attention for 3D Human Shape and Pose EstimationZiniu Wan, Zhengjia Li, Maoqing Tian, Jianbo Liu et al.ICCV 2021 · 105 citations
- 3D Human Pose Estimation Using Spatio-Temporal Networks with Explicit Occlusion TrainingYu Cheng, Bo Yang, Bo Wang, Robby T. TanAAAI 2020 · 145 citations
- Perspose: 3D Human Pose Estimation with Perspective Encoding and Perspective RotationXiaoyang Hao, Han LiICCV 2025 · 3 citations
- Dynamic Graph Reasoning for Multi-person 3D Pose EstimationZhongwei Qiu, Qiansheng Yang, Jian Wang, Dongmei FuACM MM 2022 · 14 citations
- MixSTE: Seq2seq Mixed Spatio-Temporal Encoder for 3D Human Pose Estimation in VideoJinlu Zhang, Zhigang Tu, Jianyu Yang, Yujin Chen et al.CVPR 2022 · 356 citations
