Capturing Humans in Motion: Temporal-Attentive 3D Human Pose and Shape Estimation from Monocular Video
Wen-Li Wei, Jen-Chun Lin, Tyng-Luh Liu, Hong-Yuan Mark Liao
Abstract
Learning to capture human motion is essential to 3D human pose and shape estimation from monocular video. However, the existing methods mainly rely on recurrent or convolutional operation to model such temporal information, which limits the ability to capture non-local context relations of human motion. To address this problem, we propose a motion pose and shape network (MPS-Net) to effectively capture humans in motion to estimate accurate and temporally coherent 3D human pose and shape from a video. Specifically, we first propose a motion continuity attention (MoCA) module that leverages visual cues observed from human motion to adaptively recalibrate the range that needs attention in the sequence to better capture the motion continuity dependencies. Then, we develop a hierarchical attentive feature integration (HAFI) module to effectively combine adjacent past and future feature represen-tations to strengthen temporal correlation and refine the feature representation of the current frame. By coupling the MoCA and HAFI modules, the proposed MPS-Net excels in estimating 3D human pose and shape in the video. Though conceptually simple, our MPS-Net not only outperforms the state-of-the-art methods on the 3DPW, MPI-INF-3DHP, and Human3.6M benchmark datasets, but also uses fewer network parameters. The video demos can be found at https://mps-net.github.io/MPS-Net/.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 363e6bc4-078f-4ac1-a51f-c90a77ef7f90Cited by top-tier papers27
- MotionBERT: A Unified Perspective on Learning Human Motion RepresentationsWentao Zhu, Xiaoxuan Ma, Zhaoyang Liu, Libin Liu et al.ICCV 2023 · 322 citations
- GLA-GCN: Global-local Adaptive Graph Convolutional Network for 3D Human Pose Estimation from Monocular VideoBruce X. B. Yu, Zhi Zhang, Yongxu Liu, Sheng-Hua Zhong et al.ICCV 2023 · 131 citations
- WHAM: Reconstructing World-Grounded Humans with Accurate 3D MotionSoyong Shin, Juyong Kim, Eni Halilaj, Michael J. BlackCVPR 2024 · 66 citations
- TORE: Token Reduction for Efficient Human Mesh Recovery with TransformerZhiyang Dou, Qingxuan Wu, Cheng Lin, Zeyu Cao et al.ICCV 2023 · 56 citations
- Co-Evolution of Pose and Mesh for 3D Human Body Estimation from VideoYingxuan You, Hong Liu, Ti Wang, Wenhao Li et al.ICCV 2023 · 35 citations
Builds on9
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- AMASS: Archive of Motion Capture As Surface ShapesNaureen Mahmood, Nima Ghorbani, Nikolaus F. Troje, Gerard Pons-Moll et al.ICCV 2019 · 1,784 citations
- Learning to Reconstruct 3D Human Pose and Shape via Model-Fitting in the LoopNikos Kolotouros, Georgios Pavlakos, Michael J. Black, Kostas DaniilidisICCV 2019 · 1,139 citations
- PyMAF: 3D Human Pose and Shape Regression with Pyramidal Mesh Alignment Feedback LoopHongwen Zhang, Yating Tian, Xinchi Zhou, Wanli Ouyang et al.ICCV 2021 · 376 citations
- Robust motion in-betweeningFélix G. Harvey, Mike Yurick, Derek Nowrouzezahrai, Christopher J. PalSIGGRAPH 2020 · 269 citations
Related papers
- Encoder-decoder with Multi-level Attention for 3D Human Shape and Pose EstimationZiniu Wan, Zhengjia Li, Maoqing Tian, Jianbo Liu et al.ICCV 2021 · 105 citations
- Attention Mechanism Exploits Temporal Contexts: Real-Time 3D Human Pose ReconstructionRuixu Liu, Ju Shen, He Wang, Chen Chen et al.CVPR 2020
- MotionRefineNet: Fine-Grained Pose Sequence Smoothing and RefinementHaolun Li, Weihuang Liu, Jiateng Liu, Zhenhua Tang et al.ACM MM 2025 · 9 citations
- Human Mesh Recovery From Monocular Images via a Skeleton-Disentangled RepresentationYu Sun, Yun Ye, Wu Liu, Wenpeng Gao et al.ICCV 2019 · 196 citations
- MHFormer: Multi-Hypothesis Transformer for 3D Human Pose EstimationWenhao Li, Hong Liu, Hao Tang, Pichao Wang et al.CVPR 2022 · 403 citations
