A Dual-Masked Auto-Encoder for Robust Motion Capture with Spatial-Temporal Skeletal Token Completion
Junkun Jiang, Jie Chen, Yike Guo
Abstract
Multi-person motion capture can be challenging due to ambiguities caused by severe occlusion, fast body movement, and complex interactions. Existing frameworks build on 2D pose estimations and triangulate to 3D coordinates via reasoning the appearance, trajectory, and geometric consistencies among multi-camera observations. However, 2D joint detection is usually incomplete and with wrong identity assignments due to limited observation angle, which leads to noisy 3D triangulation results. To overcome this issue, we propose to explore the short-range autoregressive characteristics of skeletal motion using transformer. First, we propose an adaptive, identity-aware triangulation module to reconstruct 3D joints and identify the missing joints for each identity. To generate complete 3D skeletal motion, we then propose a Dual-Masked Auto-Encoder (D-MAE) which encodes the joint status with both skeletal-structural and temporal position encoding for trajectory completion. D-MAE's flexible masking and encoding mechanism enable arbitrary skeleton definitions to be conveniently deployed under the same framework. In order to demonstrate the proposed model's capability in dealing with severe data loss scenarios, we contribute a high-accuracy and challenging motion capture dataset of multi-person interactions with severe occlusion. Evaluations on both benchmark and our new dataset demonstrate the efficiency of our proposed model, as well as its advantage against the other state-of-the-art methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 11a1d27f-159f-4d12-a441-3664c3a6d76aCited by top-tier papers3
- A Unified Masked Autoencoder with Patchified Skeletons for Motion SynthesisEsteve Valls Mascaro, Hyemin Ahn, Dongheui LeeAAAI 2024 · 11 citations
- Deep Compositional Phase Diffusion for Long Motion Sequence GenerationHo Yin Au, Jie Chen, Junkun Jiang, Jingyu XiangNeurIPS 2025 · 7 citations
- MBTI: Masked Blending Transformers with Implicit Positional Encoding for Frame-rate Agnostic Motion EstimationJungwoo Huh, Yeseung Park, Seongjean Kim, Jungsu Kim et al.ICCV 2025
Builds on10
- Fourier Features Let Networks Learn High Frequency Functions in Low Dimensional DomainsMatthew Tancik, Pratul P. Srinivasan, Ben Mildenhall, Sara Fridovich-Keil et al.NeurIPS 2020 · 4,036 citations
- 3D Human Pose Estimation with Spatial and Temporal TransformersCe Zheng, Sijie Zhu, Matías Mendieta, Taojiannan Yang et al.ICCV 2021 · 648 citations
- MHFormer: Multi-Hypothesis Transformer for 3D Human Pose EstimationWenhao Li, Hong Liu, Hao Tang, Pichao Wang et al.CVPR 2022 · 403 citations
- Robust motion in-betweeningFélix G. Harvey, Mike Yurick, Derek Nowrouzezahrai, Christopher J. PalSIGGRAPH 2020 · 269 citations
- TesseTrack: End-to-End Learnable Multi-Person Articulated 3D Pose TrackingN. Dinesh Reddy, Laurent Guigues, Leonid Pishchulin, Jayan Eledath et al.CVPR 2021
Related papers
- Dynamic Mesh Recovery from Partial Point Cloud SequenceHojun Jang, Minkwan Kim, Jinseok Bae, Young Min KimICCV 2023 · 5 citations
- Masked Motion Predictors are Strong 3D Action Representation LearnersYunyao Mao, Jiajun Deng, Wengang Zhou, Yao Fang et al.ICCV 2023 · 73 citations
- Auxiliary Tasks Benefit 3D Skeleton-based Human Motion PredictionChenxin Xu, Robby T. Tan, Yuhong Tan, Siheng Chen et al.ICCV 2023 · 35 citations
- MotionBERT: A Unified Perspective on Learning Human Motion RepresentationsWentao Zhu, Xiaoxuan Ma, Zhaoyang Liu, Libin Liu et al.ICCV 2023 · 322 citations
- Capturing Closely Interacted Two-Person Motions with Reaction PriorsQi Fang, Yinghui Fan, Yanjun Li, Junting Dong et al.CVPR 2024
