A Dual-Masked Auto-Encoder for Robust Motion Capture with Spatial-Temporal Skeletal Token Completion
Junkun Jiang, Jie Chen, Yike Guo
摘要
Multi-person motion capture can be challenging due to ambiguities caused by severe occlusion, fast body movement, and complex interactions. Existing frameworks build on 2D pose estimations and triangulate to 3D coordinates via reasoning the appearance, trajectory, and geometric consistencies among multi-camera observations. However, 2D joint detection is usually incomplete and with wrong identity assignments due to limited observation angle, which leads to noisy 3D triangulation results. To overcome this issue, we propose to explore the short-range autoregressive characteristics of skeletal motion using transformer. First, we propose an adaptive, identity-aware triangulation module to reconstruct 3D joints and identify the missing joints for each identity. To generate complete 3D skeletal motion, we then propose a Dual-Masked Auto-Encoder (D-MAE) which encodes the joint status with both skeletal-structural and temporal position encoding for trajectory completion. D-MAE's flexible masking and encoding mechanism enable arbitrary skeleton definitions to be conveniently deployed under the same framework. In order to demonstrate the proposed model's capability in dealing with severe data loss scenarios, we contribute a high-accuracy and challenging motion capture dataset of multi-person interactions with severe occlusion. Evaluations on both benchmark and our new dataset demonstrate the efficiency of our proposed model, as well as its advantage against the other state-of-the-art methods.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- A Unified Masked Autoencoder with Patchified Skeletons for Motion SynthesisEsteve Valls Mascaro, Hyemin Ahn, Dongheui LeeAAAI 2024 · 被引用 11 次
- Deep Compositional Phase Diffusion for Long Motion Sequence GenerationHo Yin Au, Jie Chen, Junkun Jiang, Jingyu XiangNeurIPS 2025 · 被引用 7 次
- MBTI: Masked Blending Transformers with Implicit Positional Encoding for Frame-rate Agnostic Motion EstimationJungwoo Huh, Yeseung Park, Seongjean Kim, Jungsu Kim 等ICCV 2025
它引用的顶会 Paper10
- Fourier Features Let Networks Learn High Frequency Functions in Low Dimensional DomainsMatthew Tancik, Pratul P. Srinivasan, Ben Mildenhall, Sara Fridovich-Keil 等NeurIPS 2020 · 被引用 4,036 次
- 3D Human Pose Estimation with Spatial and Temporal TransformersCe Zheng, Sijie Zhu, Matías Mendieta, Taojiannan Yang 等ICCV 2021 · 被引用 648 次
- MHFormer: Multi-Hypothesis Transformer for 3D Human Pose EstimationWenhao Li, Hong Liu, Hao Tang, Pichao Wang 等CVPR 2022 · 被引用 403 次
- Robust motion in-betweeningFélix G. Harvey, Mike Yurick, Derek Nowrouzezahrai, Christopher J. PalSIGGRAPH 2020 · 被引用 269 次
- TesseTrack: End-to-End Learnable Multi-Person Articulated 3D Pose TrackingN. Dinesh Reddy, Laurent Guigues, Leonid Pishchulin, Jayan Eledath 等CVPR 2021
相关 Paper
- Dynamic Mesh Recovery from Partial Point Cloud SequenceHojun Jang, Minkwan Kim, Jinseok Bae, Young Min KimICCV 2023 · 被引用 5 次
- Masked Motion Predictors are Strong 3D Action Representation LearnersYunyao Mao, Jiajun Deng, Wengang Zhou, Yao Fang 等ICCV 2023 · 被引用 73 次
- Auxiliary Tasks Benefit 3D Skeleton-based Human Motion PredictionChenxin Xu, Robby T. Tan, Yuhong Tan, Siheng Chen 等ICCV 2023 · 被引用 35 次
- MotionBERT: A Unified Perspective on Learning Human Motion RepresentationsWentao Zhu, Xiaoxuan Ma, Zhaoyang Liu, Libin Liu 等ICCV 2023 · 被引用 322 次
- Capturing Closely Interacted Two-Person Motions with Reaction PriorsQi Fang, Yinghui Fan, Yanjun Li, Junting Dong 等CVPR 2024
