We Are More Than Our Joints: Predicting How 3D Bodies Move
Yan Zhang, Michael J. Black, Siyu Tang
Abstract
A key step towards understanding human behavior is the prediction of 3D human motion. Successful solutions have many applications in human tracking, HCI, and graphics. Most previous work focuses on predicting a time series of future 3D joint locations given a sequence 3D joints from the past. This Euclidean formulation generally works better than predicting pose in terms of joint rotations. Body joint locations, however, do not fully constrain 3D human pose, leaving degrees of freedom (like rotation about a limb) undefined. Note that 3D joints can be viewed as a sparse point cloud. Thus the problem of human motion prediction can be seen as a problem of point cloud prediction. With this observation, we instead predict a sparse set of locations on the body surface that correspond to motion capture markers. Given such markers, we fit a parametric body model to recover the 3D body of the person. These sparse surface markers also carry detailed information about human movement that is not present in the joints, increasing the naturalness of the predicted motions. Using the AMASS dataset, we train MOJO (More than Our JOints), which is a novel variational autoencoder with a latent DCT space that generates motions from latent frequencies. MOJO preserves the full temporal resolution of the input motion, and sampling from the latent frequencies explicitly introduces high-frequency components into the generated motion. We note that motion prediction methods accumulate errors over time, resulting in joints or markers that diverge from true human bodies. To address this, we fit the SMPL-X body model to the predictions at each time step, projecting the solution back onto the space of valid bodies, before propagating the new markers in time. Quantitative and qualitative experiments show that our approach produces state-of-the-art results and realistic 3D body animations. The code is available for research purposes at https://yz-cnsdqz.github.io/MOJO/MOJO.html.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c37c4a59-cee1-4440-b2bd-ec50732c8997Cited by top-tier papers63
- MotionGPT: Human Motion as a Foreign LanguageBiao Jiang, Xin Chen, Wen Liu, Jingyi Yu et al.NeurIPS 2023 · 698 citations
- Action-Conditioned 3D Human Motion Synthesis with Transformer VAEMathis Petrovich, Michael J. Black, Gül VarolICCV 2021 · 672 citations
- HuMoR: 3D Human Motion Model for Robust Pose EstimationDavis Rempe, Tolga Birdal, Aaron Hertzmann, Jimei Yang et al.ICCV 2021 · 398 citations
- Stochastic Scene-Aware Motion PredictionMohamed Hassan, Duygu Ceylan, Ruben Villegas, Jun Saito et al.ICCV 2021 · 240 citations
- HUMANISE: Language-conditioned Human Motion Generation in 3D ScenesZan Wang, Yixin Chen, Tengyu Liu, Yixin Zhu et al.NeurIPS 2022 · 207 citations
Builds on11
- AMASS: Archive of Motion Capture As Surface ShapesNaureen Mahmood, Nima Ghorbani, Nikolaus F. Troje, Gerard Pons-Moll et al.ICCV 2019 · 1,784 citations
- Learning Trajectory Dependencies for Human Motion PredictionWei Mao, Miaomiao Liu, Mathieu Salzmann, Hongdong LiICCV 2019 · 534 citations
- Character controllers using motion VAEsHung Yu Ling, Fabio Zinno, George Cheng, Michiel van de PanneSIGGRAPH 2020 · 261 citations
- Human Motion Prediction via Spatio-Temporal InpaintingAlejandro Hernandez Ruiz, Jürgen Gall, Francesc MorenoICCV 2019 · 233 citations
- Structured Prediction Helps 3D Human Motion ModellingEmre Aksan, Manuel Kaufmann, Otmar HilligesICCV 2019 · 204 citations
Related papers
- SOMA: Solving Optical Marker-Based MoCap AutomaticallyNima Ghorbani, Michael J. BlackICCV 2021 · 48 citations
- Generating 3D People in Scenes Without PeopleYan Zhang, Mohamed Hassan, Heiko Neumann, Michael J. Black et al.CVPR 2020
- Contextually Plausible and Diverse 3D Human Motion PredictionSadegh Aliakbarian, Fatemeh Sadat Saleh, Lars Petersson, Stephen Gould et al.ICCV 2021 · 44 citations
- A Unified 3D Human Motion Synthesis Model via Conditional Variational Auto-Encoder∗Yujun Cai, Yiwei Wang, Yiheng Zhu, Tat-Jen Cham et al.ICCV 2021 · 83 citations
- WANDR: Intention-guided Human Motion GenerationMarkos Diomataris, Nikos Athanasiou, Omid Taheri, Xi Wang et al.CVPR 2024
