Neural Marionette: Unsupervised Learning of Motion Skeleton and Latent Dynamics from Volumetric Video
Jinseok Bae, Hojun Jang, Cheol-Hui Min, Hyungun Choi, Young Min Kim
Abstract
We present Neural Marionette, an unsupervised approach that discovers the skeletal structure from a dynamic sequence and learns to generate diverse motions that are consistent with the observed motion dynamics. Given a video stream of point cloud observation of an articulated body under arbitrary motion, our approach discovers the unknown low-dimensional skeletal relationship that can effectively represent the movement. Then the discovered structure is utilized to encode the motion priors of dynamic sequences in a latent structure, which can be decoded to the relative joint rotations to represent the full skeletal motion. Our approach works without any prior knowledge of the underlying motion or skeletal structure, and we demonstrate that the discovered structure is even comparable to the hand-labeled ground truth skeleton in representing a 4D sequence of motion. The skeletal structure embeds the general semantics of possible motion space that can generate motions for diverse scenarios. We verify that the learned motion prior is generalizable to the multi-modal sequence generation, interpolation of two poses, and motion retargeting to a different skeletal structure.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5bdd8c3f-9342-4b7a-b06f-226fbb3f9380Cited by top-tier papers3
- Dynamic Mesh Recovery from Partial Point Cloud SequenceHojun Jang, Minkwan Kim, Jinseok Bae, Young Min KimICCV 2023 · 5 citations
- Transfer4D: A Framework for Frugal Motion Capture and Deformation TransferShubh Maheshwari, Rahul Narain, Ramya HebbalaguppeCVPR 2023
- RigGS: Rigging of 3D Gaussians for Modeling Articulated Objects in VideosYuxin Yao, Zhi Deng, Junhui HouCVPR 2025
Builds on14
- Learning Trajectory Dependencies for Human Motion PredictionWei Mao, Miaomiao Liu, Mathieu Salzmann, Hongdong LiICCV 2019 · 534 citations
- FreiHAND: A Dataset for Markerless Capture of Hand Pose and Shape From Single RGB ImagesChristian Zimmermann, Duygu Ceylan, Jimei Yang, Bryan C. Russell et al.ICCV 2019 · 493 citations
- GENESIS: Generative Scene Inference and Sampling with Object-Centric Latent RepresentationsMartin Engelcke, Adam R. Kosiorek, Oiwi Parker Jones, Ingmar PosnerICLR 2020 · 334 citations
- Causal Discovery in Physical Systems from VideosYunzhu Li, Antonio Torralba, Anima Anandkumar, Dieter Fox et al.NeurIPS 2020 · 133 citations
- NPMs: Neural Parametric Models for 3D Deformable ShapesPablo R. Palafox, Aljaz Bozic, Justus Thies, Matthias Nießner et al.ICCV 2021 · 129 citations
Related papers
- DeepPhase: periodic autoencoders for learning motion phase manifoldsSebastian Starke, Ian Mason, Taku KomuraSIGGRAPH 2022 · 142 citations
- RigMo: Unifying Rig and Motion Learning for Generative AnimationHao Zhang, Jiahao Luo, Bohui Wan, Yizhou Zhao et al.CVPR 2026 · 6 citations
- Nonparametric Object and Parts Modeling With Lie Group DynamicsDavid S. Hayden, Jason Pacheco, John W. Fisher IIICVPR 2020
- Learning Diverse Stochastic Human-Action Generators by Learning Smooth Latent TransitionsZhenyi Wang, Ping Yu, Yang Zhao, Ruiyi Zhang et al.AAAI 2020 · 73 citations
- DIMO: Diverse 3D Motion Generation for Arbitrary ObjectsLinzhan Mou, Jiahui Lei, Chen Wang, Lingjie Liu et al.ICCV 2025 · 2 citations
