Learning to Control Physically-simulated 3D Characters via Generating and Mimicking 2D Motions
Jianan Li, Xiao Chen, Tao Huang, Tien-Tsin Wong
摘要
Video data is more cost-effective than motion capture data for learning 3D character motion controllers, yet synthesizing realistic and diverse behaviors directly from videos remains challenging. Previous approaches typically rely on off-the-shelf motion reconstruction techniques to obtain 3D trajectories for physics-based imitation. These reconstruction methods struggle with generalizability, as they either require 3D training data (potentially scarce) or fail to produce physically plausible poses, hindering their application to challenging scenarios like human-object interaction (HOI) or non-human characters. We tackle this challenge by introducing Mimic2DM, a novel motion imitation framework that learns the control policy directly and solely from widely available 2D keypoint trajectories extracted from videos. By minimizing the reprojection error, we train a general single-view 2D motion tracking policy capable of following arbitrary 2D reference motions in physics simulation, using only 2D motion data. The policy, when trained on diverse 2D motions captured from different or slightly different viewpoints, can further acquire 3D motion tracking capabilities by aggregating multiple views. Moreover, we develop a transformer-based autoregressive 2D motion generator and integrate it into a hierarchical control framework, where the generator produces high-quality 2D reference trajectories to guide the tracking policy. We show that the proposed approach is versatile and can effectively learn to synthesize physically plausible and diverse motions across a range of domains, including dancing, soccer dribbling, and animal movements, without any reliance on explicit 3D motion data. Project Website: https://jiann-li.github.io/mimic2dm/
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper29
- AMASS: Archive of Motion Capture As Surface ShapesNaureen Mahmood, Nima Ghorbani, Nikolaus F. Troje, Gerard Pons-Moll 等ICCV 2019 · 被引用 1,784 次
- ViTPose: Simple Vision Transformer Baselines for Human Pose EstimationYufei Xu, Jing Zhang, Qiming Zhang, Dacheng TaoNeurIPS 2022 · 被引用 1,105 次
- AI Choreographer: Music Conditioned 3D Dance Generation with AIST++Ruilong Li, Shan Yang, David A. Ross, Angjoo KanazawaICCV 2021 · 被引用 701 次
- AMP: adversarial motion priors for stylized physics-based character controlXue Bin Peng, Ze Ma, Pieter Abbeel, Sergey Levine 等SIGGRAPH 2021 · 被引用 392 次
- Perpetual Humanoid Control for Real-time Simulated AvatarsZhengyi Luo, Jinkun Cao, Alexander Winkler, Kris Kitani 等ICCV 2023 · 被引用 256 次
相关 Paper
- Motion-2-To-3: Leveraging 2D Motion Data for 3D Motion GenerationsRuoxi Guo, Huaijin Pi, Zehong Shen, Qing Shuai 等ICCV 2025 · 被引用 2 次
- AnyLift: Scaling Motion Reconstruction from Internet Videos via 2D DiffusionHongjie Li, Heng Yu, Jiaman Li, Hong-Xing Yu 等CVPR 2026 · 被引用 2 次
- NIL: No-data Imitation LearningMert Albaba, Chenhao Li, Markos Diomataris, Omid Taheri 等CVPR 2026
- Mocap-2-to-3: Multi-view Lifting for Monocular Motion Recovery with 2D PretrainingZhumei Wang, Zechen Hu, Ruoxi Guo, Huaijin Pi 等CVPR 2026 · 被引用 1 次
- ASE: large-scale reusable adversarial skill embeddings for physically simulated charactersXue Bin Peng, Yunrong Guo, Lina Halper, Sergey Levine 等SIGGRAPH 2022 · 被引用 217 次
