MMG-VL: A Vision-Language Driven Approach for Multi-Person Motion Generation
Songyuan Yang, Wanrong Huang, Yinuo Liu, Kedi Zhang, Xihuai He, Shaowu Yang, Huibin Tan
摘要
Generating realistic and coordinated 3D human motion for multiple individuals within complex environments remains a significant challenge. Existing text-to-motion methods are often "blind" to the physical scene, leading to implausible motions, while scene-conditioned (HSI) approaches demand cumbersome full 3D data and largely neglect multiperson dynamics. To address these limitations, we introduce the VL2Motion paradigm and its embodiment, MMG-VL, a hierarchical framework that generates coordinated multiperson motions from the most accessible inputs: a single 2D image and natural language. MMG-VL first employs a Scene-Aware Intent Planner (SAIP) to interpret the visual context and decompose the user's command into a set of spatially-grounded, multi-person action blueprints. Subsequently, a Coordinated Motion Synthesizer (CMS) translates these blueprints into high-fidelity 3D motion sequences. The synergy between these stages is driven by two novel loss functions: a Spatial-Semantic Grounding Loss (LSSG) to ensure the planner's output is grounded in visual reality, and a Coordinated Environmental Realism Loss (LCER) that enforces physical constraints and coherent group dynamics during synthesis. To facilitate this research, we introduce Hu-manVL, the first large-scale dataset featuring multi-person activities in multi-room scenes, providing aligned images, text, blueprints, 3D motions, and scene geometry. Extensive experiments demonstrate that MMG-VL significantly outperforms existing methods in generating spatially coherent, physically realistic, and coordinated multi-person motions, paving the way for more scalable and intuitive creation of dynamic virtual worlds.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper13
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- AMASS: Archive of Motion Capture As Surface ShapesNaureen Mahmood, Nima Ghorbani, Nikolaus F. Troje, Gerard Pons-Moll 等ICCV 2019 · 被引用 1,784 次
- Generating Diverse and Natural 3D Human Motions from TextChuan Guo, Shihao Zou, Xinxin Zuo, Sen Wang 等CVPR 2022 · 被引用 462 次
- Action2Motion: Conditioned Generation of 3D Human MotionsChuan Guo, Xinxin Zuo, Sen Wang, Shihao Zou 等ACM MM 2020 · 被引用 394 次
- ReMoDiffuse: Retrieval-Augmented Motion Diffusion ModelMingyuan Zhang, Xinying Guo, Liang Pan, Zhongang Cai 等ICCV 2023 · 被引用 301 次
相关 Paper
- Move as you Say, Interact as you can: Language-Guided Human Motion Generation with Scene AffordanceZan Wang, Yixin Chen, Baoxiong Jia, Puhao Li 等CVPR 2024 · 被引用 38 次
- Event-Driven Storytelling with Multiple Lifelike Humans in a 3D SceneDonggeun Lim, Jinseok Bae, Inwoo Hwang, Seungmin Lee 等ICCV 2025
- Text2Interact: High-Fidelity and Diverse Text-to-Two-Person Interaction GenerationQingxuan Wu, Zhiyang Dou, chuan guo, Yiming Huang 等ICLR 2026 · 被引用 10 次
- HUMANISE: Language-conditioned Human Motion Generation in 3D ScenesZan Wang, Yixin Chen, Tengyu Liu, Yixin Zhu 等NeurIPS 2022 · 被引用 207 次
- MotionCtrl: A Real-Time Controllable Vision-Language-Motion ModelBin Cao, Sipeng Zheng, Ye Wang, Lujie Xia 等ICCV 2025 · 被引用 1 次
