Human Motion Generation in 3D Scenes from Open-Ended Textual Instructions with MLLM Planning
Siyi Qian, Jian Fang, Yuzhou Mao, Yayun Zou, Wentao Zhang, Haiwei Xue
摘要
Generating human motion in scenes from text aims to synthesize semantically aligned and scene-aware motions. Existing methods have made significant progress by incorporating spatial reasoning and structured generation strategies to connect text descriptions with human-scene interactions. However, they typically rely on simple textual inputs and struggle to comprehend open-ended instructions. There are three key challenges: (1) difficulty in understanding complex instructions due to limited and templated training text annotations; (2) inability to generate natural motions that align with arbitrary trajectories described in text; (3) lack of motion diversity that matches the intended semantics. To address these challenges, we propose PSMo, which consists of two components: the Semantic Planner and the Scene-Aware Motion Generator. The Semantic Planner leverages a Multimodal Large Language Model (MLLM) to parse open-ended instructions, and plans fine-grained motion states aligned with arbitrary trajectories. The scene-aware motion generator adopts the diffusion model with trajectory constraints and a sequential tiling strategy. To enhance motion diversity, we introduce a retrieval-augmented strategy and Scene-Aware Retrieval Attention, which integrates multi-modal features into the generation process. Extensive experiments demonstrate that our method produces high-quality and natural motions under open-ended instructions in scenes.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- AMD: Anatomical Motion Diffusion with Interpretable Motion Decomposition and FusionBeibei Jing, Youjia Zhang, Zikai Song, Junqing Yu 等AAAI 2024 · 被引用 6 次
- LGTM: Local-to-Global Text-Driven Human Motion Diffusion ModelHaowen Sun, Ruikun Zheng, Haibin Huang, Chongyang Ma 等SIGGRAPH 2024 · 被引用 12 次
- FrankenMotion: Part-level Human Motion Generation and CompositionChuqiao Li, Xianghui Xie, Yong Cao, Andreas Geiger 等CVPR 2026 · 被引用 10 次
- OpenHOI: Open-World Hand-Object Interaction Synthesis with Multimodal Large Language ModelZhenhao Zhang, Ye Shi, Lingxiao Yang, Suting Ni 等NeurIPS 2025 · 被引用 25 次
- MotionMaster: Generalizable Text-Driven Motion Generation and EditingNan Jiang, Yunhao Li, Lexi Pang, Zimo He 等CVPR 2026
