Human Motion Generation in 3D Scenes from Open-Ended Textual Instructions with MLLM Planning
Siyi Qian, Jian Fang, Yuzhou Mao, Yayun Zou, Wentao Zhang, Haiwei Xue
Abstract
Generating human motion in scenes from text aims to synthesize semantically aligned and scene-aware motions. Existing methods have made significant progress by incorporating spatial reasoning and structured generation strategies to connect text descriptions with human-scene interactions. However, they typically rely on simple textual inputs and struggle to comprehend open-ended instructions. There are three key challenges: (1) difficulty in understanding complex instructions due to limited and templated training text annotations; (2) inability to generate natural motions that align with arbitrary trajectories described in text; (3) lack of motion diversity that matches the intended semantics. To address these challenges, we propose PSMo, which consists of two components: the Semantic Planner and the Scene-Aware Motion Generator. The Semantic Planner leverages a Multimodal Large Language Model (MLLM) to parse open-ended instructions, and plans fine-grained motion states aligned with arbitrary trajectories. The scene-aware motion generator adopts the diffusion model with trajectory constraints and a sequential tiling strategy. To enhance motion diversity, we introduce a retrieval-augmented strategy and Scene-Aware Retrieval Attention, which integrates multi-modal features into the generation process. Extensive experiments demonstrate that our method produces high-quality and natural motions under open-ended instructions in scenes.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get de2b2ea1-8b15-4821-83b6-2d2fe7c574b0Related papers
- AMD: Anatomical Motion Diffusion with Interpretable Motion Decomposition and FusionBeibei Jing, Youjia Zhang, Zikai Song, Junqing Yu et al.AAAI 2024 · 6 citations
- LGTM: Local-to-Global Text-Driven Human Motion Diffusion ModelHaowen Sun, Ruikun Zheng, Haibin Huang, Chongyang Ma et al.SIGGRAPH 2024 · 12 citations
- FrankenMotion: Part-level Human Motion Generation and CompositionChuqiao Li, Xianghui Xie, Yong Cao, Andreas Geiger et al.CVPR 2026 · 10 citations
- OpenHOI: Open-World Hand-Object Interaction Synthesis with Multimodal Large Language ModelZhenhao Zhang, Ye Shi, Lingxiao Yang, Suting Ni et al.NeurIPS 2025 · 25 citations
- MotionMaster: Generalizable Text-Driven Motion Generation and EditingNan Jiang, Yunhao Li, Lexi Pang, Zimo He et al.CVPR 2026
