HUMANISE: Language-conditioned Human Motion Generation in 3D Scenes
Zan Wang, Yixin Chen, Tengyu Liu, Yixin Zhu, Wei Liang, Siyuan Huang
Abstract
Learning to generate diverse scene-aware and goal-oriented human motions in 3D scenes remains challenging due to the mediocre characteristics of the existing datasets on Human-Scene Interaction (HSI); they only have limited scale/quality and lack semantics. To fill in the gap, we propose a large-scale and semantic-rich synthetic HSI dataset, denoted as HUMANISE, by aligning the captured human motion sequences with various 3D indoor scenes. We automatically annotate the aligned motions with language descriptions that depict the action and the unique interacting objects in the scene; e.g., sit on the armchair near the desk. HUMANISE thus enables a new generation task, language-conditioned human motion generation in 3D scenes. The proposed task is challenging as it requires joint modeling of the 3D scene, human motion, and natural language. To tackle this task, we present a novel scene-and-language conditioned generative model that can produce 3D human motions of the desirable action interacting with the specified objects. Our experiments demonstrate that our model generates diverse and semantically consistent human motions in 3D scenes.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers77
- InterDiff: Generating 3D Human-Object Interactions with Physics-Informed DiffusionSirui Xu, Zhengyuan Li, Yu-Xiong Wang, Liang-Yan GuiICCV 2023 · 201 citations
- HumanTOMATO: Text-aligned Whole-body Motion GenerationShunlin Lu, Ling-Hao Chen, Ailing Zeng, Jing Lin et al.ICML 2024 · 124 citations
- Unified Human-Scene Interaction via Prompted Chain-of-ContactsZeqi Xiao, Tai Wang, Jingbo Wang, Jinkun Cao et al.ICLR 2024 · 113 citations
- Full-Body Articulated Human-Object InteractionNan Jiang, Tengyu Liu, Zhexuan Cao, Jieming Cui et al.ICCV 2023 · 80 citations
- InterDreamer: Zero-Shot Text to 3D Dynamic Human-Object InteractionSirui Xu, Ziyin Wang, Yu-Xiong Wang, Liangyan GuiNeurIPS 2024 · 78 citations
Builds on23
- AMASS: Archive of Motion Capture As Surface ShapesNaureen Mahmood, Nima Ghorbani, Nikolaus F. Troje, Gerard Pons-Moll et al.ICCV 2019 · 1,784 citations
- Action-Conditioned 3D Human Motion Synthesis with Transformer VAEMathis Petrovich, Michael J. Black, Gül VarolICCV 2021 · 672 citations
- Action2Motion: Conditioned Generation of 3D Human MotionsChuan Guo, Xinxin Zuo, Sen Wang, Shihao Zou et al.ACM MM 2020 · 394 citations
- Resolving 3D Human Pose Ambiguities With 3D Scene ConstraintsMohamed Hassan, Vasileios Choutas, Dimitrios Tzionas, Michael J. BlackICCV 2019 · 384 citations
- Stochastic Scene-Aware Motion PredictionMohamed Hassan, Duygu Ceylan, Ruben Villegas, Jun Saito et al.ICCV 2021 · 240 citations
Related papers
- Move as you Say, Interact as you can: Language-Guided Human Motion Generation with Scene AffordanceZan Wang, Yixin Chen, Baoxiong Jia, Puhao Li et al.CVPR 2024 · 38 citations
- Generating Human Motion in 3D Scenes from Text DescriptionsZhi Cen, Huaijin Pi, Sida Peng, Zehong Shen et al.CVPR 2024
- MIME: Human-Aware 3D Scene GenerationHongwei Yi, Chun-Hao P. Huang, Shashank Tripathi, Lea Hering et al.CVPR 2023
- Synthesizing Diverse Human Motions in 3D Indoor ScenesKaifeng Zhao, Yan Zhang, Shaofei Wang, Thabo Beeler et al.ICCV 2023 · 116 citations
- IntentMotion: Learning Intent-Aware Human Motion from Language in 3D ScenesWenfeng Song, Shi Zheng, Xinyu Zhang, Xingliang Jin et al.AAAI 2026
