Narrator: Towards Natural Control of Human-Scene Interaction Generation via Relationship Reasoning
Haibiao Xuan, Xiongzheng Li, Jinsong Zhang, Hongwen Zhang, Yebin Liu, Kun Li
Abstract
Naturally controllable human-scene interaction (HSI) generation has an important role in various fields, such as VR/AR content creation and human-centered AI. However, existing methods are unnatural and unintuitive in their controllability, which heavily limits their application in practice. Therefore, we focus on a challenging task of naturally and controllably generating realistic and diverse HSIs from textual descriptions. From human cognition, the ideal generative model should correctly reason about spatial relationships and interactive actions. To that end, we propose Narrator, a novel relationship reasoning-based generative approach using a conditional variation autoencoder for naturally controllable generation given a 3D scene and a textual description. Also, we model global and local spatial relationships in a 3D scene and a textual description respectively based on the scene graph, and introduce a part-level action mechanism to represent interactions as atomic body part states. In particular, benefiting from our relationship reasoning, we further propose a simple yet effective multi-human generation strategy, which is the first exploration for controllable multi-human scene interaction generation. Our extensive experiments and perceptual studies show that Narrator can controllably generate diverse interactions and significantly outperform existing works.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext de64c921-397b-4d5d-af7b-c1b583a75d76Cited by top-tier papers8
- Human-Object Interaction via Automatically Designed VLM-Guided Motion PolicyZekai Deng, Ye Shi, Kaiyang Ji, Lan Xu et al.ICLR 2026 · 11 citations
- HOIAnimator: Generating Text-Prompt Human-Object Animations Using Novel Perceptive Diffusion ModelsWenfeng Song, Xinyu Zhang, Shuai Li, Yang Gao et al.CVPR 2024 · 6 citations
- Diffusion Implicit Policy for Unpaired Scene-aware Motion SynthesisJingyu Gong, Chong Zhang, Fengqi Liu, Ke Fan et al.AAAI 2026 · 5 citations
- HOSIG: Full-Body Human-Object-Scene Interaction Generation with Hierarchical Scene PerceptionWei Yao, Yunlian Sun, Hongwen Zhang, Yebin Liu et al.AAAI 2026 · 4 citations
- InfBaGel: Human-Object-Scene Interaction Generation with Dynamic Perception and Iterative RefinementYude Zou, Junji Gong, Xing Gao, Zixuan Li et al.ICLR 2026 · 3 citations
Builds on17
- DeepGCNs: Can GCNs Go As Deep As CNNs?Guohao Li, Matthias Müller, Ali K. Thabet, Bernard GhanemICCV 2019 · 1,586 citations
- Generating Diverse and Natural 3D Human Motions from TextChuan Guo, Shihao Zou, Xinxin Zuo, Sen Wang et al.CVPR 2022 · 462 citations
- Resolving 3D Human Pose Ambiguities With 3D Scene ConstraintsMohamed Hassan, Vasileios Choutas, Dimitrios Tzionas, Michael J. BlackICCV 2019 · 384 citations
- Synthesis of Compositional Animations from Textual DescriptionsAnindita Ghosh, Noshaba Cheema, Cennet Oguz, Christian Theobalt et al.ICCV 2021 · 226 citations
- AvatarCLIP: zero-shot text-driven generation and animation of 3D avatarsFangzhou Hong, Mingyuan Zhang, Liang Pan, Zhongang Cai et al.SIGGRAPH 2022 · 213 citations
Related papers
- Generating 3D People in Scenes Without PeopleYan Zhang, Mohamed Hassan, Heiko Neumann, Michael J. Black et al.CVPR 2020
- HOI-PAGE: Zero-Shot Human-Object Interaction Generation with Part Affordance GuidanceLei Li, Angela DaiICML 2026
- Graph-to-3D: End-to-End Generation and Manipulation of 3D Scenes Using Scene GraphsHelisa Dhamo, Fabian Manhardt, Nassir Navab, Federico TombariICCV 2021 · 98 citations
- Auto-Regressive Diffusion for Generating 3D Human-Object InteractionsZichen Geng, Zeeshan Hayder, Wei Liu, Ajmal Saeed MianAAAI 2025 · 8 citations
- CommonScenes: Generating Commonsense 3D Indoor Scenes with Scene GraphsGuangyao Zhai, Evin Pinar Örnek, Shun-Cheng Wu, Yan Di et al.NeurIPS 2023 · 76 citations
