GenZI: Zero-Shot 3D Human-Scene Interaction Generation
Lei Li, Angela Dai
Abstract
riding a motorcycle, sitting" "hands knocking on the closed door, facing the door, standing" "stretching legs and sitting on the saddle on a standing cow" "picking up dumbbells on a shelf, standing, bending over" Figure 1 . Given an arbitrary 3D scene, GenZI can synthesize virtual humans interacting with the 3D environment at specified locations from a brief text description. Our approach does not require any 3D human-scene interaction training data or 3D learning. By distilling interaction priors from powerful 2D vision-language models, we optimize for 3D human-scene interaction synthesis in a flexible fashion, with simple language-based control and high generality to various types of scene environments.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a45c66f9-e5e9-4b2b-a10c-34f766333a61Cited by top-tier papers17
- InterDreamer: Zero-Shot Text to 3D Dynamic Human-Object InteractionSirui Xu, Ziyin Wang, Yu-Xiong Wang, Liangyan GuiNeurIPS 2024 · 78 citations
- Target-Aware Video Diffusion ModelsTaeksoo Kim, Hanbyul JooICLR 2026 · 7 citations
- Decoupled Generative Modeling for Human-Object Interaction SynthesisHwanhee Jung, Seunggwan Lee, Jeongyoon Yoon, SeungHyeon Kim et al.CVPR 2026 · 4 citations
- HOSIG: Full-Body Human-Object-Scene Interaction Generation with Hierarchical Scene PerceptionWei Yao, Yunlian Sun, Hongwen Zhang, Yebin Liu et al.AAAI 2026 · 4 citations
- PrimHOI: Compositional Human-Object Interaction via Reusable PrimitivesKai Jia, Tengyu Liu, Yixin Zhu, Mingtao Pei et al.ICCV 2025 · 4 citations
Builds on20
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
Related papers
- InteractAnything: Zero-shot Human Object Interaction Synthesis via LLM Feedback and Object Affordance ParsingJinlu Zhang, Yixin Chen, Zan Wang, Jie Yang et al.CVPR 2025
- InteractVLM: 3D Interaction Reasoning from 2D Foundational ModelsSai Kumar Dwivedi, Dimitrije Antic, Shashank Tripathi, Omid Taheri et al.CVPR 2025
- HUMANISE: Language-conditioned Human Motion Generation in 3D ScenesZan Wang, Yixin Chen, Tengyu Liu, Yixin Zhu et al.NeurIPS 2022 · 207 citations
- Generating Human Motion in 3D Scenes from Text DescriptionsZhi Cen, Huaijin Pi, Sida Peng, Zehong Shen et al.CVPR 2024
- CG-HOI: Contact-Guided 3D Human-Object Interaction GenerationChristian Diller, Angela DaiCVPR 2024
