GenZI: Zero-Shot 3D Human-Scene Interaction Generation
Lei Li, Angela Dai
摘要
riding a motorcycle, sitting" "hands knocking on the closed door, facing the door, standing" "stretching legs and sitting on the saddle on a standing cow" "picking up dumbbells on a shelf, standing, bending over" Figure 1 . Given an arbitrary 3D scene, GenZI can synthesize virtual humans interacting with the 3D environment at specified locations from a brief text description. Our approach does not require any 3D human-scene interaction training data or 3D learning. By distilling interaction priors from powerful 2D vision-language models, we optimize for 3D human-scene interaction synthesis in a flexible fashion, with simple language-based control and high generality to various types of scene environments.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper17
- InterDreamer: Zero-Shot Text to 3D Dynamic Human-Object InteractionSirui Xu, Ziyin Wang, Yu-Xiong Wang, Liangyan GuiNeurIPS 2024 · 被引用 78 次
- Target-Aware Video Diffusion ModelsTaeksoo Kim, Hanbyul JooICLR 2026 · 被引用 7 次
- Decoupled Generative Modeling for Human-Object Interaction SynthesisHwanhee Jung, Seunggwan Lee, Jeongyoon Yoon, SeungHyeon Kim 等CVPR 2026 · 被引用 4 次
- HOSIG: Full-Body Human-Object-Scene Interaction Generation with Hierarchical Scene PerceptionWei Yao, Yunlian Sun, Hongwen Zhang, Yebin Liu 等AAAI 2026 · 被引用 4 次
- PrimHOI: Compositional Human-Object Interaction via Reusable PrimitivesKai Jia, Tengyu Liu, Yixin Zhu, Mingtao Pei 等ICCV 2025 · 被引用 4 次
它引用的顶会 Paper20
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li 等NeurIPS 2022 · 被引用 8,965 次
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
相关 Paper
- InteractAnything: Zero-shot Human Object Interaction Synthesis via LLM Feedback and Object Affordance ParsingJinlu Zhang, Yixin Chen, Zan Wang, Jie Yang 等CVPR 2025
- InteractVLM: 3D Interaction Reasoning from 2D Foundational ModelsSai Kumar Dwivedi, Dimitrije Antic, Shashank Tripathi, Omid Taheri 等CVPR 2025
- HUMANISE: Language-conditioned Human Motion Generation in 3D ScenesZan Wang, Yixin Chen, Tengyu Liu, Yixin Zhu 等NeurIPS 2022 · 被引用 207 次
- Generating Human Motion in 3D Scenes from Text DescriptionsZhi Cen, Huaijin Pi, Sida Peng, Zehong Shen 等CVPR 2024
- CG-HOI: Contact-Guided 3D Human-Object Interaction GenerationChristian Diller, Angela DaiCVPR 2024
