AvatarGO: Zero-shot 4D Human-Object Interaction Generation and Animation
Yukang Cao, Liang Pan, Kai Han, Kwan-Yee K. Wong, Ziwei Liu
Abstract
Recent advancements in diffusion models have led to significant improvements in the generation and animation of 4D full-body human-object interactions (HOI). Nevertheless, existing methods primarily focus on SMPL-based motion generation, which is limited by the scarcity of realistic large-scale interaction data. This constraint affects their ability to create everyday HOI scenes. This paper addresses this challenge using a zero-shot approach with a pre-trained diffusion model. Despite this potential, achieving our goals is difficult due to the diffusion model's lack of understanding of ''where'' and ''how'' objects interact with the human body. To tackle these issues, we introduce AvatarGO, a novel framework designed to generate animatable 4D HOI scenes directly from textual inputs. Specifically, 1) for the ''where'' challenge, we propose LLM-guided contact retargeting, which employs Lang-SAM to identify the contact body part from text prompts, ensuring precise representation of human-object spatial relations. 2) For the ''how'' challenge, we introduce correspondence-aware motion optimization that constructs motion fields for both human and object models using the linear blend skinning function from SMPL-X. Our framework not only generates coherent compositional motions, but also exhibits greater robustness in handling penetration issues. Extensive experiments with existing methods validate AvatarGO's superior generation and animation capabilities on a variety of human-object pairs and diverse poses. As the first attempt to synthesize 4D avatars with object interactions, we hope AvatarGO could open new doors for human-centric 4D content creation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 652d9fe8-70c1-43f9-bd48-bda0dff05d2aCited by top-tier papers10
- Lyra: Generative 3D Scene Reconstruction via Video Diffusion Model Self-DistillationSherwin Bahmani, Tianchang Shen, Jiawei Ren, Jiahui Huang et al.ICLR 2026 · 33 citations
- Humoto: A 4D Dataset of Mocap Human Object InteractionsJiaxin Lu, Chun-Hao Paul Huang, Uttaran Bhattacharya, Qixing Huang et al.ICCV 2025 · 4 citations
- Splat4D: Diffusion-Enhanced 4D Gaussian Splatting for Temporally and Spatially Consistent Content CreationMinghao Yin, Yukang Cao, Songyou Peng, Kai HanSIGGRAPH 2025 · 2 citations
- ViHOI: Human-Object Interaction Synthesis with Visual PriorsSongjin Cai, Linjie Zhong, Ling Guo, Changxing DingCVPR 2026 · 2 citations
- Towards Scalable Spatial Intelligence Via 2D-To-3D Data LiftingXingyu Miao, Haoran Duan, Quanhao Qian, Jiuniu Wang et al.ICCV 2025 · 1 citation
Builds on43
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 6,759 citations
- 3D Gaussian Splatting for Real-Time Radiance Field RenderingBernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, George DrettakisSIGGRAPH 2023 · 5,687 citations
- Zero-1-to-3: Zero-shot One Image to 3D ObjectRuoshi Liu, Rundi Wu, Basile Van Hoorick, Pavel Tokmakov et al.ICCV 2023 · 1,662 citations
Related papers
- HOI-PAGE: Zero-Shot Human-Object Interaction Generation with Part Affordance GuidanceLei Li, Angela DaiICML 2026
- DAViD: Modeling Dynamic Affordance of 3D Objects Using Pre-Trained Video Diffusion ModelsHyeonwoo Kim, Sangwon Baik, Hanbyul JooICCV 2025 · 1 citation
- InterDiff: Generating 3D Human-Object Interactions with Physics-Informed DiffusionSirui Xu, Zhengyuan Li, Yu-Xiong Wang, Liang-Yan GuiICCV 2023 · 201 citations
- AnchorHOI: Zero-shot Generation of 4D Human-Object Interaction via Anchor-based Prior DistillationSisi Dai, Kai XuAAAI 2026
- Target-Aware Video Diffusion ModelsTaeksoo Kim, Hanbyul JooICLR 2026 · 7 citations
