DreamWaltz: Make a Scene with Complex 3D Animatable Avatars
Yukun Huang, Jianan Wang, Ailing Zeng, He Cao, Xianbiao Qi, Yukai Shi, Zheng-Jun Zha, Lei Zhang
Abstract
We present DreamWaltz, a novel framework for generating and animating complex 3D avatars given text guidance and parametric human body prior. While recent methods have shown encouraging results for text-to-3D generation of common objects, creating high-quality and animatable 3D avatars remains challenging. To create high-quality 3D avatars, DreamWaltz proposes 3D-consistent occlusion-aware Score Distillation Sampling (SDS) to optimize implicit neural representations with canonical poses. It provides view-aligned supervision via 3D-aware skeleton conditioning which enables complex avatar generation without artifacts and multiple faces. For animation, our method learns an animatable and generalizable avatar representation which could map arbitrary poses to the canonical pose representation. Extensive evaluations demonstrate that DreamWaltz is an effective and robust approach for creating 3D avatars that can take on complex shapes and appearances as well as novel poses for animation. The proposed framework further enables the creation of complex scenes with diverse compositions, including avatar-avatar, avatar-object and avatar-scene interactions. See https://dreamwaltz3d.github.io/ for more vivid 3D avatar and animation results.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c80a2532-3755-488b-89ba-d97ea828e602Cited by top-tier papers27
- GaussianDreamer: Fast Generation from Text to 3D Gaussians by Bridging 2D and 3D Diffusion ModelsTaoran Yi, Jiemin Fang, Junjie Wang, Guanjun Wu et al.CVPR 2024 · 106 citations
- AvatarVerse: High-Quality & Stable 3D Avatar Creation from Text and PoseHuichao Zhang, Bowen Chen, Hao Yang, Liao Qu et al.AAAI 2024 · 73 citations
- CharacterGen: Efficient 3D Character Generation from Single Images with Multi-View Pose CanonicalizationHao-Yang Peng, Jia-Peng Zhang, Meng-Hao Guo, Yan-Pei Cao et al.SIGGRAPH 2024 · 30 citations
- LHM: Large Animatable Human Reconstruction Model for Single Image to 3D in SecondsLingteng Qiu, Xiaodong Gu, Peihao Li, Qi Zuo et al.ICCV 2025 · 19 citations
- X-Oscar: A Progressive Framework for High-quality Text-guided 3D Animatable Avatar GenerationYiwei Ma, Zhekai Lin, Jiayi Ji, Yijun Fan et al.ICML 2024 · 9 citations
Builds on26
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
Related papers
- DreamFace: Progressive Generation of Animatable 3D Faces under Text GuidanceLongwen Zhang, Qiwei Qiu, Hongyang Lin, Qixuan Zhang et al.SIGGRAPH 2023 · 68 citations
- DreamHuman: Animatable 3D Avatars from TextNikos Kolotouros, Thiemo Alldieck, Andrei Zanfir, Eduard Gabriel Bazavan et al.NeurIPS 2023 · 136 citations
- DreamAvatar: Text-and-Shape Guided 3D Human Avatar Generation via Diffusion ModelsYukang Cao, Yan-Pei Cao, Kai Han, Ying Shan et al.CVPR 2024
- Text-based Animatable 3D Avatars with Morphable Model AlignmentYiqian Wu, Malte Prinzler, Xiaogang Jin, Siyu TangSIGGRAPH 2025 · 1 citation
- HandDreamer: Zero-Shot Text to 3D Hand Model Generation using Corrective Hand Shape GuidanceGreen Rosh, Prateek Kukreja, Vishakha SR, Pawan Prasad B HCVPR 2026
