AvatarCLIP: zero-shot text-driven generation and animation of 3D avatars
Fangzhou Hong, Mingyuan Zhang, Liang Pan, Zhongang Cai, Lei Yang, Ziwei Liu
Abstract
3D avatar creation plays a crucial role in the digital age. However, the whole production process is prohibitively time-consuming and labor-intensive. To democratize this technology to a larger audience, we propose AvatarCLIP, a zero-shot text-driven framework for 3D avatar generation and animation. Unlike professional software that requires expert knowledge, AvatarCLIP empowers layman users to customize a 3D avatar with the desired shape and texture, and drive the avatar with the described motions using solely natural languages. Our key insight is to take advantage of the powerful vision-language model CLIP for supervising neural human generation, in terms of 3D geometry, texture and animation. Specifically, driven by natural language descriptions, we initialize 3D human geometry generation with a shape VAE network. Based on the generated 3D human shapes, a volume rendering model is utilized to further facilitate geometry sculpting and texture generation. Moreover, by leveraging the priors learned in the motion VAE, a CLIP-guided reference-based motion synthesis method is proposed for the animation of the generated 3D avatar. Extensive qualitative and quantitative experiments validate the effectiveness and generalizability of AvatarCLIP on a wide range of avatars. Remarkably, AvatarCLIP can generate unseen 3D avatars with novel animations, achieving superior zero-shot capability. Codes are available at https://github.com/hongfz16/AvatarCLIP.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers139
- One-2-3-45: Any Single Image to 3D Mesh in 45 Seconds without Per-Shape OptimizationMinghua Liu, Chao Xu, Haian Jin, Linghao Chen et al.NeurIPS 2023 · 755 citations
- Instruct-NeRF2NeRF: Editing 3D Scenes with InstructionsAyaan Haque, Matthew Tancik, Alexei A. Efros, Aleksander Holynski et al.ICCV 2023 · 544 citations
- DreamFusion: Text-to-3D using 2D DiffusionBen Poole, Ajay Jain, Jonathan T. Barron, Ben MildenhallICLR 2023 · 463 citations
- Magic123: One Image to High-Quality 3D Object Generation Using Both 2D and 3D Diffusion PriorsGuocheng Qian, Jinjie Mai, Abdullah Hamdi, Jian Ren et al.ICLR 2024 · 444 citations
- Human Motion Diffusion as a Generative PriorYoni Shafir, Guy Tevet, Roy Kapon, Amit Haim BermanoICLR 2024 · 371 citations
Builds on35
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray et al.ICML 2021 · 6,356 citations
- NeuS: Learning Neural Implicit Surfaces by Volume Rendering for Multi-view ReconstructionPeng Wang, Lingjie Liu, Yuan Liu, Christian Theobalt et al.NeurIPS 2021 · 2,500 citations
- AMASS: Archive of Motion Capture As Surface ShapesNaureen Mahmood, Nima Ghorbani, Nikolaus F. Troje, Gerard Pons-Moll et al.ICCV 2019 · 1,784 citations
- StyleCLIP: Text-Driven Manipulation of StyleGAN ImageryOr Patashnik, Zongze Wu, Eli Shechtman, Daniel Cohen-Or et al.ICCV 2021 · 1,437 citations
Related papers
- AvatarFusion: Zero-shot Generation of Clothing-Decoupled 3D Avatars Using 2D DiffusionShuo Huang, Zongxin Yang, Liangting Li, Yi Yang et al.ACM MM 2023 · 19 citations
- CLIPTexture: Text-Driven Texture SynthesisYiren SongACM MM 2022 · 7 citations
- CLIP-Forge: Towards Zero-Shot Text-to-Shape GenerationAditya Sanghi, Hang Chu, Joseph G. Lambourne, Ye Wang et al.CVPR 2022 · 206 citations
- ClipFace: Text-guided Editing of Textured 3D Morphable ModelsShivangi Aneja, Justus Thies, Angela Dai, Matthias NießnerSIGGRAPH 2023 · 40 citations
- DreamHuman: Animatable 3D Avatars from TextNikos Kolotouros, Thiemo Alldieck, Andrei Zanfir, Eduard Gabriel Bazavan et al.NeurIPS 2023 · 136 citations
