Zero-1-to-A: Zero-Shot One Image to Animatable Head Avatars Using Video Diffusion
Zhenglin Zhou, Fan Ma, Hehe Fan, Tat-Seng Chua
Abstract
Animatable head avatar generation typically requires extensive data for training. To reduce the data requirements, a natural solution is to leverage existing data-free static avatar generation methods, such as pre-trained diffusion models with score distillation sampling (SDS), which align avatars with pseudo ground-truth outputs from the diffusion model. However, directly distilling 4D avatars from video diffusion often leads to over-smooth results due to spatial and temporal inconsistencies in the generated video. To address this issue, we propose Zero-1-to-A, a robust method that synthesizes a spatial and temporal consistency dataset for 4D avatar reconstruction using the video diffusion model. Specifically, Zero-1-to-A iteratively constructs video datasets and optimizes animatable avatars in a progressive manner, ensuring that avatar quality increases smoothly and consistently throughout the learning process. This progressive learning involves two stages: (1) Spatial Consistency Learning fixes expressions and learns from front-to-side views, and (2) Temporal Consistency Learning fixes views and learns from relaxed to exaggerated expressions, generating 4D avatars in a simple-to-complex manner. Extensive experiments demonstrate that Zero-1-to-A improves fidelity, animation quality, and rendering speed compared to existing diffusion-based methods, providing a solution for lifelike avatar creation. Code is publicly available at: https://github.com/ZhenglinZhou/Zero-1-to-A .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 03c848a5-1d74-40db-bd3e-dde15e21c89dCited by top-tier papers5
- FastGHA: Generalized Few-Shot 3D Gaussian Head Avatars with Real-Time AnimationXinya Ji, Sebastian Weiss, Manuel Kansy, Jacek Naruniec et al.ICLR 2026 · 6 citations
- AUHead: Realistic Emotional Talking Head Generation via Action Units ControlJiayi Lyu, Leigang Qu, Wenjing Zhang, Hanyu Jiang et al.ICLR 2026 · 2 citations
- Anti-Avatar: Protect Against Unauthorized 3D Head Avatar Generation via Dual-Space DivergenceLingzhuang Meng, Mingwen Shao, Xiang Lv, Mengyao Wu et al.AAAI 2026 · 1 citation
- Multi-view Consistent 3D Gaussian Head Avatars 'without' Multi-view GenerationAviral Chharia, Fernando De la TorreCVPR 2026
- Uncertainty-Aware 3D Reconstruction for Dynamic Underwater ScenesRui Liu, Zhibo Duan, Jianzhe Gao, Yi Yang et al.ICLR 2026
Builds on49
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 6,759 citations
Related papers
- Animate3D: Animating Any 3D Model with Multi-view Video DiffusionYanqin Jiang, Chaohui Yu, Chenjie Cao, Fan Wang et al.NeurIPS 2024 · 65 citations
- GeoDiff4D: Geometry-Aware Diffusion for 4D Head Avatar ReconstructionChao Xu, Xiaochen Zhao, Xiang Deng, Jingxiang Sun et al.CVPR 2026
- FaceCraft4D: Animated 3D Facial Avatar Generation from a Single ImageFei Yin, Mallikarjun B. R., Chun-Han Yao, Rafal K. Mantiuk et al.ICCV 2025 · 2 citations
- EG4D: Explicit Generation of 4D Object without Score DistillationQi Sun, Zhiyang Guo, Ziyu Wan, Jing Nathan Yan et al.ICLR 2025
- ConsistentAvatar: Learning to Diffuse Fully Consistent Talking Head Avatar with Temporal GuidanceHaijie Yang, Zhenyu Zhang, Hao Tang, Jianjun Qian et al.ACM MM 2024 · 3 citations
