GAS: Generative Avatar Synthesis from a Single Image
Yixing Lu, Junting Dong, Youngjoong Kwon, Qin Zhao, Bo Dai, Fernando De la Torre
Abstract
We present a unified and generalizable framework for synthesizing view-consistent and temporally coherent avatars from a single image, addressing the challenging task of single-image avatar generation. Existing diffusion-based methods often condition on sparse human templates (e.g., depth or normal maps), which leads to multi-view and temporal inconsistencies due to the mismatch between these signals and the true appearance of the subject. Our approach bridges this gap by combining the reconstruction power of regression-based 3D human reconstruction with the generative capabilities of a diffusion model. In a first step, an initial 3D reconstructed human through a generalized NeRF provides comprehensive conditioning, ensuring high-quality synthesis faithful to the reference appearance and structure. Subsequently, the derived geometry and appearance from the generalized NeRF serve as input to a video-based diffusion model. This strategic integration is pivotal for enforcing both multi-view and temporal consistency throughout the avatar's generation. Empirical results underscore the superior generalization ability of our proposed method, demonstrating its effectiveness across diverse in-domain and out-of-domain in-the-wild datasets.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e11064bf-e627-4581-9231-1ad3a1ef8378Cited by top-tier papers6
- LHM: Large Animatable Human Reconstruction Model for Single Image to 3D in SecondsLingteng Qiu, Xiaodong Gu, Peihao Li, Qi Zuo et al.ICCV 2025 · 19 citations
- HoliGS: Holistic Gaussian Splatting for Embodied View SynthesisXiaoyuan Wang, Yizhou Zhao, Botao Ye, Xiaojun Shan et al.NeurIPS 2025 · 8 citations
- SIGMAN: Scaling 3D Human Gaussian Generation with Millions of AssetsYuhang Yang, Fengqi Liu, Yixing Lu, Qin Zhao et al.ICCV 2025 · 5 citations
- Bringing Your Portrait to 3D PresenceJiawei Zhang, Lei Chu, Jiahao Li, Zhenyu Zang et al.CVPR 2026 · 3 citations
- HumanNOVA: Photorealistic, Universal and Rapid 3D Human Avatar Modeling from a Single ImageHezhen Hu, Wangbo Zhao, Lanqing Guo, Hanwen Jiang et al.CVPR 2026
Builds on39
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- 3D Gaussian Splatting for Real-Time Radiance Field RenderingBernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, George DrettakisSIGGRAPH 2023 · 5,687 citations
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
- Video Diffusion ModelsJonathan Ho, Tim Salimans, Alexey A. Gritsenko, William Chan et al.NeurIPS 2022 · 2,948 citations
Related papers
- FaceCraft4D: Animated 3D Facial Avatar Generation from a Single ImageFei Yin, Mallikarjun B. R., Chun-Han Yao, Rafal K. Mantiuk et al.ICCV 2025 · 2 citations
- PERSONA: Personalized Whole-Body 3D Avatar with Pose-Driven Deformations from a Single ImageGeonhee Sim, Gyeongsik MoonICCV 2025 · 6 citations
- GAF: Gaussian Avatar Reconstruction from Monocular Videos via Multi-view DiffusionJiapeng Tang, Davide Davoli, Tobias Kirschstein, Liam Schoneveld et al.CVPR 2025
- Human-3Diffusion: Realistic Avatar Creation via Explicit 3D Consistent Diffusion ModelsYuxuan Xue, Xianghui Xie, Riccardo Marin, Gerard Pons-MollNeurIPS 2024 · 49 citations
- Expressive Talking Human from Single-Image with Imperfect PriorsJun Xiang, Yudong Guo, Leipeng Hu, Boyang Guo et al.ICCV 2025 · 3 citations
