GUAVA: Generalizable Upper Body 3D Gaussian Avatar
Dongbin Zhang, Yunfei Liu, Lijian Lin, Ye Zhu, Yang Li, Minghan Qin, Yu Li, Haoqian Wang
Abstract
Reconstructing a high-quality, animatable 3D human avatar with expressive facial and hand motions from a single image has gained significant attention due to its broad application potential. 3D human avatar reconstruction typically requires multi-view or monocular videos and training on individual IDs, which is both complex and time-consuming. Furthermore, limited by SMPLX's expressiveness, these methods often focus on body motion but struggle with facial expressions. To address these challenges, we first introduce an expressive human model (EHM) to enhance facial expression capabilities and develop an accurate tracking method. Based on this template model, we propose GUAVA, the first framework for fast animatable upper-body 3D Gaussian avatar reconstruction. We leverage inverse texture mapping and projection sampling techniques to infer Ubody (upper-body) Gaussians from a single image. The rendered images are refined through a neural refiner. Experimental results demonstrate that GUAVA significantly outperforms previous methods in rendering quality and offers significant speed improvements, with reconstruction times in the sub-second range (0.1s), and supports real-time animation and rendering.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d9bd285f-ad70-4e74-920c-cf3d4cd56efeCited by top-tier papers4
- STAvatar: Soft Binding and Temporal Density Control for Monocular 3D Head Avatars ReconstructionJiankuo Zhao, Xiangyu Zhu, Zidu Wang, Zhen LeiCVPR 2026 · 4 citations
- Large-scale Codec Avatars: The Unreasonable Effectiveness of Large-scale Avatar PretrainingJunxuan Li, Rawal Khirodkar, Egor Zakharov, Jihyun Lee et al.CVPR 2026 · 3 citations
- PEAR: Pixel-aligned Expressive humAn mesh RecoveryJiahao Wu, Yunfei Liu, Lijian Lin, Ye Zhu et al.SIGGRAPH 2026 · 2 citations
- OMG-Avatar: One-shot Multi-LOD Gaussian Head AvatarJianqiang Ren, Lin Liu, Steven HoiCVPR 2026 · 1 citation
Builds on64
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Directly Denoising Diffusion ModelsDan Zhang, Jingjing Wang, Feng LuoICML 2024 · 11,724 citations
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 6,759 citations
- 3D Gaussian Splatting for Real-Time Radiance Field RenderingBernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, George DrettakisSIGGRAPH 2023 · 5,687 citations
Related papers
- Generalizable and Animatable Gaussian Head AvatarXuangeng Chu, Tatsuya HaradaNeurIPS 2024 · 115 citations
- LHM: Large Animatable Human Reconstruction Model for Single Image to 3D in SecondsLingteng Qiu, Xiaodong Gu, Peihao Li, Qi Zuo et al.ICCV 2025 · 19 citations
- FlashAvatar: High-Fidelity Head Avatar with Efficient Gaussian EmbeddingJun Xiang, Xuan Gao, Yudong Guo, Juyong ZhangCVPR 2024 · 51 citations
- LAM: Large Avatar Model for One-shot Animatable Gaussian HeadYisheng He, Xiaodong Gu, Xiaodan Ye, Chao Xu et al.SIGGRAPH 2025 · 14 citations
- FastAvatar: Towards Unified and Fast 3D Avatar Reconstruction with Large Gaussian Reconstruction TransformersYue Wu, Xuanhong Chen, Yufan Wu, Wen Li et al.ICLR 2026 · 7 citations
