Chupa: Carving 3D Clothed Humans from Skinned Shape Priors using 2D Diffusion Probabilistic Models
Byungjun Kim, Patrick Kwon, Kwangho Lee, Myunggi Lee, Sookwan Han, Daesik Kim, Hanbyul Joo
Abstract
We propose a 3D generation pipeline that uses diffusion models to generate realistic human digital avatars. Due to the wide variety of human identities, poses, and stochastic details, the generation of 3D human meshes has been a challenging problem. To address this, we decompose the problem into 2D normal map generation and normal map-based 3D reconstruction. Specifically, we first simultaneously generate realistic normal maps for the front and backside of a clothed human, dubbed dual normal maps, using a pose-conditional diffusion model. For 3D reconstruction, we "carve" the prior SMPL-X mesh to a detailed 3D mesh according to the normal maps through mesh optimization. To further enhance the high-frequency details, we present a diffusion resampling scheme on both body and facial regions, thus encouraging the generation of realistic digital avatars. We also seamlessly incorporate a recent text-to-image diffusion model to support text-based human identity control. Our method, namely, Chupa, is capable of generating realistic 3D clothed humans with better perceptual quality and identity variety.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 252b8e75-d23e-4a58-9783-4136f0f7a6ceCited by top-tier papers12
- UP2You: Fast Reconstruction of Yourself from Unconstrained Photo CollectionsZeyu Cai, Ziyang Li, Xiaoben Li, Boqian Li et al.ICLR 2026 · 11 citations
- SyncHuman: Synchronizing 2D and 3D Generative Models for Single-view Human ReconstructionWenyue Chen, Peng Li, Wangguandong Zheng, Chengfeng Zhao et al.NeurIPS 2025 · 8 citations
- E3Gen: Efficient, Expressive and Editable Avatars GenerationWeitian Zhang, Yichao Yan, Yunhui Liu, Xingdong Sheng et al.ACM MM 2024 · 4 citations
- Disentangled Clothed Avatar Generation with Layered RepresentationWeitian Zhang, Yichao Yan, Sijing Wu, Manwen Liao et al.ICCV 2025 · 3 citations
- Joint2Human: High-quality 3D Human Generation via Compact Spherical Embedding of 3D JointsMuxin Zhang, Qiao Feng, Zhuo Su, Chao Wen et al.CVPR 2024 · 2 citations
Builds on49
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
Related papers
- DreamAvatar: Text-and-Shape Guided 3D Human Avatar Generation via Diffusion ModelsYukang Cao, Yan-Pei Cao, Kai Han, Ying Shan et al.CVPR 2024
- DINAR: Diffusion Inpainting of Neural Textures for One-Shot Human AvatarsDavid Svitov, Dmitrii Gudkov, Renat Bashirov, Victor LempitskyICCV 2023 · 37 citations
- Semantic Human Mesh Reconstruction with TexturesXiaoyu Zhan, Jianxin Yang, Yuanqi Li, Jie Guo et al.CVPR 2024 · 7 citations
- DiffHuman: Probabilistic Photorealistic 3D Reconstruction of HumansAkash Sengupta, Thiemo Alldieck, Nikos Kolotouros, Enric Corona et al.CVPR 2024 · 10 citations
- High-Quality Full-Head 3D Avatar Generation from Any Single Portrait ImageYujie Gao, Chencheng Wang, Xianbing Sun, Jiahui Zhan et al.AAAI 2026
