HyperHuman: Hyper-Realistic Human Generation with Latent Structural Diffusion
Xian Liu, Jian Ren, Aliaksandr Siarohin, Ivan Skorokhodov, Yanyu Li, Dahua Lin, Xihui Liu, Ziwei Liu, Sergey Tulyakov
Abstract
Despite significant advances in large-scale text-to-image models, achieving hyper-realistic human image generation remains a desirable yet unsolved task. Existing models like Stable Diffusion and DALL-E 2 tend to generate human images with incoherent parts or unnatural poses. To tackle these challenges, our key insight is that human image is inherently structural over multiple granularities, from the coarse-level body skeleton to fine-grained spatial geometry. Therefore, capturing such correlations between the explicit appearance and latent structure in one model is essential to generate coherent and natural human images. To this end, we propose a unified framework, HyperHuman, that generates in-the-wild human images of high realism and diverse layouts. Specifically, 1) we first build a large-scale human-centric dataset, named HumanVerse, which consists of 340M images with comprehensive annotations like human pose, depth, and surface normal. 2) Next, we propose a Latent Structural Diffusion Model that simultaneously denoises the depth and surface normal along with the synthesized RGB image. Our model enforces the joint learning of image appearance, spatial relationship, and geometry in a unified network, where each branch in the model complements to each other with both structural awareness and textural richness. 3) Finally, to further boost the visual quality, we propose a Structure-Guided Refiner to compose the predicted conditions for more detailed generation of higher resolution. Extensive experiments demonstrate that our framework yields the state-of-the-art performance, generating hyper-realistic human images under diverse scenarios. Project Page: https://snap-research.github.io/HyperHuman/
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f95a3758-8f83-46c7-80f1-81569c16fb6bCited by top-tier papers23
- CLAY: A Controllable Large-scale Generative Model for Creating High-quality 3D AssetsLongwen Zhang, Ziyu Wang, Qixuan Zhang, Qiwei Qiu et al.SIGGRAPH 2024 · 148 citations
- MoLE: Enhancing Human-centric Text-to-image Diffusion via Mixture of Low-rank ExpertsJie Zhu, Yixiong Chen, Mingyu Ding, Ping Luo et al.NeurIPS 2024 · 17 citations
- Towards Physical Understanding in Video Generation: A 3D Point Regularization ApproachYunuo Chen, Junli Cao, Vidit Goel, Sergei Korolev et al.NeurIPS 2025 · 9 citations
- DiffusionRegPose: Enhancing Multi-Person Pose Estimation Using a Diffusion-Based End-to-End Regression ApproachDayi Tan, Hansheng Chen, Wei Tian, Lu XiongCVPR 2024 · 6 citations
- Adapting Diffusion Models for Improved Prompt Compliance and Controllable Image SynthesisDeepak Sridhar, Abhishek Peri, Rohith Rachala, Nuno VasconcelosNeurIPS 2024 · 5 citations
Builds on32
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
Related papers
- HumanNorm: Learning Normal Diffusion Model for High-quality and Realistic 3D Human GenerationXin Huang, Ruizhi Shao, Qi Zhang, Hongwen Zhang et al.CVPR 2024
- GeneMAN: Generalizable Single-Image 3D Human Reconstruction from Multi-Source Human DataWentao Wang, Hang Ye, Fangzhou Hong, Xue Yang et al.NeurIPS 2025 · 6 citations
- CosmicMan: A Text-to-Image Foundation Model for HumansShikai Li, Jianglin Fu, Kaiyuan Liu, Wentao Wang et al.CVPR 2024
- HumanRef: Single Image to 3D Human Generation via Reference-Guided DiffusionJingbo Zhang, Xiaoyu Li, Qi Zhang, Yanpei Cao et al.CVPR 2024 · 15 citations
- Towards Spatially Consistent Image Generation: On Incorporating Intrinsic Scene Properties into Diffusion ModelsHyundo Lee, Suhyung Choi, Inwoo Hwang, Byoung-Tak ZhangAAAI 2026
