Towards Effective Usage of Human-Centric Priors in Diffusion Models for Text-based Human Image Generation
Junyan Wang, Zhenhong Sun, Zhiyu Tan, Xuanbai Chen, Weihua Chen, Hao Li, Cheng Zhang, Yang Song
Abstract
Vanilla text-to-image diffusion models struggle with generating accurate human images, commonly resulting in imperfect anatomies such as unnatural postures or disproportionate limbs. Existing methods address this issue mostly by fine-tuning the model with extra images or adding additional control - human-centric priors such as pose or depth maps - during the image generation phase. This paper explores the integration of these human-centric priors directly into the model fine-tuning stage, essentially eliminating the need for extra conditions at the inference stage. We realize this idea by proposing a human-centric alignment loss to strengthen human-related information from the textual prompts within the cross-attention maps. To ensure semantic detail richness and human structural accuracy during fine-tuning, we introduce scale-aware and step-wise constraints within the diffusion process, according to an indepth analysis of the cross-attention layer. Extensive experiments show that our method largely improves over state-of-the-art text-to-image models to synthesize high-quality human images based on user-written prompts. Project page: https://hcplayercvpr2024.github.io.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers6
- Chain of World: World Model Thinking in Latent MotionFuxiang Yang, Donglin Di, Lulu Tang, Xuancheng Zhang et al.CVPR 2026 · 11 citations
- Taxadiffusion: Progressively Trained Diffusion Model for Fine-Grained Species GenerationAmin Karimi Monsefi, Mridul Khurana, Rajiv Ramnath, Anuj Karpatne et al.ICCV 2025 · 1 citation
- FashionMAC: Deformation-Free Fashion Image Generation with Fine-Grained Model Appearance CustomizationRong Zhang, Jinxiao Li, Jingnan Wang, Zhiwen Zuo et al.AAAI 2026
- Prototype-Based Image Prompting for Weakly Supervised Histopathological Image SegmentationQingchen Tang, Lei Fan, Maurice Pagnucco, Yang SongCVPR 2025
- Interpretable Image Classification via Non-parametric Part Prototype LearningZhijie Zhu, Lei Fan, Maurice Pagnucco, Yang SongCVPR 2025
Builds on31
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
Related papers
- Composing People Together: Iterative Pose-Image Generation for Multi-Person Interaction ScenesWenxuan Peng, Bharath Hariharan, Hadar Averbuch-ElorSIGGRAPH 2026
- Coarse-to-Fine Latent Diffusion for Pose-Guided Person Image SynthesisYanzuo Lu, Manlin Zhang, Andy J. Ma, Xiaohua Xie et al.CVPR 2024 · 26 citations
- PHAC: Promptable Human Amodal CompletionSeung Young Noh, Ju Yong ChangCVPR 2026
- Text Embedding is Not All You Need: Attention Control for Text-to-Image Semantic Alignment with Text Self-Attention MapsJeeyung Kim, Erfan Esmaeili, Qiang QiuCVPR 2025
- SweetDreamer: Aligning Geometric Priors in 2D diffusion for Consistent Text-to-3DWeiyu Li, Rui Chen, Xuelin Chen, Ping TanICLR 2024 · 155 citations
