HumanSD: A Native Skeleton-Guided Diffusion Model for Human Image Generation
Xuan Ju, Ailing Zeng, Chenchen Zhao, Jianan Wang, Lei Zhang, Qiang Xu
摘要
Controllable human image generation (HIG) has numerous real-life applications. State-of-the-art solutions, such as ControlNet and T2I-Adapter, introduce an additional learnable branch on top of the frozen pre-trained stable diffusion (SD) model, which can enforce various conditions, including skeleton guidance of HIG. While such a plug-and-play approach is appealing, the inevitable and uncertain conflicts between the original images produced from the frozen SD branch and the given condition incur significant challenges for the learnable branch, which essentially conducts image feature editing for condition enforcement.In this work, we propose a native skeleton-guided diffusion model for controllable HIG called HumanSD. Instead of performing image editing with dual-branch diffusion, we fine-tune the original SD model using a novel heatmap-guided denoising loss. This strategy effectively and efficiently strengthens the given skeleton condition during model training while mitigating the catastrophic forgetting effects. HumanSD is fine-tuned on the assembly of three large-scale human-centric datasets with text-image-pose information, two of which are established in this work. Experimental results show that HumanSD outperforms ControlNet in terms of pose control and image quality, particularly when the given skeleton guidance is sophisticated. Code and data are available at: https://idea-research.github.io/HumanSD/.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper49
- PnP Inversion: Boosting Diffusion-based Editing with 3 Lines of CodeXuan Ju, Ailing Zeng, Yuxuan Bian, Shaoteng Liu 等ICLR 2024 · 被引用 166 次
- HumanMAC: Masked Motion Completion for Human Motion PredictionLing-Hao Chen, Jiawei Zhang, Yewen Li, Yiren Pang 等ICCV 2023 · 被引用 106 次
- HyperHuman: Hyper-Realistic Human Generation with Latent Structural DiffusionXian Liu, Jian Ren, Aliaksandr Siarohin, Ivan Skorokhodov 等ICLR 2024 · 被引用 81 次
- EditVerse: Unifying Image and Video Editing and Generation with In-Context LearningXuan Ju, Tianyu Wang, Yuqian Zhou, He Zhang 等ICLR 2026 · 被引用 56 次
- HumanGaussian: Text-Driven 3D Human Generation with Gaussian SplattingXian Liu, Xiaohang Zhan, Jiaxiang Tang, Ying Shan 等CVPR 2024 · 被引用 42 次
它引用的顶会 Paper22
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 被引用 11,743 次
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 被引用 6,759 次
相关 Paper
- Stable-Pose: Leveraging Transformers for Pose-Guided Text-to-Image GenerationJiajun Wang, Morteza Ghahremani, Yitong Li, Björn Ommer 等NeurIPS 2024 · 被引用 22 次
- MotionEditor: Editing Video Motion via Content-Aware DiffusionShuyuan Tu, Qi Dai, Zhi-Qi Cheng, Han Hu 等CVPR 2024 · 被引用 21 次
- Person in Place: Generating Associative Skeleton-Guidance Maps for Human-Object Interaction Image EditingChangHee Yang, Chanhee Kang, Kyeongbo Kong, Hanni Oh 等CVPR 2024
- DisPose: Disentangling Pose Guidance for Controllable Human Image AnimationHongxiang Li, Yaowei Li, Yuhang Yang, Junjie Cao 等ICLR 2025
- DynFusion: Rethinking Condition Fusion for Adaptive Multi-Conditional Text-to-Image GenerationZheng Fang, Lichuan Xiang, Xu Cai, Bing Wang 等CVPR 2026
