Controllable Person Image Synthesis with Pose-Constrained Latent Diffusion
Xiao Han, Xiatian Zhu, Jiankang Deng, Yi-Zhe Song, Tao Xiang
Abstract
Controllable person image synthesis aims at rendering a source image based on user-specified changes in body pose or appearance. Prior art approaches leverage pixel-level denoising diffusion models conditioned on the coarse skeleton via cross-attention. This leads to two limitations: low efficiency and inaccurate condition information. To address both issues, a novel Pose-Constrained Latent Diffusion model (PoCoLD) is introduced. Rather than using the skeleton as a sparse pose representation, we exploit DensePose which offers much richer body structure information. To effectively capitalize DensePose at a low cost, we propose an efficient pose-constrained attention module that is capable of modeling the complex interplay between appearance and pose. Extensive experiments show that our PoCoLD outperforms the state-of-the-art competitors in image synthesis fidelity. Critically, it runs 2× faster and consumes 3.6× smaller memory than the latest diffusion-model-based alternative during inference. Our code and models are available at https://github.com/BrandonHanx/PoCoLD .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 45b493db-c6bf-4cbb-810d-bd3832516800Cited by top-tier papers11
- IMAGPose: A Unified Conditional Framework for Pose-Guided Person GenerationFei Shen, Jinhui TangNeurIPS 2024 · 172 citations
- Coarse-to-Fine Latent Diffusion for Pose-Guided Person Image SynthesisYanzuo Lu, Manlin Zhang, Andy J. Ma, Xiaohua Xie et al.CVPR 2024 · 26 citations
- Cross-view Masked Diffusion Transformers for Person Image SynthesisTrung X. Pham, Kang Zhang, Chang D. YooICML 2024 · 12 citations
- Rethink Sparse Signals for Pose-Guided Text-to-Image GenerationWenjie Xuan, Jing Zhang, Juhua Liu, Bo Du et al.ICCV 2025 · 2 citations
- A Temporal and Content Co-Awareness Latent Diffusion for Controllable Hand Image GenerationShuang Hao, Pengfei Ren, Haifeng Sun, Pan Ting et al.CVPR 2026
Builds on22
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
Related papers
- DisPose: Disentangling Pose Guidance for Controllable Human Image AnimationHongxiang Li, Yaowei Li, Yuhang Yang, Junjie Cao et al.ICLR 2025
- Person Image Synthesis via Denoising Diffusion ModelAnkan Kumar Bhunia, Salman H. Khan, Hisham Cholakkal, Rao Muhammad Anwer et al.CVPR 2023
- DiffPortrait3D: Controllable Diffusion for Zero-Shot Portrait View SynthesisYuming Gu, Hongyi Xu, You Xie, Guoxian Song et al.CVPR 2024
- PersonaCraft: Personalized and Controllable Full-Body Multi-Human Scene Generation Using Occlusion-Aware 3D-Conditioned DiffusionGwanghyun Kim, Suh Yoon Jeon, Seunggyu Lee, Se Young ChunICCV 2025 · 2 citations
- HumanSD: A Native Skeleton-Guided Diffusion Model for Human Image GenerationXuan Ju, Ailing Zeng, Chenchen Zhao, Jianan Wang et al.ICCV 2023 · 137 citations
