UniHuman: A Unified Model For Editing Human Images in the Wild
Nannan Li, Qing Liu, Krishna Kumar Singh, Yilin Wang, Jianming Zhang, Bryan A. Plummer, Zhe Lin
Abstract
Human image editing includes tasks like changing a person's pose their clothing or editing the image according to a text prompt. However prior work often tackles these tasks separately overlooking the benefit of mutual reinforcement from learning them jointly. In this paper we propose UniHuman a unified model that addresses multiple facets of human image editing in real-world settings. To enhance the model's generation quality and generalization capacity we leverage guidance from human visual encoders and introduce a lightweight pose-warping module that can exploit different pose representations accommodating unseen textures and patterns. Furthermore to bridge the disparity between existing human editing benchmarks with real-world data we curated 400K high-quality human image-text pairs for training and collected 2K human images for out-of-domain testing both encompassing diverse clothing styles backgrounds and age groups. Experiments on both in-domain and out-of-domain test sets demonstrate that UniHuman outperforms task-specific models by a significant margin. In user studies UniHuman is preferred by the users in an average of 77% of cases. Our project is available at https://github.com/NannanLi999/UniHuman.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext afdf2809-961d-4edf-a4a0-2710815786ffCited by top-tier papers3
- CompleteMe: Reference-Based Human Image CompletionYu-Ju Tsai, Brian L. Price, Qing Liu, Luis Figueroa et al.ICCV 2025 · 2 citations
- Refine-by-Align: Reference-Guided Artifacts Refinement through Semantic AlignmentYizhi Song, Liu He, Zhifei Zhang, Soo Ye Kim et al.ICLR 2025
- Enhancing Virtual Try-On with Synthetic Pairs and Error-Aware Noise SchedulingNannan Li, Kevin J. Shih, Bryan A. PlummerCVPR 2025
Builds on23
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 6,759 citations
- UniT: Multimodal Multitask Learning with a Unified TransformerRonghang Hu, Amanpreet SinghICCV 2021 · 354 citations
- Text2Human: text-driven controllable human image generationYuming Jiang, Shuai Yang, Haonan Qiu, Wayne Wu et al.SIGGRAPH 2022 · 140 citations
Related papers
- Towards Scalable Human-aligned Benchmark for Text-guided Image EditingSuho Ryu, Kihyun Kim, Eugene Baek, Dongsoo Shin et al.CVPR 2025
- Clothe and PoseNakul Sharma, Aayush Bansal, Minh VoCVPR 2026
- PISE: Person Image Synthesis and Editing With Decoupled GANJinsong Zhang, Kun Li, Yu-Kun Lai, Jingyu YangCVPR 2021
- UMFuse: Unified Multi View Fusion for Human Editing applicationsRishabh Jain, Mayur Hemani, Duygu Ceylan, Krishna Kumar Singh et al.ICCV 2023 · 1 citation
- UniVG: A Generalist Diffusion Model for Unified Image Generation and EditingTsu-Jui Fu, Yusu Qian, Chen Chen, Wenze Hu et al.ICCV 2025 · 2 citations
