UniHuman: A Unified Model For Editing Human Images in the Wild
Nannan Li, Qing Liu, Krishna Kumar Singh, Yilin Wang, Jianming Zhang, Bryan A. Plummer, Zhe Lin
摘要
Human image editing includes tasks like changing a person's pose their clothing or editing the image according to a text prompt. However prior work often tackles these tasks separately overlooking the benefit of mutual reinforcement from learning them jointly. In this paper we propose UniHuman a unified model that addresses multiple facets of human image editing in real-world settings. To enhance the model's generation quality and generalization capacity we leverage guidance from human visual encoders and introduce a lightweight pose-warping module that can exploit different pose representations accommodating unseen textures and patterns. Furthermore to bridge the disparity between existing human editing benchmarks with real-world data we curated 400K high-quality human image-text pairs for training and collected 2K human images for out-of-domain testing both encompassing diverse clothing styles backgrounds and age groups. Experiments on both in-domain and out-of-domain test sets demonstrate that UniHuman outperforms task-specific models by a significant margin. In user studies UniHuman is preferred by the users in an average of 77% of cases. Our project is available at https://github.com/NannanLi999/UniHuman.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- CompleteMe: Reference-Based Human Image CompletionYu-Ju Tsai, Brian L. Price, Qing Liu, Luis Figueroa 等ICCV 2025 · 被引用 2 次
- Refine-by-Align: Reference-Guided Artifacts Refinement through Semantic AlignmentYizhi Song, Liu He, Zhifei Zhang, Soo Ye Kim 等ICLR 2025
- Enhancing Virtual Try-On with Synthetic Pairs and Error-Aware Noise SchedulingNannan Li, Kevin J. Shih, Bryan A. PlummerCVPR 2025
它引用的顶会 Paper23
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 被引用 6,759 次
- UniT: Multimodal Multitask Learning with a Unified TransformerRonghang Hu, Amanpreet SinghICCV 2021 · 被引用 354 次
- Text2Human: text-driven controllable human image generationYuming Jiang, Shuai Yang, Haonan Qiu, Wayne Wu 等SIGGRAPH 2022 · 被引用 140 次
相关 Paper
- Towards Scalable Human-aligned Benchmark for Text-guided Image EditingSuho Ryu, Kihyun Kim, Eugene Baek, Dongsoo Shin 等CVPR 2025
- Clothe and PoseNakul Sharma, Aayush Bansal, Minh VoCVPR 2026
- PISE: Person Image Synthesis and Editing With Decoupled GANJinsong Zhang, Kun Li, Yu-Kun Lai, Jingyu YangCVPR 2021
- UMFuse: Unified Multi View Fusion for Human Editing applicationsRishabh Jain, Mayur Hemani, Duygu Ceylan, Krishna Kumar Singh 等ICCV 2023 · 被引用 1 次
- UniVG: A Generalist Diffusion Model for Unified Image Generation and EditingTsu-Jui Fu, Yusu Qian, Chen Chen, Wenze Hu 等ICCV 2025 · 被引用 2 次
