CapHuman: Capture Your Moments in Parallel Universes
Chao Liang, Fan Ma, Linchao Zhu, Yingying Deng, Yi Yang
Abstract
… a pop singer, sing, play the guitar, piano, take part in the show … a scientist, work with Hawking, Hinton, present in a conference … an astronaut, travel over the universe, collaborate with Obama Reference Image 3D facial prior Figure 1. Given only one reference facial photograph, our CapHuman can generate photo-realistic specific individual portraits with contentrich representations and diverse head positions, poses, facial expressions, and illuminations in different contexts.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 648ec141-bbaf-4fab-9453-6e935d972dc1Cited by top-tier papers12
- DreamRenderer: Taming Multi-Instance Attribute Control in Large-Scale Text-to-Image ModelsDewei Zhou, Mingwei Li, Zongxin Yang, Yi YangICCV 2025 · 5 citations
- DataStealing: Steal Data from Diffusion Models in Federated Learning with Multiple TrojansYuan Gan, Jiaxu Miao, Yi YangNeurIPS 2024 · 5 citations
- MR-FIQA: Face Image Quality Assessment with Multi-Reference Representations from Synthetic Data GenerationFu-Zhao Ou, Chongyi Li, Shiqi Wang, Sam KwongICCV 2025 · 4 citations
- DynamicID: Zero-Shot Multi-ID Image Personalization With Flexible Facial EditabilityXirui Hu, Jiahao Wang, Hao Chen, Weizhan Zhang et al.ICCV 2025 · 3 citations
- UniPortrait: A Unified Framework for Identity-Preserving Single- and Multi-Human Image PersonalizationJunjie He, Yifeng Geng, Liefeng BoICCV 2025 · 2 citations
Builds on24
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 6,759 citations
Related papers
- SnapGen: Taming High-Resolution Text-to-Image Models for Mobile Devices with Efficient Architectures and TrainingJierun Chen, Dongting Hu, Xijie Huang, Huseyin Coskun et al.CVPR 2025
- CAP4D: Creating Animatable 4D Portrait Avatars with Morphable Multi-View Diffusion ModelsFelix Taubner, Ruihang Zhang, Mathieu Tuli, David B. LindellCVPR 2025
- StarGAN v2: Diverse Image Synthesis for Multiple DomainsYunjey Choi, Youngjung Uh, Jaejun Yoo, Jung-Woo HaCVPR 2020
- Enhancing 3D Fidelity of Text-to-3D using Cross-View CorrespondencesSeungwook Kim, Kejie Li, Xueqing Deng, Yichun Shi et al.CVPR 2024
- Learning Disentangled Identifiers for Action-Customized Text-to-Image GenerationSiteng Huang, Biao Gong, Yutong Feng, Xi Chen et al.CVPR 2024
