PhotoMaker: Customizing Realistic Human Photos via Stacked ID Embedding
Zhen Li, Mingdeng Cao, Xintao Wang, Zhongang Qi, Ming-Ming Cheng, Ying Shan
Abstract
Recent advances in text-to-image generation have made remarkable progress in synthesizing realistic human photos conditioned on given text prompts. However, existing personalized generation methods cannot simultaneously satisfy the requirements of high efficiency, promising identity * Interns in ARC Lab, Tencent PCG † Corresponding authors (ID) fidelity, and flexible text controllability. In this work, we introduce PhotoMaker, an efficient personalized textto-image generation method, which mainly encodes an arbitrary number of input ID images into a stack ID embedding for preserving ID information. Such an embedding, serving as a unified ID representation, can not only encapsulate the characteristics of the same input ID comprehensively, but also accommodate the characteristics of different IDs for subsequent integration. This paves the way for
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ca5e19c3-6fad-4588-bb09-8c0b056aed5cCited by top-tier papers151
- StoryDiffusion: Consistent Self-Attention for Long-Range Image and Video GenerationYupeng Zhou, Daquan Zhou, Ming-Ming Cheng, Jiashi Feng et al.NeurIPS 2024 · 291 citations
- PuLID: Pure and Lightning ID Customization via Contrastive AlignmentZinan Guo, Yanze Wu, Zhuowei Chen, Lang Chen et al.NeurIPS 2024 · 184 citations
- EditVerse: Unifying Image and Video Editing and Generation with In-Context LearningXuan Ju, Tianyu Wang, Yuqian Zhou, He Zhang et al.ICLR 2026 · 56 citations
- Towards Understanding the Working Mechanism of Text-to-Image Diffusion ModelMingyang Yi, Aoxue Li, Yi Xin, Zhenguo LiNeurIPS 2024 · 55 citations
- MultiBooth: Towards Generating All Your Concepts in an Image from TextChenyang Zhu, Kai Li, Yue Ma, Chunming He et al.AAAI 2025 · 52 citations
Builds on32
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
Related papers
- Nested Attention: Semantic-aware Attention Values for Concept PersonalizationOr Patashnik, Rinon Gal, Daniil Ostashev, Sergey Tulyakov et al.SIGGRAPH 2025 · 6 citations
- PortraitBooth: A Versatile Portrait Model for Fast Identity-Preserved PersonalizationXu Peng, Junwei Zhu, Boyuan Jiang, Ying Tai et al.CVPR 2024 · 28 citations
- DreamIdentity: Enhanced Editability for Efficient Face-Identity Preserved Image GenerationZhuowei Chen, Shancheng Fang, Wei Liu, Qian He et al.AAAI 2024 · 26 citations
- ID-Patch: Robust ID Association for Group Photo PersonalizationYimeng Zhang, Tiancheng Zhi, Jing Liu, Shen Sang et al.CVPR 2025
- CoRe: Context-Regularized Text Embedding Learning for Text-to-Image PersonalizationFeize Wu, Yun Pang, Junyi Zhang, Lianyu Pang et al.AAAI 2025 · 13 citations
