The Chosen One: Consistent Characters in Text-to-Image Diffusion Models
Omri Avrahami, Amir Hertz, Yael Vinker, Moab Arar, Shlomi Fruchter, Ohad Fried, Daniel Cohen-Or, Dani Lischinski
摘要
Recent advances in text-to-image generation models have unlocked vast potential for visual creativity. However, the users that use these models struggle with the generation of consistent characters, a crucial aspect for numerous real-world applications such as story visualization, game development, asset design, advertising, and more. Current methods typically rely on multiple pre-existing images of the target character or involve labor-intensive manual processes. In this work, we propose a fully automated solution for consistent character generation, with the sole input being a text prompt. We introduce an iterative procedure that, at each stage, identifies a coherent set of images sharing a similar identity and extracts a more consistent identity from this set. Our quantitative analysis demonstrates that our method strikes a better balance between prompt alignment and identity consistency compared to the baseline methods, and these findings are reinforced by a user study. To conclude, we showcase several practical applications of our approach.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper27
- Training-Free Consistent Text-to-Image GenerationYoad Tewel, Omri Kaduri, Rinon Gal, Yoni Kasten 等SIGGRAPH 2024 · 被引用 57 次
- Story-Iter: A Training-free Iterative Paradigm for Long Story VisualizationJiawei Mao, Xiaoke Huang, Yunfei Xie, Yuanqi Chang 等ICLR 2026 · 被引用 18 次
- Compositional Image Decomposition with Diffusion ModelsJocelin Su, Nan Liu, Yanbo Wang, Joshua B. Tenenbaum 等ICML 2024 · 被引用 16 次
- OneActor: Consistent Subject Generation via Cluster-Conditioned GuidanceJiahao Wang, Caixia Yan, Haonan Lin, Weizhan Zhang 等NeurIPS 2024 · 被引用 16 次
- Group Editing: Edit Multiple Images in One GoYue Ma, Xinyu Wang, Qianli Ma, Qinghe Wang 等CVPR 2026 · 被引用 15 次
它引用的顶会 Paper43
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 被引用 11,743 次
相关 Paper
- IdentityStory: Taming Your Identity-Preserving Generator for Human-Centric Story GenerationDonghao Zhou, Jingyu Lin, Guibao Shen, Quande Liu 等AAAI 2026 · 被引用 3 次
- One-Prompt-One-Story: Free-Lunch Consistent Text-to-Image Generation Using a Single PromptTao Liu, Kai Wang, Senmao Li, Joost van de Weijer 等ICLR 2025
- Infinite-Story: A Training-Free Consistent Text-to-Image GenerationJihun Park, Kyoungmin Lee, Jongmin Gim, Hyeonseo Jo 等AAAI 2026 · 被引用 1 次
- IP-Prompter: Training-Free Theme-Specific Image Generation via Dynamic Visual PromptingYuxin Zhang, Minyan Luo, Weiming Dong, Xiao Yang 等SIGGRAPH 2025 · 被引用 2 次
- SerialGen: Personalized Image Generation by First Standardization Then PersonalizationCong Xie, Han Zou, Ruiqi Yu, Yan Zhang 等CVPR 2025
