Expressive Text-to-Image Generation with Rich Text
Songwei Ge, Taesung Park, Jun-Yan Zhu, Jia-Bin Huang
摘要
Plain text has become a prevalent interface for text-based image synthesis and editing. Its limited customization options, however, hinder users from accurately describing desired outputs. For example, plain text makes it hard to specify continuous quantities, such as the precise RGB color value or importance of each word. Creating detailed text prompts for complex scenes is tedious for humans to write and challenging for text encoders to interpret. Furthermore, describing a reference concept or texture in plain text is non-trivial. To address these challenges, we propose using a rich-text editor supporting formats such as font style, size, color, texture fill, footnote, and embedded image. We extract each word's attributes from rich text to enable local style control, explicit token reweighting, precise color rendering, and detailed region synthesis with reference concepts or texture. We achieve these capabilities through a region-based diffusion process. We first obtain each word's mask that characterizes the region guided by the word. For each region, we enforce its text attributes by creating customized prompts, applying guidance within the region, and maintaining its fidelity against plain-text generations or input images through region-based injections. We present various examples of image generation and editing from rich text and demonstrate that our method outperforms strong baselines with quantitative evaluations.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper49
- Preserve Your Own Correlation: A Noise Prior for Video Diffusion ModelsSongwei Ge, Seungjun Nah, Guilin Liu, Tyler Poon 等ICCV 2023 · 被引用 319 次
- Boosting Consistency in Story Visualization with Rich-Contextual Conditional Diffusion ModelsFei Shen, Hu Ye, Sibo Liu, Jun Zhang 等AAAI 2025 · 被引用 74 次
- Cross-Image Attention for Zero-Shot Appearance TransferYuval Alaluf, Daniel Garibi, Or Patashnik, Hadar Averbuch-Elor 等SIGGRAPH 2024 · 被引用 72 次
- Grounded Text-to-Image Synthesis with Attention RefocusingQuynh Phung, Songwei Ge, Jia-Bin HuangCVPR 2024 · 被引用 59 次
- GenQuery: Supporting Expressive Visual Search with Generative ModelsKihoon Son, DaEun Choi, Tae Soo Kim, Young-Ho Kim 等CHI 2024 · 被引用 48 次
它引用的顶会 Paper40
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao 等ICCV 2023 · 被引用 13,211 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 被引用 11,743 次
相关 Paper
- Prompt-to-Prompt Image Editing with Cross-Attention ControlAmir Hertz, Ron Mokady, Jay Tenenbaum, Kfir Aberman 等ICLR 2023 · 被引用 361 次
- Prompt Tuning Inversion for Text-Driven Image Editing Using Diffusion ModelsWenkai Dong, Song Xue, Xiaoyue Duan, Shumin HanICCV 2023 · 被引用 104 次
- DiffEditor: Boosting Accuracy and Flexibility on Diffusion-Based Image EditingChong Mou, Xintao Wang, Jiechong Song, Ying Shan 等CVPR 2024 · 被引用 36 次
- TextCtrl: Diffusion-based Scene Text Editing with Prior Guidance ControlWeichao Zeng, Yan Shu, Zhenhang Li, Dongbao Yang 等NeurIPS 2024 · 被引用 55 次
- Self-Prompting Diffusion Transformer for Open-Vocabulary Scene Text Edit via In-Context LearningHongxi Li, Tong Wang, WU CHENGJING, Tianbao Liu 等ICML 2026
