GlyphControl: Glyph Conditional Control for Visual Text Generation
Yukang Yang, Dongnan Gui, Yuhui Yuan, Weicong Liang, Haisong Ding, Han Hu, Kai Chen
摘要
Recently, there has been an increasing interest in developing diffusion-based text-to-image generative models capable of generating coherent and well-formed visual text. In this paper, we propose a novel and efficient approach called GlyphControl to address this task. Unlike existing methods that rely on character-aware text encoders like ByT5 and require retraining of text-to-image models, our approach leverages additional glyph conditional information to enhance the performance of the off-the-shelf Stable-Diffusion model in generating accurate visual text. By incorporating glyph instructions, users can customize the content, location, and size of the generated text according to their specific requirements. To facilitate further research in visual text generation, we construct a training benchmark dataset called LAION-Glyph. We evaluate the effectiveness of our approach by measuring OCR-based metrics, CLIP score, and FID of the generated visual text. Our empirical evaluations demonstrate that GlyphControl outperforms the recent DeepFloyd IF approach in terms of OCR accuracy, CLIP score, and FID, highlighting the efficacy of our method.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper48
- AnyText: Multilingual Visual Text Generation and EditingYuxiang Tuo, Wangmeng Xiang, Jun-Yan He, Yifeng Geng 等ICLR 2024 · 被引用 148 次
- Large Multilingual Models Pivot Zero-Shot Multimodal Learning across LanguagesJinyi Hu, Yuan Yao, Chongyi Wang, Shan Wang 等ICLR 2024 · 被引用 79 次
- Harmonizing Visual Text Comprehension and GenerationZhen Zhao, Jingqun Tang, Binghong Wu, Chunhui Lin 等NeurIPS 2024 · 被引用 69 次
- TextCtrl: Diffusion-based Scene Text Editing with Prior Guidance ControlWeichao Zeng, Yan Shu, Zhenhang Li, Dongbao Yang 等NeurIPS 2024 · 被引用 55 次
- GlyphDraw2: Automatic Generation of Complex Glyph Posters with Diffusion Models and Large Language ModelsJian Ma, Yonglin Deng, Chen Chen, Nanyang Du 等AAAI 2025 · 被引用 28 次
它引用的顶会 Paper17
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 被引用 13,211 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 被引用 11,743 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
相关 Paper
- UniGlyph: Unified Segmentation-Conditioned Diffusion for Precise Visual Text SynthesisYuanrui Wang, Cong Han, Yafei Li, Zhipeng Jin 等ICCV 2025
- Precise Parameter Localization for Textual Generation in Diffusion ModelsLukasz Staniszewski, Bartosz Cywinski, Franziska Boenisch, Kamil Deja 等ICLR 2025
- FreeText: Training-Free Text Rendering via Attention Localization and Spectral Glyph InjectionRuiQiang Zhang, Hengyi Wang, Chang Liu, Guanjie Wang 等ICML 2026 · 被引用 3 次
- SpaText: Spatio-Textual Representation for Controllable Image GenerationOmri Avrahami, Thomas Hayes, Oran Gafni, Sonal Gupta 等CVPR 2023
- LayoutLLM-T2I: Eliciting Layout Guidance from LLM for Text-to-Image GenerationLeigang Qu, Shengqiong Wu, Hao Fei, Liqiang Nie 等ACM MM 2023 · 被引用 91 次
