TextDiffuser: Diffusion Models as Text Painters
Jingye Chen, Yupan Huang, Tengchao Lv, Lei Cui, Qifeng Chen, Furu Wei
Abstract
Diffusion models have gained increasing attention for their impressive generation abilities but currently struggle with rendering accurate and coherent text. To address this issue, we introduce TextDiffuser, focusing on generating images with visually appealing text that is coherent with backgrounds. TextDiffuser consists of two stages: first, a Transformer model generates the layout of keywords extracted from text prompts, and then diffusion models generate images conditioned on the text prompt and the generated layout. Additionally, we contribute the first large-scale text images dataset with OCR annotations, MARIO-10M, containing 10 million image-text pairs with text recognition, detection, and character-level segmentation annotations. We further collect the MARIO-Eval benchmark to serve as a comprehensive tool for evaluating text rendering quality. Through experiments and user studies, we show that TextDiffuser is flexible and controllable to create high-quality text images using text prompts alone or together with text template images, and conduct text inpainting to reconstruct incomplete images with text. The code, model, and dataset will be available at https://aka.ms/textdiffuser.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext dc212c52-a2b6-4597-97af-de4a17d291e7Cited by top-tier papers58
- Flow-GRPO: Training Flow Matching Models via Online RLJie Liu, Gongye Liu, Jiajun Liang, Yangguang Li et al.NeurIPS 2025 · 647 citations
- GenArtist: Multimodal LLM as an Agent for Unified Image Generation and EditingZhenyu Wang, Aoxue Li, Zhenguo Li, Xihui LiuNeurIPS 2024 · 162 citations
- Harmonizing Visual Text Comprehension and GenerationZhen Zhao, Jingqun Tang, Binghong Wu, Chunhui Lin et al.NeurIPS 2024 · 69 citations
- GRPO-Guard: Mitigating Implicit Over-Optimization in Flow Matching via Regulated ClippingJing Wang, Jiajun Liang, Jie Liu, Henglin Liu et al.CVPR 2026 · 46 citations
- Advantage Weighted Matching: Aligning RL with Pretraining in Diffusion ModelsShuchen Xue, Chongjian GE, Shilong Zhang, Yichen Li et al.ICML 2026 · 43 citations
Builds on41
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
Related papers
- AnyText: Multilingual Visual Text Generation and EditingYuxiang Tuo, Wangmeng Xiang, Jun-Yan He, Yifeng Geng et al.ICLR 2024 · 148 citations
- TextAtlas5M: A Large-Scale Dataset for Long Text Image GenerationDongxing Mao, Alex Jinpeng Wang, weiming Han, Jiawei Zhang et al.ICML 2026
- EasyText: Controllable Diffusion Transformer for Multilingual Text RenderingRunnan Lu, Yuxuan Zhang, Jiaming Liu, Haofan Wang et al.AAAI 2026 · 20 citations
- Paint by Inpaint: Learning to Add Image Objects by Removing Them FirstNavve Wasserman, Noam Rotstein, Roy Ganz, Ron KimmelCVPR 2025
- CreatiLayout: Siamese Multimodal Diffusion Transformer for Creative Layout-to-Image GenerationHui Zhang, Dexiang Hong, Yitong Wang, Jie Shao et al.ICCV 2025 · 7 citations
