Empowering Backbone Models for Visual Text Generation with Input Granularity Control and Glyph-Aware Training
Wenbo Li, Guohao Li, Zhibin Lan, Xue Xu, Wanru Zhuang, Jiachen Liu, Xinyan Xiao, Jinsong Su
Abstract
Diffusion-based text-to-image models have demonstrated impressive achievements in diversity and aesthetics but struggle to generate images with legible visual texts. Existing backbone models have limitations such as misspelling, failing to generate texts, and lack of support for Chinese text, but their development shows promising potential. In this paper, we propose a series of methods, aiming to empower backbone models to generate visual texts in English and Chinese. We first conduct a preliminary study revealing that Byte Pair Encoding (BPE) tokenization and the insufficient learning of cross-attention modules restrict the performance of the backbone models. Based on these observations, we make the following improvements: (1) We design a mixed granularity input strategy to provide more suitable text representations; (2) We propose to augment the conventional training objective with three glyph-aware training losses, which enhance the learning of cross-attention modules and encourage the model to focus on visual texts. Through experiments, we demonstrate that our methods can effectively empower backbone models to generate semantic relevant, aesthetically appealing, and accurate visual text images, while maintaining their fundamental image generation quality.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- Type-R: Automatically Retouching Typos for Text-to-Image GenerationWataru Shimoda, Naoto Inoue, Daichi Haraguchi, Hayato Mitani et al.CVPR 2025
- MirrorQA: Benchmarking Multimodal LLMs on Mirror-Orientation ReasoningJingping Liu, Xingchen Peng, Yan Zhou, Ziyan Liu et al.ACL 2026
- BizGen: Advancing Article-level Visual Text Rendering for Infographics GenerationYuyang Peng, Shishi Xiao, Keming Wu, Qisheng Liao et al.CVPR 2025
Builds on13
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 6,759 citations
Related papers
- ViType: High-Fidelity Visual Text Rendering via Glyph-Aware Multimodal DiffusionLishuai Gao, Jun-Yan He, Yingsen Zeng, Yujie Zhong et al.AAAI 2026
- FreeText: Training-Free Text Rendering via Attention Localization and Spectral Glyph InjectionRuiQiang Zhang, Hengyi Wang, Chang Liu, Guanjie Wang et al.ICML 2026 · 3 citations
- GlyphMastero: A Glyph Encoder for High-Fidelity Scene Text EditingTong Wang, Ting Liu, Xiaochao Qu, Chengjing Wu et al.CVPR 2025
- UniGlyph: Unified Segmentation-Conditioned Diffusion for Precise Visual Text SynthesisYuanrui Wang, Cong Han, Yafei Li, Zhipeng Jin et al.ICCV 2025
- Character-Aware Models Improve Visual Text RenderingRosanne Liu, Dan Garrette, Chitwan Saharia, William Chan et al.ACL 2023 · 25 citations
