Character-Aware Models Improve Visual Text Rendering
Rosanne Liu, Dan Garrette, Chitwan Saharia, William Chan, Adam Roberts, Sharan Narang, Irina Blok, RJ Mical, Mohammad Norouzi, Noah Constant
Abstract
Current image generation models struggle to reliably produce well-formed visual text. In this paper, we investigate a key contributing factor: popular text-to-image models lack character-level input features, making it much harder to predict a word's visual makeup as a series of glyphs. To quantify this effect, we conduct a series of experiments comparing character-aware vs. character-blind text encoders. In the text-only domain, we find that character-aware models provide large gains on a novel spelling task (WikiSpell). Applying our learnings to the visual domain, we train a suite of image generation models, and show that character-aware variants outperform their character-blind counterparts across a range of novel text rendering tasks (our DrawText benchmark). Our models set a much higher state-of-the-art on visual spelling, with 30+ point accuracy gains over competitors on rare words, despite training on far fewer examples.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4a99761d-4bb7-4bb8-835e-4923dacae59cCited by top-tier papers15
- TIFA: Accurate and Interpretable Text-to-Image Faithfulness Evaluation with Question AnsweringYushi Hu, Benlin Liu, Jungo Kasai, Yizhong Wang et al.ICCV 2023 · 400 citations
- Reinforcement Learning for Fine-tuning Text-to-Image Diffusion ModelsYing Fan, Olivia Watkins, Yuqing Du, Hao Liu et al.NeurIPS 2023 · 372 citations
- GlyphControl: Glyph Conditional Control for Visual Text GenerationYukang Yang, Dongnan Gui, Yuhui Yuan, Weicong Liang et al.NeurIPS 2023 · 163 citations
- AnyText: Multilingual Visual Text Generation and EditingYuxiang Tuo, Wangmeng Xiang, Jun-Yan He, Yifeng Geng et al.ICLR 2024 · 148 citations
- Davidsonian Scene Graph: Improving Reliability in Fine-grained Evaluation for Text-to-Image GenerationJaemin Cho, Yushi Hu, Jason M. Baldridge, Roopal Garg et al.ICLR 2024 · 139 citations
Builds on5
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray et al.ICML 2021 · 6,356 citations
- CLIPScore: A Reference-free Evaluation Metric for Image CaptioningJack Hessel, Ari Holtzman, Maxwell Forbes, Ronan Le Bras et al.EMNLP 2021 · 937 citations
- Charformer: Fast Character Transformers via Gradient-based Subword TokenizationYi Tay, Vinh Q. Tran, Sebastian Ruder, Jai Prakash Gupta et al.ICLR 2022 · 198 citations
Related papers
- Layout-Agnostic Scene Text Image Synthesis with Diffusion ModelsQilong Zhangli, Jindong Jiang, Di Liu, Licheng Yu et al.CVPR 2024 · 7 citations
- Empowering Backbone Models for Visual Text Generation with Input Granularity Control and Glyph-Aware TrainingWenbo Li, Guohao Li, Zhibin Lan, Xue Xu et al.EMNLP 2024 · 2 citations
- Exploring Font-independent Features for Scene Text RecognitionYizhi Wang, Zhouhui LianACM MM 2020 · 19 citations
- ViType: High-Fidelity Visual Text Rendering via Glyph-Aware Multimodal DiffusionLishuai Gao, Jun-Yan He, Yingsen Zeng, Yujie Zhong et al.AAAI 2026
- TextPecker: Rewarding Structural Anomaly Quantification for Enhancing Visual Text RenderingHanshen Zhu, Yuliang Liu, Xuecheng Wu, An-Lan Wang et al.CVPR 2026 · 15 citations
