Brush Your Text: Synthesize Any Scene Text on Images via Diffusion Model
Lingjun Zhang, Xinyuan Chen, Yaohui Wang, Yue Lu, Yu Qiao
Abstract
Recently, diffusion-based image generation methods are credited for their remarkable text-to-image generation capabilities, while still facing challenges in accurately generating multilingual scene text images. To tackle this problem, we propose Diff-Text, which is a training-free scene text generation framework for any language. Our model outputs a photo-realistic image given a text of any language along with a textual description of a scene. The model leverages rendered sketch images as priors, thus arousing the potential multilingual-generation ability of the pre-trained Stable Diffusion. Based on the observation from the influence of the cross-attention map on object placement in generated images, we propose a localized attention constraint into the cross-attention layer to address the unreasonable positioning problem of scene text. Additionally, we introduce contrastive image-level prompts to further refine the position of the textual region and achieve more accurate scene text generation. Experiments demonstrate that our method outperforms the existing method in both the accuracy of text recognition and the naturalness of foreground-background blending.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 74fbed28-a7eb-4579-9440-df13b01d80c3Cited by top-tier papers14
- EasyText: Controllable Diffusion Transformer for Multilingual Text RenderingRunnan Lu, Yuxuan Zhang, Jiaming Liu, Haofan Wang et al.AAAI 2026 · 20 citations
- ViMo: A Generative Visual GUI World Model for App AgentsDezhao Luo, Bohan Tang, Kang Li, Georgios Papoudakis et al.ICLR 2026 · 19 citations
- TextPecker: Rewarding Structural Anomaly Quantification for Enhancing Visual Text RenderingHanshen Zhu, Yuliang Liu, Xuecheng Wu, An-Lan Wang et al.CVPR 2026 · 15 citations
- CineVision: An Interactive Pre-visualization Storyboard System for Director-Cinematographer CollaborationZheng Wei, Hongtao Wu, Lvmin Zhang, Xian Xu et al.UIST 2025 · 8 citations
- Rethinking Layered Graphic Design Generation with a Top-Down ApproachJingye Chen, Zhaowen Wang, Nanxuan Zhao, Li Zhang et al.ICCV 2025 · 4 citations
Builds on22
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 6,759 citations
Related papers
- RealText: Realistic Text Image Generation based on Glyph and Scene Aware InpaintingZihou Liu, Dongming Zhang, Jing Zhang, Jun Li et al.ACM MM 2025
- StyleTextGen: Style-Conditioned Multilingual Scene Text GenerationZeyu Chen, Fangmin Zhao, Yan Shu, Yichao Liu et al.CVPR 2026 · 4 citations
- AnyText: Multilingual Visual Text Generation and EditingYuxiang Tuo, Wangmeng Xiang, Jun-Yan He, Yifeng Geng et al.ICLR 2024 · 148 citations
- Dense Text-to-Image Generation with Attention ModulationYunji Kim, Jiyoung Lee, Jin-Hwa Kim, Jung-Woo Ha et al.ICCV 2023 · 204 citations
- Text to Sketch Generation with Multi-StylesTengjie Li, Shikui Tu, Lei XuNeurIPS 2025 · 1 citation
