Brush Your Text: Synthesize Any Scene Text on Images via Diffusion Model
Lingjun Zhang, Xinyuan Chen, Yaohui Wang, Yue Lu, Yu Qiao
摘要
Recently, diffusion-based image generation methods are credited for their remarkable text-to-image generation capabilities, while still facing challenges in accurately generating multilingual scene text images. To tackle this problem, we propose Diff-Text, which is a training-free scene text generation framework for any language. Our model outputs a photo-realistic image given a text of any language along with a textual description of a scene. The model leverages rendered sketch images as priors, thus arousing the potential multilingual-generation ability of the pre-trained Stable Diffusion. Based on the observation from the influence of the cross-attention map on object placement in generated images, we propose a localized attention constraint into the cross-attention layer to address the unreasonable positioning problem of scene text. Additionally, we introduce contrastive image-level prompts to further refine the position of the textual region and achieve more accurate scene text generation. Experiments demonstrate that our method outperforms the existing method in both the accuracy of text recognition and the naturalness of foreground-background blending.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper14
- EasyText: Controllable Diffusion Transformer for Multilingual Text RenderingRunnan Lu, Yuxuan Zhang, Jiaming Liu, Haofan Wang 等AAAI 2026 · 被引用 20 次
- ViMo: A Generative Visual GUI World Model for App AgentsDezhao Luo, Bohan Tang, Kang Li, Georgios Papoudakis 等ICLR 2026 · 被引用 19 次
- TextPecker: Rewarding Structural Anomaly Quantification for Enhancing Visual Text RenderingHanshen Zhu, Yuliang Liu, Xuecheng Wu, An-Lan Wang 等CVPR 2026 · 被引用 15 次
- CineVision: An Interactive Pre-visualization Storyboard System for Director-Cinematographer CollaborationZheng Wei, Hongtao Wu, Lvmin Zhang, Xian Xu 等UIST 2025 · 被引用 8 次
- Rethinking Layered Graphic Design Generation with a Top-Down ApproachJingye Chen, Zhaowen Wang, Nanxuan Zhao, Li Zhang 等ICCV 2025 · 被引用 4 次
它引用的顶会 Paper22
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li 等NeurIPS 2022 · 被引用 8,965 次
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 被引用 6,759 次
相关 Paper
- RealText: Realistic Text Image Generation based on Glyph and Scene Aware InpaintingZihou Liu, Dongming Zhang, Jing Zhang, Jun Li 等ACM MM 2025
- StyleTextGen: Style-Conditioned Multilingual Scene Text GenerationZeyu Chen, Fangmin Zhao, Yan Shu, Yichao Liu 等CVPR 2026 · 被引用 4 次
- AnyText: Multilingual Visual Text Generation and EditingYuxiang Tuo, Wangmeng Xiang, Jun-Yan He, Yifeng Geng 等ICLR 2024 · 被引用 148 次
- Dense Text-to-Image Generation with Attention ModulationYunji Kim, Jiyoung Lee, Jin-Hwa Kim, Jung-Woo Ha 等ICCV 2023 · 被引用 204 次
- Text to Sketch Generation with Multi-StylesTengjie Li, Shikui Tu, Lei XuNeurIPS 2025 · 被引用 1 次
