How Control Information Influences Multilingual Text Image Generation and Editing?
Boqiang Zhang, Zuan Gao, Yadong Qu, Hongtao Xie
Abstract
Visual text generation has significantly advanced through diffusion models aimed at producing images with readable and realistic text. Recent works primarily use a ControlNet-based framework, employing standard font text images to control diffusion models. Recognizing the critical role of control information in generating high-quality text, we investigate its influence from three perspectives: input encoding, role at different stages, and output features. Our findings reveal that: 1) Input control information has unique characteristics compared to conventional inputs like Canny edges and depth maps. 2) Control information plays distinct roles at different stages of the denoising process. 3) Output control features significantly differ from the base and skip features of the U-Net decoder in the frequency domain. Based on these insights, we propose TextGen, a novel framework designed to enhance generation quality by optimizing control information. We improve input and output features using Fourier analysis to emphasize relevant information and reduce noise. Additionally, we employ a two-stage generation framework to align the different roles of control information at different stages. Furthermore, we introduce an effective and lightweight dataset for training. Our method achieves state-of-the-art performance in both Chinese and English text generation. The code and dataset available at https://github.com/CyrilSterling/TextGen.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 13077e02-e874-48a9-89c9-ef772cc12901Cited by top-tier papers4
- POCA: Pareto-Optimal Curriculum Alignment for Visual Text GenerationYaohou Fan, Qingzhong Wang, Yongsong Huang, Junyi Liu et al.CVPR 2026 · 2 citations
- Igd: Instructional Graphic Design With Multimodal Layer GeneratioYadong Qu, Hongtao Xie, Yongdong Zhang, Shancheng Fang et al.ICCV 2025 · 1 citation
- PosterMaker: Towards High-Quality Product Poster Generation with Accurate Text RenderingYifan Gao, Zihang Lin, Chuanbin Liu, Min Zhou et al.CVPR 2025
- VerbalTS: Generating Time Series from TextsShuqi Gu, Chuyue Li, Baoyu Jing, Kan RenICML 2025
Builds on20
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
Related papers
- IntrinsicControlNet: Cross-Distribution Image Generation with Real and UnrealJiayuan Lu, Rengan Xie, Zixuan Xie, Zhizhen Wu et al.ICCV 2025 · 4 citations
- Uni-ControlNet: All-in-One Control to Text-to-Image Diffusion ModelsShihao Zhao, Dongdong Chen, Yen-Chun Chen, Jianmin Bao et al.NeurIPS 2023 · 505 citations
- Layout-Agnostic Scene Text Image Synthesis with Diffusion ModelsQilong Zhangli, Jindong Jiang, Di Liu, Licheng Yu et al.CVPR 2024 · 7 citations
- LOOSECONTROL: Lifting ControlNet for Generalized Depth ConditioningShariq Farooq Bhat, Niloy J. Mitra, Peter WonkaSIGGRAPH 2024 · 26 citations
- Simplifying Control Mechanism in Text-to-Image Diffusion ModelsZhida Feng, Li Chen, Yuenan Sun, Jiaxiang Liu et al.AAAI 2025
