Conditional Text Image Generation with Diffusion Models
Yuanzhi Zhu, Zhaohai Li, Tianwei Wang, Mengchao He, Cong Yao
摘要
Current text recognition systems, including those for handwritten scripts and scene text, have relied heavily on image synthesis and augmentation, since it is difficult to realize real-world complexity and diversity through collecting and annotating enough real text images. In this paper, we explore the problem of text image generation, by taking advantage of the powerful abilities of Diffusion Models in generating photo-realistic and diverse image samples with given conditions, and propose a method called Conditional Text Image Generation with Diffusion Models (CTIG-DM for short). To conform to the characteristics of text images, we devise three conditions: image condition, text condition, and style condition, which can be used to control the attributes, contents, and styles of the samples in the image generation process. Specifically, four text image generation modes, namely: (1) synthesis mode, (2) augmentation mode, (3) recovery mode, and (4) imitation mode, can be derived by combining and configuring these three conditions. Extensive experiments on both handwritten and scene text demonstrate that the proposed CTIG-DM is able to produce image samples that simulate real-world complexity and diversity, and thus can boost the performance of existing text recognizers. Besides, CTIG-DM shows its appealing potential in domain adaptation and generating images containing Out-Of-Vocabulary (OOV) words. * In Fig. 1 , the real images are the first and last of each row. In Fig. 2 , the real images are even numbered.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper21
- FontDiffuser: One-Shot Font Generation via Denoising Diffusion with Multi-Scale Content Aggregation and Style Contrastive LearningZhenhua Yang, Dezhi Peng, Yuxin Kong, Yuyi Zhang 等AAAI 2024 · 被引用 90 次
- Thing2Reality: Enabling Spontaneous Creation of 3D Objects from 2D Content using Generative AI in XR MeetingsErzhen Hu, Mingyi Li, Jungtaek Hong, Xun Qian 等UIST 2025 · 被引用 13 次
- ScanTD: 360° Scanpath Prediction based on Time-Series DiffusionYujia Wang, Fang-Lue Zhang, Neil A. DodgsonACM MM 2024 · 被引用 11 次
- It Doesn't Look Like Anything to Me: Using Diffusion Model to Subvert Visual Phishing DetectorsQingying Hao, Nirav Diwan, Ying Yuan, Giovanni Apruzzese 等USENIX Security 2024 · 被引用 11 次
- Towards Effective Data-Free Knowledge Distillation via Diverse Diffusion AugmentationMuquan Li, Dongyang Zhang, Tao He, Xiurui Xie 等ACM MM 2024 · 被引用 5 次
它引用的顶会 Paper24
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 被引用 13,211 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Improved Denoising Diffusion Probabilistic ModelsAlexander Quinn Nichol, Prafulla DhariwalICML 2021 · 被引用 5,234 次
相关 Paper
- Diffusion-based Blind Text Image Super-ResolutionYuzhe Zhang, Jiawei Zhang, Hao Li, Zhouxia Wang 等CVPR 2024 · 被引用 21 次
- Layout-Agnostic Scene Text Image Synthesis with Diffusion ModelsQilong Zhangli, Jindong Jiang, Di Liu, Licheng Yu 等CVPR 2024 · 被引用 7 次
- ControlStyle: Text-Driven Stylized Image Generation Using Diffusion PriorsJingwen Chen, Yingwei Pan, Ting Yao, Tao MeiACM MM 2023 · 被引用 45 次
- ObjCtrl: Object-based Control Relaxation for Conditional Text-to-Image GenerationXinlong Zhang, Zejian Li, Wei Li, Xiaoyu Zhang 等ACM MM 2025
- Control3D: Towards Controllable Text-to-3D GenerationYang Chen, Yingwei Pan, Yehao Li, Ting Yao 等ACM MM 2023 · 被引用 54 次
