Diffusion-based Blind Text Image Super-Resolution
Yuzhe Zhang, Jiawei Zhang, Hao Li, Zhouxia Wang, Luwei Hou, Dongqing Zou, Liheng Bian
Abstract
Recovering degraded low-resolution text images is chal-lenging, especially for Chinese text images with complex strokes and severe degradation in real-world scenarios. En-suring both text fidelity and style realness is crucial for high-quality text image super-resolution. Recently, diffusion models have achieved great success in natural image synthesis and restoration due to their powerful data distribution modeling abilities and data generation capabili-ties. In this work, we propose an Image Diffusion Model (IDM) to restore text images with realistic styles. For diffusion models, they are not only suitable for modeling realis-tic image distribution but also appropriate for learning text distribution. Since text prior is important to guarantee the correctness of the restored text structure according to existing arts, we also propose a Text Diffusion Model (TDM) for text recognition which can guide IDM to generate text images with correct structures. We further propose a Mixture of Multi-modality module (MoM) to make these two diffusion models cooperate with each other in all the diffusion steps. Extensive experiments on synthetic and real-world datasets demonstrate that our Diffusion-based Blind Text Image Super-Resolution (DiffTSR) can restore text images with more accurate text structures as well as more realistic appearances simultaneously. Code is available at https://github.com/YuzheZhang-1999/DiffTSR.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8fd1401e-ddb1-455c-82c3-4f03d67cd9eaCited by top-tier papers11
- SPRINT: Sparse-Dense Residual Fusion for Efficient Diffusion TransformersDogyun Park, Moayed Haji-Ali, Yanyu Li, Willi Menapace et al.ICLR 2026 · 6 citations
- Enhancing Diffusion Model Stability for Image Restoration via Gradient ManagementHongjie Wu, Mingqin Zhang, Linchao He, Ji-Zhe Zhou et al.ACM MM 2025 · 5 citations
- Fine-Structure Preserved Real-World Image Super-Resolution Via Transfer Vae TrainingQiaosi Yi, Shuai Liu, Rongyuan Wu, Lingchen Sun et al.ICCV 2025 · 4 citations
- Blockwise Flow Matching: Improving Flow Matching Models For Efficient High-Quality GenerationDogyun Park, Taehoon Lee, Minseok Joo, Hyunwoo J. KimNeurIPS 2025 · 4 citations
- Restore Text First, Enhance Image Later: Two-Stage Scene Text Image Super-Resolution with Glyph Structure GuidanceMinxing Luo, Linlong Fan, Qiushi Wang, Ge Wu et al.CVPR 2026 · 2 citations
Builds on24
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 6,759 citations
Related papers
- Text-Aware Real-World Image Super-Resolution via Diffusion Model with Joint Segmentation DecodersQiming Hu, Linlong Fan, Yiyan Luo, Yuhang Yu et al.NeurIPS 2025 · 7 citations
- Conditional Text Image Generation with Diffusion ModelsYuanzhi Zhu, Zhaohai Li, Tianwei Wang, Mengchao He et al.CVPR 2023
- StyleSRN: Scene Text Image Super-Resolution with Text Style EmbeddingShengrong Yuan, Runmin Wang, Ke Hao, Xuqi Ma et al.ICCV 2025 · 2 citations
- STIRER: A Unified Model for Low-Resolution Scene Text Image Recovery and RecognitionMinyi Zhao, Shijie Xuyang, Jihong Guan, Shuigeng ZhouACM MM 2023 · 9 citations
- CDFormer: When Degradation Prediction Embraces Diffusion Model for Blind Image Super-ResolutionQingguo Liu, Chenyi Zhuang, Pan Gao, Jie QinCVPR 2024 · 19 citations
