Text Image Inpainting via Global Structure-Guided Diffusion Models
Shipeng Zhu, Pengfei Fang, Chenjie Zhu, Zuoyan Zhao, Qiang Xu, Hui Xue
Abstract
Real-world text can be damaged by corrosion issues caused by environmental or human factors, which hinder the preservation of the complete styles of texts, e.g., texture and structure. These corrosion issues, such as graffiti signs and incomplete signatures, bring difficulties in understanding the texts, thereby posing significant challenges to downstream applications, e.g., scene text recognition and signature identification. Notably, current inpainting techniques often fail to adequately address this problem and have difficulties restoring accurate text images along with reasonable and consistent styles. Formulating this as an open problem of text image inpainting, this paper aims to build a benchmark to facilitate its study. In doing so, we establish two specific text inpainting datasets which contain scene text images and handwritten text images, respectively. Each of them includes images revamped by real-life and synthetic datasets, featuring pairs of original images, corrupted images, and other assistant information. On top of the datasets, we further develop a novel neural framework, Global Structure-guided Diffusion Model (GSDM), as a potential solution. Leveraging the global structure of the text as a prior, the proposed GSDM develops an efficient diffusion model to recover clean texts. The efficacy of our approach is demonstrated by thorough empirical study, including a substantial boost in both recognition accuracy and image quality. These findings not only highlight the effectiveness of our method but also underscore its potential to enhance the broader field of text image understanding and processing. Code and datasets are available at: https://github.com/blackprotoss/GSDM .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 434e822a-cd8f-4b8d-ae5e-f6d4ef5a5386Cited by top-tier papers4
- PEAN: A Diffusion-Based Prior-Enhanced Attention Network for Scene Text Image Super-ResolutionZuoyan Zhao, Hui Xue, Pengfei Fang, Shipeng ZhuACM MM 2024 · 16 citations
- Reproducing the Past: A Dataset for Benchmarking Inscription RestorationShipeng Zhu, Hui Xue, Na Nie, Chenjie Zhu et al.ACM MM 2024 · 4 citations
- EpiAgent: An Agent-Centric System for Ancient Inscription RestorationShipeng Zhu, Ang Chen, Na Nie, Pengfei Fang et al.CVPR 2026 · 2 citations
- Structure-Guided Diffusion Models for High-Fidelity Portrait Shadow RemovalWanchang Yu, Qing Zhang, Rongjia Zheng, Wei-Shi ZhengICCV 2025 · 1 citation
Builds on18
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
- Denoising Diffusion Restoration ModelsBahjat Kawar, Michael Elad, Stefano Ermon, Jiaming SongNeurIPS 2022 · 1,439 citations
- RePaint: Inpainting using Denoising Diffusion Probabilistic ModelsAndreas Lugmayr, Martin Danelljan, Andrés Romero, Fisher Yu et al.CVPR 2022 · 1,425 citations
Related papers
- TextCtrl: Diffusion-based Scene Text Editing with Prior Guidance ControlWeichao Zeng, Yan Shu, Zhenhang Li, Dongbao Yang et al.NeurIPS 2024 · 55 citations
- Improving Scene Text Image Super-resolution via Dual Prior Modulation NetworkShipeng Zhu, Zuoyan Zhao, Pengfei Fang, Hui XueAAAI 2023 · 40 citations
- Towards Accurate Scene Text Recognition With Semantic Reasoning NetworksDeli Yu, Xuan Li, Chengquan Zhang, Tao Liu et al.CVPR 2020
- OmniText: A Training-Free Generalist for Controllable Text-Image ManipulationAgus Gunawan, Samuel Teodoro, Yun Chen, Soo Ye Kim et al.ICLR 2026 · 3 citations
- TextSSR: Diffusion-Based Data Synthesis for Scene Text RecognitionXingsong Ye, Yongkun Du, Yunbo Tao, Zhineng ChenICCV 2025 · 4 citations
