PEAN: A Diffusion-Based Prior-Enhanced Attention Network for Scene Text Image Super-Resolution
Zuoyan Zhao, Hui Xue, Pengfei Fang, Shipeng Zhu
Abstract
Scene text image super-resolution (STISR) aims at simultaneously increasing the resolution and readability of low-resolution scene text images, thus boosting the performance of the downstream recognition task. Two factors in scene text images, visual structure and semantic information, affect the recognition performance significantly. To mitigate the effects from these factors, this paper proposes a Prior-Enhanced Attention Network (PEAN). Specifically, an attention-based modulation module is leveraged to understand scene text images by neatly perceiving the local and global dependence of images, despite the shape of the text. Meanwhile, a diffusion-based module is developed to enhance the text prior, hence offering better guidance for the SR network to generate SR images with higher semantic accuracy. Additionally, a multi-task learning paradigm is employed to optimize the network, enabling the model to generate legible SR images. As a result, PEAN establishes new SOTA results on the TextZoom benchmark. Experiments are also conducted to analyze the importance of the enhanced text prior as a means of improving the performance of the SR network. Code is available at https://github.com/jdfxzzy/PEAN.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2be54045-8425-47ae-8fd8-1fd681c39da8Cited by top-tier papers1
Ask how each one uses itBuilds on34
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
Related papers
- Improving Scene Text Image Super-resolution via Dual Prior Modulation NetworkShipeng Zhu, Zuoyan Zhao, Pengfei Fang, Hui XueAAAI 2023 · 40 citations
- StyleSRN: Scene Text Image Super-Resolution with Text Style EmbeddingShengrong Yuan, Runmin Wang, Ke Hao, Xuqi Ma et al.ICCV 2025 · 2 citations
- A Text Attention Network for Spatial Deformation Robust Scene Text Image Super-resolutionJianqi Ma, Zhetong Liang, Lei ZhangCVPR 2022 · 95 citations
- Gradient-Based Graph Attention for Scene Text Image Super-resolutionXiangyuan Zhu, Kehua Guo, Hui Fang, Rui Ding et al.AAAI 2023 · 18 citations
- Scene Text Image Super-Resolution via Parallelly Contextual Attention NetworkCairong Zhao, Shuyang Feng, Brian Nlong Zhao, Zhijun Ding et al.ACM MM 2021 · 61 citations
