Improving Scene Text Image Super-resolution via Dual Prior Modulation Network
Shipeng Zhu, Zuoyan Zhao, Pengfei Fang, Hui Xue
Abstract
Scene text image super-resolution (STISR) aims to simultaneously increase the resolution and legibility of the text images, and the resulting images will significantly affect the performance of downstream tasks. Although numerous progress has been made, existing approaches raise two crucial issues: (1) They neglect the global structure of the text, which bounds the semantic determinism of the scene text. (2) The priors, e.g., text prior or stroke prior, employed in existing works, are extracted from pre-trained text recognizers. That said, such priors suffer from the domain gap including low resolution and blurriness caused by poor imaging conditions, leading to incorrect guidance. Our work addresses these gaps and proposes a plug-and-play module dubbed Dual Prior Modulation Network (DPMN), which leverages dual image-level priors to bring performance gain over existing approaches. Specifically, two types of prior-guided refinement modules, each using the text mask or graphic recognition result of the low-quality SR image from the preceding layer, are designed to improve the structural clarity and semantic accuracy of the text, respectively. The following attention mechanism hence modulates two quality-enhanced images to attain a superior SR result. Extensive experiments validate that our method improves the image quality and boosts the performance of downstream tasks over five typical approaches on the benchmark. Substantial visualizations and ablation studies demonstrate the advantages of the proposed DPMN. Code is available at: https://github.com/jdfxzzy/DPMN.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 40070ce1-4172-4c12-bdce-15e15de2deb7Cited by top-tier papers8
- Text Image Inpainting via Global Structure-Guided Diffusion ModelsShipeng Zhu, Pengfei Fang, Chenjie Zhu, Zuoyan Zhao et al.AAAI 2024 · 25 citations
- PEAN: A Diffusion-Based Prior-Enhanced Attention Network for Scene Text Image Super-ResolutionZuoyan Zhao, Hui Xue, Pengfei Fang, Shipeng ZhuACM MM 2024 · 16 citations
- One-stage Low-resolution Text Recognition with High-resolution Knowledge TransferHang Guo, Tao Dai, Mingyan Zhu, Guanghao Meng et al.ACM MM 2023 · 5 citations
- GlyphSR: A Simple Glyph-Aware Framework for Scene Text Image Super-ResolutionBaole Wei, Yuxuan Zhou, Liangcai Gao, Zhi TangAAAI 2025 · 4 citations
- Reproducing the Past: A Dataset for Benchmarking Inscription RestorationShipeng Zhu, Hui Xue, Na Nie, Chenjie Zhu et al.ACM MM 2024 · 4 citations
Builds on14
- Uformer: A General U-Shaped Transformer for Image RestorationZhendong Wang, Xiaodong Cun, Jianmin Bao, Wengang Zhou et al.CVPR 2022 · 1,970 citations
- Be Your Own Teacher: Improve the Performance of Convolutional Neural Networks via Self DistillationLinfeng Zhang, Jiebo Song, Anni Gao, Jingwei Chen et al.ICCV 2019 · 1,069 citations
- Conformer: Local Features Coupling Global Representations for Visual RecognitionZhiliang Peng, Wei Huang, Shanzhi Gu, Lingxi Xie et al.ICCV 2021 · 723 citations
- From Two to One: A New Scene Text Recognizer with Visual Language Modeling NetworkYuxin Wang, Hongtao Xie, Shancheng Fang, Jing Wang et al.ICCV 2021 · 184 citations
- A Text Attention Network for Spatial Deformation Robust Scene Text Image Super-resolutionJianqi Ma, Zhetong Liang, Lei ZhangCVPR 2022 · 95 citations
Related papers
- StyleSRN: Scene Text Image Super-Resolution with Text Style EmbeddingShengrong Yuan, Runmin Wang, Ke Hao, Xuqi Ma et al.ICCV 2025 · 2 citations
- STIRER: A Unified Model for Low-Resolution Scene Text Image Recovery and RecognitionMinyi Zhao, Shijie Xuyang, Jihong Guan, Shuigeng ZhouACM MM 2023 · 9 citations
- Text Gestalt: Stroke-Aware Scene Text Image Super-resolutionJingye Chen, Haiyang Yu, Jianqi Ma, Bin Li et al.AAAI 2022 · 63 citations
- Gradient-Based Graph Attention for Scene Text Image Super-resolutionXiangyuan Zhu, Kehua Guo, Hui Fang, Rui Ding et al.AAAI 2023 · 18 citations
- Scene Text Telescope: Text-Focused Scene Image Super-ResolutionJingye Chen, Bin Li, Xiangyang XueCVPR 2021
