Text Gestalt: Stroke-Aware Scene Text Image Super-resolution
Jingye Chen, Haiyang Yu, Jianqi Ma, Bin Li, Xiangyang Xue
Abstract
In the last decade, the blossom of deep learning has witnessed the rapid development of scene text recognition. However, the recognition of low-resolution scene text images remains a challenge. Even though some super-resolution methods have been proposed to tackle this problem, they usually treat text images as general images while ignoring the fact that the visual quality of strokes (the atomic unit of text) plays an essential role for text recognition. According to Gestalt Psychology, humans are capable of composing parts of details into the most similar objects guided by prior knowledge. Likewise, when humans observe a low-resolution text image, they will inherently use partial stroke-level details to recover the appearance of holistic characters. Inspired by Gestalt Psychology, we put forward a Stroke-Aware Scene Text Image Super-Resolution method containing a Stroke-Focused Module (SFM) to concentrate on stroke-level internal structures of characters in text images. Specifically, we attempt to design rules for decomposing English characters and digits at stroke-level, then pre-train a text recognizer to provide stroke-level attention maps as positional clues with the purpose of controlling the consistency between the generated super-resolution image and high-resolution ground truth. The extensive experimental results validate that the proposed method can indeed generate more distinguishable images on TextZoom and manually constructed Chinese character dataset Degraded-IC13. Furthermore, since the proposed SFM is only used to provide stroke-level guidance when training, it will not bring any time overhead during the test phase. Code is available at https://github.com/FudanVI/FudanOCR/tree/main/text-gestalt.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ad37f526-dfe0-4299-a70d-51a4d258d8a2Cited by top-tier papers9
- Improving Scene Text Image Super-resolution via Dual Prior Modulation NetworkShipeng Zhu, Zuoyan Zhao, Pengfei Fang, Hui XueAAAI 2023 · 40 citations
- A Benchmark for Chinese-English Scene Text Image Super-resolutionJianqi Ma, Zhetong Liang, Wangmeng Xiang, Xi Yang et al.ICCV 2023 · 24 citations
- Pixel Adapter: A Graph-Based Post-Processing Approach for Scene Text Image Super-ResolutionWenyu Zhang, Xin Deng, Baojun Jia, Xingtong Yu et al.ACM MM 2023 · 19 citations
- Gradient-Based Graph Attention for Scene Text Image Super-resolutionXiangyuan Zhu, Kehua Guo, Hui Fang, Rui Ding et al.AAAI 2023 · 18 citations
- PEAN: A Diffusion-Based Prior-Enhanced Attention Network for Scene Text Image Super-ResolutionZuoyan Zhao, Hui Xue, Pengfei Fang, Shipeng ZhuACM MM 2024 · 16 citations
Builds on3
- Toward Real-World Single Image Super-Resolution: A New Benchmark and a New ModelJianrui Cai, Hui Zeng, Hongwei Yong, Zisheng Cao et al.ICCV 2019 · 713 citations
- ST-SiameseNet: Spatio-Temporal Siamese Networks for Human Mobility Signature IdentificationHuimin Ren, Menghai Pan, Yanhua Li, Xun Zhou et al.KDD 2020 · 33 citations
- SEED: Semantics Enhanced Encoder-Decoder Framework for Scene Text RecognitionZhi Qiao, Yu Zhou, Dongbao Yang, Yucan Zhou et al.CVPR 2020
Related papers
- Scene Text Telescope: Text-Focused Scene Image Super-ResolutionJingye Chen, Bin Li, Xiangyang XueCVPR 2021
- StyleSRN: Scene Text Image Super-Resolution with Text Style EmbeddingShengrong Yuan, Runmin Wang, Ke Hao, Xuqi Ma et al.ICCV 2025 · 2 citations
- STIRER: A Unified Model for Low-Resolution Scene Text Image Recovery and RecognitionMinyi Zhao, Shijie Xuyang, Jihong Guan, Shuigeng ZhouACM MM 2023 · 9 citations
- A Text Attention Network for Spatial Deformation Robust Scene Text Image Super-resolutionJianqi Ma, Zhetong Liang, Lei ZhangCVPR 2022 · 95 citations
- Scene Text Image Super-Resolution via Parallelly Contextual Attention NetworkCairong Zhao, Shuyang Feng, Brian Nlong Zhao, Zhijun Ding et al.ACM MM 2021 · 61 citations
