Text Gestalt: Stroke-Aware Scene Text Image Super-resolution
Jingye Chen, Haiyang Yu, Jianqi Ma, Bin Li, Xiangyang Xue
摘要
In the last decade, the blossom of deep learning has witnessed the rapid development of scene text recognition. However, the recognition of low-resolution scene text images remains a challenge. Even though some super-resolution methods have been proposed to tackle this problem, they usually treat text images as general images while ignoring the fact that the visual quality of strokes (the atomic unit of text) plays an essential role for text recognition. According to Gestalt Psychology, humans are capable of composing parts of details into the most similar objects guided by prior knowledge. Likewise, when humans observe a low-resolution text image, they will inherently use partial stroke-level details to recover the appearance of holistic characters. Inspired by Gestalt Psychology, we put forward a Stroke-Aware Scene Text Image Super-Resolution method containing a Stroke-Focused Module (SFM) to concentrate on stroke-level internal structures of characters in text images. Specifically, we attempt to design rules for decomposing English characters and digits at stroke-level, then pre-train a text recognizer to provide stroke-level attention maps as positional clues with the purpose of controlling the consistency between the generated super-resolution image and high-resolution ground truth. The extensive experimental results validate that the proposed method can indeed generate more distinguishable images on TextZoom and manually constructed Chinese character dataset Degraded-IC13. Furthermore, since the proposed SFM is only used to provide stroke-level guidance when training, it will not bring any time overhead during the test phase. Code is available at https://github.com/FudanVI/FudanOCR/tree/main/text-gestalt.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- Improving Scene Text Image Super-resolution via Dual Prior Modulation NetworkShipeng Zhu, Zuoyan Zhao, Pengfei Fang, Hui XueAAAI 2023 · 被引用 40 次
- A Benchmark for Chinese-English Scene Text Image Super-resolutionJianqi Ma, Zhetong Liang, Wangmeng Xiang, Xi Yang 等ICCV 2023 · 被引用 24 次
- Pixel Adapter: A Graph-Based Post-Processing Approach for Scene Text Image Super-ResolutionWenyu Zhang, Xin Deng, Baojun Jia, Xingtong Yu 等ACM MM 2023 · 被引用 19 次
- Gradient-Based Graph Attention for Scene Text Image Super-resolutionXiangyuan Zhu, Kehua Guo, Hui Fang, Rui Ding 等AAAI 2023 · 被引用 18 次
- PEAN: A Diffusion-Based Prior-Enhanced Attention Network for Scene Text Image Super-ResolutionZuoyan Zhao, Hui Xue, Pengfei Fang, Shipeng ZhuACM MM 2024 · 被引用 16 次
它引用的顶会 Paper3
- Toward Real-World Single Image Super-Resolution: A New Benchmark and a New ModelJianrui Cai, Hui Zeng, Hongwei Yong, Zisheng Cao 等ICCV 2019 · 被引用 713 次
- ST-SiameseNet: Spatio-Temporal Siamese Networks for Human Mobility Signature IdentificationHuimin Ren, Menghai Pan, Yanhua Li, Xun Zhou 等KDD 2020 · 被引用 33 次
- SEED: Semantics Enhanced Encoder-Decoder Framework for Scene Text RecognitionZhi Qiao, Yu Zhou, Dongbao Yang, Yucan Zhou 等CVPR 2020
相关 Paper
- Scene Text Telescope: Text-Focused Scene Image Super-ResolutionJingye Chen, Bin Li, Xiangyang XueCVPR 2021
- StyleSRN: Scene Text Image Super-Resolution with Text Style EmbeddingShengrong Yuan, Runmin Wang, Ke Hao, Xuqi Ma 等ICCV 2025 · 被引用 2 次
- STIRER: A Unified Model for Low-Resolution Scene Text Image Recovery and RecognitionMinyi Zhao, Shijie Xuyang, Jihong Guan, Shuigeng ZhouACM MM 2023 · 被引用 9 次
- A Text Attention Network for Spatial Deformation Robust Scene Text Image Super-resolutionJianqi Ma, Zhetong Liang, Lei ZhangCVPR 2022 · 被引用 95 次
- Scene Text Image Super-Resolution via Parallelly Contextual Attention NetworkCairong Zhao, Shuyang Feng, Brian Nlong Zhao, Zhijun Ding 等ACM MM 2021 · 被引用 61 次
