One-stage Low-resolution Text Recognition with High-resolution Knowledge Transfer
Hang Guo, Tao Dai, Mingyan Zhu, Guanghao Meng, Bin Chen, Zhi Wang, Shu-Tao Xia
Abstract
Recognizing characters from low-resolution (LR) text images poses a significant challenge due to the information deficiency as well as the noise and blur in low-quality images. Current solutions for low-resolution text recognition (LTR) typically rely on a two-stage pipeline that involves super-resolution as the first stage followed by the second-stage recognition. Although this pipeline is straightforward and intuitive, it has to use an additional super-resolution network, which causes inefficiencies during training and testing. Moreover, the recognition accuracy of the second stage heavily depends on the reconstruction quality of the first stage, causing ineffectiveness.In this work, we attempt to address these challenges from a novel perspective: adapting the recognizer to low-resolution inputs by transferring the knowledge from the high-resolution. Guided by this idea, we propose an efficient and effective knowledge distillation framework to achieve multi-level knowledge transfer.Specifically, the visual focus loss is proposed to extract the character position knowledge with resolution gap reduction and character region focus, the semantic contrastive loss is employed to exploit the contextual semantic knowledge with contrastive learning, and the soft logits loss facilitates both local word-level and global sequence-level learning from the soft teacher label.Extensive experiments show that the proposed one-stage pipeline significantly outperforms super-resolution based two-stage frameworks in terms of effectiveness and efficiency, accompanied by favorable robustness.Code is available at https://github.com/csguoh/KD-LTR.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 91dffbb0-b436-4676-9410-d8aa46c17738Cited by top-tier papers1
Ask how each one uses itBuilds on13
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- From Two to One: A New Scene Text Recognizer with Visual Language Modeling NetworkYuxin Wang, Hongtao Xie, Shancheng Fang, Jing Wang et al.ICCV 2021 · 184 citations
- A Text Attention Network for Spatial Deformation Robust Scene Text Image Super-resolutionJianqi Ma, Zhetong Liang, Lei ZhangCVPR 2022 · 95 citations
- PIMNet: A Parallel, Iterative and Mimicking Network for Scene Text RecognitionZhi Qiao, Yu Zhou, Jin Wei, Wei Wang et al.ACM MM 2021 · 81 citations
- Text Gestalt: Stroke-Aware Scene Text Image Super-resolutionJingye Chen, Haiyang Yu, Jianqi Ma, Bin Li et al.AAAI 2022 · 63 citations
Related papers
- Two-Stage Multi-Scale Resolution-Adaptive Network for Low-Resolution Face RecognitionHaihan Wang, Shangfei Wang, Lin FangACM MM 2022 · 8 citations
- Look One and More: Distilling Hybrid Order Relational Knowledge for Cross-Resolution Image RecognitionShiming Ge, Kangkai Zhang, Haolin Liu, Yingying Hua et al.AAAI 2020 · 30 citations
- Self-supervised Character-to-Character Distillation for Text RecognitionTongkun Guan, Wei Shen, Xue Yang, Qi Feng et al.ICCV 2023 · 36 citations
- Bright to Dark: Stage-wise Bilevel Knowledge Transfer for Seeing Text in the DarkChengpei Xu, Wenhao Zhou, Long Ma, Weimin Wang et al.ACM MM 2025 · 1 citation
- Text is Text, No Matter What: Unifying Text Recognition using Knowledge DistillationAyan Kumar Bhunia, Aneeshan Sain, Pinaki Nath Chowdhury, Yi-Zhe SongICCV 2021 · 33 citations
