GTC: Guided Training of CTC towards Efficient and Accurate Scene Text Recognition
Wenyang Hu, Xiaocong Cai, Jun Hou, Shuai Yi, Zhiping Lin
Abstract
Connectionist Temporal Classification (CTC) and attention mechanism are two main approaches used in recent scene text recognition works. Compared with attention-based methods, CTC decoder has a much shorter inference time, yet a lower accuracy. To design an efficient and effective model, we propose the guided training of CTC (GTC), where CTC model learns a better alignment and feature representations from a more powerful attentional guidance. With the benefit of guided training, CTC model achieves robust and accurate prediction for both regular and irregular scene text while maintaining a fast inference speed. Moreover, to further leverage the potential of CTC decoder, a graph convolutional network (GCN) is proposed to learn the local correlations of extracted features. Extensive experiments on standard benchmarks demonstrate that our end-to-end model achieves a new state-of-the-art for regular and irregular scene text recognition and needs 6 times shorter inference time than attention-based methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers19
- PGNet: Real-time Arbitrarily-Shaped Text Spotting with Point Gathering NetworkPengfei Wang, Chengquan Zhang, Fei Qi, Shanshan Liu et al.AAAI 2021 · 100 citations
- PIMNet: A Parallel, Iterative and Mimicking Network for Scene Text RecognitionZhi Qiao, Yu Zhou, Jin Wei, Wei Wang et al.ACM MM 2021 · 81 citations
- Visual Semantics Allow for Textual Reasoning Better in Scene Text RecognitionYue He, Chen Chen, Jing Zhang, Juhua Liu et al.AAAI 2022 · 62 citations
- SPIN: Structure-Preserving Inner Offset Network for Scene Text RecognitionChengwei Zhang, Yunlu Xu, Zhanzhan Cheng, Shiliang Pu et al.AAAI 2021 · 27 citations
- LISTER: Neighbor Decoding for Length-Insensitive Scene Text RecognitionChangxu Cheng, Peng Wang, Cheng Da, Qi Zheng et al.ICCV 2023 · 24 citations
Related papers
- SVTRv2: CTC Beats Encoder-Decoder Models in Scene Text RecognitionYongkun Du, Zhineng Chen, Hongtao Xie, Caiyan Jia et al.ICCV 2025 · 22 citations
- Primitive Representation Learning for Scene Text RecognitionRuijie Yan, Liangrui Peng, Shanyu Xiao, Gang YaoCVPR 2021
- What Machines See Is Not What They Get: Fooling Scene Text Recognition Models With Adversarial Text ImagesXing Xu, Jiefu Chen, Jinhui Xiao, Lianli Gao et al.CVPR 2020
- TextScanner: Reading Characters in Order for Robust Scene Text RecognitionZhaoyi Wan, Minghang He, Haoran Chen, Xiang Bai et al.AAAI 2020 · 158 citations
- Towards Accurate Scene Text Recognition With Semantic Reasoning NetworksDeli Yu, Xuan Li, Chengquan Zhang, Tao Liu et al.CVPR 2020
