PIMNet: A Parallel, Iterative and Mimicking Network for Scene Text Recognition
Zhi Qiao, Yu Zhou, Jin Wei, Wei Wang, Yuan Zhang, Ning Jiang, Hongbin Wang, Weiping Wang
Abstract
Nowadays, scene text recognition has attracted more and more attention due to its various applications. Most state-of-the-art methods adopt an encoder-decoder framework with attention mechanism, which generates text autoregressively from left to right. Despite the convincing performance, the speed is limited because of the one-by-one decoding strategy. As opposed to autoregressive models, non-autoregressive models predict the results in parallel with a much shorter inference time, but the accuracy falls behind the autoregressive counterpart considerably. In this paper, we propose a Parallel, Iterative and Mimicking Network (PIMNet) to balance accuracy and efficiency. Specifically, PIMNet adopts a parallel attention mechanism to predict the text faster and an iterative generation mechanism to make the predictions more accurate. In each iteration, the context information is fully explored. To improve learning of the hidden layer, we exploit the mimicking learning in the training phase, where an additional autoregressive decoder is adopted and the parallel decoder mimics the autoregressive decoder with fitting outputs of the hidden layer. With the shared backbone between the two decoders, the proposed PIMNet can be trained end-to-end without pre-training. During inference, the branch of the autoregressive decoder is removed for a faster speed. Extensive experiments on public benchmarks demonstrate the effectiveness and efficiency of PIMNet. Our code is available in the supplementary material.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e9378330-bf56-420d-af83-8ac5caf8f8b7Cited by top-tier papers20
- TextCtrl: Diffusion-based Scene Text Editing with Prior Guidance ControlWeichao Zeng, Yan Shu, Zhenhang Li, Dongbao Yang et al.NeurIPS 2024 · 55 citations
- Self-supervised Character-to-Character Distillation for Text RecognitionTongkun Guan, Wei Shen, Xue Yang, Qi Feng et al.ICCV 2023 · 36 citations
- LISTER: Neighbor Decoding for Length-Insensitive Scene Text RecognitionChangxu Cheng, Peng Wang, Cheng Da, Qi Zheng et al.ICCV 2023 · 24 citations
- Towards Robust Real-Time Scene Text Detection: From Semantic to Instance Representation LearningXugong Qin, Pengyuan Lyu, Chengquan Zhang, Yu Zhou et al.ACM MM 2023 · 21 citations
- Pushing the Performance Limit of Scene Text Recognizer without Human AnnotationCaiyuan Zheng, Hui Li, Seon-Min Rhee, Seungju Han et al.CVPR 2022 · 20 citations
Builds on17
- MobileBERT: a Compact Task-Agnostic BERT for Resource-Limited DevicesZhiqing Sun, Hongkun Yu, Xiaodan Song, Renjie Liu et al.ACL 2020 · 660 citations
- What Is Wrong With Scene Text Recognition Model Comparisons? Dataset and Model AnalysisJeonghun Baek, Geewook Kim, Junyeop Lee, Sungrae Park et al.ICCV 2019 · 551 citations
- Decoupled Attention Network for Text RecognitionTianwei Wang, Yuanzhi Zhu, Lianwen Jin, Canjie Luo et al.AAAI 2020 · 289 citations
- TextDragon: An End-to-End Framework for Arbitrary Shaped Text SpottingWei Feng, Wenhao He, Fei Yin, Xu-Yao Zhang et al.ICCV 2019 · 212 citations
- Video Cloze Procedure for Self-Supervised Spatio-Temporal LearningDezhao Luo, Chang Liu, Yu Zhou, Dongbao Yang et al.AAAI 2020 · 167 citations
Related papers
- GTC: Guided Training of CTC towards Efficient and Accurate Scene Text RecognitionWenyang Hu, Xiaocong Cai, Jun Hou, Shuai Yi et al.AAAI 2020 · 151 citations
- Primitive Representation Learning for Scene Text RecognitionRuijie Yan, Liangrui Peng, Shanyu Xiao, Gang YaoCVPR 2021
- Efficient and Accurate Arbitrary-Shaped Text Detection With Pixel Aggregation NetworkWenhai Wang, Enze Xie, Xiaoge Song, Yuhang Zang et al.ICCV 2019 · 490 citations
- Trust Prophet or Not? Taking a Further Verification Step toward Accurate Scene Text RecognitionAnna Zhu, Ke Xiao, Bo Zhou, Runmin WangACM MM 2024 · 6 citations
- Towards Accurate Scene Text Recognition With Semantic Reasoning NetworksDeli Yu, Xuan Li, Chengquan Zhang, Tao Liu et al.CVPR 2020
