On Vocabulary Reliance in Scene Text Recognition
Zhaoyi Wan, Jielei Zhang, Liang Zhang, Jiebo Luo, Cong Yao
Abstract
The pursuit of high performance on public benchmarks has been the driving force for research in scene text recognition, and notable progress has been achieved. However, a close investigation reveals a startling fact that the state-ofthe-art methods perform well on images with words within vocabulary but generalize poorly to images with words outside vocabulary. We call this phenomenon "vocabulary reliance". In this paper, we establish an analytical framework to conduct an in-depth study on the problem of vocabulary reliance in scene text recognition. Key findings include: (1) Vocabulary reliance is ubiquitous, i.e., all existing algorithms more or less exhibit such characteristic; (2) Attention-based decoders prove weak in generalizing to words outside vocabulary and segmentation-based decoders perform well in utilizing visual features; (3) Context modeling is highly coupled with the prediction layers. These findings provide new insights and can benefit future research in scene text recognition. Furthermore, we propose a simple yet effective mutual learning strategy to allow models of two families (attention-based and segmentationbased) to learn collaboratively. This remedy alleviates the problem of vocabulary reliance and improves the overall scene text recognition performance.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers12
- Context-Based Contrastive Learning for Scene Text RecognitionXinyun Zhang, Binwu Zhu, Xufeng Yao, Qi Sun et al.AAAI 2022 · 70 citations
- Revisiting Scene Text Recognition: A Data PerspectiveQing Jiang, Jiapeng Wang, Dezhi Peng, Chongyu Liu et al.ICCV 2023 · 70 citations
- Visual Semantics Allow for Textual Reasoning Better in Scene Text RecognitionYue He, Chen Chen, Jing Zhang, Juhua Liu et al.AAAI 2022 · 62 citations
- Joint Visual Semantic Reasoning: Multi-Stage Decoder for Text RecognitionAyan Kumar Bhunia, Aneeshan Sain, Amandeep Kumar, Shuvozit Ghose et al.ICCV 2021 · 60 citations
- Text is Text, No Matter What: Unifying Text Recognition using Knowledge DistillationAyan Kumar Bhunia, Aneeshan Sain, Pinaki Nath Chowdhury, Yi-Zhe SongICCV 2021 · 33 citations
Builds on4
- Real-Time Scene Text Detection with Differentiable BinarizationMinghui Liao, Zhaoyi Wan, Cong Yao, Kai Chen et al.AAAI 2020 · 818 citations
- TextScanner: Reading Characters in Order for Robust Scene Text RecognitionZhaoyi Wan, Minghang He, Haoran Chen, Xiang Bai et al.AAAI 2020 · 158 citations
- GTC: Guided Training of CTC towards Efficient and Accurate Scene Text RecognitionWenyang Hu, Xiaocong Cai, Jun Hou, Shuai Yi et al.AAAI 2020 · 151 citations
- Symmetry-Constrained Rectification Network for Scene Text RecognitionMingkun Yang, Yushuo Guan, Minghui Liao, Xin He et al.ICCV 2019 · 136 citations
Related papers
- Trust Prophet or Not? Taking a Further Verification Step toward Accurate Scene Text RecognitionAnna Zhu, Ke Xiao, Bo Zhou, Runmin WangACM MM 2024 · 6 citations
- Self-Supervised Implicit Glyph Attention for Text RecognitionTongkun Guan, Chaochen Gu, Jingzheng Tu, Xue Yang et al.CVPR 2023
- SCATTER: Selective Context Attentional Scene Text RecognizerRon Litman, Oron Anschel, Shahar Tsiper, Roee Litman et al.CVPR 2020
- Primitive Representation Learning for Scene Text RecognitionRuijie Yan, Liangrui Peng, Shanyu Xiao, Gang YaoCVPR 2021
- CLIPTER: Looking at the Bigger Picture in Scene Text RecognitionAviad Aberdam, David Bensaïd, Alona Golts, Roy Ganz et al.ICCV 2023 · 29 citations
