Revisiting Scene Text Recognition: A Data Perspective
Qing Jiang, Jiapeng Wang, Dezhi Peng, Chongyu Liu, Lianwen Jin
Abstract
This paper aims to re-assess scene text recognition (STR) from a data-oriented perspective. We begin by revisiting the six commonly used benchmarks in STR and observe a trend of performance saturation, whereby only 2.91% of the benchmark images cannot be accurately recognized by an ensemble of 13 representative models. While these results are impressive and suggest that STR could be considered solved, however, we argue that this is primarily due to the less challenging nature of the common benchmarks, thus concealing the underlying issues that STR faces. To this end, we consolidate a large-scale real STR dataset, namely Union14M, which comprises 4 million labeled images and 10 million unlabeled images, to assess the performance of STR models in more complex real-world scenarios. Our experiments demonstrate that the 13 models can only achieve an average accuracy of 66.53% on the 4 million labeled images, indicating that STR still faces numerous challenges in the real world. By analyzing the error patterns of the 13 models, we identify seven open challenges in STR and develop a challenge-driven benchmark consisting of eight distinct subsets to facilitate further progress in the field. Our exploration demonstrates that STR is far from being solved and leveraging data may be a promising solution. In this regard, we find that utilizing the 10 million unlabeled images through self-supervised pre-training can significantly improve the robustness of STR model in real-world scenarios and leads to state-of-the-art performance. Code and dataset is available at https: //github.com/Mountchicken/Union14M .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 245c92af-7498-4e0e-9973-c0cceda9ebbbCited by top-tier papers22
- Harmonizing Visual Text Comprehension and GenerationZhen Zhao, Jingqun Tang, Binghong Wu, Chunhui Lin et al.NeurIPS 2024 · 69 citations
- SVTRv2: CTC Beats Encoder-Decoder Models in Scene Text RecognitionYongkun Du, Zhineng Chen, Hongtao Xie, Caiyan Jia et al.ICCV 2025 · 22 citations
- OTE: Exploring Accurate Scene Text Recognition Using One TokenJianjun Xu, Yuxin Wang, Hongtao Xie, Yongdong ZhangCVPR 2024 · 21 citations
- Unified Hallucination Detection for Multimodal Large Language ModelsXiang Chen, Chenxi Wang, Yida Xue, Ningyu Zhang et al.ACL 2024 · 20 citations
- Multi-modal In-Context Learning Makes an Ego-evolving Scene Text RecognizerZhen Zhao, Jingqun Tang, Chunhui Lin, Binghong Wu et al.CVPR 2024 · 15 citations
Builds on18
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- SimMIM: a Simple Framework for Masked Image ModelingZhenda Xie, Zheng Zhang, Yue Cao, Yutong Lin et al.CVPR 2022 · 1,129 citations
- Decoupled Attention Network for Text RecognitionTianwei Wang, Yuanzhi Zhu, Lianwen Jin, Canjie Luo et al.AAAI 2020 · 289 citations
- From Two to One: A New Scene Text Recognizer with Visual Language Modeling NetworkYuxin Wang, Hongtao Xie, Shancheng Fang, Jing Wang et al.ICCV 2021 · 184 citations
Related papers
- What's Wrong with Synthetic Data for Scene Text Recognition? A Strong Synthetic Engine with Diverse Simulations and Self-EvolutionXingsong Ye, Yongkun Du, JiaXin Zhang, Chen Li et al.CVPR 2026 · 2 citations
- Pushing the Performance Limit of Scene Text Recognizer without Human AnnotationCaiyuan Zheng, Hui Li, Seon-Min Rhee, Seungju Han et al.CVPR 2022 · 20 citations
- An Empirical Study of Scaling Law for Scene Text RecognitionMiao Rang, Zhenni Bi, Chuanjian Liu, Yunhe Wang et al.CVPR 2024 · 11 citations
- What if We Only Use Real Datasets for Scene Text Recognition? Toward Scene Text Recognition With Fewer LabelsJeonghun Baek, Yusuke Matsui, Kiyoharu AizawaCVPR 2021
- Out of Length Text Recognition with Sub-String MatchingYongkun Du, Zhineng Chen, Caiyan Jia, Xieping Gao et al.AAAI 2025 · 9 citations
