Context-Based Contrastive Learning for Scene Text Recognition
Xinyun Zhang, Binwu Zhu, Xufeng Yao, Qi Sun, Ruiyu Li, Bei Yu
Abstract
Pursuing accurate and robust recognizers has been a long-lasting goal for scene text recognition (STR) researchers. Recently, attention-based methods have demonstrated their effectiveness and achieved impressive results on public benchmarks. The attention mechanism enables models to recognize scene text with severe visual distortions by leveraging contextual information. However, recent studies revealed that the implicit over-reliance of context leads to catastrophic out-of-vocabulary performance. On the contrary to the superior accuracy of the seen text, models are prone to misrecognize unseen text even with good image quality. We propose a novel framework, Context-based contrastive learning (ConCLR), to alleviate this issue. Our proposed method first generates characters with different contexts via simple image concatenation operations and then optimizes contrastive loss on their embeddings. By pulling together clusters of identical characters within various contexts and pushing apart clusters of different characters in embedding space, ConCLR suppresses the side-effect of overfitting to specific contexts and learns a more robust representation. Experiments show that ConCLR significantly improves out-of-vocabulary generalization and achieves state-of-the-art performance on public benchmarks together with attention-based recognizers.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers9
- Self-supervised Character-to-Character Distillation for Text RecognitionTongkun Guan, Wei Shen, Xue Yang, Qi Feng et al.ICCV 2023 · 36 citations
- Text-DIAE: A Self-Supervised Degradation Invariant Autoencoder for Text Recognition and Document EnhancementMohamed Ali Souibgui, Sanket Biswas, Andrés Mafla, Ali Furkan Biten et al.AAAI 2023 · 31 citations
- Symmetrical Linguistic Feature Distillation with CLIP for Scene Text RecognitionZixiao Wang, Hongtao Xie, Yuxin Wang, Jianjun Xu et al.ACM MM 2023 · 31 citations
- CLIPTER: Looking at the Bigger Picture in Scene Text RecognitionAviad Aberdam, David Bensaïd, Alona Golts, Roy Ganz et al.ICCV 2023 · 29 citations
- OTE: Exploring Accurate Scene Text Recognition Using One TokenJianjun Xu, Yuxin Wang, Hongtao Xie, Yongdong ZhangCVPR 2024 · 21 citations
Builds on14
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec et al.NeurIPS 2020 · 9,171 citations
- Supervised Contrastive LearningPrannay Khosla, Piotr Teterwak, Chen Wang, Aaron Sarna et al.NeurIPS 2020 · 7,049 citations
- What Is Wrong With Scene Text Recognition Model Comparisons? Dataset and Model AnalysisJeonghun Baek, Geewook Kim, Junyeop Lee, Sungrae Park et al.ICCV 2019 · 551 citations
- Scene Text Visual Question AnsweringAli Furkan Biten, Rubèn Tito, Andrés Mafla, Lluís Gómez i Bigorda et al.ICCV 2019 · 482 citations
Related papers
- Relational Contrastive Learning for Scene Text RecognitionJinglei Zhang, Tiancheng Lin, Yi Xu, Kai Chen et al.ACM MM 2023 · 14 citations
- Sequence-to-Sequence Contrastive Learning for Text RecognitionAviad Aberdam, Ron Litman, Shahar Tsiper, Oron Anschel et al.CVPR 2021
- Boosting Semi-Supervised Scene Text Recognition via Viewing and SummarizingYadong Qu, Yuxin Wang, Bangbang Zhou, Zixiao Wang et al.NeurIPS 2024 · 6 citations
- On Vocabulary Reliance in Scene Text RecognitionZhaoyi Wan, Jielei Zhang, Liang Zhang, Jiebo Luo et al.CVPR 2020
- Perceiving Stroke-Semantic Context: Hierarchical Contrastive Learning for Robust Scene Text RecognitionHao Liu, Bin Wang, Zhimin Bao, Mobai Xue et al.AAAI 2022 · 49 citations
