Boosting Semi-Supervised Scene Text Recognition via Viewing and Summarizing
Yadong Qu, Yuxin Wang, Bangbang Zhou, Zixiao Wang, Hongtao Xie, Yongdong Zhang
Abstract
Existing scene text recognition (STR) methods struggle to recognize challenging texts, especially for artistic and severely distorted characters. The limitation lies in the insufficient exploration of character morphologies, including the monotonousness of widely used synthetic training data and the sensitivity of the model to character morphologies. To address these issues, inspired by the human learning process of viewing and summarizing, we facilitate the contrastive learning-based STR framework in a self-motivated manner by leveraging synthetic and real unlabeled data without any human cost. In the viewing process, to compensate for the simplicity of synthetic data and enrich character morphology diversity, we propose an Online Generation Strategy to generate background-free samples with diverse character styles. By excluding background noise distractions, the model is encouraged to focus on character morphology and generalize the ability to recognize complex samples when trained with only simple synthetic data. To boost the summarizing process, we theoretically demonstrate the derivation error in the previous character contrastive loss, which mistakenly causes the sparsity in the intra-class distribution and exacerbates ambiguity on challenging samples. Therefore, a new Character Unidirectional Alignment Loss is proposed to correct this error and unify the representation of the same characters in all samples by aligning the character features in the student model with the reference features in the teacher model. Extensive experiment results show that our method achieves SOTA performance (94.7% and 70.9% average accuracy on common benchmarks and Union14M-Benchmark). Code will be available at https://github.com/qqqyd/ViSu.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 12bdd761-e27e-4e25-b9db-80a13d8dfbd8Cited by top-tier papers4
- What's Wrong with Synthetic Data for Scene Text Recognition? A Strong Synthetic Engine with Diverse Simulations and Self-EvolutionXingsong Ye, Yongkun Du, JiaXin Zhang, Chen Li et al.CVPR 2026 · 2 citations
- Igd: Instructional Graphic Design With Multimodal Layer GeneratioYadong Qu, Hongtao Xie, Yongdong Zhang, Shancheng Fang et al.ICCV 2025 · 1 citation
- Appearance Discrepancy-guided Sequence Hybrid Masking for Robust Scene Text RecognitionShihao Zou, Wei Wei, Leyang Xu, Kaihe Xu et al.AAAI 2026
- SynTab-LLaVA: Enhancing Multimodal Table Understanding with Decoupled SynthesisBangbang Zhou, Zuan Gao, Zixiao Wang, Boqiang Zhang et al.CVPR 2025
Builds on9
- Big Self-Supervised Models are Strong Semi-Supervised LearnersTing Chen, Simon Kornblith, Kevin Swersky, Mohammad Norouzi et al.NeurIPS 2020 · 2,611 citations
- From Two to One: A New Scene Text Recognizer with Visual Language Modeling NetworkYuxin Wang, Hongtao Xie, Shancheng Fang, Jing Wang et al.ICCV 2021 · 184 citations
- Symmetry-Constrained Rectification Network for Scene Text RecognitionMingkun Yang, Yushuo Guan, Minghui Liao, Xin He et al.ICCV 2019 · 136 citations
- Revisiting Scene Text Recognition: A Data PerspectiveQing Jiang, Jiapeng Wang, Dezhi Peng, Chongyu Liu et al.ICCV 2023 · 70 citations
- Symmetrical Linguistic Feature Distillation with CLIP for Scene Text RecognitionZixiao Wang, Hongtao Xie, Yuxin Wang, Jianjun Xu et al.ACM MM 2023 · 31 citations
Related papers
- Context-Based Contrastive Learning for Scene Text RecognitionXinyun Zhang, Binwu Zhu, Xufeng Yao, Qi Sun et al.AAAI 2022 · 70 citations
- Pushing the Performance Limit of Scene Text Recognizer without Human AnnotationCaiyuan Zheng, Hui Li, Seon-Min Rhee, Seungju Han et al.CVPR 2022 · 20 citations
- Reading and Writing: Discriminative and Generative Modeling for Self-Supervised Text RecognitionMingkun Yang, Minghui Liao, Pu Lu, Jing Wang et al.ACM MM 2022 · 69 citations
- Perceiving Stroke-Semantic Context: Hierarchical Contrastive Learning for Robust Scene Text RecognitionHao Liu, Bin Wang, Zhimin Bao, Mobai Xue et al.AAAI 2022 · 49 citations
- Relational Contrastive Learning for Scene Text RecognitionJinglei Zhang, Tiancheng Lin, Yi Xu, Kai Chen et al.ACM MM 2023 · 14 citations
