Pushing the Performance Limit of Scene Text Recognizer without Human Annotation
Caiyuan Zheng, Hui Li, Seon-Min Rhee, Seungju Han, Jae-Joon Han, Peng Wang
摘要
Scene text recognition (STR) attracts much attention over the years because of its wide application. Most methods train STR model in a fully supervised manner which requires large amounts of labeled data. Although synthetic data contributes a lot to STR, it suffers from the real-to-synthetic domain gap the restricts model performance. In this work, we aim to boost STR models by leveraging both synthetic data and the numerous real unlabeled images, exempting human annotation cost thoroughly. A robust con-sistency regularization based semi-supervised framework is proposed for STR, which can effectively solve the instability issue due to domain inconsistency between synthetic and real images. A character-level consistency regularization is designed to mitigate the misalignment between characters in sequence recognition. Extensive experiments on standard text recognition benchmarks demonstrate the effectiveness of the proposed method. It can steadily improve existing STR models, and boost an STR model to achieve new state-of-the-art results. To our best knowledge, this is the first consistency regularization based framework that applies successfully to STR.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- CLIPTER: Looking at the Bigger Picture in Scene Text RecognitionAviad Aberdam, David Bensaïd, Alona Golts, Roy Ganz 等ICCV 2023 · 被引用 29 次
- Boosting Semi-Supervised Scene Text Recognition via Viewing and SummarizingYadong Qu, Yuxin Wang, Bangbang Zhou, Zixiao Wang 等NeurIPS 2024 · 被引用 6 次
- SemiETS: Integrating Spatial and Content Consistencies for Semi-Supervised End-to-end Text SpottingDongliang Luo, Hanshen Zhu, Ziyang Zhang, Dingkang Liang 等CVPR 2025
它引用的顶会 Paper19
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec 等NeurIPS 2020 · 被引用 9,171 次
- FixMatch: Simplifying Semi-Supervised Learning with Consistency and ConfidenceKihyuk Sohn, David Berthelot, Nicholas Carlini, Zizhao Zhang 等NeurIPS 2020 · 被引用 5,129 次
- RandAugment: Practical Automated Data Augmentation with a Reduced Search SpaceEkin Dogus Cubuk, Barret Zoph, Jonathon Shlens, Quoc LeNeurIPS 2020 · 被引用 4,453 次
- Unsupervised Data Augmentation for Consistency TrainingQizhe Xie, Zihang Dai, Eduard H. Hovy, Thang Luong 等NeurIPS 2020 · 被引用 2,774 次
相关 Paper
- Fine-grained Pseudo Labels for Scene Text RecognitionXiaoyu Li, Xiaoxue Chen, Zuming Huang, Lele Xie 等ACM MM 2023 · 被引用 2 次
- What if We Only Use Real Datasets for Scene Text Recognition? Toward Scene Text Recognition With Fewer LabelsJeonghun Baek, Yusuke Matsui, Kiyoharu AizawaCVPR 2021
- Reading and Writing: Discriminative and Generative Modeling for Self-Supervised Text RecognitionMingkun Yang, Minghui Liao, Pu Lu, Jing Wang 等ACM MM 2022 · 被引用 69 次
- What Is Wrong With Scene Text Recognition Model Comparisons? Dataset and Model AnalysisJeonghun Baek, Geewook Kim, Junyeop Lee, Sungrae Park 等ICCV 2019 · 被引用 551 次
- Context-Based Contrastive Learning for Scene Text RecognitionXinyun Zhang, Binwu Zhu, Xufeng Yao, Qi Sun 等AAAI 2022 · 被引用 70 次
