Learn to Augment: Joint Data Augmentation and Network Optimization for Text Recognition
Canjie Luo, Yuanzhi Zhu, Lianwen Jin, Yongpan Wang
Abstract
Handwritten text and scene text suffer from various shapes and distorted patterns. Thus training a robust recognition model requires a large amount of data to cover diversity as much as possible. In contrast to data collection and annotation, data augmentation is a low cost way. In this paper, we propose a new method for text image augmentation. Different from traditional augmentation methods such as rotation, scaling and perspective transformation, our proposed augmentation method is designed to learn proper and efficient data augmentation which is more effective and specific for training a robust recognizer. By using a set of custom fiducial points, the proposed augmentation method is flexible and controllable. Furthermore, we bridge the gap between the isolated processes of data augmentation and network optimization by joint learning. An agent network learns from the output of the recognition network and controls the fiducial points to generate more proper training samples for the recognition network. Extensive experiments on various benchmarks, including regular scene text, irregular scene text and handwritten text, show that the proposed augmentation and the joint learning methods significantly boost the performance of the recognition networks. A general toolkit for geometric augmentation is available 1 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers18
- PIMNet: A Parallel, Iterative and Mimicking Network for Scene Text RecognitionZhi Qiao, Yu Zhou, Jin Wei, Wei Wang et al.ACM MM 2021 · 81 citations
- Joint Visual Semantic Reasoning: Multi-Stage Decoder for Text RecognitionAyan Kumar Bhunia, Aneeshan Sain, Amandeep Kumar, Shuvozit Ghose et al.ICCV 2021 · 60 citations
- Perceiving Stroke-Semantic Context: Hierarchical Contrastive Learning for Robust Scene Text RecognitionHao Liu, Bin Wang, Zhimin Bao, Mobai Xue et al.AAAI 2022 · 49 citations
- SimAN: Exploring Self-Supervised Representation Learning of Scene Text via Similarity-Aware NormalizationCanjie Luo, Lianwen Jin, Jingdong ChenCVPR 2022 · 39 citations
- Text is Text, No Matter What: Unifying Text Recognition using Knowledge DistillationAyan Kumar Bhunia, Aneeshan Sain, Pinaki Nath Chowdhury, Yi-Zhe SongICCV 2021 · 33 citations
Builds on1
Related papers
- JokerGAN: Memory-Efficient Model for Handwritten Text Generation with Text Line AwarenessJan Zdenek, Hideki NakayamaACM MM 2021 · 22 citations
- Robust and Informative Text Augmentation (RITA) via Constrained Worst-Case Transformations for Low-Resource Named Entity RecognitionHyunwoo Sohn, Baekkwan ParkKDD 2022 · 3 citations
- FlipDA: Effective and Robust Data Augmentation for Few-Shot LearningJing Zhou, Yanan Zheng, Jie Tang, Li Jian et al.ACL 2022 · 91 citations
- Text AutoAugment: Learning Compositional Augmentation Policy for Text ClassificationShuhuai Ren, Jinchao Zhang, Lei Li, Xu Sun et al.EMNLP 2021 · 22 citations
- ODM: A Text-Image Further Alignment Pre-training Approach for Scene Text Detection and SpottingChen Duan, Pei Fu, Shan Guo, Qianyi Jiang et al.CVPR 2024
