Self-Attention Based Text Knowledge Mining for Text Detection
Qi Wan, Haoqin Ji, Linlin Shen
Abstract
Pre-trained models play an important role in deep learning based text detectors. However, most methods ignore the gap between natural images and scene text images and directly apply ImageNet for pre-training. To address such a problem, some of them firstly pre-train the model using a large amount of synthetic data and then fine-tune it on target datasets, which is task-specific and has limited generalization capability. In this paper, we focus on providing general pre-trained models for text detectors. Considering the importance of exploring text contents for text detection, we propose STKM (Self-attention based Text Knowledge Mining), which consists of a CNN Encoder and a Self-attention Decoder, to learn general prior knowledge for text detection from SynthText. Given only image level text labels, Self-attention Decoder directly decodes features extracted from CNN Encoder to texts without requirement of detection, which guides the CNN backbone to explicitly learn discriminative semantic representations ignored by previous approaches. After that, the text knowledge learned by the backbone can be transferred to various text detectors to significantly improve their detection performance (e.g., 5.89% higher F-measure for EAST on ICDAR15 dataset) without bells and whistles. Pre-trained model is available at: https://github.com/CVI-SZU/STKM
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers6
- Vision-Language Pre-Training for Boosting Scene Text DetectorsSibo Song, Jianqiang Wan, Zhibo Yang, Jun Tang et al.CVPR 2022 · 38 citations
- TPSNet: Reverse Thinking of Thin Plate Splines for Arbitrary Shape Scene Text RepresentationWei Wang, Yu Zhou, Jiahao Lyu, Dayan Wu et al.ACM MM 2022 · 35 citations
- Towards Robust Real-Time Scene Text Detection: From Semantic to Instance Representation LearningXugong Qin, Pengyuan Lyu, Chengquan Zhang, Yu Zhou et al.ACM MM 2023 · 21 citations
- Turning a CLIP Model into a Scene Text DetectorWenwen Yu, Yuliang Liu, Wei Hua, Deqiang Jiang et al.CVPR 2023
- ODM: A Text-Image Further Alignment Pre-training Approach for Scene Text Detection and SpottingChen Duan, Pei Fu, Shan Guo, Qianyi Jiang et al.CVPR 2024
Builds on4
- Efficient and Accurate Arbitrary-Shaped Text Detection With Pixel Aggregation NetworkWenhai Wang, Enze Xie, Xiaoge Song, Yuhang Zang et al.ICCV 2019 · 490 citations
- TextDragon: An End-to-End Framework for Arbitrary Shaped Text SpottingWei Feng, Wenhao He, Fei Yin, Xu-Yao Zhang et al.ICCV 2019 · 212 citations
- Deep Relational Reasoning Graph Network for Arbitrary Shape Text DetectionShi-Xue Zhang, Xiaobin Zhu, Jie-Bo Hou, Chang Liu et al.CVPR 2020
- ABCNet: Real-Time Scene Text Spotting With Adaptive Bezier-Curve NetworkYuliang Liu, Hao Chen, Chunhua Shen, Tong He et al.CVPR 2020
Related papers
- Masked Text Modeling: A Self-Supervised Pre-training Method for Scene Text DetectionKeran Wang, Hongtao Xie, Yuxin Wang, Dongming Zhang et al.ACM MM 2023 · 13 citations
- Unsupervised Object Detection Pretraining with Joint Object Priors Generation and Detector LearningYizhou Wang, Meilin Chen, Shixiang Tang, Feng Zhu et al.NeurIPS 2022 · 2 citations
- Self-supervised Scene Text Segmentation with Object-centric Layered Representations Augmented by Text RegionsYibo Wang, Yunhu Ye, Yuanpeng Mao, Yanwei Yu et al.ACM MM 2022 · 2 citations
- TextBlock: Towards Scene Text Spotting without Fine-grained DetectionJin Wei, Yuan Zhang, Yu Zhou, Gangyan Zeng et al.ACM MM 2022 · 12 citations
- Instance Localization for Self-Supervised Detection PretrainingCeyuan Yang, Zhirong Wu, Bolei Zhou, Stephen LinCVPR 2021
