Self-Attention Based Text Knowledge Mining for Text Detection
Qi Wan, Haoqin Ji, Linlin Shen
摘要
Pre-trained models play an important role in deep learning based text detectors. However, most methods ignore the gap between natural images and scene text images and directly apply ImageNet for pre-training. To address such a problem, some of them firstly pre-train the model using a large amount of synthetic data and then fine-tune it on target datasets, which is task-specific and has limited generalization capability. In this paper, we focus on providing general pre-trained models for text detectors. Considering the importance of exploring text contents for text detection, we propose STKM (Self-attention based Text Knowledge Mining), which consists of a CNN Encoder and a Self-attention Decoder, to learn general prior knowledge for text detection from SynthText. Given only image level text labels, Self-attention Decoder directly decodes features extracted from CNN Encoder to texts without requirement of detection, which guides the CNN backbone to explicitly learn discriminative semantic representations ignored by previous approaches. After that, the text knowledge learned by the backbone can be transferred to various text detectors to significantly improve their detection performance (e.g., 5.89% higher F-measure for EAST on ICDAR15 dataset) without bells and whistles. Pre-trained model is available at: https://github.com/CVI-SZU/STKM
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Vision-Language Pre-Training for Boosting Scene Text DetectorsSibo Song, Jianqiang Wan, Zhibo Yang, Jun Tang 等CVPR 2022 · 被引用 38 次
- TPSNet: Reverse Thinking of Thin Plate Splines for Arbitrary Shape Scene Text RepresentationWei Wang, Yu Zhou, Jiahao Lyu, Dayan Wu 等ACM MM 2022 · 被引用 35 次
- Towards Robust Real-Time Scene Text Detection: From Semantic to Instance Representation LearningXugong Qin, Pengyuan Lyu, Chengquan Zhang, Yu Zhou 等ACM MM 2023 · 被引用 21 次
- Turning a CLIP Model into a Scene Text DetectorWenwen Yu, Yuliang Liu, Wei Hua, Deqiang Jiang 等CVPR 2023
- ODM: A Text-Image Further Alignment Pre-training Approach for Scene Text Detection and SpottingChen Duan, Pei Fu, Shan Guo, Qianyi Jiang 等CVPR 2024
它引用的顶会 Paper4
- Efficient and Accurate Arbitrary-Shaped Text Detection With Pixel Aggregation NetworkWenhai Wang, Enze Xie, Xiaoge Song, Yuhang Zang 等ICCV 2019 · 被引用 490 次
- TextDragon: An End-to-End Framework for Arbitrary Shaped Text SpottingWei Feng, Wenhao He, Fei Yin, Xu-Yao Zhang 等ICCV 2019 · 被引用 212 次
- Deep Relational Reasoning Graph Network for Arbitrary Shape Text DetectionShi-Xue Zhang, Xiaobin Zhu, Jie-Bo Hou, Chang Liu 等CVPR 2020
- ABCNet: Real-Time Scene Text Spotting With Adaptive Bezier-Curve NetworkYuliang Liu, Hao Chen, Chunhua Shen, Tong He 等CVPR 2020
相关 Paper
- Masked Text Modeling: A Self-Supervised Pre-training Method for Scene Text DetectionKeran Wang, Hongtao Xie, Yuxin Wang, Dongming Zhang 等ACM MM 2023 · 被引用 13 次
- Unsupervised Object Detection Pretraining with Joint Object Priors Generation and Detector LearningYizhou Wang, Meilin Chen, Shixiang Tang, Feng Zhu 等NeurIPS 2022 · 被引用 2 次
- Self-supervised Scene Text Segmentation with Object-centric Layered Representations Augmented by Text RegionsYibo Wang, Yunhu Ye, Yuanpeng Mao, Yanwei Yu 等ACM MM 2022 · 被引用 2 次
- TextBlock: Towards Scene Text Spotting without Fine-grained DetectionJin Wei, Yuan Zhang, Yu Zhou, Gangyan Zeng 等ACM MM 2022 · 被引用 12 次
- Instance Localization for Self-Supervised Detection PretrainingCeyuan Yang, Zhirong Wu, Bolei Zhou, Stephen LinCVPR 2021
