TextBlock: Towards Scene Text Spotting without Fine-grained Detection
Jin Wei, Yuan Zhang, Yu Zhou, Gangyan Zeng, Zhi Qiao, Youhui Guo, Haiying Wu, Hongbin Wang, Weiping Wang
Abstract
Scene text spotting systems which integrate text detection and recognition modules have witnessed a lot of success in recent years. Existing works mostly follow the framework of word/character-level fine-grained detection and isolated-instance recognition, which overemphasize the role of detector and ignore the rich context information in recognition. After rethinking the conventional framework, and inspired by the glimpse-focus spotting pipeline of human beings, we ask:1) "can machine spot text without accurate detection just like human beings?", and if yes, 2) "is text block another alternative for scene text spotting other than word or character?". Based on these questions, we propose a new perspective of coarse-grained detection with multi-instance recognition for text spotting. Specifically, a pioneering network termed TextBlock is developed, and a heuristic text block generation method as well as a multi-instance block-level recognition module are proposed. In this way, the burden of detection is relieved, and the contextual semantic information is well explored for recognition. To train the block-level recognizer, a synthetic dataset including about 800K images is formed. As a by-product of attention, fine-grained detection can be recovered with the recognizer. Equipped with a detector without many bells and whistles (e.g., Faster R-CNN), TextBlock achieves competitive or even better performance compared with previous sophisticated text spotters on several public benchmarks. As a primary attempt, we expect this framework will have a potential impact on scene text spotting research in the future.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Cited by top-tier papers4
- Towards Robust Real-Time Scene Text Detection: From Semantic to Instance Representation LearningXugong Qin, Pengyuan Lyu, Chengquan Zhang, Yu Zhou et al.ACM MM 2023 · 21 citations
- Focus, Distinguish, and Prompt: Unleashing CLIP for Efficient and Flexible Scene Text RetrievalGangyan Zeng, Yuan Zhang, Jin Wei, Dongbao Yang et al.ACM MM 2024 · 8 citations
- Linguistics-aware Masked Image Modeling for Self-supervised Scene Text RecognitionYifei Zhang, Chang Liu, Jin Wei, Xiaomeng Yang et al.CVPR 2025
- Explicit Relational Reasoning Network for Scene Text DetectionYuchen Su, Zhineng Chen, Yongkun Du, Zhilong Ji et al.AAAI 2025
Related papers
- Arbitrary Reading Order Scene Text Spotter with Local Semantics GuidanceJiahao Lyu, Wei Wang, Dongbao Yang, Jinwen Zhong et al.AAAI 2025 · 6 citations
- SPTS: Single-Point Text SpottingDezhi Peng, Xinyu Wang, Yuliang Liu, Jiaxin Zhang et al.ACM MM 2022 · 65 citations
- Decoupling Recognition from Detection: Single Shot Self-Reliant Scene Text SpotterJingjing Wu, Pengyuan Lyu, Guangming Lu, Chengquan Zhang et al.ACM MM 2022 · 20 citations
- CRNet: A Center-aware Representation for Detecting Text of Arbitrary ShapesYu Zhou, Hongtao Xie, Shancheng Fang, Yan Li et al.ACM MM 2020 · 31 citations
- ESTextSpotter: Towards Better Scene Text Spotting with Explicit Synergy in TransformerMingxin Huang, Jiaxin Zhang, Dezhi Peng, Hao Lu et al.ICCV 2023 · 44 citations
