Few Could Be Better Than All: Feature Sampling and Grouping for Scene Text Detection
Jingqun Tang, Wenqing Zhang, Hongye Liu, Mingkun Yang, Bo Jiang, Guanglong Hu, Xiang Bai
Abstract
Recently, transformer-based methods have achieved promising progresses in object detection, as they can eliminate the post-processes like NMS and enrich the deep representations. However, these methods cannot well cope with scene text due to its extreme variance of scales and aspect ratios. In this paper, we present a simple yet effective transformer-based architecture for scene text detection. Different from previous approaches that learn robust deep representations of scene text in a holistic manner, our method performs scene text detection based on a few representative features, which avoids the disturbance by background and reduces the computational cost. Specifically, we first select a few representative features at all scales that are highly relevant to foreground text. Then, we adopt a transformer for modeling the relationship of the sampled features, which effectively divides them into reasonable groups. As each feature group corresponds to a text instance, its bounding box can be easily obtained without any post-processing operation. Using the basic feature pyramid network for feature extraction, our method consistently achieves state-of-the-art results on several popular datasets for scene text detection.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e6cc8ebc-51ae-48ec-b696-abb2fc65e7d8Cited by top-tier papers24
- DPText-DETR: Towards Better Scene Text Detection with Dynamic Points in TransformerMaoyuan Ye, Jing Zhang, Shanshan Zhao, Juhua Liu et al.AAAI 2023 · 123 citations
- Harmonizing Visual Text Comprehension and GenerationZhen Zhao, Jingqun Tang, Binghong Wu, Chunhui Lin et al.NeurIPS 2024 · 69 citations
- Attentive Eraser: Unleashing Diffusion Model's Object Removal Potential via Self-Attention Redirection GuidanceWenhao Sun, Xue-Mei Dong, Benlei Cui, Jingqun TangAAAI 2025 · 50 citations
- ESTextSpotter: Towards Better Scene Text Spotting with Explicit Synergy in TransformerMingxin Huang, Jiaxin Zhang, Dezhi Peng, Hao Lu et al.ICCV 2023 · 44 citations
- ParGo: Bridging Vision-Language with Partial and Global ViewsAn-Lan Wang, Bin Shan, Wei Shi, Kun-Yu Lin et al.AAAI 2025 · 42 citations
Builds on15
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li et al.ICLR 2021 · 7,353 citations
- CenterNet: Keypoint Triplets for Object DetectionKaiwen Duan, Song Bai, Lingxi Xie, Honggang Qi et al.ICCV 2019 · 3,348 citations
- Conditional DETR for Fast Training ConvergenceDepu Meng, Xiaokang Chen, Zejia Fan, Gang Zeng et al.ICCV 2021 · 974 citations
- Real-Time Scene Text Detection with Differentiable BinarizationMinghui Liao, Zhaoyi Wan, Cong Yao, Kai Chen et al.AAAI 2020 · 818 citations
Related papers
- PBFormer: Capturing Complex Scene Text Shape with Polynomial Band TransformerRuijin Liu, Ning Lu, Dapeng Chen, Cheng Li et al.ACM MM 2023 · 2 citations
- CRNet: A Center-aware Representation for Detecting Text of Arbitrary ShapesYu Zhou, Hongtao Xie, Shancheng Fang, Yan Li et al.ACM MM 2020 · 31 citations
- MOST: A Multi-Oriented Scene Text Detector With Localization RefinementMinghang He, Minghui Liao, Zhibo Yang, Humen Zhong et al.CVPR 2021
- SPTS: Single-Point Text SpottingDezhi Peng, Xinyu Wang, Yuliang Liu, Jiaxin Zhang et al.ACM MM 2022 · 65 citations
- Efficient and Accurate Arbitrary-Shaped Text Detection With Pixel Aggregation NetworkWenhai Wang, Enze Xie, Xiaoge Song, Yuhang Zang et al.ICCV 2019 · 490 citations
