Scene Text Segmentation with Text-Focused Transformers
Haiyang Yu, Xiaocong Wang, Ke Niu, Bin Li, Xiangyang Xue
Abstract
Text segmentation is a crucial aspect of various text-related tasks, including text erasing, text editing, and font style transfer. In recent years, multiple text segmentation datasets, such as TextSeg focusing on Latin text segmentation and BTS on bilingual text segmentation, have been proposed. However, existing methods either disregard the annotations of text location or directly use pre-trained text detectors. In general, these methods cannot fully utilize the annotations of text location in the datasets. To explicitly incorporate text location information to guide text segmentation, we propose an end-to-end text-focused segmentation framework, where text detection and segmentation are jointly optimized. In the proposed framework, we first extract multi-level global visual features through residual convolution blocks and then predict the mask of text areas using a text detection head. Subsequently, we develop a text-focused module that compels the model to pay more attention to text areas. Specifically, we introduce two types of attention masks to extract corresponding features: text-aware and instance-aware features. Finally, we employ hierarchical Transformer encoders to fuse multi-level features and predict the text mask with a text segmentation head. To evaluate the effectiveness of our method, we conduct experiments on six text segmentation benchmarks. The experimental results demonstrate that the proposed method outperforms the previous state-of-the-art (SOTA) methods by a clear margin in most cases. The code and supplementary materials are available at https://github.com/FudanVI/FudanOCR/tree/main/text-focused-Transformers https://github.com/FudanVI/FudanOCR/tree/main/text-focused-Transformers.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 715f63c1-a7d8-4dc1-a8db-95f0b5f98038Cited by top-tier papers4
- UPOCR: Towards Unified Pixel-Level OCR InterfaceDezhi Peng, Zhenhua Yang, Jiaxin Zhang, Chongyu Liu et al.ICML 2024 · 14 citations
- ConText: Driving In-context Learning for Text Removal and SegmentationFei Zhang, Pei Zhang, Baosong Yang, Fei Huang et al.ICML 2025
- ST-SAM: Multimodal Scene Text Segmentation with Dense Visual and Sparse Textual Prompts via SAMJin Wei, Yaqiang Wu, Jiayi Yan, Zeng Li et al.AAAI 2026
- Modeling Bottom-up Information Quality during Language ProcessingCui Ding, Yanning Yin, Lena Ann Jäger, Ethan WilcoxEMNLP 2025
Related papers
- BTS: A Bi-lingual Benchmark for Text Segmentation in the WildXixi Xu, Zhongang Qi, Jianqi Ma, Honglun Zhang et al.CVPR 2022 · 14 citations
- SwinTextSpotter: Scene Text Spotting via Better Synergy between Text Detection and Text RecognitionMingxin Huang, Yuliang Liu, Zhenghao Peng, Chongyu Liu et al.CVPR 2022 · 150 citations
- Weakly-Supervised Text Instance SegmentationXinyan Zu, Haiyang Yu, Bin Li, Xiangyang XueACM MM 2023 · 8 citations
- Rethinking Text Segmentation: A Novel Dataset and a Text-Specific Refinement ApproachXingqian Xu, Zhifei Zhang, Zhaowen Wang, Brian L. Price et al.CVPR 2021
- Towards Weakly-Supervised Text Spotting using a Multi-Task TransformerYair Kittenplon, Inbal Lavi, Sharon Fogel, Yarin Bar et al.CVPR 2022 · 60 citations
