BTS: A Bi-lingual Benchmark for Text Segmentation in the Wild
Xixi Xu, Zhongang Qi, Jianqi Ma, Honglun Zhang, Ying Shan, Xiaohu Qie
Abstract
As a prerequisite of many text-related tasks such as text erasing and text style transfer, text segmentation arouses more and more attention recently. Current researches mainly focus on only English characters and digits, while few work studies Chinese characters due to the lack of pub-lic large-scale and high-quality Chinese datasets, which limits the practical application scenarios of text segmentation. Different from English which has a limited alphabet of letters, Chinese has much more basic characters with com-plex structures, making the problem more difficult to deal with. To better analyze this problem, we propose the Bi-lingual Text Segmentation (BTS) dataset, a benchmark that covers various common Chinese scenes including 14,250 diverse and fine-annotated text images. BTS mainly focuses on Chinese characters, and also contains English words and digits. We also introduce Prior Guided Text Segmen-tation Network (PGTSNet), the first baseline to handle bi-lingual and complex-structured text segmentation. A plug-in text region highlighting module and a text perceptual dis-criminator are proposed in PGTSNet to supervise the model with text prior, and guide for more stable and finer text seg-mentation. A variation loss is also employed for suppressing background noise under complex scene. Extensive ex-periments are conducted not only to demonstrate the neces-sity and superiority of the proposed dataset BTS, but also to show the effectiveness of the proposed PGTSNet compared with a variety of state-of-the-art text segmentation methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers8
- TextDiffuser: Diffusion Models as Text PaintersJingye Chen, Yupan Huang, Tengchao Lv, Lei Cui et al.NeurIPS 2023 · 290 citations
- UPOCR: Towards Unified Pixel-Level OCR InterfaceDezhi Peng, Zhenhua Yang, Jiaxin Zhang, Chongyu Liu et al.ICML 2024 · 14 citations
- Weakly-Supervised Text Instance SegmentationXinyan Zu, Haiyang Yu, Bin Li, Xiangyang XueACM MM 2023 · 8 citations
- Text-Aware Real-World Image Super-Resolution via Diffusion Model with Joint Segmentation DecodersQiming Hu, Linlong Fan, Yiyan Luo, Yuhang Yu et al.NeurIPS 2025 · 7 citations
- Multi-Scenario Overlapping Text Segmentation with Depth AwarenessYang Liu, Xudong Xie, Yuliang Liu, Xiang BaiICCV 2025 · 4 citations
Builds on11
- CCNet: Criss-Cross Attention for Semantic SegmentationZilong Huang, Xinggang Wang, Lichao Huang, Chang Huang et al.ICCV 2019 · 2,972 citations
- Real-Time Scene Text Detection with Differentiable BinarizationMinghui Liao, Zhaoyi Wan, Cong Yao, Kai Chen et al.AAAI 2020 · 818 citations
- SSAP: Single-Shot Instance Segmentation With Affinity PyramidNaiyu Gao, Yanhu Shan, Yupei Wang, Xin Zhao et al.ICCV 2019 · 246 citations
- Convolutional Character NetworksLinjie Xing, Zhi Tian, Weilin Huang, Matthew R. ScottICCV 2019 · 176 citations
- GTC: Guided Training of CTC towards Efficient and Accurate Scene Text RecognitionWenyang Hu, Xiaocong Cai, Jun Hou, Shuai Yi et al.AAAI 2020 · 151 citations
Related papers
- Scene Text Segmentation with Text-Focused TransformersHaiyang Yu, Xiaocong Wang, Ke Niu, Bin Li et al.ACM MM 2023 · 9 citations
- Rethinking Text Segmentation: A Novel Dataset and a Text-Specific Refinement ApproachXingqian Xu, Zhifei Zhang, Zhaowen Wang, Brian L. Price et al.CVPR 2021
- A Benchmark for Chinese-English Scene Text Image Super-resolutionJianqi Ma, Zhetong Liang, Wangmeng Xiang, Xi Yang et al.ICCV 2023 · 24 citations
- Self-Supervised Cross-Language Scene Text EditingFuxiang Yang, Tonghua Su, Xiang Zhou, Donglin Di et al.ACM MM 2023 · 2 citations
- Self-supervised Scene Text Segmentation with Object-centric Layered Representations Augmented by Text RegionsYibo Wang, Yunhu Ye, Yuanpeng Mao, Yanwei Yu et al.ACM MM 2022 · 2 citations
