Weakly-Supervised Text Instance Segmentation
Xinyan Zu, Haiyang Yu, Bin Li, Xiangyang Xue
Abstract
Text segmentation is a challenging computer vision task with many downstream applications. Current text segmentation models need to be trained with pixel-level annotations, which requires a lot of labor cost. In this paper, we take the first attempt to perform weakly-supervised text instance segmentation through bridging text recognition and text segmentation. We observe that text recognition models are able to produce the attention localization of each text instance. Based on this observation, we propose a two-stage Text Adaptive Refinement (TAR) module to generate the pseudo labels based on the attention map of a text recognizer. Meanwhile, we develop a text segmentation module to take the rough attention location as input to predict segmentation masks, which are supervised by the aforementioned pseudo labels. In addition, we introduce a mask-augmented contrastive learning by treating the segmentation result as an augmented version of the input text image, thus improving the visual representation and further enhancing the performance of both recognition and segmentation. The experimental results demonstrate that the proposed method outperforms the state-of-the-art (SOTA) weakly-supervised generic segmentation methods by 18.95% and 17.80% in fgIoU on ICDAR13-FST and TextSeg. On MLT-S, COCO-TS and Total-Text, the proposed method achieves about 82% of the fully-supervised methods' performance. When evaluated on instance segmentation, the proposed method exceeds existing SOTA methods by 23.32% and 21.34% on ICDAR13-FST and TextSeg, respectively. Code and Supplementary Materials are available at https://github.com/FudanVI/FudanOCR/tree/main/weakly-text-segmentation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- TextDiffuser: Diffusion Models as Text PaintersJingye Chen, Yupan Huang, Tengchao Lv, Lei Cui et al.NeurIPS 2023 · 290 citations
- Deep Instruction Tuning for Segment Anything ModelXiaorui Huang, Gen Luo, Chaoyang Zhu, Bo Tong et al.ACM MM 2024 · 3 citations
Builds on7
- Scene Text Visual Question AnsweringAli Furkan Biten, Rubèn Tito, Andrés Mafla, Lluís Gómez i Bigorda et al.ICCV 2019 · 482 citations
- ShapeMask: Learning to Segment Novel Objects by Refining Shape PriorsWeicheng Kuo, Anelia Angelova, Jitendra Malik, Tsung-Yi LinICCV 2019 · 127 citations
- BTS: A Bi-lingual Benchmark for Text Segmentation in the WildXixi Xu, Zhongang Qi, Jianqi Ma, Honglun Zhang et al.CVPR 2022 · 14 citations
- Read Like Humans: Autonomous, Bidirectional and Iterative Language Modeling for Scene Text RecognitionShancheng Fang, Hongtao Xie, Yuxin Wang, Zhendong Mao et al.CVPR 2021
- Single-Stage Semantic Segmentation From Image LabelsNikita Araslanov, Stefan RothCVPR 2020
Related papers
- MANGO: A Mask Attention Guided One-Stage Scene Text SpotterLiang Qiao, Ying Chen, Zhanzhan Cheng, Yunlu Xu et al.AAAI 2021 · 91 citations
- Reading and Writing: Discriminative and Generative Modeling for Self-Supervised Text RecognitionMingkun Yang, Minghui Liao, Pu Lu, Jing Wang et al.ACM MM 2022 · 69 citations
- Scene Text Segmentation with Text-Focused TransformersHaiyang Yu, Xiaocong Wang, Ke Niu, Bin Li et al.ACM MM 2023 · 9 citations
- TagCLIP: A Local-to-Global Framework to Enhance Open-Vocabulary Multi-Label Classification of CLIP without TrainingYuqi Lin, Minghao Chen, Kaipeng Zhang, Hengjia Li et al.AAAI 2024 · 39 citations
- Boosting Weakly-Supervised Temporal Action Localization with Text InformationGuozhang Li, De Cheng, Xinpeng Ding, Nannan Wang et al.CVPR 2023
