You Can even Annotate Text with Voice: Transcription-only-Supervised Text Spotting
Jingqun Tang, Su Qiao, Benlei Cui, Yuhang Ma, Sheng Zhang, Dimitrios Kanoulas
Abstract
End-to-end scene text spotting has recently gained great attention in the research community. The majority of existing methods rely heavily on the location annotations of text instances (e.g., word-level boxes, word-level masks, and char-level boxes). We demonstrate that scene text spotting can be accomplished solely via text transcription, significantly reducing the need for costly location annotations. We propose a query-based paradigm to learn implicit location features via the interaction of text queries and image embeddings. These features are then made explicit during the text recognition stage via an attention activation map. Due to the difficulty of training the weakly-supervised model from scratch, we address the issue of model convergence via a circular curriculum learning strategy. Additionally, we propose a coarse-to-fine cross-attention localization mechanism for more precisely locating text instances. Notably, we provide a solution for text spotting via audio annotation, which further reduces the time required for annotation. Moreover, it establishes a link between audio, text, and image modalities in scene text spotting. Using only transcription annotations as supervision on both real and synthetic data, we achieve competitive results on several popular scene text benchmarks. The proposed method offers a reasonable trade-off between model accuracy and annotation time, allowing simplification of large-scale text spotting applications.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 493bd63b-b8ff-4c14-a505-a4d839e5fb03Cited by top-tier papers8
- Harmonizing Visual Text Comprehension and GenerationZhen Zhao, Jingqun Tang, Binghong Wu, Chunhui Lin et al.NeurIPS 2024 · 69 citations
- Attentive Eraser: Unleashing Diffusion Model's Object Removal Potential via Self-Attention Redirection GuidanceWenhao Sun, Xue-Mei Dong, Benlei Cui, Jingqun TangAAAI 2025 · 50 citations
- Diffusion Probe: Generated Image Result Prediction Using CNN ProbesBukun Huang, Benlei Cui, Zhizeng Ye, Xuemei Dong et al.CVPR 2026 · 13 citations
- OFFSET: Segmentation-based Focus Shift Revision for Composed Image RetrievalZhiwei Chen, Yupeng Hu, Zixu Li, Zhiheng Fu et al.ACM MM 2025 · 10 citations
- InstructOCR: Instruction Boosting Scene Text SpottingChen Duan, Qianyi Jiang, Pei Fu, Jiamin Chen et al.AAAI 2025 · 7 citations
Builds on17
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech RepresentationsAlexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, Michael AuliNeurIPS 2020 · 9,451 citations
- TextDragon: An End-to-End Framework for Arbitrary Shaped Text SpottingWei Feng, Wenhao He, Fei Yin, Xu-Yao Zhang et al.ICCV 2019 · 212 citations
- Convolutional Character NetworksLinjie Xing, Zhi Tian, Weilin Huang, Matthew R. ScottICCV 2019 · 176 citations
- All You Need Is Boundary: Toward Arbitrary-Shaped Text SpottingHao Wang, Pu Lu, Hui Zhang, Mingkun Yang et al.AAAI 2020 · 145 citations
- Towards Unconstrained End-to-End Text SpottingSiyang Qin, Alessandro Bissacco, Michalis Raptis, Yasuhisa Fujii et al.ICCV 2019 · 138 citations
Related papers
- SPTS: Single-Point Text SpottingDezhi Peng, Xinyu Wang, Yuliang Liu, Jiaxin Zhang et al.ACM MM 2022 · 65 citations
- Towards Weakly-Supervised Text Spotting using a Multi-Task TransformerYair Kittenplon, Inbal Lavi, Sharon Fogel, Yarin Bar et al.CVPR 2022 · 60 citations
- Scene Text Retrieval via Joint Text Detection and Similarity LearningHao Wang, Xiang Bai, Mingkun Yang, Shenggao Zhu et al.CVPR 2021
- SemiETS: Integrating Spatial and Content Consistencies for Semi-Supervised End-to-end Text SpottingDongliang Luo, Hanshen Zhu, Ziyang Zhang, Dingkang Liang et al.CVPR 2025
- Self-Supervised Implicit Glyph Attention for Text RecognitionTongkun Guan, Chaochen Gu, Jingzheng Tu, Xue Yang et al.CVPR 2023
