You Can even Annotate Text with Voice: Transcription-only-Supervised Text Spotting
Jingqun Tang, Su Qiao, Benlei Cui, Yuhang Ma, Sheng Zhang, Dimitrios Kanoulas
摘要
End-to-end scene text spotting has recently gained great attention in the research community. The majority of existing methods rely heavily on the location annotations of text instances (e.g., word-level boxes, word-level masks, and char-level boxes). We demonstrate that scene text spotting can be accomplished solely via text transcription, significantly reducing the need for costly location annotations. We propose a query-based paradigm to learn implicit location features via the interaction of text queries and image embeddings. These features are then made explicit during the text recognition stage via an attention activation map. Due to the difficulty of training the weakly-supervised model from scratch, we address the issue of model convergence via a circular curriculum learning strategy. Additionally, we propose a coarse-to-fine cross-attention localization mechanism for more precisely locating text instances. Notably, we provide a solution for text spotting via audio annotation, which further reduces the time required for annotation. Moreover, it establishes a link between audio, text, and image modalities in scene text spotting. Using only transcription annotations as supervision on both real and synthetic data, we achieve competitive results on several popular scene text benchmarks. The proposed method offers a reasonable trade-off between model accuracy and annotation time, allowing simplification of large-scale text spotting applications.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- Harmonizing Visual Text Comprehension and GenerationZhen Zhao, Jingqun Tang, Binghong Wu, Chunhui Lin 等NeurIPS 2024 · 被引用 69 次
- Attentive Eraser: Unleashing Diffusion Model's Object Removal Potential via Self-Attention Redirection GuidanceWenhao Sun, Xue-Mei Dong, Benlei Cui, Jingqun TangAAAI 2025 · 被引用 50 次
- Diffusion Probe: Generated Image Result Prediction Using CNN ProbesBukun Huang, Benlei Cui, Zhizeng Ye, Xuemei Dong 等CVPR 2026 · 被引用 13 次
- OFFSET: Segmentation-based Focus Shift Revision for Composed Image RetrievalZhiwei Chen, Yupeng Hu, Zixu Li, Zhiheng Fu 等ACM MM 2025 · 被引用 10 次
- InstructOCR: Instruction Boosting Scene Text SpottingChen Duan, Qianyi Jiang, Pei Fu, Jiamin Chen 等AAAI 2025 · 被引用 7 次
它引用的顶会 Paper17
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech RepresentationsAlexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, Michael AuliNeurIPS 2020 · 被引用 9,451 次
- TextDragon: An End-to-End Framework for Arbitrary Shaped Text SpottingWei Feng, Wenhao He, Fei Yin, Xu-Yao Zhang 等ICCV 2019 · 被引用 212 次
- Convolutional Character NetworksLinjie Xing, Zhi Tian, Weilin Huang, Matthew R. ScottICCV 2019 · 被引用 176 次
- All You Need Is Boundary: Toward Arbitrary-Shaped Text SpottingHao Wang, Pu Lu, Hui Zhang, Mingkun Yang 等AAAI 2020 · 被引用 145 次
- Towards Unconstrained End-to-End Text SpottingSiyang Qin, Alessandro Bissacco, Michalis Raptis, Yasuhisa Fujii 等ICCV 2019 · 被引用 138 次
相关 Paper
- SPTS: Single-Point Text SpottingDezhi Peng, Xinyu Wang, Yuliang Liu, Jiaxin Zhang 等ACM MM 2022 · 被引用 65 次
- Towards Weakly-Supervised Text Spotting using a Multi-Task TransformerYair Kittenplon, Inbal Lavi, Sharon Fogel, Yarin Bar 等CVPR 2022 · 被引用 60 次
- Scene Text Retrieval via Joint Text Detection and Similarity LearningHao Wang, Xiang Bai, Mingkun Yang, Shenggao Zhu 等CVPR 2021
- SemiETS: Integrating Spatial and Content Consistencies for Semi-Supervised End-to-end Text SpottingDongliang Luo, Hanshen Zhu, Ziyang Zhang, Dingkang Liang 等CVPR 2025
- Self-Supervised Implicit Glyph Attention for Text RecognitionTongkun Guan, Chaochen Gu, Jingzheng Tu, Xue Yang 等CVPR 2023
