Chinese Street View Text: Large-Scale Chinese Text Reading With Partially Supervised Learning
Yipeng Sun, Jiaming Liu, Wei Liu, Junyu Han, Errui Ding, Jingtuo Liu
Abstract
Most existing text reading benchmarks make it difficult to evaluate the performance of more advanced deep learning models in large vocabularies due to the limited amount of training data. To address this issue, we introduce a new large-scale text reading benchmark dataset named Chinese Street View Text (C-SVT) with 430, 000 street view images, which is at least 14 times as large as the existing Chinese text reading benchmarks. To recognize Chinese text in the wild while keeping large-scale datasets labeling cost-effective, we propose to annotate one part of the C-SVT dataset (30,000 images) in locations and text labels as full annotations and add 400, 000 more images, where only the corresponding text-of-interest in the regions is given as weak annotations. To exploit the rich information from the weakly annotated data, we design a text reading network in a partially supervised learning framework, which enables to localize and recognize text, learn from fully and weakly annotated data simultaneously. To localize the best matched text proposals from weakly labeled images, we propose an online proposal matching module incorporated in the whole model, spotting the keyword regions by sharing parameters for end-to-end training. Compared with fully supervised training algorithms, this model can improve the end-to-end recognition performance remarkably by 4.03% in F-score at the same labeling cost. The proposed model can also achieve state-of-the-art results on the ICDAR 2017-RCTW dataset, which demonstrates the effectiveness of the proposed partially supervised learning framework.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a8231646-ef77-458e-b14a-17336b79d26cCited by top-tier papers10
- Detect Anything via Next Point PredictionQing Jiang, Junan Huo, Xingyu Chen, Yuda Xiong et al.CVPR 2026 · 79 citations
- Revisiting Scene Text Recognition: A Data PerspectiveQing Jiang, Jiapeng Wang, Dezhi Peng, Chongyu Liu et al.ICCV 2023 · 70 citations
- SimAN: Exploring Self-Supervised Representation Learning of Scene Text via Similarity-Aware NormalizationCanjie Luo, Lianwen Jin, Jingdong ChenCVPR 2022 · 39 citations
- CLIPTER: Looking at the Bigger Picture in Scene Text RecognitionAviad Aberdam, David Bensaïd, Alona Golts, Roy Ganz et al.ICCV 2023 · 29 citations
- Exploring Better Text Image Translation with Multimodal CodebookZhibin Lan, Jiawei Yu, Xiang Li, Wen Zhang et al.ACL 2023 · 12 citations
Related papers
- An Exemplar-based Framework for Chinese Text RecognitionZhao Zhou, Xiangcheng Du, Yingbin Zheng, Xingjiao Wu et al.AAAI 2025
- Towards Unconstrained End-to-End Text SpottingSiyang Qin, Alessandro Bissacco, Michalis Raptis, Yasuhisa Fujii et al.ICCV 2019 · 138 citations
- Beyond Cropped Regions: New Benchmark and Corresponding Baseline for Chinese Scene Text Retrieval in Diverse LayoutsGengluo Li, Huawen Shen, Yu ZhouICML 2025
- Semantic-Aware Video Text DetectionWei Feng, Fei Yin, Xu-Yao Zhang, Cheng-Lin LiuCVPR 2021
- A Benchmark for Chinese-English Scene Text Image Super-resolutionJianqi Ma, Zhetong Liang, Wangmeng Xiang, Xi Yang et al.ICCV 2023 · 24 citations
