SVTRv2: CTC Beats Encoder-Decoder Models in Scene Text Recognition
Yongkun Du, Zhineng Chen, Hongtao Xie, Caiyan Jia, Yu-Gang Jiang
Abstract
Connectionist temporal classification (CTC)-based scene text recognition (STR) methods, e.g., SVTR, are widely employed in OCR applications, mainly due to their simple architecture, which only contains a visual model and a CTC-aligned linear classifier, and therefore fast inference. However, they generally exhibit worse accuracy than encoder-decoder-based methods (EDTRs) due to struggling with text irregularity and linguistic missing. To address these challenges, we propose SVTRv2, a CTC model endowed with the ability to handle text irregularities and model linguistic context. First, a multi-size resizing strategy is proposed to resize text instances to appropriate predefined sizes, effectively avoiding severe text distortion. Meanwhile, we introduce a feature rearrangement module to ensure that visual features accommodate the requirement of CTC, thus alleviating the alignment puzzle. Second, we propose a semantic guidance module. It integrates linguistic context into the visual features, allowing CTC model to leverage language information for accuracy improvement. This module can be omitted at the inference stage and would not increase the time cost. We extensively evaluate SVTRv2 in both standard and recent challenging benchmarks, where SVTRv2 is fairly compared to popular STR models across multiple scenarios, including different types of text irregularity, languages, long text, and whether employing pretraining. SVTRv2 surpasses most EDTRs across the scenarios in terms of accuracy and inference speed. Code: https://github.com/Topdu/OpenOCR.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c71cc00e-2cbd-4e03-9a10-4ec20849637fCited by top-tier papers4
- TextSSR: Diffusion-Based Data Synthesis for Scene Text RecognitionXingsong Ye, Yongkun Du, Yunbo Tao, Zhineng ChenICCV 2025 · 4 citations
- Multi-Scenario Overlapping Text Segmentation with Depth AwarenessYang Liu, Xudong Xie, Yuliang Liu, Xiang BaiICCV 2025 · 4 citations
- MDiff4STR: Mask Diffusion Model for Scene Text RecognitionYongkun Du, Miaomiao Zhao, Songlin Fan, Zhineng Chen et al.AAAI 2026 · 1 citation
- Context-Aware Reasoner: Enhancing Contextual Reasoning in Multimodal Large Language ModelsZhe Zheng, Wenqi Zhang, Xiaohe Zhou, Guiyang Hou et al.ICML 2026
Builds on26
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- BEiT: BERT Pre-Training of Image TransformersHangbo Bao, Li Dong, Songhao Piao, Furu WeiICLR 2022 · 3,632 citations
- What Is Wrong With Scene Text Recognition Model Comparisons? Dataset and Model AnalysisJeonghun Baek, Geewook Kim, Junyeop Lee, Sungrae Park et al.ICCV 2019 · 551 citations
- Focal Modulation NetworksJianwei Yang, Chunyuan Li, Xiyang Dai, Jianfeng GaoNeurIPS 2022 · 494 citations
- Decoupled Attention Network for Text RecognitionTianwei Wang, Yuanzhi Zhu, Lianwen Jin, Canjie Luo et al.AAAI 2020 · 289 citations
Related papers
- GTC: Guided Training of CTC towards Efficient and Accurate Scene Text RecognitionWenyang Hu, Xiaocong Cai, Jun Hou, Shuai Yi et al.AAAI 2020 · 151 citations
- Out of Length Text Recognition with Sub-String MatchingYongkun Du, Zhineng Chen, Caiyan Jia, Xieping Gao et al.AAAI 2025 · 9 citations
- What Machines See Is Not What They Get: Fooling Scene Text Recognition Models With Adversarial Text ImagesXing Xu, Jiefu Chen, Jinhui Xiao, Lianli Gao et al.CVPR 2020
- SCATTER: Selective Context Attentional Scene Text RecognizerRon Litman, Oron Anschel, Shahar Tsiper, Roee Litman et al.CVPR 2020
- Visual Semantics Allow for Textual Reasoning Better in Scene Text RecognitionYue He, Chen Chen, Jing Zhang, Juhua Liu et al.AAAI 2022 · 62 citations
