Self-Distillation Regularized Connectionist Temporal Classification Loss for Text Recognition: A Simple Yet Effective Approach
Ziyin Zhang, Ning Lu, Minghui Liao, Yongshuai Huang, Cheng Li, Min Wang, Wei Peng
摘要
Text recognition methods are gaining rapid development. Some advanced techniques, e.g., powerful modules, language models, and un- and semi-supervised learning schemes, consecutively push the performance on public benchmarks forward. However, the problem of how to better optimize a text recognition model from the perspective of loss functions is largely overlooked. CTC-based methods, widely used in practice due to their good balance between performance and inference speed, still grapple with accuracy degradation. This is because CTC loss emphasizes the optimization of the entire sequence target while neglecting to learn individual characters. We propose a self-distillation scheme for CTC-based model to address this issue. It incorporates a framewise regularization term in CTC loss to emphasize individual supervision, and leverages the maximizing-a-posteriori of latent alignment to solve the inconsistency problem that arises in distillation between CTC-based models. We refer to the regularized CTC loss as Distillation Connectionist Temporal Classification (DCTC) loss. DCTC loss is module-free, requiring no extra parameters, longer inference lag, or additional training data or phases. Extensive experiments on public benchmarks demonstrate that DCTC can boost text recognition model accuracy by up to 2.6%, without any of these drawbacks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- SVTRv2: CTC Beats Encoder-Decoder Models in Scene Text RecognitionYongkun Du, Zhineng Chen, Hongtao Xie, Caiyan Jia 等ICCV 2025 · 被引用 22 次
- Out of Length Text Recognition with Sub-String MatchingYongkun Du, Zhineng Chen, Caiyan Jia, Xieping Gao 等AAAI 2025 · 被引用 9 次
- TextSSR: Diffusion-Based Data Synthesis for Scene Text RecognitionXingsong Ye, Yongkun Du, Yunbo Tao, Zhineng ChenICCV 2025 · 被引用 4 次
- MDiff4STR: Mask Diffusion Model for Scene Text RecognitionYongkun Du, Miaomiao Zhao, Songlin Fan, Zhineng Chen 等AAAI 2026 · 被引用 1 次
- Linguistics-aware Masked Image Modeling for Self-supervised Scene Text RecognitionYifei Zhang, Chang Liu, Jin Wei, Xiaomeng Yang 等CVPR 2025
它引用的顶会 Paper11
- TrOCR: Transformer-Based Optical Character Recognition with Pre-trained ModelsMinghao Li, Tengchao Lv, Jingye Chen, Lei Cui 等AAAI 2023 · 被引用 607 次
- What Is Wrong With Scene Text Recognition Model Comparisons? Dataset and Model AnalysisJeonghun Baek, Geewook Kim, Junyeop Lee, Sungrae Park 等ICCV 2019 · 被引用 551 次
- Reading and Writing: Discriminative and Generative Modeling for Self-Supervised Text RecognitionMingkun Yang, Minghui Liao, Pu Lu, Jing Wang 等ACM MM 2022 · 被引用 69 次
- Self-supervised Character-to-Character Distillation for Text RecognitionTongkun Guan, Wei Shen, Xue Yang, Qi Feng 等ICCV 2023 · 被引用 36 次
- Text is Text, No Matter What: Unifying Text Recognition using Knowledge DistillationAyan Kumar Bhunia, Aneeshan Sain, Pinaki Nath Chowdhury, Yi-Zhe SongICCV 2021 · 被引用 33 次
相关 Paper
- CR-CTC: Consistency regularization on CTC for improved speech recognitionZengwei Yao, Wei Kang, Xiaoyu Yang, Fangjun Kuang 等ICLR 2025
- Revisiting the Entropy Semiring for Neural Speech RecognitionOscar Chang, Dongseong Hwang, Olivier SiohanICLR 2023 · 被引用 1 次
- GTC: Guided Training of CTC towards Efficient and Accurate Scene Text RecognitionWenyang Hu, Xiaocong Cai, Jun Hou, Shuai Yi 等AAAI 2020 · 被引用 151 次
- Align With Purpose: Optimize Desired Properties in CTC Models with a General Plug-and-Play FrameworkEliya Segev, Maya Alroy, Ronen Katsir, Noam Wies 等ICLR 2024 · 被引用 2 次
- W-CTC: a Connectionist Temporal Classification Loss with Wild CardsXingyu Cai, Jiahong Yuan, Yuchen Bian, Guangxu Xun 等ICLR 2022 · 被引用 16 次
