Self-supervised Character-to-Character Distillation for Text Recognition
Tongkun Guan, Wei Shen, Xue Yang, Qi Feng, Zekun Jiang, Xiaokang Yang
Abstract
When handling complicated text images (e.g., irregular structures, low resolution, heavy occlusion, and uneven illumination), existing supervised text recognition methods are data-hungry. Although these methods employ large-scale synthetic text images to reduce the dependence on annotated real images, the domain gap still limits the recognition performance. Therefore, exploring the robust text feature representations on unlabeled real images by self-supervised learning is a good solution. However, existing self-supervised text recognition methods conduct sequence-to-sequence representation learning by roughly splitting the visual features along the horizontal axis, which limits the flexibility of the augmentations, as large geometric-based augmentations may lead to sequence-to-sequence feature inconsistency. Motivated by this, we propose a novel self-supervised Character-to-Character Distillation method, CCD, which enables versatile augmentations to facilitate general text representation learning. Specifically, we delineate the character structures of unlabeled real images by designing a self-supervised character segmentation module. Following this, CCD easily enriches the diversity of local characters while keeping their pairwise alignment under flexible augmentations, using the transformation matrix between two augmented views from images. Experiments demonstrate that CCD achieves state-of-the-art results, with average performance gains of 1.38% in text recognition, 1.7% in text segmentation, 0.24 dB (PSNR) and 0.0321 (SSIM) in text super-resolution. Code is available at https://github.com/TongkunGuan/CCD.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 86b5ccc7-7951-4e00-94ee-c0a616be94eaCited by top-tier papers13
- SVTRv2: CTC Beats Encoder-Decoder Models in Scene Text RecognitionYongkun Du, Zhineng Chen, Hongtao Xie, Caiyan Jia et al.ICCV 2025 · 22 citations
- Self-Distillation Regularized Connectionist Temporal Classification Loss for Text Recognition: A Simple Yet Effective ApproachZiyin Zhang, Ning Lu, Minghui Liao, Yongshuai Huang et al.AAAI 2024 · 20 citations
- Choose What You Need: Disentangled Representation Learning for Scene Text Recognition, Removal and EditingBoqiang Zhang, Hongtao Xie, Zuan Gao, Yuxin WangCVPR 2024 · 9 citations
- CodePercept: Code-Grounded Visual STEM Perception for MLLMsTongkun Guan, Zhibo Yang, Jianqiang Wan, Mingkun Yang et al.CVPR 2026 · 6 citations
- Decoder Pre-Training with only Text for Scene Text RecognitionShuai Zhao, Yongkun Du, Zhineng Chen, Yu-Gang JiangACM MM 2024 · 6 citations
Builds on35
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec et al.NeurIPS 2020 · 9,171 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- Data-Efficient Image Recognition with Contrastive Predictive CodingOlivier J. HénaffICML 2020 · 1,553 citations
Related papers
- Reading and Writing: Discriminative and Generative Modeling for Self-Supervised Text RecognitionMingkun Yang, Minghui Liao, Pu Lu, Jing Wang et al.ACM MM 2022 · 69 citations
- Self-supervised Scene Text Segmentation with Object-centric Layered Representations Augmented by Text RegionsYibo Wang, Yunhu Ye, Yuanpeng Mao, Yanwei Yu et al.ACM MM 2022 · 2 citations
- Pushing the Performance Limit of Scene Text Recognizer without Human AnnotationCaiyuan Zheng, Hui Li, Seon-Min Rhee, Seungju Han et al.CVPR 2022 · 20 citations
- Graph-Based Cross-Domain Knowledge Distillation for Cross-Dataset Text-to-Image Person RetrievalBingjun Luo, Jinpeng Wang, Zewen Wang, Junjie Zhu et al.AAAI 2025 · 8 citations
- One-stage Low-resolution Text Recognition with High-resolution Knowledge TransferHang Guo, Tao Dai, Mingyan Zhu, Guanghao Meng et al.ACM MM 2023 · 5 citations
