Linguistics-aware Masked Image Modeling for Self-supervised Scene Text Recognition
Yifei Zhang, Chang Liu, Jin Wei, Xiaomeng Yang, Yu Zhou, Can Ma, Xiangyang Ji
Abstract
Abstract Text images are unique in their dual nature, encompassing both visual and linguistic information. The visual component encompasses structural and appearance-based features, while the linguistic dimension incorporates contextual and semantic elements. In scenarios with degraded visual quality, linguistic patterns serve as crucial supplements for comprehension, highlighting the necessity of integrating both aspects for robust scene text recognition (STR). Contemporary STR approaches often use language models or semantic reasoning modules to capture linguistic features, typically requiring large-scale annotated datasets. Self-supervised learning, which lacks annotations, presents challenges in disentangling linguistic features related to the global context. Typically, sequence contrastive learning emphasizes the alignment of local features, while masked image modeling (MIM) tends to exploit local structures to reconstruct visual patterns, resulting in limited linguistic knowledge. In this paper, we propose a Linguistics-aware Masked Image Modeling (LMIM) approach, which channels the linguistic information into the decoding process of MIM through a separate branch. Specifically, we design a linguistics alignment module to extract vision-independent features as linguistic guidance using inputs with different visual appearances. As features extend beyond mere visual structures, LMIM must consider the global context to achieve reconstruction. Extensive experiments on var-This CVPR paper is the Open Access version, provided by the Computer Vision Foundation. Except for this watermark, it is identical to the accepted version; the final published version of the proceedings is available on IEEE Xplore. ious benchmarks quantitatively demonstrate our state-ofthe-art performance, and attention visualizations qualitatively show the simultaneous capture of both visual and linguistic information. The code is available at https: //github.com/zhangyifei01/LMIM .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- Gather and Trace: Rethinking Video TextVQA from an Instance-oriented PerspectiveYan Zhang, Gangyan Zeng, Daiqing Wu, Huawen Shen et al.ACM MM 2025 · 2 citations
- Masked Representation Modeling for Domain-Adaptive SegmentationWenlve Zhou, Zhiheng Zhou, Tiantao Xian, Yikui Zhai et al.CVPR 2026
- Beyond Cropped Regions: New Benchmark and Corresponding Baseline for Chinese Scene Text Retrieval in Diverse LayoutsGengluo Li, Huawen Shen, Yu ZhouICML 2025
- ST-SAM: Multimodal Scene Text Segmentation with Dense Visual and Sparse Textual Prompts via SAMJin Wei, Yaqiang Wu, Jiayi Yan, Zeng Li et al.AAAI 2026
Builds on49
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- BEiT: BERT Pre-Training of Image TransformersHangbo Bao, Li Dong, Songhao Piao, Furu WeiICLR 2022 · 3,632 citations
- SimMIM: a Simple Framework for Masked Image ModelingZhenda Xie, Zheng Zhang, Yue Cao, Yutong Lin et al.CVPR 2022 · 1,129 citations
- Masked Feature Prediction for Self-Supervised Visual Pre-TrainingChen Wei, Haoqi Fan, Saining Xie, Chao-Yuan Wu et al.CVPR 2022 · 524 citations
Related papers
- Appearance Discrepancy-guided Sequence Hybrid Masking for Robust Scene Text RecognitionShihao Zou, Wei Wei, Leyang Xu, Kaihe Xu et al.AAAI 2026
- From Two to One: A New Scene Text Recognizer with Visual Language Modeling NetworkYuxin Wang, Hongtao Xie, Shancheng Fang, Jing Wang et al.ICCV 2021 · 184 citations
- VL-Reader: Vision and Language Reconstructor is an Effective Scene Text RecognizerHumen Zhong, Zhibo Yang, Zhaohai Li, Peng Wang et al.ACM MM 2024 · 3 citations
- Multi-Modal Representation Learning with Text-Driven Soft MasksJaeyoo Park, Bohyung HanCVPR 2023
- Trust Prophet or Not? Taking a Further Verification Step toward Accurate Scene Text RecognitionAnna Zhu, Ke Xiao, Bo Zhou, Runmin WangACM MM 2024 · 6 citations
