Appearance Discrepancy-guided Sequence Hybrid Masking for Robust Scene Text Recognition
Shihao Zou, Wei Wei, Leyang Xu, Kaihe Xu, Wenfeng Xie
摘要
Masked Image Modeling (MIM) has been widely recognized as a powerful self-supervised paradigm for learning general-purpose visual representations. However, standard MIM based on random masking tends to underperform in domain-specific tasks like Scene Text Recognition (STR), due to challenges such as information sparsity and appearance discrepancies caused by partial occlusion or distortion. To address this issue, we propose a novel pre-training framework called Appearance Discrepancy-guided Sequence Hybrid Masking (DSHM), specifically designed to learn robust representations for STR. To this end, we introduce an Appearance Discrepancy Metric that quantifies the discrepancy level of each image patch by measuring its deviation from anisotropic local discrepancy and intra-instance global style discrepancy. The resulting discrepancy scores are utilized in two key components: (1) A Sequence Hybrid Masking strategy, which prioritizes masking high-discrepancy patches in coherent block forms, thereby elevating the pretext task from simple pixel-level completion to more complex structural reasoning; (2) Discrepancy-Conditioned Tokens (DC-Tokens), which encode prior knowledge about patch difficulty into the decoder, enabling an adaptive reconstruction process and improving the model robustness under scenarios with partial occlusion or text distortion. We achieve competitive performance on multiple benchmark datasets, including common benchmarks, Union14M benchmarks, and Chinese benchmarks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper24
- SimMIM: a Simple Framework for Masked Image ModelingZhenda Xie, Zheng Zhang, Yue Cao, Yutong Lin 等CVPR 2022 · 被引用 1,129 次
- From Two to One: A New Scene Text Recognizer with Visual Language Modeling NetworkYuxin Wang, Hongtao Xie, Shancheng Fang, Jing Wang 等ICCV 2021 · 被引用 184 次
- TextScanner: Reading Characters in Order for Robust Scene Text RecognitionZhaoyi Wan, Minghang He, Haoran Chen, Xiang Bai 等AAAI 2020 · 被引用 158 次
- GTC: Guided Training of CTC towards Efficient and Accurate Scene Text RecognitionWenyang Hu, Xiaocong Cai, Jun Hou, Shuai Yi 等AAAI 2020 · 被引用 151 次
- Revisiting Scene Text Recognition: A Data PerspectiveQing Jiang, Jiapeng Wang, Dezhi Peng, Chongyu Liu 等ICCV 2023 · 被引用 70 次
相关 Paper
- Masked Text Modeling: A Self-Supervised Pre-training Method for Scene Text DetectionKeran Wang, Hongtao Xie, Yuxin Wang, Dongming Zhang 等ACM MM 2023 · 被引用 13 次
- Hybrid Distillation: Connecting Masked Autoencoders with Contrastive LearnersBowen Shi, Xiaopeng Zhang, Yaoming Wang, Jin Li 等ICLR 2024 · 被引用 10 次
- Linguistics-aware Masked Image Modeling for Self-supervised Scene Text RecognitionYifei Zhang, Chang Liu, Jin Wei, Xiaomeng Yang 等CVPR 2025
- MimCo: Masked Image Modeling Pre-training with Contrastive TeacherQiang Zhou, Chaohui Yu, Hao Luo, Zhibin Wang 等ACM MM 2022 · 被引用 16 次
- MDiff4STR: Mask Diffusion Model for Scene Text RecognitionYongkun Du, Miaomiao Zhao, Songlin Fan, Zhineng Chen 等AAAI 2026 · 被引用 1 次
