VL-Reader: Vision and Language Reconstructor is an Effective Scene Text Recognizer
Humen Zhong, Zhibo Yang, Zhaohai Li, Peng Wang, Jun Tang, Wenqing Cheng, Cong Yao
摘要
Text recognition is an inherent integration of vision and language, encompassing the visual texture in stroke patterns and the semantic context among the character sequences. Towards advanced text recognition, there are three key challenges: (1) an encoder capable of representing the visual and semantic distributions; (2) a decoder that ensures the alignment between vision and semantics; and (3) consistency in the framework during pre-training, if it exists, and fine-tuning. Inspired by masked autoencoding, a successful pre-training strategy in both vision and language, we propose an innovative scene text recognition approach, named VL-Reader. The novelty of the VL-Reader lies in the pervasive interplay between vision and language throughout the entire process. Concretely, we first introduce a Masked Visual-Linguistic Reconstruction (MVLR) objective, which aims at simultaneously modeling visual and linguistic information. Then, we design a Masked Visual-Linguistic Decoder (MVLD) to further leverage masked vision-language context and achieve bi-modal feature interaction. The architecture of VL-Reader maintains consistency from pre-training to fine-tuning. In the pre-training stage, VL-Reader reconstructs both masked visual and text tokens, while in the fine-tuning stage, the network degrades to reconstruct all characters from an image without any masked regions. VL-reader achieves an average accuracy of 97.1% on six typical datasets, surpassing the SOTA by 1.1%. The improvement was even more significant on challenging datasets. The results demonstrate that vision and language reconstructor can serve as an effective scene text recognizer.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- SVTRv2: CTC Beats Encoder-Decoder Models in Scene Text RecognitionYongkun Du, Zhineng Chen, Hongtao Xie, Caiyan Jia 等ICCV 2025 · 被引用 22 次
- MDiff4STR: Mask Diffusion Model for Scene Text RecognitionYongkun Du, Miaomiao Zhao, Songlin Fan, Zhineng Chen 等AAAI 2026 · 被引用 1 次
它引用的顶会 Paper19
- RandAugment: Practical Automated Data Augmentation with a Reduced Search SpaceEkin Dogus Cubuk, Barret Zoph, Jonathon Shlens, Quoc LeNeurIPS 2020 · 被引用 4,453 次
- TrOCR: Transformer-Based Optical Character Recognition with Pre-trained ModelsMinghao Li, Tengchao Lv, Jingye Chen, Lei Cui 等AAAI 2023 · 被引用 607 次
- Decoupled Attention Network for Text RecognitionTianwei Wang, Yuanzhi Zhu, Lianwen Jin, Canjie Luo 等AAAI 2020 · 被引用 289 次
- From Two to One: A New Scene Text Recognizer with Visual Language Modeling NetworkYuxin Wang, Hongtao Xie, Shancheng Fang, Jing Wang 等ICCV 2021 · 被引用 184 次
- TextScanner: Reading Characters in Order for Robust Scene Text RecognitionZhaoyi Wan, Minghang He, Haoran Chen, Xiang Bai 等AAAI 2020 · 被引用 158 次
相关 Paper
- Unicoder-VL: A Universal Encoder for Vision and Language by Cross-Modal Pre-TrainingGen Li, Nan Duan, Yuejian Fang, Ming Gong 等AAAI 2020 · 被引用 966 次
- Vision-Language Pre-Training for Boosting Scene Text DetectorsSibo Song, Jianqiang Wan, Zhibo Yang, Jun Tang 等CVPR 2022 · 被引用 38 次
- Trust Prophet or Not? Taking a Further Verification Step toward Accurate Scene Text RecognitionAnna Zhu, Ke Xiao, Bo Zhou, Runmin WangACM MM 2024 · 被引用 6 次
- ROSITA: Enhancing Vision-and-Language Semantic Alignments via Cross- and Intra-modal Knowledge IntegrationYuhao Cui, Zhou Yu, Chunqi Wang, Zhongzhou Zhao 等ACM MM 2021 · 被引用 46 次
- Linguistics-aware Masked Image Modeling for Self-supervised Scene Text RecognitionYifei Zhang, Chang Liu, Jin Wei, Xiaomeng Yang 等CVPR 2025
