Trust Prophet or Not? Taking a Further Verification Step toward Accurate Scene Text Recognition
Anna Zhu, Ke Xiao, Bo Zhou, Runmin Wang
摘要
Inducing linguistic knowledge for scene text recognition (STR) is a new trend that could provide semantics for performance boost. However, most autoregressive STR models optimize one-step ahead prediction (i.e., 1-gram prediction) for character sequence, which only utilizes the previous semantic context. Most non-autoregressive models only apply linguistic knowledge individually on the output sequence to refine the results in parallel, which do not fully utilize the visual clues concurrently. In this paper, we propose a novel language-based STR model, called ProphetSTR. It adopts an n-stream attention mechanism in the decoder to simultaneously predict the next n characters based on the previous predictions at each time step. It behaves like a prophet, encouraging the model to predict more accurate results by utilizing the previous semantic information and the near future clues. If the prediction results for the same character at successive time steps are inconsistent, we should not trust any of them. Otherwise, they are reliable predictions. Therefore, we propose a multi-modality verification module, masking the unreliable semantic features and inputting with visual and trusted semantic ones simultaneously for masked prediction recovery in parallel. It learns to align different modalities implicitly and considers both visual context and linguistic knowledge, which could generate more reliable results. Furthermore, we propose a multi-scale weight-sharing encoder for multi-granularity image representation. Extensive experiments demonstrate that ProphetSTR achieves state-of-the-art performances on many benchmarks. Further ablative studies prove the effectiveness of our proposed components.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- One2Seq: One-Token Wise Decoder for Efficient Scene Text RecognitionZhibin Ma, Pengwen Dai, Wei Zhuo, Xugong QinAAAI 2026
- Self-Supervised Implicit Glyph Attention for Text RecognitionTongkun Guan, Chaochen Gu, Jingzheng Tu, Xue Yang 等CVPR 2023
- OTE: Exploring Accurate Scene Text Recognition Using One TokenJianjun Xu, Yuxin Wang, Hongtao Xie, Yongdong ZhangCVPR 2024 · 被引用 21 次
- SCATTER: Selective Context Attentional Scene Text RecognizerRon Litman, Oron Anschel, Shahar Tsiper, Roee Litman 等CVPR 2020
- VL-Reader: Vision and Language Reconstructor is an Effective Scene Text RecognizerHumen Zhong, Zhibo Yang, Zhaohai Li, Peng Wang 等ACM MM 2024 · 被引用 3 次
