One2Seq: One-Token Wise Decoder for Efficient Scene Text Recognition
Zhibin Ma, Pengwen Dai, Wei Zhuo, Xugong Qin
摘要
Auto-regressive (AR)-based decoders, owing to their flexibility in handling variable-length outputs and their strong capability in modeling character-level dependencies, have emerged as the predominant decoding paradigm in the field of scene text recognition (STR). However, AR-based decoders suffer from attention drift, slow decoding speed, and difficulty capturing global dependencies, restricting their performance in various scenarios. In this paper, we propose a novel paradigm for AR-based decoding, called One-Token to Sequence (One2Seq), to address the above issues. Unlike existing methods, we encode the semantic features into a single context token and design a One-Token Wise Decoder to perform the decoding, which alleviates the attention drift caused by the accumulation of semantic information. Moreover, we proposed Positioal-aware Hash Embedding to embed the decoded characters, ensuring the order information is obtained in the context token. By continuously updating this token, One2Seq fully leverages the decoded semantic information while avoiding the computational overhead associated with the growing query sequence. Furthermore, to leverage global information for decoding, we propose Dynamic Global Infusion to dynamically integrates global visual features into the context token. Equipped with the enriched context token, the model has an enhanced ability to extract discriminative local features under the guidance of global context, thereby enhancing recognition accuracy. Extensive experiments reveal that, with its ingenious design, One2Seq exhibits marked superiority on both accuracy and decoding speed compared to existing STR models.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper12
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- RandAugment: Practical Automated Data Augmentation with a Reduced Search SpaceEkin Dogus Cubuk, Barret Zoph, Jonathon Shlens, Quoc LeNeurIPS 2020 · 被引用 4,453 次
- Decoupled Attention Network for Text RecognitionTianwei Wang, Yuanzhi Zhu, Lianwen Jin, Canjie Luo 等AAAI 2020 · 被引用 289 次
- From Two to One: A New Scene Text Recognizer with Visual Language Modeling NetworkYuxin Wang, Hongtao Xie, Shancheng Fang, Jing Wang 等ICCV 2021 · 被引用 184 次
- Revisiting Scene Text Recognition: A Data PerspectiveQing Jiang, Jiapeng Wang, Dezhi Peng, Chongyu Liu 等ICCV 2023 · 被引用 70 次
相关 Paper
- OTE: Exploring Accurate Scene Text Recognition Using One TokenJianjun Xu, Yuxin Wang, Hongtao Xie, Yongdong ZhangCVPR 2024 · 被引用 21 次
- Trust Prophet or Not? Taking a Further Verification Step toward Accurate Scene Text RecognitionAnna Zhu, Ke Xiao, Bo Zhou, Runmin WangACM MM 2024 · 被引用 6 次
- SCATTER: Selective Context Attentional Scene Text RecognizerRon Litman, Oron Anschel, Shahar Tsiper, Roee Litman 等CVPR 2020
- Self-Supervised Implicit Glyph Attention for Text RecognitionTongkun Guan, Chaochen Gu, Jingzheng Tu, Xue Yang 等CVPR 2023
- SVTRv2: CTC Beats Encoder-Decoder Models in Scene Text RecognitionYongkun Du, Zhineng Chen, Hongtao Xie, Caiyan Jia 等ICCV 2025 · 被引用 22 次
