One2Seq: One-Token Wise Decoder for Efficient Scene Text Recognition
Zhibin Ma, Pengwen Dai, Wei Zhuo, Xugong Qin
Abstract
Auto-regressive (AR)-based decoders, owing to their flexibility in handling variable-length outputs and their strong capability in modeling character-level dependencies, have emerged as the predominant decoding paradigm in the field of scene text recognition (STR). However, AR-based decoders suffer from attention drift, slow decoding speed, and difficulty capturing global dependencies, restricting their performance in various scenarios. In this paper, we propose a novel paradigm for AR-based decoding, called One-Token to Sequence (One2Seq), to address the above issues. Unlike existing methods, we encode the semantic features into a single context token and design a One-Token Wise Decoder to perform the decoding, which alleviates the attention drift caused by the accumulation of semantic information. Moreover, we proposed Positioal-aware Hash Embedding to embed the decoded characters, ensuring the order information is obtained in the context token. By continuously updating this token, One2Seq fully leverages the decoded semantic information while avoiding the computational overhead associated with the growing query sequence. Furthermore, to leverage global information for decoding, we propose Dynamic Global Infusion to dynamically integrates global visual features into the context token. Equipped with the enriched context token, the model has an enhanced ability to extract discriminative local features under the guidance of global context, thereby enhancing recognition accuracy. Extensive experiments reveal that, with its ingenious design, One2Seq exhibits marked superiority on both accuracy and decoding speed compared to existing STR models.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e60959eb-561d-4448-9db8-bdd1d1a7186dBuilds on12
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- RandAugment: Practical Automated Data Augmentation with a Reduced Search SpaceEkin Dogus Cubuk, Barret Zoph, Jonathon Shlens, Quoc LeNeurIPS 2020 · 4,453 citations
- Decoupled Attention Network for Text RecognitionTianwei Wang, Yuanzhi Zhu, Lianwen Jin, Canjie Luo et al.AAAI 2020 · 289 citations
- From Two to One: A New Scene Text Recognizer with Visual Language Modeling NetworkYuxin Wang, Hongtao Xie, Shancheng Fang, Jing Wang et al.ICCV 2021 · 184 citations
- Revisiting Scene Text Recognition: A Data PerspectiveQing Jiang, Jiapeng Wang, Dezhi Peng, Chongyu Liu et al.ICCV 2023 · 70 citations
Related papers
- OTE: Exploring Accurate Scene Text Recognition Using One TokenJianjun Xu, Yuxin Wang, Hongtao Xie, Yongdong ZhangCVPR 2024 · 21 citations
- Trust Prophet or Not? Taking a Further Verification Step toward Accurate Scene Text RecognitionAnna Zhu, Ke Xiao, Bo Zhou, Runmin WangACM MM 2024 · 6 citations
- SCATTER: Selective Context Attentional Scene Text RecognizerRon Litman, Oron Anschel, Shahar Tsiper, Roee Litman et al.CVPR 2020
- Self-Supervised Implicit Glyph Attention for Text RecognitionTongkun Guan, Chaochen Gu, Jingzheng Tu, Xue Yang et al.CVPR 2023
- SVTRv2: CTC Beats Encoder-Decoder Models in Scene Text RecognitionYongkun Du, Zhineng Chen, Hongtao Xie, Caiyan Jia et al.ICCV 2025 · 22 citations
