OTE: Exploring Accurate Scene Text Recognition Using One Token
Jianjun Xu, Yuxin Wang, Hongtao Xie, Yongdong Zhang
Abstract
In this paper, we propose a novel framework to fully exploit the potential of a single vector for scene text recognition (STR). Different from previous sequence-to-sequence methods that rely on a sequence of visual tokens to rep-resent scene text images, we prove that just one token is enough to characterize the entire text image and achieve ac-curate text recognition. Based on this insight, we introduce a new paradigm for STR, called One Token rEcognizer (OTE). Specifically, we implement an image-to-vector en-coder to extract the fine-grained global semantics, elimi-nating the need for sequential features. Furthermore, an elegant yet potent vector-to-sequence decoder is designed to adaptively diffuse global semantics to corresponding character locations, enabling both autoregressive and non-autoregressive decoding schemes. by executing decoding within a high-level representational space, our vector-to-sequence (V2S) approach avoids the alignment issues between visual tokens and character embeddings prevalent in traditional sequence-to-sequence methods. Remarkably, due to introducing character-wise fine-grained information, such global tokens also boost the performance of scene text retrieval tasks. Extensive experiments on synthetic and real datasets demonstrate the effectiveness of our method by achieving new state-of-the-art results on various public STR benchmarks. Our code is available at h t t <tex xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink"></tex> https://github.com/Xu-Jianjun/OTE.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a0d1e937-00e5-4cd0-94ca-93d12f067c08Cited by top-tier papers7
- SVTRv2: CTC Beats Encoder-Decoder Models in Scene Text RecognitionYongkun Du, Zhineng Chen, Hongtao Xie, Caiyan Jia et al.ICCV 2025 · 22 citations
- Out of Length Text Recognition with Sub-String MatchingYongkun Du, Zhineng Chen, Caiyan Jia, Xieping Gao et al.AAAI 2025 · 9 citations
- Decoder Pre-Training with only Text for Scene Text RecognitionShuai Zhao, Yongkun Du, Zhineng Chen, Yu-Gang JiangACM MM 2024 · 6 citations
- TextSSR: Diffusion-Based Data Synthesis for Scene Text RecognitionXingsong Ye, Yongkun Du, Yunbo Tao, Zhineng ChenICCV 2025 · 4 citations
- Prefill-Time Intervention for Mitigating Hallucination in Large Vision-Language ModelsChengsheng Zhang, Chenghao Sun, Xinyan Jiang, Wei Li et al.CVPR 2026 · 2 citations
Builds on14
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- FCOS: Fully Convolutional One-Stage Object DetectionZhi Tian, Chunhua Shen, Hao Chen, Tong HeICCV 2019 · 6,042 citations
- Decoupled Attention Network for Text RecognitionTianwei Wang, Yuanzhi Zhu, Lianwen Jin, Canjie Luo et al.AAAI 2020 · 289 citations
- From Two to One: A New Scene Text Recognizer with Visual Language Modeling NetworkYuxin Wang, Hongtao Xie, Shancheng Fang, Jing Wang et al.ICCV 2021 · 184 citations
Related papers
- One2Seq: One-Token Wise Decoder for Efficient Scene Text RecognitionZhibin Ma, Pengwen Dai, Wei Zhuo, Xugong QinAAAI 2026
- Trust Prophet or Not? Taking a Further Verification Step toward Accurate Scene Text RecognitionAnna Zhu, Ke Xiao, Bo Zhou, Runmin WangACM MM 2024 · 6 citations
- SPTS: Single-Point Text SpottingDezhi Peng, Xinyu Wang, Yuliang Liu, Jiaxin Zhang et al.ACM MM 2022 · 65 citations
- Visual Semantics Allow for Textual Reasoning Better in Scene Text RecognitionYue He, Chen Chen, Jing Zhang, Juhua Liu et al.AAAI 2022 · 62 citations
- Exploring Font-independent Features for Scene Text RecognitionYizhi Wang, Zhouhui LianACM MM 2020 · 19 citations
