ESTextSpotter: Towards Better Scene Text Spotting with Explicit Synergy in Transformer
Mingxin Huang, Jiaxin Zhang, Dezhi Peng, Hao Lu, Can Huang, Yuliang Liu, Xiang Bai, Lianwen Jin
Abstract
In recent years, end-to-end scene text spotting approaches are evolving to the Transformer-based framework. While previous studies have shown the crucial importance of the intrinsic synergy between text detection and recognition, recent advances in Transformer-based methods usually adopt an implicit synergy strategy with shared query, which can not fully realize the potential of these two interactive tasks. In this paper, we argue that the explicit synergy considering distinct characteristics of text detection and recognition can significantly improve the performance text spotting. To this end, we introduce a new model named Explicit Synergy-based Text Spotting Transformer framework (ESTextSpotter), which achieves explicit synergy by modeling discriminative and interactive features for text detection and recognition within a single decoder. Specifically, we decompose the conventional shared query into task-aware queries for text polygon and content, respectively. Through the decoder with the proposed vision-language communication module, the queries interact with each other in an explicit manner while preserving discriminative patterns of text detection and recognition, thus improving performance significantly. Additionally, we propose a task-aware query initialization scheme to ensure stable training. Experimental results demonstrate that our model significantly outperforms previous state-of-theart methods. Code is available at https://github. com/mxin262/ESTextSpotter .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c2bbfd1b-8f14-42cf-8b24-b777c245a199Cited by top-tier papers9
- GoMatching: A Simple Baseline for Video Text Spotting via Long and Short Term MatchingHaibin He, Maoyuan Ye, Jing Zhang, Juhua Liu et al.NeurIPS 2024 · 16 citations
- UPOCR: Towards Unified Pixel-Level OCR InterfaceDezhi Peng, Zhenhua Yang, Jiaxin Zhang, Chongyu Liu et al.ICML 2024 · 14 citations
- InstructOCR: Instruction Boosting Scene Text SpottingChen Duan, Qianyi Jiang, Pei Fu, Jiamin Chen et al.AAAI 2025 · 7 citations
- DNTextSpotter: Arbitrary-Shaped Scene Text Spotting via Improved Denoising TrainingQian Qiao, Yu Xie, Jun Gao, Tianxiang Wu et al.ACM MM 2024 · 7 citations
- Arbitrary Reading Order Scene Text Spotter with Local Semantics GuidanceJiahao Lyu, Wei Wang, Dongbao Yang, Jinwen Zhong et al.AAAI 2025 · 6 citations
Builds on31
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li et al.ICLR 2021 · 7,353 citations
- DAB-DETR: Dynamic Anchor Boxes are Better Queries for DETRShilong Liu, Feng Li, Hao Zhang, Xiao Yang et al.ICLR 2022 · 1,218 citations
- DN-DETR: Accelerate DETR Training by Introducing Query DeNoisingFeng Li, Hao Zhang, Shilong Liu, Jian Guo et al.CVPR 2022 · 879 citations
- Real-Time Scene Text Detection with Differentiable BinarizationMinghui Liao, Zhaoyi Wan, Cong Yao, Kai Chen et al.AAAI 2020 · 818 citations
- Efficient and Accurate Arbitrary-Shaped Text Detection With Pixel Aggregation NetworkWenhai Wang, Enze Xie, Xiaoge Song, Yuhang Zang et al.ICCV 2019 · 490 citations
Related papers
- SwinTextSpotter: Scene Text Spotting via Better Synergy between Text Detection and Text RecognitionMingxin Huang, Yuliang Liu, Zhenghao Peng, Chongyu Liu et al.CVPR 2022 · 150 citations
- DeepSolo: Let Transformer Decoder with Explicit Points Solo for Text SpottingMaoyuan Ye, Jing Zhang, Shanshan Zhao, Juhua Liu et al.CVPR 2023
- Towards Weakly-Supervised Text Spotting using a Multi-Task TransformerYair Kittenplon, Inbal Lavi, Sharon Fogel, Yarin Bar et al.CVPR 2022 · 60 citations
- Text Spotting TransformersXiang Zhang, Yongwen Su, Subarna Tripathi, Zhuowen TuCVPR 2022 · 125 citations
- SPTS: Single-Point Text SpottingDezhi Peng, Xinyu Wang, Yuliang Liu, Jiaxin Zhang et al.ACM MM 2022 · 65 citations
