Bridging the Gap Between End-to-End and Two-Step Text Spotting
Mingxin Huang, Hongliang Li, Yuliang Liu, Xiang Bai, Lianwen Jin
Abstract
Modularity plays a crucial role in the development and maintenance of complex systems. While end-to-end text spotting efficiently mitigates the issues of error accumulation and sub-optimal performance seen in traditional twostep methodologies, the two-step methods continue to be favored in many competitions and practical settings due to their superior modularity. In this paper, we introduce Bridging Text Spotting, a novel approach that resolves the error accumulation and suboptimal performance issues in two-step methods while retaining modularity. To achieve this, we adopt a well-trained detector and recognizer that are developed and trained independently and then lock their parameters to preserve their already acquired capabilities. Subsequently, we introduce a Bridge that connects the locked detector and recognizer through a zero-initialized neural network. This zero-initialized neural network, initialized with weights set to zeros, ensures seamless integration of the large receptive field features in detection into the locked recognizer. Furthermore, since the fixed detector and recognizer cannot naturally acquire end-to-end optimization features, we adopt the Adapter to facilitate their efficient learning of these features. We demonstrate the effectiveness of the proposed method through extensive experiments: Connecting the latest detector and recognizer through Bridging Text Spotting, we achieved an accuracy of 83.3% on Total-Text, 69.8% on CTW1500, and 89.5% on ICDAR 2015. The code is available at https://github.com/mxin262/ Bridging-Text-Spotting.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- Text-Aware Image Restoration with Diffusion ModelsJaewon Min, Jin Hyeon Kim, Paul Hyunbin Cho, Jaeeun Lee et al.ICLR 2026 · 7 citations
- MSTAR: Box-free Multi-query Scene Text Retrieval with Attention RecyclingLiang Yin, Xudong Xie, Zhang Li, Xiang Bai et al.NeurIPS 2025 · 2 citations
- Beyond Cropped Regions: New Benchmark and Corresponding Baseline for Chinese Scene Text Retrieval in Diverse LayoutsGengluo Li, Huawen Shen, Yu ZhouICML 2025
Builds on26
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li et al.ICLR 2021 · 7,353 citations
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 6,759 citations
- Real-Time Scene Text Detection with Differentiable BinarizationMinghui Liao, Zhaoyi Wan, Cong Yao, Kai Chen et al.AAAI 2020 · 818 citations
- TextDragon: An End-to-End Framework for Arbitrary Shaped Text SpottingWei Feng, Wenhao He, Fei Yin, Xu-Yao Zhang et al.ICCV 2019 · 212 citations
- Convolutional Character NetworksLinjie Xing, Zhi Tian, Weilin Huang, Matthew R. ScottICCV 2019 · 176 citations
Related papers
- Text Grouping Adapter: Adapting Pre-Trained Text Detector for Layout AnalysisTianci Bi, Xiaoyi Zhang, Zhizheng Zhang, Wenxuan Xie et al.CVPR 2024
- Text Perceptron: Towards End-to-End Arbitrary-Shaped Text SpottingLiang Qiao, Sanli Tang, Zhanzhan Cheng, Yunlu Xu et al.AAAI 2020 · 128 citations
- SwinTextSpotter: Scene Text Spotting via Better Synergy between Text Detection and Text RecognitionMingxin Huang, Yuliang Liu, Zhenghao Peng, Chongyu Liu et al.CVPR 2022 · 150 citations
- All You Need Is Boundary: Toward Arbitrary-Shaped Text SpottingHao Wang, Pu Lu, Hui Zhang, Mingkun Yang et al.AAAI 2020 · 145 citations
- ESTextSpotter: Towards Better Scene Text Spotting with Explicit Synergy in TransformerMingxin Huang, Jiaxin Zhang, Dezhi Peng, Hao Lu et al.ICCV 2023 · 44 citations
