Arbitrary Reading Order Scene Text Spotter with Local Semantics Guidance
Jiahao Lyu, Wei Wang, Dongbao Yang, Jinwen Zhong, Yu Zhou
摘要
Scene text spotting has attracted the enthusiasm of relative researchers in recent years. Most existing scene text spotters follow the detection-then-recognition paradigm, where the vanilla detection module hardly determines the reading order and leads to failure recognition. After rethinking the auto-regressive scene text recognition method, we find that a well-trained recognizer can implicitly perceive the local semantics of all characters in a complete word or a sentence without a character-level detection module. Local semantic knowledge not only includes text content but also spatial information in the right reading order. Motivated by the above analysis, we propose the Local Semantics Guided scene text Spotter (LSGSpotter), which auto-regressively decodes the position and content of characters guided by the local semantics. Specifically, two effective modules are proposed in LS-GSpotter. On the one hand, we design a Start Point Localization Module (SPLM) for locating text start points to determine the right reading order. On the other hand, a Multi-scale Adaptive Attention Module (MAAM) is proposed to adaptively aggregate text features in a local area. In conclusion, LSGSpotter achieves the arbitrary reading order spotting task without the limitation of sophisticated detection, while alleviating the cost of computational resources with the grid sampling strategy. Extensive experiment results show LSGSpotter achieves state-of-the-art performance on the InverseText benchmark. Moreover, our spotter demonstrates superior performance on English benchmarks for arbitrary-shaped text, achieving improvements of 0.7% and 2.5% on Total-Text and SCUT-CTW1500, respectively. These results validate our text spotter is effective for scene texts in arbitrary reading order and shape.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- Towards Real-World Document Parsing via Realistic Scene Synthesis and Document-Aware TrainingGengluo Li, Pengyuan Lyu, Chengquan Zhang, Huawen Shen 等CVPR 2026 · 被引用 9 次
- StyleTextGen: Style-Conditioned Multilingual Scene Text GenerationZeyu Chen, Fangmin Zhao, Yan Shu, Yichao Liu 等CVPR 2026 · 被引用 4 次
- Gather and Trace: Rethinking Video TextVQA from an Instance-oriented PerspectiveYan Zhang, Gangyan Zeng, Daiqing Wu, Huawen Shen 等ACM MM 2025 · 被引用 2 次
- Uni-DocDiff: A Unified Document Restoration Model Based on DiffusionFangmin Zhao, Weichao Zeng, Zhenhang Li, Dongbao Yang 等ACM MM 2025 · 被引用 1 次
- DocVLM: Make Your VLM an Efficient ReaderMor Shpigel Nacson, Aviad Aberdam, Roy Ganz, Elad Ben-Avraham 等CVPR 2025
它引用的顶会 Paper23
- Real-Time Scene Text Detection with Differentiable BinarizationMinghui Liao, Zhaoyi Wan, Cong Yao, Kai Chen 等AAAI 2020 · 被引用 818 次
- LayoutLM: Pre-training of Text and Layout for Document Image UnderstandingYiheng Xu, Minghao Li, Lei Cui, Shaohan Huang 等KDD 2020 · 被引用 575 次
- TextDragon: An End-to-End Framework for Arbitrary Shaped Text SpottingWei Feng, Wenhao He, Fei Yin, Xu-Yao Zhang 等ICCV 2019 · 被引用 212 次
- Convolutional Character NetworksLinjie Xing, Zhi Tian, Weilin Huang, Matthew R. ScottICCV 2019 · 被引用 176 次
- SwinTextSpotter: Scene Text Spotting via Better Synergy between Text Detection and Text RecognitionMingxin Huang, Yuliang Liu, Zhenghao Peng, Chongyu Liu 等CVPR 2022 · 被引用 150 次
相关 Paper
- TDI TextSpotter: Taking Data Imbalance into Account in Scene Text SpottingYu Zhou, Hongtao Xie, Shancheng Fang, Jing Wang 等ACM MM 2021 · 被引用 4 次
- SPTS: Single-Point Text SpottingDezhi Peng, Xinyu Wang, Yuliang Liu, Jiaxin Zhang 等ACM MM 2022 · 被引用 65 次
- Text Perceptron: Towards End-to-End Arbitrary-Shaped Text SpottingLiang Qiao, Sanli Tang, Zhanzhan Cheng, Yunlu Xu 等AAAI 2020 · 被引用 128 次
- Towards Unified Scene Text Spotting Based on Sequence GenerationTaeho Kil, Seonghyeon Kim, Sukmin Seo, Yoonsik Kim 等CVPR 2023
- SemiETS: Integrating Spatial and Content Consistencies for Semi-Supervised End-to-end Text SpottingDongliang Luo, Hanshen Zhu, Ziyang Zhang, Dingkang Liang 等CVPR 2025
