Text Spotting Transformers
Xiang Zhang, Yongwen Su, Subarna Tripathi, Zhuowen Tu
Abstract
In this paper, we present TExt Spotting TRansformers (TESTR), a generic end-to-end text spotting framework using Transformers for text detection and recognition in the wild. TESTR builds upon a single encoder and dual decoders for the joint text-box control point regression and character recognition. Other than most existing literature, our method is free from Region-of-Interest operations and heuristics-driven post-processing procedures; TESTR is particularly effective when dealing with curved text-boxes where special cares are needed for the adaptation of the tra-ditional bounding-box representations. We show our canonical representation of control points suitable for text in-stances in both Bezier curve and polygon annotations. In addition, we design a bounding-box guided polygon detection (box-to-polygon) process. Experiments on curved and arbitrarily shaped datasets demonstrate state-of-the-art performances of the proposed TESTR algorithm.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ce22165a-a47b-4837-a7d0-a7755f531c5dCited by top-tier papers20
- DPText-DETR: Towards Better Scene Text Detection with Dynamic Points in TransformerMaoyuan Ye, Jing Zhang, Shanshan Zhao, Juhua Liu et al.AAAI 2023 · 123 citations
- ESTextSpotter: Towards Better Scene Text Spotting with Explicit Synergy in TransformerMingxin Huang, Jiaxin Zhang, Dezhi Peng, Hao Lu et al.ICCV 2023 · 44 citations
- OMNIPARSER: A Unified Framework for Text Spotting, Key Information Extraction and Table RecognitionJianqiang Wan, Sibo Song, Wenwen Yu, Yuliang Liu et al.CVPR 2024 · 29 citations
- SRFormer: Text Detection Transformer with Incorporated Segmentation and RegressionQingwen Bu, Sungrae Park, Minsoo Khang, Yichuan ChengAAAI 2024 · 16 citations
- GoMatching: A Simple Baseline for Video Text Spotting via Long and Short Term MatchingHaibin He, Maoyuan Ye, Jing Zhang, Juhua Liu et al.NeurIPS 2024 · 16 citations
Builds on13
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li et al.ICLR 2021 · 7,353 citations
- TextDragon: An End-to-End Framework for Arbitrary Shaped Text SpottingWei Feng, Wenhao He, Fei Yin, Xu-Yao Zhang et al.ICCV 2019 · 212 citations
- Convolutional Character NetworksLinjie Xing, Zhi Tian, Weilin Huang, Matthew R. ScottICCV 2019 · 176 citations
- All You Need Is Boundary: Toward Arbitrary-Shaped Text SpottingHao Wang, Pu Lu, Hui Zhang, Mingkun Yang et al.AAAI 2020 · 145 citations
Related papers
- SPTS: Single-Point Text SpottingDezhi Peng, Xinyu Wang, Yuliang Liu, Jiaxin Zhang et al.ACM MM 2022 · 65 citations
- Towards Weakly-Supervised Text Spotting using a Multi-Task TransformerYair Kittenplon, Inbal Lavi, Sharon Fogel, Yarin Bar et al.CVPR 2022 · 60 citations
- SwinTextSpotter: Scene Text Spotting via Better Synergy between Text Detection and Text RecognitionMingxin Huang, Yuliang Liu, Zhenghao Peng, Chongyu Liu et al.CVPR 2022 · 150 citations
- DeepSolo: Let Transformer Decoder with Explicit Points Solo for Text SpottingMaoyuan Ye, Jing Zhang, Shanshan Zhao, Juhua Liu et al.CVPR 2023
- PBFormer: Capturing Complex Scene Text Shape with Polynomial Band TransformerRuijin Liu, Ning Lu, Dapeng Chen, Cheng Li et al.ACM MM 2023 · 2 citations
