A Multiplexed Network for End-to-End, Multilingual OCR
Jing Huang, Guan Pang, Rama Kovvuri, Mandy Toh, Kevin J. Liang, Praveen Krishnan, Xi Yin, Tal Hassner
摘要
Recent advances in OCR have shown that an end-toend (E2E) training pipeline that includes both detection and recognition leads to the best results. However, many existing methods focus primarily on Latin-alphabet languages, often even only case-insensitive English characters. In this paper, we propose an E2E approach, Multiplexed Multilingual Mask TextSpotter, that performs script identification at the word level and handles different scripts with different recognition heads, all while maintaining a unified loss that simultaneously optimizes script identification and multiple recognition heads. Experiments show that our method outperforms the single-head model with similar number of parameters in end-to-end recognition tasks, and achieves state-of-the-art results on MLT17 and MLT19 joint text detection and script identification benchmarks. We believe that our work is a step towards the end-to-end trainable and scalable multilingual multi-purpose OCR system. Our code and model will be released.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- ESTextSpotter: Towards Better Scene Text Spotting with Explicit Synergy in TransformerMingxin Huang, Jiaxin Zhang, Dezhi Peng, Hao Lu 等ICCV 2023 · 被引用 44 次
- MRN: Multiplexed Routing Network for Incremental Multilingual Text RecognitionTianlun Zheng, Zhineng Chen, Bingchen Huang, Wei Zhang 等ICCV 2023 · 被引用 17 次
- Kernel Adaptive Convolution for Scene Text Detection via Distance Map PredictionJinzhi Zheng, Heng Fan, Libo ZhangCVPR 2024 · 被引用 14 次
- TextOCR: Towards Large-Scale End-to-End Reasoning for Arbitrary-Shaped Scene TextAmanpreet Singh, Guan Pang, Mandy Toh, Jing Huang 等CVPR 2021
它引用的顶会 Paper7
- LEEP: A New Measure to Evaluate Transferability of Learned RepresentationsCuong V. Nguyen, Tal Hassner, Matthias W. Seeger, Cédric ArchambeauICML 2020 · 被引用 279 次
- Transferability and Hardness of Supervised Classification TasksAnh Tuan Tran, Cuong V. Nguyen, Tal HassnerICCV 2019 · 被引用 201 次
- Convolutional Character NetworksLinjie Xing, Zhi Tian, Weilin Huang, Matthew R. ScottICCV 2019 · 被引用 176 次
- Towards Unconstrained End-to-End Text SpottingSiyang Qin, Alessandro Bissacco, Michalis Raptis, Yasuhisa Fujii 等ICCV 2019 · 被引用 138 次
- Text Perceptron: Towards End-to-End Arbitrary-Shaped Text SpottingLiang Qiao, Sanli Tang, Zhanzhan Cheng, Yunlu Xu 等AAAI 2020 · 被引用 128 次
相关 Paper
- Towards Weakly-Supervised Text Spotting using a Multi-Task TransformerYair Kittenplon, Inbal Lavi, Sharon Fogel, Yarin Bar 等CVPR 2022 · 被引用 60 次
- SwinTextSpotter: Scene Text Spotting via Better Synergy between Text Detection and Text RecognitionMingxin Huang, Yuliang Liu, Zhenghao Peng, Chongyu Liu 等CVPR 2022 · 被引用 150 次
- MANGO: A Mask Attention Guided One-Stage Scene Text SpotterLiang Qiao, Ying Chen, Zhanzhan Cheng, Yunlu Xu 等AAAI 2021 · 被引用 91 次
- SPTS: Single-Point Text SpottingDezhi Peng, Xinyu Wang, Yuliang Liu, Jiaxin Zhang 等ACM MM 2022 · 被引用 65 次
- TrOCR: Transformer-Based Optical Character Recognition with Pre-trained ModelsMinghao Li, Tengchao Lv, Jingye Chen, Lei Cui 等AAAI 2023 · 被引用 607 次
