TextOCR: Towards Large-Scale End-to-End Reasoning for Arbitrary-Shaped Scene Text
Amanpreet Singh, Guan Pang, Mandy Toh, Jing Huang, Wojciech Galuba, Tal Hassner
摘要
A crucial component for the scene text based reasoning required for TextVQA and TextCaps datasets involve detecting and recognizing text present in the images using an optical character recognition (OCR) system. The current systems are crippled by the unavailability of ground truth text annotations for these datasets as well as lack of scene text detection and recognition datasets on real images disallowing the progress in the field of OCR and evaluation of scene text based reasoning in isolation from OCR systems. In this work, we propose TextOCR, an arbitrary-shaped scene text detection and recognition with 900k annotated words collected on real images from TextVQA dataset. We show that current state-of-the-art text-recognition (OCR) models fail to perform well on TextOCR and that training on TextOCR helps achieve state-of-the-art performance on multiple other OCR datasets as well. We use a TextOCR trained OCR model to create PixelM4C model which can do scene text based reasoning on an image in an end-to-end fashion, allowing us to revisit several design choices to achieve new state-of-the-art performance on TextVQA dataset.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper56
- VisionLLM v2: An End-to-End Generalist Multimodal Large Language Model for Hundreds of Vision-Language TasksJiannan Wu, Muyan Zhong, Sen Xing, Zeqiang Lai 等NeurIPS 2024 · 被引用 179 次
- Towards End-to-End Unified Scene Text Detection and Layout AnalysisShangbang Long, Siyang Qin, Dmitry Panteleev, Alessandro Bissacco 等CVPR 2022 · 被引用 86 次
- LaTr: Layout-Aware Transformer for Scene-Text VQAAli Furkan Biten, Ron Litman, Yusheng Xie, Srikar Appalaraju 等CVPR 2022 · 被引用 82 次
- Detect Anything via Next Point PredictionQing Jiang, Junan Huo, Xingyu Chen, Yuda Xiong 等CVPR 2026 · 被引用 79 次
- Revisiting Scene Text Recognition: A Data PerspectiveQing Jiang, Jiapeng Wang, Dezhi Peng, Chongyu Liu 等ICCV 2023 · 被引用 70 次
它引用的顶会 Paper7
- The Hateful Memes Challenge: Detecting Hate Speech in Multimodal MemesDouwe Kiela, Hamed Firooz, Aravind Mohan, Vedanuj Goswami 等NeurIPS 2020 · 被引用 1,022 次
- Scene Text Visual Question AnsweringAli Furkan Biten, Rubèn Tito, Andrés Mafla, Lluís Gómez i Bigorda 等ICCV 2019 · 被引用 482 次
- All You Need Is Boundary: Toward Arbitrary-Shaped Text SpottingHao Wang, Pu Lu, Hui Zhang, Mingkun Yang 等AAAI 2020 · 被引用 145 次
- Towards Unconstrained End-to-End Text SpottingSiyang Qin, Alessandro Bissacco, Michalis Raptis, Yasuhisa Fujii 等ICCV 2019 · 被引用 138 次
- A Multiplexed Network for End-to-End, Multilingual OCRJing Huang, Guan Pang, Rama Kovvuri, Mandy Toh 等CVPR 2021
相关 Paper
- TAP: Text-Aware Pre-Training for Text-VQA and Text-CaptionZhengyuan Yang, Yijuan Lu, Jianfeng Wang, Xi Yin 等CVPR 2021
- Beyond OCR + VQA: Involving OCR into the Flow for Robust and Accurate TextVQAGangyan Zeng, Yuan Zhang, Yu Zhou, Xiaomeng YangACM MM 2021 · 被引用 38 次
- Simple is not Easy: A Simple Strong Baseline for TextVQA and TextCapsQi Zhu, Chenyu Gao, Peng Wang, Qi WuAAAI 2021 · 被引用 59 次
- From Token to Word: OCR Token Evolution via Contrastive Learning and Semantic Matching for Text-VQAZan-Xia Jin, Mike Zheng Shou, Fang Zhou, Satoshi Tsutsui 等ACM MM 2022 · 被引用 11 次
- Towards Reasoning Ability in Scene Text Visual Question AnsweringQingqing Wang, Liqiang Xiao, Yue Lu, Yaohui Jin 等ACM MM 2021 · 被引用 12 次
