InstructOCR: Instruction Boosting Scene Text Spotting
Chen Duan, Qianyi Jiang, Pei Fu, Jiamin Chen, Shengxi Li, Zining Wang, Shan Guo, Junfeng Luo
Abstract
In the field of scene text spotting, previous OCR methods primarily relied on image encoders and pre-trained text information, but they often overlooked the advantages of incorporating human language instructions. To address this gap, we propose InstructOCR, an innovative instruction-based scene text spotting model that leverages human language instructions to enhance the understanding of text within images. Our framework employs both text and image encoders during training and inference, along with instructions meticulously designed based on text attributes. This approach enables the model to interpret text more accurately and flexibly. Extensive experiments demonstrate the effectiveness of our model and we achieve state-of-the-art results on widely used benchmarks. Furthermore, the proposed framework can be seamlessly applied to scene text VQA tasks. By leveraging instruction strategies during pre-training, the performance on downstream VQA tasks can be significantly improved, with a 2.6% increase on the TextVQA dataset and a 2.1% increase on the ST-VQA dataset. These experimental results provide insights into the benefits of incorporating human language instructions for OCR-related tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a4ce70c1-8e9e-439f-abec-c492a402ed57Cited by top-tier papers1
Ask how each one uses itBuilds on34
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech et al.NeurIPS 2022 · 6,707 citations
- Scene Text Visual Question AnsweringAli Furkan Biten, Rubèn Tito, Andrés Mafla, Lluís Gómez i Bigorda et al.ICCV 2019 · 482 citations
- Pix2seq: A Language Modeling Framework for Object DetectionTing Chen, Saurabh Saxena, Lala Li, David J. Fleet et al.ICLR 2022 · 435 citations
Related papers
- PreSTU: Pre-Training for Scene-Text UnderstandingJihyung Kil, Soravit Changpinyo, Xi Chen, Hexiang Hu et al.ICCV 2023 · 39 citations
- TAP: Text-Aware Pre-Training for Text-VQA and Text-CaptionZhengyuan Yang, Yijuan Lu, Jianfeng Wang, Xi Yin et al.CVPR 2021
- ODM: A Text-Image Further Alignment Pre-training Approach for Scene Text Detection and SpottingChen Duan, Pei Fu, Shan Guo, Qianyi Jiang et al.CVPR 2024
- You Can even Annotate Text with Voice: Transcription-only-Supervised Text SpottingJingqun Tang, Su Qiao, Benlei Cui, Yuhang Ma et al.ACM MM 2022 · 22 citations
- SPTS: Single-Point Text SpottingDezhi Peng, Xinyu Wang, Yuliang Liu, Jiaxin Zhang et al.ACM MM 2022 · 65 citations
