Beyond OCR + VQA: Involving OCR into the Flow for Robust and Accurate TextVQA
Gangyan Zeng, Yuan Zhang, Yu Zhou, Xiaomeng Yang
Abstract
Text-based visual question answering (TextVQA) requires analyzing both the visual contents and texts in an image to answer a question, which is more practical than general visual question answering (VQA). Existing efforts tend to regard optical character recognition (OCR) as a pre-processing and then combine it with a VQA framework. It makes the performance of multimodal reasoning and question answering highly depend on the accuracy of OCR. In this work, we address this issue with two perspectives. First, we take advantages of multimodal cues to complete the semantic information of texts. A visually enhanced text embedding is proposed to enable understanding of texts without accurately recognizing them. Second, we further leverage rich contextual information to modify the answer texts even if the OCR module does not correctly recognize them. In addition, the visual objects are endued with semantic representations to enable objects in the same semantic space as OCR tokens. Equipped with these techniques, the cumulative error propagation caused by poor OCR performance is effectively suppressed. Extensive experiments on TextVQA and ST-VQA datasets demonstrate that our approach achieves the state-of-the-art performance in terms of accuracy and robustness.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 477ac2de-f52e-486d-8642-4e368b8b62dfCited by top-tier papers8
- ViSTA: Vision and Scene Text Aggregation for Cross-Modal RetrievalMengjun Cheng, Yipeng Sun, Longchao Wang, Xiongwei Zhu et al.CVPR 2022 · 86 citations
- LaTr: Layout-Aware Transformer for Scene-Text VQAAli Furkan Biten, Ron Litman, Yusheng Xie, Srikar Appalaraju et al.CVPR 2022 · 82 citations
- TPSNet: Reverse Thinking of Thin Plate Splines for Arbitrary Shape Scene Text RepresentationWei Wang, Yu Zhou, Jiahao Lyu, Dayan Wu et al.ACM MM 2022 · 35 citations
- Towards Robust Real-Time Scene Text Detection: From Semantic to Instance Representation LearningXugong Qin, Pengyuan Lyu, Chengquan Zhang, Yu Zhou et al.ACM MM 2023 · 21 citations
- Separate and Locate: Rethink the Text in Text-based Visual Question AnsweringChengyang Fang, Jiangnan Li, Liang Li, Can Ma et al.ACM MM 2023 · 18 citations
Related papers
- From Token to Word: OCR Token Evolution via Contrastive Learning and Semantic Matching for Text-VQAZan-Xia Jin, Mike Zheng Shou, Fang Zhou, Satoshi Tsutsui et al.ACM MM 2022 · 11 citations
- Cascade Reasoning Network for Text-based Visual Question AnsweringFen Liu, Guanghui Xu, Qi Wu, Qing Du et al.ACM MM 2020 · 61 citations
- From Strings to Things: Knowledge-Enabled VQA Model That Can Read and ReasonAjeet Kumar Singh, Anand Mishra, Shashank Shekhar, Anirban ChakrabortyICCV 2019 · 54 citations
- Track the Answer: Extending TextVQA from Image to Video with Spatio-Temporal CluesYan Zhang, Gangyan Zeng, Huawen Shen, Daiqing Wu et al.AAAI 2025 · 1 citation
- Simple is not Easy: A Simple Strong Baseline for TextVQA and TextCapsQi Zhu, Chenyu Gao, Peng Wang, Qi WuAAAI 2021 · 59 citations
