Position-Augmented Transformers with Entity-Aligned Mesh for TextVQA
Xuanyu Zhang, Qing Yang
Abstract
In addition to visual components, many images usually contain valuable text information, which is essential for understanding the scene. Thus, we study the TextVQA task that requires reading texts in images to answer corresponding questions. However, most of previous works utilize sophisticated graph structure and manually crafted features to model the position relationship between visual entities and texts in images. And traditional multimodal transformers cannot effectively capture relative position information and original image features. To address these issues in an intuitive but effective way, we propose a novel model, position-augmented transformers with entity-aligned mesh, for the TextVQA task. Different from traditional attention mechanism in transformers, we explicitly introduce continuous relative position information of objects and OCR tokens without complex rules. Furthermore, we replace the complicated graph structure with intuitive entity-aligned mesh according to perspective mapping. In this mesh, the information of discrete entities and image patches at different positions can interact with each other. Extensive experiments on two benchmark datasets (TextVQA and ST-VQA) show that our proposed model is superior to several state-of-the-art methods.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 6ea7ba33-1d53-4f7d-8c8a-49596b3bbefeCited by top-tier papers1
Ask how each one uses itRelated papers
- Iterative Answer Prediction With Pointer-Augmented Multimodal Transformers for TextVQARonghang Hu, Amanpreet Singh, Trevor Darrell, Marcus RohrbachCVPR 2020
- Locate Then Generate: Bridging Vision and Language with Bounding Box for Scene-Text VQAYongxin Zhu, Zhen Liu, Yukang Liang, Xin Li et al.AAAI 2023 · 11 citations
- TAP: Text-Aware Pre-Training for Text-VQA and Text-CaptionZhengyuan Yang, Yijuan Lu, Jianfeng Wang, Xi Yin et al.CVPR 2021
- Beyond OCR + VQA: Involving OCR into the Flow for Robust and Accurate TextVQAGangyan Zeng, Yuan Zhang, Yu Zhou, Xiaomeng YangACM MM 2021 · 38 citations
- Separate and Locate: Rethink the Text in Text-based Visual Question AnsweringChengyang Fang, Jiangnan Li, Liang Li, Can Ma et al.ACM MM 2023 · 18 citations
