Lune

ACM MM2021Top-tier venue

Position-Augmented Transformers with Entity-Aligned Mesh for TextVQA

Xuanyu Zhang, Qing Yang

2021Year
14Citations
1Top-tier citations

Abstract

In addition to visual components, many images usually contain valuable text information, which is essential for understanding the scene. Thus, we study the TextVQA task that requires reading texts in images to answer corresponding questions. However, most of previous works utilize sophisticated graph structure and manually crafted features to model the position relationship between visual entities and texts in images. And traditional multimodal transformers cannot effectively capture relative position information and original image features. To address these issues in an intuitive but effective way, we propose a novel model, position-augmented transformers with entity-aligned mesh, for the TextVQA task. Different from traditional attention mechanism in transformers, we explicitly introduce continuous relative position information of objects and OCR tokens without complex rules. Furthermore, we replace the complicated graph structure with intuitive entity-aligned mesh according to perspective mapping. In this mesh, the information of discrete entities and image patches at different positions can interact with each other. Extensive experiments on two benchmark datasets (TextVQA and ST-VQA) show that our proposed model is superior to several state-of-the-art methods.

Ask about this paper

Ask your agent about it.

Lune has read the top-tier papers around this one, so every answer names the papers it rests on.

Questions to start from

Your agent calls

Lunesearch_papers

Ask in Lune

Free to start. No credit card required.

lune papers get 6ea7ba33-1d53-4f7d-8c8a-49596b3bbefe

Cited by top-tier papers1

Ask how each one uses it

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines