Lune

ACM MM2021顶会

Position-Augmented Transformers with Entity-Aligned Mesh for TextVQA

Xuanyu Zhang, Qing Yang

2021年份
14被引次数
1顶会引用

摘要

In addition to visual components, many images usually contain valuable text information, which is essential for understanding the scene. Thus, we study the TextVQA task that requires reading texts in images to answer corresponding questions. However, most of previous works utilize sophisticated graph structure and manually crafted features to model the position relationship between visual entities and texts in images. And traditional multimodal transformers cannot effectively capture relative position information and original image features. To address these issues in an intuitive but effective way, we propose a novel model, position-augmented transformers with entity-aligned mesh, for the TextVQA task. Different from traditional attention mechanism in transformers, we explicitly introduce continuous relative position information of objects and OCR tokens without complex rules. Furthermore, we replace the complicated graph structure with intuitive entity-aligned mesh according to perspective mapping. In this mesh, the information of discrete entities and image patches at different positions can interact with each other. Extensive experiments on two benchmark datasets (TextVQA and ST-VQA) show that our proposed model is superior to several state-of-the-art methods.

问问这篇 Paper

问问你的智能体。

Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。

可以从这些问题问起

智能体调用

Lunesearch_papers

在 Lune 里问

免费开始,无需绑卡

lune papers get 6ea7ba33-1d53-4f7d-8c8a-49596b3bbefe

引用它的顶会 Paper1

问问它们各自怎么用它

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖