Hierarchical Multi-Task Learning for Diagram Question Answering with Multi-Modal Transformer
Zhaoquan Yuan, Xiao Peng, Xiao Wu, Changsheng Xu
Abstract
Diagram question answering (DQA) is an effective way to evaluate the reasoning ability for diagram semantic understanding, which is a very challenging task and largely understudied compared with natural images. Existing separate two-stage methods for DQA are limited in ineffective feedback mechanisms. To address this problem, in this paper, we propose a novel structural parsing-integrated Hierarchical Multi-Task Learning (HMTL) model for diagram question answering based on a multi-modal transformer framework. In the proposed paradigm of multi-task learning, the two tasks of diagram structural parsing and question answering are in the different semantic levels and equipped with different transformer blocks, which constituents a hierarchical architecture. The structural parsing module encodes the information of constituents and their relationships in diagrams, while the diagram question answering module decodes the structural signals and combines question-answers to infer correct answers. Visual diagrams and textual question-answers are interplayed in the multi-modal transformer, which achieves cross-modal semantic comprehension and reasoning. Extensive experiments on the benchmark AI2D and FOODWEBS datasets demonstrate the effectiveness of our proposed HMTL over other state-of-the-art methods.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Cited by top-tier papers1
Ask how each one uses itRelated papers
- GlFoMR: A Glance-then-Focus Multimodal Reasoning Framework for Diagram Question AnsweringYaxian Wang, Bifan Wei, Jun Liu, Lingling Zhang et al.SIGIR 2025
- ISAAQ - Mastering Textbook Questions with Pre-trained Transformers and Bottom-Up and Top-Down AttentionJosé Manuél Gómez-Pérez, Raúl OrtegaEMNLP 2020 · 1 citation
- STL-CQA: Structure-based Transformers with Localization and Encoding for Chart Question AnsweringHrituraj Singh, Sumit ShekharEMNLP 2020 · 42 citations
- SGEITL: Scene Graph Enhanced Image-Text Learning for Visual Commonsense ReasoningZhecan Wang, Haoxuan You, Liunian Harold Li, Alireza Zareian et al.AAAI 2022 · 40 citations
- Iterative Answer Prediction With Pointer-Augmented Multimodal Transformers for TextVQARonghang Hu, Amanpreet Singh, Trevor Darrell, Marcus RohrbachCVPR 2020
