ISAAQ - Mastering Textbook Questions with Pre-trained Transformers and Bottom-Up and Top-Down Attention
José Manuél Gómez-Pérez, Raúl Ortega
摘要
Textbook Question Answering is a complex task in the intersection of Machine Comprehension and Visual Question Answering that requires reasoning with multimodal information from text and diagrams. For the first time, this paper taps on the potential of transformer language models and bottom-up and top-down attention to tackle the language and visual understanding challenges this task entails. Rather than training a language-visual transformer from scratch we rely on pretrained transformers, fine-tuning and ensembling. We add bottom-up and top-down attention to identify regions of interest corresponding to diagram constituents and their relationships, improving the selection of relevant visual information for each question and answer options. Our system ISAAQ reports unprecedented success in all TQA question types, with accuracies of 81.36%, 71.11% and 55.12% on true/false, text-only and diagram multiple choice questions. ISAAQ also demonstrates its broad applicability, obtaining state-of-theart results in other demanding datasets.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- CoG-DQA: Chain-of-Guiding Learning with Large Language Models for Diagram Question AnsweringShaowei Wang, Lingling Zhang, Longji Zhu, Tao Qin 等CVPR 2024 · 被引用 5 次
- MacVQA: Adaptive Memory Allocation and Global Noise Filtering for Continual Visual Question AnsweringZhifei Li, Yiran Wang, Chenyi Xiong, Yujing Xia 等AAAI 2026
- Diagram-Driven Course Questions GenerationXinyu Zhang, Lingling Zhang, Yanrui Wu, Muye Huang 等EMNLP 2025
它引用的顶会 Paper1
相关 Paper
- Hierarchical Multi-Task Learning for Diagram Question Answering with Multi-Modal TransformerZhaoquan Yuan, Xiao Peng, Xiao Wu, Changsheng XuACM MM 2021 · 被引用 10 次
- Iterative Answer Prediction With Pointer-Augmented Multimodal Transformers for TextVQARonghang Hu, Amanpreet Singh, Trevor Darrell, Marcus RohrbachCVPR 2020
- GlFoMR: A Glance-then-Focus Multimodal Reasoning Framework for Diagram Question AnsweringYaxian Wang, Bifan Wei, Jun Liu, Lingling Zhang 等SIGIR 2025
- STL-CQA: Structure-based Transformers with Localization and Encoding for Chart Question AnsweringHrituraj Singh, Sumit ShekharEMNLP 2020 · 被引用 42 次
- Position-Augmented Transformers with Entity-Aligned Mesh for TextVQAXuanyu Zhang, Qing YangACM MM 2021 · 被引用 14 次
