ICLR2024
EQA-MX: Embodied Question Answering using Multimodal Expression
Md Mofijul Islam, Alexi Gladstone, Riashat Islam, Tariq Iqbal
18 citations
Abstract
Figure 1: Compared to the QA tasks in existing VQA (Antol et al., 2015) and EQA (Das et al., 2018a) datasets, the models to answer EQA tasks in our EQA-MX dataset require the reasoning of questions with multimodal expressions (verbal and nonverbal gestures).