ICLR2024

EQA-MX: Embodied Question Answering using Multimodal Expression

Md Mofijul Islam, Alexi Gladstone, Riashat Islam, Tariq Iqbal

被引用 18 次

摘要

Figure 1: Compared to the QA tasks in existing VQA (Antol et al., 2015) and EQA (Das et al., 2018a) datasets, the models to answer EQA tasks in our EQA-MX dataset require the reasoning of questions with multimodal expressions (verbal and nonverbal gestures).