EQA-MX: Embodied Question Answering using Multimodal Expression
Md Mofijul Islam, Alexi Gladstone, Riashat Islam, Tariq Iqbal
摘要
Humans predominantly use verbal utterances and nonverbal gestures (e.g., eye gaze and pointing gestures) in their natural interactions. For instance, pointing gestures and verbal information is often required to comprehend questions such as "what object is that?" Thus, this question-answering (QA) task involves complex reasoning of multimodal expressions (verbal utterances and nonverbal gestures). However, prior works have explored QA tasks in non-embodied settings, where questions solely contain verbal utterances from a single verbal and visual perspective. In this paper, we have introduced 8 novel embodied question answering (EQA) tasks to develop learning models to comprehend embodied questions with multimodal expressions. We have developed a novel large-scale dataset, EQA-MX, with over 8 million diverse embodied QA data samples involving multimodal expressions from multiple visual and verbal perspectives. To learn salient multimodal representations from discrete verbal embeddings and continuous wrapping of multiview visual representations, we propose a vector-quantization (VQ) based multimodal representation learning model, VQ-Fusion, for the EQA tasks. Our extensive experimental results suggest that VQ-Fusion can improve the performance of existing state-of-the-art visual-language models up to 13% across EQA tasks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Energy-Based Transformers are Scalable Learners and ThinkersAlexi Gladstone, Ganesh Nanduru, Md Mofijul Islam, Peixuan Han 等ICLR 2026 · 被引用 38 次
- Extending Embodied Question Answering from Perception to DecisionXicheng Gong, Qiwei Li, Peiran Xu, Yadong MuCVPR 2026 · 被引用 1 次
它引用的顶会 Paper12
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- ViLT: Vision-and-Language Transformer Without Convolution or Region SupervisionWonjae Kim, Bokyung Son, Ildoo KimICML 2021 · 被引用 2,258 次
- MMAct: A Large-Scale Dataset for Cross Modal Human Action UnderstandingQuan Kong, Ziming Wu, Ziwei Deng, Martin Klinkigt 等ICCV 2019 · 被引用 108 次
- SpatialSense: An Adversarially Crowdsourced Benchmark for Spatial Relation RecognitionKaiyu Yang, Olga Russakovsky, Jia DengICCV 2019 · 被引用 81 次
相关 Paper
- SocialGesture: Delving into Multi-person Gesture UnderstandingXu Cao, Pranav Virupaksha, Wenqi Jia, Bolin Lai 等CVPR 2025
- PATRON: Perspective-Aware Multitask Model for Referring Expression Grounding Using Embodied Multimodal CuesMd Mofijul Islam, Alexi Gladstone, Tariq IqbalAAAI 2023 · 被引用 10 次
- AVQA: A Dataset for Audio-Visual Question Answering on VideosPinci Yang, Xin Wang, Xuguang Duan, Hong Chen 等ACM MM 2022 · 被引用 60 次
- PaLM-E: An Embodied Multimodal Language ModelDanny Driess, Fei Xia, Mehdi S. M. Sajjadi, Corey Lynch 等ICML 2023 · 被引用 2,601 次
- VTQA: Visual Text Question Answering via Entity Alignment and Cross-Media ReasoningKang Chen, Xiangqian WuCVPR 2024
