Extending Embodied Question Answering from Perception to Decision
Xicheng Gong, Qiwei Li, Peiran Xu, Yadong Mu
Abstract
Embodied Question Answering (EQA) connects perception, reasoning, and interaction within embodied environments. However, existing datasets and benchmarks remain fragmented, each focusing on a limited subset of reasoning skills such as spatial understanding or procedural reasoning, without offering a unified large-scale framework for comprehensive evaluation. We present EQA-Decision, a large-scale embodied QA dataset that systematically covers four complementary dimensions of embodied reasoning: static scene construction, spatial understanding, task dynamics reasoning, and instant decision. The dataset contains over four million question-answer pairs with hierarchical annotations across diverse embodied scenarios. In addition, we develop RoboDecision, a strong baseline model aligned with the EQA-Decision Benchmark, providing a unified framework that jointly evaluates perception, reasoning, and action-level decision-making in embodied environments. Results demonstrate that EQA-Decision effectively benchmarks and enhances VLM capabilities in spatial and interaction reasoning, providing a solid foundation for advancing embodied intelligence research.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b483f73e-d853-415d-9e5d-56bfc35c67e7Builds on17
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- PaLM-E: An Embodied Multimodal Language ModelDanny Driess, Fei Xia, Mehdi S. M. Sajjadi, Corey Lynch et al.ICML 2023 · 2,601 citations
- Ego4D: Around the World in 3, 000 Hours of Egocentric VideoKristen Grauman, Andrew Westbury, Eugene Byrne, Zachary Chavis et al.CVPR 2022 · 525 citations
- EmbodiedGPT: Vision-Language Pre-Training via Embodied Chain of ThoughtYao Mu, Qinglong Zhang, Mengkang Hu, Wenhai Wang et al.NeurIPS 2023 · 453 citations
- DreamVLA: A Vision-Language-Action Model Dreamed with Comprehensive World KnowledgeWenyao Zhang, Hongsi Liu, Zekun Qi, Yunnan Wang et al.NeurIPS 2025 · 244 citations
Related papers
- Beyond the Destination: A Novel Benchmark for Exploration-Aware Embodied Question AnsweringKaixuan Jiang, Yang Liu, Weixing Chen, Jingzhou Luo et al.ICCV 2025 · 4 citations
- ECBench: Can Multi-modal Foundation Models Understand the Egocentric World? A Holistic Embodied Cognition BenchmarkRonghao Dang, Yuqian Yuan, Wenqi Zhang, Yifei Xin et al.CVPR 2025
- SQA3D: Situated Question Answering in 3D ScenesXiaojian Ma, Silong Yong, Zilong Zheng, Qing Li et al.ICLR 2023 · 16 citations
- Understanding Dynamic Scenes in Ego Centric 4D Point CloudsJunsheng Huang, Shengyu Hao, Bocheng Hu, Hongwei Wang et al.AAAI 2026 · 4 citations
- CityEQA: A Hierarchical LLM Agent on Embodied Question Answering Benchmark in City SpaceYong Zhao, Kai Xu, Zhengqiu Zhu, Yue Hu et al.EMNLP 2025 · 3 citations
