Generating Explanations for Embodied Action Decision from Visual Observation
Xiaohan Wang, Yuehu Liu, Xinhang Song, Beibei Wang, Shuqiang Jiang
摘要
Getting trust is crucial for embodied agents (such as robots and autonomous vehicles) to collaborate with human beings, especially non-experts. The most direct way for mutual understanding is through natural language explanation. Existing researches consider generating visual explanations for object recognition, while the exploration of explaining embodied decisions remains vacant. In this paper, we study generating action decisions and explanations based on visual observation. Distinct to explanations for recognition, justifying an action needs to show why it's better than other actions. Besides, the understanding of scene structure is required since the agent needs to interact with the environment (e.g. navigation, moving objects). We introduce a new dataset THOR-EAE (Embodied Action Explanation) collected based on AI2-THOR simulator. The dataset consists of over 840,000 egocentric images of indoor embodied observation which are annotated with the optimal action labels and explanation sentences. An explainable decision-making criterion is developed considering scene layout and action attributes for efficient annotation. We propose a graph action justification model, exploiting graph neural networks for obstacle-surroundings relations representation and justifying the actions under the guidance of decision results. Experimental results on THOR-EAE dataset showcase its challenge and the effectiveness of the proposed method.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper5
- Lookahead Exploration with Neural Radiance Representation for Continuous Vision-Language NavigationZihan Wang, Xiangyang Li, Jiahao Yang, Yeqi Liu 等CVPR 2024 · 被引用 13 次
- Trial-Oriented Visual RearrangementYuyi Liu, Xinhang Song, Tianliang Qi, Shuqiang JiangICCV 2025 · 被引用 1 次
- EgoExoLearn: A Dataset for Bridging Asynchronous Ego- and Exo-centric View of Procedural Activities in Real WorldYifei Huang, Guo Chen, Jilan Xu, Mingfang Zhang 等CVPR 2024
- Rethinking Visual Rearrangement from A Diffusion PerspectiveTianliang Qi, Xinhang Song, Yuyi Liu, Shuqiang JiangCVPR 2026
- A Category Agnostic Model for Visual RearrangmentYuyi Liu, Xinhang Song, Weijie Li, Xiaohan Wang 等CVPR 2024
相关 Paper
- Towards Explainable Action Recognition by Salient Qualitative Spatial Object Relation ChainsHua Hua, Dongxu Li, Ruiqi Li, Peng Zhang 等AAAI 2022 · 被引用 7 次
- AI-VQA: Visual Question Answering based on Agent Interaction with InterpretabilityRengang Li, Cong Xu, Zhenhua Guo, Baoyu Fan 等ACM MM 2022 · 被引用 7 次
- When to Explain: Modeling User Need for Explanations in Real-World Autonomous DrivingShihong Ling, Yaohan Ding, Yu Liu, Yue Wan 等CHI 2026 · 被引用 1 次
- Generating Explanations to Understand and Repair Embedding-Based Entity AlignmentXiaobin Tian, Zequn Sun, Wei HuICDE 2024 · 被引用 6 次
- ION: Instance-level Object NavigationWeijie Li, Xinhang Song, Yubing Bai, Sixian Zhang 等ACM MM 2021 · 被引用 25 次
