Generating Explanations for Embodied Action Decision from Visual Observation
Xiaohan Wang, Yuehu Liu, Xinhang Song, Beibei Wang, Shuqiang Jiang
Abstract
Getting trust is crucial for embodied agents (such as robots and autonomous vehicles) to collaborate with human beings, especially non-experts. The most direct way for mutual understanding is through natural language explanation. Existing researches consider generating visual explanations for object recognition, while the exploration of explaining embodied decisions remains vacant. In this paper, we study generating action decisions and explanations based on visual observation. Distinct to explanations for recognition, justifying an action needs to show why it's better than other actions. Besides, the understanding of scene structure is required since the agent needs to interact with the environment (e.g. navigation, moving objects). We introduce a new dataset THOR-EAE (Embodied Action Explanation) collected based on AI2-THOR simulator. The dataset consists of over 840,000 egocentric images of indoor embodied observation which are annotated with the optimal action labels and explanation sentences. An explainable decision-making criterion is developed considering scene layout and action attributes for efficient annotation. We propose a graph action justification model, exploiting graph neural networks for obstacle-surroundings relations representation and justifying the actions under the guidance of decision results. Experimental results on THOR-EAE dataset showcase its challenge and the effectiveness of the proposed method.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get c93eeeda-84b7-465a-8047-cb66f26685d6Cited by top-tier papers5
- Lookahead Exploration with Neural Radiance Representation for Continuous Vision-Language NavigationZihan Wang, Xiangyang Li, Jiahao Yang, Yeqi Liu et al.CVPR 2024 · 13 citations
- Trial-Oriented Visual RearrangementYuyi Liu, Xinhang Song, Tianliang Qi, Shuqiang JiangICCV 2025 · 1 citation
- EgoExoLearn: A Dataset for Bridging Asynchronous Ego- and Exo-centric View of Procedural Activities in Real WorldYifei Huang, Guo Chen, Jilan Xu, Mingfang Zhang et al.CVPR 2024
- Rethinking Visual Rearrangement from A Diffusion PerspectiveTianliang Qi, Xinhang Song, Yuyi Liu, Shuqiang JiangCVPR 2026
- A Category Agnostic Model for Visual RearrangmentYuyi Liu, Xinhang Song, Weijie Li, Xiaohan Wang et al.CVPR 2024
Related papers
- Towards Explainable Action Recognition by Salient Qualitative Spatial Object Relation ChainsHua Hua, Dongxu Li, Ruiqi Li, Peng Zhang et al.AAAI 2022 · 7 citations
- AI-VQA: Visual Question Answering based on Agent Interaction with InterpretabilityRengang Li, Cong Xu, Zhenhua Guo, Baoyu Fan et al.ACM MM 2022 · 7 citations
- When to Explain: Modeling User Need for Explanations in Real-World Autonomous DrivingShihong Ling, Yaohan Ding, Yu Liu, Yue Wan et al.CHI 2026 · 1 citation
- Generating Explanations to Understand and Repair Embedding-Based Entity AlignmentXiaobin Tian, Zequn Sun, Wei HuICDE 2024 · 6 citations
- ION: Instance-level Object NavigationWeijie Li, Xinhang Song, Yubing Bai, Sixian Zhang et al.ACM MM 2021 · 25 citations
