DEFT: Demystifying VLN Failures via a Unified Dual-View Explainability Framework for LLM-based Agents
Yawen Wang, Yihan Dai, Jianming Chen, Junjie Wang, Qing Wang
Abstract
Large Language Models (LLMs) have emerged as central planners in Vision-and-Language Navigation (VLN), yet their complexity increasingly obscures their internal decisionmaking. Existing interpretability methods typically isolate temporal criticality from feature salience, creating an alignment gap and failing to account for the behavioral instability of black-box agents. To address this, we propose DEFT, a unified dual-view framework that demystifies agent behavior by jointly analyzing when a decision is pivotal and what visual evidence grounds it. Featuring a dualhead architecture with a shared latent representation, DEFT employs a Mask Head for counterfactual-based criticality detection and an Action Head that leverages an ensemble of surrogates to recover robust visual cues. Extensive experiments on MatterPort3D across three LLM-based agents demonstrate that DEFT outperforms baselines in both temporal and feature fidelity. User studies further validate its utility, showing 78% alignment with human intuition.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5857848a-6ff1-4f1b-ac73-623a8eb3c0adBuilds on11
- Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought PromptingMiles Turpin, Julian Michael, Ethan Perez, Samuel R. BowmanNeurIPS 2023 · 1,792 citations
- Think Global, Act Local: Dual-scale Graph Transformer for Vision-and-Language NavigationShizhe Chen, Pierre-Louis Guhur, Makarand Tapaswi, Cordelia Schmid et al.CVPR 2022 · 150 citations
- EDGE: Explaining Deep Reinforcement Learning PoliciesWenbo Guo, Xian Wu, Usmann Khan, Xinyu XingNeurIPS 2021 · 79 citations
- Making Sense of Dependence: Efficient Black-box Explanations Using Dependence MeasurePaul Novello, Thomas Fel, David VigourouxNeurIPS 2022 · 48 citations
- Rethinking the Role of Gradient-based Attribution Methods for Model InterpretabilitySuraj Srinivas, François FleuretICLR 2021 · 46 citations
Related papers
- Run, Ruminate, and Regulate: A Dual-process Thinking System for Vision-and-Language NavigationYu Zhong, Zihao Zhang, Rui Zhang, Lingdong Huang et al.AAAI 2026
- ProFocus: Proactive Perception and Focused Reasoning in Vision-and-Language NavigationWei Xue, Mingcheng Li, Xuecheng Wu, Jingqun Tang et al.CVPR 2026 · 4 citations
- VLN-MME: Diagnosing MLLMs as Language-guided Visual Navigation AgentsXunyi Zhao, Gengze Zhou, Qi WuACL 2026 · 3 citations
- UNeMo: Collaborative Visual-Language Reasoning and Navigation via a Multimodal World ModelChangxin Huang, Lv Tang, Zhaohuan Zhan, Lisha Yu et al.AAAI 2026 · 2 citations
- Not All Inconsistency Is Equal: Decomposing LVLM Uncertainty into Belief Divergence and Belief ConflictJie Shi, Xiaodong Yue, Wei Liu, Yufei Chen et al.AAAI 2026 · 1 citation
