DEFT: Demystifying VLN Failures via a Unified Dual-View Explainability Framework for LLM-based Agents
Yawen Wang, Yihan Dai, Jianming Chen, Junjie Wang, Qing Wang
摘要
Large Language Models (LLMs) have emerged as central planners in Vision-and-Language Navigation (VLN), yet their complexity increasingly obscures their internal decisionmaking. Existing interpretability methods typically isolate temporal criticality from feature salience, creating an alignment gap and failing to account for the behavioral instability of black-box agents. To address this, we propose DEFT, a unified dual-view framework that demystifies agent behavior by jointly analyzing when a decision is pivotal and what visual evidence grounds it. Featuring a dualhead architecture with a shared latent representation, DEFT employs a Mask Head for counterfactual-based criticality detection and an Action Head that leverages an ensemble of surrogates to recover robust visual cues. Extensive experiments on MatterPort3D across three LLM-based agents demonstrate that DEFT outperforms baselines in both temporal and feature fidelity. User studies further validate its utility, showing 78% alignment with human intuition.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper11
- Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought PromptingMiles Turpin, Julian Michael, Ethan Perez, Samuel R. BowmanNeurIPS 2023 · 被引用 1,792 次
- Think Global, Act Local: Dual-scale Graph Transformer for Vision-and-Language NavigationShizhe Chen, Pierre-Louis Guhur, Makarand Tapaswi, Cordelia Schmid 等CVPR 2022 · 被引用 150 次
- EDGE: Explaining Deep Reinforcement Learning PoliciesWenbo Guo, Xian Wu, Usmann Khan, Xinyu XingNeurIPS 2021 · 被引用 79 次
- Making Sense of Dependence: Efficient Black-box Explanations Using Dependence MeasurePaul Novello, Thomas Fel, David VigourouxNeurIPS 2022 · 被引用 48 次
- Rethinking the Role of Gradient-based Attribution Methods for Model InterpretabilitySuraj Srinivas, François FleuretICLR 2021 · 被引用 46 次
相关 Paper
- Run, Ruminate, and Regulate: A Dual-process Thinking System for Vision-and-Language NavigationYu Zhong, Zihao Zhang, Rui Zhang, Lingdong Huang 等AAAI 2026
- ProFocus: Proactive Perception and Focused Reasoning in Vision-and-Language NavigationWei Xue, Mingcheng Li, Xuecheng Wu, Jingqun Tang 等CVPR 2026 · 被引用 4 次
- VLN-MME: Diagnosing MLLMs as Language-guided Visual Navigation AgentsXunyi Zhao, Gengze Zhou, Qi WuACL 2026 · 被引用 3 次
- UNeMo: Collaborative Visual-Language Reasoning and Navigation via a Multimodal World ModelChangxin Huang, Lv Tang, Zhaohuan Zhan, Lisha Yu 等AAAI 2026 · 被引用 2 次
- Not All Inconsistency Is Equal: Decomposing LVLM Uncertainty into Belief Divergence and Belief ConflictJie Shi, Xiaodong Yue, Wei Liu, Yufei Chen 等AAAI 2026 · 被引用 1 次
