Embodied Interpretability: Linking Causal Understanding to Generalization in Vision-Language-Action Models
Hanxin Zhang, Mingshuo Xu, Abdulqader Dhafer, Shigang Yue, Hongbiao Dong, Zhou Daniel Hao
Abstract
Vision–Language–Action (VLA) policies often fail under distribution shift, suggesting that decisions may depend on spurious visual correlations rather than task-relevant causes. We formulate visual–action attribution as an interventional estimation problem. Accordingly, we introduce the Interventional Significance Score (ISS) , an interventional masking procedure for estimating the causal influence of visual regions on action predictions, and the Nuisance Mass Ratio (NMR) , a scalar measure of attribution to task-irrelevant features. We analyze the statistical properties of ISS and show that it admits unbiased estimation, and we characterize conditions under which action prediction error provides a valid proxy for causal influence. Experiments across diverse manipulation tasks indicate that NMR predicts generalization behavior and that ISS yields more faithful explanations than existing interpretability methods. These results suggest that interventional attribution provides a simple diagnostic approach for identifying causal misalignment in embodied policies.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f2e28f7a-9403-459b-aab5-565770fc9d8dBuilds on9
- Sigmoid Loss for Language Image Pre-TrainingXiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, Lucas BeyerICCV 2023 · 2,932 citations
- Exploring the Limits of Vision-Language-Action Manipulation in Cross-task GeneralizationJiaming Zhou, Ke Ye, Jiayi Liu, Teli Ma et al.NeurIPS 2025 · 43 citations
- Towards Reliable Code-as-Policies: A Neuro-Symbolic Framework for Embodied Task PlanningSanghyun Ahn, Wonje Choi, Junyong Lee, Jinwoo Park et al.NeurIPS 2025 · 14 citations
- UP-VLA: A Unified Understanding and Prediction Model for Embodied AgentJianke Zhang, Yanjiang Guo, Yucheng Hu, Xiaoyu Chen et al.ICML 2025
- OTTER: A Vision-Language-Action Model with Text-Aware Visual Feature ExtractionHuang Huang, Fangchen Liu, Letian Fu, Tingfan Wu et al.ICML 2025
Related papers
- Breaking Cross-modal Alignment in Embodied Intelligence: A Multimodal Adversarial Attack Framework for Vision-Language-Action ModelsZhihui Zhao, Xiaorong Dong, Yaowen Zheng, Xiaohui Chen et al.WWW 2026
- Interpretability Transfer from Language to Vision via Sparse AutoencodersAlexey Kravets, Da Li, Chuan Li, Da Chen et al.ICML 2026
- Mechanisms of Object Localization in Vision-Language ModelsTimothy Schaumlöffel, Martina G. Vilas, Gemma RoigCVPR 2026 · 1 citation
- Visual Explanations of Image-Text Representations via Multi-Modal Information Bottleneck AttributionYing Wang, Tim G. J. Rudner, Andrew Gordon WilsonNeurIPS 2023 · 51 citations
- Structural Graph Probing of Vision-Language ModelsHaoyu He, Yue Zhuo, Yu Zheng, Qi R. WangCVPR 2026
