Can VLMs Diagnose and Recover from VLA Manipulation Faults?
Bowen Yan, Jiahao Xiao, Kehui Liu, Jianbo Zhang, Zicheng Zhang, Qi Jia, Zhongjie Jia, Haoming Song, Chunyi Li, Bin Zhao, Guangtao Zhai
Abstract
Existing VLA models frequently fail in robotic manipulation tasks, with poorly structured fault types that often require expert diagnosis. While VLMs offer strong explanatory capabilities, their effectiveness in assisting VLAs is limited by their unclear role in diagnostics and inadequate collaboration mechanisms. To address this, we introduce VLA-FixBench, a fault evaluation dataset that spans perception, planning, and control failures, and provides annotations for task stages, fault types, and spatiotemporal repair strategies. We further propose FaultEval, a static-to-dynamic-to-real evaluation framework that benchmarks 20 VLMs across multiple fault-related dimensions. Building on these insights, we design a VLM–VLA collaboration mechanism that localizes spatiotemporal deviations and rolls back task execution to enable targeted recovery. Experiments show that FaultEval reliably characterizes VLM-based closed-loop diagnosis and repair. The upper-bound analysis using human expert intervention shows that an idealized feedback loop can improve task success rates by 13% on LIBERO and 35% on real-world robots. Our code, benchmark, and project page will be publicly released at: https://kakigo.github.io/VLA-FixBench/
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on2
- AHA: A Vision-Language-Model for Detecting and Reasoning Over Failures in Robotic ManipulationJiafei Duan, Wilbert Pumacay, Nishanth Kumar, Yi Ru Wang et al.ICLR 2025 · 4 citations
- mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language ModelsJiabo Ye, Haiyang Xu, Haowei Liu, Anwen Hu et al.ICLR 2025
Related papers
- Diagnose, Correct, and Learn from Manipulation Failures via Visual SymbolsXianchao Zeng, Xinyu Zhou, Youcheng Li, Jiayou Shi et al.CVPR 2026 · 18 citations
- LIBERO-Plus: A Progressive Robustness Benchmark for Visual-Language-Action ModelsSenyu Fei, Siyin Wang, Junhao Shi, Zihao Dai et al.CVPR 2026
- RoboFailRing: Retrieval-Augmented and Language Grounding Failure Detection for VLM-enabled Robotic ManipulationChenduo Ying, Linkang Du, Yuanchao Shu, Peng ChengACL 2026
- FLARE: A Failure-Aware Framework for Autonomous Correction and Recovery in Visual-Language Robotic ManipulationGanlong Zhao, Zijia Tang, Xingping Chen, Zhanghui Kuang et al.CVPR 2026 · 10 citations
- INSIGHT Bench: Towards Grounded IN-SItu Guidance for Robotic ManipulaTionSeonho Kim, Junhyeong Hong, Kyungjae Lee, Yoonseon OhCVPR 2026
