Can VLMs Diagnose and Recover from VLA Manipulation Faults?
Bowen Yan, Jiahao Xiao, Kehui Liu, Jianbo Zhang, Zicheng Zhang, Qi Jia, Zhongjie Jia, Haoming Song, Chunyi Li, Bin Zhao, Guangtao Zhai
摘要
Existing VLA models frequently fail in robotic manipulation tasks, with poorly structured fault types that often require expert diagnosis. While VLMs offer strong explanatory capabilities, their effectiveness in assisting VLAs is limited by their unclear role in diagnostics and inadequate collaboration mechanisms. To address this, we introduce VLA-FixBench, a fault evaluation dataset that spans perception, planning, and control failures, and provides annotations for task stages, fault types, and spatiotemporal repair strategies. We further propose FaultEval, a static-to-dynamic-to-real evaluation framework that benchmarks 20 VLMs across multiple fault-related dimensions. Building on these insights, we design a VLM–VLA collaboration mechanism that localizes spatiotemporal deviations and rolls back task execution to enable targeted recovery. Experiments show that FaultEval reliably characterizes VLM-based closed-loop diagnosis and repair. The upper-bound analysis using human expert intervention shows that an idealized feedback loop can improve task success rates by 13% on LIBERO and 35% on real-world robots. Our code, benchmark, and project page will be publicly released at: https://kakigo.github.io/VLA-FixBench/
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper2
- AHA: A Vision-Language-Model for Detecting and Reasoning Over Failures in Robotic ManipulationJiafei Duan, Wilbert Pumacay, Nishanth Kumar, Yi Ru Wang 等ICLR 2025 · 被引用 4 次
- mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language ModelsJiabo Ye, Haiyang Xu, Haowei Liu, Anwen Hu 等ICLR 2025
相关 Paper
- Diagnose, Correct, and Learn from Manipulation Failures via Visual SymbolsXianchao Zeng, Xinyu Zhou, Youcheng Li, Jiayou Shi 等CVPR 2026 · 被引用 18 次
- LIBERO-Plus: A Progressive Robustness Benchmark for Visual-Language-Action ModelsSenyu Fei, Siyin Wang, Junhao Shi, Zihao Dai 等CVPR 2026
- RoboFailRing: Retrieval-Augmented and Language Grounding Failure Detection for VLM-enabled Robotic ManipulationChenduo Ying, Linkang Du, Yuanchao Shu, Peng ChengACL 2026
- FLARE: A Failure-Aware Framework for Autonomous Correction and Recovery in Visual-Language Robotic ManipulationGanlong Zhao, Zijia Tang, Xingping Chen, Zhanghui Kuang 等CVPR 2026 · 被引用 10 次
- INSIGHT Bench: Towards Grounded IN-SItu Guidance for Robotic ManipulaTionSeonho Kim, Junhyeong Hong, Kyungjae Lee, Yoonseon OhCVPR 2026
