VLATest: Testing and Evaluating Vision-Language-Action Models for Robotic Manipulation
Zhijie Wang, Zhehua Zhou, Jiayang Song, Yuheng Huang, Zhan Shu, Lei Ma
Abstract
The rapid advancement of generative AI and multi-modal foundation models has shown significant potential in advancing robotic manipulation. Vision-language-action (VLA) models, in particular, have emerged as a promising approach for visuomotor control by leveraging large-scale vision-language data and robot demonstrations. However, current VLA models are typically evaluated using a limited set of hand-crafted scenes, leaving their general performance and robustness in diverse scenarios largely unexplored. To address this gap, we present VLATest, a fuzzing framework designed to generate robotic manipulation scenes for testing VLA models. Based on VLATest, we conducted an empirical study to assess the performance of seven representative VLA models. Our study results revealed that current VLA models lack the robustness necessary for practical deployment. Additionally, we investigated the impact of various factors, including the number of confounding objects, lighting conditions, camera poses, unseen objects, and task instruction mutations, on the VLA model's performance. Our findings highlight the limitations of existing VLA models, emphasizing the need for further research to develop reliable and trustworthy VLA applications.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5adf352d-a8ef-44be-8a11-785103f4b2caCited by top-tier papers5
- Diagnose, Correct, and Learn from Manipulation Failures via Visual SymbolsXianchao Zeng, Xinyu Zhou, Youcheng Li, Jiayou Shi et al.CVPR 2026 · 18 citations
- Exploring the Adversarial Vulnerabilities of Vision-Language-Action Models in RoboticsTaowen Wang, Cheng Han, James Liang, Wenhao Yang et al.ICCV 2025 · 8 citations
- On Robustness of Vision-Language-Action Model against Multi-Modal PerturbationsJianing Guo, Zhenhong Wu, Chang Tu, Yiyao Ma et al.ICLR 2026 · 7 citations
- LIBERO-Plus: A Progressive Robustness Benchmark for Visual-Language-Action ModelsSenyu Fei, Siyin Wang, Junhao Shi, Zihao Dai et al.CVPR 2026
- Localizing, Structuring, and Rendering: Bridging 3D and 2D Vision-Language-Action Models for Robotic ManipulationYunlong Zhao, Xiaoheng Deng, Yichao Cao, Yi Chen et al.CVPR 2026
Builds on33
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- Tree of Thoughts: Deliberate Problem Solving with Large Language ModelsShunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran et al.NeurIPS 2023 · 5,068 citations
- TruthfulQA: Measuring How Models Mimic Human FalsehoodsStephanie Lin, Jacob Hilton, Owain EvansACL 2022 · 3,228 citations
- Sigmoid Loss for Language Image Pre-TrainingXiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, Lucas BeyerICCV 2023 · 2,932 citations
Related papers
- Vision-Language Foundation Models as Effective Robot ImitatorsXinghang Li, Minghuan Liu, Hanbo Zhang, Cunjun Yu et al.ICLR 2024 · 375 citations
- VideoVLA: Video Generators Can Be Generalizable Robot ManipulatorsYichao Shen, Fangyun Wei, Zhiying Du, Yaobo Liang et al.NeurIPS 2025 · 73 citations
- VLA-Arena: An Open-Source Framework for Benchmarking Vision-Language-Action ModelsBorong Zhang, Jiahao Li, Jiachen Shen, Yuhao Zhang et al.ICML 2026 · 25 citations
- RoboInter: A Holistic Intermediate Representation Suite Towards Robotic ManipulationHao Li, Ziqin Wang, Zi-han Ding, Shuai Yang et al.ICLR 2026 · 17 citations
- VirtualEnv: A Platform for Embodied AI ResearchKabir Swain, Sijie Han, Ayush Raina, Jin Zhang et al.AAAI 2026
