Exploring the Adversarial Vulnerabilities of Vision-Language-Action Models in Robotics
Taowen Wang, Cheng Han, James Liang, Wenhao Yang, Dongfang Liu, Luna Xinyu Zhang, Qifan Wang, Jiebo Luo, Ruixiang Tang
摘要
Recently in robotics, Vision-Language-Action (VLA) models have emerged as a transformative approach, enabling robots to execute complex tasks by integrating visual and linguistic inputs within an end-to-end learning framework. Despite their significant capabilities, VLA models introduce new attack surfaces. This paper systematically evaluates their robustness. Recognizing the unique demands of robotic execution, our attack objectives target the inherent spatial and functional characteristics of robotic systems. In particular, we introduce two untargeted attack objectives that leverage spatial foundations to destabilize robotic actions, and a targeted attack objective that manipulates the robotic trajectory. Additionally, we design an adversarial patch generation approach that places a small, colorful patch within the camera's view, effectively executing the attack in both digital and physical environments. Our evaluation reveals a marked degradation in task success rates, with up to a 100% reduction across a suite of simulated robotic tasks, highlighting critical security gaps in current VLA architectures. By unveiling these vulnerabilities and proposing actionable evaluation metrics, we advance both the understanding and enhancement of safety for VLA-based robotic systems, underscoring the necessity for continuously developing robust defense strategies prior to physical-world deployments 1 .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- BadVLA: Towards Backdoor Attacks on Vision-Language-Action Models via Objective-Decoupled OptimizationXueyang Zhou, Guiyao Tie, Guowen Zhang, Hechang Wang 等NeurIPS 2025 · 被引用 50 次
- When Robots Obey the Patch: Universal Transferable Patch Attacks on Vision-Language-Action ModelsHui Lu, Yi Yu, Yiming Yang, Chenyu Yi 等CVPR 2026 · 被引用 12 次
- Bootstrap Dynamic-Aware 3D Visual Representation for Scalable Robot LearningQiwei Liang, Boyang Cai, Minghao Lai, Sitong Zhuang 等CVPR 2026 · 被引用 6 次
- TRAP: Hijacking VLA CoT-Reasoning via Adversarial PatchesZhengxian Huang, Wenjun Zhu, Haoxuan Qiu, Xiaoyu Ji 等ICML 2026 · 被引用 5 次
- FlowHijack: A Dynamics-Aware Backdoor Attack on Flow-Matching Vision-Language-Action ModelsXinyuan An, Tao Luo, Gengyun Peng, Yaobing Wang 等CVPR 2026 · 被引用 1 次
它引用的顶会 Paper13
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- Accessorize to a Crime: Real and Stealthy Attacks on State-of-the-Art Face RecognitionMahmood Sharif, Sruti Bhagavatula, Lujo Bauer, Michael K. ReiterCCS 2016 · 被引用 1,765 次
- Feature Squeezing: Detecting Adversarial Examples in Deep Neural NetworksWeilin Xu, David Evans, Yanjun QiNDSS 2018 · 被引用 1,633 次
- On Adaptive Attacks to Adversarial Example DefensesFlorian Tramèr, Nicholas Carlini, Wieland Brendel, Aleksander MadryNeurIPS 2020 · 被引用 1,026 次
- Vision-Language Foundation Models as Effective Robot ImitatorsXinghang Li, Minghuan Liu, Hanbo Zhang, Cunjun Yu 等ICLR 2024 · 被引用 375 次
相关 Paper
- Phantom Menace: Exploring and Enhancing the Robustness of VLA Models Against Physical Sensor AttacksXuancun Lu, Jiaxiang Chen, Shilin Xiao, Zizhi Jin 等AAAI 2026
- Breaking Cross-modal Alignment in Embodied Intelligence: A Multimodal Adversarial Attack Framework for Vision-Language-Action ModelsZhihui Zhao, Xiaorong Dong, Yaowen Zheng, Xiaohui Chen 等WWW 2026
- Spatial-Spectral Homogeneous Attacks on Physical-World Large Vision-Language ModelsDaizong Liu, Baoquan Chen, Wei HuAAAI 2026
- Adversary is on the Road: Attacks on Visual SLAM using Unnoticeable Adversarial PatchBaodong Chen, Wei Wang, Pascal Sikorski, Ting ZhuUSENIX Security 2024 · 被引用 8 次
- LIBERO-Plus: A Progressive Robustness Benchmark for Visual-Language-Action ModelsSenyu Fei, Siyin Wang, Junhao Shi, Zihao Dai 等CVPR 2026
