VLA-Arena: An Open-Source Framework for Benchmarking Vision-Language-Action Models
Borong Zhang, Jiahao Li, Jiachen Shen, Yuhao Zhang, Yishuai Cai, Yuanpei Chen, Juntao Dai, Jiaming Ji, Yaodong Yang
摘要
While Vision-Language-Action models (VLAs) are rapidly advancing toward generalist robot policies, quantitatively characterizing their capability boundaries and failure modes remains challenging. To address this, we introduce VLA-Arena , a comprehensive benchmark. It features a novel structured task design framework to quantify difficulty across three orthogonal axes: (1) Task Structure , (2) Language Command , and (3) Visual Observation . This allows us to systematically design tasks with fine-grained difficulty levels, enabling a precise measurement of model capability frontiers. For task structure, VLA-Arena comprises 11 task suites organized into four dimensions: Safety , Distractor , Extrapolation , and Long Horizon , totaling 170 tasks. Each suite spans three difficulty levels (L0–L2), with fine-tuning restricted to L0 to rigorously assess generalization. Orthogonal to this, language (W0-W4) and visual (V0-V4) perturbations can be applied to any task as diagnostic probes to distinguish robust grounding from superficial pattern matching. Our extensive evaluation of state-of-the-art VLAs reveals critical limitations: memorization over generalization, superficial visual perception, and a neglect of safety constraints. Additionally, model rank reversals across L0–L2 validate that each level provides non-redundant insights. To foster research addressing these model limitations and ensure reproducibility, we provide the complete VLA-Arena framework, including an end-to-end toolchain from task definition to automated evaluation and the VLA-Arena-S/M/L datasets for fine-tuning. Our benchmark, datasets, models, and leaderboard are publicly available at https://vla-arena.github.io.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- SpecPrune-VLA: Accelerating Vision-Language-Action Models via Action-Aware Self-Speculative PruningHanzhen Wang, Jiaming Xu, Yushun Xiang, Jiayi Pan 等ICML 2026 · 被引用 32 次
- Dismantling the Illusion of Vision-Language-Action Models Competence via Explicit Distributional ShiftsXueyang Zhou, Yangming Xu, Guiyao Tie, Chaoran Hu 等ICML 2026
它引用的顶会 Paper11
- Motus: A Unified Latent Action World ModelHongzhe Bi, Hengkai Tan, Shenghao Xie, Zeyuan Wang 等CVPR 2026 · 被引用 271 次
- OneTwoVLA: A Unified Vision-Language-Action Model with Adaptive ReasoningFanqi Lin, Ruiqian Nai, Yingdong Hu, Jiacheng You 等ICLR 2026 · 被引用 129 次
- What Can RL Bring to VLA Generalization? An Empirical StudyJijia Liu, Feng Gao, Bingwen Wei, Xinlei Chen 等NeurIPS 2025 · 被引用 120 次
- DexGraspVLA: A Vision-Language-Action Framework Towards General Dexterous GraspingYifan Zhong, Xuchuan Huang, Ruochong Li, Ceyao Zhang 等AAAI 2026 · 被引用 89 次
- Discrete Diffusion VLA: Bringing Discrete Diffusion to Action Decoding in Vision-Language-Action PoliciesZhixuan Liang, Yizhuo Li, Tianshuo Yang, CHENGYUE WU 等ICML 2026 · 被引用 86 次
相关 Paper
- LIBERO-Plus: A Progressive Robustness Benchmark for Visual-Language-Action ModelsSenyu Fei, Siyin Wang, Junhao Shi, Zihao Dai 等CVPR 2026
- RobotArena ∞: Scalable Robot Benchmarking via Real-to-Sim TranslationYash Jangir, Yidi Zhang, Kashu Yamazaki, Chenyu Zhang 等ICLR 2026 · 被引用 22 次
- VLANeXt: Recipes for Building Strong VLA ModelsXiao-Ming Wu, Bin Fan, Kang Liao, Jian-Jian Jiang 等ICML 2026 · 被引用 10 次
- VisualWebArena: Evaluating Multimodal Agents on Realistic Visual Web TasksJing Yu Koh, Robert Lo, Lawrence Jang, Vikram Duvvur 等ACL 2024 · 被引用 25 次
- VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning TasksShiduo Zhang, Zhe Xu, Peiju Liu, Xiaopeng Yu 等ICCV 2025 · 被引用 12 次
