Characterizing Vision-Language-Action Models across XPUs: Constraints and Acceleration for On-Robot Deployment
Kaijun Zhou, Qiwei Chen, Da Peng, Zhiyang Li, Xijun Li, Jinyu Gu
Abstract
Vision-Language-Action (VLA) models are promising for generalist robot control, but on-robot deployment is bottlenecked by real-time inference under tight cost and energy budgets. Most prior evaluations rely on desktop-grade GPUs, obscuring the trade-offs and opportunities offered by heterogeneous edge accelerators (GPUs/XPUs/NPUs). We present a systematic framework for low-cost VLA deployment via model--hardware co-characterization. First, we build a cross-accelerator leaderboard and evaluate model--hardware pairs under CET (Cost, Energy, Time), showing that ``right-sized'' edge devices can be more cost-/energy-efficient than flagship GPUs while meeting control-rate constraints. Second, using fine-grained SM tracing and Roofline analysis, we uncover a consistent two-phase inference pattern: a compute-bound VLM backbone followed by a memory-bound Action Expert, which induces phase-dependent underutilization and hardware inefficiency. Finally, guided by these insights, we propose DP-Cache and V-AEFusion to reduce diffusion redundancy and enable asynchronous pipeline parallelism, achieving up to (2.9) speedup on GPUs and (6) on edge NPUs with only marginal success degradation. The example leaderboard website is: https://vla-leaderboard-01.vercel.app/.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1516cf47-0fa1-45ad-bcff-237e5c378633Builds on13
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra et al.NeurIPS 2022 · 5,493 citations
- Consistency ModelsYang Song, Prafulla Dhariwal, Mark Chen, Ilya SutskeverICML 2023 · 1,720 citations
- Real-Time Execution of Action Chunking Flow PoliciesKevin Black, Manuel Y. Galliker, Sergey LevineNeurIPS 2025 · 280 citations
- DeeR-VLA: Dynamic Inference of Multimodal Large Language Models for Efficient Robot ExecutionYang Yue, Yulin Wang, Bingyi Kang, Yizeng Han et al.NeurIPS 2024 · 153 citations
- VLA-Cache: Efficient Vision-Language-Action Manipulation via Adaptive Token CachingSiyu Xu, Yunke Wang, Chenghao Xia, Dihao Zhu et al.NeurIPS 2025 · 95 citations
Related papers
- Towards Efficient Embodied Reasoning: Mixture-of-Depth Compute Allocation for Vision-Language-Action ModelWeiying Xie, Qingchen Zeng, Zihan Meng, Jiayun Tian et al.KDD 2026
- EfficientVLA: Training-Free Acceleration and Compression for Vision-Language-Action ModelsYantai Yang, Yuhao Wang, Zichen Wen, Luo Zhongwei et al.NeurIPS 2025 · 94 citations
- SP-VLA: A Joint Model Scheduling and Token Pruning Approach for VLA Model AccelerationYe Li, Yuan Meng, Zewen Sun, Kangye Ji et al.ICLR 2026 · 60 citations
- TINY BUT MIGHTY: A SOFTWARE-HARDWARE CO-DESIGN APPROACH FOR EFFICIENT MULTIMODAL INFERENCE ON BATTERY-POWERED SMALL DEVICESYilong Li, Yijing Zeng, Shuai Zhang, Hao Zhang et al.ICLR 2026 · 1 citation
- Flex: Fast, Accurate DNN Inference on Low-Cost Edges Using Heterogeneous Accelerator ExecutionTanmoy Sen, Haiying Shen, Anand Padmanabha IyerEuroSys 2025 · 2 citations
