SP-VLA: A Joint Model Scheduling and Token Pruning Approach for VLA Model Acceleration
Ye Li, Yuan Meng, Zewen Sun, Kangye Ji, Chen Tang, Jiajun Fan, Xinzhu Ma, Shu-Tao Xia, Zhi Wang, Wenwu Zhu
摘要
Vision-Language-Action (VLA) models have attracted increasing attention for their strong control capabilities. However, their high computational cost and low execution frequency hinder their suitability for real-time tasks such as robotic manipulation and autonomous navigation. Existing VLA acceleration methods primarily focus on structural optimization, overlooking the fact that these models operate in sequential decision-making environments. As a result, temporal redundancy in sequential action generation and spatial redundancy in visual input remain unaddressed. To this end, we propose SP-VLA, a unified framework that accelerates VLA models by jointly scheduling models and pruning tokens. Specifically, we design an action-aware model scheduling mechanism that reduces temporal redundancy by dynamically switching between VLA model and a lightweight generator. Inspired by the human motion pattern of focusing on key decision points while relying on intuition for other actions, we categorize VLA actions into deliberative and intuitive, assigning the former to the VLA model and the latter to the lightweight generator, enabling frequency-adaptive execution through collaborative model scheduling. To address spatial redundancy, we further develop a spatio-semantic dual-aware token pruning method. Tokens are classified into spatial and semantic types and pruned based on their dual-aware importance to accelerate VLA inference. These two mechanisms work jointly to guide the VLA in focusing on critical actions and salient visual information, achieving effective acceleration while maintaining high accuracy. Extensive experiments show that our method achieves 1.5 lossless acceleration in LIBERO and 2.4 in SimplerEnv, with up to 6% average performance gain. Inference frequency and latency improve by 2.2 in SimplerEnv and 1.4 in LIBERO.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- SpecPrune-VLA: Accelerating Vision-Language-Action Models via Action-Aware Self-Speculative PruningHanzhen Wang, Jiaming Xu, Yushun Xiang, Jiayi Pan 等ICML 2026 · 被引用 32 次
- Action-aware Dynamic Pruning for Efficient Vision-Language-Action ManipulationXiaohuan Pei, Yuxing Chen, Siyu Xu, Yunke Wang 等ICLR 2026 · 被引用 31 次
- AVA-VLA: Improving Vision-Language-Action models with Active Visual AttentionLei Xiao, Jifeng Li, Juntao Gao, Feiyang Ye 等CVPR 2026 · 被引用 26 次
- Incentivizing Consistent, Effective and Scalable Reasoning Capability in Audio LLMs via Reasoning Process RewardsJiajun Fan, Roger Ren, Jingyuan Li, Rahul Pandey 等ICLR 2026 · 被引用 15 次
- Sparse ActionGen: Accelerating Diffusion Policy with Real-time PruningKangye Ji, Jianbo Zhou, Yuan Meng, Ye Li 等ICML 2026 · 被引用 4 次
它引用的顶会 Paper10
- Learning-to-Cache: Accelerating Diffusion Transformer via Layer CachingXinyin Ma, Gongfan Fang, Michael Bi Mi, Xinchao WangNeurIPS 2024 · 被引用 167 次
- DeeR-VLA: Dynamic Inference of Multimodal Large Language Models for Efficient Robot ExecutionYang Yue, Yulin Wang, Bingyi Kang, Yizeng Han 等NeurIPS 2024 · 被引用 153 次
- EfficientVLA: Training-Free Acceleration and Compression for Vision-Language-Action ModelsYantai Yang, Yuhao Wang, Zichen Wen, Luo Zhongwei 等NeurIPS 2025 · 被引用 94 次
- LLaVA-Prumerge: Adaptive Token Reduction for Efficient Large Multimodal ModelsYuzhang Shang, Mu Cai, Bingxin Xu, Yong Jae Lee 等ICCV 2025 · 被引用 37 次
- Cache Me if You Can: Accelerating Diffusion Models through Block CachingFelix Wimbauer, Bichen Wu, Edgar Schönfeld, Xiaoliang Dai 等CVPR 2024 · 被引用 25 次
相关 Paper
- VLA-Cache: Efficient Vision-Language-Action Manipulation via Adaptive Token CachingSiyu Xu, Yunke Wang, Chenghao Xia, Dihao Zhu 等NeurIPS 2025 · 被引用 95 次
- SemanticVLA: Semantic-Aligned Sparsification and Enhancement for Efficient Robotic ManipulationWei Li, Renshan Zhang, Rui Shao, Zhijian Fang 等AAAI 2026 · 被引用 13 次
- EcoVLA: Environment-Aware Adaptive Pruning with Interleaved Inference Orchestration for Vision-Language-Action ModelsYuting Huang, Leilei Ding, Zhipeng Tang, Zenghuan Zhu 等ICML 2026 · 被引用 1 次
- CogVLA: Cognition-Aligned Vision-Language-Action Models via Instruction-Driven Routing & SparsificationWei Li, Renshan Zhang, Rui Shao, Jie He 等NeurIPS 2025 · 被引用 87 次
- See What Matters: Differentiable Grid Sample Pruning for Generalizable Vision-Language-Action ModelYixu Feng, Zinan Zhao, Yanxiang Ma, Chenghao Xia 等ICML 2026 · 被引用 6 次
