Sparse ActionGen: Accelerating Diffusion Policy with Real-time Pruning
Kangye Ji, Jianbo Zhou, Yuan Meng, Ye Li, Hanyun Cui, Zhi Wang
Abstract
Diffusion Policy has dominated action generation due to its strong capabilities for modeling multi-modal action distributions, but its multi-step denoising processes make it impractical for real-time visuomotor control. Existing caching-based acceleration methods typically rely on schedules that fail to adapt to the dynamics of robot-environment interactions, thereby leading to suboptimal performance. In this paper, we propose parse ctionen for extremely sparse action generation. To accommodate the iterative interactions, SAG customizes a rollout-adaptive prune-then-reuse mechanism that first identifies prunable computations globally and then reuses cached activations to substitute them during action diffusion. To capture the rollout dynamics, SAG parameterizes an observation-conditioned diffusion pruner for environment-aware adaptation and instantiates it with a highly parameter- and inference-efficient design for real-time prediction. Furthermore, SAG introduces a one-for-all reusing strategy that reuses activations across both timesteps and blocks in a zig-zag manner, minimizing the global redundancy. Extensive experiments on multiple robotic benchmarks demonstrate that SAG achieves up to 4 generation speedup without sacrificing performance. Project Page: https://sparse-actiongen.github.io/.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 17d27021-f2b5-478e-9e27-2d78222ea80fCited by top-tier papers1
Ask how each one uses itBuilds on14
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
- Learning-to-Cache: Accelerating Diffusion Transformer via Layer CachingXinyin Ma, Gongfan Fang, Michael Bi Mi, Xinchao WangNeurIPS 2024 · 167 citations
- EfficientVLA: Training-Free Acceleration and Compression for Vision-Language-Action ModelsYantai Yang, Yuhao Wang, Zichen Wen, Luo Zhongwei et al.NeurIPS 2025 · 94 citations
- DeepCache: Accelerating Diffusion Models for FreeXinyin Ma, Gongfan Fang, Xinchao WangCVPR 2024 · 87 citations
- SP-VLA: A Joint Model Scheduling and Token Pruning Approach for VLA Model AccelerationYe Li, Yuan Meng, Zewen Sun, Kangye Ji et al.ICLR 2026 · 60 citations
Related papers
- Block-wise Adaptive Caching for Accelerating Diffusion PolicyKangye Ji, Yuan Meng, Hanyun Cui, Ye Li et al.ICLR 2026 · 9 citations
- Falcon: Fast Visuomotor Policies via Partial DenoisingHaojun Chen, Minghao Liu, Chengdong Ma, Xiaojian Ma et al.ICML 2025
- VLA-Cache: Efficient Vision-Language-Action Manipulation via Adaptive Token CachingSiyu Xu, Yunke Wang, Chenghao Xia, Dihao Zhu et al.NeurIPS 2025 · 95 citations
- STEP: Warm-Started Visuomotor Policies with Spatiotemporal Consistency PredictionJinhao Li, Yuxuan Cong, Yingqiao Wang, Hao Xia et al.ICML 2026 · 5 citations
- Turbo4DGen: Ultra-Fast Acceleration for 4D GenerationYuanbin Man, Ying Huang, Zhile Ren, Miao YinICML 2026
