DiTPA: A DiT-Based Action Planner Accelerator Exploiting Action-Denoising-Multimodality Redundancy for Embodied Artificial Intelligence
Xin Zhao, Longke Yan, Jiancong Li, Yongkun Wu, Fengbin Tu
Abstract
Recent advances in multimodal vision-languageaction (VLA) models have endowed embodied artificial intelligence (embodied AI) systems with remarkable perception, reasoning, and planning capabilities. Among these VLA models, diffusion transformers (DiTs) have become the backbone for action planning due to their strong and continuous generation capability. However, multimodal DiT-based action planners typically need to generate hundreds of actions to complete a single task, and each action requires about 10-50 denoising steps. This results in extremely low action frequencies, preventing real-time deployment in embodied AI applications. In this work, we systematically analyze the inference and data distribution characteristics of DiT-based action planners and observe significant computational redundancy across action, denoising and multimodality. Motivated by these findings, we propose DiTPA, a softwarehardware co-designed DiT-based action planner accelerator to fully exploit these three levels of redundancy. We first introduce a DiTPA software framework, consisting of (1) an orientationconditioned action prediction mechanism to reuse actions with minimal orientation variation (action redundancy), (2) an alternating denoising with feature reuse technique that replaces lowimpact iterations with low-cost residual computation for noise updates (denoising redundancy), and (3) a calibrated multimodal approximate computing strategy that eliminates redundant multimodal operations based on modality lifespan and attention sparsity (multimodality redundancy). At the hardware level, the DiTPA accelerator supports this redundancy-aware framework through an action predictor, a reconfigurable processing element (PE) array, and a multimodal scheduler. Together, these innovations convert high computational redundancy into substantial performance and energy efficiency gains. Owing to softwarehardware co-design, DiTPA obtains average action frequency of 217.65Hz and task execution time of 1.73s on the LIBERO-Long benchmark, with only 1.05 W power consumption. It achieves 386.93×, 13.22×, 9.54× speedups and 2356.77 ×, 8.71 ×, 11.59 × energy-efficiency improvements over NVIDIA A40, EXION and Ditto, while maintaining the task success rate. The code is opensourced in https://github.com/fengbintu/ISCA2026-DiTPA.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 07260465-ba54-4c75-b21d-d5ebb312eeeaRelated papers
- EfficientVLA: Training-Free Acceleration and Compression for Vision-Language-Action ModelsYantai Yang, Yuhao Wang, Zichen Wen, Luo Zhongwei et al.NeurIPS 2025 · 94 citations
- EXION: Exploiting Inter-and Intra-Iteration Output Sparsity for Diffusion ModelsJaehoon Heo, Adiwena Putra, Jieon Yoon, Sungwoong Yune et al.HPCA 2025 · 9 citations
- RADiT: Redundancy-Aware Diffusion Transformer Acceleration Leveraging Timestep SimilarityYoungjun Park, Sangyeon Kim, Yeonggeon Kim, Gisan Ji et al.DAC 2025 · 1 citation
- QuantVLA: Scale-Calibrated Post-Training Quantization for Vision-Language-Action ModelsJingxuan Zhang, Yunta Hsieh, Zhongwei Wan, Haokun Lin et al.CVPR 2026 · 24 citations
- Dita: Scaling Diffusion Transformer for Generalist Vision-Language-Action PolicyZhi Hou, Tianyi Zhang, Yuwen Xiong, Haonan Duan et al.ICCV 2025 · 9 citations
