DiTPA: A DiT-Based Action Planner Accelerator Exploiting Action-Denoising-Multimodality Redundancy for Embodied Artificial Intelligence
Xin Zhao, Longke Yan, Jiancong Li, Yongkun Wu, Fengbin Tu
摘要
Recent advances in multimodal vision-languageaction (VLA) models have endowed embodied artificial intelligence (embodied AI) systems with remarkable perception, reasoning, and planning capabilities. Among these VLA models, diffusion transformers (DiTs) have become the backbone for action planning due to their strong and continuous generation capability. However, multimodal DiT-based action planners typically need to generate hundreds of actions to complete a single task, and each action requires about 10-50 denoising steps. This results in extremely low action frequencies, preventing real-time deployment in embodied AI applications. In this work, we systematically analyze the inference and data distribution characteristics of DiT-based action planners and observe significant computational redundancy across action, denoising and multimodality. Motivated by these findings, we propose DiTPA, a softwarehardware co-designed DiT-based action planner accelerator to fully exploit these three levels of redundancy. We first introduce a DiTPA software framework, consisting of (1) an orientationconditioned action prediction mechanism to reuse actions with minimal orientation variation (action redundancy), (2) an alternating denoising with feature reuse technique that replaces lowimpact iterations with low-cost residual computation for noise updates (denoising redundancy), and (3) a calibrated multimodal approximate computing strategy that eliminates redundant multimodal operations based on modality lifespan and attention sparsity (multimodality redundancy). At the hardware level, the DiTPA accelerator supports this redundancy-aware framework through an action predictor, a reconfigurable processing element (PE) array, and a multimodal scheduler. Together, these innovations convert high computational redundancy into substantial performance and energy efficiency gains. Owing to softwarehardware co-design, DiTPA obtains average action frequency of 217.65Hz and task execution time of 1.73s on the LIBERO-Long benchmark, with only 1.05 W power consumption. It achieves 386.93×, 13.22×, 9.54× speedups and 2356.77 ×, 8.71 ×, 11.59 × energy-efficiency improvements over NVIDIA A40, EXION and Ditto, while maintaining the task success rate. The code is opensourced in https://github.com/fengbintu/ISCA2026-DiTPA.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- EfficientVLA: Training-Free Acceleration and Compression for Vision-Language-Action ModelsYantai Yang, Yuhao Wang, Zichen Wen, Luo Zhongwei 等NeurIPS 2025 · 被引用 94 次
- EXION: Exploiting Inter-and Intra-Iteration Output Sparsity for Diffusion ModelsJaehoon Heo, Adiwena Putra, Jieon Yoon, Sungwoong Yune 等HPCA 2025 · 被引用 9 次
- RADiT: Redundancy-Aware Diffusion Transformer Acceleration Leveraging Timestep SimilarityYoungjun Park, Sangyeon Kim, Yeonggeon Kim, Gisan Ji 等DAC 2025 · 被引用 1 次
- QuantVLA: Scale-Calibrated Post-Training Quantization for Vision-Language-Action ModelsJingxuan Zhang, Yunta Hsieh, Zhongwei Wan, Haokun Lin 等CVPR 2026 · 被引用 24 次
- Dita: Scaling Diffusion Transformer for Generalist Vision-Language-Action PolicyZhi Hou, Tianyi Zhang, Yuwen Xiong, Haonan Duan 等ICCV 2025 · 被引用 9 次
