EXION: Exploiting Inter-and Intra-Iteration Output Sparsity for Diffusion Models
Jaehoon Heo, Adiwena Putra, Jieon Yoon, Sungwoong Yune, Hangyeol Lee, Ji-Hoon Kim, Joo-Young Kim
Abstract
Over the past few years, diffusion models have emerged as novel solutions in AI industries, offering the capability to generate diverse, multi-modal outputs such as images, videos, and motions from input prompts (i.e., text). Despite their impressive capabilities, diffusion models face significant challenges in computing, including excessive latency and energy consumption due to the numerous iterations in model architecture. Although prior works specialized in transformer acceleration can be applied to diffusion models, given that transformers are key components, the problem of the iterative nature of diffusion models still needs to be addressed.In this paper, we present EXION, the first software-hardware co-designed diffusion accelerator that solves the computation challenges of excessive iterations by exploiting the unique inter-and intra-iteration output sparsity in diffusion models. To this end, we propose two software-level optimizations in EXION. First, we introduce the FFN-Reuse algorithm that identifies and skips redundant computations in FFN layers across different iterations (i.e., inter-iteration sparsity). Second, we use a modified eager prediction method that employs two-step leading-one detection to accurately predict the attention score in diffusion models, skipping unnecessary computations within an iteration (i.e., intra-iteration sparsity). We also introduce a novel data compaction mechanism named ConMerge, which can enhance hardware utilization by condensing and merging large and sparse matrices into small and compact forms. Finally, EXION has a dedicated hardware architecture that supports the above sparsity-inducing algorithms, translating high output sparsity into improved energy efficiency and performance. To verify the feasibility of the EXION accelerator, we first demonstrate that it has no impact on accuracy in various types of multi-modal diffusion models, including text-to-motion, -audio, -image, and -video. We then instantiate EXION in both server-and edge-level settings and compare its performance against GPUs with similar specifications. Our evaluation shows that EXION achieves dramatic improvements in performance and energy efficiency by 3.2-379.3 × and 45.1-3067.6 × compared to a server GPU and by 42.6-1090.9 × and 196.9-4668.2 × compared to an edge GPU.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 04b6d994-7c7d-4c15-ac77-d1c8fa1d823aCited by top-tier papers3
- ToProVAR: Efficient Visual Autoregressive Modeling via Tri-Dimensional Entropy-Aware Semantic Analysis and Sparsity OptimizationJiayu Chen, Ruoyu Lin, Zihao Zheng, Jingxin Li et al.ICLR 2026 · 5 citations
- VisRes Bench: On Evaluating the Visual Reasoning Capabilities of VLMsBrigitta Malagurski Törtei, Yasser Dahou, Ngoc Dung Huynh, Wamiq Reyaz Para et al.CVPR 2026 · 3 citations
- V-Rex: Real-Time Streaming Video LLM Acceleration via Dynamic KV Cache RetrievalDonghyuk Kim, Sejeong Yang, Wonjin Shin, Joo-Young KimHPCA 2026
Builds on27
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
- Consistency ModelsYang Song, Prafulla Dhariwal, Mark Chen, Ilya SutskeverICML 2023 · 1,720 citations
Related papers
- RADiT: Redundancy-Aware Diffusion Transformer Acceleration Leveraging Timestep SimilarityYoungjun Park, Sangyeon Kim, Yeonggeon Kim, Gisan Ji et al.DAC 2025 · 1 citation
- DiTPA: A DiT-Based Action Planner Accelerator Exploiting Action-Denoising-Multimodality Redundancy for Embodied Artificial IntelligenceXin Zhao, Longke Yan, Jiancong Li, Yongkun Wu et al.ISCA 2026
- AIG-CIM: A Scalable Chiplet Module with Tri-Gear Heterogeneous Compute-in-Memory for Diffusion AccelerationYiqi Jing, Meng Wu, Jiaqi Zhou, Yiyang Sun et al.DAC 2024 · 7 citations
- Sparse-vDiT: Unleashing the Power of Sparse Attention to Accelerate Video Diffusion TransformersPengtao Chen, Xianfang Zeng, Maosen Zhao, Mingzhu Shen et al.AAAI 2026
- SLA: Beyond Sparsity in Diffusion Transformers via Fine-Tunable Sparse–Linear AttentionJintao Zhang, Haoxu Wang, Kai Jiang, Shuo Yang et al.ICLR 2026 · 57 citations
