Cambricon-D: Full-Network Differential Acceleration for Diffusion Models
Weihao Kong, Yifan Hao, Qi Guo, Yongwei Zhao, Xinkai Song, Xiaqing Li, Mo Zou, Zidong Du, Rui Zhang, Chang Liu, Yuanbo Wen, Pengwei Jin
Abstract
Diffusion models have made significant progress in current image generation tasks, thus becoming a prominent area of research. Diffusion models necessitate repetitive iterations on minimally altered input data across timesteps, each timestep requiring the recalculation of the entire model, resulting in a remarkable computational redundancy and substantial hardware expenditures.Performing differential computing on input data seems to be a feasible approach for addressing such computational redundancy and improving hardware efficacy. However, non-linear operations (particularly activation functions) necessitate the merging of deltas (i.e., differential values) with raw inputs repeatedly to ensure computational correctness, leading to significant memory access for loading raw inputs, which fragmentedly blocks the forwarding of deltas throughout the network and undermines performance.To solve this problem, we propose Cambricon-D, a fullnetwork differential computing architecture with concise memory access. While maintaining the computational efficiency brought by differential computing, Cambricon-D employs a sign-mask dataflow, which requires only the loading of 1-bit signs (instead of large bitwidth raw inputs), thereby facilitating the seamless forwarding of deltas and effectively mitigating memory access overheads. Experimental results show that, compared to Diffy, Cambricon-D’s dataflow reduces 66% 82% off-chip memory access. In total, Cambricon-D achieves 1.46× 2.38× speedup over A100 on various diffusion models with different resolutions.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 1d2ebeec-e662-4b45-b986-733f311d5991Cited by top-tier papers5
- EXION: Exploiting Inter-and Intra-Iteration Output Sparsity for Diffusion ModelsJaehoon Heo, Adiwena Putra, Jieon Yoon, Sungwoong Yune et al.HPCA 2025 · 9 citations
- Ditto: Accelerating Diffusion Model via Temporal Value SimilaritySungbin Kim, Hyunwuk Lee, Wonho Cho, Mincheol Park et al.HPCA 2025 · 9 citations
- Neo: Real-Time On-Device 3D Gaussian Splatting with Reuse-and-Update Sorting AccelerationChanghun Oh, Seongryong Oh, Jinwoo Hwang, Yoonsung Kim et al.ASPLOS 2026 · 6 citations
- Assassyn: A Unified Abstraction for Architectural Simulation and ImplementationJian Weng, Boyang Han, Derui Gao, Ruijie Gao et al.ISCA 2025 · 1 citation
- ReFrame: Layer Caching for Accelerated Inference in Real-Time RenderingLufei Liu, Tor M. AamodtICML 2025
Related papers
- Clockwork Diffusion: Efficient Generation With Model-Step DistillationAmirhossein Habibian, Amir Ghodrati, Noor Fathima, Guillaume Sautière et al.CVPR 2024
- Cambricon-P: A Bitflow Architecture for Arbitrary Precision ComputingYifan Hao, Yongwei Zhao, Chenxiao Liu, Zidong Du et al.MICRO 2022 · 9 citations
- Accelerating Parallel Diffusion Model Serving with Residual CompressionJiajun Luo, Yicheng Xiao, Jianru Xu, Yangxiu You et al.NeurIPS 2025 · 3 citations
- Cambricon-M: A Fibonacci-Coded Charge-Domain SRAM-Based CIM Accelerator for DNN InferenceHongrui Guo, Mo Zou, Yifan Hao, Zidong Du et al.MICRO 2024 · 3 citations
- MHDiff: Memory- and Hardware-Efficient Diffusion Acceleration via Focal Pixel Aware QuantizationChunyu Qi, Xuhang Wang, Ruiyang Chen, Yuanzheng Yao et al.DAC 2025
