MHDiff: Memory- and Hardware-Efficient Diffusion Acceleration via Focal Pixel Aware Quantization
Chunyu Qi, Xuhang Wang, Ruiyang Chen, Yuanzheng Yao, Naifeng Jing, Chen Zhang, Jun Wang, Zhihui Fu, Xiaoyao Liang, Zhuoran Song
摘要
Diffusion models have demonstrated superior performance in image generation tasks, thus becoming the mainstream model for generative visual tasks. Diffusion models need to execute multiple timesteps sequentially, resulting in a dramatic increase in workload. Existing accelerators leverage the data similarity between adjacent timesteps and perform mixed-precision differential quantization to accelerate diffusion models. However, merging differential values with raw inputs in each layer of each timestep to ensure computational correctness requires significant memory access for loading raw inputs, which creates a heavy memory burden. Moreover, mixed-precision computations may lead to low hardware utilization if not well designed. Unlike these works, we propose MHDiff, a tailored framework that identifies the focal pixels at the first layer and finetunes them to fit all layers, then represents focal pixels with high-precision while using low-precision for others, thereby accelerating diffusion models while minimizing memory burden. To improve hardware utilization, MHDiff employs a packing module that merges low-precision values into high-precision values to create full high-precision matrices and designs a processing element (PE) array to efficiently process the packed matrices. Extensive experiment results demonstrate that MHDiff can achieve satisfactory performance with negligible quality loss.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- SQ-DM: Accelerating Diffusion Models with Aggressive Quantization and Temporal SparsityZichen Fan, Steve Dai, Rangharajan Venkatesan, Dennis Sylvester 等DAC 2025 · 被引用 3 次
- Ditto: Accelerating Diffusion Model via Temporal Value SimilaritySungbin Kim, Hyunwuk Lee, Wonho Cho, Mincheol Park 等HPCA 2025 · 被引用 9 次
- Modulated Diffusion: Accelerating Generative Modeling with Modulated QuantizationWeizhi Gao, Zhichao Hou, Junqi Yin, Feiyi Wang 等ICML 2025
- ALTER: All-in-One Layer Pruning and Temporal Expert Routing for Efficient Diffusion GenerationXiaomeng Yang, Lei Lu, Qihui Fan, Changdi Yang 等NeurIPS 2025 · 被引用 4 次
- Fewer Denoising Steps or Cheaper Per-Step Inference: Towards Compute-Optimal Diffusion Model DeploymentZhenbang Du, Yonggan Fu, Lifu Wang, Jiayi Qian 等ICCV 2025
