Clockwork Diffusion: Efficient Generation With Model-Step Distillation
Amirhossein Habibian, Amir Ghodrati, Noor Fathima, Guillaume Sautière, Risheek Garrepalli, Fatih Porikli, Jens Petersen
摘要
This work aims to improve the efficiency of text-to-image diffusion models. While diffusion models use computationally expensive UNet-based denoising operations in every generation step, we identify that not all operations are equally relevant for the final output quality. In particular, we observe that UNet layers operating on highres feature maps are relatively sensitive to small perturbations. In contrast, low-res feature maps influence the semantic layout of the final image and can often be perturbed with no noticeable change in the output. Based on this observation, we propose Clockwork Diffusion, a method that periodically reuses computation from preceding denoising steps to approximate low-res feature maps at one or more subsequent steps. For multiple baselines, and for both text-to-image generation and image editing, we demonstrate that Clockwork leads to comparable or improved perceptual scores with drastically reduced computational complexity. As an example, for Stable Diffusion v1.5 with 8 DPM++ steps we save 32% of FLOPs with negligible FID and CLIP change. We release code at https://github.com/Qualcomm- AI-research/clockwork-diffusion
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- Astraea: A Token-wise Acceleration Framework for Video Diffusion TransformersHaosong Liu, Yuge Cheng, Wenxuan Miao, Zihan Liu 等ICLR 2026 · 被引用 9 次
- Adaptive Caching for Faster Video Generation With Diffusion TransformersKumara Kahatapitiya, Haozhe Liu, Sen He, Ding Liu 等ICCV 2025 · 被引用 5 次
- Importance-Based Token Merging for Efficient Image and Video GenerationHaoyu Wu, Jingyi Xu, Hieu Le, Dimitris SamarasICCV 2025 · 被引用 3 次
- Toward Early Quality Assessment of Text-to-Image Diffusion ModelsHuanlei Guo, Hongxin Wei, Bingyi JingCVPR 2026 · 被引用 2 次
- DICE: Staleness-Centric Optimizations for Parallel Diffusion MoE InferenceJiajun Luo, Lizhuo Luo, Jianru Xu, Jiajun Song 等ICCV 2025 · 被引用 1 次
它引用的顶会 Paper26
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 被引用 13,211 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 被引用 11,743 次
相关 Paper
- SnapFusion: Text-to-Image Diffusion Model on Mobile Devices within Two SecondsYanyu Li, Huan Wang, Qing Jin, Ju Hu 等NeurIPS 2023 · 被引用 300 次
- Faster Diffusion: Rethinking the Role of the Encoder for Diffusion Model InferenceSenmao Li, Taihang Hu, Joost van de Weijer, Fahad Shahbaz Khan 等NeurIPS 2024 · 被引用 50 次
- Schedule On the Fly: Diffusion Time Prediction for Faster and Better Image GenerationZilyu Ye, Zhiyang Chen, Tiancheng Li, Zemin Huang 等CVPR 2025
- Text Embedding Knows How to Quantize Text-Guided Diffusion ModelsHongjae Lee, Myungjun Son, Dongjea Kang, Seung-Won JungICCV 2025 · 被引用 2 次
- Reusing Computation in Text-to-Image Diffusion for Efficient Generation of Image SetsDale Decatur, Thibault Groueix, Wang Yifan, Rana Hanocka 等ICCV 2025
