DiffFit: Unlocking Transferability of Large Diffusion Models via Simple Parameter-Efficient Fine-Tuning
Enze Xie, Lewei Yao, Han Shi, Zhili Liu, Daquan Zhou, Zhaoqiang Liu, Jiawei Li, Zhenguo Li
摘要
Diffusion models have proven to be highly effective in generating high-quality images. However, adapting large pre-trained diffusion models to new domains remains an open challenge, which is critical for real-world applications. This paper proposes DiffFit, a parameter-efficient strategy to fine-tune large pre-trained diffusion models that enable fast adaptation to new domains. DiffFit is embarrassingly simple that only fine-tunes the bias term and newly-added scaling factors in specific layers, yet resulting in significant training speed-up and reduced model storage costs. Compared with full fine-tuning, DiffFit achieves Correspondence to xie.enze, li.zhenguo@huawei.com 2× training speed-up and only needs to store approximately 0.12% of the total model parameters. Intuitive theoretical analysis has been provided to justify the efficacy of scaling factors on fast adaptation. On 8 downstream datasets, DiffFit achieves superior or competitive performances compared to the full fine-tuning while being more efficient. Remarkably, we show that DiffFit can adapt a pre-trained low-resolution generative model to a high-resolution one by adding minimal cost. Among diffusion-based methods, DiffFit sets a new state-of-the-art FID of 3.02 on ImageNet 512×512 benchmark by fine-tuning only 25 epochs from a public pre-trained ImageNet 256×256 checkpoint while being 30× more training efficient than the closest competitor.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper46
- PixArt-α: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image SynthesisJunsong Chen, Jincheng Yu, Chongjian Ge, Lewei Yao 等ICLR 2024 · 被引用 831 次
- Rewarded soups: towards Pareto-optimal alignment by interpolating weights fine-tuned on diverse rewardsAlexandre Ramé, Guillaume Couairon, Corentin Dancette, Jean-Baptiste Gaya 等NeurIPS 2023 · 被引用 295 次
- DiT-3D: Exploring Plain Diffusion Transformers for 3D Shape GenerationShentong Mo, Enze Xie, Ruihang Chu, Lanqing Hong 等NeurIPS 2023 · 被引用 157 次
- ScaleCrafter: Tuning-free Higher-Resolution Visual Generation with Diffusion ModelsYingqing He, Shaoshu Yang, Haoxin Chen, Xiaodong Cun 等ICLR 2024 · 被引用 125 次
- Latent Space Editing in Transformer-Based Flow MatchingVincent Tao Hu, Wei Zhang, Meng Tang, Pascal Mettes 等AAAI 2024 · 被引用 42 次
它引用的顶会 Paper41
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 被引用 13,211 次
相关 Paper
- DEFT: Decompositional Efficient Fine-Tuning for Text-to-Image ModelsKomal Kumar, Rao Muhammad Anwer, Fahad Shahbaz Khan, Salman H. Khan 等NeurIPS 2025 · 被引用 2 次
- DogFit: Domain-guided Fine-tuning for Efficient Transfer Learning of Diffusion ModelsYara Bahram, Mohammadhadi Shateri, Eric GrangerAAAI 2026 · 被引用 4 次
- SPRINT: Sparse-Dense Residual Fusion for Efficient Diffusion TransformersDogyun Park, Moayed Haji-Ali, Yanyu Li, Willi Menapace 等ICLR 2026 · 被引用 6 次
- Flash Diffusion: Accelerating Any Conditional Diffusion Model for Few Steps Image GenerationClément Chadebec, Onur Tasar, Eyal Benaroche, Benjamin AubinAAAI 2025 · 被引用 52 次
- IntLoRA: Integral Low-rank Adaptation of Quantized Diffusion ModelsHang Guo, Yawei Li, Tao Dai, Shu-Tao Xia 等ICML 2025
