Decouple-Then-Merge: Finetune Diffusion Models as Multi-Task Learning
Qianli Ma, Xuefei Ning, Dongrui Liu, Li Niu, Linfeng Zhang
Abstract
Diffusion models are trained by learning a sequence of models that reverse each step of noise corruption. Typically, the model parameters are fully shared across multiple timesteps to enhance training efficiency. However, since the denoising tasks differ at each timestep, the gradients computed at different timesteps may conflict, potentially degrading the overall performance of image generation. To solve this issue, this work proposes a Decouple-then-Merge (DeMe) framework, which begins with a pretrained model and finetunes separate models tailored to specific timesteps. We introduce several improved techniques during the finetuning stage to promote effective knowledge sharing while minimizing training interference across timesteps. Finally, after finetuning, these separate models can be merged into a single model in the parameter space, ensuring efficient and practical inference. Experimental results show significant generation quality improvements upon 6 benchmarks including Stable Diffusion on COCO30K, ImageNet1K, PartiPrompts, and DDPM on LSUN Church, LSUN Bedroom, and CIFAR10. Code is included in the supplementary material and will be released on Github.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 04dbba40-5fe1-4525-9ad9-e30f73aca0b6Cited by top-tier papers4
- Group Editing: Edit Multiple Images in One GoYue Ma, Xinyu Wang, Qianli Ma, Qinghe Wang et al.CVPR 2026 · 15 citations
- Paper2Rebuttal: A Multi-Agent Framework for Transparent Author Response AssistanceQianli Ma, Chang Guo, Zhiheng Tian, Siyu Wang et al.ACL 2026 · 6 citations
- A Unified Density Operator View of Flow Control and MergingRiccardo De Santi, Malte Franke, Ya-Ping Hsieh, Andreas KrauseICML 2026 · 2 citations
- NaRA: Noise-Aware LoRA for Parameter-Efficient Fine-Tuning of Diffusion LLMsShuaidi Wang, Zhan Zhuang, HUANG Ruping, Yu ZhangICML 2026 · 1 citation
Builds on30
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Directly Denoising Diffusion ModelsDan Zhang, Jingjing Wang, Feng LuoICML 2024 · 11,724 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
Related papers
- Divide-and-Denoise: A Game-Theoretic Method for Fairly Composing Diffusion ModelsAbhi Gupta, Polina Barabanshchikova, Vikas Garg, Samuel Kaski et al.ICML 2026
- Improving Training Efficiency of Diffusion Models via Multi-Stage Framework and Tailored Multi-Decoder ArchitectureHuijie Zhang, Yifu Lu, Ismail Alkhouri, Saiprasad Ravishankar et al.CVPR 2024 · 10 citations
- DEFT: Decompositional Efficient Fine-Tuning for Text-to-Image ModelsKomal Kumar, Rao Muhammad Anwer, Fahad Shahbaz Khan, Salman H. Khan et al.NeurIPS 2025 · 2 citations
- DDT: Decoupled Diffusion TransformerShuai Wang, Zhi Tian, Weilin Huang, Limin WangCVPR 2026 · 102 citations
- Resolving Multi-Condition Confusion for Finetuning-Free Personalized Image GenerationQihan Huang, Siming Fu, Jinlong Liu, Hao Jiang et al.AAAI 2025 · 43 citations
