Upcycling Text-to-Image Diffusion Models for Multi-Task Capabilities
Ruchika Chavhan, Abhinav Mehrotra, Malcolm Chadwick, Alberto Gil Couto Pimentel Ramos, Luca Morreale, Mehdi Noroozi, Sourav Bhattacharya
摘要
Text-to-image synthesis has witnessed remarkable advancements in recent years. Many attempts have been made to adopt text-to-image models to support multiple tasks. However, existing approaches typically require resource-intensive retraining or additional parameters to accommodate for the new tasks, which makes the model inefficient for on-device deployment. We propose Multi-Task Upcycling (MTU), a simple yet effective recipe that extends the capabilities of a pretrained text-to-image diffusion model to support a variety of image-to-image generation tasks. MTU replaces Feed-Forward Network (FFN) layers in the diffusion model with smaller FFNs, referred to as experts, and combines them with a dynamic routing mechanism. To the best of our knowledge, MTU is the first multi-task diffusion modeling approach that seamlessly blends multi-tasking with on-device compatibility, by mitigating the issue of parameter inflation. We show that the performance of MTU is on par with the single-task fine-tuned diffusion models across several tasks including image editing, super-resolution, and inpainting, while maintaining similar latency and computational load (GFLOPs) as the single-task fine-tuned models.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Multi-Task Label Discovery via Hierarchical Task Tokens for Partially Annotated Dense PredictionsJingdong Zhang, Hanrong Ye, Xin Li, Wenping Wang 等ACM MM 2025 · 被引用 1 次
- Semantic-level Backdoor Attack against Text-to-Image Diffusion ModelsTianxin Chen, Wenbo Jiang, Hongqiao Chen, Zhirun Zheng 等ICML 2026 · 被引用 1 次
- MTPano: Multi-Task Panoramic Scene Understanding via Label-Free Integration of Dense Prediction PriorsJingdong Zhang, Xiaohang Zhan, Lingzhi Zhang, Yizhou Wang 等SIGGRAPH 2026
它引用的顶会 Paper29
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 被引用 11,743 次
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 被引用 6,759 次
相关 Paper
- UniVG: A Generalist Diffusion Model for Unified Image Generation and EditingTsu-Jui Fu, Yusu Qian, Chen Chen, Wenze Hu 等ICCV 2025 · 被引用 2 次
- Remix-DiT: Mixing Diffusion Transformers for Multi-Expert DenoisingGongfan Fang, Xinyin Ma, Xinchao WangNeurIPS 2024 · 被引用 7 次
- CoDi: Conditional Diffusion Distillation for Higher-Fidelity and Faster Image GenerationKangfu Mei, Mauricio Delbracio, Hossein Talebi, Zhengzhong Tu 等CVPR 2024 · 被引用 11 次
- One Diffusion to Generate Them AllDuong H. Le, Tuan Pham, Sangho Lee, Christopher Clark 等CVPR 2025
- FunEditor: Achieving Complex Image Edits via Function Aggregation with Diffusion ModelsMohammadreza Samadi, Fred X. Han, Mohammad Salameh, Hao Wu 等AAAI 2025 · 被引用 1 次
