D i F low : A System for Micro-Serving Text-to-image Di ffusion Work flows
Lingyun Yang, Suyi Li, Tianyu Feng, Xiaoxiao Jiang, Zhipeng Di, Weiyi Lu, Kan Liu, Yinghao Yu, Tao Lan, Guodong Yang, Lin Qu, Liping Zhang, Wei Wang
摘要
Text-to-image generation executes a diffusion workflow comprising multiple models centered on a base diffusion model. Existing serving systems treat each workflow as an opaque monolith, provisioning, placing, and scaling all constituent models together, which obscures internal dataflow, prevents model sharing, and enforces coarse-grained resource management. In this paper, we make a case for micro-serving diffusion workflows with DiFlow, a system that decomposes a workflow into loosely coupled model-execution nodes that can be independently managed and scheduled. By explicitly managing individual model inference, DiFlow unlocks cluster-scale optimizations, including per-model scaling, model sharing, and adaptive model parallelism. Collectively, DiFlow outperforms existing diffusion workflow serving systems, sustaining up to 3× higher request rates and tolerating up to 8× higher burst traffic. We have open-sourced DiFlow at https://github.com/diflow-project/diflow.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- ChituDiffusion: A Data-Characteristic-Aware Serving System for Diffusion ModelsChengzhang Wu, Liyan Zheng, Haojie Wang, Kezhao Huang 等PPoPP 2026
- MixFusion: A Patch-Level Parallel Serving System for Mixed-Resolution Diffusion ModelsDesen Sun, Zepeng Zhao, Yuke WangPPoPP 2026 · 被引用 1 次
- Katz: Efficient Workflow Serving for Diffusion Models with Many AdaptersSuyi Li, Lingyun Yang, Xiaoxiao Jiang, Hanfeng Lu 等USENIX ATC 2025 · 被引用 14 次
- MoDM: Efficient Serving for Image Generation via Mixture-of-Diffusion ModelsYuchen Xia, Divyam Sharma, Yichao Yuan, Souvik Kundu 等ASPLOS 2026
- NeuStream: Bridging Deep Learning Serving and Stream ProcessingHaochen Yuan, Yuanqing Wang, Wenhao Xie, Yu Cheng 等EuroSys 2025 · 被引用 1 次
