DEMUS: Large Multimodal Model Serving at the Edge with Diffusion-Based Scheduling
Han Zhang, Xiangkai Ma, Tiantian Wang, Mingkai Lin, Wenzhong Li
摘要
Deploying large multimodal models (LMMs) on the edge is a promising solution for latency-sensitive applications, while also addressing the high bandwidth costs and privacy concerns associated with cloud-based serving. LMM inference involves a compute-intensive context stage and a memory-intensive generation stage. However, existing LMM serving systems typically colocate both stages within a single instance, which precludes independent optimization and leads to inefficient resource utilization and increased latency, especially in resource-constrained edge settings. In this paper, we present DEMUS, a Decoupled edge serving framework for large Multimodal models with diffusion-based Scheduling. DEMUS decouples the two stages to optimize them independently without interference, and schedules context and generation tasks to corresponding instances. It combines diffusion-based generation with model predictive control to effectively schedule tasks across instances on heterogeneous nodes, aiming to maximize overall system throughput under latency constraints. We implement and evaluate DEMUS on a heterogeneous cluster with 16 edge server nodes and 45 GPUs. Extensive experiments shows that DEMUS achieves 2.1 3.7× higher throughput while meeting latency SLOs, and generalizes well in dynamic environments and unseen settings.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- SpecEdge: Scalable Edge-Assisted Serving Framework for Interactive LLMsJinwoo Park, Seunggeun Cho, Dongsu HanNeurIPS 2025 · 被引用 20 次
- Lemix: Unified Scheduling for Llm Training and Inference on Multi-Gpu SystemsYufei Li, Zexin Li, Yinglun Zhu, Cong LiuRTSS 2025 · 被引用 4 次
- HexGen-2: Disaggregated Generative Inference of LLMs in Heterogeneous EnvironmentYouhe Jiang, Ran Yan, Binhang YuanICLR 2025
- Helix: Serving Large Language Models over Heterogeneous GPUs and Network via Max-FlowYixuan Mei, Yonghao Zhuang, Xupeng Miao, Juncheng Yang 等ASPLOS 2025 · 被引用 33 次
- Multi-Tier Multi-Node Scheduling of LLM for Collaborative AI ComputingMulei Ma, Chenyu Gong, Liekang Zeng, Yang YangINFOCOM 2025 · 被引用 12 次
