DEMUS: Large Multimodal Model Serving at the Edge with Diffusion-Based Scheduling
Han Zhang, Xiangkai Ma, Tiantian Wang, Mingkai Lin, Wenzhong Li
Abstract
Deploying large multimodal models (LMMs) on the edge is a promising solution for latency-sensitive applications, while also addressing the high bandwidth costs and privacy concerns associated with cloud-based serving. LMM inference involves a compute-intensive context stage and a memory-intensive generation stage. However, existing LMM serving systems typically colocate both stages within a single instance, which precludes independent optimization and leads to inefficient resource utilization and increased latency, especially in resource-constrained edge settings. In this paper, we present DEMUS, a Decoupled edge serving framework for large Multimodal models with diffusion-based Scheduling. DEMUS decouples the two stages to optimize them independently without interference, and schedules context and generation tasks to corresponding instances. It combines diffusion-based generation with model predictive control to effectively schedule tasks across instances on heterogeneous nodes, aiming to maximize overall system throughput under latency constraints. We implement and evaluate DEMUS on a heterogeneous cluster with 16 edge server nodes and 45 GPUs. Extensive experiments shows that DEMUS achieves 2.1 3.7× higher throughput while meeting latency SLOs, and generalizes well in dynamic environments and unseen settings.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 29a58ffe-adcc-4903-95b5-36b908826d47Related papers
- SpecEdge: Scalable Edge-Assisted Serving Framework for Interactive LLMsJinwoo Park, Seunggeun Cho, Dongsu HanNeurIPS 2025 · 20 citations
- Lemix: Unified Scheduling for Llm Training and Inference on Multi-Gpu SystemsYufei Li, Zexin Li, Yinglun Zhu, Cong LiuRTSS 2025 · 4 citations
- HexGen-2: Disaggregated Generative Inference of LLMs in Heterogeneous EnvironmentYouhe Jiang, Ran Yan, Binhang YuanICLR 2025
- Helix: Serving Large Language Models over Heterogeneous GPUs and Network via Max-FlowYixuan Mei, Yonghao Zhuang, Xupeng Miao, Juncheng Yang et al.ASPLOS 2025 · 33 citations
- Multi-Tier Multi-Node Scheduling of LLM for Collaborative AI ComputingMulei Ma, Chenyu Gong, Liekang Zeng, Yang YangINFOCOM 2025 · 12 citations
