Lune

INFOCOM2026Top-tier venue

DEMUS: Large Multimodal Model Serving at the Edge with Diffusion-Based Scheduling

Han Zhang, Xiangkai Ma, Tiantian Wang, Mingkai Lin, Wenzhong Li

2026Year

Abstract

Deploying large multimodal models (LMMs) on the edge is a promising solution for latency-sensitive applications, while also addressing the high bandwidth costs and privacy concerns associated with cloud-based serving. LMM inference involves a compute-intensive context stage and a memory-intensive generation stage. However, existing LMM serving systems typically colocate both stages within a single instance, which precludes independent optimization and leads to inefficient resource utilization and increased latency, especially in resource-constrained edge settings. In this paper, we present DEMUS, a Decoupled edge serving framework for large Multimodal models with diffusion-based Scheduling. DEMUS decouples the two stages to optimize them independently without interference, and schedules context and generation tasks to corresponding instances. It combines diffusion-based generation with model predictive control to effectively schedule tasks across instances on heterogeneous nodes, aiming to maximize overall system throughput under latency constraints. We implement and evaluate DEMUS on a heterogeneous cluster with 16 edge server nodes and 45 GPUs. Extensive experiments shows that DEMUS achieves 2.1 3.7× higher throughput while meeting latency SLOs, and generalizes well in dynamic environments and unseen settings.

Ask about this paper

Ask your agent about it.

Lune has read the top-tier papers around this one, so every answer names the papers it rests on.

Questions to start from

Your agent calls

Lunesearch_papers

Ask in Lune

Free to start. No credit card required.

lune papers get 29a58ffe-adcc-4903-95b5-36b908826d47

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines