Lune

EuroSys2026顶会

Efficient Multimodal Serving via Module Multiplexing

Zicong Hong, Yuyan Chen, Haoyue Zhang, Peng Li, Wuhui Chen, Song Guo, Xiaowei Shen

2026年份

摘要

Multimodal learning enables models to process and reason over diverse information sources, unlocking human-like perceptual and cognitive capabilities. As such models gain adoption, efficiently serving them on GPUs has become increasingly important. However, the modular architecture of multimodal models poses significant challenges to existing unimodal serving systems, which treat models as monolithic and overlook inter-module heterogeneity. This results in severe GPU underutilization. To address this, we propose Eevee, a multimodal serving system based on a new scheduling paradigm we call module multiplexing. Unlike prior approaches that execute all modules sequentially with uniform batch sizes, Eevee schedules modality-specific modules concurrently on the same GPU with independently tuned batching and resource allocation. This design enables fine-grained GPU sharing, boosting intra-GPU parallelism and improving request-level throughput. We implement a prototype of Eevee and evaluate it on several representative multimodal models (e.g., CLIP, BLIP, LLaVA, InternVL). Our results show that Eevee significantly outperforms state-of-the-art serving systems in both throughput and GPU utilization.

问问这篇 Paper

问问你的智能体。

Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。

可以从这些问题问起

智能体调用

Lunesearch_papers

在 Lune 里问

免费开始,无需绑卡

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖