Efficient Multimodal Serving via Module Multiplexing
Zicong Hong, Yuyan Chen, Haoyue Zhang, Peng Li, Wuhui Chen, Song Guo, Xiaowei Shen
Abstract
Multimodal learning enables models to process and reason over diverse information sources, unlocking human-like perceptual and cognitive capabilities. As such models gain adoption, efficiently serving them on GPUs has become increasingly important. However, the modular architecture of multimodal models poses significant challenges to existing unimodal serving systems, which treat models as monolithic and overlook inter-module heterogeneity. This results in severe GPU underutilization. To address this, we propose Eevee, a multimodal serving system based on a new scheduling paradigm we call module multiplexing. Unlike prior approaches that execute all modules sequentially with uniform batch sizes, Eevee schedules modality-specific modules concurrently on the same GPU with independently tuned batching and resource allocation. This design enables fine-grained GPU sharing, boosting intra-GPU parallelism and improving request-level throughput. We implement a prototype of Eevee and evaluate it on several representative multimodal models (e.g., CLIP, BLIP, LLaVA, InternVL). Our results show that Eevee significantly outperforms state-of-the-art serving systems in both throughput and GPU utilization.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Related papers
- ElasticMM: Efficient Multimodal LLMs Serving with Elastic Multimodal ParallelismZedong Liu, Shenggan Cheng, Guangming Tan, Yang You et al.NeurIPS 2025 · 12 citations
- SpaceServe: Spatial Multiplexing of Complementary Encoders and Decoders for Multimodal LLMsZhicheng Li, Shuoming Zhang, Jiacheng Zhao, Siqi Li et al.NeurIPS 2025 · 5 citations
- Towards High-Goodput LLM Serving with Prefill-decode MultiplexingYukang Chen, Weihao Cui, Han Zhao, Ziyi Xu et al.ASPLOS 2026 · 18 citations
- OmniScale: Scaling Any Modality Model Training with Model-Centric Distributed Recipe ZooQianli Ma, Yaowei Zheng, Zhelun Shi, Zhongkai Zhao et al.AAAI 2026
- Symbiotic MLLM Serving: Dynamically Balancing Parallelism Across GPUs and Resources Within GPUsZhicheng Li, Jiacheng Zhao, Yangyu Zhang, Zhaolin Duan et al.ISCA 2026
