Symbiotic MLLM Serving: Dynamically Balancing Parallelism Across GPUs and Resources Within GPUs
Zhicheng Li, Jiacheng Zhao, Yangyu Zhang, Zhaolin Duan, Xinyu Liu, Siqi Li, Shuoming Zhang, Shuaijiang Li, Donglin Yu, Yuan Wen, Chunwei Xia, Xiyu Shi, Huimin Cui
Abstract
Multimodal Large Language Models (MLLMs) present a key serving challenge. Existing systems usually extend text-only LLM stacks. They either co-locate the vision encoder with the decoder on the same GPUs or move it to a fixed GPU pool. Both choices rely on ahead-of-time, static decisions for parallelism and GPU allocation. They ignore a key asymmetry in MLLMs: the encoder has a small memory footprint but inputdependent, often quadratic, compute cost, while the decoder is both memory-and compute-intensive. With mixed highand low-resolution requests, this leads to interference between encoder and decoder, persistent SM and HBM slack, poor GPU utilization, and frequent SLO violations. In this paper, we propose Resonator, an efficient MLLM serving system that uses fine-grained runtime scheduling along two axes. First, an intra-GPU sharing engine manages SM and HBM sharing between encoder and decoder. It exploits stagelevel complementarity and kernel-level slack to run computebound encoder kernels and memory-bound LLM kernels in a work-conserving way on the same GPU. Second, an interGPU parallelism engine selects the encoder's data-parallel and tensor-parallel plan at runtime based on input resolution, batch composition, and a lightweight Performance Atlas, and switches plans with near-zero overhead. Evaluation on three state-of-theart MLLMs against strong baselines shows that Resonator improves mean TTFT by up to , TPOT by up to , mean end-to-end latency by up to , and overall throughput by up to , while adapting efficiently to highly concurrent and dynamic workloads.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Related papers
- SpaceServe: Spatial Multiplexing of Complementary Encoders and Decoders for Multimodal LLMsZhicheng Li, Shuoming Zhang, Jiacheng Zhao, Siqi Li et al.NeurIPS 2025 · 5 citations
- ElasticMM: Efficient Multimodal LLMs Serving with Elastic Multimodal ParallelismZedong Liu, Shenggan Cheng, Guangming Tan, Yang You et al.NeurIPS 2025 · 12 citations
- MegaScale-Omni: A Hyper-Scale, Workload-Resilient System for MultiModal LLM Training in ProductionChunyu Xue, Yangrui Chen, Jianyu Jiang, Ningxin Zheng et al.EuroSys 2026
- Optimus: Accelerating Large-Scale Multi-Modal LLM Training by Bubble ExploitationWeiqi Feng, Yangrui Chen, Shaoyu Wang, Yanghua Peng et al.USENIX ATC 2025 · 34 citations
- Accelerating Multi-modal LLM Training with Adaptive Model Placement and ParallelizationYiming Yin, Shaohuai Shi, Qiang Wang, Xiaowen ChuINFOCOM 2026
