Lune

NeurIPS2025Top-tier venue

SpaceServe: Spatial Multiplexing of Complementary Encoders and Decoders for Multimodal LLMs

Zhicheng Li, Shuoming Zhang, Jiacheng Zhao, Siqi Li, Xiyu Shi, Yangyu Zhang, Shuaijiang Li, Donglin Yu, Zheming Yang, Yuan Wen, Huimin Cui

2025Year
5Citations

Abstract

Recent multimodal large language models (MLLMs) marry modality-specific vision or audio encoders with a shared text decoder . While the encoder is compute-intensive but memory-light, the decoder is the opposite, yet state-of-the-art serving stacks still time-multiplex these complementary kernels, idling SMs or HBM in turn. We introduce SpaceServe , a serving system that space-multiplexes MLLMs: it decouples all modality encoders from the decoder, and co-locates them on the same GPU using fine-grained SM partitioning available in modern run-times. A cost-model-guided Space-Inference Scheduler (SIS) dynamically assigns SM slices, while a Time-Windowed Shortest-Remaining-First (TWSRFT) policy batches encoder requests to minimise completion latency and smooth decoder arrivals. Evaluation shows that SpaceServe reduces time-per-output-token by 4.81 × on average and up to 28.9 × on Nvidia A100 GPUs. SpaceServe is available at https://github.com/gofreelee/SpaceServe

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext 3dc426ed-e281-414a-86c4-9b784b11dba8

Builds on18

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines