Lune

NeurIPS2025顶会

SpaceServe: Spatial Multiplexing of Complementary Encoders and Decoders for Multimodal LLMs

Zhicheng Li, Shuoming Zhang, Jiacheng Zhao, Siqi Li, Xiyu Shi, Yangyu Zhang, Shuaijiang Li, Donglin Yu, Zheming Yang, Yuan Wen, Huimin Cui

2025年份
5被引次数

摘要

Recent multimodal large language models (MLLMs) marry modality-specific vision or audio encoders with a shared text decoder . While the encoder is compute-intensive but memory-light, the decoder is the opposite, yet state-of-the-art serving stacks still time-multiplex these complementary kernels, idling SMs or HBM in turn. We introduce SpaceServe , a serving system that space-multiplexes MLLMs: it decouples all modality encoders from the decoder, and co-locates them on the same GPU using fine-grained SM partitioning available in modern run-times. A cost-model-guided Space-Inference Scheduler (SIS) dynamically assigns SM slices, while a Time-Windowed Shortest-Remaining-First (TWSRFT) policy batches encoder requests to minimise completion latency and smooth decoder arrivals. Evaluation shows that SpaceServe reduces time-per-output-token by 4.81 × on average and up to 28.9 × on Nvidia A100 GPUs. SpaceServe is available at https://github.com/gofreelee/SpaceServe

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

lune papers fulltext 3dc426ed-e281-414a-86c4-9b784b11dba8

它引用的顶会 Paper18

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖