Efficiently Serving Large Multimodal Models Using EPD Disaggregation
Gursimran Singh, Xinglu Wang, Yifan Hu, Timothy Tin Long Yu, Linzi Xing, Wei Jiang, Zhefeng Wang, Xiaolong Bai, Yi Li, Ying Xiong, Yong Zhang, Zhenan Fan
Abstract
Large Multimodal Models (LMMs) extend Large Language Models (LLMs) by handling diverse inputs such as images, audio, and video, but at the cost of adding a multimodal encoding stage that increases both computational and memory overhead. This step negatively affects key Service Level Objectives (SLOs), such as time to first token (TTFT) and time per output token (TPOT). We introduce Encode-Prefill-Decode (EPD) Disaggregation, a novel framework that separates the encoding, prefill, and decode stages onto dedicated resources. Unlike current systems, which bundle encoding and prefill together, our approach decouples these steps, unlocking new opportunities and optimizations. These include a mechanism to cache multimedia tokens for efficient transfer, a novel way to parallelize the encoding load within a request, a module for optimal resource allocation for disaggregated serving, and a novel role-switching method to handle changing workload characteristics. Experimental evaluations with popular LMMs show substantial gains in memory efficiency (up to 15× lower peak memory utilization), batch sizes (up to 22× larger), 10× more images per request, and 2.2× larger KV caches. Furthermore, it leads to significant improvements in SLO attainment (up to 90-100% improvement) and TTFT (up to 71% reduction), compared to systems that do not disaggregate. The code is available at https://github.com/vbdi/epdserve .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4e1b6cc4-aacc-4ee2-b998-6a875f3bb7a2Cited by top-tier papers3
- ElasticMM: Efficient Multimodal LLMs Serving with Elastic Multimodal ParallelismZedong Liu, Shenggan Cheng, Guangming Tan, Yang You et al.NeurIPS 2025 · 12 citations
- Efficient Multi-round LLM Inference over Disaggregated ServingWenhao He, Youhe Jiang, Penghao Zhao, Quanqing Xu et al.ICML 2026 · 7 citations
- SpaceServe: Spatial Multiplexing of Complementary Encoders and Decoders for Multimodal LLMsZhicheng Li, Shuoming Zhang, Jiacheng Zhao, Siqi Li et al.NeurIPS 2025 · 5 citations
Builds on6
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng et al.SOSP 2023 · 1,016 citations
- DéjàVu: KV-cache Streaming for Fast, Fault-tolerant Generative LLM ServingFoteini Strati, Sara McAllister, Amar Phanishayee, Jakub Tarnawski et al.ICML 2024 · 59 citations
- Intern VL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic TasksZhe Chen, Jiannan Wu, Wenhai Wang, Weijie Su et al.CVPR 2024
- Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video AnalysisChaoyou Fu, Yuhan Dai, Yongdong Luo, Lei Li et al.CVPR 2025
Related papers
- DistServe: Disaggregating Prefill and Decoding for Goodput-optimized Large Language Model ServingYinmin Zhong, Shengyu Liu, Junda Chen, Jianbo Hu et al.OSDI 2024 · 646 citations
- Not All Prefills Are Equal: PPD Disaggregation for Multi-turn LLM ServingZongze Li, Jingyu Liu, Zach Xu, Yineng Zhang et al.ICML 2026 · 4 citations
- Tropical: Enhancing SLO Attainment in Disaggregated LLM Serving via SLO-Aware MultiplexingJinming Ma, Jiefei Chen, Xiuhong Li, Jiangfei Duan et al.DAC 2025 · 2 citations
- WindServe: Efficient Phase-Disaggregated LLM Serving with Stream-based Dynamic SchedulingJingqi Feng, Yukai Huang, Rui Zhang, Sicheng Liang et al.ISCA 2025 · 16 citations
- Semantic Cache Distillation: Efficient State Transfer via Reuse and Selective PatchingQianli Ma, Zhiqing Tang, Hanshuai Cui, Zhi Yao et al.ICML 2026 · 1 citation
