Mooncake: Trading More Storage for Less Computation - A KVCache-centric Architecture for Serving LLM Chatbot
Ruoyu Qin, Zheming Li, Weiran He, Jialei Cui, Feng Ren, Mingxing Zhang, Yongwei Wu, Weimin Zheng, Xinran Xu
Abstract
MOONCAKE is the serving platform for Kimi, an LLM chatbot service developed by Moonshot AI. This platform features a KVCache-centric disaggregated architecture that not only separates prefill and decoding clusters but also efficiently utilizes the underexploited CPU, DRAM, SSD and NIC resources of the GPU cluster to establish a disaggregated KV-Cache. At the core of MOONCAKE is its KVCache-centric global cache and a scheduler designed to maximize throughput while adhering to stringent latency-related Service Level Objectives (SLOs).
Our experiments demonstrate that MOONCAKE excels in scenarios involving long-context inputs. In tests using real traces, MOONCAKE increases the effective request capacity by 59%∼498% when compared to baseline methods, all while complying with SLOs. Currently, MOONCAKE is operational across thousands of nodes, processing over 100 billion tokens daily. In practical deployments, MOONCAKE's innovative architecture enables Kimi to handle 115% and 107% more requests on NVIDIA A800 and H800 clusters, respectively, compared to previous systems.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 87ebfd0a-84c6-45b8-b792-002aa342063aCited by top-tier papers58
- ServeGen: Workload Characterization and Generation of Large Language Model Serving in ProductionYuxing Xiang, Xue Li, Kun Qian, Yan Zhang et al.NSDI 2026 · 58 citations
- Strata: Hierarchical Context Caching for Long Context Language Model ServingZhiqiang Xie, Ziyi Xu, Mark Zhao, Yuwei An et al.OSDI 2026 · 40 citations
- DEEPSERVE: Serverless Large Language Model Serving at ScaleJunhao Hu, Jiang Xu, Zhixia Liu, Yulong He et al.USENIX ATC 2025 · 38 citations
- Prism: Cost-Efficient Multi-LLM Serving via GPU Memory BallooningShan Yu, Yifan Qiao, Mingyuan Ma, Yangmin Li et al.OSDI 2026 · 33 citations
- Seer: Online Context Learning for Fast Synchronous LLM Reinforcement LearningRuoyu Qin, Weiran He, Weixiao Huang, Yangkun Zhang et al.OSDI 2026 · 32 citations
Builds on21
- Generative Agents: Interactive Simulacra of Human BehaviorJoon Sung Park, Joseph C. O'Brien, Carrie Jun Cai, Meredith Ringel Morris et al.UIST 2023 · 1,882 citations
- SGLang: Efficient Execution of Structured Language Model ProgramsLianmin Zheng, Liangsheng Yin, Zhiqiang Xie, Chuyue Sun et al.NeurIPS 2024 · 1,586 citations
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng et al.SOSP 2023 · 1,016 citations
- KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache QuantizationColeman Hooper, Sehoon Kim, Hiva Mohammadzadeh, Michael W. Mahoney et al.NeurIPS 2024 · 738 citations
- Orca: A Distributed Serving System for Transformer-Based Generative ModelsGyeong-In Yu, Joo Seong Jeong, Geon-Woo Kim, Soojeong Kim et al.OSDI 2022 · 690 citations
Related papers
- DualPath: Accelerating Agentic LLM Inference by Harvesting Disaggregated KV-Cache Storage I/OYongtong Wu, Shaoyuan Chen, Rilin Huang, Yixuan Tan et al.SIGCOMM 2026
- Libra: Flexible Request Partitioning and Scheduling for Serving Unbalanced and Dynamic LLM WorkloadsChaoyi Ruan, Yinhe Chen, Dongqi Tian, Yandong Shi et al.NSDI 2026 · 5 citations
- Efficient LLM Serving on Commodity GPU Clusters with Data-Reduced Cross-Instance OrchestrationJiangsu Du, Hongbin Zhang, Taosheng Wei, Zhenyi Zheng et al.OSDI 2026
- Compute or Load KV Cache? Why Not Both?Shuowei Jin, Xueshen Liu, Qingzhao Zhang, Zhuoqing MaoICML 2025
- SYMPHONY: Enabling Compute-Memory Disaggregation in LLM Serving SystemsSaurabh Agarwal, Bodun Hu, Anyong Mao, Aditya Akella et al.NSDI 2026 · 4 citations
