MDK: Rethinking the Data Center Memory Reclamation Problem
Shaurya Patel, Suli Yang, Yawen Wang, Kan Wu, Alexandra (Sasha) Fedorova, Margo Seltzer, Kimberly Keeton
Abstract
The traditional memory management problem maximizes application performance when constrained by a fixed-size memory. Today's data centers face a different problem: their goal is to maximize the number of jobs on a server without violating performance Service Level Objectives (SLOs). Since a key constraint for placing additional jobs is memory, data center systems proactively reclaim memory from running jobs to create space for new jobs. This difference fundamentally flips the optimization problem that memory management policies need to address.
Designing practical policies requires a set of tools: 1) an optimal policy that provides a bound on what any policy can achieve, 2) metrics to compare policies, and 3) efficient techniques for evaluating potential policies. However, we find that foundational tools from the traditional setting, such as the optimal policy OPT, Miss Ratio Curves (MRCs), and efficient ways to generate MRCs, do not apply in this new setting. The data center setting demands a new set of tools.
We present the Memory Designer's Kit, MDK, a framework for designing and evaluating data center memory management policies. MDK includes an offline provably optimal policy; Memory Performance Curves (MPCs), which show how memory savings vary when constrained by performance; and an efficient technique that is up to 208× faster than simulation for producing MPCs. We demonstrate MDK's utility by developing three data center policies that improve average memory savings by up to 10% relative to a state-of-the-art policy.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext fd1a5d1b-4ab8-484c-91ee-dcfb97e0ee79Builds on23
- Pond: CXL-Based Memory Pooling Systems for Cloud PlatformsHuaicheng Li, Daniel S. Berger, Lisa Hsu, Daniel Ernst et al.ASPLOS 2023 · 328 citations
- TPP: Transparent Page Placement for CXL-Enabled Tiered-MemoryHasan Al Maruf, Hao Wang, Abhishek Dhanotia, Johannes Weiner et al.ASPLOS 2023 · 255 citations
- A large scale analysis of hundreds of in-memory cache clusters at TwitterJuncheng Yang, Yao Yue, K. V. RashmiOSDI 2020 · 245 citations
- Learning Relaxed Belady for Content Distribution Network CachingZhenyu Song, Daniel S. Berger, Kai Li, Wyatt LloydNSDI 2020 · 193 citations
- An Imitation Learning Approach for Cache ReplacementEvan Zheran Liu, Milad Hashemi, Kevin Swersky, Parthasarathy Ranganathan et al.ICML 2020 · 108 citations
Related papers
- CBMM: Financial Advice for Kernel Memory ManagersMark Mansi, Bijan Tabatabai, Michael M. SwiftUSENIX ATC 2022
- M3: end-to-end memory management in elastic system software stacksDavid Lion, Adrian Chiu, Ding YuanEuroSys 2021 · 3 citations
- TMO: transparent memory offloading in datacentersJohannes Weiner, Niket Agarwal, Dan Schatzberg, Leon Yang et al.ASPLOS 2022 · 103 citations
- Svalinn: Overload Control in Large-Scale Servers with Multiple Resource BottlenecksBhaskar Subhash Pardeshi, Peidi Song, Ahmed SaeedOSDI 2026
- Characterizing a Memory Allocator at Warehouse ScaleZhuangzhuang Zhou, Vaibhav Gogte, Nilay Vaish, Chris Kennelly et al.ASPLOS 2024 · 17 citations
