MDK: Rethinking the Data Center Memory Reclamation Problem
Shaurya Patel, Suli Yang, Yawen Wang, Kan Wu, Alexandra (Sasha) Fedorova, Margo Seltzer, Kimberly Keeton
摘要
The traditional memory management problem maximizes application performance when constrained by a fixed-size memory. Today's data centers face a different problem: their goal is to maximize the number of jobs on a server without violating performance Service Level Objectives (SLOs). Since a key constraint for placing additional jobs is memory, data center systems proactively reclaim memory from running jobs to create space for new jobs. This difference fundamentally flips the optimization problem that memory management policies need to address.
Designing practical policies requires a set of tools: 1) an optimal policy that provides a bound on what any policy can achieve, 2) metrics to compare policies, and 3) efficient techniques for evaluating potential policies. However, we find that foundational tools from the traditional setting, such as the optimal policy OPT, Miss Ratio Curves (MRCs), and efficient ways to generate MRCs, do not apply in this new setting. The data center setting demands a new set of tools.
We present the Memory Designer's Kit, MDK, a framework for designing and evaluating data center memory management policies. MDK includes an offline provably optimal policy; Memory Performance Curves (MPCs), which show how memory savings vary when constrained by performance; and an efficient technique that is up to 208× faster than simulation for producing MPCs. We demonstrate MDK's utility by developing three data center policies that improve average memory savings by up to 10% relative to a state-of-the-art policy.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper23
- Pond: CXL-Based Memory Pooling Systems for Cloud PlatformsHuaicheng Li, Daniel S. Berger, Lisa Hsu, Daniel Ernst 等ASPLOS 2023 · 被引用 328 次
- TPP: Transparent Page Placement for CXL-Enabled Tiered-MemoryHasan Al Maruf, Hao Wang, Abhishek Dhanotia, Johannes Weiner 等ASPLOS 2023 · 被引用 255 次
- A large scale analysis of hundreds of in-memory cache clusters at TwitterJuncheng Yang, Yao Yue, K. V. RashmiOSDI 2020 · 被引用 245 次
- Learning Relaxed Belady for Content Distribution Network CachingZhenyu Song, Daniel S. Berger, Kai Li, Wyatt LloydNSDI 2020 · 被引用 193 次
- An Imitation Learning Approach for Cache ReplacementEvan Zheran Liu, Milad Hashemi, Kevin Swersky, Parthasarathy Ranganathan 等ICML 2020 · 被引用 108 次
相关 Paper
- CBMM: Financial Advice for Kernel Memory ManagersMark Mansi, Bijan Tabatabai, Michael M. SwiftUSENIX ATC 2022
- M3: end-to-end memory management in elastic system software stacksDavid Lion, Adrian Chiu, Ding YuanEuroSys 2021 · 被引用 3 次
- TMO: transparent memory offloading in datacentersJohannes Weiner, Niket Agarwal, Dan Schatzberg, Leon Yang 等ASPLOS 2022 · 被引用 103 次
- Svalinn: Overload Control in Large-Scale Servers with Multiple Resource BottlenecksBhaskar Subhash Pardeshi, Peidi Song, Ahmed SaeedOSDI 2026
- Characterizing a Memory Allocator at Warehouse ScaleZhuangzhuang Zhou, Vaibhav Gogte, Nilay Vaish, Chris Kennelly 等ASPLOS 2024 · 被引用 17 次
