TMO: transparent memory offloading in datacenters
Johannes Weiner, Niket Agarwal, Dan Schatzberg, Leon Yang, Hao Wang, Blaise Sanouillet, Bikash Sharma, Tejun Heo, Mayank Jain, Chunqiang Tang, Dimitrios Skarlatos
摘要
The unrelenting growth of the memory needs of emerging datacenter applications, along with ever increasing cost and volatility of DRAM prices, has led to DRAM being a major infrastructure expense. Alternative technologies, such as NVMe SSDs and upcoming NVM devices, offer higher capacity than DRAM at a fraction of the cost and power. One promising approach is to transparently offload colder memory to cheaper memory technologies via kernel or hypervisor techniques. The key challenge, however, is to develop a datacenter-scale solution that is robust in dealing with diverse workloads and large performance variance of different offload devices such as compressed memory, SSD, and NVM. This paper presents TMO, Meta’s transparent memory offloading solution for heterogeneous datacenter environments. TMO introduces a new Linux kernel mechanism that directly measures in realtime the lost work due to resource shortage across CPU, memory, and I/O. Guided by this information and without any prior application knowledge, TMO automatically adjusts how much memory to offload to heterogeneous devices (e.g., compressed memory or SSD) according to the device’s performance characteristics and the application’s sensitivity to memory-access slowdown. TMO holistically identifies offloading opportunities from not only the application containers but also the sidecar containers that provide infrastructure-level functions. To maximize memory savings, TMO targets both anonymous memory and file cache, and balances the swap-in rate of anonymous memory and the reload rate of file pages that were recently evicted from the file cache. TMO has been running in production for more than a year, and has saved between 20-32% of the total memory across millions of servers in our large datacenter fleet. We have successfully upstreamed TMO into the Linux kernel.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper46
- Pond: CXL-Based Memory Pooling Systems for Cloud PlatformsHuaicheng Li, Daniel S. Berger, Lisa Hsu, Daniel Ernst 等ASPLOS 2023 · 被引用 328 次
- TPP: Transparent Page Placement for CXL-Enabled Tiered-MemoryHasan Al Maruf, Hao Wang, Abhishek Dhanotia, Johannes Weiner 等ASPLOS 2023 · 被引用 255 次
- Demystifying CXL Memory with Genuine CXL-Ready Systems and DevicesYan Sun, Yifan Yuan, Zeduo Yu, Reese Kuper 等MICRO 2023 · 被引用 133 次
- Managing Memory Tiers with CXL in Virtualized EnvironmentsYuhong Zhong, Daniel S. Berger, Carl A. Waldspurger, Ryan Wee 等OSDI 2024 · 被引用 77 次
- Towards an Adaptable Systems Architecture for Memory Tiering at Warehouse-ScalePadmapriya Duraisamy, Wei Xu, Scott Hare, Ravi Rajwar 等ASPLOS 2023 · 被引用 74 次
它引用的顶会 Paper11
- An Empirical Guide to the Behavior and Use of Scalable Persistent MemoryJian Yang, Juno Kim, Morteza Hoseinzadeh, Joseph Izraelevitz 等FAST 2020 · 被引用 470 次
- Effectively Prefetching Remote Memory with LeapHasan Al Maruf, Mosharaf ChowdhuryUSENIX ATC 2020 · 被引用 186 次
- Can far memory improve job throughput?Emmanuel Amaro, Christopher Branner-Augmon, Zhihong Luo, Amy Ousterhout 等EuroSys 2020 · 被引用 163 次
- The CacheLib Caching Engine: Design and Experiences at ScaleBenjamin Berg, Daniel S. Berger, Sara McAllister, Isaac Grosof 等OSDI 2020 · 被引用 145 次
- HeMem: Scalable Tiered Memory Management for Big Data Applications and Real NVMAmanda Raybuck, Tim Stamler, Wei Zhang, Mattan Erez 等SOSP 2021 · 被引用 93 次
相关 Paper
- PageFlex: Flexible and Efficient User-space Delegation of Linux Paging Policies with eBPFAnil Yelam, Kan Wu, Zhiyuan Guo, Suli Yang 等USENIX ATC 2025 · 被引用 15 次
- IOCost: block IO control for containers in datacentersTejun Heo, Dan Schatzberg, Andrew Newell, Song Liu 等ASPLOS 2022 · 被引用 22 次
- HyperAlloc: Efficient VM Memory De/Inflation via Hypervisor-Shared Page-Frame AllocatorsLars Wrenger, Kenny Albes, Marco Wurps, Christian Dietrich 等EuroSys 2025 · 被引用 1 次
- CBMM: Financial Advice for Kernel Memory ManagersMark Mansi, Bijan Tabatabai, Michael M. SwiftUSENIX ATC 2022
- Nomad: Non-Exclusive Memory Tiering via Transactional Page MigrationLingfeng Xiang, Zhen Lin, Weishu Deng, Hui Lu 等OSDI 2024 · 被引用 59 次
