Beyond malloc efficiency to fleet efficiency: a hugepage-aware memory allocator
A. H. Hunter, Chris Kennelly, Paul Turner, Darryl Gove, Tipp Moseley, Parthasarathy Ranganathan
摘要
Memory allocation represents significant compute cost at the warehouse scale and its optimization can yield considerable cost savings. One classical approach is to increase the efficiency of an allocator to minimize the cycles spent in the allocator code. However, memory allocation decisions also impact overall application performance via data placement, offering opportunities to improve fleetwide productivity by completing more units of application work using fewer hardware resources. Here, we focus on hugepage coverage. We present TEMERAIRE, a hugepage-aware enhancement of TCMALLOC to reduce CPU overheads in the application's code. We discuss the design and implementation of TEMERAIRE including strategies for hugepage-aware memory layouts to maximize hugepage coverage and to minimize fragmentation overheads. We present application studies for 8 applications, improving requests-per-second (RPS) by 7.7% and reducing RAM usage 2.4%. We present the results of a 1% experiment at fleet scale as well as the longitudinal rollout in Google's warehouse scale computers. This yielded 6% fewer TLB miss stalls, and 26% reduction in memory wasted due to fragmentation. We conclude with a discussion of additional techniques for improving the allocator development process and potential optimization strategies for future memory allocators.
- How do we pick object sizes and organize metadata to
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper16
- Towards an Adaptable Systems Architecture for Memory Tiering at Warehouse-ScalePadmapriya Duraisamy, Wei Xu, Scott Hare, Ravi Rajwar 等ASPLOS 2023 · 被引用 74 次
- Propeller: A Profile Guided, Relinking Optimizer for Warehouse-Scale ApplicationsHan Shen, Krzysztof Pszeniczny, Rahman Lavaee, Snehasish Kumar 等ASPLOS 2023 · 被引用 37 次
- Carbink: Fault-Tolerant Far MemoryYang Zhou, Hassan M. G. Wassel, Sihang Liu, Jiaqi Gao 等OSDI 2022 · 被引用 35 次
- Trident: Harnessing Architectural Resources for All Page Sizes in x86 ProcessorsVenkat Sri Sai Ram, Ashish Panwar, Arkaprava BasuMICRO 2021 · 被引用 27 次
- Contiguitas: The Pursuit of Physical Memory Contiguity in DatacentersKaiyang Zhao, Kaiwen Xue, Ziqi Wang, Dan Schatzberg 等ISCA 2023 · 被引用 26 次
它引用的顶会 Paper2
- Accelerometer: Understanding Acceleration Opportunities for Data Center Overheads at HyperscaleAkshitha Sriraman, Abhishek DhanotiaASPLOS 2020 · 被引用 78 次
- Learning-based Memory Allocation for C++ Server WorkloadsMartin Maas, David G. Andersen, Michael Isard, Mohammad Mahdi Javanmard 等ASPLOS 2020 · 被引用 59 次
相关 Paper
- Characterizing a Memory Allocator at Warehouse ScaleZhuangzhuang Zhou, Vaibhav Gogte, Nilay Vaish, Chris Kennelly 等ASPLOS 2024 · 被引用 17 次
- Necro-reaper: Pruning away Dead Memory Traffic in Warehouse-Scale ComputersSotiris Apostolakis, Chris Kennelly, Xinliang David Li, Parthasarathy RanganathanASPLOS 2025 · 被引用 4 次
- OBASE: Object-Based Address-Space Engineering to Improve Memory TieringVinay Banakar, Suli Yang, Kan Wu, Andrea C. Arpaci-Dusseau 等OSDI 2026
- PTEMagnet: fine-grained physical memory reservation for faster page walks in public cloudsArtemiy Margaritov, Dmitrii Ustiugov, Amna Shahab, Boris GrotASPLOS 2021 · 被引用 19 次
- Merchandiser: Data Placement on Heterogeneous Memory for Task-Parallel HPC Applications with Load-Balance AwarenessZhen Xie, Jie Liu, Jiajia Li, Dong LiPPoPP 2023 · 被引用 18 次
