Beyond malloc efficiency to fleet efficiency: a hugepage-aware memory allocator
A. H. Hunter, Chris Kennelly, Paul Turner, Darryl Gove, Tipp Moseley, Parthasarathy Ranganathan
Abstract
Memory allocation represents significant compute cost at the warehouse scale and its optimization can yield considerable cost savings. One classical approach is to increase the efficiency of an allocator to minimize the cycles spent in the allocator code. However, memory allocation decisions also impact overall application performance via data placement, offering opportunities to improve fleetwide productivity by completing more units of application work using fewer hardware resources. Here, we focus on hugepage coverage. We present TEMERAIRE, a hugepage-aware enhancement of TCMALLOC to reduce CPU overheads in the application's code. We discuss the design and implementation of TEMERAIRE including strategies for hugepage-aware memory layouts to maximize hugepage coverage and to minimize fragmentation overheads. We present application studies for 8 applications, improving requests-per-second (RPS) by 7.7% and reducing RAM usage 2.4%. We present the results of a 1% experiment at fleet scale as well as the longitudinal rollout in Google's warehouse scale computers. This yielded 6% fewer TLB miss stalls, and 26% reduction in memory wasted due to fragmentation. We conclude with a discussion of additional techniques for improving the allocator development process and potential optimization strategies for future memory allocators.
- How do we pick object sizes and organize metadata to
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f54e711b-5edc-40c9-9016-63c4c1a2f2c6Cited by top-tier papers16
- Towards an Adaptable Systems Architecture for Memory Tiering at Warehouse-ScalePadmapriya Duraisamy, Wei Xu, Scott Hare, Ravi Rajwar et al.ASPLOS 2023 · 74 citations
- Propeller: A Profile Guided, Relinking Optimizer for Warehouse-Scale ApplicationsHan Shen, Krzysztof Pszeniczny, Rahman Lavaee, Snehasish Kumar et al.ASPLOS 2023 · 37 citations
- Carbink: Fault-Tolerant Far MemoryYang Zhou, Hassan M. G. Wassel, Sihang Liu, Jiaqi Gao et al.OSDI 2022 · 35 citations
- Trident: Harnessing Architectural Resources for All Page Sizes in x86 ProcessorsVenkat Sri Sai Ram, Ashish Panwar, Arkaprava BasuMICRO 2021 · 27 citations
- Contiguitas: The Pursuit of Physical Memory Contiguity in DatacentersKaiyang Zhao, Kaiwen Xue, Ziqi Wang, Dan Schatzberg et al.ISCA 2023 · 26 citations
Builds on2
- Accelerometer: Understanding Acceleration Opportunities for Data Center Overheads at HyperscaleAkshitha Sriraman, Abhishek DhanotiaASPLOS 2020 · 78 citations
- Learning-based Memory Allocation for C++ Server WorkloadsMartin Maas, David G. Andersen, Michael Isard, Mohammad Mahdi Javanmard et al.ASPLOS 2020 · 59 citations
Related papers
- Characterizing a Memory Allocator at Warehouse ScaleZhuangzhuang Zhou, Vaibhav Gogte, Nilay Vaish, Chris Kennelly et al.ASPLOS 2024 · 17 citations
- Necro-reaper: Pruning away Dead Memory Traffic in Warehouse-Scale ComputersSotiris Apostolakis, Chris Kennelly, Xinliang David Li, Parthasarathy RanganathanASPLOS 2025 · 4 citations
- OBASE: Object-Based Address-Space Engineering to Improve Memory TieringVinay Banakar, Suli Yang, Kan Wu, Andrea C. Arpaci-Dusseau et al.OSDI 2026
- PTEMagnet: fine-grained physical memory reservation for faster page walks in public cloudsArtemiy Margaritov, Dmitrii Ustiugov, Amna Shahab, Boris GrotASPLOS 2021 · 19 citations
- Merchandiser: Data Placement on Heterogeneous Memory for Task-Parallel HPC Applications with Load-Balance AwarenessZhen Xie, Jie Liu, Jiajia Li, Dong LiPPoPP 2023 · 18 citations
