Learning-based Memory Allocation for C++ Server Workloads
Martin Maas, David G. Andersen, Michael Isard, Mohammad Mahdi Javanmard, Kathryn S. McKinley, Colin Raffel
Abstract
Modern C++ servers have memory footprints that vary widely over time, causing persistent heap fragmentation of up to 2× from long-lived objects allocated during peak memory usage. This fragmentation is exacerbated by the use of huge (2MB) pages, a requirement for high performance on large heap sizes. Reducing fragmentation automatically is challenging because C++ memory managers cannot move objects.
This paper presents a new approach to huge page fragmentation. It combines modern machine learning techniques with a novel memory manager (Llama) that manages the heap based on object lifetimes and huge pages (divided into blocks and lines). A neural network-based language model predicts lifetime classes using symbolized calling contexts. The model learns context-sensitive per-allocation site lifetimes from previous runs, generalizes over different binary versions, and extrapolates from samples to unobserved calling contexts. Instead of size classes, Llama's heap is organized by lifetime classes that are dynamically adjusted based on observed behavior at a block granularity.
Llama reduces memory fragmentation by up to 78% while only using huge pages on several production servers. We address ML-specific questions such as tolerating mispredictions and amortizing expensive predictions across application execution. Although our results focus on memory allocation, the questions we identify apply to other system-level problems with strict latency and resource requirements where machine learning could be applied.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ca8b11ee-bc08-4cff-b5b4-c49239d749d7Cited by top-tier papers24
- Pond: CXL-Based Memory Pooling Systems for Cloud PlatformsHuaicheng Li, Daniel S. Berger, Lisa Hsu, Daniel Ernst et al.ASPLOS 2023 · 328 citations
- LinnOS: Predictability on Unpredictable Flash Storage with a Light Neural NetworkMingzhe Hao, Levent Toksoz, Nanqinqin Li, Edward Edberg Halim et al.OSDI 2020 · 97 citations
- GL-Cache: Group-level learning for efficient and high-performance cachingJuncheng Yang, Ziming Mao, Yao Yue, K. V. RashmiFAST 2023 · 60 citations
- Beyond malloc efficiency to fleet efficiency: a hugepage-aware memory allocatorA. H. Hunter, Chris Kennelly, Paul Turner, Darryl Gove et al.OSDI 2021 · 51 citations
- FlexMem: Adaptive Page Profiling and Migration for Tiered MemoryDong Xu, Junhee Ryu, Kwangsik Shin, Pengfei Su et al.USENIX ATC 2024 · 41 citations
Related papers
- GMLake: Efficient and Transparent GPU Memory Defragmentation for Large-scale DNN Training with Virtual Memory StitchingCong Guo, Rui Zhang, Jiale Xu, Jingwen Leng et al.ASPLOS 2024 · 30 citations
- Jenga: Effective Memory Management for Serving LLM with HeterogeneityChen Zhang, Kuntai Du, Shu Liu, Woosuk Kwon et al.SOSP 2025
- vAttention: Dynamic Memory Management for Serving LLMs without PagedAttentionRamya Prabhu, Ajay Nayak, Jayashree Mohan, Ramachandran Ramjee et al.ASPLOS 2025 · 38 citations
- STAlloc: Enhancing Memory Efficiency in Large-Scale Model Training with Spatio-Temporal PlanningZixiao Huang, Junhao Hu, Hao Lin, Chunyang Zhu et al.EuroSys 2026 · 1 citation
- ConServe: Contiguity-Preserving Memory Management for Multi-Turn LLM ServingBingyao LiISCA 2026
