Are dynamic memory managers on GPUs slow?: a survey and benchmarks
Martin Winter, Mathias Parger, Daniel Mlakar, Markus Steinberger
Abstract
Dynamic memory management on GPUs is generally understood to be a challenging topic. On current GPUs, hundreds of thousands of threads might concurrently allocate new memory or free previously allocated memory. This leads to problems with thread contention, synchronization overhead and fragmentation. Various approaches have been proposed in the last ten years and we set out to evaluate them on a level playing field on modern hardware to answer the question, if dynamic memory managers are as slow as commonly thought of. In this survey paper, we provide a consistent framework to evaluate all publicly available memory managers in a large set of scenarios. We summarize each approach and thoroughly evaluate allocation performance (thread-based as well as warp-based), and look at performance scaling, fragmentation and real-world performance considering a synthetic workload as well as updating dynamic graphs. We discuss the strengths and weaknesses of each approach and provide guidelines for the respective best usage scenario. We provide a unified interface to integrate any of the tested memory managers into an application and switch between them for benchmarking purposes. Given our results, we can dispel some of the dread associated with dynamic memory managers on the GPU.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get de429583-5749-4fe1-9d1e-56046c88620fCited by top-tier papers8
- Efficient and Scalable Graph Pattern Mining on GPUsXuhao Chen, ArvindOSDI 2022 · 53 citations
- GTS: GPU-based Tree Index for Fast Similarity SearchYifan Zhu, Ruiyao Ma, Baihua Zheng, Xiangyu Ke et al.SIGMOD 2024 · 8 citations
- Efficient Maximal Biclique Enumeration on GPUsZhe Pan, Shuibing He, Xu Li, Xuechen Zhang et al.SC 2023 · 8 citations
- SIMR: Single Instruction Multiple Request Processing for Energy-Efficient Data Center MicroservicesMahmoud Khairy, Ahmad Alawneh, Aaron Barnes, Timothy G. RogersMICRO 2022 · 8 citations
- Occamy: Memory-efficient GPU Compiler for DNN InferenceJaeho Lee, Shinnung Jeong, Seungbin Song, Kunwoo Kim et al.DAC 2023 · 3 citations
Related papers
- Gallatin: A General-Purpose GPU Memory ManagerHunter McCoy, Prashant PandeyPPoPP 2024 · 7 citations
- Towards Sufficient GPU-accelerated Dynamic Graph Management: Survey and ExperimentYinnian Lin, Lei Zou, Xunbin SuVLDB 2025 · 2 citations
- MoonBright: A GPU Memory Allocator with Device-Side Page Table Materialization and Deferred TLB CoherenceYangyu Zhang, Lei Chen, Chunwei Xia, Shuaijiang Li et al.OSDI 2026
- Memory Harvesting in Multi-GPU Systems with Hierarchical Unified Virtual MemorySangjin Choi, Taeksoo Kim, Jinwoo Jeong, Rachata Ausavarungnirun et al.USENIX ATC 2022 · 28 citations
- SuperCollider: Scalable and Effective Data Race Detection for CUDAMark Stephenson, Sana Damani, Mohamed Tarek Ibn Ziad, Anis Ladram et al.PLDI 2026
