Overlapping host-to-device copy and computation using hidden unified memory
Jaehoon Jung, Daeyoung Park, Youngdong Do, Jungho Park, Jaejin Lee
Abstract
In this paper, we propose a runtime, called HUM, which hides host-to-device memory copy time without any code modification. It overlaps the host-to-device memory copy with host computation or CUDA kernel computation by exploiting Unified Memory and fault mechanisms. HUM provides wrapper functions of CUDA commands and executes host-to-device memory copy commands in an asynchronous manner. We also propose two runtime techniques. One checks if it is correct to make the synchronous host-to-device memory copy command asynchronous. If not, HUM makes the host computation or the kernel computation wait until the memory copy completes. The other subdivides consecutive host-to-device memory copy commands into smaller memory copy requests and schedules the requests from different commands in a round-robin manner. As a result, the kernel execution can be scheduled as early as possible to maximize the overlap. We evaluate HUM using 51 applications from Parboil, Rodinia, and CUDA Code Samples and compare their performance under HUM with that of hand-optimized implementations. The evaluation result shows that executing the applications under HUM is, on average, 1.21 times faster than executing them under original CUDA. The speedup is comparable to the average speedup 1.22 of the hand-optimized implementations for Unified Memory.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 22bad307-650e-47b6-a506-73762077edd6Cited by top-tier papers3
- Occamy: Memory-efficient GPU Compiler for DNN InferenceJaeho Lee, Shinnung Jeong, Seungbin Song, Kunwoo Kim et al.DAC 2023 · 3 citations
- Efficiently Joining Large Relations on Multi-GPU SystemsTobias Maltenberger, Ilin Tolovski, Tilmann RablVLDB 2025 · 3 citations
- Designing GPU Data Structures for Efficient Memory OversubscriptionVipin Patel, Srinjoy Sarkar, Swarnendu Biswas, Mainak ChaudhuriOOPSLA 2026
Related papers
- SUV: Static Analysis Guided Unified Virtual MemoryPratheek B, Guilherme Cox, Ján Veselý, Arkaprava BasuMICRO 2024 · 7 citations
- DeepUM: Tensor Migration and Prefetching in Unified MemoryJaehoon Jung, Jinpyo Kim, Jaejin LeeASPLOS 2023 · 33 citations
- ARIADNE: Adaptive UVM Management for Efficient GPU Memory OversubscriptionHyunkyun Shin, Seongtae Bang, Hyungwon Park, Daehoon KimHPCA 2026 · 2 citations
- Batch-Aware Unified Memory Management in GPUs for Irregular WorkloadsHyojong Kim, Jaewoong Sim, Prasun Gera, Ramyad Hadidi et al.ASPLOS 2020 · 89 citations
- How to Copy Memory? Coordinated Asynchronous Copy as a First-Class OS ServiceJingkai He, Yunpeng Dong, Dong Du, Mo Zou et al.SOSP 2025 · 2 citations
