Offload Annotations: Bringing Heterogeneous Computing to Existing Libraries and Workloads
Gina Yuan, Shoumik Palkar, Deepak Narayanan, Matei Zaharia
Abstract
As specialized hardware accelerators such as GPUs become increasingly popular, developers are looking for ways to target these platforms with high-level APIs. One promising approach is kernel libraries such as PyTorch or cuML, which provide interfaces that mirror CPU-only counterparts such as NumPy or Scikit-Learn. Unfortunately, these libraries are hard to develop and to adopt incrementally: they only support a subset of their CPU equivalents, only work with datasets that fit in device memory, and require developers to reason about data placement and transfers manually. To address these shortcomings, we present a new approach called offload annotations (OAs) that enables heterogeneous GPU computing in existing workloads with few or no code modifications. An annotator annotates the types and functions in a CPU library with equivalent kernel library functions and provides an offloading API to specify how the inputs and outputs of the function can be partitioned into chunks that fit in device memory and transferred between devices. A runtime then maps existing CPU functions to equivalent GPU kernels and schedules execution, data transfers and paging. In data science workloads using CPU libraries such as NumPy and Pandas, OAs enable speedups of up to 1200× and a median speedup of 6.3× by transparently offloading functions to a GPU using existing kernel libraries. In many cases, OAs match the performance of handwritten heterogeneous implementations. Finally, OAs can automatically page data in these workloads to scale to datasets larger than GPU memory, which would need to be done manually with most current GPU libraries.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3f053eca-b2d7-4787-8a78-e16d72deedfaCited by top-tier papers3
- Tensors: An abstraction for general data processingDimitrios Koutsoukos, Supun Nakandala, Konstantinos Karanasos, Karla Saur et al.VLDB 2021 · 38 citations
- PaSh: light-touch data-parallel shell processingNikos Vasilakis, Konstantinos Kallas, Konstantinos Mamouras, Achilles Benetopoulos et al.EuroSys 2021 · 12 citations
- DiSh: Dynamic Shell-Script DistributionTammam Mustafa, Konstantinos Kallas, Pratyush Das, Nikos VasilakisNSDI 2023 · 11 citations
Related papers
- SmartDispatch: Dynamic Substitution of NumPy-Style APIs on Heterogeneous CPU-GPU SystemsJinku Cui, Yueming Hao, Shuyin Jiao, Jiajia Li et al.FSE 2026
- Composing Distributed Computations Through Task and Kernel FusionRohan Yadav, Shiv Sundram, Wonchan Lee, Michael Garland et al.ASPLOS 2025
- Legate Sparse: Distributed Sparse Computing in PythonRohan Yadav, Wonchan Lee, Melih Elibol, Manolis Papadakis et al.SC 2023 · 8 citations
- Bringing UMAP Closer to the Speed of Light with GPU AccelerationCorey J. Nolet, Victor Lafargue, Edward Raff, Thejaswi Nanditale et al.AAAI 2021 · 38 citations
- PystachIO: Efficient Distributed GPU Query Processing with PyTorch over Fast Networks & Fast StorageJigao Luo, Nils Boeschen, Muhammad El-Hindi, Carsten BinnigVLDB 2026
