Dynamically Fusing Python HPC Kernels
Nader Al Awar, Muhammad Hannan Naeem, James Almgren-Bell, George Biros, Milos Gligoric
Abstract
Recent trends in high-performance computing show an increase in the adoption of performance portable frameworks such as Kokkos and interpreted languages such as Python. PyKokkos follows these trends and enables programmers to write performance-portable kernels in Python which greatly increases productivity. One issue that programmers still face is how to organize parallel code, as splitting code into separate kernels simplifies testing and debugging but may result in suboptimal performance. To enable programmers to organize kernels in any way they prefer while ensuring good performance, we present PyFuser, a program analysis framework for automatic fusion of performance portable PyKokkos kernels. PyFuser dynamically traces kernel calls and lazily fuses them once the result is requested by the application. PyFuser generates fused kernels that execute faster due to better reuse of data, improved compiler optimizations, and reduced kernel launch overhead, while not requiring any changes to existing PyKokkos code. We also introduce automated code transformations that further optimize the fused kernels generated by PyFuser. Our experiments show that on average PyFuser achieves speedups compared to unfused kernels of 3.8× on NVIDIA and AMD GPUs, as well as Intel and AMD CPUs. CCS Concepts: • Software and its engineering → Just-in-time compilers; • Computing methodologies → Parallel programming languages.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5d7f9ff1-821c-469c-a260-6b4331ef89ccBuilds on2
- Productivity, portability, performance: data-centric PythonAlexandros Nikolaos Ziogas, Timo Schneider, Tal Ben-Nun, Alexandru Calotoiu et al.SC 2021 · 32 citations
- Parla: A Python Orchestration System for Heterogeneous ArchitecturesHochan Lee, William Ruys, Ian Henriksen, Arthur Michener Peters et al.SC 2022 · 6 citations
Related papers
- High-Performance GPU-to-CPU Transpilation and Optimization via High-Level Parallel ConstructsWilliam S. Moses, Ivan R. Ivanov, Jens Domke, Toshio Endo et al.PPoPP 2023 · 27 citations
- Breaking the computation and communication abstraction barrier in distributed machine learning workloadsAbhinav Jangda, Jun Huang, Guodong Liu, Amir Hossein Nodehi Sabet et al.ASPLOS 2022 · 68 citations
- MCFuser: High-Performance and Rapid Fusion of Memory-Bound Compute-Intensive OperatorsZheng Zhang, Donglin Yang, Xiaobo Zhou, Dazhao ChengSC 2024 · 9 citations
- Dynamic Generation of Python Bindings for HPC KernelsSteven Zhu, Nader Al Awar, Mattan Erez, Milos GligoricASE 2021 · 3 citations
- Discovering Parallelisms in Python ProgramsSiwei Wei, Guyang Song, Senlin Zhu, Ruoyi Ruan et al.FSE 2023 · 1 citation
