Dynamically Fusing Python HPC Kernels
Nader Al Awar, Muhammad Hannan Naeem, James Almgren-Bell, George Biros, Milos Gligoric
摘要
Recent trends in high-performance computing show an increase in the adoption of performance portable frameworks such as Kokkos and interpreted languages such as Python. PyKokkos follows these trends and enables programmers to write performance-portable kernels in Python which greatly increases productivity. One issue that programmers still face is how to organize parallel code, as splitting code into separate kernels simplifies testing and debugging but may result in suboptimal performance. To enable programmers to organize kernels in any way they prefer while ensuring good performance, we present PyFuser, a program analysis framework for automatic fusion of performance portable PyKokkos kernels. PyFuser dynamically traces kernel calls and lazily fuses them once the result is requested by the application. PyFuser generates fused kernels that execute faster due to better reuse of data, improved compiler optimizations, and reduced kernel launch overhead, while not requiring any changes to existing PyKokkos code. We also introduce automated code transformations that further optimize the fused kernels generated by PyFuser. Our experiments show that on average PyFuser achieves speedups compared to unfused kernels of 3.8× on NVIDIA and AMD GPUs, as well as Intel and AMD CPUs. CCS Concepts: • Software and its engineering → Just-in-time compilers; • Computing methodologies → Parallel programming languages.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper2
- Productivity, portability, performance: data-centric PythonAlexandros Nikolaos Ziogas, Timo Schneider, Tal Ben-Nun, Alexandru Calotoiu 等SC 2021 · 被引用 32 次
- Parla: A Python Orchestration System for Heterogeneous ArchitecturesHochan Lee, William Ruys, Ian Henriksen, Arthur Michener Peters 等SC 2022 · 被引用 6 次
相关 Paper
- High-Performance GPU-to-CPU Transpilation and Optimization via High-Level Parallel ConstructsWilliam S. Moses, Ivan R. Ivanov, Jens Domke, Toshio Endo 等PPoPP 2023 · 被引用 27 次
- Breaking the computation and communication abstraction barrier in distributed machine learning workloadsAbhinav Jangda, Jun Huang, Guodong Liu, Amir Hossein Nodehi Sabet 等ASPLOS 2022 · 被引用 68 次
- MCFuser: High-Performance and Rapid Fusion of Memory-Bound Compute-Intensive OperatorsZheng Zhang, Donglin Yang, Xiaobo Zhou, Dazhao ChengSC 2024 · 被引用 9 次
- Dynamic Generation of Python Bindings for HPC KernelsSteven Zhu, Nader Al Awar, Mattan Erez, Milos GligoricASE 2021 · 被引用 3 次
- Discovering Parallelisms in Python ProgramsSiwei Wei, Guyang Song, Senlin Zhu, Ruoyi Ruan 等FSE 2023 · 被引用 1 次
