High-Performance GPU-to-CPU Transpilation and Optimization via High-Level Parallel Constructs
William S. Moses, Ivan R. Ivanov, Jens Domke, Toshio Endo, Johannes Doerfert, Oleksandr Zinenko
摘要
While parallelism remains the main source of performance, architectural implementations and programming models change with each new hardware generation, often leading to costly application re-engineering. Most tools for performance portability require manual and costly application porting to yet another programming model.
We propose an alternative approach that automatically translates programs written in one programming model (CUDA), into another (CPU threads) based on Polygeist/MLIR. Our approach includes a representation of parallel constructs that allows conventional compiler transformations to apply transparently and without modification and enables parallelism-specific optimizations. We evaluate our framework by transpiling and optimizing the CUDA Rodinia benchmark suite for a multicore CPU and achieve a 76% geomean speedup over handwritten OpenMP code. Further, we show how CUDA kernels from PyTorch can efficiently run and scale on the CPU-only Supercomputer Fugaku without user intervention. Our PyTorch compatibility layer making use of transpiled CUDA PyTorch kernels outperforms the PyTorch CPU native backend by 2.7×.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Unleashing CPU Potential for Executing GPU Programs Through Compiler/Runtime OptimizationsRuobing Han, Jisheng Zhao, Hyesoon KimMICRO 2024 · 被引用 4 次
- Directed Testing in MLIR: Unleashing Its Potential by Overcoming the Limitations of Random FuzzingWeiyuan Tong, Zixu Wang, Zhanyong Tang, Jianbin Fang 等FSE 2025 · 被引用 1 次
- Interleaved Learning and Exploration: A Self-Adaptive Fuzz Testing Framework for MLIRZeyu Sun, Jingjing Liang, Weiyi Wang, Chenyao Suo 等ASE 2025 · 被引用 1 次
- Strided Difference Bound MatricesArjun Pitchanathan, Albert Cohen, Oleksandr Zinenko, Tobias GrosserCAV 2024
- Scaling GPU-to-CPU Migration for Efficient Distributed Execution on CPU ClustersRuobing Han, Hyesoon KimPPoPP 2026
它引用的顶会 Paper3
- Co-design for A64FX manycore processor and "Fugaku"Mitsuhisa Sato, Yutaka Ishikawa, Hirofumi Tomita, Yuetsu Kodama 等SC 2020 · 被引用 112 次
- Reverse-mode automatic differentiation and optimization of GPU kernels via enzymeWilliam S. Moses, Valentin Churavy, Ludger Paehler, Jan Hückelheim 等SC 2021 · 被引用 50 次
- Specifying and testing GPU workgroup progress modelsTyler Sorensen, Lucas F. Salvador, Harmit Raval, Hugues Evrard 等OOPSLA 2021 · 被引用 11 次
相关 Paper
- GraCE: Unlocking CUDA Graphs with Compiler Support for ML WorkloadsAbhishek Ghosh, Ajay Nayak, Ashish Panwar, Arkaprava BasuOSDI 2026
- Dynamically Fusing Python HPC KernelsNader Al Awar, Muhammad Hannan Naeem, James Almgren-Bell, George Biros 等ISSTA 2025
- PyTorch 2: Faster Machine Learning Through Dynamic Python Bytecode Transformation and Graph CompilationJason Ansel, Edward Z. Yang, Horace He, Natalia Gimelshein 等ASPLOS 2024 · 被引用 693 次
- BabelTower: Learning to Auto-parallelized Program TranslationYuanbo Wen, Qi Guo, Qiang Fu, Xiaqing Li 等ICML 2022 · 被引用 27 次
- PerfDojo: Automated ML Library Generation for Heterogeneous ArchitecturesAndrei Ivanov, Siyuan Shen, Gioele Gottardo, Marcin Chrapek 等SC 2025 · 被引用 2 次
