SYCL++: A Unified Programming Framework for Heterogeneous Supercomputers at Scale
Zitao Shen, Yuyang Jin, Kinman Lei, Zixuan Ma, Wenqiang Wang, Yinuo Wang, Wenhao Zhou, Zhenchuan Chen, Di Wei, Qi Zhang, Fei Wang, Ying Liu
摘要
With the rise of heterogeneous HPC systems, traditional scientific applications face increasingly severe challenges. Today’s supercomputers adopt vastly different designs in compute units, memory hierarchies, and interconnects, making application porting an expensive task that demands significant expert effort. This deep coupling results in fragmented, unmaintainable codebases. In particular, it creates substantial barriers for domain scientists, hindering algorithmic innovation across such diverse architectures. Consequently, there is an urgent need for a unified compiler and programming framework that enables true single-source performance portability across this diverse landscape. However, achieving this unification is challenging due to a central contradiction: while diverse architectures often converge on high-level abstractions, their low-level architectures frequently diverge in parallel execution models, memory management mechanisms, and network topologies. To address this, we propose SYCL++1, a unified programming framework that decouples the expression of computational semantics from platform-specific optimization strategies. SYCL++ introduces declarative primitives that allow a single, clean codebase to be efficiently mapped to different HPC systems. We validate our approach using a real-world CGFDM3D seismic simulation application. Our evaluation demonstrates that SYCL++ achieves performance portability across three distinct platforms, delivering up to 2.2 × speedups over baseline SYCL with minimal code divergence. The framework exhibits near linear scaling to 1,597,440 Sunway cores and 8,192 DCUs, enabling high-resolution simulation of realistic earthquake scenarios. We demonstrate this capability by modeling the March 2025 Myanmar earthquake at 125 m resolution. To our knowledge, this is among the first demonstrations of a unified HPC application achieving performance portability across three distinct leading supercomputers.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- CCAMP: an integrated translation and optimization framework for OpenACC and OpenMPJacob Lambert, Seyong Lee, Jeffrey S. Vetter, Allen D. MalonySC 2020 · 被引用 17 次
- Taming the Zoo: The Unified GraphIt Compiler Framework for Novel ArchitecturesAjay Brahmakshatriya, Emily Furst, Victor A. Ying, Claire Hsu 等ISCA 2021 · 被引用 12 次
- UniQ: A Unified Programming Model for Efficient Quantum Circuit SimulationChen Zhang, Haojie Wang, Zixuan Ma, Lei Xie 等SC 2022 · 被引用 11 次
- SYnergy: Fine-grained Energy-Efficient Heterogeneous Computing for Scalable Energy SavingKaijie Fan, Marco D'Antonio, Lorenzo Carpentieri, Biagio Cosenza 等SC 2023 · 被引用 13 次
- 5 ExaFlop/s HPL-MxP Benchmark with Linear Scalability on the 40-Million-Core Sunway SupercomputerRongfen Lin, Xinhui Yuan, Wei Xue, Wanwang Yin 等SC 2023 · 被引用 11 次
