POLAR-PIC: A Holistic Framework for Matrixized PIC with Co-Designed Compute, Layout, and Communication
Yizhuo Rao, Xingjian Cui, Shangzhi Pang, Jiabin Xie, Guangnan Feng, Ziyan Zhang, Jinhui Wei, Languang Gao, Zhenyu Wang, Zhiguang Chen, Yutong Lu
摘要
Particle-in-Cell (PIC) simulations are fundamental to plasma physics but often suffer from limited scalability due to particle–grid interaction bottlenecks and particle redistribution costs. Specifically, the particle–grid interaction computations have not taken full advantage of the emerging Matrix Processing Units (MPUs), the particle motion introduces irregular memory accesses, and the bulk-synchronous redistribution further destroys long-term data locality thereby limiting parallel efficiency. To address these inefficiencies, we present POLAR-PIC, a co-designed framework for large-scale PIC simulations that (i) reformulates Field Interpolation into an MPU-friendly outer-product form, (ii) maintains a physically ordered particle layout to preserve memory contiguity, and (iii) overlaps particle communication with Deposition to hide redistribution overhead. The evaluation on the pilot system of an Exascale supercomputer demonstrates that POLAR-PIC accelerates the entire particle-processing phase by up to 10.9 × in uniform plasma and 4.4 × in real-world laser-ion acceleration scenarios compared to the native WarpX reference pipeline on LX2. Ablation studies reveal that the speedups achieved by Interpolation and Deposition are 8.0 × and 13.2 × , respectively, and the asynchronous communication design sustains a overlap ratio. In cross-platform comparisons, POLAR-PIC achieves of theoretical peak efficiency on the CPU-based LS system, while WarpX reaches on NVIDIA A800 GPUs. Notably, the scalability evaluation demonstrates that POLAR-PIC maintains weak scaling efficiency on over 2 million cores under high-migration dynamic workloads, highlighting the importance of holistic co-design for future matrix-centric HPC systems.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper6
- ConvStencil: Transform Stencil Computation to Matrix Multiplication on Tensor CoresYuetao Chen, Kun Li, Yuhao Wang, Donglin Bai 等PPoPP 2024 · 被引用 25 次
- CereSZ: Enabling and Scaling Error-bounded Lossy Compression on Cerebras CS-2Shihui Song, Yafan Huang, Peng Jiang, Xiaodong Yu 等HPDC 2024 · 被引用 14 次
- LoRAStencil: Low-Rank Adaptation of Stencil Computation on Tensor CoresYiwei Zhang, Kun Li, Liang Yuan, Jiawen Cheng 等SC 2024 · 被引用 13 次
- HStencil: Matrix-Vector Stencil Computation with Interleaved Outer Product and MLAHan Huang, Jiabin Xie, Guangnan Feng, Xianwei Zhang 等SC 2025 · 被引用 5 次
- Matrix‑PIC: Harnessing Matrix Outer-product for High‑Performance Particle‑in‑Cell SimulationsYizhuo Rao, Xingjian Cui, Jiabin Xie, Shangzhi Pang 等EuroSys 2026 · 被引用 1 次
相关 Paper
- Designing a GPU-Accelerated Communication Layer for Efficient Fluid-Structure Interaction Computations on Heterogeneous SystemsAristotle X. Martin, Geng Liu, Bálint Joó, Runxin Wu 等SC 2024 · 被引用 1 次
- A massively parallel and scalable multi-CPU material point methodXinlei Wang, Yuxing Qiu, Stuart R. Slattery, Yu Fang 等SIGGRAPH 2020 · 被引用 82 次
- LIA: A Single-GPU LLM Inference Acceleration with Cooperative AMX-Enabled CPU-GPU Computation and CXL OffloadingHyungyo Kim, Nachuan Wang, Qirong Xia, Jinghan Huang 等ISCA 2025 · 被引用 20 次
- Graphite: A NUMA-aware HPC System for Graph Analytics Based on a new MPI * X Parallelism ModelMohammad Hasanzadeh-Mofrad, Rami G. Melhem, Muhammad Yousuf Ahmad, Mohammad HammoudVLDB 2020 · 被引用 142 次
- 5 ExaFlop/s HPL-MxP Benchmark with Linear Scalability on the 40-Million-Core Sunway SupercomputerRongfen Lin, Xinhui Yuan, Wei Xue, Wanwang Yin 等SC 2023 · 被引用 11 次
