POLAR-PIC: A Holistic Framework for Matrixized PIC with Co-Designed Compute, Layout, and Communication
Yizhuo Rao, Xingjian Cui, Shangzhi Pang, Jiabin Xie, Guangnan Feng, Ziyan Zhang, Jinhui Wei, Languang Gao, Zhenyu Wang, Zhiguang Chen, Yutong Lu
Abstract
Particle-in-Cell (PIC) simulations are fundamental to plasma physics but often suffer from limited scalability due to particle–grid interaction bottlenecks and particle redistribution costs. Specifically, the particle–grid interaction computations have not taken full advantage of the emerging Matrix Processing Units (MPUs), the particle motion introduces irregular memory accesses, and the bulk-synchronous redistribution further destroys long-term data locality thereby limiting parallel efficiency. To address these inefficiencies, we present POLAR-PIC, a co-designed framework for large-scale PIC simulations that (i) reformulates Field Interpolation into an MPU-friendly outer-product form, (ii) maintains a physically ordered particle layout to preserve memory contiguity, and (iii) overlaps particle communication with Deposition to hide redistribution overhead. The evaluation on the pilot system of an Exascale supercomputer demonstrates that POLAR-PIC accelerates the entire particle-processing phase by up to 10.9 × in uniform plasma and 4.4 × in real-world laser-ion acceleration scenarios compared to the native WarpX reference pipeline on LX2. Ablation studies reveal that the speedups achieved by Interpolation and Deposition are 8.0 × and 13.2 × , respectively, and the asynchronous communication design sustains a overlap ratio. In cross-platform comparisons, POLAR-PIC achieves of theoretical peak efficiency on the CPU-based LS system, while WarpX reaches on NVIDIA A800 GPUs. Notably, the scalability evaluation demonstrates that POLAR-PIC maintains weak scaling efficiency on over 2 million cores under high-migration dynamic workloads, highlighting the importance of holistic co-design for future matrix-centric HPC systems.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b77c6ce7-8454-4144-814a-fe1ba4c29928Builds on6
- ConvStencil: Transform Stencil Computation to Matrix Multiplication on Tensor CoresYuetao Chen, Kun Li, Yuhao Wang, Donglin Bai et al.PPoPP 2024 · 25 citations
- CereSZ: Enabling and Scaling Error-bounded Lossy Compression on Cerebras CS-2Shihui Song, Yafan Huang, Peng Jiang, Xiaodong Yu et al.HPDC 2024 · 14 citations
- LoRAStencil: Low-Rank Adaptation of Stencil Computation on Tensor CoresYiwei Zhang, Kun Li, Liang Yuan, Jiawen Cheng et al.SC 2024 · 13 citations
- HStencil: Matrix-Vector Stencil Computation with Interleaved Outer Product and MLAHan Huang, Jiabin Xie, Guangnan Feng, Xianwei Zhang et al.SC 2025 · 5 citations
- Matrix‑PIC: Harnessing Matrix Outer-product for High‑Performance Particle‑in‑Cell SimulationsYizhuo Rao, Xingjian Cui, Jiabin Xie, Shangzhi Pang et al.EuroSys 2026 · 1 citation
Related papers
- Designing a GPU-Accelerated Communication Layer for Efficient Fluid-Structure Interaction Computations on Heterogeneous SystemsAristotle X. Martin, Geng Liu, Bálint Joó, Runxin Wu et al.SC 2024 · 1 citation
- A massively parallel and scalable multi-CPU material point methodXinlei Wang, Yuxing Qiu, Stuart R. Slattery, Yu Fang et al.SIGGRAPH 2020 · 82 citations
- LIA: A Single-GPU LLM Inference Acceleration with Cooperative AMX-Enabled CPU-GPU Computation and CXL OffloadingHyungyo Kim, Nachuan Wang, Qirong Xia, Jinghan Huang et al.ISCA 2025 · 20 citations
- Graphite: A NUMA-aware HPC System for Graph Analytics Based on a new MPI * X Parallelism ModelMohammad Hasanzadeh-Mofrad, Rami G. Melhem, Muhammad Yousuf Ahmad, Mohammad HammoudVLDB 2020 · 142 citations
- 5 ExaFlop/s HPL-MxP Benchmark with Linear Scalability on the 40-Million-Core Sunway SupercomputerRongfen Lin, Xinhui Yuan, Wei Xue, Wanwang Yin et al.SC 2023 · 11 citations
