Matrix‑PIC: Harnessing Matrix Outer-product for High‑Performance Particle‑in‑Cell Simulations
Yizhuo Rao, Xingjian Cui, Jiabin Xie, Shangzhi Pang, Guangnan Feng, Jinhui Wei, Zhiguang Chen, Yutong Lu
Abstract
Particle-in-Cell (PIC) simulations devote most cycles to particle-grid interactions, and their fine-grained atomic updates become a severe bottleneck on traditional many-core CPUs. The evolution of CPU architectures, particularly the integration of specialized Matrix Processing Units (MPUs) designed for efficient matrix outer-product operations, presents a paradigm shift and an opportunity to alleviate these bottlenecks. Capitalizing on this architectural advancement, this work focuses on adapting the critical current deposition step in PIC simulations to this new matrix-centric computational model.
We introduce MatrixPIC, a novel framework representing, to our knowledge, the first holistic co-design of the deposition kernel, data layout, and an incremental sorting mechanism, all tailored specifically for the hybrid MPU-VPU SIMD execution model on modern CPUs. It provides a validated design paradigm for the scientific computing community to leverage the next generation of high-density hardware.
The key technical innovations are: (i) the refactoring of the current deposition algorithm into block-matrix updates that map efficiently onto the MPU's native operational paradigm; (ii) a hybrid execution pipeline that synergistically orchestrates MPU kernels for high-density accumulation with VPU stages for data preparation and control flow; and (iii) an 𝑂 (1)-amortized incremental particle sorter, which utilizes a gapped packed-memory array to establish and maintain the crucial data locality required for MPU efficiency.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 07de338a-fb9c-4bc7-9a1d-ab897f3c503bCited by top-tier papers1
Ask how each one uses itRelated papers
- Unlocking High Performance with Low-Bit NPUs and CPUs for Highly Optimized HPL-MxP on Cloud Brain IIWeicheng Xue, Kai Yang, Yongxiang Liu, Dengdong Fan et al.SC 2024 · 3 citations
- A massively parallel and scalable multi-CPU material point methodXinlei Wang, Yuxing Qiu, Stuart R. Slattery, Yu Fang et al.SIGGRAPH 2020 · 82 citations
- A Scalable Hybrid Total FETI Method for Massively Parallel FEM SimulationsKehao Lin, Chunbao Zhou, Yan Zeng, Ningming Nie et al.PPoPP 2023 · 2 citations
- Scaling the hartree-fock matrix build on summitGiuseppe M. J. Barca, David L. Poole, Jorge L. Galvez Vallejo, Melisa Alkan et al.SC 2020 · 29 citations
- C3ache: Towards Hierarchical Cache-Centric Computing for Sparse Matrix Multiplication on GPGPUsXiaojie Li, Mingyu Wang, Baiqing Zhong, Haiqiu Huang et al.MICRO 2025 · 1 citation
