Matrix‑PIC: Harnessing Matrix Outer-product for High‑Performance Particle‑in‑Cell Simulations
Yizhuo Rao, Xingjian Cui, Jiabin Xie, Shangzhi Pang, Guangnan Feng, Jinhui Wei, Zhiguang Chen, Yutong Lu
摘要
Particle-in-Cell (PIC) simulations devote most cycles to particle-grid interactions, and their fine-grained atomic updates become a severe bottleneck on traditional many-core CPUs. The evolution of CPU architectures, particularly the integration of specialized Matrix Processing Units (MPUs) designed for efficient matrix outer-product operations, presents a paradigm shift and an opportunity to alleviate these bottlenecks. Capitalizing on this architectural advancement, this work focuses on adapting the critical current deposition step in PIC simulations to this new matrix-centric computational model.
We introduce MatrixPIC, a novel framework representing, to our knowledge, the first holistic co-design of the deposition kernel, data layout, and an incremental sorting mechanism, all tailored specifically for the hybrid MPU-VPU SIMD execution model on modern CPUs. It provides a validated design paradigm for the scientific computing community to leverage the next generation of high-density hardware.
The key technical innovations are: (i) the refactoring of the current deposition algorithm into block-matrix updates that map efficiently onto the MPU's native operational paradigm; (ii) a hybrid execution pipeline that synergistically orchestrates MPU kernels for high-density accumulation with VPU stages for data preparation and control flow; and (iii) an 𝑂 (1)-amortized incremental particle sorter, which utilizes a gapped packed-memory array to establish and maintain the crucial data locality required for MPU efficiency.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它相关 Paper
- Unlocking High Performance with Low-Bit NPUs and CPUs for Highly Optimized HPL-MxP on Cloud Brain IIWeicheng Xue, Kai Yang, Yongxiang Liu, Dengdong Fan 等SC 2024 · 被引用 3 次
- A massively parallel and scalable multi-CPU material point methodXinlei Wang, Yuxing Qiu, Stuart R. Slattery, Yu Fang 等SIGGRAPH 2020 · 被引用 82 次
- A Scalable Hybrid Total FETI Method for Massively Parallel FEM SimulationsKehao Lin, Chunbao Zhou, Yan Zeng, Ningming Nie 等PPoPP 2023 · 被引用 2 次
- Scaling the hartree-fock matrix build on summitGiuseppe M. J. Barca, David L. Poole, Jorge L. Galvez Vallejo, Melisa Alkan 等SC 2020 · 被引用 29 次
- C3ache: Towards Hierarchical Cache-Centric Computing for Sparse Matrix Multiplication on GPGPUsXiaojie Li, Mingyu Wang, Baiqing Zhong, Haiqiu Huang 等MICRO 2025 · 被引用 1 次
