Lune

EuroSys2026顶会

Matrix‑PIC: Harnessing Matrix Outer-product for High‑Performance Particle‑in‑Cell Simulations

Yizhuo Rao, Xingjian Cui, Jiabin Xie, Shangzhi Pang, Guangnan Feng, Jinhui Wei, Zhiguang Chen, Yutong Lu

2026年份
1被引次数
1顶会引用

摘要

Particle-in-Cell (PIC) simulations devote most cycles to particle-grid interactions, and their fine-grained atomic updates become a severe bottleneck on traditional many-core CPUs. The evolution of CPU architectures, particularly the integration of specialized Matrix Processing Units (MPUs) designed for efficient matrix outer-product operations, presents a paradigm shift and an opportunity to alleviate these bottlenecks. Capitalizing on this architectural advancement, this work focuses on adapting the critical current deposition step in PIC simulations to this new matrix-centric computational model.

We introduce MatrixPIC, a novel framework representing, to our knowledge, the first holistic co-design of the deposition kernel, data layout, and an incremental sorting mechanism, all tailored specifically for the hybrid MPU-VPU SIMD execution model on modern CPUs. It provides a validated design paradigm for the scientific computing community to leverage the next generation of high-density hardware.

The key technical innovations are: (i) the refactoring of the current deposition algorithm into block-matrix updates that map efficiently onto the MPU's native operational paradigm; (ii) a hybrid execution pipeline that synergistically orchestrates MPU kernels for high-density accumulation with VPU stages for data preparation and control flow; and (iii) an 𝑂 (1)-amortized incremental particle sorter, which utilizes a gapped packed-memory array to establish and maintain the crucial data locality required for MPU efficiency.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

引用它的顶会 Paper1

问问它们各自怎么用它

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖