Optimizing parallel PREM compilation over nested loop structures
Zhao Gu, Rodolfo Pellizzoni
摘要
We consider automatic parallelization of a computational kernel executed according to the PRedictable Execution Model (PREM), where each thread is divided into execution and memory phases. We target a scratchpad-based architecture, where memory phases are executed by a dedicated DMA component. We employ data analysis and loop tiling to split the kernel execution into segments, and schedule them based on a DAG representation of data and execution dependencies. Our main observation is that properly selecting tile sizes is key to optimize the makespan of the kernel. We thus propose a heuristic that efficiently searches for optimized tile size and core assignments over deeply nested loops, and demonstrate its applicability and performance compared to the state-of-the-art in PREM compilation using the PolyBench-NN benchmark suite.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- Analytical characterization and design space exploration for optimization of CNNsRui Li, Yufan Xu, Aravind Sukumaran-Rajam, Atanas Rountev 等ASPLOS 2021 · 被引用 52 次
- A study of predictable execution models implementation for industrial data-flow applications on a multi-core platform with shared banked memoryMatheus Schuh, Claire Maiza, Joël Goossens, Pascal Raymond 等RTSS 2020 · 被引用 10 次
- Welder: Scheduling Deep Learning Memory Access via Tile-graphYining Shi, Zhi Yang, Jilong Xue, Lingxiao Ma 等OSDI 2023 · 被引用 64 次
- Efficient tiled sparse matrix multiplication through matrix signaturesSüreyya Emre Kurt, Aravind Sukumaran-Rajam, Fabrice Rastello, P. SadayappanSC 2020 · 被引用 20 次
- Exploiting Computation Reuse for Stencil AcceleratorsYuze Chi, Jason CongDAC 2020 · 被引用 11 次
