Optimizing parallel PREM compilation over nested loop structures
Zhao Gu, Rodolfo Pellizzoni
Abstract
We consider automatic parallelization of a computational kernel executed according to the PRedictable Execution Model (PREM), where each thread is divided into execution and memory phases. We target a scratchpad-based architecture, where memory phases are executed by a dedicated DMA component. We employ data analysis and loop tiling to split the kernel execution into segments, and schedule them based on a DAG representation of data and execution dependencies. Our main observation is that properly selecting tile sizes is key to optimize the makespan of the kernel. We thus propose a heuristic that efficiently searches for optimized tile size and core assignments over deeply nested loops, and demonstrate its applicability and performance compared to the state-of-the-art in PREM compilation using the PolyBench-NN benchmark suite.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Related papers
- Analytical characterization and design space exploration for optimization of CNNsRui Li, Yufan Xu, Aravind Sukumaran-Rajam, Atanas Rountev et al.ASPLOS 2021 · 52 citations
- A study of predictable execution models implementation for industrial data-flow applications on a multi-core platform with shared banked memoryMatheus Schuh, Claire Maiza, Joël Goossens, Pascal Raymond et al.RTSS 2020 · 10 citations
- Welder: Scheduling Deep Learning Memory Access via Tile-graphYining Shi, Zhi Yang, Jilong Xue, Lingxiao Ma et al.OSDI 2023 · 64 citations
- Efficient tiled sparse matrix multiplication through matrix signaturesSüreyya Emre Kurt, Aravind Sukumaran-Rajam, Fabrice Rastello, P. SadayappanSC 2020 · 20 citations
- Exploiting Computation Reuse for Stencil AcceleratorsYuze Chi, Jason CongDAC 2020 · 11 citations
