Automatic Parallelism Management
Sam Westrick, Matthew Fluet, Mike Rainey, Umut A. Acar
摘要
On any modern computer architecture today, parallelism comes with a modest cost, born from the creation and management of threads or tasks. Today, programmers battle this cost by manually optimizing/tuning their codes to minimize the cost of parallelism without harming its benefit, performance. This is a difficult battle: programmers must reason about architectural constant factors hidden behind layers of software abstractions, including thread schedulers and memory managers, and their impact on performance, also at scale. In languages that support higher-order functions, the battle hardens: higher order functions can make it difficult, if not impossible, to reason about the cost and benefits of parallelism.
Motivated by these challenges and the numerous advantages of high-level languages, we believe that it has become essential to manage parallelism automatically so as to minimize its cost and maximize its benefit. This is a challenging problem, even when considered on a case-by-case, application-specific basis. But if a solution were possible, then it could combine the many correctness benefits of high-level languages with performance by managing parallelism without the programmer effort needed to ensure performance. This paper proposes techniques for such automatic management of parallelism by combining static (compilation) and run-time techniques. Specifically, we consider the Parallel ML language with task parallelism, and describe a compiler pipeline that embeds "potential parallelism" directly into the call-stack and avoids the cost of task creation by default. We then pair this compilation pipeline with a run-time system that dynamically converts potential parallelism into actual parallel tasks. Together, the compiler and run-time system guarantee that the cost of parallelism remains low without losing its benefit. We prove that our techniques have no asymptotic impact on the work and span of parallel programs and thus preserve their asymptotic properties. We implement the proposed techniques by extending the MPL compiler for Parallel ML and show that it can eliminate the burden of manual optimization while delivering good practical performance.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Efficient Dual-Numbers Reverse AD via Well-Known Program TransformationsTom Smeding, Matthijs VákárPOPL 2023 · 被引用 10 次
- Beehive: A Scalable Disaggregated Memory Runtime Exploiting Asynchrony of Multithreaded ProgramsQuanxi Li, Hong Huang, Ying Liu, Yanwen Xia 等NSDI 2025 · 被引用 7 次
- Composing Distributed Computations Through Task and Kernel FusionRohan Yadav, Shiv Sundram, Wonchan Lee, Michael Garland 等ASPLOS 2025
它引用的顶会 Paper9
- Disentanglement in nested-parallel programsSam Westrick, Rohan Yadav, Matthew Fluet, Umut A. AcarPOPL 2020 · 被引用 19 次
- Responsive parallelism with futures and stateStefan K. Muller, Kyle Singer, Noah Goldstein, Umut A. Acar 等PLDI 2020 · 被引用 12 次
- Parallel determinacy race detection for futuresYifan Xu, Kyle Singer, I-Ting Angelina LeePPoPP 2020 · 被引用 11 次
- Task parallel assembly language for uncompromising parallelismMike Rainey, Ryan R. Newton, Kyle C. Hale, Nikos Hardavellas 等PLDI 2021 · 被引用 9 次
- Provably space-efficient parallel functional programmingJatin Arora, Sam Westrick, Umut A. AcarPOPL 2021 · 被引用 9 次
相关 Paper
- Efficient Parallel Functional Programming with EffectsJatin Arora, Sam Westrick, Umut A. AcarPLDI 2023 · 被引用 6 次
- Parallelism in a Region Inference ContextMartin Elsman, Troels HenriksenPLDI 2023 · 被引用 2 次
- Compiling Loop-Based Nested Parallelism for Irregular WorkloadsYian Su, Mike Rainey, Nick Wanninger, Nadharm Dhiantravan 等ASPLOS 2024 · 被引用 5 次
- T4: Compiling Sequential Code for Effective Speculative Parallelization in HardwareVictor A. Ying, Mark C. Jeffrey, Daniel SánchezISCA 2020 · 被引用 25 次
- Responsive Parallelism with Dynamic and First-Class PrioritiesMarelle León, My Dinh, Stefan K. MullerPLDI 2026
