Task parallel assembly language for uncompromising parallelism
Mike Rainey, Ryan R. Newton, Kyle C. Hale, Nikos Hardavellas, Simone Campanoni, Peter A. Dinda, Umut A. Acar
Abstract
Achieving parallel performance and scalability involves making compromises between parallel and sequential computation. If not contained, the overheads of parallelism can easily outweigh its benefits, sometimes by orders of magnitude. Today, we expect programmers to implement this compromise by optimizing their code manually. This process is labor intensive, requires deep expertise, and reduces code quality. Recent work on heartbeat scheduling shows a promising approach that manifests the potentially vast amounts of available, latent parallelism, at a regular rate, based on even beats in time. The idea is to amortize the overheads of parallelism over the useful work performed between the beats. Heartbeat scheduling is promising in theory, but the reality is complicated: it has no known practical implementation.
In this paper, we propose a practical approach to heartbeat scheduling that involves equipping the assembly language with a small set of primitives. These primitives leverage existing kernel and hardware support for interrupts to allow parallelism to remain latent, until a heartbeat, when it can be manifested with low cost. Our Task Parallel Assembly Language (TPAL) is a compact, RISC-like assembly language.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 253cf2ff-176f-499f-b552-d1e3ec723f24Cited by top-tier papers3
- Automatic Parallelism ManagementSam Westrick, Matthew Fluet, Mike Rainey, Umut A. AcarPOPL 2024 · 7 citations
- SpEQ: Translation of Sparse Codes using EquivalencesAvery Laird, Bangtian Liu, Nikolaj S. Bjørner, Maryam Mehri DehnaviPLDI 2024 · 6 citations
- Compiling Loop-Based Nested Parallelism for Irregular WorkloadsYian Su, Mike Rainey, Nick Wanninger, Nadharm Dhiantravan et al.ASPLOS 2024 · 5 citations
Builds on3
- Disentanglement in nested-parallel programsSam Westrick, Rohan Yadav, Matthew Fluet, Umut A. AcarPOPL 2020 · 19 citations
- Provably space-efficient parallel functional programmingJatin Arora, Sam Westrick, Umut A. AcarPOPL 2021 · 9 citations
- Compiler-based timing for extremely fine-grain preemptive parallelismSouradip Ghosh, Michael Cuevas, Simone Campanoni, Peter A. DindaSC 2020 · 7 citations
Related papers
- T4: Compiling Sequential Code for Effective Speculative Parallelization in HardwareVictor A. Ying, Mark C. Jeffrey, Daniel SánchezISCA 2020 · 25 citations
- LibPreemptible: Enabling Fast, Adaptive, and Hardware-Assisted User-Space SchedulingYueying Li, Nikita Lazarev, David Koufaty, Tenny Yin et al.HPCA 2024 · 14 citations
- The Benefits and Limitations of User Interrupts for Preemptive Userspace SchedulingLinsong Guo, Danial Zuberi, Tal Garfinkel, Amy OusterhoutNSDI 2025 · 10 citations
- Phloem: Automatic Acceleration of Irregular Applications with Fine-Grain Pipeline ParallelismQuan M. Nguyen, Daniel SánchezHPCA 2023 · 7 citations
- Computation Is Fast, Use T hreadlet !: Efficient Threading for μs-Scale Computing via OS/Hardware Co-DesignYiming Yao, Xiaohe Qin, Yi Fan, Yuanlong Li et al.SOSP 2026
