Beyond Static Parallel Loops: Supporting Dynamic Task Parallelism on Manycore Architectures with Software-Managed Scratchpad Memories
Lin Cheng, Max Ruttenberg, Dai Cheol Jung, Dustin Richmond, Michael B. Taylor, Mark Oskin, Christopher Batten
Abstract
Manycore architectures integrate hundreds of cores on a single chip by using simple cores and simple memory systems usually based on software-managed scratchpad memories (SPMs). However, such architectures are notoriously challenging to program, since the programmers need to manually manage all aspects of data movement and synchronization for both correctness and performance. We argue that this manycore programmability challenge is one of the key barriers to achieving the promise of manycore architectures. At the same time, the dynamic task parallel programming model is enjoying considerable success in addressing the programmability challenge of multi-core processors with tens of complex cores and hardware cache coherence.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get f6ace63a-c1f6-41d3-a8b0-e5b927c12614Cited by top-tier papers2
- Scalable, Programmable and Dense: The HammerBlade Open-Source RISC-V ManycoreDai Cheol Jung, Max Ruttenberg, Paul Gao, Scott Davidson et al.ISCA 2024 · 11 citations
- Streaming Tensor Programs: A Streaming Abstraction for Dynamic ParallelismGina Sohn, Genghan Zhang, Konstantin Hoßfeld, Jungwoo Kim et al.ASPLOS 2026 · 1 citation
Related papers
- Efficiently Supporting Dynamic Task Parallelism on Heterogeneous Cache-Coherent SystemsMoyang Wang, Tuan Ta, Lin Cheng, Christopher BattenISCA 2020 · 11 citations
- Symbiotic Task Scheduling and Data PrefetchingGilead Posluns, Mark C. JeffreyMICRO 2025 · 1 citation
- TD-NUCA: Runtime Driven Management of NUCA Caches in Task Dataflow Programming ModelsPaul Caheny, Lluc Alvarez, Marc Casas, Miquel MoretóSC 2022 · 4 citations
- Task-Based Tensor Computations on Modern GPUsRohan Yadav, Michael Garland, Alex Aiken, Michael BauerPLDI 2025 · 5 citations
- DynAMO: Improving Parallelism Through Dynamic Placement of Atomic Memory OperationsVíctor Soria Pardos, Adrià Armejach, Tiago Mück, Darío Suárez Gracia et al.ISCA 2023 · 5 citations
