Efficiently Supporting Dynamic Task Parallelism on Heterogeneous Cache-Coherent Systems
Moyang Wang, Tuan Ta, Lin Cheng, Christopher Batten
Abstract
Manycore processors, with tens to hundreds of tiny cores but no hardware-based cache coherence, can offer tremendous peak throughput on highly parallel programs while being complexity and energy efficient. Manycore processors can be combined with a few high-performance big cores for executing operating systems, legacy code, and serial regions. These systems use heterogeneous cache coherence (HCC) with hardware-based cache coherence between big cores and software-centric cache coherence between tiny cores. Unfortunately, programming these heterogeneous cache-coherent systems to enable collaborative execution is challenging, especially when considering dynamic task parallelism. This paper seeks to address this challenge using a combination of light-weight software and hardware techniques. We provide a detailed description of how to implement a work-stealing runtime to enable dynamic task parallelism on heterogeneous cache-coherent systems. We also propose direct task stealing (DTS), a new technique based on user-level interrupts to bypass the memory system and thus improve the performance and energy efficiency of work stealing. Our results demonstrate that executing dynamic task-parallel applications on a 64-core system (4 big, 60 tiny) with complexity-effective HCC and DTS can achieve: speedup over a single big core; speedup over an area-equivalent eight bigcore system with hardware-based cache coherence; and 21% better performance and similar energy efficiency compared to a 64-core system (4 big, 60 tiny) with full-system hardware-based cache coherence.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 58549d22-4f05-4f0f-9ad9-4ff87b36fa52Cited by top-tier papers3
- HeteroGen: Automatic Synthesis of Heterogeneous Cache Coherence ProtocolsNicolai Oswald, Vijay Nagarajan, Daniel J. Sorin, Vasilis Gavrielatos et al.HPCA 2022 · 15 citations
- ALTOCUMULUS: Scalable Scheduling for Nanosecond-Scale Remote Procedure CallsJiechen Zhao, Iris Uwizeyimana, Karthik Ganesan, Mark C. Jeffrey et al.MICRO 2022 · 11 citations
- CORD: Low-Latency, Bandwidth-Efficient and Scalable Release Consistency via Directory OrderingYanpeng Yu, Nicolai Oswald, Anurag KhandelwalISCA 2025 · 3 citations
Related papers
- Beyond Static Parallel Loops: Supporting Dynamic Task Parallelism on Manycore Architectures with Software-Managed Scratchpad MemoriesLin Cheng, Max Ruttenberg, Dai Cheol Jung, Dustin Richmond et al.ASPLOS 2023 · 4 citations
- BWoS: Formally Verified Block-based Work Stealing for Parallel ProcessingJiawei Wang, Bohdan Trach, Ming Fu, Diogo Behrens et al.OSDI 2023 · 5 citations
- Symbiotic Task Scheduling and Data PrefetchingGilead Posluns, Mark C. JeffreyMICRO 2025 · 1 citation
- Machine Learning-based Thermally-Safe Cache Contention Mitigation in Clustered ManycoresMohammed Bakr Sikal, Heba Khdr, Martin Rapp, Jörg HenkelDAC 2023 · 7 citations
- WiDir: A Wireless-Enabled Directory Cache Coherence ProtocolAntonio Franques, Apostolos Kokolis, Sergi Abadal, Vimuth Fernando et al.HPCA 2021 · 11 citations
