Efficiently Supporting Dynamic Task Parallelism on Heterogeneous Cache-Coherent Systems
Moyang Wang, Tuan Ta, Lin Cheng, Christopher Batten
摘要
Manycore processors, with tens to hundreds of tiny cores but no hardware-based cache coherence, can offer tremendous peak throughput on highly parallel programs while being complexity and energy efficient. Manycore processors can be combined with a few high-performance big cores for executing operating systems, legacy code, and serial regions. These systems use heterogeneous cache coherence (HCC) with hardware-based cache coherence between big cores and software-centric cache coherence between tiny cores. Unfortunately, programming these heterogeneous cache-coherent systems to enable collaborative execution is challenging, especially when considering dynamic task parallelism. This paper seeks to address this challenge using a combination of light-weight software and hardware techniques. We provide a detailed description of how to implement a work-stealing runtime to enable dynamic task parallelism on heterogeneous cache-coherent systems. We also propose direct task stealing (DTS), a new technique based on user-level interrupts to bypass the memory system and thus improve the performance and energy efficiency of work stealing. Our results demonstrate that executing dynamic task-parallel applications on a 64-core system (4 big, 60 tiny) with complexity-effective HCC and DTS can achieve: speedup over a single big core; speedup over an area-equivalent eight bigcore system with hardware-based cache coherence; and 21% better performance and similar energy efficiency compared to a 64-core system (4 big, 60 tiny) with full-system hardware-based cache coherence.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper3
- HeteroGen: Automatic Synthesis of Heterogeneous Cache Coherence ProtocolsNicolai Oswald, Vijay Nagarajan, Daniel J. Sorin, Vasilis Gavrielatos 等HPCA 2022 · 被引用 15 次
- ALTOCUMULUS: Scalable Scheduling for Nanosecond-Scale Remote Procedure CallsJiechen Zhao, Iris Uwizeyimana, Karthik Ganesan, Mark C. Jeffrey 等MICRO 2022 · 被引用 11 次
- CORD: Low-Latency, Bandwidth-Efficient and Scalable Release Consistency via Directory OrderingYanpeng Yu, Nicolai Oswald, Anurag KhandelwalISCA 2025 · 被引用 3 次
相关 Paper
- Beyond Static Parallel Loops: Supporting Dynamic Task Parallelism on Manycore Architectures with Software-Managed Scratchpad MemoriesLin Cheng, Max Ruttenberg, Dai Cheol Jung, Dustin Richmond 等ASPLOS 2023 · 被引用 4 次
- BWoS: Formally Verified Block-based Work Stealing for Parallel ProcessingJiawei Wang, Bohdan Trach, Ming Fu, Diogo Behrens 等OSDI 2023 · 被引用 5 次
- Symbiotic Task Scheduling and Data PrefetchingGilead Posluns, Mark C. JeffreyMICRO 2025 · 被引用 1 次
- Machine Learning-based Thermally-Safe Cache Contention Mitigation in Clustered ManycoresMohammed Bakr Sikal, Heba Khdr, Martin Rapp, Jörg HenkelDAC 2023 · 被引用 7 次
- WiDir: A Wireless-Enabled Directory Cache Coherence ProtocolAntonio Franques, Apostolos Kokolis, Sergi Abadal, Vimuth Fernando 等HPCA 2021 · 被引用 11 次
