Delegated Replies: Alleviating Network Clogging in Heterogeneous Architectures
Xia Zhao, Lieven Eeckhout, Magnus Jahre
Abstract
Heterogeneous architectures with latency-sensitive CPU cores and bandwidth-intensive accelerators are attractive as they deliver high performance at favorable cost. These architectures typically have significantly more compute cores than memory nodes. The many bandwidth-intensive accelerators hence overwhelm the few memory nodes, resulting in suboptimal accelerator performance — as their bandwidth needs are not met — and poor CPU performance — because memory node blocking creates high latencies. We call this phenomenon network clogging. Since network clogging is a widespread issue in heterogeneous architectures, we first investigate if existing state-of-the-art approaches can address it. We find that the most effective prior approach, called Realistic Probing (RP), is suboptimal because it searches the local caches of other cores for missing data.We propose Delegated Replies which lets memory nodes speculatively delegate the responsibility of replying to last-level cache hits to the private cache that last accessed the requested cache block, hence avoiding the search that fundamentally limits RP. Moreover, Delegated Replies uses the (typically) under-utilized request network for delegation; it is the reply network links of the memory nodes that commonly clog because replies include complete cache blocks in addition to metadata. We evaluate Delegated Replies in the context of heterogeneous architectures with latency-sensitive CPU cores and bandwidth-intensive GPU cores and find that it improves GPU (CPU) performance by 14.2% (5.2%) and 25.7% (8.8%) on average compared to RP and our baseline, respectively.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3ed3d46b-aec9-43c5-b8f5-96fe83d325feBuilds on4
- Locality-Centric Data and Threadblock Management for Massive GPUsMahmoud Khairy, Vadim Nikiforov, David W. Nellans, Timothy G. RogersMICRO 2020 · 38 citations
- Analyzing and Leveraging Decoupled L1 Caches in GPUsMohamed Assem Ibrahim, Onur Kayiran, Yasuko Eckert, Gabriel H. Loh et al.HPCA 2021 · 30 citations
- Selective Replication in Memory-Side GPU CachesXia Zhao, Magnus Jahre, Lieven EeckhoutMICRO 2020 · 15 citations
- EquiNox: Equivalent NoC Injection Routers for Silicon Interposer-Based Throughput ProcessorsYunfan Li, Lizhong ChenHPCA 2020 · 3 citations
Related papers
- Latency-Aware Caching with Delayed Hits: From Bursty Traffic to Pipeline ArchitecturesNadav Keren, Gil Einziger, Gabriel ScalosubNSDI 2026 · 1 citation
- HELM: Characterizing Unified Memory Accesses to Improve GPU Performance under Memory OversubscriptionNathan Jones, Tyler N. Allen, Rong GeSC 2025 · 5 citations
- A Simple Cache Coherence Scheme for Integrated CPU-GPU SystemsArdhi Wiratama Baskara Yudha, Reza Pulungan, Henry Hoffmann, Yan SolihinDAC 2020 · 3 citations
- Orchestrating Data Placement and Query Execution in Heterogeneous CPU-GPU DBMSBobbi W. Yogatama, Weiwei Gong, Xiangyao YuVLDB 2022 · 45 citations
- RELIEF: Relieving Memory Pressure In SoCs Via Data Movement-Aware Accelerator SchedulingSudhanshu Gupta, Sandhya DwarkadasHPCA 2024 · 4 citations
