Tiny but mighty: designing and realizing scalable latency tolerance for manycore SoCs
Marcelo Orenes-Vera, Aninda Manocha, Jonathan Balkind, Fei Gao, Juan L. Aragón, David Wentzlaff, Margaret Martonosi
摘要
Modern computing systems employ significant heterogeneity and specialization to meet performance targets at manageable power. However, memory latency bottlenecks remain problematic, particularly for sparse neural network and graph analytic applications where indirect memory accesses (IMAs) challenge the memory hierarchy.
Decades of prior art have proposed hardware and software mechanisms to mitigate IMA latency, but they fail to analyze real-chip considerations, especially when used in SoCs and manycores. In this paper, we revisit many of these techniques while taking into account manycore integration and verification.
We present the first system implementation of latency tolerance hardware that provides significant speedups without requiring any memory hierarchy or processor tile modifications. This is achieved through a Memory Access Parallel-Load Engine (MAPLE), integrated through the Network-on-Chip (NoC) in a scalable manner. Our hardware-software co-design allows programs to perform longlatency memory accesses asynchronously from the core, avoiding pipeline stalls, and enabling greater memory parallelism (MLP).
In April 2021 we taped out a manycore chip that includes tens of MAPLE instances for efficient data supply. MAPLE demonstrates a full RTL implementation of out-of-core latency-mitigation hardware, with virtual memory support and automated compilation targetting it. This paper evaluates MAPLE integrated with a dualcore FPGA prototype running applications with full SMP Linux, and demonstrates geomean speedups of 2.35× and 2.27× over softwarebased prefetching and decoupling, respectively. Compared to stateof-the-art hardware, it provides geomean speedups of 1.82× and 1.72× over prefetching and decoupling techniques.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- Dalorex: A Data-Local Program Execution and Architecture for Memory-bound ApplicationsMarcelo Orenes-Vera, Esin Tureci, David Wentzlaff, Margaret MartonosiHPCA 2023 · 被引用 23 次
- HotTiles: Accelerating SpMM with Heterogeneous Accelerator ArchitecturesGerasimos Gerogiannis, Sriram Aananthakrishnan, Josep Torrellas, Ibrahim HurHPCA 2024 · 被引用 19 次
- SMAPPIC: Scalable Multi-FPGA Architecture Prototype Platform in the CloudGrigory Chirkov, David WentzlaffASPLOS 2023 · 被引用 13 次
- Cohort: Software-Oriented Acceleration for Heterogeneous SoCsTianrui Wei, Nazerke Turtayeva, Marcelo Orenes-Vera, Omkar Lonkar 等ASPLOS 2023 · 被引用 12 次
- AutoCC: Automatic Discovery of Covert Channels in Time-Shared HardwareMarcelo Orenes-Vera, Hyunsung Yun, Nils Wistoff, Gernot Heiser 等MICRO 2023 · 被引用 7 次
它引用的顶会 Paper5
- Prodigy: Improving the Memory Latency of Data-Indirect Irregular Workloads Using Hardware-Software Co-DesignNishil Talati, Kyle May, Armand Behroozi, Yichen Yang 等HPCA 2021 · 被引用 62 次
- AutoSVA: Democratizing Formal Verification of RTL Module InteractionsMarcelo Orenes-Vera, Aninda Manocha, David Wentzlaff, Margaret MartonosiDAC 2021 · 被引用 30 次
- BYOC: A "Bring Your Own Core" Framework for Heterogeneous-ISA ResearchJonathan Balkind, Katie Lim, Michael Schaffner, Fei Gao 等ASPLOS 2020 · 被引用 29 次
- Pipette: Improving Core Utilization on Irregular Applications through Intra-Core Pipeline ParallelismQuan M. Nguyen, Daniel SánchezMICRO 2020 · 被引用 28 次
- Slipstream Processors Revisited: Exploiting Branch SetsVinesh Srinivasan, Rangeen Basu Roy Chowdhury, Eric RotenbergISCA 2020 · 被引用 12 次
相关 Paper
- Magellan: A High-Performance Loop-Guided Prefetcher for Indirect Memory AccessGelin Fu, Tian Xia, Mingzhuo Yin, Prashant J. Nair 等ISCA 2025 · 被引用 2 次
- Decoupled Vector RunaheadAjeya Naithani, Jaime Roelandts, Sam Ainsworth, Timothy M. Jones 等MICRO 2023 · 被引用 15 次
- Differential-Matching Prefetcher for Indirect Memory AccessGelin Fu, Tian Xia, Zhongpei Luo, Ruiyang Chen 等HPCA 2024 · 被引用 15 次
- RICH Prefetcher: Storing Rich Information in Memory to Trade Capacity and Bandwidth for Latency HidingNingzhi Ai, Wenjian He, Hu He, Jing Xia 等MICRO 2025 · 被引用 2 次
- Hermes: Accelerating Long-Latency Load Requests via Perceptron-Based Off-Chip Load PredictionRahul Bera, Konstantinos Kanellopoulos, Shankar Balachandran, David Novo 等MICRO 2022 · 被引用 37 次
