Tiny but mighty: designing and realizing scalable latency tolerance for manycore SoCs
Marcelo Orenes-Vera, Aninda Manocha, Jonathan Balkind, Fei Gao, Juan L. Aragón, David Wentzlaff, Margaret Martonosi
Abstract
Modern computing systems employ significant heterogeneity and specialization to meet performance targets at manageable power. However, memory latency bottlenecks remain problematic, particularly for sparse neural network and graph analytic applications where indirect memory accesses (IMAs) challenge the memory hierarchy.
Decades of prior art have proposed hardware and software mechanisms to mitigate IMA latency, but they fail to analyze real-chip considerations, especially when used in SoCs and manycores. In this paper, we revisit many of these techniques while taking into account manycore integration and verification.
We present the first system implementation of latency tolerance hardware that provides significant speedups without requiring any memory hierarchy or processor tile modifications. This is achieved through a Memory Access Parallel-Load Engine (MAPLE), integrated through the Network-on-Chip (NoC) in a scalable manner. Our hardware-software co-design allows programs to perform longlatency memory accesses asynchronously from the core, avoiding pipeline stalls, and enabling greater memory parallelism (MLP).
In April 2021 we taped out a manycore chip that includes tens of MAPLE instances for efficient data supply. MAPLE demonstrates a full RTL implementation of out-of-core latency-mitigation hardware, with virtual memory support and automated compilation targetting it. This paper evaluates MAPLE integrated with a dualcore FPGA prototype running applications with full SMP Linux, and demonstrates geomean speedups of 2.35× and 2.27× over softwarebased prefetching and decoupling, respectively. Compared to stateof-the-art hardware, it provides geomean speedups of 1.82× and 1.72× over prefetching and decoupling techniques.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers9
- Dalorex: A Data-Local Program Execution and Architecture for Memory-bound ApplicationsMarcelo Orenes-Vera, Esin Tureci, David Wentzlaff, Margaret MartonosiHPCA 2023 · 23 citations
- HotTiles: Accelerating SpMM with Heterogeneous Accelerator ArchitecturesGerasimos Gerogiannis, Sriram Aananthakrishnan, Josep Torrellas, Ibrahim HurHPCA 2024 · 19 citations
- SMAPPIC: Scalable Multi-FPGA Architecture Prototype Platform in the CloudGrigory Chirkov, David WentzlaffASPLOS 2023 · 13 citations
- Cohort: Software-Oriented Acceleration for Heterogeneous SoCsTianrui Wei, Nazerke Turtayeva, Marcelo Orenes-Vera, Omkar Lonkar et al.ASPLOS 2023 · 12 citations
- AutoCC: Automatic Discovery of Covert Channels in Time-Shared HardwareMarcelo Orenes-Vera, Hyunsung Yun, Nils Wistoff, Gernot Heiser et al.MICRO 2023 · 7 citations
Builds on5
- Prodigy: Improving the Memory Latency of Data-Indirect Irregular Workloads Using Hardware-Software Co-DesignNishil Talati, Kyle May, Armand Behroozi, Yichen Yang et al.HPCA 2021 · 62 citations
- AutoSVA: Democratizing Formal Verification of RTL Module InteractionsMarcelo Orenes-Vera, Aninda Manocha, David Wentzlaff, Margaret MartonosiDAC 2021 · 30 citations
- BYOC: A "Bring Your Own Core" Framework for Heterogeneous-ISA ResearchJonathan Balkind, Katie Lim, Michael Schaffner, Fei Gao et al.ASPLOS 2020 · 29 citations
- Pipette: Improving Core Utilization on Irregular Applications through Intra-Core Pipeline ParallelismQuan M. Nguyen, Daniel SánchezMICRO 2020 · 28 citations
- Slipstream Processors Revisited: Exploiting Branch SetsVinesh Srinivasan, Rangeen Basu Roy Chowdhury, Eric RotenbergISCA 2020 · 12 citations
Related papers
- Magellan: A High-Performance Loop-Guided Prefetcher for Indirect Memory AccessGelin Fu, Tian Xia, Mingzhuo Yin, Prashant J. Nair et al.ISCA 2025 · 2 citations
- Decoupled Vector RunaheadAjeya Naithani, Jaime Roelandts, Sam Ainsworth, Timothy M. Jones et al.MICRO 2023 · 15 citations
- Differential-Matching Prefetcher for Indirect Memory AccessGelin Fu, Tian Xia, Zhongpei Luo, Ruiyang Chen et al.HPCA 2024 · 15 citations
- RICH Prefetcher: Storing Rich Information in Memory to Trade Capacity and Bandwidth for Latency HidingNingzhi Ai, Wenjian He, Hu He, Jing Xia et al.MICRO 2025 · 2 citations
- Hermes: Accelerating Long-Latency Load Requests via Perceptron-Based Off-Chip Load PredictionRahul Bera, Konstantinos Kanellopoulos, Shankar Balachandran, David Novo et al.MICRO 2022 · 37 citations
