Vector Runahead
Ajeya Naithani, Sam Ainsworth, Timothy M. Jones, Lieven Eeckhout
摘要
The memory wall places a significant limit on performance for many modern workloads. These applications feature complex chains of dependent, indirect memory accesses, which cannot be picked up by even the most advanced microar-chitectural prefetchers. The result is that current out-of-order superscalar processors spend the majority of their time stalled. While it is possible to build special-purpose architectures to exploit the fundamental memory-level parallelism, a microarchi-tectural technique to automatically improve their performance in conventional processors has remained elusive.Runahead execution is a tempting proposition for hiding latency in program execution. However, to achieve high memory-level parallelism, a standard runahead execution skips ahead of cache misses. In modern workloads, this means it only prefetches the first cache-missing load in each dependent chain. We argue that this is not a fundamental limitation. If runahead were instead to stall on cache misses to generate dependent chain loads, then it could regain performance if it could stall on many at once. With this insight, we present Vector Runahead, a technique that prefetches entire load chains and speculatively reorders scalar operations from multiple loop iterations into vector format to bring in many independent loads at once. Vectorization of the runahead instruction stream increases the effective fetch/decode bandwidth with reduced resource requirements, to achieve high degrees of memory-level parallelism at a much faster rate. Across a variety of memory-latency-bound indirect workloads, Vector Runahead achieves a 1.79× performance speedup on a large out-of-order superscalar system, significantly improving on state-of-the-art techniques.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper12
- Crescent: taming memory irregularities for accelerating deep point cloud analyticsYu Feng, Gunnar Hammonds, Yiming Gan, Yuhao ZhuISCA 2022 · 被引用 44 次
- Hermes: Accelerating Long-Latency Load Requests via Perceptron-Based Off-Chip Load PredictionRahul Bera, Konstantinos Kanellopoulos, Shankar Balachandran, David Novo 等MICRO 2022 · 被引用 37 次
- Decoupled Vector RunaheadAjeya Naithani, Jaime Roelandts, Sam Ainsworth, Timothy M. Jones 等MICRO 2023 · 被引用 15 次
- Scalar Vector RunaheadJaime Roelandts, Ajeya Naithani, Sam Ainsworth, Timothy M. Jones 等MICRO 2024 · 被引用 11 次
- Constable: Improving Performance and Power Efficiency by Safely Eliminating Load Instruction ExecutionRahul Bera, Adithya Ranganathan, Joydeep Rakshit, Sujit Mahto 等ISCA 2024 · 被引用 8 次
它引用的顶会 Paper6
- Bouquet of Instruction Pointers: Instruction Pointer Classifier-based Spatial Hardware PrefetchingSamuel Pakalapati, Biswabandan PandaISCA 2020 · 被引用 97 次
- Classifying Memory Access Patterns for PrefetchingGrant Ayers, Heiner Litz, Christos Kozyrakis, Parthasarathy RanganathanASPLOS 2020 · 被引用 83 次
- Livia: Data-Centric Computing Throughout the Memory HierarchyElliot Lockerman, Axel Feldmann, Mohammad Bakhshalipour, Alexandru Stanescu 等ASPLOS 2020 · 被引用 55 次
- Precise Runahead ExecutionAjeya Naithani, Josué Feliu, Almutaz Adileh, Lieven EeckhoutHPCA 2020 · 被引用 32 次
- Pipette: Improving Core Utilization on Irregular Applications through Intra-Core Pipeline ParallelismQuan M. Nguyen, Daniel SánchezMICRO 2020 · 被引用 28 次
相关 Paper
- Register file prefetchingSudhanshu Shukla, Sumeet Bandishte, Jayesh Gaur, Sreenivas SubramoneyISCA 2022 · 被引用 8 次
- Interleaved Multi-VectorizingZhuhe Fang, Beilei Zheng, Chuliang WengVLDB 2020 · 被引用 19 次
- CRISP: critical slice prefetchingHeiner Litz, Grant Ayers, Parthasarathy RanganathanASPLOS 2022 · 被引用 33 次
- Reliability-Aware RunaheadAjeya Naithani, Lieven EeckhoutHPCA 2022 · 被引用 2 次
- Reducing Load Latency with Cache Level PredictionMajid Jalili, Mattan ErezHPCA 2022 · 被引用 17 次
