Decoupled Vector Runahead
Ajeya Naithani, Jaime Roelandts, Sam Ainsworth, Timothy M. Jones, Lieven Eeckhout
摘要
We present Decoupled Vector Runahead (DVR), an in-core prefetching technique, executing separately to the main application thread, that exploits massive amounts of memory-level parallelism to improve the performance of applications featuring indirect memory accesses. DVR dynamically infers loop bounds at run-time, recognizing striding loads, and vectorizing subsequent instructions that are part of an indirect chain. It proactively issues memory accesses for the resulting loads far into the future, even when the out-of-order core has not yet stalled, bringing their data into the L1 cache, and thus providing timely prefetches for the main thread. DVR can adjust the degree of vectorization at run-time, vectorize the same chain of indirect memory accesses across multiple invocations of an inner loop, and efficiently handle branch divergence along the vectorized chain. DVR runs as an on-demand, speculative, in-order, lightweight hardware subthread alongside the main thread within the core and incurs a minimal hardware overhead of only 1139 bytes. Relative to a large superscalar 5-wide out-of-order baseline and Vector Runahead — a recent microarchitectural technique to accelerate indirect memory accesses on out-of-order processors — DVR delivers 2.4 × and 2 × higher performance, respectively, for a set of graph analytics, database, and HPC workloads.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Scalar Vector RunaheadJaime Roelandts, Ajeya Naithani, Sam Ainsworth, Timothy M. Jones 等MICRO 2024 · 被引用 11 次
- Constable: Improving Performance and Power Efficiency by Safely Eliminating Load Instruction ExecutionRahul Bera, Adithya Ranganathan, Joydeep Rakshit, Sujit Mahto 等ISCA 2024 · 被引用 8 次
- DX100: Programmable Data Access Accelerator for IndirectionAlireza Khadem, Kamalavasan Kamalakkannan, Zhenyan Zhu, Akash Poptani 等ISCA 2025 · 被引用 2 次
- NVR: Vector Runahead on NPUs for Sparse Memory AccessHui Wang, Zhengpeng Zhao, Jing Wang, Yushu Du 等DAC 2025 · 被引用 1 次
- ICP: Exploiting Instruction Correlation for Prefetching Irregular Memory AccessesMengming Li, Chenlu Miao, Buqing Xu, Qijun Zhang 等ISCA 2026
它引用的顶会 Paper14
- A hierarchical neural model of data prefetchingZhan Shi, Akanksha Jain, Kevin Swersky, Milad Hashemi 等ASPLOS 2021 · 被引用 100 次
- Bouquet of Instruction Pointers: Instruction Pointer Classifier-based Spatial Hardware PrefetchingSamuel Pakalapati, Biswabandan PandaISCA 2020 · 被引用 97 次
- Pythia: A Customizable Hardware Prefetching Framework Using Online Reinforcement LearningRahul Bera, Konstantinos Kanellopoulos, Anant Nori, Taha Shahroodi 等MICRO 2021 · 被引用 95 次
- Classifying Memory Access Patterns for PrefetchingGrant Ayers, Heiner Litz, Christos Kozyrakis, Parthasarathy RanganathanASPLOS 2020 · 被引用 83 次
- Berti: an Accurate Local-Delta Data PrefetcherAgustín Navarro-Torres, Biswabandan Panda, Jesús Alastruey-Benedé, Pablo Ibáñez 等MICRO 2022 · 被引用 82 次
相关 Paper
- Vector RunaheadAjeya Naithani, Sam Ainsworth, Timothy M. Jones, Lieven EeckhoutISCA 2021 · 被引用 27 次
- Reducing Load Latency with Cache Level PredictionMajid Jalili, Mattan ErezHPCA 2022 · 被引用 17 次
- Differential-Matching Prefetcher for Indirect Memory AccessGelin Fu, Tian Xia, Zhongpei Luo, Ruiyang Chen 等HPCA 2024 · 被引用 15 次
- Hermes: Accelerating Long-Latency Load Requests via Perceptron-Based Off-Chip Load PredictionRahul Bera, Konstantinos Kanellopoulos, Shankar Balachandran, David Novo 等MICRO 2022 · 被引用 37 次
- Interleaved Multi-VectorizingZhuhe Fang, Beilei Zheng, Chuliang WengVLDB 2020 · 被引用 19 次
