Integrating Prefetcher Selection with Dynamic Request Allocation Improves Prefetching Efficiency
Mengming Li, Qijun Zhang, Yongqing Ren, Zhiyao Xie
摘要
Hardware prefetching plays a critical role in hiding the off-chip DRAM latency. The complexity of applications results in a wide variety of memory access patterns, prompting the development of numerous cache-prefetching algorithms. Consequently, commercial processors often employ a hybrid of these algorithms to enhance the overall prefetching performance. Nonetheless, since these prefetchers share hardware resources, conflicts arising from competing prefetching requests can negate the benefits of hardware prefetching. Under such circumstances, several prefetcher selection algorithms have been proposed to mitigate conflicts between prefetchers. However, these prior solutions suffer from two limitations. First, the input demand request allocation is inaccurate. Second, the prefetcher selection criteria are coarse-grained. In this paper, we address both limitations by introducing an efficient and widely applicable prefetcher selection algorithm—Alecto1, which tailors the demand requests for each prefetcher. Every demand request is first sent to Alecto to identify suitable prefetchers before being routed to prefetchers for training and prefetching. Our analysis shows that Alecto is adept at not only harmonizing prefetching accuracy, coverage, and timeliness but also significantly enhancing the utilization of the prefetcher table, which is vital for temporal prefetching. Alecto outperforms the state-of-the-art RL-based prefetcher selection algorithm—Bandit by in single-core, and in eight-core. For memory-intensive benchmarks, Alecto outperforms Bandit by . Alecto consistently delivers state-of-the-art performance in scheduling various types of cache prefetchers. In addition to the performance improvement, Alecto can reduce the energy consumption associated with accessing the prefetchers’ table by ( energy reduction on the entire memory hierarchy), while only adding less than 1 KB of storage overhead.1The name Alecto stands for the combination of selection and allocation.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Profile-Guided Temporal PrefetchingMengming Li, Qijun Zhang, Yichuan Gao, Wenji Fang 等ISCA 2025 · 被引用 4 次
- ICP: Exploiting Instruction Correlation for Prefetching Irregular Memory AccessesMengming Li, Chenlu Miao, Buqing Xu, Qijun Zhang 等ISCA 2026
- PF-LLM: Large Language Model Hinted Hardware PrefetchingCeyu Xu, Xiangfeng Sun, Weihang Li, Chen Bai 等ASPLOS 2026
它引用的顶会 Paper9
- Bouquet of Instruction Pointers: Instruction Pointer Classifier-based Spatial Hardware PrefetchingSamuel Pakalapati, Biswabandan PandaISCA 2020 · 被引用 97 次
- Pythia: A Customizable Hardware Prefetching Framework Using Online Reinforcement LearningRahul Bera, Konstantinos Kanellopoulos, Anant Nori, Taha Shahroodi 等MICRO 2021 · 被引用 95 次
- Classifying Memory Access Patterns for PrefetchingGrant Ayers, Heiner Litz, Christos Kozyrakis, Parthasarathy RanganathanASPLOS 2020 · 被引用 83 次
- Berti: an Accurate Local-Delta Data PrefetcherAgustín Navarro-Torres, Biswabandan Panda, Jesús Alastruey-Benedé, Pablo Ibáñez 等MICRO 2022 · 被引用 82 次
- Merging Similar Patterns for Hardware PrefetchingShizhi Jiang, Qiusong Yang, Yiwei CiMICRO 2022 · 被引用 27 次
相关 Paper
- I-POP: Ignite Positive PrefetchersYiquan Lin, Wenhai Lin, Yiquan Chen, Jiexiong Xu 等HPCA 2026 · 被引用 1 次
- CLIP: Load Criticality based Data Prefetching for Bandwidth-constrained Many-core SystemsBiswabandan PandaMICRO 2023 · 被引用 21 次
- ReSemble: Reinforced Ensemble Framework for Data PrefetchingPengmiao Zhang, Rajgopal Kannan, Ajitesh Srivastava, Anant V. Nori 等SC 2022 · 被引用 17 次
- A Cost-Effective Entangling Prefetcher for InstructionsAlberto Ros, Alexandra JimboreanISCA 2021 · 被引用 31 次
- A New Formulation of Neural Data PrefetchingQuang Duong, Akanksha Jain, Calvin LinISCA 2024 · 被引用 16 次
