Athena: Synergizing Data Prefetching and Off-Chip Prediction via Online Reinforcement Learning
Rahul Bera, Zhenrong Lang, Caroline Hengartner, Konstantinos Kanellopoulos, Rakesh Kumar, Mohammad Sadrosadati, Onur Mutlu
Abstract
Prefetching and off-chip prediction are two techniques proposed to hide long memory access latencies in high-performance processors. In this work, we demonstrate that: (1) prefetching and off-chip prediction often provide complementary performance benefits, yet (2) naively combining these two mechanisms often fails to realize their full performance potential, and (3) existing prefetcher control policies (both heuristic-and learning-based) leave significant room for performance improvement behind.
Our goal is to design a holistic framework that can autonomously learn to coordinate an off-chip predictor with multiple prefetchers employed at various cache levels, delivering consistent performance benefits across a wide range of workloads and system configurations.
To this end, we propose a new technique called Athena, which models the coordination between prefetchers and off-chip predictor (OCP) as a reinforcement learning (RL) problem. Athena acts as the RL agent that observes multiple system-level features (e.g., prefetcher/OCP accuracy, bandwidth usage) over an epoch of program execution, and uses them as state information to select a coordination action (i.e., enabling the prefetcher and/or OCP, and adjusting prefetcher aggressiveness). At the end of every epoch, Athena receives a numerical reward that measures the change in multiple system-level metrics (e.g., number of cycles taken to execute an epoch). Athena uses this reward to autonomously and continuously learn a policy to coordinate prefetchers with OCP.
Athena makes a key observation that using performance improvement as the sole RL reward, as used in prior work, is unreliable, as it confounds the effects of the agent's actions with inherent variations in workload behavior. To address this limitation, Athena introduces a composite reward framework that separates (1) system-level metrics directly influenced by Athena's actions (e.g., last-level cache misses) from ( 2) metrics primarily driven by workload phase changes (e.g., number of mispredicted branches). This allows Athena to autonomously learn a coordination policy by isolating the true impact of its actions from inherent variations in workload behavior.
Our extensive evaluation using a diverse set of 100 memoryintensive workloads shows that Athena consistently outperforms multiple prior state-of-the-art coordination policies across a wide range of system configurations with various combinations of underlying prefetchers at various cache levels, OCPs, and main memory bandwidths, while incurring only modest storage overhead and design complexity. The source code of Athena is freely available at https://github.com/CMU-SAFARI/Athena.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c7f83e2d-b170-4d44-bb2c-4de5d4e7c329Builds on30
- ConfuciuX: Autonomous Hardware Resource Assignment for DNN Accelerators using Reinforcement LearningSheng-Chun Kao, Geonhwa Jeong, Tushar KrishnaMICRO 2020 · 115 citations
- An Imitation Learning Approach for Cache ReplacementEvan Zheran Liu, Milad Hashemi, Kevin Swersky, Parthasarathy Ranganathan et al.ICML 2020 · 108 citations
- Bouquet of Instruction Pointers: Instruction Pointer Classifier-based Spatial Hardware PrefetchingSamuel Pakalapati, Biswabandan PandaISCA 2020 · 97 citations
- Pythia: A Customizable Hardware Prefetching Framework Using Online Reinforcement LearningRahul Bera, Konstantinos Kanellopoulos, Anant Nori, Taha Shahroodi et al.MICRO 2021 · 95 citations
- Berti: an Accurate Local-Delta Data PrefetcherAgustín Navarro-Torres, Biswabandan Panda, Jesús Alastruey-Benedé, Pablo Ibáñez et al.MICRO 2022 · 82 citations
Related papers
- Hermes: Accelerating Long-Latency Load Requests via Perceptron-Based Off-Chip Load PredictionRahul Bera, Konstantinos Kanellopoulos, Shankar Balachandran, David Novo et al.MICRO 2022 · 37 citations
- I-POP: Ignite Positive PrefetchersYiquan Lin, Wenhai Lin, Yiquan Chen, Jiexiong Xu et al.HPCA 2026 · 1 citation
- ReSemble: Reinforced Ensemble Framework for Data PrefetchingPengmiao Zhang, Rajgopal Kannan, Ajitesh Srivastava, Anant V. Nori et al.SC 2022 · 17 citations
- CHROME: Concurrency-Aware Holistic Cache Management Framework with Online Reinforcement LearningXiaoyang Lu, Hamed Najafi, Jason Liu, Xian-He SunHPCA 2024 · 26 citations
- Micro-MAMA: Multi-Agent Reinforcement Learning for Multicore PrefetchingCharles Block, Gerasimos Gerogiannis, Josep TorrellasMICRO 2025 · 6 citations
