Hermes: Accelerating Long-Latency Load Requests via Perceptron-Based Off-Chip Load Prediction
Rahul Bera, Konstantinos Kanellopoulos, Shankar Balachandran, David Novo, Ataberk Olgun, Mohammad Sadrosadati, Onur Mutlu
Abstract
Long-latency load requests continue to limit the performance of modern high-performance processors. To increase the latency tolerance of a processor, architects have primarily relied on two key techniques: sophisticated data prefetchers and large on-chip caches. In this work, we show that: (1) even a sophisticated state-of-the-art prefetcher can only predict half of the off-chip load requests on average across a wide range of workloads, and (2) due to the increasing size and complexity of on-chip caches, a large fraction of the latency of an off-chip load request is spent accessing the on-chip cache hierarchy to solely determine that it needs to go off-chip. The goal of this work is to accelerate off-chip load requests by removing the on-chip cache access latency from their critical path. To this end, we propose a new technique called Hermes, whose key idea is to: (1) accurately predict which load requests might go off-chip, and (2) speculatively fetch the data required by the predicted off-chip loads directly from the main memory, while also concurrently accessing the cache hierarchy for such loads. To enable Hermes, we develop a new lightweight, perceptron-based off-chip load prediction technique that learns to identify off-chip load requests using multiple program features (e.g., sequence of program counters, byte offset of a load request). For every load request generated by the processor, the predictor observes a set of program features to predict whether or not the load would go off-chip. If the load is predicted to go off-chip, Hermes issues a speculative load request directly to the main memory controller once the load’s physical address is generated. If the prediction is correct, the load eventually misses the cache hierarchy and waits for the ongoing speculative load request to finish, and thus Hermes completely hides the on-chip cache hierarchy access latency from the critical path of the correctly-predicted off-chip load. Our extensive evaluation using a wide range of workloads shows that Hermes provides consistent performance improvement on top of a state-of-the-art baseline system across a wide range of configurations with varying core count, main memory bandwidth, high-performance data prefetchers, and on-chip cache hierarchy access latencies, while incurring only modest storage overhead. The source code of Hermes is freely available at: https://github.com/CMU-SAFARI/Hermes.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers18
- Micro-Armed Bandit: Lightweight & Reusable Reinforcement Learning for Microarchitecture Decision-MakingGerasimos Gerogiannis, Josep TorrellasMICRO 2023 · 23 citations
- A Two Level Neural Approach Combining Off-Chip Prediction with Adaptive Prefetch FilteringAlexandre Valentin Jamet, Georgios Vavouliotis, Daniel A. Jiménez, Lluc Alvarez et al.HPCA 2024 · 18 citations
- A New Formulation of Neural Data PrefetchingQuang Duong, Akanksha Jain, Calvin LinISCA 2024 · 16 citations
- Decoupled Vector RunaheadAjeya Naithani, Jaime Roelandts, Sam Ainsworth, Timothy M. Jones et al.MICRO 2023 · 15 citations
- Compiler-Directed Whole-System PersistenceJianping Zeng, Tong Zhang, Changhee JungISCA 2024 · 13 citations
Builds on6
- Bouquet of Instruction Pointers: Instruction Pointer Classifier-based Spatial Hardware PrefetchingSamuel Pakalapati, Biswabandan PandaISCA 2020 · 97 citations
- Pythia: A Customizable Hardware Prefetching Framework Using Online Reinforcement LearningRahul Bera, Konstantinos Kanellopoulos, Anant Nori, Taha Shahroodi et al.MICRO 2021 · 95 citations
- CRISP: critical slice prefetchingHeiner Litz, Grant Ayers, Parthasarathy RanganathanASPLOS 2022 · 33 citations
- Stream Floating: Enabling Proactive and Decentralized Cache OptimizationsZhengrong Wang, Jian Weng, Jason Lowe-Power, Jayesh Gaur et al.HPCA 2021 · 27 citations
- Vector RunaheadAjeya Naithani, Sam Ainsworth, Timothy M. Jones, Lieven EeckhoutISCA 2021 · 27 citations
Related papers
- Reducing Load Latency with Cache Level PredictionMajid Jalili, Mattan ErezHPCA 2022 · 17 citations
- Athena: Synergizing Data Prefetching and Off-Chip Prediction via Online Reinforcement LearningRahul Bera, Zhenrong Lang, Caroline Hengartner, Konstantinos Kanellopoulos et al.HPCA 2026
- PF-LLM: Large Language Model Hinted Hardware PrefetchingCeyu Xu, Xiangfeng Sun, Weihang Li, Chen Bai et al.ASPLOS 2026
- CLIP: Load Criticality based Data Prefetching for Bandwidth-constrained Many-core SystemsBiswabandan PandaMICRO 2023 · 21 citations
- APT-GET: profile-guided timely software prefetchingSaba Jamilan, Tanvir Ahmed Khan, Grant Ayers, Baris Kasikci et al.EuroSys 2022 · 25 citations
