Enhancing Instruction Prefetching via Cache and TLB Management
Alexandre Valentin Jamet, Georgios Vavouliotis, Martí Torrents, Dimitrios Chasapis, Marc Casas
Abstract
Modern server workloads have massive instruction footprints, exacerbating pressure on the processor front-end, and making techniques like L1 instruction (L1I) prefetching essential to alleviate this bottleneck. Although L1I prefetchers deliver significant performance gains, their full potential remains underutilized due to two main factors: i) L1I prefetch requests that cross page boundaries require address translation before being issued, and the latency involved in retrieving these translations undermines the timeliness of the L1I prefetches; ii) the reuse potential of code lines fetched in the cache hierarchy by L1I prefetches is highly variable-while a few lines are accessed multiple times, many are dead-on-arrival.
This paper proposes the Instruction Prefetch Centric Cache and TLB Management (IP-CaT), a microarchitectural scheme orchestrating TLB and cache management to maximize the benefits of L1I prefetching. IP-CaT comprises two modules: i) the translation Prefetch Buffer (tPB), a small buffer located alongside the last-level TLB (sTLB) that accommodates translations fetched by L1I page-cross prefetches to reduce the address translation cost of L1I prefetching and ii) the Trimodal Instruction Prefetch Replacement Policy (TIPRP), a decision-tree based replacement policy for the L2 cache (L2C) specialized in the management of lines fetched by L1I prefetches.
Our evaluation shows that IP-CaT delivers significant performance benefits when integrated with three state-of-the-art L1I prefetchers (EPI [1], FNL+MMA [2], Barc ¸a [3]). For example, IP-CaT+EPI achieves a 6.1% geomean speedup over EPI across a set of 105 contemporary server workloads. We also show that IP-CaT outperforms the state-of-the-art instruction TLB prefetcher [4], the leading TLB replacement policy [5], and codeaware, prefetch-aware, and general-purpose cache replacement policies (Emissary [6], SHiP++ [7], Mockingjay [8]).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a2d22fe4-0cc0-4bd5-8a8e-6cc078b25f93Builds on19
- Berti: an Accurate Local-Delta Data PrefetcherAgustín Navarro-Torres, Biswabandan Panda, Jesús Alastruey-Benedé, Pablo Ibáñez et al.MICRO 2022 · 82 citations
- Beyond malloc efficiency to fleet efficiency: a hugepage-aware memory allocatorA. H. Hunter, Chris Kennelly, Paul Turner, Darryl Gove et al.OSDI 2021 · 51 citations
- Effective Mimicry of Belady's MIN PolicyIshan Shah, Akanksha Jain, Calvin LinHPCA 2022 · 44 citations
- Hermes: Accelerating Long-Latency Load Requests via Perceptron-Based Off-Chip Load PredictionRahul Bera, Konstantinos Kanellopoulos, Shankar Balachandran, David Novo et al.MICRO 2022 · 37 citations
- Exploiting Page Table Locality for Agile TLB PrefetchingGeorgios Vavouliotis, Lluc Alvarez, Vasileios Karakostas, Konstantinos Nikas et al.ISCA 2021 · 34 citations
Related papers
- Instruction-Aware Cooperative TLB and Cache Replacement PoliciesDimitrios Chasapis, Georgios Vavouliotis, Daniel A. Jiménez, Marc CasasASPLOS 2025 · 5 citations
- Morrigan: A Composite Instruction TLB PrefetcherGeorgios Vavouliotis, Lluc Alvarez, Boris Grot, Daniel A. Jiménez et al.MICRO 2021 · 19 citations
- Divide and Conquer Frontend BottleneckAli Ansari, Pejman Lotfi-Kamran, Hamid Sarbazi-AzadISCA 2020 · 28 citations
- Twig: Profile-Guided BTB Prefetching for Data Center ApplicationsTanvir Ahmed Khan, Nathan Brown, Akshitha Sriraman, Niranjan K. Soundararajan et al.MICRO 2021 · 33 citations
- PDIP: Priority Directed Instruction PrefetchingBhargav Reddy Godala, Sankara Prasad Ramesh, Gilles A. Pokam, Jared Stark et al.ASPLOS 2024 · 17 citations
