Exploiting Page Table Locality for Agile TLB Prefetching
Georgios Vavouliotis, Lluc Alvarez, Vasileios Karakostas, Konstantinos Nikas, Nectarios Koziris, Daniel A. Jiménez, Marc Casas
Abstract
Frequent Translation Lookaside Buffer (TLB) misses incur high performance and energy costs due to page walks required for fetching the corresponding address translations. Prefetching page table entries (PTEs) ahead of demand TLB accesses can mitigate the address translation performance bottleneck, but each prefetch requires traversing the page table, triggering additional accesses to the memory hierarchy. Therefore, TLB prefetching is a costly technique that may undermine performance when the prefetches are not accurate.In this paper we exploit the locality in the last level of the page table to reduce the cost and enhance the effectiveness of TLB prefetching by fetching cache-line adjacent PTEs "for free". We propose Sampling-Based Free TLB Prefetching (SBFP), a dynamic scheme that predicts the usefulness of these "free" PTEs and prefetches only the ones most likely to prevent TLB misses. We demonstrate that combining SBFP with novel and state-of-the-art TLB prefetchers significantly improves miss coverage and reduces most memory accesses due to page walks.Moreover, we propose Agile TLB Prefetcher (ATP), a novel composite TLB prefetcher particularly designed to maximize the benefits of SBFP. ATP efficiently combines three low-cost TLB prefetchers and disables TLB prefetching for those execution phases that do not benefit from it. Unlike state-of-the-art TLB prefetchers that correlate patterns with only one feature (e.g., strides, PC, distances), ATP correlates patterns with multiple features and dynamically enables the most appropriate TLB prefetcher per TLB miss.To alleviate the address translation performance bottleneck, we propose a unified solution that combines ATP and SBFP. Across an extensive set of industrial workloads provided by Qualcomm, ATP coupled with SBFP improves geometric speedup by 16.2%, and eliminates on average 37% of the memory references due to page walks. Considering the SPEC CPU 2006 and SPEC CPU 2017 benchmark suites, ATP with SBFP increases geometric speedup by 11.1%, and eliminates page walk memory references by 26%. Applied to big data workloads (GAP suite, XSBench), ATP with SBFP yields a geometric speedup of 11.8% while reducing page walk memory references by 5%. Over the best state-of-the-art TLB prefetcher for each benchmark suite, ATP with SBFP achieves speedups of 8.7%, 3.4%, and 4.2% for the Qualcomm, SPEC, and GAP+XSBench workloads, respectively.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 41569e30-1e82-4cbd-8bd1-e1f36e760649Cited by top-tier papers17
- Contiguitas: The Pursuit of Physical Memory Contiguity in DatacentersKaiyang Zhao, Kaiwen Xue, Ziqi Wang, Dan Schatzberg et al.ISCA 2023 · 26 citations
- Page Size Aware Cache PrefetchingGeorgios Vavouliotis, Gino Chacon, Lluc Alvarez, Paul V. Gratz et al.MICRO 2022 · 24 citations
- Micro-Armed Bandit: Lightweight & Reusable Reinforcement Learning for Microarchitecture Decision-MakingGerasimos Gerogiannis, Josep TorrellasMICRO 2023 · 23 citations
- Morrigan: A Composite Instruction TLB PrefetcherGeorgios Vavouliotis, Lluc Alvarez, Boris Grot, Daniel A. Jiménez et al.MICRO 2021 · 19 citations
- A Two Level Neural Approach Combining Off-Chip Prediction with Adaptive Prefetch FilteringAlexandre Valentin Jamet, Georgios Vavouliotis, Daniel A. Jiménez, Lluc Alvarez et al.HPCA 2024 · 18 citations
Related papers
- Every walk's a hit: making page walks single-access cache hitsChang Hyun Park, Ilias Vougioukas, Andreas Sandberg, David Black-SchafferASPLOS 2022 · 34 citations
- LATPC: Accelerating GPU Address Translation Using Locality-Aware TLB Prefetching and MSHR CompressionYeonan Ha, Jiho Park, Hanna Cha, Jiwon Lee et al.MICRO 2025 · 2 citations
- Divide and Conquer Frontend BottleneckAli Ansari, Pejman Lotfi-Kamran, Hamid Sarbazi-AzadISCA 2020 · 28 citations
- Instruction-Aware Cooperative TLB and Cache Replacement PoliciesDimitrios Chasapis, Georgios Vavouliotis, Daniel A. Jiménez, Marc CasasASPLOS 2025 · 5 citations
- Tailored Page SizesFaruk Guvenilir, Yale N. PattISCA 2020 · 22 citations
