Every walk's a hit: making page walks single-access cache hits
Chang Hyun Park, Ilias Vougioukas, Andreas Sandberg, David Black-Schaffer
Abstract
As memory capacity has outstripped TLB coverage, large data applications suffer from frequent page table walks. We investigate two complementary techniques for addressing this cost: reducing the number of accesses required and reducing the latency of each access. The first approach is accomplished by opportunistically "flattening" the page table: merging two levels of traditional 4 KB page table nodes into a single 2 MB node, thereby reducing the table's depth and the number of indirections required to traverse it. The second is accomplished by biasing the cache replacement algorithm to keep page table entries during periods of high TLB miss rates, as these periods also see high data miss rates and are therefore more likely to benefit from having the smaller page table in the cache than to suffer from increased data cache misses.
We evaluate these approaches for both native and virtualized systems and across a range of realistic memory fragmentation scenarios, describe the limited changes needed in our kernel implementation and hardware design, identify and address challenges related to self-referencing page tables and kernel memory allocation, and compare results across server and mobile systems using both academic and industrial simulators for robustness.
We find that flattening does reduce the number of accesses required on a page walk (to 1.0), but its performance impact (+2.3%) is small due to Page Walker Caches (already 1.5 accesses). Prioritizing caching has a larger effect (+6.8%), and the combination improves performance by +9.2%. Flattening is more effective on virtualized systems (4.4 to 2.8 accesses, +7.1% performance), due to 2D page walks. By combining the two techniques we demonstrate a state-ofthe-art +14.0% performance gain and -8.7% dynamic cache energy and -4.7% dynamic DRAM energy for virtualized execution with very simple hardware and software changes.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 64b5b7a4-8f7f-417f-a95e-b04ccb3165d6Cited by top-tier papers20
- Barre Chord: Efficient Virtual Memory Translation for Multi-Chip-Module GPUsYuan Feng, Seonjin Na, Hyesoon Kim, Hyeran JeonISCA 2024 · 20 citations
- Victima: Drastically Increasing Address Translation Reach by Leveraging Underutilized Cache ResourcesKonstantinos Kanellopoulos, Hong Chul Nam, Nisa Bostanci, Rahul Bera et al.MICRO 2023 · 16 citations
- IDYLL: Enhancing Page Translation in Multi-GPUs via Light Weight PTE InvalidationsBingyao Li, Yanan Guo, Yueqi Wang, Aamer Jaleel et al.MICRO 2023 · 16 citations
- Utopia: Fast and Efficient Address Translation via Hybrid Restrictive & Flexible Virtual-to-Physical Address MappingsKonstantinos Kanellopoulos, Rahul Bera, Kosta Stojiljkovic, F. Nisa Bostanci et al.MICRO 2023 · 15 citations
- Translation-optimized Memory Compression for CapacityGagandeep Panwar, Muhammad Laghari, David Bears, Yuqing Liu et al.MICRO 2022 · 12 citations
Builds on8
- Mitosis: Transparently Self-Replicating Page-Tables for Large-Memory MachinesReto Achermann, Ashish Panwar, Abhishek Bhattacharjee, Timothy Roscoe et al.ASPLOS 2020 · 62 citations
- Elastic Cuckoo Page Tables: Rethinking Virtual Memory Translation for ParallelismDimitrios Skarlatos, Apostolos Kokolis, Tianyin Xu, Josep TorrellasASPLOS 2020 · 55 citations
- Enhancing and Exploiting Contiguity for Fast Memory VirtualizationChloe Alverti, Stratos Psomadakis, Vasileios Karakostas, Jayneel Gandhi et al.ISCA 2020 · 41 citations
- A Comprehensive Analysis of Superpage Management Mechanisms and PoliciesWeixi Zhu, Alan L. Cox, Scott RixnerUSENIX ATC 2020 · 40 citations
- Perforated Page: Supporting Fragmented Memory Allocation for Large PagesChang Hyun Park, Sanghoon Cha, Bokyeong Kim, Youngjin Kwon et al.ISCA 2020 · 35 citations
Related papers
- Exploiting Page Table Locality for Agile TLB PrefetchingGeorgios Vavouliotis, Lluc Alvarez, Vasileios Karakostas, Konstantinos Nikas et al.ISCA 2021 · 34 citations
- Tailored Page SizesFaruk Guvenilir, Yale N. PattISCA 2020 · 22 citations
- PTEMagnet: fine-grained physical memory reservation for faster page walks in public cloudsArtemiy Margaritov, Dmitrii Ustiugov, Amna Shahab, Boris GrotASPLOS 2021 · 19 citations
- Fast local page-tables for virtualized NUMA servers with vMitosisAshish Panwar, Reto Achermann, Arkaprava Basu, Abhishek Bhattacharjee et al.ASPLOS 2021 · 29 citations
- Parallel virtualized memory translation with nested elastic cuckoo page tablesJovan Stojkovic, Dimitrios Skarlatos, Apostolos Kokolis, Tianyin Xu et al.ASPLOS 2022 · 15 citations
