Every walk's a hit: making page walks single-access cache hits
Chang Hyun Park, Ilias Vougioukas, Andreas Sandberg, David Black-Schaffer
摘要
As memory capacity has outstripped TLB coverage, large data applications suffer from frequent page table walks. We investigate two complementary techniques for addressing this cost: reducing the number of accesses required and reducing the latency of each access. The first approach is accomplished by opportunistically "flattening" the page table: merging two levels of traditional 4 KB page table nodes into a single 2 MB node, thereby reducing the table's depth and the number of indirections required to traverse it. The second is accomplished by biasing the cache replacement algorithm to keep page table entries during periods of high TLB miss rates, as these periods also see high data miss rates and are therefore more likely to benefit from having the smaller page table in the cache than to suffer from increased data cache misses.
We evaluate these approaches for both native and virtualized systems and across a range of realistic memory fragmentation scenarios, describe the limited changes needed in our kernel implementation and hardware design, identify and address challenges related to self-referencing page tables and kernel memory allocation, and compare results across server and mobile systems using both academic and industrial simulators for robustness.
We find that flattening does reduce the number of accesses required on a page walk (to 1.0), but its performance impact (+2.3%) is small due to Page Walker Caches (already 1.5 accesses). Prioritizing caching has a larger effect (+6.8%), and the combination improves performance by +9.2%. Flattening is more effective on virtualized systems (4.4 to 2.8 accesses, +7.1% performance), due to 2D page walks. By combining the two techniques we demonstrate a state-ofthe-art +14.0% performance gain and -8.7% dynamic cache energy and -4.7% dynamic DRAM energy for virtualized execution with very simple hardware and software changes.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper20
- Barre Chord: Efficient Virtual Memory Translation for Multi-Chip-Module GPUsYuan Feng, Seonjin Na, Hyesoon Kim, Hyeran JeonISCA 2024 · 被引用 20 次
- Victima: Drastically Increasing Address Translation Reach by Leveraging Underutilized Cache ResourcesKonstantinos Kanellopoulos, Hong Chul Nam, Nisa Bostanci, Rahul Bera 等MICRO 2023 · 被引用 16 次
- IDYLL: Enhancing Page Translation in Multi-GPUs via Light Weight PTE InvalidationsBingyao Li, Yanan Guo, Yueqi Wang, Aamer Jaleel 等MICRO 2023 · 被引用 16 次
- Utopia: Fast and Efficient Address Translation via Hybrid Restrictive & Flexible Virtual-to-Physical Address MappingsKonstantinos Kanellopoulos, Rahul Bera, Kosta Stojiljkovic, F. Nisa Bostanci 等MICRO 2023 · 被引用 15 次
- Translation-optimized Memory Compression for CapacityGagandeep Panwar, Muhammad Laghari, David Bears, Yuqing Liu 等MICRO 2022 · 被引用 12 次
它引用的顶会 Paper8
- Mitosis: Transparently Self-Replicating Page-Tables for Large-Memory MachinesReto Achermann, Ashish Panwar, Abhishek Bhattacharjee, Timothy Roscoe 等ASPLOS 2020 · 被引用 62 次
- Elastic Cuckoo Page Tables: Rethinking Virtual Memory Translation for ParallelismDimitrios Skarlatos, Apostolos Kokolis, Tianyin Xu, Josep TorrellasASPLOS 2020 · 被引用 55 次
- Enhancing and Exploiting Contiguity for Fast Memory VirtualizationChloe Alverti, Stratos Psomadakis, Vasileios Karakostas, Jayneel Gandhi 等ISCA 2020 · 被引用 41 次
- A Comprehensive Analysis of Superpage Management Mechanisms and PoliciesWeixi Zhu, Alan L. Cox, Scott RixnerUSENIX ATC 2020 · 被引用 40 次
- Perforated Page: Supporting Fragmented Memory Allocation for Large PagesChang Hyun Park, Sanghoon Cha, Bokyeong Kim, Youngjin Kwon 等ISCA 2020 · 被引用 35 次
相关 Paper
- Exploiting Page Table Locality for Agile TLB PrefetchingGeorgios Vavouliotis, Lluc Alvarez, Vasileios Karakostas, Konstantinos Nikas 等ISCA 2021 · 被引用 34 次
- Tailored Page SizesFaruk Guvenilir, Yale N. PattISCA 2020 · 被引用 22 次
- PTEMagnet: fine-grained physical memory reservation for faster page walks in public cloudsArtemiy Margaritov, Dmitrii Ustiugov, Amna Shahab, Boris GrotASPLOS 2021 · 被引用 19 次
- Fast local page-tables for virtualized NUMA servers with vMitosisAshish Panwar, Reto Achermann, Arkaprava Basu, Abhishek Bhattacharjee 等ASPLOS 2021 · 被引用 29 次
- Parallel virtualized memory translation with nested elastic cuckoo page tablesJovan Stojkovic, Dimitrios Skarlatos, Apostolos Kokolis, Tianyin Xu 等ASPLOS 2022 · 被引用 15 次
