Architectural Support for Optimizing Huge Page Selection Within the OS
Aninda Manocha, Zi Yan, Esin Tureci, Juan L. Aragón, David W. Nellans, Margaret Martonosi
摘要
Irregular, memory-intensive applications often incur high translation lookaside bu!er (TLB) miss rates that result in signi"cant address translation overheads. Employing huge pages is an e!ective way to reduce these overheads, however in real systems the number of available huge pages can be limited when system memory is nearly full and/or fragmented. Thus, huge pages must be used selectively to back application memory. This work demonstrates that choosing memory regions that incur the most TLB misses for huge page promotion best reduces address translation overheads. We call these regions High reUse TLB-sensitive data (HUBs). Unlike prior work which relies on expensive per-page software counters to identify promotion regions, we propose new architectural support to identify these regions dynamically at application runtime.
We propose a promotion candidate cache (PCC) that identi"es HUB candidates based on hardware page table walks after a lastlevel TLB miss. This small, "xed-size structure tracks huge pagealigned regions (consisting of 𝐿 base pages), ranks them based on observed page table walk frequency, and only keeps the most frequently accessed ones. Evaluated on applications of various memory intensity, our approach successfully identi"es application pages incurring the highest address translation overheads. Our approach demonstrates that with the help of a PCC, the OS only needs to promote 4% of the application footprint to achieve more than 75% of the peak achievable performance, yielding 1.19-1.33→ speedups over 4KB base pages alone. In real systems where memory is typically fragmented, the PCC outperforms Linux's page promotion policy by 14% (when 50% of total memory is fragmented) and 16% (when 90% of total memory is fragmented) respectively.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper8
- A Comprehensive Analysis of Superpage Management Mechanisms and PoliciesWeixi Zhu, Alan L. Cox, Scott RixnerUSENIX ATC 2020 · 被引用 40 次
- Perforated Page: Supporting Fragmented Memory Allocation for Large PagesChang Hyun Park, Sanghoon Cha, Bokyeong Kim, Youngjin Kwon 等ISCA 2020 · 被引用 35 次
- Every walk's a hit: making page walks single-access cache hitsChang Hyun Park, Ilias Vougioukas, Andreas Sandberg, David Black-SchafferASPLOS 2022 · 被引用 34 次
- Trident: Harnessing Architectural Resources for All Page Sizes in x86 ProcessorsVenkat Sri Sai Ram, Ashish Panwar, Arkaprava BasuMICRO 2021 · 被引用 27 次
- Contiguitas: The Pursuit of Physical Memory Contiguity in DatacentersKaiyang Zhao, Kaiwen Xue, Ziqi Wang, Dan Schatzberg 等ISCA 2023 · 被引用 26 次
相关 Paper
- Gemina: A Coordinated and High-Performance Memory Deduplication EngineZhehua Zhang, Suzhen Wu, Wenyan You, Chunfeng Du 等HPCA 2025
- Exploiting Page Table Locality for Agile TLB PrefetchingGeorgios Vavouliotis, Lluc Alvarez, Vasileios Karakostas, Konstantinos Nikas 等ISCA 2021 · 被引用 34 次
- Victima: Drastically Increasing Address Translation Reach by Leveraging Underutilized Cache ResourcesKonstantinos Kanellopoulos, Hong Chul Nam, Nisa Bostanci, Rahul Bera 等MICRO 2023 · 被引用 16 次
- Dead Page and Dead Block Predictors: Cleaning TLBs and Caches TogetherChandrashis Mazumdar, Prachatos Mitra, Arkaprava BasuHPCA 2021 · 被引用 20 次
- Enhancing and Exploiting Contiguity for Fast Memory VirtualizationChloe Alverti, Stratos Psomadakis, Vasileios Karakostas, Jayneel Gandhi 等ISCA 2020 · 被引用 41 次
