Architectural Support for Optimizing Huge Page Selection Within the OS
Aninda Manocha, Zi Yan, Esin Tureci, Juan L. Aragón, David W. Nellans, Margaret Martonosi
Abstract
Irregular, memory-intensive applications often incur high translation lookaside bu!er (TLB) miss rates that result in signi"cant address translation overheads. Employing huge pages is an e!ective way to reduce these overheads, however in real systems the number of available huge pages can be limited when system memory is nearly full and/or fragmented. Thus, huge pages must be used selectively to back application memory. This work demonstrates that choosing memory regions that incur the most TLB misses for huge page promotion best reduces address translation overheads. We call these regions High reUse TLB-sensitive data (HUBs). Unlike prior work which relies on expensive per-page software counters to identify promotion regions, we propose new architectural support to identify these regions dynamically at application runtime.
We propose a promotion candidate cache (PCC) that identi"es HUB candidates based on hardware page table walks after a lastlevel TLB miss. This small, "xed-size structure tracks huge pagealigned regions (consisting of 𝐿 base pages), ranks them based on observed page table walk frequency, and only keeps the most frequently accessed ones. Evaluated on applications of various memory intensity, our approach successfully identi"es application pages incurring the highest address translation overheads. Our approach demonstrates that with the help of a PCC, the OS only needs to promote 4% of the application footprint to achieve more than 75% of the peak achievable performance, yielding 1.19-1.33→ speedups over 4KB base pages alone. In real systems where memory is typically fragmented, the PCC outperforms Linux's page promotion policy by 14% (when 50% of total memory is fragmented) and 16% (when 90% of total memory is fragmented) respectively.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e5ef2887-1c22-4adf-93f2-d3242c3e87e0Cited by top-tier papers1
Ask how each one uses itBuilds on8
- A Comprehensive Analysis of Superpage Management Mechanisms and PoliciesWeixi Zhu, Alan L. Cox, Scott RixnerUSENIX ATC 2020 · 40 citations
- Perforated Page: Supporting Fragmented Memory Allocation for Large PagesChang Hyun Park, Sanghoon Cha, Bokyeong Kim, Youngjin Kwon et al.ISCA 2020 · 35 citations
- Every walk's a hit: making page walks single-access cache hitsChang Hyun Park, Ilias Vougioukas, Andreas Sandberg, David Black-SchafferASPLOS 2022 · 34 citations
- Trident: Harnessing Architectural Resources for All Page Sizes in x86 ProcessorsVenkat Sri Sai Ram, Ashish Panwar, Arkaprava BasuMICRO 2021 · 27 citations
- Contiguitas: The Pursuit of Physical Memory Contiguity in DatacentersKaiyang Zhao, Kaiwen Xue, Ziqi Wang, Dan Schatzberg et al.ISCA 2023 · 26 citations
Related papers
- Gemina: A Coordinated and High-Performance Memory Deduplication EngineZhehua Zhang, Suzhen Wu, Wenyan You, Chunfeng Du et al.HPCA 2025
- Exploiting Page Table Locality for Agile TLB PrefetchingGeorgios Vavouliotis, Lluc Alvarez, Vasileios Karakostas, Konstantinos Nikas et al.ISCA 2021 · 34 citations
- Victima: Drastically Increasing Address Translation Reach by Leveraging Underutilized Cache ResourcesKonstantinos Kanellopoulos, Hong Chul Nam, Nisa Bostanci, Rahul Bera et al.MICRO 2023 · 16 citations
- Dead Page and Dead Block Predictors: Cleaning TLBs and Caches TogetherChandrashis Mazumdar, Prachatos Mitra, Arkaprava BasuHPCA 2021 · 20 citations
- Enhancing and Exploiting Contiguity for Fast Memory VirtualizationChloe Alverti, Stratos Psomadakis, Vasileios Karakostas, Jayneel Gandhi et al.ISCA 2020 · 41 citations
