Fast local page-tables for virtualized NUMA servers with vMitosis
Ashish Panwar, Reto Achermann, Arkaprava Basu, Abhishek Bhattacharjee, K. Gopinath, Jayneel Gandhi
摘要
Increasing memory heterogeneity mandates careful data placement to hide the non-uniform memory access (NUMA) effects on applications. While NUMA optimizations have focused on application data for decades, they have ignored the placement of kernel data structures due to their small memory footprint; this is evident in typical OSes that pin kernel data structures in memory. In this paper, we show that careful placement of kernel data structures is gaining importance in the context of page-tables: their sub-optimal placement causes severe slowdown (up to 3.1×) on virtualized NUMA servers.
In response, we present vMitosis ś a system for explicit management of two-level page-tables, i.e., the guest and extended pagetables, on virtualized NUMA servers. vMitosis enables faster address translation by migrating and replicating page-tables. It supports two prevalent virtualization configurations: first, where the hypervisor exposes the NUMA architecture to the guest OS, and second, where such information is hidden from the guest OS. vMitosis is implemented in Linux/KVM, and our evaluation on a recent 1.5TiB 4-socket server shows that it effectively eliminates NUMA effects on 2D page-table walks, resulting in a speedup of 1.8-3.1× for Thin (single-socket) and 1.06 -1.6× for Wide (multi-socket) workloads.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper14
- MEMTIS: Efficient Memory Tiering with Dynamic Page Classification and Page Size DeterminationTaehyung Lee, Sumit Kumar Monga, Changwoo Min, Young Ik EomSOSP 2023 · 被引用 67 次
- Trident: Harnessing Architectural Resources for All Page Sizes in x86 ProcessorsVenkat Sri Sai Ram, Ashish Panwar, Arkaprava BasuMICRO 2021 · 被引用 27 次
- Designing Virtual Memory System of MCM GPUsPratheek B, Neha Jawalkar, Arkaprava BasuMICRO 2022 · 被引用 21 次
- Victima: Drastically Increasing Address Translation Reach by Leveraging Underutilized Cache ResourcesKonstantinos Kanellopoulos, Hong Chul Nam, Nisa Bostanci, Rahul Bera 等MICRO 2023 · 被引用 16 次
- IDYLL: Enhancing Page Translation in Multi-GPUs via Light Weight PTE InvalidationsBingyao Li, Yanan Guo, Yueqi Wang, Aamer Jaleel 等MICRO 2023 · 被引用 16 次
它引用的顶会 Paper5
- Mitosis: Transparently Self-Replicating Page-Tables for Large-Memory MachinesReto Achermann, Ashish Panwar, Abhishek Bhattacharjee, Timothy Roscoe 等ASPLOS 2020 · 被引用 62 次
- Elastic Cuckoo Page Tables: Rethinking Virtual Memory Translation for ParallelismDimitrios Skarlatos, Apostolos Kokolis, Tianyin Xu, Josep TorrellasASPLOS 2020 · 被引用 55 次
- Enhancing and Exploiting Contiguity for Fast Memory VirtualizationChloe Alverti, Stratos Psomadakis, Vasileios Karakostas, Jayneel Gandhi 等ISCA 2020 · 被引用 41 次
- A Comprehensive Analysis of Superpage Management Mechanisms and PoliciesWeixi Zhu, Alan L. Cox, Scott RixnerUSENIX ATC 2020 · 被引用 40 次
- Tailored Page SizesFaruk Guvenilir, Yale N. PattISCA 2020 · 被引用 22 次
相关 Paper
- Every walk's a hit: making page walks single-access cache hitsChang Hyun Park, Ilias Vougioukas, Andreas Sandberg, David Black-SchafferASPLOS 2022 · 被引用 34 次
- WASP: Workload-Aware Self-Replicating Page-Tables for NUMA ServersHongliang Qu, Zhibin YuASPLOS 2024 · 被引用 7 次
- Scalable and Effective Page-table and TLB management on NUMA SystemsBin Gao, Qingxuan Kang, Hao-Wei Tee, Kyle Timothy Ng Chu 等USENIX ATC 2024 · 被引用 5 次
- Learning to Walk: Architecting Learned Virtual Memory TranslationKaiyang Zhao, Yuang Chen, Xenia Xu, Dan Schatzberg 等MICRO 2025 · 被引用 2 次
- PaCaR: Improved Buffered I/O Locality on NUMA Systems with Page Cache ReplicationJérôme Coquisart, Julien Sopena, Redha GouicemEuroSys 2026
