Scalable and Effective Page-table and TLB management on NUMA Systems
Bin Gao, Qingxuan Kang, Hao-Wei Tee, Kyle Timothy Ng Chu, Alireza Sanaee, Djordje Jevdjic
Abstract
Memory management operations that modify page-tables, typically performed during memory allocation/deallocation, are infamous for their poor performance in highly threaded applications, largely due to process-wide TLB shootdowns that the OS must issue due to the lack of hardware support for TLB coherence. We study these operations in NUMA settings, where we observe up to 40x overhead for basic operations such as munmap or mprotect. The overhead further increases if pagetable replication is used, where complete coherent copies of the page-tables are maintained across all NUMA nodes. While eager system-wide replication is extremely effective at localizing page-table reads during address translation, we find that it creates additional penalties upon any page-table changes due to the need to maintain all replicas coherent.
In this paper, we propose a novel page-table management mechanism, called Hydra, to enable transparent, on-demand, and partial page-table replication across NUMA nodes in order to perform address translation locally, while avoiding the overheads and scalability issues of system-wide full page-table replication. We then show that Hydra's precise knowledge of page-table sharers can be leveraged to significantly reduce the number of TLB shootdowns issued upon any memory-management operation. As a result, Hydra not only avoids replication-related slowdowns, but also provides significant speedup over the baseline on memory allocation/deallocation and access control operations. We implement Hydra in Linux on x86_64, evaluate it on 4-and 8-socket systems, and show that Hydra achieves the full benefits of eager page-table replication on a wide range of applications, while also achieving a 12% and 36% runtime improvement on Webserver and Memcached respectively due to a significant reduction in TLB shootdowns.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 91e290b2-bfc0-4d7e-978e-248233755e4cCited by top-tier papers4
- ParaSync: Exploiting Fine-Grained Parallelism for Efficient File SynchronizationZhihao Zhang, Lu Tang, Huiba Li, Yue Yu et al.FAST 2026 · 1 citation
- Compaction-Free Memory Defragmentation for Virtualization via Infinite Guest Physical Address SpacePeixin Zeng, Hao Huang, Yanqi Pan, Wen Xia et al.OSDI 2026
- ScaleSwap: A Scalable OS Swap System for All-Flash Swap ArraysTaehwan Ahn, Chanhyeong Yu, Sangjin Lee, Yongseok SonFAST 2026
- MAC: Metadata Acceleration for Sustainable Performance in Big-Data Systems with CXL DRAMDusol Lee, Yan Sun, Houxiang Ji, Vinit Gupta et al.OSDI 2026
Builds on9
- Pond: CXL-Based Memory Pooling Systems for Cloud PlatformsHuaicheng Li, Daniel S. Berger, Lisa Hsu, Daniel Ernst et al.ASPLOS 2023 · 328 citations
- TPP: Transparent Page Placement for CXL-Enabled Tiered-MemoryHasan Al Maruf, Hao Wang, Abhishek Dhanotia, Johannes Weiner et al.ASPLOS 2023 · 255 citations
- Understanding host network stack overheadsQizhe Cai, Shubham Chaudhary, Midhul Vuppalapati, Jaehyun Hwang et al.SIGCOMM 2021 · 150 citations
- Mitosis: Transparently Self-Replicating Page-Tables for Large-Memory MachinesReto Achermann, Ashish Panwar, Abhishek Bhattacharjee, Timothy Roscoe et al.ASPLOS 2020 · 62 citations
- Don't shoot down TLB shootdowns!Nadav Amit, Amy Tai, Michael WeiEuroSys 2020 · 32 citations
Related papers
- WASP: Workload-Aware Self-Replicating Page-Tables for NUMA ServersHongliang Qu, Zhibin YuASPLOS 2024 · 7 citations
- PaCaR: Improved Buffered I/O Locality on NUMA Systems with Page Cache ReplicationJérôme Coquisart, Julien Sopena, Redha GouicemEuroSys 2026
- Fast local page-tables for virtualized NUMA servers with vMitosisAshish Panwar, Reto Achermann, Arkaprava Basu, Abhishek Bhattacharjee et al.ASPLOS 2021 · 29 citations
- Learning to Walk: Architecting Learned Virtual Memory TranslationKaiyang Zhao, Yuang Chen, Xenia Xu, Dan Schatzberg et al.MICRO 2025 · 2 citations
- Hydra : Resilient and Highly Available Remote MemoryYoungmoon Lee, Hasan Al Maruf, Mosharaf Chowdhury, Asaf Cidon et al.FAST 2022
