StarNUMA: Mitigating NUMA Challenges with Memory Pooling
Albert Cho, Alexandros Daglis
Abstract
Large multi-socket machines are mission-critical high-performance systems for workloads requiring massive memory shared by hundreds of processors. Beyond eight sockets, such systems typically feature multi-hop inter-socket networks, exacerbating the Non-Uniform Memory Access (NUMA) challenge. NUMA effects stem from major disparity in latency and bandwidth characteristics of local and remote memory, often in the 4–10× range. While judicious data placement across the distributed memory's fragments can ameliorate NUMA effects, we observe that in challenging workloads with irregular access patterns, a large fraction of accessed pages are “vagabond”: being actively shared by multiple sockets, they lack a fitting home socket location. On 16-socket systems, such pages incur up to 75% remote memory accesses, which encounter significant latency overheads and bandwidth bottlenecks. StarNUMA introduces a new architectural block for multi-socket architectures to ameliorate the challenge posed by vagabond pages. By leveraging the capabilities of the emerging CXL interconnect, StarNUMA augments a typical NUMA architecture with a memory pool that is directly accessible by every socket in a single high-bandwidth interconnect hop. We show that placement of vagabond pages in StarNUMA's memory pool effectively curbs the latency overheads and queuing delays of the bandwidth-constrained multi-hop inter-socket network, reducing the average memory access time of 16-socket systems by 48%. In turn, faster memory access yields performance improvements of 1.54× on average, and up to 2.17×.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 2aa96489-9f87-42f8-badb-5abe7e6be222Cited by top-tier papers2
- COAXIAL: A CXL-Centric Memory System for Scalable ServersAlbert Cho, Anish Saxena, Moinuddin Qureshi, Alexandros DaglisSC 2024 · 12 citations
- PIPM: Partial and Incremental Page Migration for Multi-host CXL Disaggregated Shared MemoryGangqi Huang, Heiner Litz, Yuanchao XuASPLOS 2026 · 1 citation
Related papers
- CARINA: An Efficient CXL-Oriented Embedding Serving System for Recommendation ModelsPeiqi Yin, Qihui Zhou, Xiao Yan, Chao Wang et al.SIGMOD 2025 · 2 citations
- Mitosis: Transparently Self-Replicating Page-Tables for Large-Memory MachinesReto Achermann, Ashish Panwar, Abhishek Bhattacharjee, Timothy Roscoe et al.ASPLOS 2020 · 62 citations
- Beluga: A CXL-Based Memory Architecture for Scalable and Efficient LLM KVCache ManagementXinjun Yang, Qingda Hu, Junru Li, Feifei Li et al.SIGMOD 2026 · 24 citations
- Demystifying CXL Memory with Genuine CXL-Ready Systems and DevicesYan Sun, Yifan Yuan, Zeduo Yu, Reese Kuper et al.MICRO 2023 · 133 citations
- PaCaR: Improved Buffered I/O Locality on NUMA Systems with Page Cache ReplicationJérôme Coquisart, Julien Sopena, Redha GouicemEuroSys 2026
