Lune

OSDI2025Top-tier venue

Tiered Memory Management Beyond Hotness

Jinshu Liu, Hamid Hadian, Hanchen Xu, Huaicheng Li

2025Year
13Citations
9Top-tier citations

Abstract

Tiered memory systems often rely on access frequency ("hotness") to guide data placement. However, hot data is not always performance-critical, limiting the effectiveness of hotness-based policies. We introduce amortized offcore latency (AOL), a novel metric that precisely captures the true performance impact of memory accesses by accounting for memory access latency and memory-level parallelism (MLP). Leveraging AOL, we present two powerful tiering mechanisms: Soar, a profile-guided allocation policy that places objects based on their performance contribution, and Alto, a lightweight page migration regulation policy to eliminate unnecessary migrations. Soar and Alto outperform four state-of-the-art tiering designs across a diverse set of workloads by up to 12.4×, while underperforming in a few cases by no more than 3%. graph, cloud, and HPC workloads on both NUMA and real CXL platforms, varying fast-to-slow tier ratios and bandwidth contention levels. Soar outperforms Nomad, NBT, Colloid, and TPP by 14-547%, 4-79%, -1-68%, and 31-1242%, respectively; Alto improves performance by -2-81%, 1-31%, -3-18%, and 2-471%. Negative improvements indicate that Soar/Alto underperform relative to baselines in a few cases (5 out of 182 in total). While Soar and Alto achieve strong results broadly, their performance gains are less pronounced under high bandwidth contention due to AOL inflation from queuing delays. Raising AOL thresholds can restore their performance gains but requires contention-aware tuning. We highlight this to clarify the scope of our approach and leave AOL tuning as future work.

In summary, we make the following contributions: • We quantitatively demonstrate that hotness is an unreliable proxy for performance-criticality: the performance impact of memory accesses can vary by up to 4× across workloads.

• We introduce AOL, a performance metric that combines memory access latency and MLP, and leverages CPU stall cycles to accurately estimate tiered memory performance.

• We propose AOL-powered memory management policies: Soar for near-optimal data placement and Alto for adaptive migration control.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext 4ffb2ab9-438b-46ba-a3c7-dcc4f74515e2

Cited by top-tier papers9

Ask how each one uses it

Builds on24

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines