Lune

OSDI2025顶会

Tiered Memory Management Beyond Hotness

Jinshu Liu, Hamid Hadian, Hanchen Xu, Huaicheng Li

出版方
2025年份
13被引次数
9顶会引用

摘要

Tiered memory systems often rely on access frequency ("hotness") to guide data placement. However, hot data is not always performance-critical, limiting the effectiveness of hotness-based policies. We introduce amortized offcore latency (AOL), a novel metric that precisely captures the true performance impact of memory accesses by accounting for memory access latency and memory-level parallelism (MLP). Leveraging AOL, we present two powerful tiering mechanisms: Soar, a profile-guided allocation policy that places objects based on their performance contribution, and Alto, a lightweight page migration regulation policy to eliminate unnecessary migrations. Soar and Alto outperform four state-of-the-art tiering designs across a diverse set of workloads by up to 12.4×, while underperforming in a few cases by no more than 3%. graph, cloud, and HPC workloads on both NUMA and real CXL platforms, varying fast-to-slow tier ratios and bandwidth contention levels. Soar outperforms Nomad, NBT, Colloid, and TPP by 14-547%, 4-79%, -1-68%, and 31-1242%, respectively; Alto improves performance by -2-81%, 1-31%, -3-18%, and 2-471%. Negative improvements indicate that Soar/Alto underperform relative to baselines in a few cases (5 out of 182 in total). While Soar and Alto achieve strong results broadly, their performance gains are less pronounced under high bandwidth contention due to AOL inflation from queuing delays. Raising AOL thresholds can restore their performance gains but requires contention-aware tuning. We highlight this to clarify the scope of our approach and leave AOL tuning as future work.

In summary, we make the following contributions: • We quantitatively demonstrate that hotness is an unreliable proxy for performance-criticality: the performance impact of memory accesses can vary by up to 4× across workloads.

• We introduce AOL, a performance metric that combines memory access latency and MLP, and leverages CPU stall cycles to accurately estimate tiered memory performance.

• We propose AOL-powered memory management policies: Soar for near-optimal data placement and Alto for adaptive migration control.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

lune papers fulltext 4ffb2ab9-438b-46ba-a3c7-dcc4f74515e2

引用它的顶会 Paper9

问问它们各自怎么用它

它引用的顶会 Paper24

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖