Lune

ISCA2026顶会

Revelator: Rapid Data Fetching Via System-Software-Guided Hash-Based Speculative Address Translation

Konstantinos Kanellopoulos, Konstantinos Sgouras, Harsh Songara, Andreas Kosmas Kakolyris, Vlad-Petru Nitu, Spiros Galanopoulos, Rahul Bera, Konstantina Koliogeorgi, Rakesh Kumar, Onur Mutlu

2026年份

摘要

Address translation is a significant performance bottleneck in modern computing systems. Predicting the physical address (PA) of the requested data before address translation completes is a promising technique to hide the latency of address translation. However, accurately predicting the PA based on the virtual address (VA) is highly challenging due to the inherent unpredictability of the VA-to-PA mappings that conventional operating systems introduce. Prior works try to introduce predictability into VA-to-PA mappings but exhibit two key shortcomings: (i) they rely on the availability of large pages or VA-to-PA contiguity, and (ii) they store speculation-related metadata in costly hardware components whose power and area overheads may not always justify their benefits.

We introduce Revelator, a new hardware-OS cooperative technique that uses hashing to enable highly accurate speculative address translation with small system modifications. At the OS level, Revelator employs a tiered hash-based memory allocation policy for both program data and last-level page table entries (PTEs) to establish predictable VA-to-PA and VA-to-PTE mappings. At the hardware level, after an L2 TLB miss, a lightweight speculation engine uses the OS's hash functions to predict the VAto-PA and VA-to-PTE mappings and prefetch the corresponding cache blocks before address translation completes, thereby both hiding address translation latency and accelerating page table walks (PTWs). Revelator comes with two key benefits: (i) it does not rely on large pages or VA-to-PA contiguity to achieve high speculation accuracy, and (ii) it requires small OS and hardware modifications.

Our evaluation across 11 data-intensive workloads shows that, (i) in single-core systems, Revelator provides an average speedup of 15.3% over the state-of-the-art speculative address translation technique under high memory fragmentation, and (ii) in virtualized environments, by accurately predicting both guest and host physical addresses, Revelator provides a 13.6% average speedup over Nested Paging. In multicore systems, Revelator's benefits scale with core count, enabling a speedup of 1.40× (1.50×) over Transparent Huge Pages (THP) across 30 server workload mixes from Google in a 16-core system under medium (high) memory fragmentation. These gains come at very small hardware cost. Our RTL synthesis shows that Revelator incurs only 0.02% area and 0.03% power overheads on top of a high-end servergrade CPU. Revelator is freely available at github.com/CMU- SAFARI/Virtuoso.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

lune papers fulltext 8e755d6d-138f-48de-964f-c6796f912e19

它引用的顶会 Paper44

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖