Revelator: Rapid Data Fetching Via System-Software-Guided Hash-Based Speculative Address Translation
Konstantinos Kanellopoulos, Konstantinos Sgouras, Harsh Songara, Andreas Kosmas Kakolyris, Vlad-Petru Nitu, Spiros Galanopoulos, Rahul Bera, Konstantina Koliogeorgi, Rakesh Kumar, Onur Mutlu
摘要
Address translation is a significant performance bottleneck in modern computing systems. Predicting the physical address (PA) of the requested data before address translation completes is a promising technique to hide the latency of address translation. However, accurately predicting the PA based on the virtual address (VA) is highly challenging due to the inherent unpredictability of the VA-to-PA mappings that conventional operating systems introduce. Prior works try to introduce predictability into VA-to-PA mappings but exhibit two key shortcomings: (i) they rely on the availability of large pages or VA-to-PA contiguity, and (ii) they store speculation-related metadata in costly hardware components whose power and area overheads may not always justify their benefits.
We introduce Revelator, a new hardware-OS cooperative technique that uses hashing to enable highly accurate speculative address translation with small system modifications. At the OS level, Revelator employs a tiered hash-based memory allocation policy for both program data and last-level page table entries (PTEs) to establish predictable VA-to-PA and VA-to-PTE mappings. At the hardware level, after an L2 TLB miss, a lightweight speculation engine uses the OS's hash functions to predict the VAto-PA and VA-to-PTE mappings and prefetch the corresponding cache blocks before address translation completes, thereby both hiding address translation latency and accelerating page table walks (PTWs). Revelator comes with two key benefits: (i) it does not rely on large pages or VA-to-PA contiguity to achieve high speculation accuracy, and (ii) it requires small OS and hardware modifications.
Our evaluation across 11 data-intensive workloads shows that, (i) in single-core systems, Revelator provides an average speedup of 15.3% over the state-of-the-art speculative address translation technique under high memory fragmentation, and (ii) in virtualized environments, by accurately predicting both guest and host physical addresses, Revelator provides a 13.6% average speedup over Nested Paging. In multicore systems, Revelator's benefits scale with core count, enabling a speedup of 1.40× (1.50×) over Transparent Huge Pages (THP) across 30 server workload mixes from Google in a 16-core system under medium (high) memory fragmentation. These gains come at very small hardware cost. Our RTL synthesis shows that Revelator incurs only 0.02% area and 0.03% power overheads on top of a high-end servergrade CPU. Revelator is freely available at github.com/CMU- SAFARI/Virtuoso.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper44
- Meltdown: Reading Kernel Memory from User SpaceMoritz Lipp, Michael Schwarz, Daniel Gruss, Thomas Prescher 等USENIX Security 2018 · 被引用 1,456 次
- Effectively Prefetching Remote Memory with LeapHasan Al Maruf, Mosharaf ChowdhuryUSENIX ATC 2020 · 被引用 186 次
- Rethinking software runtimes for disaggregated memoryIrina Calciu, M. Talha Imran, Ivan Puddu, Sanidhya Kashyap 等ASPLOS 2021 · 被引用 116 次
- One-sided RDMA-Conscious Extendible Hashing for Disaggregated MemoryPengfei Zuo, Jiazhao Sun, Liu Yang, Shuangwu Zhang 等USENIX ATC 2021 · 被引用 113 次
- Prodigy: Improving the Memory Latency of Data-Indirect Irregular Workloads Using Hardware-Software Co-DesignNishil Talati, Kyle May, Armand Behroozi, Yichen Yang 等HPCA 2021 · 被引用 62 次
相关 Paper
- Enhancing and Exploiting Contiguity for Fast Memory VirtualizationChloe Alverti, Stratos Psomadakis, Vasileios Karakostas, Jayneel Gandhi 等ISCA 2020 · 被引用 41 次
- Victima: Drastically Increasing Address Translation Reach by Leveraging Underutilized Cache ResourcesKonstantinos Kanellopoulos, Hong Chul Nam, Nisa Bostanci, Rahul Bera 等MICRO 2023 · 被引用 16 次
- A Case for Speculative Address Translation with Rapid Validation for GPUsJunhyeok Park, Osang Kwon, Yongho Lee, Seongwook Kim 等MICRO 2024 · 被引用 13 次
- CARAT: a case for virtual memory through compiler- and runtime-based address translationBrian Suchy, Simone Campanoni, Nikos Hardavellas, Peter A. DindaPLDI 2020 · 被引用 16 次
- SoftWalker: Supporting Software Page Table Walk for Irregular GPU ApplicationsSungbin Jang, Junhyeok Park, Yongho Lee, Osang Kwon 等MICRO 2025 · 被引用 4 次
