DyLeCT: Achieving Huge-page-like Translation Performance for Hardware-compressed Memory
Gagandeep Panwar, Muhammad Laghari, Esha Choukse, Xun Jian
摘要
To expand effective memory capacity, hardware memory compression transparently compresses and packs memory values more densely together in DRAM. This requires introducing a new layer of hardware-managed address translation in the memory controller (MC). However, for large and irregular workloads that already suffer from frequent virtual address translation misses in the TLB, adding an additional layer of address translation can double the translation misses (e.g., by adding a new miss in the MC per TLB miss). While TLB misses can be drastically reduced by using huge pages, no prior work has explored huge-page-like translation reach for hardware memory compression. While compressing and moving an entire huge page worth of data at a time can lead to huge-page-like address translation, moving a huge page worth of data together can consume an exorbitant amount of memory bandwidth.This paper explores how to achieve huge-page-like translation performance in this new address translation layer, while keeping compression at the page (instead of huge page) granularity. We propose dynamically shortening the translation entries of hot pages to only a few bits per entry by migrating hot pages to the limited number of DRAM locations whose addresses can be encoded using a few bits; colder pages still use the bigger fulllength translations so that colder pages can be placed anywhere in memory to fully utilize all the space in memory. Each short translation is tiny (e.g., 2 bits); as such, a 128KB translation cache filled mostly with short translations can achieve similar (e.g., 2GB) total translation reach as a TLB filled entirely with huge page entries. Evaluations show our idea – Dynamic Length Compressed-Memory Translations (DyLeCT) – improves average performance by 10.25% over the prior art.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Memory Allocation Under Hardware CompressionMuhammad Laghari, Yuqing Liu, Gagandeep Panwar, David Bears 等MICRO 2024 · 被引用 2 次
- Assassyn: A Unified Abstraction for Architectural Simulation and ImplementationJian Weng, Boyang Han, Derui Gao, Ruijie Gao 等ISCA 2025 · 被引用 1 次
- The XOR Cache: A Catalyst for CompressionZhewen Pan, Joshua San MiguelISCA 2025 · 被引用 1 次
它引用的顶会 Paper15
- DRAMA: Exploiting DRAM Addressing for Cross-CPU AttacksPeter Pessl, Daniel Gruss, Clémentine Maurice, Michael Schwarz 等USENIX Security 2016 · 被引用 500 次
- TMO: transparent memory offloading in datacentersJohannes Weiner, Niket Agarwal, Dan Schatzberg, Leon Yang 等ASPLOS 2022 · 被引用 103 次
- MEMTIS: Efficient Memory Tiering with Dynamic Page Classification and Page Size DeterminationTaehyung Lee, Sumit Kumar Monga, Changwoo Min, Young Ik EomSOSP 2023 · 被引用 67 次
- Elastic Cuckoo Page Tables: Rethinking Virtual Memory Translation for ParallelismDimitrios Skarlatos, Apostolos Kokolis, Tianyin Xu, Josep TorrellasASPLOS 2020 · 被引用 55 次
- Beyond malloc efficiency to fleet efficiency: a hugepage-aware memory allocatorA. H. Hunter, Chris Kennelly, Paul Turner, Darryl Gove 等OSDI 2021 · 被引用 51 次
相关 Paper
- Translation-optimized Memory Compression for CapacityGagandeep Panwar, Muhammad Laghari, David Bears, Yuqing Liu 等MICRO 2022 · 被引用 12 次
- Elastic Translations: Fast Virtual Memory with Multiple Translation SizesStratos Psomadakis, Chloe Alverti, Vasileios Karakostas, Christos Katsakioris 等MICRO 2024 · 被引用 4 次
- Mosaic Pages: Big TLB Reach with Small PagesKrishnan Gosakan, Jaehyun Han, William Kuszmaul, Ibrahim N. Mubarek 等ASPLOS 2023 · 被引用 21 次
- Direct Memory Translation for Virtualized CloudsJiyuan Zhang, Weiwei Jia, Siyuan Chai, Peizhe Liu 等ASPLOS 2024 · 被引用 5 次
- Random-Access Hardware Sequence CompressionNolan Chu, Yoon Lee, Gagandeep Panwar, Xun Steve JianISCA 2026
