Translation-optimized Memory Compression for Capacity
Gagandeep Panwar, Muhammad Laghari, David Bears, Yuqing Liu, Chandler Jearls, Esha Choukse, Kirk W. Cameron, Ali Raza Butt, Xun Jian
摘要
The demand for memory is ever increasing. Many prior works have explored hardware memory compression to increase effective memory capacity. However, prior works compress and pack/migrate data at a small - memory block-level - granularity; this introduces an additional block-level translation after the page-level virtual address translation. In general, the smaller the granularity of address translation, the higher the translation overhead. As such, this additional block-level translation exacerbates the well-known address translation problem for large and/or irregular workloads. A promising solution is to only save memory from cold (i.e., less recently accessed) pages without saving memory from hot (i.e., more recently accessed) pages (e.g., keep the hot pages uncompressed); this avoids block-level translation overhead for hot pages. However, it still faces two challenges. First, after a compressed cold page becomes hot again, migrating the page to a full 4KB DRAM location still adds another level (albeit page-level, instead of block-level) of translation on top of existing virtual address translation. Second, only compressing cold data require compressing them very aggressively to achieve high overall memory savings; decompressing very aggressively compressed data is very slow (e.g., assuming the latest Deflate ASIC in industry). This paper presents Translation-optimized Memory Compression for Capacity (TMCC) to tackle the two challenges above. To address the first challenge, we propose compressing page table blocks in hardware to opportunistically embed compression translations into them in a software-transparent manner to effectively prefetch compression translations during a page walk, instead of serially fetching them after the walk. To address the second challenge, we perform a large design space exploration across many hardware configurations and diverse workloads to derive and implement in HDL an ASIC Deflate that is specialized for memory; for memory pages, it is 4X as fast as the state-of-the art ASIC Deflate, with little to no sacrifice in compression ratio. Our evaluations show that for large and/or irregular workloads, TMCC can either improve performance by 14% without sacrificing effective capacity or provide 2.2x the effective capacity without sacrificing performance compared to a state-of-the-art hardware memory compression for capacity.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- LLM.265: Video Codecs are Secretly Tensor CodecsCeyu Xu, Yongji Wu, Xinyu Yang, Beidi Chen 等MICRO 2025 · 被引用 13 次
- DyLeCT: Achieving Huge-page-like Translation Performance for Hardware-compressed MemoryGagandeep Panwar, Muhammad Laghari, Esha Choukse, Xun JianISCA 2024 · 被引用 4 次
- Memory Allocation Under Hardware CompressionMuhammad Laghari, Yuqing Liu, Gagandeep Panwar, David Bears 等MICRO 2024 · 被引用 2 次
它引用的顶会 Paper9
- DRAMA: Exploiting DRAM Addressing for Cross-CPU AttacksPeter Pessl, Daniel Gruss, Clémentine Maurice, Michael Schwarz 等USENIX Security 2016 · 被引用 500 次
- TMO: transparent memory offloading in datacentersJohannes Weiner, Niket Agarwal, Dan Schatzberg, Leon Yang 等ASPLOS 2022 · 被引用 103 次
- Mitosis: Transparently Self-Replicating Page-Tables for Large-Memory MachinesReto Achermann, Ashish Panwar, Abhishek Bhattacharjee, Timothy Roscoe 等ASPLOS 2020 · 被引用 62 次
- Elastic Cuckoo Page Tables: Rethinking Virtual Memory Translation for ParallelismDimitrios Skarlatos, Apostolos Kokolis, Tianyin Xu, Josep TorrellasASPLOS 2020 · 被引用 55 次
- Memory-harvesting VMs in cloud platformsAlexander Fuerst, Stanko Novakovic, Iñigo Goiri, Gohar Irfan Chaudhry 等ASPLOS 2022 · 被引用 39 次
相关 Paper
- Random-Access Hardware Sequence CompressionNolan Chu, Yoon Lee, Gagandeep Panwar, Xun Steve JianISCA 2026
- MetaZip: a high-throughput and efficient accelerator for DEFLATERuihao Gao, Xueqi Li, Yewen Li, Xun Wang 等DAC 2022 · 被引用 9 次
- LATPC: Accelerating GPU Address Translation Using Locality-Aware TLB Prefetching and MSHR CompressionYeonan Ha, Jiho Park, Hanna Cha, Jiwon Lee 等MICRO 2025 · 被引用 2 次
- Victima: Drastically Increasing Address Translation Reach by Leveraging Underutilized Cache ResourcesKonstantinos Kanellopoulos, Hong Chul Nam, Nisa Bostanci, Rahul Bera 等MICRO 2023 · 被引用 16 次
- Learning to Walk: Architecting Learned Virtual Memory TranslationKaiyang Zhao, Yuang Chen, Xenia Xu, Dan Schatzberg 等MICRO 2025 · 被引用 2 次
