Memory Allocation Under Hardware Compression
Muhammad Laghari, Yuqing Liu, Gagandeep Panwar, David Bears, Chandler Jearls, Raghavendra Srinivas, Esha Choukse, Kirk W. Cameron, Ali Raza Butt, Xun Jian
摘要
As the scaling of memory density slows physically, a promising solution is to scale memory logically by enhancing the CPU's memory controller to encode and store data more densely in memory. This is known as hardware memory compression. Hardware memory compression decouples OS-managed physical memory from actual memory (i.e., DRAM); the memory controller spends a dynamically varying amount of DRAM on each physical page, depending on the compressibility of the page's content. The newly-decoupled actual memory effectively forms a new layer of memory beyond the traditional layers of virtual, pseudo-physical, and physical memory. We note unlike these traditional memory layers, each with its own specialized allocation interface (e.g., malloc/mmap for virtual memory, page tables+MMU for physical memory), this new layer of memory introduced by hardware memory compression still awaits its own unique memory allocation interface; its absence makes the allocation of actual memory imprecise and, sometimes, even impossible. Imprecisely allocating less actual memory, and/or unable to allocate more, can harm performance. Even imprecisely allocating more actual memory to some jobs can be harmful as it can result in allocating less actual memory to other jobs in highly-occupied memory systems, where compression is useful. To restore precise memory allocation, we design a new memory allocation specialized for this new layer of memory and, subsequently, architect a new MMU-like component in the memory controller and tackle the corresponding design challenges. We create a full-system FPGA prototype of a hardware-compressed memory system with precise memory allocation. Our evaluations using the prototype show that jobs perform stably under colocation. The performance variation is only 1%-2%; in comparison, it is 19%-89% under the prior art.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper6
- TMO: transparent memory offloading in datacentersJohannes Weiner, Niket Agarwal, Dan Schatzberg, Leon Yang 等ASPLOS 2022 · 被引用 103 次
- BCD deduplication: effective memory compression using partial cache-line deduplicationSungbo Park, Ingab Kang, Yaebin Moon, Jung Ho Ahn 等ASPLOS 2021 · 被引用 19 次
- GBDI: Going Beyond Base-Delta-Immediate Compression with Global BasesAlexandra Angerd, Angelos Arelakis, Vasilis Spiliopoulos, Erik Sintorn 等HPCA 2022 · 被引用 13 次
- Translation-optimized Memory Compression for CapacityGagandeep Panwar, Muhammad Laghari, David Bears, Yuqing Liu 等MICRO 2022 · 被引用 12 次
- Exploiting Inter-block Entropy to Enhance the Compressibility of Blocks with Diverse DataJinkwon Kim, Mincheol Kang, Jeongkyu Hong, Soontae KimHPCA 2022 · 被引用 5 次
相关 Paper
- DyLeCT: Achieving Huge-page-like Translation Performance for Hardware-compressed MemoryGagandeep Panwar, Muhammad Laghari, Esha Choukse, Xun JianISCA 2024 · 被引用 4 次
- Random-Access Hardware Sequence CompressionNolan Chu, Yoon Lee, Gagandeep Panwar, Xun Steve JianISCA 2026
- Buddy Compression: Enabling Larger Memory for Deep Learning and HPC Workloads on GPUsEsha Choukse, Michael B. Sullivan, Mike O'Connor, Mattan Erez 等ISCA 2020 · 被引用 58 次
- Software-defined address mapping: a case on 3D memoryJialiang Zhang, Michael M. Swift, Jing Jane LiASPLOS 2022 · 被引用 14 次
- High-Performance and Resource-Efficient Dynamic Memory Management in High-Level SynthesisQinggang Wang, Long Zheng, Zhaozeng An, Haoqin Huang 等DAC 2024 · 被引用 5 次
