Memory Allocation Under Hardware Compression
Muhammad Laghari, Yuqing Liu, Gagandeep Panwar, David Bears, Chandler Jearls, Raghavendra Srinivas, Esha Choukse, Kirk W. Cameron, Ali Raza Butt, Xun Jian
Abstract
As the scaling of memory density slows physically, a promising solution is to scale memory logically by enhancing the CPU's memory controller to encode and store data more densely in memory. This is known as hardware memory compression. Hardware memory compression decouples OS-managed physical memory from actual memory (i.e., DRAM); the memory controller spends a dynamically varying amount of DRAM on each physical page, depending on the compressibility of the page's content. The newly-decoupled actual memory effectively forms a new layer of memory beyond the traditional layers of virtual, pseudo-physical, and physical memory. We note unlike these traditional memory layers, each with its own specialized allocation interface (e.g., malloc/mmap for virtual memory, page tables+MMU for physical memory), this new layer of memory introduced by hardware memory compression still awaits its own unique memory allocation interface; its absence makes the allocation of actual memory imprecise and, sometimes, even impossible. Imprecisely allocating less actual memory, and/or unable to allocate more, can harm performance. Even imprecisely allocating more actual memory to some jobs can be harmful as it can result in allocating less actual memory to other jobs in highly-occupied memory systems, where compression is useful. To restore precise memory allocation, we design a new memory allocation specialized for this new layer of memory and, subsequently, architect a new MMU-like component in the memory controller and tackle the corresponding design challenges. We create a full-system FPGA prototype of a hardware-compressed memory system with precise memory allocation. Our evaluations using the prototype show that jobs perform stably under colocation. The performance variation is only 1%-2%; in comparison, it is 19%-89% under the prior art.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 440b6676-eaef-4f03-a8f5-79ba948c6ab5Builds on6
- TMO: transparent memory offloading in datacentersJohannes Weiner, Niket Agarwal, Dan Schatzberg, Leon Yang et al.ASPLOS 2022 · 103 citations
- BCD deduplication: effective memory compression using partial cache-line deduplicationSungbo Park, Ingab Kang, Yaebin Moon, Jung Ho Ahn et al.ASPLOS 2021 · 19 citations
- GBDI: Going Beyond Base-Delta-Immediate Compression with Global BasesAlexandra Angerd, Angelos Arelakis, Vasilis Spiliopoulos, Erik Sintorn et al.HPCA 2022 · 13 citations
- Translation-optimized Memory Compression for CapacityGagandeep Panwar, Muhammad Laghari, David Bears, Yuqing Liu et al.MICRO 2022 · 12 citations
- Exploiting Inter-block Entropy to Enhance the Compressibility of Blocks with Diverse DataJinkwon Kim, Mincheol Kang, Jeongkyu Hong, Soontae KimHPCA 2022 · 5 citations
Related papers
- DyLeCT: Achieving Huge-page-like Translation Performance for Hardware-compressed MemoryGagandeep Panwar, Muhammad Laghari, Esha Choukse, Xun JianISCA 2024 · 4 citations
- Random-Access Hardware Sequence CompressionNolan Chu, Yoon Lee, Gagandeep Panwar, Xun Steve JianISCA 2026
- Buddy Compression: Enabling Larger Memory for Deep Learning and HPC Workloads on GPUsEsha Choukse, Michael B. Sullivan, Mike O'Connor, Mattan Erez et al.ISCA 2020 · 58 citations
- Software-defined address mapping: a case on 3D memoryJialiang Zhang, Michael M. Swift, Jing Jane LiASPLOS 2022 · 14 citations
- High-Performance and Resource-Efficient Dynamic Memory Management in High-Level SynthesisQinggang Wang, Long Zheng, Zhaozeng An, Haoqin Huang et al.DAC 2024 · 5 citations
