NOMAD: Enabling Non-blocking OS-managed DRAM Cache via Tag-Data Decoupling
Youngin Kim, Hyeonjin Kim, William J. Song
摘要
This paper introduces a DRAM cache architecture that provides near-ideal access time and non-blocking miss handling. Previous DRAM cache (DC) designs are classified into two categories, HW-based and OS-managed schemes. Hardware-based designs implement non-blocking caches that can handle multiple DC misses using MSHRs, but they have drawbacks in metadata management since storing tags in on-package DRAM significantly increases the effective cycle time of DC accesses. In contrast, OS-managed schemes utilize PTEs for storing tags and caching them in TLBs, which can achieve ideal DC access time. However, they implement blocking caches that stall application threads on misses until cache fills are completed. To overcome the limitations of both HW-based and OS-managed schemes, this paper introduces a DRAM cache architecture named Non-blocking OS-managed DRAM cache (NOMAD). Unlike conventional caches that guarantee the presence of data on tag hits, NOMAD decouples tag and data management to enable non-blocking miss handling in an OS-managed DRAM cache. The front-end OS routines of NOMAD manage DC tags using PTEs and TLBs, and its back-end hardware handles data management in the DRAM cache. On a DC miss, the OS updates a tag, offloads a cache-fill command to the back-end, and immediately resumes an application thread without waiting for the cache fill to complete. Instead, the back-end hardware handles the cache fill without blocking the application thread. By decoupling tag and data management in NOMAD, a tag hit does not necessarily guarantee the presence of data in the DRAM cache. The back-end traces which DC lines are still in transfers and checks if the demanded part of a cache line has been transferred yet for every DC access. Notably, this back-end procedure does not require an OS intervention, thereby implementing a non-blocking DRAM cache. Experiment results show that NOMAD reduces application stall cycles by 76.1% and improves IPC by 16.7% over a state-of-the-art OS-managed scheme.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Bandwidth-Effective DRAM Cache for GPU s with Storage-Class MemoryJeongmin Hong, Sungjun Cho, Geonwoo Park, Wonhyuk Yang 等HPCA 2024 · 被引用 21 次
- Beehive: A Scalable Disaggregated Memory Runtime Exploiting Asynchrony of Multithreaded ProgramsQuanxi Li, Hong Huang, Ying Liu, Yanwen Xia 等NSDI 2025 · 被引用 7 次
- Efficient Caching with A Tag-enhanced DRAMMaryam Babaie, Ayaz Akram, Wendy Elsasser, Brent Haukness 等HPCA 2025 · 被引用 2 次
它引用的顶会 Paper4
- Sentinel: Efficient Tensor Migration and Allocation on Heterogeneous Memory Systems for Deep LearningJie Ren, Jiaolin Luo, Kai Wu, Minjia Zhang 等HPCA 2021 · 被引用 62 次
- Hybrid2: Combining Caching and Migration in Hybrid Memory SystemsEvangelos Vasilakis, Vassilis Papaefstathiou, Pedro Trancoso, Ioannis SourdisHPCA 2020 · 被引用 37 次
- Perforated Page: Supporting Fragmented Memory Allocation for Large PagesChang Hyun Park, Sanghoon Cha, Bokyeong Kim, Youngjin Kwon 等ISCA 2020 · 被引用 35 次
- KLOCs: kernel-level object contexts for heterogeneous memory systemsSudarsun Kannan, Yujie Ren, Abhishek BhattacharjeeASPLOS 2021 · 被引用 22 次
相关 Paper
- Genie Cache: Non-Blocking Miss Handling and Replacement in Page-Table-Based DRAM CacheYoungin Kim, William J. SongMICRO 2024 · 被引用 3 次
- Native DRAM Cache: Re-architecting DRAM as a Large-Scale Cache for Data CentersYesin Ryu, Yoojin Kim, Giyong Jung, Jung Ho Ahn 等ISCA 2024 · 被引用 5 次
- A Case for Hardware-Based Demand PagingGyusun Lee, Wenjing Jin, Wonsuk Song, Jeonghun Gong 等ISCA 2020 · 被引用 23 次
- FIGARO: Improving System Performance via Fine-Grained In-DRAM Data Relocation and CachingYaohua Wang, Lois Orosa, Xiangjun Peng, Yang Guo 等MICRO 2020 · 被引用 72 次
- DRAMHiT: A Hash Table Architected for the Speed of DRAMVikram Narayanan, David Detweiler, Tianjiao Huang, Anton BurtsevEuroSys 2023 · 被引用 8 次
