SMART: on simultaneously marching racetracks to improve the performance of racetrack-based main memory
Xiangjun Peng, Ming-Chang Yang, Ho Ming Tsui, Chi Ngai Leung, Wang Kang
摘要
RaceTrack Memory (RTM) is a promising media for modern Main Memory subsystems. However, the "shift-before-access" principle, as the nature of RTM, introduces considerable overheads to the access latency. To obtain more insights for the mitigation of shift overheads, this work characterizes and observes that the access patterns, exhibited by the state-of-the-art RTM-based Main Memory, mismatches with the granularity of shift commands (i.e., a group of RaceTracks called Domain Block Cluster (DBC)). Based on the characterization, we propose a novel mechanism called SMART, which simultaneously and proactively marches all DBCs within a subarray, so that subsequent accesses to other DBCs can be served without additional shift commands. Evaluation results show that, averaged across 15 real-world workloads, SMART significantly outperforms other state-of-the-art proposals of RTM-based Main Memory by at least 1.53X in terms of the total execution time, on two different generations of RTM technologies.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- BLOwing Trees to the Ground: Layout Optimization of Decision Trees on Racetrack MemoryChristian Hakert, Asif Ali Khan, Kuan-Hsun Chen, Fazal Hameed 等DAC 2021 · 被引用 10 次
- StreamPIM: Streaming Matrix Computation in Racetrack MemoryYuda An, Yunxiao Tang, Shushu Yi, Li Peng 等HPCA 2024 · 被引用 8 次
- Streamline Ring ORAM Accesses through Spatial and Temporal OptimizationDingyuan Cao, Mingzhe Zhang, Hang Lu, Xiaochun Ye 等HPCA 2021 · 被引用 17 次
- CRISP: critical slice prefetchingHeiner Litz, Grant Ayers, Parthasarathy RanganathanASPLOS 2022 · 被引用 33 次
- Rambda: RDMA-driven Acceleration Framework for Memory-intensive µs-scale Datacenter ApplicationsYifan Yuan, Jinghan Huang, Yan Sun, Tianchen Wang 等HPCA 2023 · 被引用 25 次
