Lune

ISCA2024顶会

Memento: An Adaptive, Compiler-Assisted Register File Cache for GPUs

Mojtaba Abaie Shoushtary, José-María Arnau, Jordi Tubella Murgadas, Antonio González

2024年份
4被引次数
2顶会引用

摘要

Modern GPUs require an enormous register file (RF) to store the context of thousands of active threads. It consumes considerable energy and contains multiple large banks to provide enough throughput. Thus, a RF caching mechanism can significantly improve the performance and energy consumption of the GPUs by avoiding reads from the large banks that consume significant energy and may cause port conflicts. This paper introduces an energy-efficient RF caching mechanism called Memento that repurposes an existing component in GPUs’ RF to operate as a cache in addition to its original functionality. In this way, Memento minimizes the overhead of adding a RF cache to GPUs. Besides, Memento leverages an issue scheduling policy that utilizes the reuse distance of the values in the RF cache and is controlled by a dynamic algorithm. The goal is to adapt the issue policy to the runtime program characteristics to maximize the GPU’s performance and the hit ratio of the RF cache. The reuse distance is approximated by the compiler using profiling and is used at run time by the proposed caching scheme. We show that Memento reduces the number of reads to the RF\mathbf{R F} banks by 46.4%46.4 \% and the dynamic energy of the RF by 28.3%28.3 \%. Besides, it improves performance by 6.1%6.1 \% while adding only 2 KB of extra storage per core to the baseline RF of 256 KB, which represents a negligible overhead of 0.78%\mathbf{0. 7 8 \%}.

问问这篇 Paper

问问你的智能体。

Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。

可以从这些问题问起

智能体调用

Lunesearch_papers

在 Lune 里问

免费开始,无需绑卡

引用它的顶会 Paper2

问问它们各自怎么用它

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖