Memento: An Adaptive, Compiler-Assisted Register File Cache for GPUs
Mojtaba Abaie Shoushtary, José-María Arnau, Jordi Tubella Murgadas, Antonio González
Abstract
Modern GPUs require an enormous register file (RF) to store the context of thousands of active threads. It consumes considerable energy and contains multiple large banks to provide enough throughput. Thus, a RF caching mechanism can significantly improve the performance and energy consumption of the GPUs by avoiding reads from the large banks that consume significant energy and may cause port conflicts. This paper introduces an energy-efficient RF caching mechanism called Memento that repurposes an existing component in GPUs’ RF to operate as a cache in addition to its original functionality. In this way, Memento minimizes the overhead of adding a RF cache to GPUs. Besides, Memento leverages an issue scheduling policy that utilizes the reuse distance of the values in the RF cache and is controlled by a dynamic algorithm. The goal is to adapt the issue policy to the runtime program characteristics to maximize the GPU’s performance and the hit ratio of the RF cache. The reuse distance is approximated by the compiler using profiling and is used at run time by the proposed caching scheme. We show that Memento reduces the number of reads to the banks by and the dynamic energy of the RF by . Besides, it improves performance by while adding only 2 KB of extra storage per core to the baseline RF of 256 KB, which represents a negligible overhead of .
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 166d2c1d-b4f3-48b8-a32e-d6f8c50934f8Cited by top-tier papers2
- Dissecting and Modeling the Architecture of Modern GPU CoresRodrigo Huerta, Mojtaba Abaie Shoushtary, José-Lorenzo Cruz, Antonio GonzálezMICRO 2025 · 8 citations
- Assassyn: A Unified Abstraction for Architectural Simulation and ImplementationJian Weng, Boyang Han, Derui Gao, Ruijie Gao et al.ISCA 2025 · 1 citation
Related papers
- BOW: Breathing Operand Windows to Exploit Bypassing in GPUsHodjat Asghari Esfeden, AmirAli Abdolrashidi, Shafiur Rahman, Daniel Wong et al.MICRO 2020 · 19 citations
- Concurrency-Aware Register Stacks for Efficient GPU Function CallsNi Kang, Ahmad Alawneh, Mengchi Zhang, Timothy G. RogersMICRO 2024 · 1 citation
- Exploiting Zero Data to Reduce Register File and Execution Unit Dynamic Power Consumption in GPGPUsAhmad M. Radaideh, Paul V. GratzDAC 2020 · 4 citations
- Warped-Compaction: Maximizing GPU Register File Bandwidth Utilization via Operand CompactionEunbi Jeong, Ipoom Jeong, Myung Kuk Yoon, Nam Sung KimHPCA 2025 · 2 citations
- Morpheus: Extending the Last Level Cache Capacity in GPU Systems Using Idle GPU Core ResourcesSina Darabi, Mohammad Sadrosadati, Negar Akbarzadeh, Joël Lindegger et al.MICRO 2022 · 24 citations
