Selective Replication in Memory-Side GPU Caches
Xia Zhao, Magnus Jahre, Lieven Eeckhout
Abstract
Data-intensive applications put immense strain on the memory systems of Graphics Processing Units (GPUs). To cater to this need, GPU memory systems distribute requests across independent units to provide high bandwidth by servicing requests (mostly) in parallel. We find that this strategy breaks down for shared data structures because the shared Last-Level Cache (LLC) organization used by contemporary GPUs stores shared data in a single LLC slice. Shared data requests are hence serialized -resulting in data-intensive applications not being provided with the bandwidth they require. A private LLC organization can provide high bandwidth, but it is often undesirable since it significantly reduces the effective LLC capacity.
In this work, we propose the Selective Replication (SelRep) LLC which selectively replicates shared read-only data across LLC slices to improve bandwidth supply while ensuring that the LLC retains sufficient capacity to keep shared data cached. The compile-time component of SelRep LLC uses dataflow analysis to identify read-only shared data structures and uses a special-purpose load instruction for these accesses. The runtime component of SelRep LLC then monitors the caching behavior of these loads. Leveraging an analytical model, SelRep LLC chooses a replication degree that carefully balances the effective LLC bandwidth benefits of replication against its capacity cost. SelRep LLC consistently provides high performance to replication-sensitive applications across different data set sizes. More specifically, SelRep LLC improves performance by 19.7% and 11.1% on average (and up to 61.6% and 31.0%) compared to the shared LLC baseline and the state-of-the-art Adaptive LLC, respectively.
967
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6c8c0bb7-ee2d-44e3-948a-9dfcba66cfbeCited by top-tier papers6
- Morpheus: Extending the Last Level Cache Capacity in GPU Systems Using Idle GPU Core ResourcesSina Darabi, Mohammad Sadrosadati, Negar Akbarzadeh, Joël Lindegger et al.MICRO 2022 · 24 citations
- SAC: Sharing-Aware Caching in Multi-Chip GPUsShiqing Zhang, Mahmood Naderan-Tahan, Magnus Jahre, Lieven EeckhoutISCA 2023 · 18 citations
- NUBA: Non-Uniform Bandwidth GPUsXia Zhao, Magnus Jahre, Yuhua Tang, Guangda Zhang et al.ASPLOS 2023 · 17 citations
- GPU Scale-Model SimulationHossein SeyyedAghaei, Mahmood Naderan-Tahan, Lieven EeckhoutHPCA 2024 · 13 citations
- Delegated Replies: Alleviating Network Clogging in Heterogeneous ArchitecturesXia Zhao, Lieven Eeckhout, Magnus JahreHPCA 2022 · 9 citations
Builds on1
Related papers
- Analyzing and Leveraging Decoupled L1 Caches in GPUsMohamed Assem Ibrahim, Onur Kayiran, Yasuko Eckert, Gabriel H. Loh et al.HPCA 2021 · 30 citations
- SMILE: LLC-based Shared Memory Expansion to Improve GPU Thread Level ParallelismTianyu Guo, Xuanteng Huang, Kan Wu, Xianwei Zhang et al.DAC 2024 · 3 citations
- The Sparsity-Aware LazyGPU ArchitectureChangxi Liu, Miao Yu, Yifan Sun, Trevor E. CarlsonISCA 2025 · 1 citation
- Snake: A Variable-length Chain-based Prefetching for GPUsSaba Mostofi, Hajar Falahati, Negin Mahani, Pejman Lotfi-Kamran et al.MICRO 2023 · 11 citations
- Scaling your Hybrid CPU-GPU DBMS to Multiple GPUsBobbi W. Yogatama, Weiwei Gong, Xiangyao YuVLDB 2024 · 9 citations
