ACRS: Adjacent Computation Resource Sharing among Partitioned GPU Sub-Cores
Penghao Song, Chongxi Wang, Chenji Han, Haoyu Zhao, Tingting Zhang, Tianyi Liu, Jian Wang
Abstract
Modern GPUs typically segment Streaming Multiprocessors (SMs) into sub-cores (e.g. 4 sub-cores) to reduce power consumption and chip area. However, this partitioned design prevents potential task distributions across sub-cores, impairing overall execution efficiency. In this paper, we explore the performance benefit of sharing hardware resources among sub-cores and identify functional units (FUs) as critical components for compute-intensive applications. Moreover, our observations reveal that instructions residing in operand collectors can be obstructed by back-end FUs, but there is a high probability that unoccupied FUs are available in adjacent sub-cores during such blockages. In response, we introduce the adjacent computation resource sharing (ACRS) framework to efficiently utilize these unoccupied units among sub-cores. ACRS has two key modules: Shared FU Issue (SF_ISSUE) and Shared FU Write Back (SF_WriteBack). SF_ISSUE monitors the status of operand collectors and functional units, and offloads instructions from blocked sub-cores to unoccupied resources. Meanwhile, SF_WriteBack routes results back to the original sub-core.To minimize wiring overhead, each sub-core is assigned a fixed target core for sharing. We design a series of matching policies and finally filter out the most effective sequential method. Evaluation results show that ACRS improves performance by up to , with an average of over the traditional partitioned architecture, while reducing energy consumption by . Besides, ACRS achieves an additional 12.3% performance improvement compared with the SOTA method.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 32cd7c3f-f242-4ffc-9c6f-1a4713f1e876Related papers
- Mitigating GPU Core Partitioning Performance EffectsAaron Barnes, Fangjia Shen, Timothy G. RogersHPCA 2023 · 18 citations
- Warped-Compaction: Maximizing GPU Register File Bandwidth Utilization via Operand CompactionEunbi Jeong, Ipoom Jeong, Myung Kuk Yoon, Nam Sung KimHPCA 2025 · 2 citations
- VISTA: Optimizing GPU Scheduling through Versatile Locality-Aware Data SharingHajar Falahati, Negin Mahani, Adrián Cristal, Osman S. UnsalDAC 2025 · 1 citation
- Concurrency-Aware Register Stacks for Efficient GPU Function CallsNi Kang, Ahmad Alawneh, Mengchi Zhang, Timothy G. RogersMICRO 2024 · 1 citation
- µShare: Non-Intrusive Kernel Co-Locating on NVIDIA GPUsWenhao Huang, Zhaolin Duan, Laiping Zhao, Yuhao Zhang et al.HPCA 2026
