Barre Chord: Efficient Virtual Memory Translation for Multi-Chip-Module GPUs
Yuan Feng, Seonjin Na, Hyesoon Kim, Hyeran Jeon
摘要
With the advancement of processor packaging technology and the looming end of Moore’s law, multi-chip-module (MCM) GPUs become a promising architecture to continue the performance scaling. However, due to the increasing concurrency, it is challenging to achieve scalable performance. In this study, we show that the limited parallelism in IOMMU is one of the critical bottlenecks and propose Barre Chord to fundamentally reduce the translation loads. By leveraging the unique GPU execution model and page mapping on MCM-GPUs, Barre translates virtual addresses in a unit of coalescing group. Once one page is translated, all the other pages within the same coalescing group can be translated with simple calculations without page table walks. Full Barre (F-Barre) further reduces translations by enabling intra-MCM translation through coalescing information sharing across GPU chiplets and contiguity-aware coalescing group expansion. With the combination of Barre and F-Barre, the Barre Chord outperforms state-of-the-art solutions by an average of 1.36× (2.09× with coalescing group expansion) with negligible area overhead (4.22% of a GPU L2 TLB).
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- A Case for Speculative Address Translation with Rapid Validation for GPUsJunhyeok Park, Osang Kwon, Yongho Lee, Seongwook Kim 等MICRO 2024 · 被引用 13 次
- HDPAT: Hierarchical Distributed Page Address Translation for Wafer-Scale GPUsDaoxuan Xu, Ying Li, Yuwei Sun, Jie Ren 等HPCA 2026 · 被引用 1 次
- Assassyn: A Unified Abstraction for Architectural Simulation and ImplementationJian Weng, Boyang Han, Derui Gao, Ruijie Gao 等ISCA 2025 · 被引用 1 次
- BCCE: Block-Centric GPU Co-Design for Real-Time Range-Top-K Query at ScaleChengying Huan, Ziheng Meng, Zhengyi Yang, Yongchao Liu 等HPDC 2026
- Revelator: Rapid Data Fetching Via System-Software-Guided Hash-Based Speculative Address TranslationKonstantinos Kanellopoulos, Konstantinos Sgouras, Harsh Songara, Andreas Kosmas Kakolyris 等ISCA 2026
它引用的顶会 Paper18
- Batch-Aware Unified Memory Management in GPUs for Irregular WorkloadsHyojong Kim, Jaewoong Sim, Prasun Gera, Ramyad Hadidi 等ASPLOS 2020 · 被引用 89 次
- Mitosis: Transparently Self-Replicating Page-Tables for Large-Memory MachinesReto Achermann, Ashish Panwar, Abhishek Bhattacharjee, Timothy Roscoe 等ASPLOS 2020 · 被引用 62 次
- Griffin: Hardware-Software Support for Efficient Page Migration in Multi-GPU SystemsTrinayan Baruah, Yifan Sun, Ali Tolga Dinçer, Saiful A. Mojumder 等HPCA 2020 · 被引用 50 次
- Enhancing and Exploiting Contiguity for Fast Memory VirtualizationChloe Alverti, Stratos Psomadakis, Vasileios Karakostas, Jayneel Gandhi 等ISCA 2020 · 被引用 41 次
- Locality-Centric Data and Threadblock Management for Massive GPUsMahmoud Khairy, Vadim Nikiforov, David W. Nellans, Timothy G. RogersMICRO 2020 · 被引用 38 次
相关 Paper
- Leveraging Chiplet-Locality for Efficient Memory Mapping in Multi-Chip Module GPUsJunhyeok Park, Sungbin Jang, Osang Kwon, Yongho Lee 等MICRO 2025 · 被引用 7 次
- Designing Virtual Memory System of MCM GPUsPratheek B, Neha Jawalkar, Arkaprava BasuMICRO 2022 · 被引用 21 次
- LRM-GPU: Alleviating Synchronization Overhead for Multi-Chiplet GPU ArchitectureBaiqing Zhong, Zhirong Ye, Xiaojie Li, Peilin Wang 等HPCA 2026
- Marching Page Walks: Batching and Concurrent Page Table Walks for Enhancing GPU ThroughputJiwon Lee, Gun Ko, Myung Kuk Yoon, Ipoom Jeong 等HPCA 2025 · 被引用 4 次
- IDYLL: Enhancing Page Translation in Multi-GPUs via Light Weight PTE InvalidationsBingyao Li, Yanan Guo, Yueqi Wang, Aamer Jaleel 等MICRO 2023 · 被引用 16 次
