GRIT: Enhancing Multi-GPU Performance with Fine-Grained Dynamic Page Placement
Yueqi Wang, Bingyao Li, Aamer Jaleel, Jun Yang, Xulong Tang
摘要
Multi-GPU systems have become popular to cater to the growing demands for high parallelism and large memory capacity. However, the delivered performance is constrained by the non-uniform memory access (NUMA) overhead arising from data sharing and communication across multiple GPUs. Recent multi-GPUs employ unified virtual memory (UVM) to simplify the programming effort. In UVM-enabled multi-GPUs, three popular page placement schemes are adopted to mitigate the NUMA overheads: i) on-touch page migration, ii) access counter-based migration, and iii) page duplication. However, we observe that the preferred page placement scheme varies across i) different applications, ii) different pages of the same application, and iii) even different execution phases of a single page, making it challenging to find a “one-size-fits-all” page placement scheme. To this end, we propose GRIT, which dynamically and automatically determines the appropriate page placement schemes at runtime in a fine-grained manner to enhance multi-GPU performance and scalability. Experimental results indicate that GRIT achieves an average of 60%, 49%, and 29% performance improvements over uniformly adopting on-touch migration, access counter-based migration, and page duplication, respectively.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- OASIS: Object-Aware Page Management for Multi-GPU SystemsYueqi Wang, Bingyao Li, Mohamed Tarek Ibn Ziad, Lieven Eeckhout 等HPCA 2025 · 被引用 12 次
- STAR: Sub-Entry Sharing-Aware TLB for Multi-Instance GPUBingyao Li, Yueqi Wang, Tianyu Wang, Lieven Eeckhout 等MICRO 2024 · 被引用 11 次
- Virtuoso: Enabling Fast and Accurate Virtual Memory Research via an Imitation-based Operating System Simulation MethodologyKonstantinos Kanellopoulos, Konstantinos Sgouras, F. Nisa Bostanci, Andreas Kosmas Kakolyris 等ASPLOS 2025 · 被引用 8 次
- WIC: Hiding Producer-Consumer Synchronization Delays with Warp-Level Interrupt-based GPU CommunicationsJiajian Zhang, Fangyu Wu, Hai Jiang, Qiufeng Wang 等USENIX ATC 2025 · 被引用 1 次
- Revelator: Rapid Data Fetching Via System-Software-Guided Hash-Based Speculative Address TranslationKonstantinos Kanellopoulos, Konstantinos Sgouras, Harsh Songara, Andreas Kosmas Kakolyris 等ISCA 2026
它引用的顶会 Paper16
- Balancing efficiency and fairness in heterogeneous GPU clusters for deep learningShubham Chaudhary, Ramachandran Ramjee, Muthian Sivathanu, Nipun Kwatra 等EuroSys 2020 · 被引用 135 次
- In-depth analyses of unified virtual memory system for GPU accelerated computingTyler N. Allen, Rong GeSC 2021 · 被引用 73 次
- Griffin: Hardware-Software Support for Efficient Page Migration in Multi-GPU SystemsTrinayan Baruah, Yifan Sun, Ali Tolga Dinçer, Saiful A. Mojumder 等HPCA 2020 · 被引用 50 次
- EnvPipe: Performance-preserving DNN Training Framework for Saving EnergySangjin Choi, Inhoe Koo, Jeongseob Ahn, Myeongjae Jeon 等USENIX ATC 2023 · 被引用 41 次
- Locality-Centric Data and Threadblock Management for Massive GPUsMahmoud Khairy, Vadim Nikiforov, David W. Nellans, Timothy G. RogersMICRO 2020 · 被引用 38 次
相关 Paper
- Coarse-Grained Duplication First, Fine-Grained Deduplication Later: Duplication-Centric Multi-GPU Memory ManagementXiangyue Huang, Yanan Guo, Yuanchao XuISCA 2026
- LIBRA: A High-Accuracy, Cost-Aware, and Coordinated Multi-GPU Page PrefetcherXiangyue Huang, Yanan Guo, Yuanchao XuISCA 2026
- ARIADNE: Adaptive UVM Management for Efficient GPU Memory OversubscriptionHyunkyun Shin, Seongtae Bang, Hyungwon Park, Daehoon KimHPCA 2026 · 被引用 2 次
- SUV: Static Analysis Guided Unified Virtual MemoryPratheek B, Guilherme Cox, Ján Veselý, Arkaprava BasuMICRO 2024 · 被引用 7 次
- Reducing Page Faults via Invalidation-Based Mapping Propagation in Multi-GPU SystemsJunsung Kim, Dongho Ha, Sungwoo Kim, Wonho Cho 等ISCA 2026
