LIBRA: A High-Accuracy, Cost-Aware, and Coordinated Multi-GPU Page Prefetcher
Xiangyue Huang, Yanan Guo, Yuanchao Xu
摘要
Multi-GPU systems increasingly rely on unified virtual memory to satisfy the growing memory demands of modern applications. However, their performance is often limited by noncoherent Non-Uniform Memory Access overheads, where remote accesses are costly and page migration can introduce substantial data-movement overhead. Existing migration techniques are either reactive, placing migration on the critical path, or predictive but designed mainly for CPU-GPU settings. In multi-GPU environments, existing methods, such as NVIDIA's Tree-Based Neighboring Prefetcher and its variants, suffer from low accuracy, overlook the trade-off between remote access and migration, and may cause ping-pong page movements across GPUs. To address these limitations, we propose LIBRA, an access-patternaware, cost-aware, and coordinated page prefetcher for multi-GPU systems. LIBRA uses stride-based prediction to identify GPU memory access patterns, estimates future access benefits to guide migration decisions, and coordinates requests across GPUs based on predicted demand and current page locations. Comprehensive evaluations demonstrate that LIBRA significantly improves performance, outperforming state-of-the-art reactive (GRIT) and predictive (Forest) migration methods by 30% and 35%, respectively.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- GRIT: Enhancing Multi-GPU Performance with Fine-Grained Dynamic Page PlacementYueqi Wang, Bingyao Li, Aamer Jaleel, Jun Yang 等HPCA 2024 · 被引用 19 次
- OASIS: Object-Aware Page Management for Multi-GPU SystemsYueqi Wang, Bingyao Li, Mohamed Tarek Ibn Ziad, Lieven Eeckhout 等HPCA 2025 · 被引用 12 次
- Forest: Access-aware GPU UVM ManagementMao Lin, Yuan Feng, Guilherme Cox, Hyeran JeonISCA 2025 · 被引用 9 次
- LIBRA: Memory Bandwidth- and Locality-Aware Parallel Tile RenderingAurora Tomás, Juan L. Aragón, Joan-Manuel Parcerisa, Antonio GonzálezMICRO 2024 · 被引用 1 次
- Coarse-Grained Duplication First, Fine-Grained Deduplication Later: Duplication-Centric Multi-GPU Memory ManagementXiangyue Huang, Yanan Guo, Yuanchao XuISCA 2026
