MDM: The GPU Memory Divergence Model
Lu Wang, Magnus Jahre, Almutaz Adileh, Lieven Eeckhout
Abstract
Analytical models enable architects to carry out early-stage design space exploration several orders of magnitude faster than cycle-accurate simulation by capturing first-order performance phenomena with a set of mathematical equations. However, this speed advantage is void if the conclusions obtained through the model are misleading due to model inaccuracies. Therefore, a practical analytical model needs to be sufficiently accurate to capture key performance trends across a broad range of applications and architectural configurations.
In this work, we focus on analytically modeling the performance of emerging memory-divergent GPU-compute applications which are common in domains such as machine learning and data analytics. The poor spatial locality of these applications leads to frequent L1 cache blocking due to the application issuing significantly more concurrent cache misses than the cache can support, which cripples the GPU's ability to use Thread-Level Parallelism (TLP) to hide memory latencies. We propose the GPU Memory Divergence Model (MDM) which faithfully captures the key performance characteristics of memory-divergent applications, including memory request batching and excessive NoC/DRAM queueing delays. We validate MDM against detailed simulation and real hardware, and report substantial improvements in (1) scope: the ability to model prevalent memory-divergent applications in addition to non-memory divergent applications;
(2) practicality: 6.1× faster by computing model inputs using binary instrumentation as opposed to functional simulation; and
(3) accuracy: 13.9% average prediction error versus 162% for the state-of-the-art GPUMech model.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- Predict; Don't React for Enabling Efficient Fine-Grain DVFS in GPUsSrikant Bharadwaj, Shomit Das, Kaushik Mazumdar, Bradford M. Beckmann et al.ASPLOS 2023 · 15 citations
- GPU Scale-Model SimulationHossein SeyyedAghaei, Mahmood Naderan-Tahan, Lieven EeckhoutHPCA 2024 · 13 citations
- Neoscope: How Resilient Is My SoC to Workload Churn?Joseph Rogers, Lieven Eeckhout, Taha Soliman, Magnus JahreISCA 2025 · 2 citations
Builds on1
Related papers
- GCoM: a detailed GPU core model for accurate analytical modeling of modern GPUsJounghoo Lee, Yeonan Ha, Suhyun Lee, Jinyoung Woo et al.ISCA 2022 · 25 citations
- AMALI: An Analytical Model for Accurately Modeling LLM Inference on Modern GPUsShiheng Cao, Junmin Wu, Junshi Chen, Hong An et al.ISCA 2025 · 5 citations
- AGILE: Lightweight and Efficient Asynchronous GPU-SSD IntegrationZhuoping Yang, Jinming Zhuang, Xingzhen Chen, Alex K. Jones et al.SC 2025 · 3 citations
- sfGPUMC: A Stateless Model Checker for GPU Weak Memory ConcurrencySoham Chakraborty, S. Krishna, Andreas Pavlogiannis, Omkar TuppeCAV 2025 · 2 citations
- A Modular Static Cost Analysis for GPU Warp-Level ParallelismGregory Blike, Hannah Zicarelli, Udaya Sathiyamoorthy, Julien Lange et al.POPL 2026 · 1 citation
