SC2022Top-tier venue
Scalable Deep Learning-Based Microarchitecture Simulation on GPUs
Santosh Pandey, Lingda Li, Thomas Flynn, Adolfy Hoisie, Hang Liu
Abstract
Cycle-accurate microarchitecture simulators are es-sential tools for designers to architect, estimate, optimize, and manufacture new processors that meet specific design expectations. However, conventional simulators based on discrete-event methods often require an exceedingly long time-to-solution for the simulation of applications and architectures at full complexity and scale. Given the excitement around wielding the machine learning (ML) hammer to tackle various architecture problems, there have been attempts to employ ML to perform architecture simulations, such as Ithemal and SimNet. However, the direct application of existing ML approaches to architecture simulation may be even slower due to overwhelming memory traffic and stringent sequential computation logic. This work proposes the first graphics processing unit (GPU)-based microarchitecture simulator that fully unleashes the poten-tial of GPUs to accelerate state-of-the-art ML-based simulators. First, considering the application traces are loaded from central processing unit (CPU) to GPU for simulation, we introduce various designs to reduce the data movement cost between CPUs and GPUs. Second, we propose a parallel simulation paradigm that partitions the application trace into sub-traces to simulate them in parallel with rigorous error analysis and effective error correction mechanisms. Combined, this scalable GPU-based simulator outperforms by orders of magnitude the traditional CPU-based simulators and the state-of-the-art ML-based simulators, i.e., SimNet and Ithemal.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6ad6eb8b-d724-48b7-bf10-f0aa34daf6bfCited by top-tier papers2
- Learning Generalizable Program and Architecture Representations for Performance ModelingLingda Li, Thomas Flynn, Adolfy HoisieSC 2024 · 5 citations
- Concorde: Fast and Accurate CPU Performance Modeling with Compositional Analytical-ML FusionArash Nasr-Esfahany, Mohammad Alizadeh, Victor Lee, Hanna Alam et al.ISCA 2025 · 1 citation
Builds on1
Related papers
- GEM: GPU-Accelerated Emulator-Inspired RTL SimulationZizheng Guo, Yanqing Zhang, Runsheng Wang, Yibo Lin et al.DAC 2025 · 3 citations
- Accel-Sim: An Extensible Simulation Framework for Validated GPU ModelingMahmoud Khairy, Zhesheng Shen, Tor M. Aamodt, Timothy G. RogersISCA 2020 · 366 citations
- GeDES: GPU-Driven Discrete Event Network SimulatorQinyong Li, Zhiwei Zhao, Geyong Min, Zi Wang et al.EuroSys 2026 · 1 citation
- GPU Scale-Model SimulationHossein SeyyedAghaei, Mahmood Naderan-Tahan, Lieven EeckhoutHPCA 2024 · 13 citations
- PyTorchSim: A Comprehensive, Fast, and Accurate NPU Simulation FrameworkWonhyuk Yang, Yunseon Shin, Okkyun Woo, Geonwoo Park et al.MICRO 2025 · 6 citations
