Concorde: Fast and Accurate CPU Performance Modeling with Compositional Analytical-ML Fusion
Arash Nasr-Esfahany, Mohammad Alizadeh, Victor Lee, Hanna Alam, Brett W. Coon, David E. Culler, Vidushi Dadu, Martin Dixon, Henry M. Levy, Santosh Pandey, Parthasarathy Ranganathan, Amir Yazdanbakhsh
Abstract
Cycle-level simulators such as gem5 are widely used in microarchitecture design, but they are prohibitively slow for large-scale design space explorations. We present Concorde, a new methodology for learning fast and accurate performance models of microarchitectures. Unlike existing simulators and learning approaches that emulate each instruction, Concorde predicts the behavior of a program based on compact performance distributions that capture the impact of different microarchitectural components. It derives these performance distributions using simple analytical models that estimate bounds on performance induced by each microarchitectural component, providing a simple yet rich representation of a program's performance characteristics across a large space of microarchitectural parameters. Experiments show that Concorde is more than five orders of magnitude faster than a reference cycle-level simulator, with about 2% average Cycles-Per-Instruction (CPI) prediction error across a range of SPEC, open-source, and proprietary benchmarks. This enables rapid design-space exploration and performance sensitivity analyses that are currently infeasible, e.g., in about an hour, we conducted a first-of-its-kind fine-grained performance attribution to different microarchitectural components across a diverse set of programs, requiring nearly 150 million CPI evaluations.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9d1d10ef-1ae1-4165-a32b-d62b8bca372cCited by top-tier papers1
Ask how each one uses itBuilds on6
- PolyGraph: Exposing the Value of Flexibility for Graph Processing AcceleratorsVidushi Dadu, Sihao Liu, Tony NowatzkiISCA 2021 · 60 citations
- Data-Driven Offline Optimization for Architecting Hardware AcceleratorsAviral Kumar, Amir Yazdanbakhsh, Milad Hashemi, Kevin Swersky et al.ICLR 2022 · 45 citations
- Mind the Gap: Attainable Data Movement and Operational Intensity Bounds for Tensor AlgorithmsQijing Huang, Po-An Tsai, Joel S. Emer, Angshuman ParasharISCA 2024 · 11 citations
- Scalable Deep Learning-Based Microarchitecture Simulation on GPUsSantosh Pandey, Lingda Li, Thomas Flynn, Adolfy Hoisie et al.SC 2022 · 7 citations
- Phloem: Automatic Acceleration of Irregular Applications with Fine-Grain Pipeline ParallelismQuan M. Nguyen, Daniel SánchezHPCA 2023 · 7 citations
Related papers
- Learning Generalizable Program and Architecture Representations for Performance ModelingLingda Li, Thomas Flynn, Adolfy HoisieSC 2024 · 5 citations
- Graph Representation Learning for Microarchitecture Design Space ExplorationXiaoling Yi, Jialin Lu, Xiankui Xiong, Dong Xu et al.DAC 2023 · 22 citations
- GCoM: a detailed GPU core model for accurate analytical modeling of modern GPUsJounghoo Lee, Yeonan Ha, Suhyun Lee, Jinyoung Woo et al.ISCA 2022 · 25 citations
- AccelWattch: A Power Modeling Framework for Modern GPUsVijay Kandiah, Scott Peverelle, Mahmoud Khairy, Junrui Pan et al.MICRO 2021 · 134 citations
- AIO: An Abstraction for Performance Analysis Across Diverse Accelerator ArchitecturesJoseph Rogers, Taha Soliman, Magnus JahreISCA 2024 · 5 citations
