DiffTune: Optimizing CPU Simulator Parameters with Learned Differentiable Surrogates
Alex Renda, Yishen Chen, Charith Mendis, Michael Carbin
Abstract
CPU simulators are useful tools for modeling CPU execution behavior. However, they suffer from inaccuracies due to the cost and complexity of setting their fine-grained parameters, such as the latencies of individual instructions. This complexity arises from the expertise required to design benchmarks and measurement frameworks that can precisely measure the values of parameters at such fine granularity. In some cases, these parameters do not necessarily have a physical realization and are therefore fundamentally approximate, or even unmeasurable.
In this paper we present DiffTune, a system for learning the parameters of x86 basic block CPU simulators from coarsegrained end-to-end measurements. Given a simulator, DiffTune learns its parameters by first replacing the original simulator with a differentiable surrogate, another function that approximates the original function; by making the surrogate differentiable, DiffTune is then able to apply gradient-based optimization techniques even when the original function is non-differentiable, such as is the case with CPU simulators. With this differentiable surrogate, DiffTune then applies gradient-based optimization to produce values of the simulator's parameters that minimize the simulator's error on a dataset of ground truth end-to-end performance measurements. Finally, the learned parameters are plugged back into the original simulator. DiffTune is able to automatically learn the entire set of microarchitecture-specific parameters within the Intel x86 simulation model of llvm-mca, a basic block CPU simulator based on LLVM's instruction scheduling model. DiffTune's learned parameters lead llvm-mca to an average error that not only matches but lowers that of its original, expert-provided parameter values.
Discussion. Note the similarity between Equation (1) and Equation (3): the two equations only differ by the use of 𝑓 and f , respectively. The close correspondence between forms makes clear that f stands in as a surrogate for the original program, 𝑓 . This is a general algorithmic approach [15] that is desirable when it is possible to choose f such that it is easier or more efficient to optimize 𝜃 using f than 𝑓 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 91246317-4ee1-428d-b1c1-80fb05bf25bfCited by top-tier papers8
- Mind mappings: enabling efficient algorithm-accelerator mapping space searchKartik Hegde, Po-An Tsai, Sitao Huang, Vikas Chandra et al.ASPLOS 2021 · 95 citations
- Unsupervised Learning for Combinatorial Optimization with Principled Objective RelaxationHaoyu Wang, Nan Wu, Hang Yang, Cong Hao et al.NeurIPS 2022 · 54 citations
- APT-GET: profile-guided timely software prefetchingSaba Jamilan, Tanvir Ahmed Khan, Grant Ayers, Baris Kasikci et al.EuroSys 2022 · 25 citations
- Amortized Synthesis of Constrained Configurations Using a Differentiable SurrogateXingyuan Sun, Tianju Xue, Szymon Rusinkiewicz, Ryan P. AdamsNeurIPS 2021 · 14 citations
- ZeroGrads: Learning Local Surrogates for Non-Differentiable GraphicsMichael Fischer, Tobias RitschelSIGGRAPH 2024 · 7 citations
Builds on4
- NEUZZ: Efficient Fuzzing with Neural Program SmoothingDongdong She, Kexin Pei, Dave Epstein, Junfeng Yang et al.S&P 2019 · 220 citations
- Black-Box Optimization with Local Generative SurrogatesSergey Shirobokov, Vladislav Belavin, Michael Kagan, Andrey Ustyuzhanin et al.NeurIPS 2020 · 60 citations
- CacheQuery: learning replacement policies from hardware cachesPepe Vila, Pierre Ganty, Marco Guarnieri, Boris KöpfPLDI 2020 · 38 citations
- PMEvo: portable inference of port mappings for out-of-order processors by evolutionary optimizationFabian Ritter, Sebastian HackPLDI 2020 · 13 citations
Related papers
- Concorde: Fast and Accurate CPU Performance Modeling with Compositional Analytical-ML FusionArash Nasr-Esfahany, Mohammad Alizadeh, Victor Lee, Hanna Alam et al.ISCA 2025 · 1 citation
- Introducing Instruction-Accurate Simulators for Performance Estimation of Autotuning WorkloadsRebecca Pelke, Nils Bosbach, Lennart M. Reimann, Rainer LeupersDAC 2025
- AdaTune: Adaptive Tensor Program Compilation Made EfficientMenghao Li, Minjia Zhang, Chi Wang, Mingqin LiNeurIPS 2020 · 39 citations
- HYPERF: End-to-End Autotuning Framework for High-Performance ComputingJuseong Park, Yongwon Shin, Junghyun Lee, Junseo Lee et al.HPDC 2025 · 2 citations
- Scalable Deep Learning-Based Microarchitecture Simulation on GPUsSantosh Pandey, Lingda Li, Thomas Flynn, Adolfy Hoisie et al.SC 2022 · 7 citations
