Introducing Instruction-Accurate Simulators for Performance Estimation of Autotuning Workloads
Rebecca Pelke, Nils Bosbach, Lennart M. Reimann, Rainer Leupers
摘要
Accelerating Machine Learning (ML) workloads requires efficient methods due to their large optimization space. Autotuning has emerged as an effective approach for systematically evaluating variations of implementations. Traditionally, autotuning requires the workloads to be executed on the target hardware (HW). We present an interface that allows executing autotuning workloads on simulators. This approach offers high scalability when the availability of the target HW is limited, as many simulations can be run in parallel on any accessible HW.
Additionally, we evaluate the feasibility of using fast instruction-accurate simulators for autotuning. We train various predictors to forecast the performance of ML workload implementations on the target HW based on simulation statistics.
Our results demonstrate that the tuned predictors are highly effective. The best workload implementation in terms of actual run time on the target HW is always within the top 3 % of predictions for the tested x86, ARM, and RISC-V-based architectures. In the best case, this approach outperforms native execution on the target HW for embedded architectures when running as few as three samples on three simulators in parallel.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper2
- Ansor: Generating High-Performance Tensor Programs for Deep LearningLianmin Zheng, Chengfan Jia, Minmin Sun, Zhao Wu 等OSDI 2020 · 被引用 551 次
- Analytical characterization and design space exploration for optimization of CNNsRui Li, Yufan Xu, Aravind Sukumaran-Rajam, Atanas Rountev 等ASPLOS 2021 · 被引用 52 次
相关 Paper
- Bootstrapping in-situ workflow auto-tuning via combining performance models of component applicationsTong Shu, Yanfei Guo, Justin M. Wozniak, Xiaoning Ding 等SC 2021 · 被引用 9 次
- TAIDL: Tensor Accelerator ISA Definition Language with Auto-generation of Scalable Test OraclesDevansh Jain, Marco Frigo, Jai Arora, Akash Pardeshi 等MICRO 2025 · 被引用 1 次
- Scalable Deep Learning-Based Microarchitecture Simulation on GPUsSantosh Pandey, Lingda Li, Thomas Flynn, Adolfy Hoisie 等SC 2022 · 被引用 7 次
- Compiler-Driven Simulation of Reconfigurable Hardware AcceleratorsZhijing Li, Yuwei Ye, Stephen Neuendorffer, Adrian SampsonHPCA 2022 · 被引用 4 次
- Fast End-to-End Performance Simulation of Accelerated Hardware-Software StacksJiacheng Ma, Jonas Kaufmann, Emilien Guandalino, Rishabh R. Iyer 等SOSP 2025
