Fast End-to-End Performance Simulation of Accelerated Hardware-Software Stacks
Jiacheng Ma, Jonas Kaufmann, Emilien Guandalino, Rishabh R. Iyer, Thomas Bourgeat, George Candea
Abstract
The increased use of hardware acceleration has created a need for efficient simulators of the end-to-end performance of accelerated hardware-software stacks: both software and hardware developers need to evaluate the impact of their design choices on overall system performance. However, accurate full-stack simulations are extremely slow, taking hours to simulate just 1 second of real execution. As a result, development of accelerated stacks is non-interactive, and this hurts productivity.
We propose a way to simulate end-to-end performance that is orders-of-magnitude faster yet still accurate. The main idea is to take a minimalist approach: We simulate only those components of the system that are not available, and run the rest natively. Even for unavailable components, we simulate cycleaccurately only aspects that are performance-critical. The key challenge is how to correctly and efficiently synchronize the natively executing components with the simulated ones.
Using this approach, we demonstrate 6× to 879× speedup compared to the state of the art, across three different hardwareaccelerated stacks. The accuracy of simulated time is high: 7% error rate on average and 14% in the worst case, assuming CPU cores are not underprovisioned. Reducing simulation time down to seconds enables interactive development of accelerated stacks, which was until now not possible.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f21d32e1-cda2-440f-a668-3c4b0b3e7a55Builds on7
- Warehouse-scale video acceleration: co-design and deployment in the wildParthasarathy Ranganathan, Daniel Stodolsky, Jeff Calow, Jeremy Dorfman et al.ASPLOS 2021 · 43 citations
- A Hardware Accelerator for Protocol BuffersSagar Karandikar, Chris Leary, Chris Kennelly, Jerry Zhao et al.MICRO 2021 · 42 citations
- SimBricks: end-to-end network system evaluation with modular simulationHejing Li, Jialin Li, Antoine KaufmannSIGCOMM 2022 · 25 citations
- EMISSARY: Enhanced Miss Awareness Replacement Policy for L2 Instruction CachingNayana Prasad Nagendra, Bhargav Reddy Godala, Ishita Chaturvedi, Atmn Patel et al.ISCA 2023 · 17 citations
- LogNIC: A High-Level Performance Model for SmartNICsZerui Guo, Jiaxin Lin, Yuebin Bai, Daehyeok Kim et al.MICRO 2023 · 17 citations
Related papers
- FirePerf: FPGA-Accelerated Full-System Hardware/Software Performance Profiling and Co-DesignSagar Karandikar, Albert J. Ou, Alon Amid, Howard Mao et al.ASPLOS 2020 · 18 citations
- Introducing Instruction-Accurate Simulators for Performance Estimation of Autotuning WorkloadsRebecca Pelke, Nils Bosbach, Lennart M. Reimann, Rainer LeupersDAC 2025
- Performance Interfaces for Hardware AcceleratorsJiacheng Ma, Rishabh R. Iyer, Sahand Kashani, Mahyar Emami et al.OSDI 2024 · 3 citations
- OmniSim: Simulating Hardware with C Speed and RTL Accuracy for High-Level Synthesis DesignsRishov Sarkar, Cong HaoMICRO 2025 · 3 citations
- Virtuoso: Enabling Fast and Accurate Virtual Memory Research via an Imitation-based Operating System Simulation MethodologyKonstantinos Kanellopoulos, Konstantinos Sgouras, F. Nisa Bostanci, Andreas Kosmas Kakolyris et al.ASPLOS 2025 · 8 citations
