FireAxe: Partitioned FPGA-Accelerated Simulation of Large-Scale RTL Designs
Joonho Whangbo, Edwin Lim, Chengyi Lux Zhang, Kevin Anderson, Abraham Gonzalez, Raghav Gupta, Nivedha Krishnakumar, Sagar Karandikar, Borivoje Nikolic, Yakun Sophia Shao, Krste Asanovic
Abstract
Pre-silicon validation and end-to-end system evaluation are integral parts of hardware development as they provide architects with insights about the complex interactions between various hardware components, system software, and application code. Although this process can be accelerated using FPGAs as a simulation host, existing platforms fall short when the resource requirements of a custom hardware design exceed a single FPGA.
We present FireAxe, an open-source FPGA-accelerated RTL simulation platform that supports push-button user-guided partitioning across multiple FPGAs, using a compiler called FireRipper. Given a partition point, FireRipper automatically maps a monolithic RTL design onto multiple FPGAs while providing hardware designers quick feedback about the partition interface and expected simulation performance. Furthermore, FireRipper enables users to choose between an exact-mode which provides cycle-exact results with RTL-level fidelity, or a fast-mode that improves simulation rate while sacrificing fidelity only at the partition boundary. Built on FireSim, FireAxe preserves the ability to elastically scale simulations from on-premises FPGAs to cloud FPGAs. For example, pulling out a core from a systemon-chip (SoC) onto a separate FPGA, we achieve simulation rates of 1.6 MHz using on-premises FPGAs connected by direct-attach cables and 1 MHz on AWS F1 FPGAs using peer-to-peer PCIe.
To show FireAxe's ability to enable pre-silicon performance validation at unprecedented scale, we show several case studies. First, we replicate full-stack system-level effects such as latency spikes from garbage collection in a Golang application on an SoC containing 4 out-of-order (OoO) cores. We also boot Linux on, to our knowledge, the largest OoO core ever cycle-exactly simulated in academia. Lastly, we simulate a system-on-chip containing 24 OoO cores mapped onto five datacenter-class FPGAs. We discover an RTL bug when trying to run Linux user-space applications that did not appear with less substantial software stacks. This was discovered in less than 2 hours using FireAxe and would have taken weeks in a commercial software RTL simulator.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 912a6e7f-6a72-4481-b012-93b4afbc67cdBuilds on7
- Reexamining Direct Cache Access to Optimize I/O Intensive Applications for Multi-hundred-gigabit NetworksAlireza Farshin, Amir Roozbeh, Gerald Q. Maguire Jr., Dejan KosticUSENIX ATC 2020 · 88 citations
- Don't Forget the I/O When Allocating Your LLCYifan Yuan, Mohammad Alian, Yipeng Wang, Ren Wang et al.ISCA 2021 · 37 citations
- RepCut: Superlinear Parallel RTL Simulation with Replication-Aided PartitioningHaoyuan Wang, Scott BeamerASPLOS 2023 · 25 citations
- Manticore: Hardware-Accelerated RTL Simulation with Static Bulk-Synchronous ParallelismMahyar Emami, Sahand Kashani, Keisuke Kamahori, Mohammad Sepehr Pourghannad et al.ASPLOS 2023 · 16 citations
- TIP: Time-Proportional Instruction ProfilingBjörn Gottschall, Lieven Eeckhout, Magnus JahreMICRO 2021 · 16 citations
Related papers
- FirePerf: FPGA-Accelerated Full-System Hardware/Software Performance Profiling and Co-DesignSagar Karandikar, Albert J. Ou, Alon Amid, Howard Mao et al.ASPLOS 2020 · 18 citations
- Parendi: Thousand-Way Parallel RTL SimulationMahyar Emami, Thomas Bourgeat, James R. LarusASPLOS 2025 · 6 citations
- Zoomie: A Software-like Debugging Tool for FPGAsTianrui Wei, Kevin Laeufer, Katie Lim, Jerry Zhao et al.ASPLOS 2024 · 1 citation
- Lotus: A Multi-FPGA Task Dataflow Architecture to Accelerate Cycle-Level SimulationFares Elsabbagh, Joel S. Emer, Daniel SánchezISCA 2026
- Graph.hls: A Compiler Framework for Composable Graph Accelerator DesignFeiyang Wu, Xuxiao Yang, Zhuohang Bian, Jing Wang et al.ISCA 2026 · 1 citation
