Lotus: A Multi-FPGA Task Dataflow Architecture to Accelerate Cycle-Level Simulation
Fares Elsabbagh, Joel S. Emer, Daniel Sánchez
Abstract
Simulation is crucial to design and build hardware. But simulating large and complex digital designs is slow. Hardware emulators are the standard accelerator for cycle-level RTL simulation, but these systems are expensive, inefficient, slow to compile for, and limited to simulating RTL. Emulators consist of many chips, typically FPGAs, to which the design is compiled. Emulators are bottlenecked by communication, and use FPGAs at a fraction of their speed. We present Lotus, a large-scale architecture that accelerates cycle-level simulation. Lotus uses multiple FPGAs like emulators, but takes a different approach: rather than mapping logic directly to FPGAs, Lotus implements thousands of simple cores, along with hardware support that enables software simulation to scale. Lotus simulates digital systems by encoding them as large dataflow graphs of tiny tasks that run on these cores. Lotus uses dataflow execution to extract abundant parallelism; task priorities to focus work on the critical path; and selective execution to avoid ineffectual work. We contribute new implementations of these techniques that scale to multiple chips and require simple hardware. We also develop a compiler to use Lotus efficiently from high-level dataflow graphs. We build an implementation of Lotus using 8 FPGAs, featuring over two thousand cores. On several large designs, this Lotus prototype achieves speeds comparable to emulators, while reducing the number of FPGAs needed by up to and improving performance per FPGA by up to . Lotus is also faster than a 128-core server. Overall, Lotus is the first system to show that software simulation can outperform emulators by leveraging large-scale parallelism.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get cd204e49-2df8-42d6-b66f-e4007e4bbfe3Related papers
- Accelerating RTL Simulation with Hardware-Software Co-DesignFares Elsabbagh, Shabnam Sheikhha, Victor A. Ying, Quan M. Nguyen et al.MICRO 2023 · 14 citations
- Manticore: Hardware-Accelerated RTL Simulation with Static Bulk-Synchronous ParallelismMahyar Emami, Sahand Kashani, Keisuke Kamahori, Mohammad Sepehr Pourghannad et al.ASPLOS 2023 · 16 citations
- GEM: GPU-Accelerated Emulator-Inspired RTL SimulationZizheng Guo, Yanqing Zhang, Runsheng Wang, Yibo Lin et al.DAC 2025 · 3 citations
- FireAxe: Partitioned FPGA-Accelerated Simulation of Large-Scale RTL DesignsJoonho Whangbo, Edwin Lim, Chengyi Lux Zhang, Kevin Anderson et al.ISCA 2024 · 6 citations
- OmniSim: Simulating Hardware with C Speed and RTL Accuracy for High-Level Synthesis DesignsRishov Sarkar, Cong HaoMICRO 2025 · 3 citations
