Manticore: Hardware-Accelerated RTL Simulation with Static Bulk-Synchronous Parallelism
Mahyar Emami, Sahand Kashani, Keisuke Kamahori, Mohammad Sepehr Pourghannad, Ritik Raj, James R. Larus
Abstract
The demise of Moore's Law and Dennard Scaling has revived interest in specialized computer architectures and accelerators. Verification and testing of this hardware depend heavily upon cycleaccurate simulation of register-transfer-level (RTL) designs. The fastest software RTL simulators can simulate designs at 1-1000 kHz, i.e., more than three orders of magnitude slower than hardware. Improved simulators can increase designers' productivity by speeding design iterations and permitting more exhaustive exploration.
One possibility is to exploit low-level parallelism, as RTL expresses considerable fine-grain concurrency. Unfortunately, stateof-the-art RTL simulators often perform best on a single core since modern processors cannot effectively exploit fine-grain parallelism.
This work presents Manticore: a parallel computer designed to accelerate RTL simulation. Manticore uses a static bulk-synchronous parallel (BSP) execution model to eliminate fine-grain synchronization overhead. It relies entirely on a compiler to schedule resources and communication, which is feasible since RTL code contains few divergent execution paths. With static scheduling, communication and synchronization no longer incur runtime overhead, making fine-grain parallelism practical. Moreover, static scheduling dramatically simplifies processor implementation, significantly increasing the number of cores that fit on a chip. Our 225-core FPGA implementation running at 475 MHz outperforms a state-of-the-art RTL simulator running on desktop and server computers in 8 out of 9 benchmarks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 14aa4e6f-3efc-425b-9e70-2c1e3b30ab25Cited by top-tier papers6
- Accelerating Zero-Knowledge Proofs Through Hardware-Algorithm Co-DesignNikola Samardzic, Simon Langowski, Srinivas Devadas, Daniel SánchezMICRO 2024 · 24 citations
- Parendi: Thousand-Way Parallel RTL SimulationMahyar Emami, Thomas Bourgeat, James R. LarusASPLOS 2025 · 6 citations
- FireAxe: Partitioned FPGA-Accelerated Simulation of Large-Scale RTL DesignsJoonho Whangbo, Edwin Lim, Chengyi Lux Zhang, Kevin Anderson et al.ISCA 2024 · 6 citations
- Don't Repeat Yourself! Coarse-Grained Circuit Deduplication to Accelerate RTL SimulationHaoyuan Wang, Thomas Nijssen, Scott BeamerASPLOS 2024 · 5 citations
- DiffTest-H: Toward Semantic-Aware Communication in Hardware-Accelerated Processor VerificationKunlin You, Yinan Xu, Kehan Feng, Luoshan Cai et al.MICRO 2025 · 1 citation
Builds on6
- Think Fast: A Tensor Streaming Processor (TSP) for Accelerating Deep Learning WorkloadsDennis Abts, Jonathan Ross, Jonathan Sparling, Mark Wong-VanHaren et al.ISCA 2020 · 91 citations
- Fast stencil-code computation on a wafer-scale processorKamil Rocki, Dirk Van Essendelft, Ilya Sharapov, Robert Schreiber et al.SC 2020 · 69 citations
- The essence of Bluespec: a core language for rule-based hardware designThomas Bourgeat, Clément Pit-Claudel, Adam Chlipala, ArvindPLDI 2020 · 55 citations
- A software-defined tensor streaming multiprocessor for large-scale machine learningDennis Abts, Garrin Kimmell, Andrew C. Ling, John Kim et al.ISCA 2022 · 46 citations
- Efficiently Exploiting Low Activity Factors to Accelerate RTL SimulationScott Beamer, David DonofrioDAC 2020 · 36 citations
Related papers
- Accelerating RTL Simulation with Hardware-Software Co-DesignFares Elsabbagh, Shabnam Sheikhha, Victor A. Ying, Quan M. Nguyen et al.MICRO 2023 · 14 citations
- RepCut: Superlinear Parallel RTL Simulation with Replication-Aided PartitioningHaoyuan Wang, Scott BeamerASPLOS 2023 · 25 citations
- Lotus: A Multi-FPGA Task Dataflow Architecture to Accelerate Cycle-Level SimulationFares Elsabbagh, Joel S. Emer, Daniel SánchezISCA 2026
- GSIM: Accelerating RTL Simulation for Large-Scale DesignsLu Chen, Dingyi Zhao, Zihao Yu, Ninghui Sun et al.DAC 2025 · 1 citation
- GEM: GPU-Accelerated Emulator-Inspired RTL SimulationZizheng Guo, Yanqing Zhang, Runsheng Wang, Yibo Lin et al.DAC 2025 · 3 citations
