Parendi: Thousand-Way Parallel RTL Simulation
Mahyar Emami, Thomas Bourgeat, James R. Larus
Abstract
Hardware development critically depends on cycle-accurate RTL simulation. However, as chip complexity increases, conventional single-threaded simulation becomes impractical due to stagnant single-core performance.
Parendi is an RTL simulator that addresses this challenge by exploiting the abundant fine-grained parallelism inherent in RTL simulation and efficiently mapping it onto the massively parallel Graphcore IPU (Intelligence Processing Unit) architecture. Parendi scales up to 5888 cores on 4 Graphcore IPU sockets. It allows us to run large RTL designs up to 4× faster than the most powerful state-of-the-art x64 multicore systems.
To achieve this performance, we developed new partitioning and compilation techniques and carefully quantified the synchronization, communication, and computation costs of parallel RTL simulation: The paper comprehensively analyzes these factors and details the strategies that Parendi uses to optimize them.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 119170f9-d642-4a81-be3a-8d18db2a6106Cited by top-tier papers1
Ask how each one uses itBuilds on7
- Think Fast: A Tensor Streaming Processor (TSP) for Accelerating Deep Learning WorkloadsDennis Abts, Jonathan Ross, Jonathan Sparling, Mark Wong-VanHaren et al.ISCA 2020 · 91 citations
- A software-defined tensor streaming multiprocessor for large-scale machine learningDennis Abts, Garrin Kimmell, Andrew C. Ling, John Kim et al.ISCA 2022 · 46 citations
- Efficiently Exploiting Low Activity Factors to Accelerate RTL SimulationScott Beamer, David DonofrioDAC 2020 · 36 citations
- RepCut: Superlinear Parallel RTL Simulation with Replication-Aided PartitioningHaoyuan Wang, Scott BeamerASPLOS 2023 · 25 citations
- Manticore: Hardware-Accelerated RTL Simulation with Static Bulk-Synchronous ParallelismMahyar Emami, Sahand Kashani, Keisuke Kamahori, Mohammad Sepehr Pourghannad et al.ASPLOS 2023 · 16 citations
Related papers
- Accelerating RTL Simulation with Hardware-Software Co-DesignFares Elsabbagh, Shabnam Sheikhha, Victor A. Ying, Quan M. Nguyen et al.MICRO 2023 · 14 citations
- GEM: GPU-Accelerated Emulator-Inspired RTL SimulationZizheng Guo, Yanqing Zhang, Runsheng Wang, Yibo Lin et al.DAC 2025 · 3 citations
- Don't Repeat Yourself! Coarse-Grained Circuit Deduplication to Accelerate RTL SimulationHaoyuan Wang, Thomas Nijssen, Scott BeamerASPLOS 2024 · 5 citations
- FireAxe: Partitioned FPGA-Accelerated Simulation of Large-Scale RTL DesignsJoonho Whangbo, Edwin Lim, Chengyi Lux Zhang, Kevin Anderson et al.ISCA 2024 · 6 citations
- Khronos: Fusing Memory Access for Improved Hardware RTL SimulationKexing Zhou, Yun Liang, Yibo Lin, Runsheng Wang et al.MICRO 2023 · 11 citations
