Manticore: Hardware-Accelerated RTL Simulation with Static Bulk-Synchronous Parallelism
Mahyar Emami, Sahand Kashani, Keisuke Kamahori, Mohammad Sepehr Pourghannad, Ritik Raj, James R. Larus
摘要
The demise of Moore's Law and Dennard Scaling has revived interest in specialized computer architectures and accelerators. Verification and testing of this hardware depend heavily upon cycleaccurate simulation of register-transfer-level (RTL) designs. The fastest software RTL simulators can simulate designs at 1-1000 kHz, i.e., more than three orders of magnitude slower than hardware. Improved simulators can increase designers' productivity by speeding design iterations and permitting more exhaustive exploration.
One possibility is to exploit low-level parallelism, as RTL expresses considerable fine-grain concurrency. Unfortunately, stateof-the-art RTL simulators often perform best on a single core since modern processors cannot effectively exploit fine-grain parallelism.
This work presents Manticore: a parallel computer designed to accelerate RTL simulation. Manticore uses a static bulk-synchronous parallel (BSP) execution model to eliminate fine-grain synchronization overhead. It relies entirely on a compiler to schedule resources and communication, which is feasible since RTL code contains few divergent execution paths. With static scheduling, communication and synchronization no longer incur runtime overhead, making fine-grain parallelism practical. Moreover, static scheduling dramatically simplifies processor implementation, significantly increasing the number of cores that fit on a chip. Our 225-core FPGA implementation running at 475 MHz outperforms a state-of-the-art RTL simulator running on desktop and server computers in 8 out of 9 benchmarks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Accelerating Zero-Knowledge Proofs Through Hardware-Algorithm Co-DesignNikola Samardzic, Simon Langowski, Srinivas Devadas, Daniel SánchezMICRO 2024 · 被引用 24 次
- Parendi: Thousand-Way Parallel RTL SimulationMahyar Emami, Thomas Bourgeat, James R. LarusASPLOS 2025 · 被引用 6 次
- FireAxe: Partitioned FPGA-Accelerated Simulation of Large-Scale RTL DesignsJoonho Whangbo, Edwin Lim, Chengyi Lux Zhang, Kevin Anderson 等ISCA 2024 · 被引用 6 次
- Don't Repeat Yourself! Coarse-Grained Circuit Deduplication to Accelerate RTL SimulationHaoyuan Wang, Thomas Nijssen, Scott BeamerASPLOS 2024 · 被引用 5 次
- DiffTest-H: Toward Semantic-Aware Communication in Hardware-Accelerated Processor VerificationKunlin You, Yinan Xu, Kehan Feng, Luoshan Cai 等MICRO 2025 · 被引用 1 次
它引用的顶会 Paper6
- Think Fast: A Tensor Streaming Processor (TSP) for Accelerating Deep Learning WorkloadsDennis Abts, Jonathan Ross, Jonathan Sparling, Mark Wong-VanHaren 等ISCA 2020 · 被引用 91 次
- Fast stencil-code computation on a wafer-scale processorKamil Rocki, Dirk Van Essendelft, Ilya Sharapov, Robert Schreiber 等SC 2020 · 被引用 69 次
- The essence of Bluespec: a core language for rule-based hardware designThomas Bourgeat, Clément Pit-Claudel, Adam Chlipala, ArvindPLDI 2020 · 被引用 55 次
- A software-defined tensor streaming multiprocessor for large-scale machine learningDennis Abts, Garrin Kimmell, Andrew C. Ling, John Kim 等ISCA 2022 · 被引用 46 次
- Efficiently Exploiting Low Activity Factors to Accelerate RTL SimulationScott Beamer, David DonofrioDAC 2020 · 被引用 36 次
相关 Paper
- Accelerating RTL Simulation with Hardware-Software Co-DesignFares Elsabbagh, Shabnam Sheikhha, Victor A. Ying, Quan M. Nguyen 等MICRO 2023 · 被引用 14 次
- RepCut: Superlinear Parallel RTL Simulation with Replication-Aided PartitioningHaoyuan Wang, Scott BeamerASPLOS 2023 · 被引用 25 次
- Lotus: A Multi-FPGA Task Dataflow Architecture to Accelerate Cycle-Level SimulationFares Elsabbagh, Joel S. Emer, Daniel SánchezISCA 2026
- GSIM: Accelerating RTL Simulation for Large-Scale DesignsLu Chen, Dingyi Zhao, Zihao Yu, Ninghui Sun 等DAC 2025 · 被引用 1 次
- GEM: GPU-Accelerated Emulator-Inspired RTL SimulationZizheng Guo, Yanqing Zhang, Runsheng Wang, Yibo Lin 等DAC 2025 · 被引用 3 次
