OmniSim: Simulating Hardware with C Speed and RTL Accuracy for High-Level Synthesis Designs
Rishov Sarkar, Cong Hao
Abstract
High-Level Synthesis (HLS) is increasingly popular for hardware design using C/C++ instead of Register-Transfer Level (RTL). To express concurrent hardware behavior in a sequential language like C/C++, HLS tools introduce constructs such as infinite loops and dataflow modules connected by FIFOs (first-in first-out). While these constructs can represent concurrency, efficiently and accurately simulating them at C level remains challenging. First, without hardware timing information, functional verification typically requires slow RTL synthesis and simulation, as the current approaches in commercial HLS tools. Second, cycle-accurate performance metrics, such as end-to-end latency or throughput, also rely on RTL simulation. No existing HLS tool fully overcomes the first limitation. For the second, prior work such as LightningSim partially improves simulation speed but lacks support for advanced dataflow features like cyclic dependencies and non-blocking FIFO accesses.
To overcome both limitations, we propose OmniSim, a framework that significantly extends the simulation capabilities of both academic and commercial HLS tools. First, OmniSim enables fast and accurate simulation of complex dataflow designs, especially those explicitly declared unsupported by commercial tools. It does so through sophisticated software multi-threading, where threads are orchestrated by querying and updating a set of FIFO tables that explicitly record exact hardware timing of each FIFO access. Second, OmniSim achieves near-C simulation speed with near-RTL accuracy for both functionality and performance, via flexibly coupled and overlapped functionality and performance simulations.
We demonstrate that OmniSim successfully simulates eleven designs previously unsupported by any HLS tool, achieving up to 35.9× speedup over traditional C/RTL co-simulation, and up to 6.61× speedup over the state-of-the-art yet less capable simulator, LightningSim, on its own benchmark suite.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext fd9262ba-8e8b-4ed0-9db0-e1c68e9cb5e2Builds on2
- FlowGNN: A Dataflow Architecture for Real-Time Workload-Agnostic Graph Neural Network InferenceRishov Sarkar, Stefan Abi-Karam, Yuqi He, Lakshmi Sathidevi et al.HPCA 2023 · 100 citations
- Allo: A Programming Model for Composable Accelerator DesignHongzheng Chen, Niansong Zhang, Shaojie Xiang, Zhichen Zeng et al.PLDI 2024 · 41 citations
Related papers
- Khronos: Fusing Memory Access for Improved Hardware RTL SimulationKexing Zhou, Yun Liang, Yibo Lin, Runsheng Wang et al.MICRO 2023 · 11 citations
- Graphiti: Formally Verified Out-of-Order Execution in Dataflow CircuitsYann Herklotz, Ayatallah Elakhras, Martina Camaioni, Paolo Ienne et al.ASPLOS 2026 · 1 citation
- GEM: GPU-Accelerated Emulator-Inspired RTL SimulationZizheng Guo, Yanqing Zhang, Runsheng Wang, Yibo Lin et al.DAC 2025 · 3 citations
- Manticore: Hardware-Accelerated RTL Simulation with Static Bulk-Synchronous ParallelismMahyar Emami, Sahand Kashani, Keisuke Kamahori, Mohammad Sepehr Pourghannad et al.ASPLOS 2023 · 16 citations
- Hestia: An Efficient Cross-Level Debugger for High-Level SynthesisRuifan Xu, Jin Luo, Yawen Zhang, Yibo Lin et al.MICRO 2024 · 4 citations
