OverGen: Improving FPGA Usability through Domain-specific Overlay Generation
Sihao Liu, Jian Weng, Dylan Kupsh, Atefeh Sohrabizadeh, Zhengrong Wang, Licheng Guo, Jiuyang Liu, Maxim Zhulin, Rishabh Mani, Lucheng Zhang, Jason Cong, Tony Nowatzki
Abstract
FPGAs have been proven to be powerful computational accelerators across many types of workloads. The mainstream programming approach is high level synthesis (HLS), which maps high-level languages (e.g. C+ #pragmas) to hardware. Unfortunately, HLS leaves a significant programmability gap in terms of reconfigurability, customization and versatility: Although HLS compilation is fast, the downstream physical design takes hours to days; FPGA reconfiguration time limits the time-multiplexing ability of hardware, and tools do not reason about cross-workload flexibility. Overlay architectures mitigate the above by mapping a programmable design (e.g. CPU, GPU, etc.) on top of FPGAs. However, the abstraction gap between overlay and FPGA leads to low efficiency/utilization. Our essential idea is to develop a hardware generation framework targeting a highly-customizable overlay, so that the abstraction gap can be lowered by tuning the design instance to applications of interest. We leverage and extend prior work on customizable spatial architectures, SoC generation, accelerator compilers, and design space explorers to create an end-to-end FPGA acceleration system. Our novel techniques address inefficient networks between on-chip memories and processing elements, as well as improving DSE by reducing the amount of recompilation required. Our framework, OverGen, is highly competitive with fixed-function HLS-based designs, even though the generated designs are programmable with fast reconfiguration. We compared to a state-of-the-art DSE-based HLS framework, AutoDSE. Without kernel-tuning for AutoDSE, OverGen gets 1.2 geomean performance, and even with manual kernel-tuning for the baseline, OverGen still gets 0.55 geomean performance--all while providing runtime flexibility across workloads.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext adfa78c1-2e01-43dc-a41d-37f85a1a2906Cited by top-tier papers4
- Infinity Stream: Portable and Programmer-Friendly In-/Near-Memory FusionZhengrong Wang, Christopher Liu, Aman Arora, Lizy Kurian John et al.ASPLOS 2023 · 20 citations
- EFFACT: A Highly Efficient Full-Stack FHE Acceleration PlatformYi Huang, Xinsheng Gong, Xiangyu Kong, Dibei Chen et al.HPCA 2025 · 10 citations
- Reconfigurable Stream Network ArchitectureChengyue Wang, Xiaofan Zhang, Jason Cong, James C. HoeISCA 2025 · 8 citations
- Assassyn: A Unified Abstraction for Architectural Simulation and ImplementationJian Weng, Boyang Han, Derui Gao, Ruijie Gao et al.ISCA 2025 · 1 citation
Builds on18
- DSAGEN: Synthesizing Programmable Spatial AcceleratorsJian Weng, Sihao Liu, Vidushi Dadu, Zhengrong Wang et al.ISCA 2020 · 140 citations
- A Hybrid Systolic-Dataflow Architecture for Inductive Matrix AlgorithmsJian Weng, Sihao Liu, Zhengrong Wang, Vidushi Dadu et al.HPCA 2020 · 80 citations
- SynCron: Efficient Synchronization Support for Near-Data-Processing ArchitecturesChristina Giannoula, Nandita Vijaykumar, Nikela Papadopoulou, Vasileios Karakostas et al.HPCA 2021 · 70 citations
- Adapt-NoC: A Flexible Network-on-Chip Design for Heterogeneous Manycore ArchitecturesHao Zheng, Ke Wang, Ahmed LouriHPCA 2021 · 67 citations
- REVAMP: a systematic framework for heterogeneous CGRA realizationThilini Kaushalya Bandara, Dhananjaya Wijerathne, Tulika Mitra, Li-Shiuan PehASPLOS 2022 · 64 citations
Related papers
- Graph.hls: A Compiler Framework for Composable Graph Accelerator DesignFeiyang Wu, Xuxiao Yang, Zhuohang Bian, Jing Wang et al.ISCA 2026 · 1 citation
- Allo: A Programming Model for Composable Accelerator DesignHongzheng Chen, Niansong Zhang, Shaojie Xiang, Zhichen Zeng et al.PLDI 2024 · 41 citations
- An Optimizing Framework on MLIR for Efficient FPGA-based Accelerator GenerationWeichuang Zhang, Jieru Zhao, Guan Shen, Quan Chen et al.HPCA 2024 · 8 citations
- Predictable accelerator design with time-sensitive affine typesRachit Nigam, Sachille Atapattu, Samuel Thomas, Zhijing Li et al.PLDI 2020 · 58 citations
- Cayman: Custom Accelerator Generation with Control Flow and Data Access OptimizationYouwei Xiao, Fan Cui, Zizhang Luo, Weijie Peng et al.DAC 2025 · 1 citation
