OverGen: Improving FPGA Usability through Domain-specific Overlay Generation
Sihao Liu, Jian Weng, Dylan Kupsh, Atefeh Sohrabizadeh, Zhengrong Wang, Licheng Guo, Jiuyang Liu, Maxim Zhulin, Rishabh Mani, Lucheng Zhang, Jason Cong, Tony Nowatzki
摘要
FPGAs have been proven to be powerful computational accelerators across many types of workloads. The mainstream programming approach is high level synthesis (HLS), which maps high-level languages (e.g. C+ #pragmas) to hardware. Unfortunately, HLS leaves a significant programmability gap in terms of reconfigurability, customization and versatility: Although HLS compilation is fast, the downstream physical design takes hours to days; FPGA reconfiguration time limits the time-multiplexing ability of hardware, and tools do not reason about cross-workload flexibility. Overlay architectures mitigate the above by mapping a programmable design (e.g. CPU, GPU, etc.) on top of FPGAs. However, the abstraction gap between overlay and FPGA leads to low efficiency/utilization. Our essential idea is to develop a hardware generation framework targeting a highly-customizable overlay, so that the abstraction gap can be lowered by tuning the design instance to applications of interest. We leverage and extend prior work on customizable spatial architectures, SoC generation, accelerator compilers, and design space explorers to create an end-to-end FPGA acceleration system. Our novel techniques address inefficient networks between on-chip memories and processing elements, as well as improving DSE by reducing the amount of recompilation required. Our framework, OverGen, is highly competitive with fixed-function HLS-based designs, even though the generated designs are programmable with fast reconfiguration. We compared to a state-of-the-art DSE-based HLS framework, AutoDSE. Without kernel-tuning for AutoDSE, OverGen gets 1.2 geomean performance, and even with manual kernel-tuning for the baseline, OverGen still gets 0.55 geomean performance--all while providing runtime flexibility across workloads.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Infinity Stream: Portable and Programmer-Friendly In-/Near-Memory FusionZhengrong Wang, Christopher Liu, Aman Arora, Lizy Kurian John 等ASPLOS 2023 · 被引用 20 次
- EFFACT: A Highly Efficient Full-Stack FHE Acceleration PlatformYi Huang, Xinsheng Gong, Xiangyu Kong, Dibei Chen 等HPCA 2025 · 被引用 10 次
- Reconfigurable Stream Network ArchitectureChengyue Wang, Xiaofan Zhang, Jason Cong, James C. HoeISCA 2025 · 被引用 8 次
- Assassyn: A Unified Abstraction for Architectural Simulation and ImplementationJian Weng, Boyang Han, Derui Gao, Ruijie Gao 等ISCA 2025 · 被引用 1 次
它引用的顶会 Paper18
- DSAGEN: Synthesizing Programmable Spatial AcceleratorsJian Weng, Sihao Liu, Vidushi Dadu, Zhengrong Wang 等ISCA 2020 · 被引用 140 次
- A Hybrid Systolic-Dataflow Architecture for Inductive Matrix AlgorithmsJian Weng, Sihao Liu, Zhengrong Wang, Vidushi Dadu 等HPCA 2020 · 被引用 80 次
- SynCron: Efficient Synchronization Support for Near-Data-Processing ArchitecturesChristina Giannoula, Nandita Vijaykumar, Nikela Papadopoulou, Vasileios Karakostas 等HPCA 2021 · 被引用 70 次
- Adapt-NoC: A Flexible Network-on-Chip Design for Heterogeneous Manycore ArchitecturesHao Zheng, Ke Wang, Ahmed LouriHPCA 2021 · 被引用 67 次
- REVAMP: a systematic framework for heterogeneous CGRA realizationThilini Kaushalya Bandara, Dhananjaya Wijerathne, Tulika Mitra, Li-Shiuan PehASPLOS 2022 · 被引用 64 次
相关 Paper
- Graph.hls: A Compiler Framework for Composable Graph Accelerator DesignFeiyang Wu, Xuxiao Yang, Zhuohang Bian, Jing Wang 等ISCA 2026 · 被引用 1 次
- Allo: A Programming Model for Composable Accelerator DesignHongzheng Chen, Niansong Zhang, Shaojie Xiang, Zhichen Zeng 等PLDI 2024 · 被引用 41 次
- An Optimizing Framework on MLIR for Efficient FPGA-based Accelerator GenerationWeichuang Zhang, Jieru Zhao, Guan Shen, Quan Chen 等HPCA 2024 · 被引用 8 次
- Predictable accelerator design with time-sensitive affine typesRachit Nigam, Sachille Atapattu, Samuel Thomas, Zhijing Li 等PLDI 2020 · 被引用 58 次
- Cayman: Custom Accelerator Generation with Control Flow and Data Access OptimizationYouwei Xiao, Fan Cui, Zizhang Luo, Weijie Peng 等DAC 2025 · 被引用 1 次
