Lune

ISCA2026顶会

CODO: An Automated Compiler for Comprehensive Dataflow Optimization

Weichuang Zhang, Yiquan Wang, Xinzhou Zhang, Chi Zhang, Yu Feng, Xiaofeng Hou, Chao Li, Jieru Zhao, Minyi Guo

2026年份

摘要

FPGAs are well-suited for dataflow architectures that process data in a streaming or pipelined manner, thus satisfying the high computational and communication demands of emerging applications. However, manually implementing an efficient dataflow architecture for large-scale applications is still challenging, even for specialists who use high-level synthesis (HLS) to simplify FPGA programming. To address this, we introduce CODO, an automated compiler that generates feasible and efficient dataflow accelerators on FPGAs. CODO features a systematic method for detecting and eliminating both coarse-grained and fine-grained dataflow violations. Building on this, CODO performs both on- and off-chip data movement optimizations to maximize transfer efficiency. To guarantee a higher design quality, CODO performs automatic scheduling to generate high-performance dataflow accelerators, ensuring a balanced performance-resource trade-off. Synthesis results show that CODO delivers 1.45×1.45 \times to 4.52×4.52 \times latency speedups on typical computation kernels and 3.7×3.7 \times to 33.8×33.8 \times speedups on DNN models compared to SOTA frameworks. In on-board evaluations, CODO achieves 7.3×7.3 \times average speedup on CNN models and 2.07×2.07 \times average speedup on the GPT-2 model over SOTA frameworks. The compiler is open-sourced at https://github.com/sjtu-zhao-lab/codo-artifact.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

lune papers fulltext 9b391b48-cafa-48ab-9536-c6e2d792b0bf

它引用的顶会 Paper11

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖