CODO: An Automated Compiler for Comprehensive Dataflow Optimization
Weichuang Zhang, Yiquan Wang, Xinzhou Zhang, Chi Zhang, Yu Feng, Xiaofeng Hou, Chao Li, Jieru Zhao, Minyi Guo
摘要
FPGAs are well-suited for dataflow architectures that process data in a streaming or pipelined manner, thus satisfying the high computational and communication demands of emerging applications. However, manually implementing an efficient dataflow architecture for large-scale applications is still challenging, even for specialists who use high-level synthesis (HLS) to simplify FPGA programming. To address this, we introduce CODO, an automated compiler that generates feasible and efficient dataflow accelerators on FPGAs. CODO features a systematic method for detecting and eliminating both coarse-grained and fine-grained dataflow violations. Building on this, CODO performs both on- and off-chip data movement optimizations to maximize transfer efficiency. To guarantee a higher design quality, CODO performs automatic scheduling to generate high-performance dataflow accelerators, ensuring a balanced performance-resource trade-off. Synthesis results show that CODO delivers to latency speedups on typical computation kernels and to speedups on DNN models compared to SOTA frameworks. In on-board evaluations, CODO achieves average speedup on CNN models and average speedup on the GPT-2 model over SOTA frameworks. The compiler is open-sourced at https://github.com/sjtu-zhao-lab/codo-artifact.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper11
- Heterogeneous Dataflow Accelerators for Multi-DNN WorkloadsHyoukjun Kwon, Liangzhen Lai, Michael Pellauer, Tushar Krishna 等HPCA 2021 · 被引用 143 次
- DFX: A Low-latency Multi-FPGA Appliance for Accelerating Transformer-based Text GenerationSeongmin Hong, Seungjae Moon, Junsoo Kim, Sungjae Lee 等MICRO 2022 · 被引用 107 次
- Think Fast: A Tensor Streaming Processor (TSP) for Accelerating Deep Learning WorkloadsDennis Abts, Jonathan Ross, Jonathan Sparling, Mark Wong-VanHaren 等ISCA 2020 · 被引用 91 次
- ScaleHLS: A New Scalable High-Level Synthesis Framework on Multi-Level Intermediate RepresentationHanchen Ye, Cong Hao, Jianyi Cheng, Hyunmin Jeong 等HPCA 2022 · 被引用 77 次
- Predictable accelerator design with time-sensitive affine typesRachit Nigam, Sachille Atapattu, Samuel Thomas, Zhijing Li 等PLDI 2020 · 被引用 58 次
相关 Paper
- Memory and Computation Coordinated Mapping of DNNs onto Complex Heterogeneous SoCSize Zheng, Siyuan Chen, Yun LiangDAC 2023 · 被引用 10 次
- A Streaming Collectives Interface Targeting Dataflow Acceleration and HPC WorkloadsNicholas Contini, Jake Queiser, Bharath Ramesh, Hari Subramoni 等SC 2025 · 被引用 1 次
- HIDA: A Hierarchical Dataflow Compiler for High-Level SynthesisHanchen Ye, Hyegang Jun, Deming ChenASPLOS 2024 · 被引用 21 次
- StreamTensor: Make Tensors Stream in Dataflow Accelerators for LLMsHanchen Ye, Deming ChenMICRO 2025 · 被引用 5 次
- Cocktailer: Analyzing and Optimizing Dynamic Control Flow in Deep LearningChen Zhang, Lingxiao Ma, Jilong Xue, Yining Shi 等OSDI 2023 · 被引用 28 次
