Principle-based Dataflow Optimization for Communication Lower Bound in Operator-Fused Tensor Accelerator
Lei Xu, Chen Yin, Zelong Yuan, Weiguang Sheng, Jianfei Jiang, Qin Wang, Naifeng Jing
Abstract
Although design space exploration (DSE) is good at finding dataflow for optimal memory access in tensor accelerators, it is very timing-consuming and lacks architecture insight. In this study, we for the first time propose several principles for dataflow optimization that provides lower bound of memory communication for tensor operators such as matrix multiplication. Through these principles we can calculate the best tiling, scheduling and mapping for both intra- and inter-operator dataflow. In addition, we can identify all the tensor-wise opertor fusion that are profitable in memory communication, so we propose FuseCU, a new architecture that supports these profitable fusion which can be applied to existing spatial architectures for data movement saving. Experimental results show that FuseCU delivers 63.6%, 62.4% and 38.7% data movement saving and and speedup compared to the TPUv4i, Gemmini and Planaria designs without increasing buffer size or bandwidth. Additionally, FuseCU is open-sourced.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 2e7daf61-3893-4732-80fb-23bf8640feb3Related papers
- Enabling Multiple Tensor-wise Operator Fusion for Transformer Models on Spatial AcceleratorsLei Xu, Zhiwen Mo, Qin Wang, Jianfei Jiang et al.DAC 2024 · 4 citations
- SFD: Towards Segment Fusion Dataflow for Spatial AcceleratorsFuyu Wang, Minghua Shen, Yufei Ding, Nong Xiao et al.HPCA 2026
- A Tensor Marshaling Unit for Sparse Tensor Algebra on General-Purpose ProcessorsMarco Siracusa, Víctor Soria Pardos, Francesco Sgherzi, Joshua Randall et al.MICRO 2023 · 11 citations
- TileFlow: A Framework for Modeling Fusion Dataflow via Tree-based AnalysisSize Zheng, Siyuan Chen, Siyuan Gao, Liancheng Jia et al.MICRO 2023 · 31 citations
- FuseME: Distributed Matrix Computation Engine based on Cuboid-based Fused Operator and Plan GenerationDonghyoung Han, Jongwuk Lee, Min-Soo KimSIGMOD 2022 · 2 citations
