TileLoom: Automatic Dataflow Planning for Tile-Based Languages on Spatial Dataflow Accelerators
Wei Li, Zhenyu Bai, Heru Wang, Pranav Dangi, Zhiqiang Zhang, Cheng Tan, Huiying Lan, Weng-Fai Wong, Tulika Mitra
Abstract
Spatial dataflow accelerators are a promising direction for next-generation computer systems because they can reduce the memory bottlenecks of traditional von Neumann machines such as CPUs and GPUs. They organize computation around explicit, compiler-managed data movement over on-chip networks, allowing operands to be forwarded directly between processing elements and reducing reliance on high-latency, bandwidth-limited global shared memory. However, their performance depends strongly on how workloads are mapped to hardware. Naive mappings can perform poorly, and most users rely on hand-tuned vendor libraries. Thus, despite their potential for high performance, energy efficiency, and cost efficiency, limited programmability remains a major barrier to wider adoption. This paper presents TileLoom, an MLIR-based end-to-end framework that compiles tile-based programs, such as Triton kernels, onto spatial dataflow architectures. Unlike compiler frameworks that focus on optimizing code generation within a single tile, TileLoom distributes tile instances across spatially distributed cores and exploits the on-chip network and distributed memories to increase data reuse and reduce communication. TileLoom introduces a hardware representation that captures interconnect topology, memory hierarchy, and compute capabilities, enabling both architecture-specific optimizations and support for diverse spatial dataflow targets. In experiments on two generations of Tenstorrent systems, TileLoom achieves performance comparable to vendor libraries on various kernels.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on9
- DSAGEN: Synthesizing Programmable Spatial AcceleratorsJian Weng, Sihao Liu, Vidushi Dadu, Zhengrong Wang et al.ISCA 2020 · 140 citations
- Think Fast: A Tensor Streaming Processor (TSP) for Accelerating Deep Learning WorkloadsDennis Abts, Jonathan Ross, Jonathan Sparling, Mark Wong-VanHaren et al.ISCA 2020 · 91 citations
- AMOS: enabling automatic mapping for tensor computations on spatial accelerators with hardware abstractionSize Zheng, Renze Chen, Anjiang Wei, Yicheng Jin et al.ISCA 2022 · 63 citations
- TensorLib: A Spatial Accelerator Generation Framework for Tensor AlgebraLiancheng Jia, Zizhang Luo, Liqiang Lu, Yun LiangDAC 2021 · 49 citations
- A software-defined tensor streaming multiprocessor for large-scale machine learningDennis Abts, Garrin Kimmell, Andrew C. Ling, John Kim et al.ISCA 2022 · 46 citations
Related papers
- Exploiting Efficient Mapping and Pipelined Execution for Accelerating SpMV on Tensor CoresKaige Zhang, Hailong Yang, Xin You, Tianyu Feng et al.PPoPP 2026
- Streaming Tensor Programs: A Streaming Abstraction for Dynamic ParallelismGina Sohn, Genghan Zhang, Konstantin Hoßfeld, Jungwoo Kim et al.ASPLOS 2026 · 1 citation
- SFD: Towards Segment Fusion Dataflow for Spatial AcceleratorsFuyu Wang, Minghua Shen, Yufei Ding, Nong Xiao et al.HPCA 2026
- TensorIR: An Abstraction for Automatic Tensorized Program OptimizationSiyuan Feng, Bohan Hou, Hongyi Jin, Wuwei Lin et al.ASPLOS 2023 · 80 citations
- Revet: A Language and Compiler for Dataflow ThreadsAlexander C. Rucker, Shiv Sundram, Coleman Smith, Matthew Vilim et al.HPCA 2024 · 3 citations
