TileFlow: A Framework for Modeling Fusion Dataflow via Tree-based Analysis
Size Zheng, Siyuan Chen, Siyuan Gao, Liancheng Jia, Guangyu Sun, Runsheng Wang, Yun Liang
摘要
With the increasing size of DNN models and the growing discrepancy between compute performance and memory bandwidth, fusing multiple layers together to reduce off-chip memory access has become a popular approach in dataflow design. However, designing such dataflows requires flexible and accurate performance models to facilitate evaluation, architecture analysis, and design space exploration. Unfortunately, current state-of-the-art performance models are limited to the dataflows of single operator acceleration, making them inapplicable to operator fusion dataflows.
In this paper, we propose a framework called TileFlow that models dataflows for operator fusion. We first characterize the design space of fusion dataflows as a 3D space encompassing compute ordering, resource binding, and loop tiling. We then introduce a tile-centric notation to express dataflow designs within this space. Inspired by the tiling structure of fusion dataflows, we present a tree-based approach to analyze two critical performance metrics: data movement volume within the accelerator memory hierarchy and accelerator compute/memory resource usage. Finally, we leverage these metrics to calculate latency and energy consumption. Our evaluation validates TileFlow's modeling accuracy against both real
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- LUT Tensor Core: A Software-Hardware Co-Design for LUT-Based Low-Bit LLM InferenceZhiwen Mo, Lei Wang, Jianyu Wei, Zhichen Zeng 等ISCA 2025 · 被引用 17 次
- MAGIS: Memory Optimization via Coordinated Graph Transformation and Scheduling for DNNRenze Chen, Zijian Ding, Size Zheng, Chengrui Zhang 等ASPLOS 2024 · 被引用 14 次
- VQ-LLM: High-performance Code Generation for Vector Quantization Augmented LLM InferenceZihan Liu, Xinhao Luo, Junxian Guo, Wentao Ni 等HPCA 2025 · 被引用 11 次
- FuseMax: Leveraging Extended Einsums to Optimize Attention Accelerator DesignNandeeka Nayak, Xinrui Wu, Toluwanimi O. Odemuyiwa, Michael Pellauer 等MICRO 2024 · 被引用 10 次
- ClusterFusion: Expanding Operator Fusion Scope for LLM Inference via Cluster-Level Collective PrimitiveXinhao Luo, Zihan Liu, Yangjie Zhou, Shihan Fang 等NeurIPS 2025 · 被引用 9 次
它引用的顶会 Paper31
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra 等NeurIPS 2022 · 被引用 5,493 次
- Ansor: Generating High-Performance Tensor Programs for Deep LearningLianmin Zheng, Chengfan Jia, Minmin Sun, Zhao Wu 等OSDI 2020 · 被引用 551 次
- Gemmini: Enabling Systematic Deep-Learning Architecture Evaluation via Full-Stack IntegrationHasan Genc, Seah Kim, Alon Amid, Ameer Haj-Ali 等DAC 2021 · 被引用 325 次
相关 Paper
- SFD: Towards Segment Fusion Dataflow for Spatial AcceleratorsFuyu Wang, Minghua Shen, Yufei Ding, Nong Xiao 等HPCA 2026
- Welder: Scheduling Deep Learning Memory Access via Tile-graphYining Shi, Zhi Yang, Jilong Xue, Lingxiao Ma 等OSDI 2023 · 被引用 64 次
- Enabling Multiple Tensor-wise Operator Fusion for Transformer Models on Spatial AcceleratorsLei Xu, Zhiwen Mo, Qin Wang, Jianfei Jiang 等DAC 2024 · 被引用 4 次
- Principle-based Dataflow Optimization for Communication Lower Bound in Operator-Fused Tensor AcceleratorLei Xu, Chen Yin, Zelong Yuan, Weiguang Sheng 等DAC 2025
- DeFiNES: Enabling Fast Exploration of the Depth-first Scheduling Space for DNN Accelerators through Analytical ModelingLinyan Mei, Koen Goetschalckx, Arne Symons, Marian VerhelstHPCA 2023 · 被引用 40 次
