StreamTensor: Make Tensors Stream in Dataflow Accelerators for LLMs
Hanchen Ye, Deming Chen
摘要
Efficient execution of deep learning workloads on dataflow architectures is crucial for overcoming memory bottlenecks and maximizing performance. While streaming intermediate results between computation kernels can significantly improve efficiency, existing approaches struggle with inter-kernel correlations, external memory access management, and buffer optimization. In this work, we propose StreamTensor, a compiler framework that automatically constructs and optimizes stream-based dataflow accelerators. StreamTensor introduces a novel iterative tensor type system to explicitly encode stream layouts, enabling seamless kernel fusion, buffer allocation, and memory optimization. By systematically exploring three hierarchical design spaces, including tensor tiling, kernel fusion, and resource allocation, StreamTensor balances computational intensity, memory efficiency, and data streaming to maximize performance. Based on FPGA evaluations on Large Language Models (LLM), StreamTensor achieves up to 0.76x and 0.64x lower latency compared to the state-of-the-art FPGA LLM accelerators and GPUs, and up to 1.99x higher energy efficiency compared to GPUs, making it a promising approach for scalable dataflow-based deep learning acceleration.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper16
- PyTorch 2: Faster Machine Learning Through Dynamic Python Bytecode Transformation and Graph CompilationJason Ansel, Edward Z. Yang, Horace He, Natalia Gimelshein 等ASPLOS 2024 · 被引用 693 次
- DSAGEN: Synthesizing Programmable Spatial AcceleratorsJian Weng, Sihao Liu, Vidushi Dadu, Zhengrong Wang 等ISCA 2020 · 被引用 140 次
- DFX: A Low-latency Multi-FPGA Appliance for Accelerating Transformer-based Text GenerationSeongmin Hong, Seungjae Moon, Junsoo Kim, Sungjae Lee 等MICRO 2022 · 被引用 107 次
- FlowGNN: A Dataflow Architecture for Real-Time Workload-Agnostic Graph Neural Network InferenceRishov Sarkar, Stefan Abi-Karam, Yuqi He, Lakshmi Sathidevi 等HPCA 2023 · 被引用 100 次
- ScaleHLS: A New Scalable High-Level Synthesis Framework on Multi-Level Intermediate RepresentationHanchen Ye, Cong Hao, Jianyi Cheng, Hyunmin Jeong 等HPCA 2022 · 被引用 77 次
相关 Paper
- Streaming Tensor Programs: A Streaming Abstraction for Dynamic ParallelismGina Sohn, Genghan Zhang, Konstantin Hoßfeld, Jungwoo Kim 等ASPLOS 2026 · 被引用 1 次
- Rammer: Enabling Holistic Deep Learning Compiler Optimizations with rTasksLingxiao Ma, Zhiqiang Xie, Zhi Yang, Jilong Xue 等OSDI 2020 · 被引用 192 次
- ROLLER: Fast and Efficient Tensor Compilation for Deep LearningHongyu Zhu, Ruofan Wu, Yijia Diao, Shanbin Ke 等OSDI 2022 · 被引用 84 次
- Welder: Scheduling Deep Learning Memory Access via Tile-graphYining Shi, Zhi Yang, Jilong Xue, Lingxiao Ma 等OSDI 2023 · 被引用 64 次
- SFD: Towards Segment Fusion Dataflow for Spatial AcceleratorsFuyu Wang, Minghua Shen, Yufei Ding, Nong Xiao 等HPCA 2026
