Mind the Gap: Attainable Data Movement and Operational Intensity Bounds for Tensor Algorithms
Qijing Huang, Po-An Tsai, Joel S. Emer, Angshuman Parashar
摘要
The architectural design-space exploration (or DSE) process-whether manual or automated-benefits greatly from knowing the limits of the metrics of interest in advance. Data movement is rapidly emerging as a critical metric for DSE due to its increasing impact on both performance and energy efficiency. Unfortunately, the commonly used algorithmic minimum (or “compulsory misses”) limit for data movement is extremely loose, limiting its utility in design-space search. In this paper, we present Orojenesis, an approach to compute data movement limits (or bounds) for tensor algorithms. Unlike algorithmic-minimum bounds, Orojenesis comprehends reuse and the ability of a buffer (such as a cache or scratchpad) to exploit reuse to reduce data movement. Orojenesis provides a bound that no dataflow or mapping can possibly exceed under varying onchip buffer capacity constraints, including mappings that fuse a sequence of tensor operations to exploit producer-consumer reuse. Orojenesis produces a plot that shows the relationship between a buffer’s size and the lower data movement limit to/from the next level in a memory hierarchy. This plot, dubbed a ski-slope diagram, allows designers to gain critical insights into the behavior of a workload as a function of storage capacity. This analysis can inform early high-level design decisions before embarking on thorough design space searches. We use Orojenesis to analyze a set of valuable tensor algorithms including batched and grouped matrix multiplications, convolutions, and sequences of operations in Large Language Models (LLMs). Our analysis reveals a range of architectural insights, including the fact that attainable data movement can be orders-of-magnitude higher than algorithmic minimum, that there exists a sweet spot between SRAM and compute resource provisioning for optimal throughput, and that up to data movement reduction can be achieved with fusion with a buffer capacity of 320 MB for the GPT-3-6.7b LLM.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper4
- RAGO: Systematic Performance Optimization for Retrieval-Augmented Generation ServingWenqi Jiang, Suvinay Subramanian, Cat Graves, Gustavo Alonso 等ISCA 2025 · 被引用 16 次
- Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data FormatChao Fang, Man Shi, Robin Geens, Arne Symons 等HPCA 2025 · 被引用 15 次
- DataMaestro: A Versatile and Efficient Data Streaming Engine Bringing Decoupled Memory Access To Dataflow AcceleratorsXiaoling Yi, Yunhao Deng, Ryan Antonio, Fanchen Kong 等DAC 2025 · 被引用 4 次
- Concorde: Fast and Accurate CPU Performance Modeling with Compositional Analytical-ML FusionArash Nasr-Esfahany, Mohammad Alizadeh, Victor Lee, Hanna Alam 等ISCA 2025 · 被引用 1 次
相关 Paper
- Principle-based Dataflow Optimization for Communication Lower Bound in Operator-Fused Tensor AcceleratorLei Xu, Chen Yin, Zelong Yuan, Weiguang Sheng 等DAC 2025
- Automated derivation of parametric data movement lower bounds for affine programsAuguste Olivry, Julien Langou, Louis-Noël Pouchet, P. Sadayappan 等PLDI 2020
- Patterns Behind Chaos: Forecasting Data Movement for Efficient Large-Scale Moe LLM InferenceZhongkai Yu, Yue Guan, Zihao Yu, Chenyang Zhou 等ISCA 2026
- VTC: DNN Compilation with Virtual Tensors for Data Movement EliminationMuyan Hu, Ahan Gupta, Jiachen Yuan, Vima Gupta 等OSDI 2026
- Sparsepipe: Sparse Inter-operator Dataflow Architecture with Cross-Iteration ReuseYunan Zhang, Po-An Tsai, Hung-Wei TsengMICRO 2024 · 被引用 2 次
