Mind the Gap: Attainable Data Movement and Operational Intensity Bounds for Tensor Algorithms
Qijing Huang, Po-An Tsai, Joel S. Emer, Angshuman Parashar
Abstract
The architectural design-space exploration (or DSE) process-whether manual or automated-benefits greatly from knowing the limits of the metrics of interest in advance. Data movement is rapidly emerging as a critical metric for DSE due to its increasing impact on both performance and energy efficiency. Unfortunately, the commonly used algorithmic minimum (or “compulsory misses”) limit for data movement is extremely loose, limiting its utility in design-space search. In this paper, we present Orojenesis, an approach to compute data movement limits (or bounds) for tensor algorithms. Unlike algorithmic-minimum bounds, Orojenesis comprehends reuse and the ability of a buffer (such as a cache or scratchpad) to exploit reuse to reduce data movement. Orojenesis provides a bound that no dataflow or mapping can possibly exceed under varying onchip buffer capacity constraints, including mappings that fuse a sequence of tensor operations to exploit producer-consumer reuse. Orojenesis produces a plot that shows the relationship between a buffer’s size and the lower data movement limit to/from the next level in a memory hierarchy. This plot, dubbed a ski-slope diagram, allows designers to gain critical insights into the behavior of a workload as a function of storage capacity. This analysis can inform early high-level design decisions before embarking on thorough design space searches. We use Orojenesis to analyze a set of valuable tensor algorithms including batched and grouped matrix multiplications, convolutions, and sequences of operations in Large Language Models (LLMs). Our analysis reveals a range of architectural insights, including the fact that attainable data movement can be orders-of-magnitude higher than algorithmic minimum, that there exists a sweet spot between SRAM and compute resource provisioning for optimal throughput, and that up to data movement reduction can be achieved with fusion with a buffer capacity of 320 MB for the GPT-3-6.7b LLM.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 133a9e78-f8c9-4b86-bf39-ec59608de8d1Cited by top-tier papers4
- RAGO: Systematic Performance Optimization for Retrieval-Augmented Generation ServingWenqi Jiang, Suvinay Subramanian, Cat Graves, Gustavo Alonso et al.ISCA 2025 · 16 citations
- Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data FormatChao Fang, Man Shi, Robin Geens, Arne Symons et al.HPCA 2025 · 15 citations
- DataMaestro: A Versatile and Efficient Data Streaming Engine Bringing Decoupled Memory Access To Dataflow AcceleratorsXiaoling Yi, Yunhao Deng, Ryan Antonio, Fanchen Kong et al.DAC 2025 · 4 citations
- Concorde: Fast and Accurate CPU Performance Modeling with Compositional Analytical-ML FusionArash Nasr-Esfahany, Mohammad Alizadeh, Victor Lee, Hanna Alam et al.ISCA 2025 · 1 citation
Related papers
- Principle-based Dataflow Optimization for Communication Lower Bound in Operator-Fused Tensor AcceleratorLei Xu, Chen Yin, Zelong Yuan, Weiguang Sheng et al.DAC 2025
- Automated derivation of parametric data movement lower bounds for affine programsAuguste Olivry, Julien Langou, Louis-Noël Pouchet, P. Sadayappan et al.PLDI 2020
- Patterns Behind Chaos: Forecasting Data Movement for Efficient Large-Scale Moe LLM InferenceZhongkai Yu, Yue Guan, Zihao Yu, Chenyang Zhou et al.ISCA 2026
- VTC: DNN Compilation with Virtual Tensors for Data Movement EliminationMuyan Hu, Ahan Gupta, Jiachen Yuan, Vima Gupta et al.OSDI 2026
- Sparsepipe: Sparse Inter-operator Dataflow Architecture with Cross-Iteration ReuseYunan Zhang, Po-An Tsai, Hung-Wei TsengMICRO 2024 · 2 citations
