Automated derivation of parametric data movement lower bounds for affine programs
Auguste Olivry, Julien Langou, Louis-Noël Pouchet, P. Sadayappan, Fabrice Rastello
Abstract
For most relevant computation, the energy and time needed for data movement dominates that for performing arithmetic operations on all computing systems today. Hence it is of critical importance to understand the minimal total data movement achievable during the execution of an algorithm. The achieved total data movement for different schedules of an algorithm can vary widely depending on how efficiently the cache is used, e.g., untiled versus effectively tiled matrix-matrix multiplication. A significant current challenge is that no existing tool is able to meaningfully quantify the potential reduction to the data movement of a computation that can be achieved by more effective use of the cache through operation rescheduling. Asymptotic parametric expressions of data movement lower bounds have previously been manually derived for a limited number of algorithms, often without scaling constants. In this paper, we present the first compile-time approach for deriving non-asymptotic parametric expressions of data movement lower bounds for arbitrary affine computations.
The approach has been implemented in a fully automatic tool (IOLB) that can generate these lower bounds for input affine programs. IOLB's use is demonstrated by exercising it on all the benchmarks of the PolyBench suite. The advantages of IOLB are many: (1) IOLB enables us to derive bounds for few dozens of algorithms for which these lower bounds have never been derived. This reflects an increase of productivity by automation.
(2) Anyone is able to obtain these lower bounds through IOLB, no expertise is required. (3) For some of the most well-studied algorithms, the lower bounds obtained by IOLB are higher than any previously reported manually derived lower bounds.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4c1e36cf-c860-4724-830d-74648619f93fCited by top-tier papers4
- On the parallel I/O optimality of linear algebra kernels: near-optimal matrix factorizationsGrzegorz Kwasniewski, Marko Kabic, Tal Ben-Nun, Alexandros Nikolaos Ziogas et al.SC 2021 · 18 citations
- FPL: fast Presburger arithmetic through transprecisionArjun Pitchanathan, Christian Ulmann, Michel Weber, Torsten Hoefler et al.OOPSLA 2021 · 10 citations
- IOOpt: automatic derivation of I/O complexity bounds for affine programsAuguste Olivry, Guillaume Iooss, Nicolas Tollenaere, Atanas Rountev et al.PLDI 2021 · 10 citations
- Symmetric Block-Cyclic Distribution: Fewer Communications Leads to Faster Dense Cholesky FactorizationOlivier Beaumont, Philippe Duchon, Lionel Eyraud-Dubois, Julien Langou et al.SC 2022 · 7 citations
Related papers
- Mind the Gap: Attainable Data Movement and Operational Intensity Bounds for Tensor AlgorithmsQijing Huang, Po-An Tsai, Joel S. Emer, Angshuman ParasharISCA 2024 · 11 citations
- Analytical characterization and design space exploration for optimization of CNNsRui Li, Yufan Xu, Aravind Sukumaran-Rajam, Atanas Rountev et al.ASPLOS 2021 · 52 citations
- Deinsum: Practically I/O Optimal Multi-Linear AlgebraAlexandros Nikolaos Ziogas, Grzegorz Kwasniewski, Tal Ben-Nun, Timo Schneider et al.SC 2022 · 2 citations
- Static Generation of Efficient OpenMP Offload Data MappingsLuke Marzen, Akash Dutta, Ali JannesariSC 2024 · 4 citations
- MagiCache: A Virtual In-Cache Computing EngineRenhao Fan, Yikai Cui, Weike Li, Mingyu Wang et al.ISCA 2025 · 3 citations
