Deinsum: Practically I/O Optimal Multi-Linear Algebra
Alexandros Nikolaos Ziogas, Grzegorz Kwasniewski, Tal Ben-Nun, Timo Schneider, Torsten Hoefler
摘要
Multilinear algebra kernel performance on modern massively-parallel systems is determined mainly by data movement. However, deriving data movement-optimal distributed schedules for programs with many high-dimensional inputs is a notoriously hard problem. State-of-the-art libraries rely on heuristics and often fall back to suboptimal tensor folding and BLAS calls. We present Deinsum, an automated framework for distributed multilinear algebra computations expressed in Einstein notation, based on rigorous mathematical tools to address this problem. Our framework automatically derives data movement-optimal tiling and generates corresponding distributed schedules, further optimizing the performance of local computations by increasing their arithmetic intensity. To show the benefits of our approach, we test it on two important tensor kernel classes: Matricized Tensor Times Khatri-Rao Products and Tensor Times Matrix chains. We show performance results and scaling on the Piz Daint supercomputer, with up to 19x speedup over state-of-the-art solutions on 512 nodes.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- High-Performance and Programmable Attentional Graph Neural Networks with Global Tensor FormulationsMaciej Besta, Pawel Renc, Robert Gerstenberger, Paolo Sylos Labini 等SC 2023 · 被引用 5 次
- FuzzyFlow: Leveraging Dataflow To Find and Squash Program Optimization BugsPhilipp Schaad, Timo Schneider, Tal Ben-Nun, Alexandru Calotoiu 等SC 2023 · 被引用 3 次
它引用的顶会 Paper3
- HASCO: Towards Agile HArdware and Software CO-design for Tensor ComputationQingcheng Xiao, Size Zheng, Bingzhe Wu, Pengcheng Xu 等ISCA 2021 · 被引用 73 次
- Productivity, portability, performance: data-centric PythonAlexandros Nikolaos Ziogas, Timo Schneider, Tal Ben-Nun, Alexandru Calotoiu 等SC 2021 · 被引用 32 次
- On the parallel I/O optimality of linear algebra kernels: near-optimal matrix factorizationsGrzegorz Kwasniewski, Marko Kabic, Tal Ben-Nun, Alexandros Nikolaos Ziogas 等SC 2021 · 被引用 18 次
相关 Paper
- EinDecomp: Decomposition of Declaratively-Specified Machine Learning and Numerical Computations for Parallel ExecutionDaniel Bourgeois, Zhimin Ding, Dimitrije Jankov, Jiehui Li 等VLDB 2025 · 被引用 4 次
- Einsum Trees: An Abstraction for Optimizing the Execution of Tensor ExpressionsAlexander Breuer, Mark Blacher, Max Engel, Joachim Giesen 等ASPLOS 2025
- Automatic Optimization of Matrix Implementations for Distributed Machine Learning and Linear AlgebraShangyu Luo, Dimitrije Jankov, Binhang Yuan, Chris JermaineSIGMOD 2021 · 被引用 9 次
- Automatic Generation of Mappings for Distributed Fourier OperationsDoru-Thom Popovici, Botao Wu, John Shalf, Martin KongSC 2025 · 被引用 3 次
- Automated Tensor-Relational Decomposition for Large-Scale Sparse Tensor ComputationYuxin Tang, Zhiyuan Xin, Zhimin Ding, Xinyu Yao 等VLDB 2026
