SC2023Top-tier venue
Automatic Generation of Distributed-Memory Mappings for Tensor Computations
Martin Kong, Raneem Abu Yosef, Atanas Rountev, P. Sadayappan
2023Year
9Citations
1Top-tier citations
Abstract
While considerable research has been directed at automatic parallelization for shared-memory platforms, little progress has been made in automatic parallelization schemes for distributed-memory systems. We introduce an innovative approach to automatically produce distributed-memory parallel code for an important subclass of affine tensor computations common to Coupled Cluster (CC) electronic structure methods, neuro-imaging applications, and deep learning models.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 201b9d22-ce96-4704-9cec-65b51324bc93Cited by top-tier papers1
Ask how each one uses itRelated papers
- HAP: SPMD DNN Training on Heterogeneous GPU Clusters with Automated Program SynthesisShiwei Zhang, Lansong Diao, Chuan Wu, Zongyan Cao et al.EuroSys 2024 · 16 citations
- SpDISTAL: Compiling Distributed Sparse Tensor ComputationsRohan Yadav, Alex Aiken, Fredrik KjolstadSC 2022 · 7 citations
- Compressed and Parallelized Structured Tensor AlgebraMahdi Ghorbani, Emilien Bauer, Tobias Grosser, Amir ShaikhhaOOPSLA 2025 · 1 citation
- Automatic Optimization of Matrix Implementations for Distributed Machine Learning and Linear AlgebraShangyu Luo, Dimitrije Jankov, Binhang Yuan, Chris JermaineSIGMOD 2021 · 9 citations
- Mosaic: Exploiting Instruction-Level Parallelism on Deep Learning Accelerators with iTex TessellationJianxing Xu, Yuanbo Wen, Zikang Liu, Ruibai Xu et al.ASPLOS 2025 · 2 citations
