Automatic Generation of Distributed-Memory Mappings for Tensor Computations
Martin Kong, Raneem Abu Yosef, Atanas Rountev, P. Sadayappan
2023年份
9被引次数
1顶会引用
摘要
While considerable research has been directed at automatic parallelization for shared-memory platforms, little progress has been made in automatic parallelization schemes for distributed-memory systems. We introduce an innovative approach to automatically produce distributed-memory parallel code for an important subclass of affine tensor computations common to Coupled Cluster (CC) electronic structure methods, neuro-imaging applications, and deep learning models.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper1
问问它们各自怎么用它相关 Paper
- HAP: SPMD DNN Training on Heterogeneous GPU Clusters with Automated Program SynthesisShiwei Zhang, Lansong Diao, Chuan Wu, Zongyan Cao 等EuroSys 2024 · 被引用 16 次
- SpDISTAL: Compiling Distributed Sparse Tensor ComputationsRohan Yadav, Alex Aiken, Fredrik KjolstadSC 2022 · 被引用 7 次
- Compressed and Parallelized Structured Tensor AlgebraMahdi Ghorbani, Emilien Bauer, Tobias Grosser, Amir ShaikhhaOOPSLA 2025 · 被引用 1 次
- Automatic Optimization of Matrix Implementations for Distributed Machine Learning and Linear AlgebraShangyu Luo, Dimitrije Jankov, Binhang Yuan, Chris JermaineSIGMOD 2021 · 被引用 9 次
- Mosaic: Exploiting Instruction-Level Parallelism on Deep Learning Accelerators with iTex TessellationJianxing Xu, Yuanbo Wen, Zikang Liu, Ruibai Xu 等ASPLOS 2025 · 被引用 2 次
