Composing Distributed Computations Through Task and Kernel Fusion
Rohan Yadav, Shiv Sundram, Wonchan Lee, Michael Garland, Michael Bauer, Alex Aiken, Fredrik Kjolstad
摘要
We introduce Diffuse, a system that dynamically performs task and kernel fusion in distributed, task-based runtime systems. The key component of Diffuse is an intermediate representation of distributed computation that enables the necessary analyses for the fusion of distributed tasks to be performed in a scalable manner. We pair task fusion with a JIT compiler to fuse together the kernels within fused tasks. We show empirically that Diffuse's intermediate representation is general enough to be a target for two real-world, task-based libraries (cuPyNumeric and Legate Sparse), letting Diffuse find optimization opportunities across function and library boundaries. Diffuse accelerates unmodified applications developed by composing task-based libraries by 1.86x on average (geo-mean), and by between 0.93x--10.7x on up to 128 GPUs. Diffuse also finds optimization opportunities missed by the original application developers, enabling high-level Python programs to match or exceed the performance of an explicitly parallel MPI library.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Task-Based Tensor Computations on Modern GPUsRohan Yadav, Michael Garland, Alex Aiken, Michael BauerPLDI 2025 · 被引用 5 次
- Iris: Expressive Traffic Analysis for the Modern InternetThea Rossman, Diana Qing, Gerry Wan, Zakir DurumericNSDI 2026 · 被引用 2 次
- Lightweight and Locality-Aware Composition of Black-Box SubroutinesManya Bansal, Dillon Sharlet, Jonathan Ragan-Kelley, Saman P. AmarasinghePLDI 2025
它引用的顶会 Paper11
- PyTorch 2: Faster Machine Learning Through Dynamic Python Bytecode Transformation and Graph CompilationJason Ansel, Edward Z. Yang, Horace He, Natalia Gimelshein 等ASPLOS 2024 · 被引用 693 次
- DNNFusion: accelerating deep neural networks execution with advanced operator fusionWei Niu, Jiexiong Guan, Yanzhi Wang, Gagan Agrawal 等PLDI 2021 · 被引用 166 次
- Unity: Accelerating DNN Training Through Joint Optimization of Algebraic Transformations and ParallelizationColin Unger, Zhihao Jia, Wei Wu, Sina Lin 等OSDI 2022 · 被引用 105 次
- Task bench: a parameterized benchmark for evaluating parallel runtime performanceElliott Slaughter, Wei Wu, Yuankun Fu, Legend Brandenburg 等SC 2020 · 被引用 51 次
- DISTAL: the distributed tensor algebra compilerRohan Yadav, Alex Aiken, Fredrik KjolstadPLDI 2022 · 被引用 29 次
相关 Paper
- Legate Sparse: Distributed Sparse Computing in PythonRohan Yadav, Wonchan Lee, Melih Elibol, Manolis Papadakis 等SC 2023 · 被引用 8 次
- SpDISTAL: Compiling Distributed Sparse Tensor ComputationsRohan Yadav, Alex Aiken, Fredrik KjolstadSC 2022 · 被引用 7 次
- Offload Annotations: Bringing Heterogeneous Computing to Existing Libraries and WorkloadsGina Yuan, Shoumik Palkar, Deepak Narayanan, Matei ZahariaUSENIX ATC 2020 · 被引用 11 次
- SmartDispatch: Dynamic Substitution of NumPy-Style APIs on Heterogeneous CPU-GPU SystemsJinku Cui, Yueming Hao, Shuyin Jiao, Jiajia Li 等FSE 2026
- Parla: A Python Orchestration System for Heterogeneous ArchitecturesHochan Lee, William Ruys, Ian Henriksen, Arthur Michener Peters 等SC 2022 · 被引用 6 次
