Composing Distributed Computations Through Task and Kernel Fusion
Rohan Yadav, Shiv Sundram, Wonchan Lee, Michael Garland, Michael Bauer, Alex Aiken, Fredrik Kjolstad
Abstract
We introduce Diffuse, a system that dynamically performs task and kernel fusion in distributed, task-based runtime systems. The key component of Diffuse is an intermediate representation of distributed computation that enables the necessary analyses for the fusion of distributed tasks to be performed in a scalable manner. We pair task fusion with a JIT compiler to fuse together the kernels within fused tasks. We show empirically that Diffuse's intermediate representation is general enough to be a target for two real-world, task-based libraries (cuPyNumeric and Legate Sparse), letting Diffuse find optimization opportunities across function and library boundaries. Diffuse accelerates unmodified applications developed by composing task-based libraries by 1.86x on average (geo-mean), and by between 0.93x--10.7x on up to 128 GPUs. Diffuse also finds optimization opportunities missed by the original application developers, enabling high-level Python programs to match or exceed the performance of an explicitly parallel MPI library.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- Task-Based Tensor Computations on Modern GPUsRohan Yadav, Michael Garland, Alex Aiken, Michael BauerPLDI 2025 · 5 citations
- Iris: Expressive Traffic Analysis for the Modern InternetThea Rossman, Diana Qing, Gerry Wan, Zakir DurumericNSDI 2026 · 2 citations
- Lightweight and Locality-Aware Composition of Black-Box SubroutinesManya Bansal, Dillon Sharlet, Jonathan Ragan-Kelley, Saman P. AmarasinghePLDI 2025
Builds on11
- PyTorch 2: Faster Machine Learning Through Dynamic Python Bytecode Transformation and Graph CompilationJason Ansel, Edward Z. Yang, Horace He, Natalia Gimelshein et al.ASPLOS 2024 · 693 citations
- DNNFusion: accelerating deep neural networks execution with advanced operator fusionWei Niu, Jiexiong Guan, Yanzhi Wang, Gagan Agrawal et al.PLDI 2021 · 166 citations
- Unity: Accelerating DNN Training Through Joint Optimization of Algebraic Transformations and ParallelizationColin Unger, Zhihao Jia, Wei Wu, Sina Lin et al.OSDI 2022 · 105 citations
- Task bench: a parameterized benchmark for evaluating parallel runtime performanceElliott Slaughter, Wei Wu, Yuankun Fu, Legend Brandenburg et al.SC 2020 · 51 citations
- DISTAL: the distributed tensor algebra compilerRohan Yadav, Alex Aiken, Fredrik KjolstadPLDI 2022 · 29 citations
Related papers
- Legate Sparse: Distributed Sparse Computing in PythonRohan Yadav, Wonchan Lee, Melih Elibol, Manolis Papadakis et al.SC 2023 · 8 citations
- SpDISTAL: Compiling Distributed Sparse Tensor ComputationsRohan Yadav, Alex Aiken, Fredrik KjolstadSC 2022 · 7 citations
- Offload Annotations: Bringing Heterogeneous Computing to Existing Libraries and WorkloadsGina Yuan, Shoumik Palkar, Deepak Narayanan, Matei ZahariaUSENIX ATC 2020 · 11 citations
- SmartDispatch: Dynamic Substitution of NumPy-Style APIs on Heterogeneous CPU-GPU SystemsJinku Cui, Yueming Hao, Shuyin Jiao, Jiajia Li et al.FSE 2026
- Parla: A Python Orchestration System for Heterogeneous ArchitecturesHochan Lee, William Ruys, Ian Henriksen, Arthur Michener Peters et al.SC 2022 · 6 citations
