Lightweight and Locality-Aware Composition of Black-Box Subroutines
Manya Bansal, Dillon Sharlet, Jonathan Ragan-Kelley, Saman P. Amarasinghe
Abstract
Subroutines are essential building blocks in software design: users encapsulate common functionality in libraries and write applications by composing calls to subroutines. Unfortunately, performance may be lost at subroutine boundaries due to reduced locality and increased memory consumption. Operator fusion helps recover the performance lost at composition boundaries. Previous solutions fuse operators by manually rewriting code into monolithic fused subroutines, or by relying on heavy-weight compilers to generate code that performs fusion. Both approaches require a semantic understanding of the entire computation, breaking the decoupling necessary for modularity and reusability of subroutines.
In this work, we attempt to identify the minimal ingredients required to fuse computations, enabling composition of subroutines without sacrificing performance or modularity. We find that, unlike previous approaches that require a semantic understanding of the computation, most opportunities for fusion require understanding only data production and consumption patterns. Exploiting this insight, we add fusion on top of black-box subroutines by proposing a lightweight enrichment of subroutine declarations to expose data-dependence patterns. We implement our approach in a system called Fern, and demonstrate Fern's benefits by showing that it is competitive with state-of-the-art, high-performance libraries with manually fused operators, can fuse across library and domain boundaries for unforeseen workloads, and can deliver speedups of up to 5× over unfused code.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 297a53a4-8d07-4a17-bb8f-06a57b66e3c8Builds on6
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra et al.NeurIPS 2022 · 5,493 citations
- PyTorch 2: Faster Machine Learning Through Dynamic Python Bytecode Transformation and Graph CompilationJason Ansel, Edward Z. Yang, Horace He, Natalia Gimelshein et al.ASPLOS 2024 · 693 citations
- OpenCilk: A Modular and Extensible Software Infrastructure for Fast Task-Parallel CodeTao B. Schardl, I-Ting Angelina LeePPoPP 2023 · 30 citations
- Mosaic: An Interoperable Compiler for Tensor AlgebraManya Bansal, Olivia Hsu, Kunle Olukotun, Fredrik KjolstadPLDI 2023 · 16 citations
- Hector: An Efficient Programming and Compilation Framework for Implementing Relational Graph Neural Networks in GPU ArchitecturesKun Wu, Mert Hidayetoglu, Xiang Song, Sitao Huang et al.ASPLOS 2024 · 4 citations
Related papers
- MCFuser: High-Performance and Rapid Fusion of Memory-Bound Compute-Intensive OperatorsZheng Zhang, Donglin Yang, Xiaobo Zhou, Dazhao ChengSC 2024 · 9 citations
- RedFuser: An Automatic Operator Fusion Framework for Cascaded Reductions on AI AcceleratorsXinsheng Tang, Yangcheng Li, Nan Wang, Zhiyi Shu et al.ASPLOS 2026
- Runtime Composition of Iterations for Fusing Loop-carried Sparse DependenceKazem Cheshmi, Michelle Strout, Maryam Mehri DehnaviSC 2023 · 9 citations
- Assertion-based optimization of Quantum programsThomas Häner, Torsten Hoefler, Matthias TroyerOOPSLA 2020 · 13 citations
- Neptune: Advanced ML Operator Fusion for Locality and Parallelism on GPUsYifan Zhao, Egan Johnson, Prasanth Chatarasi, Vikram S. Adve et al.PLDI 2026 · 1 citation
