Neptune: Advanced ML Operator Fusion for Locality and Parallelism on GPUs
Yifan Zhao, Egan Johnson, Prasanth Chatarasi, Vikram S. Adve, Sasa Misailovic
Abstract
Operator fusion has become a key optimization for deep learning, which combines multiple deep learning operators to improve data reuse and reduce global memory transfers. However, existing tensor compilers struggle to fuse complex reduction computations involving loop-carried dependencies, such as attention mechanisms. This paper introduces Neptune, a tensor compiler for advanced operator fusion for sequences of reduction operators. Neptune presents a new approach for advanced operator fusion, which intentionally breaks some existing dependencies and compensates by constructing algebraic correction expressions that allow the kernel to produce the correct result. Applying Neptune’s advanced operator fusion to a plain attention operator generates operators equivalent to FlashAttention and FlashDecoding. On ten attention-based benchmarks, Neptune, starting from a plain attention code and a high–level scheduling template, outperforms existing compilers like Triton, TVM, and Flex Attention, including Triton–based implementations of FlashAttention. Across four different GPU architectures from NVIDIA and AMD, Neptune–generated kernels have an average speedup of 1.35× over the next best alternative, with up to 2.65 × speedup on Nvidia GPUs and up to 3.32 × on AMD GPUs, demonstrating its effectiveness for deep learning workloads.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9d284519-a9f1-48f4-ba2c-56aabbcdd6bfBuilds on17
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra et al.NeurIPS 2022 · 5,493 citations
- FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningTri DaoICLR 2024 · 2,600 citations
- Ansor: Generating High-Performance Tensor Programs for Deep LearningLianmin Zheng, Chengfan Jia, Minmin Sun, Zhao Wu et al.OSDI 2020 · 551 citations
- FlexTensor: An Automatic Schedule Exploration and Optimization Framework for Tensor Computation on Heterogeneous SystemSize Zheng, Yun Liang, Shuo Wang, Renze Chen et al.ASPLOS 2020 · 171 citations
- DNNFusion: accelerating deep neural networks execution with advanced operator fusionWei Niu, Jiexiong Guan, Yanzhi Wang, Gagan Agrawal et al.PLDI 2021 · 166 citations
Related papers
- RedFuser: An Automatic Operator Fusion Framework for Cascaded Reductions on AI AcceleratorsXinsheng Tang, Yangcheng Li, Nan Wang, Zhiyi Shu et al.ASPLOS 2026
- Trinity: Three-Dimensional Tensor Program Optimization via Tile-level Equality SaturationJaehyeong Park, Youngchan Kim, Haechan An, Gieun Jeong et al.ASPLOS 2026
- Uncovering Nested Data Parallelism and Data Reuse in DNN Computation with FractalTensorSiran Liu, Chengxiang Qi, Ying Cao, Chao Yang et al.SOSP 2024 · 1 citation
- Optimizing Deep Learning Inference via Global Analysis and Tensor ExpressionsChunwei Xia, Jiacheng Zhao, Qianqi Sun, Zheng Wang et al.ASPLOS 2024 · 14 citations
- SpaceFusion: Advanced Deep Learning Operator Fusion via Space-Mapping GraphLiang Zhu, Jianguo Yao, Haibing GuanEuroSys 2025 · 3 citations
