Lune

PLDI2026Top-tier venue

Neptune: Advanced ML Operator Fusion for Locality and Parallelism on GPUs

Yifan Zhao, Egan Johnson, Prasanth Chatarasi, Vikram S. Adve, Sasa Misailovic

2026Year
1Citations

Abstract

Operator fusion has become a key optimization for deep learning, which combines multiple deep learning operators to improve data reuse and reduce global memory transfers. However, existing tensor compilers struggle to fuse complex reduction computations involving loop-carried dependencies, such as attention mechanisms. This paper introduces Neptune, a tensor compiler for advanced operator fusion for sequences of reduction operators. Neptune presents a new approach for advanced operator fusion, which intentionally breaks some existing dependencies and compensates by constructing algebraic correction expressions that allow the kernel to produce the correct result. Applying Neptune’s advanced operator fusion to a plain attention operator generates operators equivalent to FlashAttention and FlashDecoding. On ten attention-based benchmarks, Neptune, starting from a plain attention code and a high–level scheduling template, outperforms existing compilers like Triton, TVM, and Flex Attention, including Triton–based implementations of FlashAttention. Across four different GPU architectures from NVIDIA and AMD, Neptune–generated kernels have an average speedup of 1.35× over the next best alternative, with up to 2.65 × speedup on Nvidia GPUs and up to 3.32 × on AMD GPUs, demonstrating its effectiveness for deep learning workloads.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext 9d284519-a9f1-48f4-ba2c-56aabbcdd6bf

Builds on17

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines