Lune

SOSP2026Top-tier venue

A ve : Guiding Agentic GPU Optimization Using Data-Flow Invariants

Haohui Mai, Xiaoyan Guo, Xiangyun Ding, Daifeng Li, Qiuchu Yu, Chenzhun Guo, Cong Wang, Jiacheng Zhao, Christos Kozyrakis, Binhang Yuan

2026Year

Abstract

LLM coding agents can generate correct GPU kernels, yet their performance still lags behind expert libraries. Achieving peak GPU throughput requires coordinating low-level transformations such as shared-memory staging, software pipelining, and instruction scheduling. Unit tests and performance profiles provide only sparse end-to-end feedback, leaving agents unable to localize violations of the global constraints that these transformations must preserve.

We present Ave, an agentic framework that uses data-flow invariants as compile-time guardrails for kernel optimization. These invariants specify relationships among data values that must hold throughout execution. Ave's tile-based Pythonic DSL exposes hardware instructions and compiler policies while abstracting complex memory layout encodings as tiles. Tag functions assign symbolic labels derived from selected logical coordinates and expressions. The compiler propagates these labels through data and control flow, and tag assertions enforce required relationships among them at use sites. The compiler checks these assertions with a tractable, flow-sensitive, path-insensitive analysis backed by an SMT solver, returning concrete assignments that witness abstract violations and guide repairs with no runtime overhead. An in-context reinforcement learning (ICRL) planner proposes optimizations from a curated knowledge base. A lowering agent implements them and instantiates their invariants.

We evaluate Ave on AMD MI300X across GEMM, flash attention, and MoE, which together account for up to 90% of GPU time in LLM inference. With GPT-5.6 Sol, its optimized kernels achieve 89-99% of the effective throughput of stateof-the-art hand-optimized libraries and improve geometricmean throughput by 1.62-1176× over uncontaminated agentic baselines. On 200 KernelBench tasks, with GPT-5.6 Sol and in-context DSL examples, Ave produces valid kernels within

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext bdb63492-85c7-447d-8dd1-53bed60d9cb8

Builds on9

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines